morelike #
Turns one artist name into a ~100-song playlist, as service-neutral
(artist, title) pairs that can be resolved to any platform.
Live at https://morelike.dweb.workers.dev
Why it exists #
The playlist builders worth imitating were all wrappers around one large streaming service's recommendations endpoint. That endpoint was permanently restricted on 2024-11-27 to applications registered before that date — existing integrations are grandfathered, new API keys get 403 — so it can no longer be built on, and the tools already using it narrowed their exports to that one service.
So this builds the playlist from other similarity sources entirely, and keeps
the output platform-neutral: a list of (artist, title) pairs, resolved to a
player at the end rather than sourced from one at the start.
Running it #
npm install
npm run dev # Vite + the Worker in workerd, one code path with production
npm run deploy # vite build && wrangler deploy
npm run types # regenerate worker-configuration.d.ts after wrangler.jsonc changes
npx tsc -b # typecheck
LASTFM_API_KEY goes in .dev.vars locally (see .dev.vars.example) and in
wrangler secret put LASTFM_API_KEY for production.
worker-configuration.d.ts is generated and gitignored, so a fresh clone or
worktree fails npx tsc -b on worker/ and shared/ until npm run types
has been run once.
Architecture #
One Cloudflare Worker serves the built SPA and proxies the upstream APIs a browser cannot call directly.
shared/proxy-routes.ts—routes,matchRoute(). Single source of truth for the two proxied upstreams,/dz(Deezer) and/fm(Last.fm), plus/yt. ListenBrainz and MusicBrainz are called straight fromsrc/sources.ts.worker/index.ts— matches a route and proxies it, or falls through toenv.ASSETS.fetch(). Forwards none of the incoming request's headers.worker/youtube.ts— the/ytvideo-ID resolver, described under Playback.wrangler.jsonc—assets.not_found_handling: "single-page-application", with every proxy prefix listed inrun_worker_firstso the SPA fallback doesn't swallow API calls.src/About.tsxandsrc/App.tsxuseAboutRoute()— the about and roadmap page at/about. Nothing serves that path specially: the fallback above returnsindex.htmlfor anythingrun_worker_firstdoes not claim, anduseAboutRoute()readslocation.pathnameto decide what to render. Adding a second page means adding a path to that hook, not a route to the Worker.
The Last.fm key is a Worker secret that never reaches client code. Each
remaining proxy earns its place for a different reason — measured 2026-08-04
with an Origin header, Deezer is the one upstream that omits
access-control-allow-origin, and Last.fm's route holds the key rather than
fixing CORS. ListenBrainz and MusicBrainz both answer
access-control-allow-origin: * and so are called directly; see the comment at
the top of shared/proxy-routes.ts.
@cloudflare/vite-plugin runs that same Worker in workerd during npm run dev,
so development and production share one implementation of the routes rather
than a dev-server proxy config that can drift from the deployed one.
The algorithm #
Seed artist → similarity fan-out → rank fusion → second-hop expansion → per-artist catalogue pull → artist draw → two-pass track selection.
Similarity sources #
src/sources.ts fanout() queries three sources for one seed:
| source | returns | notes |
|---|---|---|
| Deezer related | ~13–20 | caps at 20 and ignores limit |
| ListenBrainz | ~100 | listening-session co-occurrence; needs a MusicBrainz id |
| Last.fm | up to 100 | scrobble-neighbour similarity, with a 0–1 match score |
Measured hits against a seed's own reference playlist: Deezer 6–9 of 20, ListenBrainz 0–15 of 100, Last.fm 14–18 of 100. No single source is close to sufficient, which is the reason for fusing them.
findDeezerArtist() takes 10 search results and prefers an exact normalized
name match. Deezer ranks search by popularity rather than name match and
returned a novelty act for "Four Tet" that contaminated every playlist built
from that seed — do not reduce this back to limit=1.
Artist suggestions (searchArtists() and src/ArtistTypeahead.tsx) search
Deezer and Last.fm in parallel. Results are interleaved, deduplicated by
normalised name, and sorted with exact matches first, then prefixes and
substrings, before limiting to six. A failed source does not discard the
other's matches; each source has a four-second timeout. Last.fm-only entries
use a neutral avatar instead of requesting a nonexistent Deezer image.
Picking a suggestion (by mouse or highlighted Enter) only fills the field;
the rail's New button or a subsequent Enter explicitly starts generation.
Last.fm's artist.getinfo and artist.search return the generic star image
for Radioclit, not the photo visible on its website; artist.getimages rejects
the current API key, so these endpoints are not an artwork fallback.
Regression checks: node --experimental-strip-types --test tests/artist-search.test.mjs.
Several seeds go through fanoutMulti(), which runs fanout() once per seed
and fuses their fused lists together — see "Several seed artists at once" in
the Roadmap.
Fusion #
src/sources.ts fuse() — reciprocal rank fusion at the artist level.
- Ranks fuse, scores do not. ListenBrainz emits raw session counts in the thousands, Last.fm a 0–1 match, Deezer nothing at all. They are not comparable quantities.
- Fusion is at the artist level, not the track level. Fusing tracks across catalogues would mean matching remaster, live and featuring variants between three naming conventions. Fusing artists means comparing names.
k = 8, not the conventional 60. k=60 comes from web search, where result lists run to thousands. Against a 13-item Deezer list it scores rank 1 at 1/61 and rank 20 at 1/80 — near-identical — erasing any preference for close neighbours.
Tracks then resolve from one catalogue (Deezer) via topTracks(), which stamps
a depth per track (0 = the artist's biggest hit, 1 = deepest pulled).
Selection #
src/recommender.ts artistsNeeded()— how many artists to draw. Derives the draw from a spread target, not from the per-artist cap. See "Where it landed".src/recommender.ts spreadTarget()— whereadventurousnesssets that target: 3.0 tracks per artist at 0, 1.2 at 100.artistsNeeded()divides the target track count by this value, so the knob now moves the draw size itself rather than biasing which artists get picked.src/recommender.ts pickArtists()— draws artists from the roster, uniformly at random regardless ofadventurousness.src/recommender.ts selectTracks()— two-pass selector. Its finaldetailsReadyparameter must be false before enrichment, or the bpm, gain and year filters drop the entire pool.knobsandfiltersinsrc/App.tsxmust keep stable identity across renders.shortlistmemoizes on them and the enrichment effect keys offshortlist, so handing either one a fresh object literal re-runs the entire detail pass on every render — several hundred/dz/trackrequests per keystroke elsewhere in the component. They are memoized on individual primitive leaves rather than onpanel.knobs/panel.filtersbecauseuseDialKitrebuilds its whole nested value tree whenever any single dial changes, so the container objects are new on every dial move even when nothing under them differs.src/App.tsxshortlistandresultmemos — all knobs re-select live from a cached candidate pool; only Generate refetches. Artist-level knobs must be applied in theshortlistmemo, not during the fetch, or they go inert.src/App.tsx's enrichment effect (commented "Only the shortlist is enriched") abortsenrichTracks()on cleanup rather than only ignoring its callbacks. Stopping at the callbacks was not enough: a pass that only stops reporting still pays for every request it started. A reload starts one such pass against DialKit's schema defaults, because the persisted panel is not restored until a render later than the first — roughly 170 wasted/dz/trackrequests per reload before the abort; measured after it, a reload issues 10 — enrichment's concurrency — all reported asnet::ERR_ABORTED, and a normal, non-superseded pass issues none. The same applies to knob drags: each intermediate value produced a pass that went on fetching for a shortlist already gone.- The same effect clears its status line 600ms after the pass ends rather than
immediately. React batches a clear issued in the same microtask as the last
per-track write, so the line jumped from one short of the total straight to
blank and the finished count never painted. Measured on a 30-track detail
pass:
30/30renders and is held for 597ms. The delay is a minimum display duration, not a wait for anything — the pass is already over — and a single frame, while enough to beat the batching, is too short to read.
Measurement #
reference.txt holds 1375 tracks across 15 real playlist exports, rebuilt by the
scripts in tools/ (see tools/README.md). Two of the original 16 exports were
byte-identical in artist content (both The Field); playlists.md keeps the
dropped URL as a comment explaining why, so a rebuild does not bring it back.
Its texture is a long tail: 57–92 distinct artists per 85–99 tracks, about 1.24
tracks per artist, with most artists holding exactly one track.
src/harness.ts bundles with esbuild and runs under Node against a live dev
server, overriding globalThis.fetch to prefix relative URLs with the server's
origin:
npx esbuild src/harness.ts --bundle --platform=node --format=esm \
--packages=external --outfile=/tmp/h.mjs
MODE=confirm SEEDS='Four Tet,The Field' node /tmp/h.mjs
Modes: single, confirm (config comparison), sweep, funnel (per-stage
retention), ranks (where reference artists sit in the fused order),
knob-inertness, shape, draw, sampler, texture.
Three rules that measurements here have repeatedly turned on:
- Score each playlist against its own seed, never against the union. The 15 playlists come from different seeds; union scoring is meaningless.
- Three shuffles cannot resolve a config comparison. Per-seed precision has a standard deviation of 3.4 points across 3-shuffle means. Use 10 or more and report a confidence interval; anything inside the band is "no measurable difference", not a winner.
- Girl Talk is excluded from tuning means, not from the file. Its artist
field is largely YouTube channel handles (
FrederickBarrVEVO,spacemutant,GiraffeSwordsman) and remixer bylines, which no similarity source can return; its ceiling is set by the reference data, not the pipeline. The exclusion is enforced by code rather than left to memory:reference.txtcarries a# excluded-from-scoring:line for it,parseReferenceFile()reads that intoPlaylistRef.excluded, andscorablePlaylists()filters it out of every mean over playlists. Per-playlist diagnostic tables (runFunnel(),runRanks(),runSingle()) still show its row — only the aggregates drop it.
What worked, what didn't #
Where the quality problem actually was #
A funnel measurement settled a long-running ambiguity: track selection loses
nothing. selectTracks() retains 100% of reference artists from shortlist to
final playlist on all 15 scorable playlists. Two multiplicative ~0.45 losses sit
upstream — the pool truncation and the artist draw — and 0.45 × 0.45 explains
the end-to-end result. All ranking work belongs upstream of track selection.
Rejected: a bigger candidate pool #
Raising ARTIST_POOL monotonically destroys precision — 10 to 14 points worse
at 450, far outside the noise band, at every per-artist cap. Reference artists
concentrate in the head of the fused ranking; artists past rank ~90 are worse
material, so a wider roster only dilutes what the draw selects from.
Rejected: a more selective artist sampler #
A rank-weighted sample without replacement, and a deterministic top-N sampler representing the ceiling of rank-only selectivity, were both statistically indistinguishable from the shipped random draw at a 90-artist pool. Rank order inside that roster carries no usable signal, so there is nothing for a smarter sampler to exploit.
Rejected: tag-vector reranking #
Cosine similarity over Last.fm top tags lowered precision on every seed tested
(Four Tet 38%→21%, The Field 35%→19%, Fela Kuti 26%→24%). artist.gettoptags
returns only ~10 broad tags per artist, so the vectors clustered candidates by
genre rather than by scene. Revisiting this needs a higher-dimensional tag
source — MusicBrainz tags carry vote counts, but that endpoint is rate-limited
to 1 request/second.
Kept: second-hop expansion #
src/sources.ts expandSecondHop() asks the closest first-hop neighbours who
their neighbours are, with contributions discounted by head rank. It more
than doubles reach — reference artists present anywhere in the candidate
list go from 20 to 44 on Four Tet.
src/harness.ts CONFIRM_CONFIGS carries a shipped defaults, secondHop off
row for a direct comparison against the shipped pipeline. MODE=confirm SHUFFLES=12 over the default seeds (Four Tet, The Field, Fela Kuti), scored
against the pruned reference set, mean and 95% confidence interval across 12
shuffles:
| config | seed | precision | recall |
|---|---|---|---|
| shipped defaults | Four Tet | 26.1% ±1.0 | 24.6% ±0.8 |
| shipped defaults | The Field | 23.6% ±1.9 | 24.2% ±1.3 |
| shipped defaults | Fela Kuti | 26.8% ±0.9 | 20.9% ±0.9 |
| secondHop off | Four Tet | 20.4% ±0.6 | 19.9% ±0.7 |
| secondHop off | The Field | 23.3% ±1.2 | 24.0% ±0.8 |
| secondHop off | Fela Kuti | 27.8% ±1.2 | 21.0% ±0.8 |
Mean across seeds: shipped defaults 25.5% precision / 23.2% recall, secondHop
off 23.8% / 21.6%. Second-hop expansion earns its place and stays on at
secondHop 0.6.
The gain concentrates in one seed. On Four Tet it is decisive and the confidence intervals do not overlap (precision 26.1 vs 20.4, recall 24.6 vs 19.9 with it off). The Field is unchanged either way. Fela Kuti is a wash — precision is nominally better with second-hop off (27.8 vs 26.8) but the intervals overlap. A seed whose first hop already returns a rich roster gains nothing from a second one, which is why this is evidence the mechanism should stay on rather than evidence it helps everywhere.
The earlier reading of "slightly harmful" (precision 22.6% → 21.8%, recall
11.8% → 9.8%) came from a 3-shuffle comparison under the pre-spread-target
pipeline — the same 11.8% recall that artistsNeeded()'s note cites as the
baseline before the spread target raised it to 19.7%, and that today sits
around 23%. A 3-shuffle mean cannot resolve a difference this size; that is
the whole reason SHUFFLES and the confidence interval exist.
The same run puts no-listenbrainz hop=0.6, hopCount=30 nominally ahead of
shipped defaults overall (26.1% precision / 23.7% recall vs 25.5% / 23.2%),
but its per-seed intervals overlap the shipped ones everywhere, so it is not a
demonstrated winner — unresolved, not an improvement.
Penalising rather than rewarding breadth of hop2 agreement, on the theory that artists many heads agree on are generic to the neighbourhood rather than specific to the scene, is still an open experiment.
Knob findings #
devotionis per-artist depth, not global obscurity. It yields deep cuts by known artists rather than hits by unknown ones. Average track depth moved from 0.04 to 0.55 when catalogues went from 5 to 25 tracks deep.adventurousnesssets artist spread, not obscurity. It moves where the draw lands on the tracks-per-artist span viaspreadTarget(): at 0 the draw is ~40 artists from an ~85-artist roster at ~2.8 tracks each; at 100 it is close to the whole roster at ~1.5 tracks each. Distinct artists in the final playlist rise from ~35 to ~65 across that span. Precision stays flat (roughly 21-27%, inside the ~3.4-point per-seed noise), and recall rises monotonically (~11-15% at 0 to ~20-22% at 100) because a wider artist set covers more of a reference playlist's artists — a property of the recall metric, not evidence that spread is better. The tail of the fused ranking is still less similar, not less famous — at adventurousness 90 the median artist audience rose to 1.6M — so use the popularity filter for obscurity.popularityPenaltyonly bites when the draw is well below the roster size. At 15 drawn artists it cut median audience from 443k to 114k; at 40 drawn from 85 it did nothing. Atadventurousness100 the draw is essentially the whole roster, so reordering it changes nothing the draw takes — penalty rows at 0, 0.5 and 2 came back byte-identical. It only has something to bite on at loweradventurousness.- "Require two sources" is close to unusable at one hop — only 4 of 109 candidates cleared it, since Deezer returns ~13–20 against ListenBrainz's 100. The agreement knob defaults to a ranking boost instead.
Data gotchas #
src/sources.ts json()retries both HTTP 429 and HTTP 200 responses carrying a Deezer error object. Without the second case, enrichment silently lost ~47% of itsyearandgainfields.src/sources.ts enrichTracks()reports progress through two callbacks with different cadences:onBatchfires every 25 tracks plus once at the end, carrying a full snapshot, because it drives aMaprebuild;onTrackfires per track, since it only feeds a status string. Individual request failures are tolerated —pool()'s per-item catch leaves a failed track's slot null rather than failing the whole pass.enrichTracks()andpool()also take an optionalAbortSignal, threaded down tofetch():pool()'s runners stop taking further items the moment the signal trips, so a cancelled pass costs at most the requests already in flight — ten, at enrichment's concurrency — rather than the whole list, and items never started are left null exactly as a failed one is. An aborted fetch rejects beforejson()'s retry logic, so a cancelled pass does not back off and ask again;pool()'s per-item catch swallows that rejection too, so a cancelled pass neither rejects nor raises the "Reading track details failed" error line.- Deezer
bpmis ~67% populated, hence theallowUnknownBpmfilter.gain,year,durationandrankare fully populated. - The era filter measures the pressing, not the recording. Deezer
release_datepoints at whichever remaster its top-tracks endpoint returns, so a 1994 song can present as 2015. MusicBrainz first-release-date would fix it. - Deezer preview URLs are signed with an
hdnea=exp=...token that expires in roughly a day, so a restored playlist'spreviewfield can be dead.
Playback #
Full tracks via YouTube. Deezer's 30-second previews were the earlier stopgap and are gone.
worker/youtube.ts searchVideoId()— resolves(artist, title)to a videoId by requesting YouTube's search page server-side and walking the embeddedytInitialDatafor the firstvideoRenderercarrying both avideoIdand alengthText(which skips live streams and premieres). Exposed asGET /yt?artist=&title=.- Results cache in Workers KV (
YT_CACHE): 30 days for a hit, 1 day for a genuine miss. A scraper failure throwsScrapeError→ HTTP 502 and is never cached, so broken markup is loud on the next request instead of being frozen in as a fact about the track. - The video is off-screen, not hidden. The player block shows the artist's
picture; the iframe still exists at the 200×113 the player is constructed
with, parked by
.player-offscreenatleft: -9999px.display: noneand a zero-sized frame are both grounds for a browser to stop the audio, so neither is available. Audio from the off-screen frame is confirmed by ear in Chrome and Firefox — headless cannot check this, because no YouTube video plays there at all. The positioning belongs to that wrapper rather than the frame becauseYT.Playerreplaces the element it is handed, taking any class on it along. The artwork itself comes from/dz/artist/{artistId}/image, which Deezer redirects to its CDN and the Worker proxies, so no artwork has to be carried through the pipeline or stored with an archived playlist. src/youtubePlayer.ts usePlaylistPlayer()— one YouTube IFrame Player instance, event callbacks reading state through refs so they never go stale. Resolves the current track plus two ahead as playback advances.ENDEDandonErrorboth advance, marking a failed row visibly rather than skipping silently.usePlaylistPlayer()also exposesresolveAll(list, onProgress), for exporting a whole playlist rather than playing it. It resolves every track the lazy path above has not already looked up and returns a snapshot of the whole id map, sharing thevideoIdsandattemptedrefs with lazy playback so nothing is looked up twice and a track that already returned a terminal 404 is not retried. Lookups are capped at four in flight (RESOLVE_ALL_CONCURRENCY): each is a scrape of youtube.com from the Worker, and a ~100-track playlist fired off at once is a burst against both YouTube and the Worker's subrequest budget. The single-track fetch is factored intolookup(), shared by both callers, so a track that resolves or fails during an export moves playback along exactly as a prefetch would.resolveAll()never picks a track, so it does not build the YouTube embed — exporting still costs zero YouTube telemetry connections, which is why the embed is only built on first play. The Worker caches ids in KV for 30 days, so a playlist's second export is immediate; the cost falls on the first. Insrc/App.tsx,exportPlaylist()drives this: the XSPF and JSPF buttons are async, are disabled while a pass runs, and report a running count through the status line — only while something is still outstanding, since an already-resolved playlist reports done on its first callback. Measured against a production build in a headless browser, 15-track playlist: 15 of 15 tracks got a YouTube first-location, 15/ytrequests with never more than 4 in flight, 0 new requests on a second export, and 0 requests to any youtube.com host.- The embed is built on the first play, not on mount: it opens telemetry requests that never settle, so a page that builds a player nobody asked for keeps outstanding connections for a visitor who only wanted the playlist.
- The videoId cache, the attempted-lookup set and the per-track status map are
keyed by track id, and playlist invalidation keys on the id sequence rather
than the
tracksarray's identity. The detail pass rebuilds that array on every batch and can reorder it, so keying on either position or identity stopped playback mid-track for anyone who pressed play before it finished.
This is deliberately brittle: parsing search markup breaks whenever YouTube changes it, and it is contrary to YouTube's terms of service. The alternatives were worse — see below.
Rejected playback and export routes #
- YouTube Data API — 100 search queries per day for the whole application. One 100-track playlist per day, not per user.
- Client-side search of YouTube or SoundCloud — YouTube search HTML,
YouTube Music's
youtubeiendpoint and SoundCloud'sapi-v2all work incurlbut send no CORS headers, so none run in a browser. This is what the Worker exists to get around. - Odesli — CORS-open and takes Deezer IDs, which the pipeline already holds, but never returned a YouTube or SoundCloud link from Deezer input across three probes.
- SoundCloud — the widget needs resolved track URLs, and API registration has been closed for years.
- Scraping YouTube or SoundCloud playlist pages while private — both serve
consent or stub pages to server-side fetches; YouTube playlist contents render
client-side and
ytInitialDatacomes back as a ~680-char stub. Making the playlists public fixed both, which is howreference.txtwas built. - YouTube playlist RSS (
feeds/videos.xml?playlist_id=) — caps at ~14 entries. Useful only as a public/private reachability check.
State and export #
src/session.ts— persists the fetched candidate pool, per-track enrichment and the shuffle token to localStorage, so a reload redisplays the same playlist without refetching. Versioned; a mismatched, corrupt or malformed payload discards cleanly. OnQuotaExceededErrorit retries without the bulk cache and then removes the key rather than leaving a stale payload that looks current, surfacing the outcome in the UI.- Knob and filter state is handled by DialKit's own
persist: trueunderlocalStorage['dialkit:morelike'], not duplicated here. src/playlists.ts— the collection: every playlist that has been on screen, automatic or saved. Anoriginofautomaticis what generating a new one used to end outright — a generate replaces the candidate cache, and this is what makes that non-destructive — capped at 20 entries, evicted oldest first. Anoriginofsavedis filed by a deliberate save and is never evicted by that cap or dropped to satisfy a storage quota before every automatic entry is gone; promoting an automatic entry to saved is a field change (promoteToSaved()), not a copy into a second store. Persisted underlocalStorage['morelike:playlists'], version-tagged and discarded wholesale on a version mismatch likesession.ts, except that an absentmorelike:playlistskey falls back to reading the previousmorelike:historyshape (PastPlaylist[], version 1) and lifting its entries in as automatic playlists named after their seed — so a browser that already has playlists archived under the old key does not read back an empty collection the first time this ships. A separate store rather than a field onsession.ts, because the two have different lifetimes: a session is one automatic slot overwritten on every change, and this collection is otherwise-automatic entries plus deliberately-kept ones living side by side. An entry holds the finished playlist and the settings (knobs, filters, weights, shuffle) it was generated under, not the candidate pool that produced it — restoring one puts its tracks on screen to play, export, reshuffle or refilter, but keeps no pool, so a fresh Generate still goes back to the network. A pool per entry would cost roughly a megabyte each and exhaust the origin's ~5MB quota within a handful of generates.addPlaylist()stripspreviewfrom filed tracks, since Deezer's signed preview token expires about a day after the fetch and would be a dead link by the time anybody looks back at it, and it is the largest field on aTrack;linkdoes not expire and stays. It also skips a playlist identical to the newest entry of the same origin, so pressing Generate twice does not file the same untouched playlist twice, and pressing Save twice on the same playlist does not file two copies of it.savePlaylists()persists as much as fits, dropping automatic entries before any saved one and the oldest of whichever kind is being dropped first, and returns what was actually stored so state and storage cannot drift apart.src/recommender.ts—toCsv(),toText(),toM3u(),toXspf(),toJspf(), and YouTube/SoundCloud search deep links.toM3u()omits tracks with no preview URL, since an#EXTINFwith no following URL is malformed;toXspf()andtoJspf()never drop a track, because both formats permit a track with no<location>, and the status line reports how many of those there are.toXspf()'s child element order inside<track>follows the XSPF spec:location,identifier,title,creator,album,duration.Track.durationis in seconds, so both emitters convert to the milliseconds XSPF/JSPF expect.identifieris the Deezer track link.locationsOf()emits locations in preference order: the whole track on YouTube where it resolved, then Deezer's 30-second preview clip, then the Deezer track page — a player takes the first location it can handle, so the full track goes first.toXspf()andtoJspf()take an optional third argument, aVideoIdsmap of track id to YouTube video id; without one an export is still valid, just one whose locations stop at Deezer.hasLocation()in the same file is what the UI counts locationless tracks with.src/App.tsx—displaced, a ref, holds the live playlistgenerate()is about to replace (its seed, tracks and the settings that produced them), andgenerate()files it first viaaddPlaylist()/savePlaylists()withorigin: 'automatic'. It is the live playlist whether or not it is the one on screen, because generating while looking back at an earlier playlist still replaces the live one.generate()takes an optional seed override, for the per-row button that sets the artist field and generates in the same tick — the panel value it just wrote is not readable until the next render.viewingselects an entry from the collection — automatic or saved — to show in place of the live one, andshownprefers it. An effect keyed on[cache, shuffle, knobs, filters]ends the look back whenever something produces a new live playlist, because a knob that visibly does nothing is worse. Enrichment is deliberately not in that list: it changesdetail, not the playlist, and would otherwise yank the view away mid-read. Save playlist, alongside Generate and Reshuffle, is a dial-panel action rather than a button in the header, since it is another action that authors the playlist rather than exports it; pressing it while viewing an entry from the collection promotes that entry to saved (promoteToSaved()), and pressing it while showing the live playlist files a new saved entry (addPlaylist()withorigin: 'saved') — an explicit act either way, never implied by Generate or Reshuffle. The saved-playlists list renders each entry's name, track count and save date, and opens (viewingId), renames (renamePlaylist()) and deletes (deletePlaylist()) it; the "Earlier" row keeps its original job of opening an automatic entry for a look back. Each track row has a button that seeds a new playlist from that row's artist; it sits in the links column so the row's grid columns are unchanged and the left of the row stays playback-only. Rows themselves are not clickable, and nothing plays without a press on a play button. The panel usesuseDialKitControllerrather thanuseDialKit, forsetValue(), which the row button needs to move the seed field;panelis nowdial.values, and the values are otherwise identical.
DialKit has no two-handle range control, so every range filter is a pair of
sliders; it also has no title for the shell it renders around multiple panels,
which is why there is one panel named morelike with folders inside.
Where it landed #
Drawing artists to a spread target rather than to the per-artist cap was the change that mattered. Measured over 10 shuffles across 7 reference seeds:
| before | after | reference | |
|---|---|---|---|
| distinct artists | 36.2 | 64.9 | 67.2 |
| tracks per artist | 2.73 | 1.54 | ~1.37 |
| artists with one track | 4.5% | 53.6% | 78.7% |
| recall | 11.8% | 19.7% (+8.0 ±2.0) | — |
| track precision | 25.5% | 24.2% (−1.3 ±2.4) | — |
| extra requests | — | none | — |
Recall rises by two-thirds and the playlist's texture lands close to the
reference, at no measurable precision cost and no additional network calls.
ARTIST_POOL stays at 90 — it is not the binding constraint.
Evidence the approach works when the pool is right: a Fela Kuti seed yields Gyedu-Blay Ambolley, Vaudou Game, Ofege, Orchestre Poly-Rythmo de Cotonou and Ebo Taylor.
Known gaps #
- Reference data carries channel-name noise beyond Girl Talk, which is already excluded from scoring. Playlists still in the scored set have it too (for example "Manu DIBANGO Officiel" never matches "Manu Dibango"), depressing recall uniformly. It does not bias comparisons between configurations.
- Push over SSH, not HTTPS.
originisgit@tangled.sh:burrito.space/morelike. The HTTPS form (https://tangled.org/burrito.space/morelike) redirects the push toknot1.tangled.shand asks for credentials interactively;tangled.shis the SSH endpoint that fronts the knot, and port 22 on the knot host itself refuses connections.
Roadmap #
The entries below are ordered as the work is to be done, and the first three build on each other.
Saved playlists in browser storage shipped: a playlist is a thing you save,
find again in a list, open for playback, rename and delete — not a version of
one live playlist. src/playlists.ts holds the model this settled on: a
playlist has a stable id, a user-editable name, a created and an updated time,
its seed and settings, and its tracks; several exist at once, distinguished by
an origin of automatic (generate()'s own non-destructive look-back, capped
at 20, evicted oldest first) or saved (kept for good). See "State and
export" below for the detail; the items after this one build on that model.
A new playlist defaults to its seed's name; uniqueName() in
src/playlists.ts appends the next free number when that name is already
taken — Four Tet, then Four Tet 2 — compared trimmed and
case-insensitively, so two playlists from the same seed never collide in the
list.
Pruning a playlist shipped too, by hand and automatically. A removal is keyed
by track id rather than position — matching how videoIds, attempted and
statuses in src/youtubePlayer.ts are already keyed — since the list can be
reordered by the playback shuffle. Hand-pruning a saved playlist is a write:
pruneTrack() in src/playlists.ts edits the stored record and bumps
updatedAt, so mergeRemote() in src/playlistSync.ts treats it as an edit
another device's stale copy can't undo. The live playlist has no record to
write to — result in src/App.tsx is a useMemo rebuilt on every
enrichment batch and every knob or filter move — so its prunes live in
removedIds, a set of ids applied wherever result (and a viewed entry's
own tracks) are built, and cleared only on a fresh Generate. A track that
turns out to be unplayable prunes itself the same way: usePlaylistPlayer()
in src/youtubePlayer.ts owns playback and only surfaces the failing id
through onUnplayable, never touching playlist state itself, and the removal
does not interrupt the advance to the next track. An automatic prune never
edits a saved playlist's stored record even while it is the one playing —
one device's resolution failure is not treated as grounds to change a record
that syncs elsewhere.
A reload reopening the playlist last played or edited, rather than the last
one generated, shipped too. saveLastViewed() in src/session.ts writes a
playlist id (or null for the live playlist) to its own
localStorage['morelike:lastViewed'] key whenever a track starts playing, or
a playlist is renamed, pruned, reshuffled, refiltered or saved; loadLastViewed()
reads it back on mount to set viewingId before anything is generated. A
separate key rather than a field on session.ts's larger payload, since the
two are written at different times and for different reasons.
The player block always showing an artist picture, even before a generate or
a restore has produced one to show, shipped too. src/App.tsx falls through,
in order: the artist actually playing; for a saved or shared playlist (which
keeps no seed id) its first track's artist; for the live playlist its first
seed; and last, a stand-in resolved once per name through findDeezerArtist()
for whatever is currently typed in the seed field, since a returning browser
can restore that field from DialKit's own persisted panel without the session
that would have restored a candidate pool alongside it.
The interface went through a pass toward looking like one thing rather than a panel bolted onto a demo: a single heading, the settings panel (seed, knobs, filters, Generate, Reshuffle, Save playlist) restyled to match the rest of the app instead of DialKit's own look, saved playlists compacted to one line each, and "About & roadmap" moved to the bottom of the left column. The settings panel then moved out of that column entirely into its own right-hand column, matching the left column's width and border on the opposite edge; on a phone each column becomes its own slide-in drawer instead of narrowing in place.
-
1. atproto sign-in, with playlists in the user's own PDS. Shipped — the app keeps working signed out, on browser storage alone. Signed in, playlists become records in an account the user controls, which is what makes them reachable from a second device and what gives sharing a real address. Sign-in itself is a browser-only OAuth client — PKCE and DPoP, no backend of this project's own — modeled on
.../szzt/apps/publisher/src/main.ts; seesrc/atproto.ts.isAuthCallback()reads the redirect back from the PDS out oflocation.hash, sinceBrowserOAuthClientdefaults toresponseMode: 'fragment'rather than the query string. An optimistic hint written tolocalStorage['morelike:auth-hint']paints the signed-in shell from a storeddid/handlebeforeoauth.init()resolves the real session, so a returning visitor never sees the signed-out form flash before the real one. The handle field suggests as it's typed, against the public, unauthenticatedapp.bsky.actor.searchActorsTypeahead, which costs no CORS preflight and needs no account of its own to query;signingIndisables the field and relabels it while a sign-in is in flight, and a rejectedoauth.signIn()— an unresolvable handle — surfaces assignInErrorrather than a silent failure.The hard part was not the login but the sync: two devices editing the same collection need a merge rule, and it settled on last-writer-wins by
updatedAt, inmergeRemote()insrc/playlistSync.ts— a newer remote record replaces the local copy, a tie or older one does not, and a device's first sync for an account adopts every remote entry unconditionally rather than treating an empty local collection as everything having been deleted. A delete is written as a tombstone record at the same rkey (deletedAt,updatedAt, no content fields) rather than record absence, since absence alone can't distinguish "never synced" from "deleted" — and a device that deleted a playlist itself won't resurrect it from a remote copy it hasn't seen the tombstone echoed back to yet. Depends on the record shape insrc/playlists.ts. The record type this writes into a signed-in user's repo,PLAYLIST_COLLECTIONinsrc/playlistSync.ts(space.burrito.morelike.playlist), is documented as a lexicon schema atlexicons/space.burrito.morelike.playlist.json, published as acom.atproto.lexicon.schemarecord bytools/publish-lexicons.js(seetools/README.md). Lexicon-1 has no floating-point field type, so every fractional value —popularityPenalty, the four fusion weights, and each track'sbpm,gainanddepth— is scaled byTHOUSANDTHS = 1000on the way in and divided back down on the way out (toThousandths()/fromThousandths()insrc/playlistSync.ts);assertAllIntegers()walks a record before every write and throws if anything slipped through unscaled. -
2. Sharing a playlist. Shipped, for a playlist that is a record in a PDS — one saved from a signed-in account (see item 1). What a share link carries had answers differing in kind rather than in degree:
- The settings. Seed, knobs, filters and the shuffle token are small
enough to sit in a URL comfortably. They do not reproduce a playlist:
generate()refetches the candidate pool from catalogues that change upstream, so the same parameters yield a different playlist later. This is a URL for "here is how I had it set", not for "here is my playlist". - The tracks. Exact by construction, and the only form that survives
upstream drift. A 100-track playlist as
(artist, title)pairs is a few KB — beyond a comfortable bare URL, workable compressed. Deezer track ids would be far more compact but would tie a shared playlist to one catalogue, which is the opposite of keeping the output platform-neutral. - A reference to a stored record sidesteps the size question, and is what the PDS entry above makes available. This is what shipped: a share link carries an address, not a playlist.
A link is
/share/<did>/<id>— the signed-in author's DID and the playlist's own id,encodeURIComponent()-escaped (adid:has colons before its first/, which a browser's URL parser percent-encodes inlocation.pathnameto keep the segment from reading as a URI scheme).useRoute()insrc/App.tsxparses it into a third route alongside/and/about, andsrc/shareResolve.tsresolveSharedPlaylist()resolves it: a DID document fetch (mappingdid:plc:*via the PLC directory anddid:web:*via its.well-knownor path form, mirroringpackages/gateway/src/did-resolver.tsandpds-resolver.tsin theszztproject) finds the author's PDS, then onecom.atproto.repo.getRecordreads the playlist. Both are plain, unauthenticatedfetch()calls — a public repo takes no auth to read — so opening a share link needs no sign-in, and pulls in neither@atproto/apinor the OAuth client that sign-in itself loads on demand (seesrc/atproto.ts'sloadOAuthClient()).Because the link is an address, not a copy, it reflects the playlist as it is now, not as it was when shared: an edit at the source shows up on reopening the same link, and a deleted playlist's link stops resolving (the record's tombstone, or its absence, both read as "no playlist here" rather than a crash). Opening someone else's playlist this way never writes to the browser's own collection — it is a view, playable in place, with an explicit "Save a copy" action for anyone who wants to keep it.
Whichever is encoded,
previewmust not be: the signed Deezer token expires within a day, which is whyaddPlaylist()insrc/playlists.tsalready strips it before filing a playlist. - The settings. Seed, knobs, filters and the shuffle token are small
enough to sit in a URL comfortably. They do not reproduce a playlist:
-
3. A custom domain. Undecided, and deliberately left open. It matters before sharing links are handed out, since a shared URL is the one artifact that outlives a rename.
-
4. Several seed artists at once, not one. Shipped. The artist field accepts a comma-separated list;
src/sources.tsfanoutMulti()runs each name through the existing single-seedfanout()(unchanged, still what the harness and a single-seed generate use directly) and fuses the seeds' already-fused candidate lists into one, the same rank-fusion question the three sources answer one level up — an artist similar to every seed outranks one similar to only one, rather than each seed's list being drawn from in proportion. Seeds run two at a time to keep the per-seed request multiplication from compounding Deezer's session limit and MusicBrainz's rate limit.Playlist.seedinsrc/playlists.tsstays a plain string — several names join with,rather than the field widening to an array, since it is already written to browser storage and to synced repos. Unmeasured:src/harness.tshas no multi-seed mode, so there is no precision/recall number for this against the reference playlists, only the qualitative check that a two-seed run's candidates visibly draw from both neighbourhoods. -
A statically deployable front-end, with what is left as modular edge functions. Every part of the pipeline that decides what goes in a playlist already runs in the browser and touches no network — all of
src/recommender.ts, plusfuse(),penalisePopularity()andpool()insrc/sources.ts. The Worker holds no recommendation logic. What keeps the front-end tied to an origin is three things, not five:/dzis a CORS shim and nothing else — stateless, holds no secret, needs no bindings./fmholdsLASTFM_API_KEY. A static front-end can never call Last.fm directly without exposing it. The alternatives are a per-user key kept beside the session inlocalStorage, or dropping the source — which is not free, since Last.fm is one of the two sources weighted on by default./ytneeds a non-browser HTTP client for the scrape inworker/youtube.ts searchVideoId()and shared state inYT_CACHE. A KV namespace can be bound to any number of separately deployed Workers by id, so splitting this out costs nothing cache-wise.
/lband/mbare already gone:listenBrainzSimilar(),findMbid()andartistMeta()insrc/sources.tscall the upstreams through theLBandMBbase constants there. That removed the onlyheadersentry inroutes, and gave MusicBrainz better rate-limit headroom, since it throttles one request per second per IP and proxying put every user behind shared Worker egress IPs. The<img>insrc/App.tsxneeds no proxy at all, images not being subject to CORS.Pointing the client at a configurable API base is one change, not ten: a prefix-aware resolver inside the module-local
json()insrc/sources.tscovers its eight callers untouched, leavinglookup()insrc/youtubePlayer.tsand that<img>to fix individually. The harness'sfetchmonkey-patch stays either way — it also does the disk caching.The constraint to hold on to: requests must stay CORS-simple.
json()sends no custom headers and noContent-Type, so none of the several hundred requests a generate fires ever preflights. OneAuthorizationorX-Api-Keyheader would double every one of them.Unverified: whether MusicBrainz considers a browser's own User-Agent acceptable as policy rather than merely in practice — their documentation does not address browser applications.