find text or images in your atproto repo pensieve.waow.tech
pds repo search atproto
TypeScript 52%
JavaScript 43%
CSS 3%
HTML 2%
<1%

README.md

pensieve #

search everything on your PDS. sign in with your atproto account, pensieve walks your public repo, embeds text, images, and video frames into one vector space, and gives you a single search box over all of it.

https://pensieve.waow.tech

how it works #

who are you?  →  atproto OAuth  →  build index  →  search  →  save to your PDS
  • identity — @atcute/oauth-node-client as a discoverable public client. the session lives behind an opaque cookie; the stored OAuth session is used for exactly one write path: saving your own index to your own PDS (repo:tech.waow.pensieve.index + blob:*/*). ALLOWED_DIDS gates building, because indexing is the only thing that spends money.
  • index — built by a durable object, so closing the tab never kills a build. com.atproto.sync.getRepo → @atcute/repo CAR walk → artifact inference (src/artifact.js) → voyage-multimodal-3.5 → turbopuffer, one namespace per DID, one vector column for text and media. images are embedded by URL via the bsky CDN thumbnail (460k px, ~$0.0003 each); video contributes its poster frame; gifs come straight from the PDS blob because the CDN renders them blank. likes, follows, reposts, and blocks are skipped.
  • freshness — the way ken does it: every build re-walks the repo and diffs (id, cid) against what turbopuffer already holds, so only new or changed records are embedded and deleted ones are pruned. blob rows are content-addressed and never re-embed. /api/me compares the stored rev to getLatestCommit so the UI can offer a refresh only when the repo actually moved.
  • search — query embedded with the same model, ANN + BM25 in parallel, reciprocal-rank fused. a media row and its parent record collapse into one result carrying the thumbnail, and constellation backlinks show who referenced a result.
  • memory — your index can be saved to your own PDS as a tech.waow.pensieve.index record (int8 vectors + pointers, chunked blobs): anyone could compute it from your public repo, saving it just spares them the cost. design and the private/shared half: docs/memory.md.

measured #

on did:plc:xbtmt2zjwlrfegqvch7fboei (29,695 records, 16 MB CAR), apple m-series, scripts/bench.mjs:

stage number
CAR walk 29,695 scanned → 9,370 kept + 869 media in 234 ms
voyage images, image_url, batch 16 0.86 Mpx/s at conc 1 → 2.85 Mpx/s at conc 4, flat past 4
voyage text, 96 inputs ~2.2–3.0 s (voyage-4-lite was ~0.5 s; unified space costs that)
turbopuffer upsert ~1.2 s / 96 rows, ~2.1 s / 256 rows
turbopuffer query ~75 ms

the pipeline runs 4 embed→upsert tasks concurrently (workers allow 6 outbound connections). a full first build of that repo is ~$0.35 of voyage; a refresh with no changes is one getLatestCommit call.

develop #

npm install
cp .dev.vars.example .dev.vars   # VOYAGE_API_KEY, TURBOPUFFER_API_KEY
npm run dev                      # open http://127.0.0.1:8787 — not localhost,
                                 # the loopback OAuth client needs the IP
node scripts/bench.mjs --did did:plc:... --passes 5
VOYAGE_API_KEY=... node scripts/bench.mjs --did did:plc:... --embed 96 --conc 4

see docs/OPERATIONS.md for deploy, secrets, and KV.

project map #

src/worker.js      routes: oauth, /api/me, /api/index, /api/build, /api/search
src/auth.js        oauth client, kv stores, cookie session
src/build.js       build orchestration: lease, slices, cancel, stall detection
src/builder.js     the IndexBuilder durable object (thin shell over build.js)
src/index.js       walk → diff → embed → upsert pipeline, search fusion
src/artifact.js    record → {title, body, url, media} inference
src/pack.js        tech.waow.pensieve.index chunks: int8 vectors, cbor entries
src/publish.js     writes the index record + blobs to the owner's PDS
src/public/        static front end (html, css, js — no build step)
lexicons/          tech.waow.pensieve.{index,memory,memoryAccess}
scripts/bench.mjs  reproducible pipeline numbers