Neighborhood Feedgen #
An explainable Bluesky custom feed: it recommends posts liked by people whose recent likes overlap with yours. It is a bounded, CPU-only two-hop walk: your likes → their likers → their recent likes.
It runs with Node 22+ and has no runtime npm dependencies. The storage engine is Node's built-in SQLite (node:sqlite).
What is implemented #
app.bsky.feed.getFeedSkeletonat/xrpc/app.bsky.feed.getFeedSkeleton- Per-viewer ranked skeletons, opaque cursor pagination, feed URI validation, health endpoint
- Persistent SQLite indexes for posts, likes, deletes, and Jetstream cursor state
- Live ingestion from Bluesky Jetstream for
app.bsky.feed.postandapp.bsky.feed.like - Bounded neighbor walk, inverse-popularity weighting, exponential recency decay, per-author cap, deterministic ordering
- Unit tests for ranking and XRPC response shape
Run locally #
cp .env.example .env # then export values, or use your process manager's env file
FEED_URI='at://did:plc:YOUR_DID/app.bsky.feed.generator/neighborhood' \
FEED_DID='did:plc:YOUR_DID' \
ALLOW_UNVERIFIED_DEV_DID=true \
npm start
Development request (the header is intentionally available only with ALLOW_UNVERIFIED_DEV_DID=true):
curl -H 'x-feedgen-user-did: did:plc:viewer' \
--get 'http://localhost:3000/xrpc/app.bsky.feed.getFeedSkeleton' \
--data-urlencode 'feed=at://did:plc:YOUR_DID/app.bsky.feed.generator/neighborhood'
Run npm test and npm run demo for an offline proof.
Seven-day targeted bootstrap #
This is the development backfill intended for a single accurate feed, not a whole-network crawl:
npm run backfill -- --actor callie.on-her.computer
It resolves the actor's DID/PDS, reads only their last seven days of public app.bsky.feed.like records, fetches likers of those seed posts, then reads only the last seven days of likes from the discovered neighbors. It hydrates only posts in that bounded two-hop subgraph. Limits and concurrency are in .env.example; reduce them when iterating.
The precise createdAt timestamps come from public repository records. This is why the AT Protocol's per-repo record endpoint is preferable here to trying to infer when an old post was liked from its post timestamp.
Publishing / production boundary #
Create an app.bsky.feed.generator record on the account that owns FEED_URI, with its did set to this service's DID and serviceEndpoint set to the public HTTPS base URL. The service must expose GET /xrpc/app.bsky.feed.getFeedSkeleton over public HTTPS.
ALLOW_UNVERIFIED_DEV_DID is not production authentication. The endpoint needs the requesting DID to rank per viewer, and the feed generator starter kit recommends verifying its signed JWT. Wire in @atproto/xrpc-server's auth verifier before public deployment, replacing src/auth.js; the rejection-by-default behavior avoids silently trusting an unverified sub claim.
Jetstream is deliberately used for the initial implementation because it carries decoded JSON records. For a large public feed, replace it with full com.atproto.sync.subscribeRepos ingestion plus repo resync/backfill, while retaining the same Store methods and ranker.
Tuning #
The most consequential knobs are FANOUT_LIMIT (default 500), NEIGHBOR_LIMIT (1500), CANDIDATE_MAX_AGE_HOURS (168), and RECENCY_HALF_LIFE_HOURS (18). The ranker exposes contributing neighbors and shared seed likes, so add an authenticated diagnostics endpoint later rather than making the serving endpoint verbose.