Glean #
A social RSS reader built on the AT Protocol. Sign in with your Bluesky/Atmosphere account, subscribe to feeds, and discover what like-minded readers are into.
Try it at glean.at.
Your subscriptions live as records on your PDS. You own them. If Glean goes away, your data doesn't.
What you get #
- RSS, Atom, and JSON Feed support with a keyboard-driven reading interface
- Highlights, notes, tags, and ratings on articles
- margin.at annotations displayed alongside glean annotations
- A trending page showing what's popular across all users
- Feed and people recommendations based on reading overlap
- Daily digest with an AI-generated summary of your unread articles
- OPML import and export
- Sign in with Bluesky / Atmosphere account; no new account needed
How recommendations work #
Glean looks at what you and other users subscribe to, read, and like to suggest feeds, articles, and people you might enjoy.
Feed suggestions come from readers who share your subscriptions. If a lot of people who follow the same blogs as you also follow a blog you haven't seen, that blog shows up as a recommendation. The system also considers which articles you've liked, whether you follow the person on Bluesky, and how popular the feed is overall.
People suggestions are split into two groups: "Your network" shows people you already follow on Bluesky who share your reading habits, and "Discover new readers" surfaces readers you don't follow but who have overlapping subscriptions and likes.
Dismissals keep things tidy. If you dismiss a recommendation, it won't come back. If a suggestion sits ignored for more than 5 days, it's automatically removed so newer recommendations can take its place.
Cold start. If you're new and have fewer than five subscriptions, Glean shows feeds from people you follow on Bluesky alongside popular feeds from the community, so there's something to explore right away.
The system improves over time: as you subscribe to feeds and like articles, Glean learns which signals matter most to you and adjusts accordingly.
Run it #
Glean is a Cloudflare Worker serving a server-rendered SvelteKit UI in front of one Cloudflare Container (the Go AppView). Relational state lives in one D1 database with flat table names. Recs KNN lives in Vectorize.
One-time setup:
make web-install # root (wrangler) + web (SvelteKit) deps
bunx wrangler d1 create glean # note the database id
bunx wrangler r2 bucket create glean-archive
bunx wrangler vectorize create glean-articles --dimensions 1536 --metric cosine
bunx wrangler vectorize create glean-feeds --dimensions 1536 --metric cosine
bunx wrangler secret put GLEAN_SESSION_KEY
bunx wrangler secret put GLEAN_CLOUDFLARE_ACCOUNT_ID
bunx wrangler secret put GLEAN_CLOUDFLARE_API_TOKEN # Vectorize permissions
# optional: GLEAN_EMBED_API_KEY, GLEAN_LLM_API_KEY, …
wrangler.jsonc declares these under secrets.required, so a deploy with
any of them missing fails immediately instead of shipping a container that
cannot boot. GLEAN_D1_PROXY_URL ships as a var in wrangler.jsonc
(http://glean.proxy).
Recent article bodies live in D1; after GLEAN_ARCHIVE_WINDOW days a
maintenance pass uploads them to R2 and drops them from the row. Reading an
archived article pulls its body from R2 on demand; if the blob is missing (a
pre-R2 article), the server re-fetches the body from the source URL once and
stores it back to R2. The container reaches both D1 and R2 through the
Worker's outbound glean.proxy host (binding proxy, bearer tokens derived
from the session key), so there are no S3 credentials or request signing to
manage.
Then deploy:
make deploy
Remove everything (worker, container, images, D1 with all data, R2, Vectorize):
make teardown
Set GLEAN_FRONTEND_URL / GLEAN_OAUTH_CLIENT_ID in wrangler.jsonc to the
public origin before the first deploy. The OAuth callback is always
$GLEAN_FRONTEND_URL/api/auth/callback.
Local development #
Same topology as production, including D1: wrangler dev serves the D1
binding locally, and the container reaches it through the same proxy. Docker
must be running.
make dev # builds the frontend, then wrangler dev
wrangler dev reads public vars from wrangler.jsonc and secrets from
.dev.vars. Required secret: GLEAN_SESSION_KEY (openssl rand -hex 32).
GLEAN_D1_PROXY_URL comes from wrangler.jsonc; wrangler dev intercepts
the glean.proxy host inside the container's network namespace, same as
production. Add API keys as needed.
Storage bounds #
D1 databases cap at 10 GB, so storage is bounded everywhere: a fetch inserts at
most the 50 newest items per feed and skips items whose publish date is
already past the retention window (maintenance purges those within the
hour), each feed keeps at most 500 articles,
articles older than GLEAN_ARTICLE_RETENTION_DAYS are deleted unless liked or
annotated, article bodies move to R2 after GLEAN_ARCHIVE_WINDOW days (the
row and its summary excerpt stay; reads hydrate from R2, and a body with no
blob is re-fetched from the source once), summaries are stored as a
~1000-character excerpt, unreferenced feeds are pruned hourly, impressions
older than 14 days are dropped, and OAuth login attempts and sessions nobody
completes or refreshes are purged after 7 and 90 days.
Configuration #
All GLEAN_* vars live in wrangler.jsonc (public) or as Worker secrets;
the Worker forwards them to the container. Their meaning and defaults are
documented in docs/specs.md; the code fallbacks live in
main.go.
Documentation #
- Technical specification: architecture, database schema, AT Protocol lexicons, API endpoints, recommendations
- Design system
Stack #
Go, SQLite, SvelteKit, Cloudflare Workers + Containers, TailwindCSS, AT Protocol OAuth.