coral #
watch named entities form, cluster, and decay on the AT Protocol firehose, in real time.
what it does #
extracts named entities (people, organizations, places, events) from Bluesky posts and tracks how they're discussed together. entities mentioned in the same post form an edge; the cluster structure shifts as conversations rise and fade. an LLM curator names the top clusters every few minutes so the trending list reads as stories ("Trump Budget Fight") instead of bare entity names.
how it works #
- jetstream NER — the Zig backend consumes the jetstream firehose, runs spacez NER in-process, extracts entities
- labeler integration drops spam before it hits the graph (via Hailey's labeler)
- entity graph tracks co-occurrences (entities in same post = edge), computes clusters via union-find
- pheromone edges — edge weights decay exponentially, reinforced on repeated co-occurrence (ant colony optimization inspired)
- surprise trending — entities ranked by statistical surprise vs baseline (z‑like), not raw counts
- LLM curator — Claude Haiku names the top clusters every 5 min (e.g., "Iran Nuclear Talks") and writes a haiku about what's happening
- frontend visualizes entity activity, cluster structure, named groups, and firehose health
history serving #
/history and /simcluster/history show durable entity and topic history.
A bounded background worker materializes responses and serves completed
generations without request-time analytics.
Checksummed response files beside the SQLite database survive restarts; failed
refreshes preserve the last completed generation. Each section displays its own
generation and source timestamps. See history operations
for limits, diagnostics, and verification.
inspiration #
the term percolation in docs/ is a metaphor, not a claim. the system uses union-find for cluster detection (Newman-Ziff style), but it doesn't have a fixed lattice or controllable occupation parameter, so the phase-transition machinery from the literature doesn't apply directly — we just borrow the cluster-merging mental model.
Xie et al. 2021 — heterogeneous activity on social networks. we use the insight (some entities are far more "active" than others, so simple thresholds mislead) without implementing their analytic threshold.
Hailey's trending topics — the NER-first approach is borrowed directly. extracting structured entities collapses the post-text surface area into something graph-shaped.
ATProto labeler system — spam filtering via com.atproto.label. we subscribe to Hailey's labeler stream and drop posts from accounts labeled as spam before NER processing.
design decisions
these are documented as arbitrary choices to be revisited:
| decision | choice | why |
|---|---|---|
| edge definition | same-post co-occurrence | simplest, captures "discussed together" |
| edge weights | pheromone decay (configurable half-life) | ant colony inspired, recent co-occurrences matter more |
| activity threshold | 0.01 mentions/sec (~3 per 5 min) | rate normalizes across quiet/busy periods |
| trending metric | surprise vs baseline (UI), trend ratio (backend) | anomaly detection, not popularity contest |
| percolation threshold | largest_cluster / active > 50% | placeholder, needs empirical calibration |
| entity position | hash(text) → (x, y) | deterministic, stable, no semantic meaning yet |
| user weighting | planned (currently off) | power users count more (Xie 2021) |
see docs/02-semantic-percolation-plan.md for full rationale.
stack #
- backend (zig): jetstream consumer + spacez NER + labeler gate + entity graph + websocket server + SQLite persistence
- ner (python): LLM curator only — periodically fetches the entity graph and POSTs named groups back
- site: static html/css/js on cloudflare pages
run locally #
cd backend && zig build run # backend (jetstream + NER + entity graph + websocket)
cd ner && uv run python dedup.py # LLM curator (Claude Haiku names clusters)
cd site && npx wrangler pages dev . # frontend
deploy #
cd backend && fly deploy
cd ner && fly deploy
cd site && npx wrangler pages deploy . --project-name=coral --branch=main --commit-dirty=true
future work #
ideas being explored (not commitments):
-
semantic positioning - currently entities hash to arbitrary grid positions. could use embeddings to place semantically similar entities near each other, making the 2D layout a meaningful projection of topic space. unclear whether to embed entity names, representative posts, or cluster centroids.
-
temporal co-activity edges - entities that spike together might be related even without same-post co-occurrence. "earthquake" and "LA" could both trend during an event without always appearing together.
-
percolation calibration - the 50% threshold is arbitrary. need to correlate cluster merges with real-world events to understand what "discourse unification" actually looks like in the data.
references #
- Newman & Ziff, Efficient Monte Carlo algorithm and high-precision results for percolation, Phys. Rev. Lett. 85 (2000)
- Xie et al., Detecting and Modelling Real Percolation and Phase Transitions of Information on Social Media, Nature Human Behaviour (2021)
- Hailey, Bluesky Trending Topics - NER approach for topic detection
- Stauffer & Aharony, Introduction to Percolation Theory - theoretical foundations
- ATProto Labels - moderation architecture