zlay #
an AT Protocol relay in zig. subscribes to every PDS on the network, verifies commit signatures, and serves the merged event stream to downstream consumers.
live instance: zlay.waow.tech — metrics dashboard
design #
-
direct PDS crawl — the bootstrap relay is called once at startup for the host list via
listHosts, then all data flows directly from each PDS. -
optimistic signature validation — on signing key cache miss, the frame passes through immediately and the DID is queued for background resolution. all subsequent commits are verified against the cached key. the cache caps at a configurable size and evicts the least recently used entry when full.
-
inline collection index — indexes
(DID, collection)pairs in the event processing pipeline using RocksDB. serveslistReposByCollectionfrom the relay process — no sidecar. the index design draws on fig's work on lightrail. -
reader per PDS + frame processing pool — each PDS gets a lightweight reader (cursor tracking, rate limiting, header decode). heavy work (full CBOR decode, validation, DB persist, broadcast) runs on a shared pool of frame workers (configurable, default 16).
-
pluggable std.Io backend — the relay is written against zig's
std.Iointerface. the default backend isIo.Threaded(one OS thread per blocking call site); production runs zio (-Dbackend=zio), which schedules the same code as cooperative fibers on a single epoll-driven scheduler thread — ~90 threads instead of ~2,800 at fleet scale, for a ~40% RSS reduction.
spec compliance #
implements the relay endpoints from the AT Protocol sync spec: subscribeRepos, listRepos, getRepoStatus, listHosts, getHostStatus, and requestCrawl.
also implements getLatestCommit, listReposByCollection, and getRepo (302 redirect to PDS) for compatibility with the indigo relay reference implementation.
dependencies #
| dependency | purpose |
|---|---|
| zat | AT Protocol primitives (CBOR, CAR, signatures, DID resolution) |
| websocket.zig | WebSocket client/server (fork with HTTP fallback + TCP split fixes) |
| pg.zig | PostgreSQL driver |
| rocksdb-zig | RocksDB bindings |
build #
requires zig 0.16 and a C/C++ toolchain (for RocksDB).
zig build # build (debug)
zig build test # run tests
zig build -Doptimize=ReleaseSafe # release build (production default)
zig build -Dbackend=zio # zio std.Io backend (production runs ReleaseSafe + zio)
just docker-build # local container, native Linux architecture
just docker-build linux/amd64 # production-compatible container
configuration #
| variable | default | description |
|---|---|---|
RELAY_PORT |
3000 |
firehose + API port |
RELAY_METRICS_PORT |
3001 |
prometheus metrics port |
RELAY_UPSTREAM |
bsky.network |
bootstrap relay for initial host list (set to "" or "none" to disable) |
RELAY_DATA_DIR |
data/events |
event log storage |
RELAY_RETENTION_HOURS |
72 |
replay window on disk. the default matches upstream indigo but is infeasible at full-network volume (~8.5G/hour); production runs 2 — deep replay belongs to stream.waow.tech's archive |
RELAY_MAX_EVENTS_GB |
100 |
byte cap on the event dir (GiB), enforced every GC tick. an overflow backstop — size it above the retention window so the clock governs and the cap never silently shortens the advertised window (production: 30 over a ~17G 2h window) |
COLLECTION_INDEX_DIR |
data/collection-index |
RocksDB collection index path |
DATABASE_URL |
— | PostgreSQL connection string |
RELAY_ADMIN_PASSWORD |
— | bearer token for admin endpoints |
RELAY_INGEST_STALL_ENABLED |
1 |
/_healthz returns 503 on a process-wide ingest stall; 0 disables. runtime-settable via /admin/ingest-stall |
RELAY_INGEST_STALL_THRESHOLD_FPS |
50 |
ingest rate below which the process counts as stalled |
RELAY_INGEST_STALL_WINDOW_SEC |
900 |
how long ingest must stay below the threshold before /_healthz flips |
RELAY_INGEST_STALL_STARTUP_GRACE_SEC |
900 |
no stall verdict this long after start (cold-start spawn + backfill) |
RELAY_WS_PING_ENABLED |
1 |
keepalive-ping upstream connections to detect half-open TCP; 0 disables. runtime-settable via /admin/ws-ping |
RELAY_WS_PING_INTERVAL_SEC |
30 |
ping a connection after this much silence |
RELAY_WS_PING_MAX_FAILURES |
4 |
reconnect after this many consecutive unanswered pings |
RESOLVER_THREADS |
4 |
background DID resolution threads |
FRAME_WORKERS |
16 |
frame processing pool worker count |
FRAME_QUEUE_CAPACITY |
4096 |
max queued frames before backpressure |
VALIDATOR_CACHE_SIZE |
250000 |
max cached signing keys before eviction |
see docs/deployment.md for production deployment and docs/backfill.md for collection index backfill.
license #
MIT