atproto relay in zig zlay.waow.tech
relay zig atproto
Zig 99%
Python <1%
Shell <1%
Dockerfile <1%
Just <1%
<1%

README.md

zlay #

an AT Protocol relay in zig. subscribes to every PDS on the network, verifies commit signatures, and serves the merged event stream to downstream consumers.

live instance: zlay.waow.tech — metrics dashboard

design #

  • direct PDS crawl — the bootstrap relay is called once at startup for the host list via listHosts, then all data flows directly from each PDS.

  • optimistic signature validation — on signing key cache miss, the frame passes through immediately and the DID is queued for background resolution. all subsequent commits are verified against the cached key. the cache caps at a configurable size and evicts the least recently used entry when full.

  • inline collection index — indexes (DID, collection) pairs in the event processing pipeline using RocksDB. serves listReposByCollection from the relay process — no sidecar. the index design draws on fig's work on lightrail.

  • reader per PDS + frame processing pool — each PDS gets a lightweight reader (cursor tracking, rate limiting, header decode). heavy work (full CBOR decode, validation, DB persist, broadcast) runs on a shared pool of frame workers (configurable, default 16).

  • pluggable std.Io backend — the relay is written against zig's std.Io interface. the default backend is Io.Threaded (one OS thread per blocking call site); production runs zio (-Dbackend=zio), which schedules the same code as cooperative fibers on a single epoll-driven scheduler thread — ~90 threads instead of ~2,800 at fleet scale, for a ~40% RSS reduction.

spec compliance #

implements the relay endpoints from the AT Protocol sync spec: subscribeRepos, listRepos, getRepoStatus, listHosts, getHostStatus, and requestCrawl.

also implements getLatestCommit, listReposByCollection, and getRepo (302 redirect to PDS) for compatibility with the indigo relay reference implementation.

dependencies #

dependency purpose
zat AT Protocol primitives (CBOR, CAR, signatures, DID resolution)
websocket.zig WebSocket client/server (fork with HTTP fallback + TCP split fixes)
pg.zig PostgreSQL driver
rocksdb-zig RocksDB bindings

build #

requires zig 0.16 and a C/C++ toolchain (for RocksDB).

zig build                          # build (debug)
zig build test                     # run tests
zig build -Doptimize=ReleaseSafe   # release build (production default)
zig build -Dbackend=zio            # zio std.Io backend (production runs ReleaseSafe + zio)
just docker-build                 # local container, native Linux architecture
just docker-build linux/amd64     # production-compatible container

configuration #

variable default description
RELAY_PORT 3000 firehose + API port
RELAY_METRICS_PORT 3001 prometheus metrics port
RELAY_UPSTREAM bsky.network bootstrap relay for initial host list (set to "" or "none" to disable)
RELAY_DATA_DIR data/events event log storage
RELAY_RETENTION_HOURS 72 replay window on disk. the default matches upstream indigo but is infeasible at full-network volume (~8.5G/hour); production runs 2 — deep replay belongs to stream.waow.tech's archive
RELAY_MAX_EVENTS_GB 100 byte cap on the event dir (GiB), enforced every GC tick. an overflow backstop — size it above the retention window so the clock governs and the cap never silently shortens the advertised window (production: 30 over a ~17G 2h window)
COLLECTION_INDEX_DIR data/collection-index RocksDB collection index path
DATABASE_URL — PostgreSQL connection string
RELAY_ADMIN_PASSWORD — bearer token for admin endpoints
RELAY_INGEST_STALL_ENABLED 1 /_healthz returns 503 on a process-wide ingest stall; 0 disables. runtime-settable via /admin/ingest-stall
RELAY_INGEST_STALL_THRESHOLD_FPS 50 ingest rate below which the process counts as stalled
RELAY_INGEST_STALL_WINDOW_SEC 900 how long ingest must stay below the threshold before /_healthz flips
RELAY_INGEST_STALL_STARTUP_GRACE_SEC 900 no stall verdict this long after start (cold-start spawn + backfill)
RELAY_WS_PING_ENABLED 1 keepalive-ping upstream connections to detect half-open TCP; 0 disables. runtime-settable via /admin/ws-ping
RELAY_WS_PING_INTERVAL_SEC 30 ping a connection after this much silence
RELAY_WS_PING_MAX_FAILURES 4 reconnect after this many consecutive unanswered pings
RESOLVER_THREADS 4 background DID resolution threads
FRAME_WORKERS 16 frame processing pool worker count
FRAME_QUEUE_CAPACITY 4096 max queued frames before backpressure
VALIDATOR_CACHE_SIZE 250000 max cached signing keys before eviction

see docs/deployment.md for production deployment and docs/backfill.md for collection index backfill.

license #

MIT