# zlay an [AT Protocol](https://atproto.com/) relay in zig. subscribes to every PDS on the network, verifies commit signatures, and serves the merged event stream to downstream consumers. **live instance**: [zlay.waow.tech](https://zlay.waow.tech/_health) — [metrics dashboard](https://zlay-metrics.waow.tech) ## design - **direct PDS crawl** — the bootstrap relay is called once at startup for the host list via `listHosts`, then all data flows directly from each PDS. - **optimistic signature validation** — on signing key cache miss, the frame passes through immediately and the DID is queued for background resolution. all subsequent commits are verified against the cached key. the cache caps at a configurable size and evicts the least recently used entry when full. - **inline collection index** — indexes `(DID, collection)` pairs in the event processing pipeline using RocksDB. serves `listReposByCollection` from the relay process — no sidecar. the index design draws on [fig](https://tangled.org/microcosm.blue)'s work on [lightrail](https://tangled.org/microcosm.blue/lightrail). - **reader per PDS + frame processing pool** — each PDS gets a lightweight reader (cursor tracking, rate limiting, header decode). heavy work (full CBOR decode, validation, DB persist, broadcast) runs on a shared pool of frame workers (configurable, default 16). - **pluggable std.Io backend** — the relay is written against zig's `std.Io` interface. the default backend is `Io.Threaded` (one OS thread per blocking call site); production runs [zio](https://tangled.org/zzstoatzz.io/zio) (`-Dbackend=zio`), which schedules the same code as cooperative fibers on a single epoll-driven scheduler thread — ~90 threads instead of ~2,800 at fleet scale, for a ~40% RSS reduction. ## spec compliance implements the relay endpoints from the [AT Protocol sync spec](https://atproto.com/specs/sync): `subscribeRepos`, `listRepos`, `getRepoStatus`, `listHosts`, `getHostStatus`, and `requestCrawl`. also implements `getLatestCommit`, `listReposByCollection`, and `getRepo` (302 redirect to PDS) for compatibility with the [indigo relay](https://github.com/bluesky-social/indigo) reference implementation. ## dependencies | dependency | purpose | |---|---| | [zat](https://tangled.org/zzstoatzz.io/zat) | AT Protocol primitives (CBOR, CAR, signatures, DID resolution) | | [websocket.zig](https://github.com/zzstoatzz/websocket.zig) | WebSocket client/server (fork with HTTP fallback + TCP split fixes) | | [pg.zig](https://github.com/karlseguin/pg.zig) | PostgreSQL driver | | [rocksdb-zig](https://github.com/Syndica/rocksdb-zig) | RocksDB bindings | ## build requires zig 0.17 and a C/C++ toolchain (for RocksDB). ```bash zig build # build (debug) zig build test # run tests zig build -Doptimize=ReleaseSafe # release build (production default) zig build -Dbackend=zio # zio std.Io backend (production runs ReleaseSafe + zio) just docker-build # local container, native Linux architecture just docker-build linux/amd64 # production-compatible container ``` ## configuration | variable | default | description | |---|---|---| | `RELAY_PORT` | `3000` | firehose + API port | | `RELAY_METRICS_PORT` | `3001` | prometheus metrics port | | `RELAY_UPSTREAM` | `bsky.network` | bootstrap relay for initial host list (set to `""` or `"none"` to disable) | | `RELAY_DATA_DIR` | `data/events` | event log storage | | `RELAY_RETENTION_HOURS` | `72` | replay window on disk. the default matches upstream indigo but is infeasible at full-network volume (~8.5G/hour); production runs `2` — deep replay belongs to stream.waow.tech's archive | | `RELAY_MAX_EVENTS_GB` | `100` | byte cap on the event dir (GiB), enforced every GC tick. an overflow backstop — size it above the retention window so the clock governs and the cap never silently shortens the advertised window (production: `30` over a ~17G 2h window) | | `COLLECTION_INDEX_DIR` | `data/collection-index` | RocksDB collection index path | | `DATABASE_URL` | — | PostgreSQL connection string | | `RELAY_ADMIN_PASSWORD` | — | bearer token for admin endpoints | | `RELAY_INGEST_STALL_ENABLED` | `1` | `/_healthz` returns 503 on a process-wide ingest stall; `0` disables. runtime-settable via `/admin/ingest-stall` | | `RELAY_INGEST_STALL_THRESHOLD_FPS` | `50` | ingest rate below which the process counts as stalled | | `RELAY_INGEST_STALL_WINDOW_SEC` | `900` | how long ingest must stay below the threshold before `/_healthz` flips | | `RELAY_INGEST_STALL_STARTUP_GRACE_SEC` | `900` | no stall verdict this long after start (cold-start spawn + backfill) | | `RELAY_WS_PING_ENABLED` | `1` | keepalive-ping upstream connections to detect half-open TCP; `0` disables. runtime-settable via `/admin/ws-ping` | | `RELAY_WS_PING_INTERVAL_SEC` | `30` | ping a connection after this much silence | | `RELAY_WS_PING_MAX_FAILURES` | `4` | reconnect after this many consecutive unanswered pings | | `RESOLVER_THREADS` | `4` | background DID resolution threads | | `FRAME_WORKERS` | `16` | frame processing pool worker count | | `FRAME_QUEUE_CAPACITY` | `4096` | max queued frames before backpressure | | `VALIDATOR_CACHE_SIZE` | `250000` | max cached signing keys before eviction | see [docs/deployment.md](docs/deployment.md) for production deployment and [docs/backfill.md](docs/backfill.md) for collection index backfill. ## license MIT