lessons from zlay (devlogs 006–010 + incident docs) #
What zlay is, and why Stream inherits from it. zlay is a separate project —
an atproto relay written in Zig (tangled.org/zat.dev/zlay), a sibling of
this one, sharing the zat atproto library and the same runtime choices
(Io.Threaded, glibc, ReleaseSafe). It has been in production longer and at
higher thread counts, so it hit the Zig-runtime and glibc failure modes first.
None of the rules below are about relay behaviour; they are about surviving this
language and runtime at scale, which is why they transfer even though Stream is
an archive rather than a relay.
Sources, all in sibling checkouts rather than this repository — ~/{forge}/{org}/{repo},
so alongside this one under zat.dev/:
zat/devlog/006–010zlay/docs/gotchas.md,allocation-audit.md,incident-*.md
read those for the full stories; these are the rules stream inherits. Each
parenthesised (006), (008), (incident-2026-05-31) is a pointer into them.
Spot-checked 2026-07-30 — still true in Stream: MALLOC_ARENA_MAX=4 is set
(deploy/README.md), and the build_info{git_sha,optimize} canary exists as
stream_build_info (runtime/metrics.zig).
- Io.Threaded + ReleaseSafe + c_allocator + 8 MB thread stacks — each bought with an incident. ReleaseFast hid heap corruption for days (008) and has a known double-free; GPA never returns memory to the OS; 2–4 MB stacks overflow in ReleaseSafe TLS/CBOR chains (134 KiB for tls.Client.init alone, 007); default 16 MB stacks map 44 GB VM at ~2.7k threads (006).
- never mix Io backends — Io.Mutex/Condition/io.sleep crash at runtime (NULL Thread.current) when called from a different backend than they were created under. cross-context comms = pure atomics (DbRequestQueue pattern) (008).
- durability before emission — never broadcast a seq a subscriber could later replay from before it is durable. atomic committed_seq watermark, NOT broadcast-inside-flush (that serialized 2,670 producers and OOM'd). recovery scans validate every record and truncate torn tails (incident-2026-05-31). applies the day the segment archive lands.
- socket close ownership — only the thread that owns a consumer's write loop closes the
socket; everyone else flips an
aliveatomic (EBADF race, 009). stream's Handler.close joins the subscriber thread before conn teardown — keep it that way. - serialize ws client writes; background loops must be cancellation-cooperative — never
catch {}an io.sleep (swallowed error.Canceled = use-after-free, 008); throttle mass reconnects in batches. - assume TCP splits everything — handshake/CRLF/HTTP parsing must tolerate byte-at-a-time delivery; one CRLF off-by-one was misblamed on the fiber scheduler for days (006, 008).
- thin reader threads — websocket read + queue only; decode/verify/fan-out on a shared pool.
0.45 vs 3.9 MiB per-reader RSS is run-vs-OOM (007). stream's single ingest thread currently
decodes+verifies inline: fine at one upstream, revisit if per-PDS crawling ever lands.
That condition has since been met, so the revisit is due. Backfill now runs
100 workers (
--backfill-workers=100) fetchinggetRepothrough the relay's 302 to the origin PDS, with per-host accounting iningest/backfill/host_store.zigand per-host parking inretry.zig. The live-ingest side is still one upstream, so the rule has not been broken — but "if per-PDS crawling ever lands" now reads as a condition already satisfied, not a hypothetical. - memory observability from day one — smaps-based RSS (mallinfo overflows at 2 GiB), malloc_info all-arena gauges, MALLOC_ARENA_MAX=4 in deploy env, build_info{git_sha,optimize} canary metric. balanced allocs + growing RSS = glibc arena fragmentation, not a leak (allocation-audit, incident-2026-05-30).
- build/deploy —
-Dtarget=x86_64-linux-gnu(musl SIGILLs C++ deps), native build on the server, immutable SHA tags, probes on the concurrent port, probe-path+image changes atomic (006, incident-2026-03-04). - the network is input — DID/handle resolution is SSRF surface: reject private hosts, DNS-preflight, no redirects, dial the checked address (010). XRPC errors are data: parse the error envelope, retry only transient with capped jittered backoff, honor retry-after.
- observe → measure → enforce for any new verification, behind a flag; beware checks that trivially pass on empty input (007's extractOps bug).
- incident discipline — one variable per experiment, verify for hours, keep a last-known-good SHA, roll back when hypothesis #3 fails (incident-2026-03-04, 009).