# lessons from zlay (devlogs 006–010 + incident docs) **What zlay is, and why Stream inherits from it.** zlay is a separate project — an atproto **relay** written in Zig (`tangled.org/zat.dev/zlay`), a sibling of this one, sharing the `zat` atproto library and the same runtime choices (`Io.Threaded`, glibc, ReleaseSafe). It has been in production longer and at higher thread counts, so it hit the Zig-runtime and glibc failure modes first. None of the rules below are about relay behaviour; they are about surviving this language and runtime at scale, which is why they transfer even though Stream is an archive rather than a relay. Sources, all in sibling checkouts rather than this repository — `~/{forge}/{org}/{repo}`, so alongside this one under `zat.dev/`: - `zat/devlog/006`–`010` - `zlay/docs/gotchas.md`, `allocation-audit.md`, `incident-*.md` read those for the full stories; these are the rules stream inherits. Each parenthesised `(006)`, `(008)`, `(incident-2026-05-31)` is a pointer into them. Spot-checked 2026-07-30 — still true in Stream: `MALLOC_ARENA_MAX=4` is set (`deploy/README.md`), and the `build_info{git_sha,optimize}` canary exists as `stream_build_info` (`runtime/metrics.zig`). 1. **Io.Threaded + ReleaseSafe + c_allocator + 8 MB thread stacks** — each bought with an incident. ReleaseFast hid heap corruption for days (008) and has a known double-free; GPA never returns memory to the OS; 2–4 MB stacks overflow in ReleaseSafe TLS/CBOR chains (134 KiB for tls.Client.init alone, 007); default 16 MB stacks map 44 GB VM at ~2.7k threads (006). 2. **never mix Io backends** — Io.Mutex/Condition/io.sleep crash at runtime (NULL Thread.current) when called from a different backend than they were created under. cross-context comms = pure atomics (DbRequestQueue pattern) (008). 3. **durability before emission** — never broadcast a seq a subscriber could later replay from before it is durable. atomic committed_seq watermark, NOT broadcast-inside-flush (that serialized 2,670 producers and OOM'd). recovery scans validate every record and truncate torn tails (incident-2026-05-31). applies the day the segment archive lands. 4. **socket close ownership** — only the thread that owns a consumer's write loop closes the socket; everyone else flips an `alive` atomic (EBADF race, 009). stream's Handler.close joins the subscriber thread before conn teardown — keep it that way. 5. **serialize ws client writes; background loops must be cancellation-cooperative** — never `catch {}` an io.sleep (swallowed error.Canceled = use-after-free, 008); throttle mass reconnects in batches. 6. **assume TCP splits everything** — handshake/CRLF/HTTP parsing must tolerate byte-at-a-time delivery; one CRLF off-by-one was misblamed on the fiber scheduler for days (006, 008). 7. **thin reader threads** — websocket read + queue only; decode/verify/fan-out on a shared pool. 0.45 vs 3.9 MiB per-reader RSS is run-vs-OOM (007). stream's single ingest thread currently decodes+verifies inline: fine at one upstream, revisit if per-PDS crawling ever lands. **That condition has since been met, so the revisit is due.** Backfill now runs 100 workers (`--backfill-workers=100`) fetching `getRepo` through the relay's 302 to the origin PDS, with per-host accounting in `ingest/backfill/host_store.zig` and per-host parking in `retry.zig`. The live-ingest side is still one upstream, so the rule has not been *broken* — but "if per-PDS crawling ever lands" now reads as a condition already satisfied, not a hypothetical. 8. **memory observability from day one** — smaps-based RSS (mallinfo overflows at 2 GiB), malloc_info all-arena gauges, MALLOC_ARENA_MAX=4 in deploy env, build_info{git_sha,optimize} canary metric. balanced allocs + growing RSS = glibc arena fragmentation, not a leak (allocation-audit, incident-2026-05-30). 9. **build/deploy** — `-Dtarget=x86_64-linux-gnu` (musl SIGILLs C++ deps), native build on the server, immutable SHA tags, probes on the concurrent port, probe-path+image changes atomic (006, incident-2026-03-04). 10. **the network is input** — DID/handle resolution is SSRF surface: reject private hosts, DNS-preflight, no redirects, dial the checked address (010). XRPC errors are data: parse the error envelope, retry only transient with capped jittered backoff, honor retry-after. 11. **observe → measure → enforce** for any new verification, behind a flag; beware checks that trivially pass on empty input (007's extractOps bug). 12. **incident discipline** — one variable per experiment, verify for hours, keep a last-known-good SHA, roll back when hypothesis #3 fails (incident-2026-03-04, 009).