diff --git a/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md b/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md index bf50f0e..672e717 100644 --- a/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md +++ b/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md @@ -96,22 +96,68 @@ an rsp-bounds assert never fired, and a 512B park snapshot saw nothing. The 4KB-window detector run was cut short by fix 7 making the question non-blocking. The no-groups build ran 72 minutes clean. -**Diagnosis progress (2026-08-04, post-certification):** the corruption is -now confirmed as **stack aliasing** — a parked fiber's stack rewritten with -ordinary execution frames (spilled compiler constants, code addresses, -frame chains), i.e. something *ran* on a stack the sched believed parked. -Bounds and list-membership invariants cannot see this (both aliases are -in-bounds on the same mapping). A creation-time tripwire (spawn asserts no -live fiber owns the assigned stack) plus a sweep-time pairwise check are -deployed in the hunt build; either names both fibers on the next -occurrence. Transport tripwires localized the garbage error values to -`HostName.connect`'s internals before the aliasing diagnosis subsumed them. - -Status: fiber-native groups pass the suite and the 4.5M-iteration storm, -but are NOT trusted under field churn until the aliasing path is named. -zlay no longer exercises them (fix 7). The hunt farm on the machine -(`/root/huntloop.sh`, artifacts in `/root/hunt/`) runs unattended and is -fully instrumented. +**Post-certification investigation (2026-08-04 evening).** Two real zio +bugs were root-caused after the canary, plus one correction to earlier +claims in this document. + +**Root cause of the group-era corruption: `io.async` was never overridden.** +zio's vtable replaced concurrent/await/cancel but not `async`, so every +`io.async` call fell through to the inner `Io.Threaded` and returned a +*Threaded* future object — which `Future.await`/`cancel` then handed back to +zio's `vtAwait`/`vtCancel`, which cast it to `*Task` and read/wrote +`refs`/`state`/`result` at offsets that mean something else entirely inside +a Threaded future. `HostName.connect` is built on `io.async` +(`connectMany`, and `lookup` inside it), so every reconnect minted a +type-confused future. This explains the garbage error unions at +`@errorName` and the `0xAA…` (zig undefined-fill) values seen in task +fields. Fixed in zio `41d3be1`; `async` now spawns fibers exactly like +`concurrent`, with the eager-inline fallback its contract permits. + +**A second, latent bug that fix exposed: lost fiber wake from a spilled +DNS lookup.** `Threaded.netLookup` writes results into the *caller's* queue +and closes it through its **own inner io**, whose futexWake is kernel-only +and readies no fibers. `HostName.connectMany` consumes that queue +concurrently with the lookup, so its fiber parked in `queue.getOne` forever. +Symptom: a relay that spawns 1,797 subscriber fibers and establishes zero +connections, while `/metrics` and the accept path stay perfectly healthy — +`relay_seq` pinned at 0. Latent before the async fix because `connectMany` +then ran on a Threaded thread, where a kernel-only wake was the right +domain. Fixed in zio `9c8d3cf` (broadcast fiber wake on spill completion). + +**CORRECTION to an earlier revision of this document:** it claimed the +corruption was "confirmed as stack aliasing" based on a parked-stack +snapshot detector. That detector had an address bug — it snapshotted at the +stack pointer captured *before* the context switch (the fiber's previous +park) and verified against the current one, so ordinary changes in park +depth read as corruption. **Every park-shot diff and GEOMETRY panic from it +is unfounded**, and no evidence of stack aliasing survives. The two bugs +above were found by other means (error-value tripwires, `relay_seq` +inspection, code reading) and are unaffected. Detector fixed in zio +`1321be0`; it is comptime-gated off in ship builds either way. + +**Still open: the prod deafness itself.** Neither bug above explains it — +the ship build uses sequential connect (no `connectMany`, no `io.async`) and +had no async override. A consumer-churn harness now exists to attack it +(see below); it has not yet reproduced prod's failure. + +## observability added for the next canary (this is the important part) + +`zio.ZioBackend.debugStats()` (zio `d0526f8`) exposes scheduler internals +cheaply and on demand — switches counter, live/ready/parked fiber counts, +futex waiter count, timer heap depth, inject queue depth, free list, +poller registration count, and DNS spill pool depth/workers/idle. Walk is +O(fibers) and happens only when called. + +**Wire these into `/metrics` before the next canary** so the operator's +prometheus captures them through any onset (last time the scrape timeline +was the only forensics that survived the pod). The discriminators they give +you, which we lacked entirely during the incident: + +- `switches` flat across scrapes → the scheduler loop is wedged +- `parked` climbing, `ready` ~0 → fibers piling up on a lost wake +- `poller_registrations` collapsing → registration erosion +- `dns_queued` climbing, `dns_idle` 0 → spill pool saturated/stuck +- all healthy while serving is deaf → the failure is above the runtime ## instrumentation that now exists (all in zio, cheap enough to keep)