From be0ca2271c726c35e57b2428ecfa68b1d1f88002 Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Tue, 4 Aug 2026 01:39:35 -0500 Subject: [PATCH] =?UTF-8?q?docs:=20handoff=20=E2=80=94=20fix=208,=20soak?= =?UTF-8?q?=20v1/v2=20evidence,=20machine=20lifecycle=20notes?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Fable 5 --- .../HANDOFF-2026-08-04-zio-backend.md | 55 ++++++++++++++----- 1 file changed, 42 insertions(+), 13 deletions(-) diff --git a/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md b/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md index a3ce403..aa7bc33 100644 --- a/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md +++ b/docs/handoffs/HANDOFF-2026-08-04-zio-backend.md @@ -77,6 +77,13 @@ f27c7ff backend: non-blocking ip6 connect on the fiber path <- fix 1½ The subscriber now resolves and tries addresses one at a time. This also removes fiber-native groups from the hot path entirely, sidestepping the open bug. +8. **poller left empty registrations as RDHUP-only** (zio, `fc6e0aa`): once + both direction tokens were consumed, the registration kept firing + token-less hangup events for every dead socket (level-triggered), and a + few hundred churned connections filled every epoll batch with noise — + observed as the ws accept loop going deaf mid-run while metrics stayed + healthy. Registrations are deleted when no direction is armed. This was + caught because a consumer was finally attached — always attach one. ## the open bug: parked-stack scribble under group churn @@ -111,20 +118,42 @@ correlate park-shot reports with the crash. ## the machine `zlay-zio-0801` (fly, personal, ord, 8cpu/16GB, NO VOLUME — stop/start loses -everything). Warm tree `/root/zlay-work` (keep it warm: untar over it, don't +everything). **The machine's init command exits after exactly 24 hours** +(observed twice: clean exit_code=0 at the 24h boundary) — a soak scheduled +near the end of a lifecycle dies with the box; check `fly machine status` +event logs before blaming the workload. `/root/provision.sh` (base64-shipped +from the session) reprovisions from a bare rootfs: toolchain, postgres with +scram auth on localhost, zlay build, cron event-log janitor (the 8GB disk +fills in under an hour at full ingest), core pattern. Warm tree `/root/zlay-work` (keep it warm: untar over it, don't recreate). Postgres seeded. 8GB disk: the event log fills it in under an hour at full ingest — wipe `/data/events/*` per run. `pkill -f` with any pattern from your own command line kills your own ssh session. -## soak results — TO FINALIZE - -Build `97440d2` (all seven fixes). Window: 90 minutes, 5-minute samples. -Early readings (t+90s): 1,882 connections (full fleet), `/metrics` in -0.26s during ramp (previous best 13s), relay_seq advancing, db queue -drained, RSS 1.16GB (prod: 2.6GB flat). - -- [ ] 90-minute window verdict: -- [ ] RSS at end of window vs start: -- [ ] metrics latency across all samples: -- [ ] recommended next: overnight soak, consumer attached, then operator - conversation about a canary +## soak evidence — TO FINALIZE with soak v3 + +**Soak v1** (build `97440d2`, fixes 1–7): 75+ minutes, ingest fully healthy +(1,500–1,900 conns, metrics 0.03–0.6s, RSS 1.16–1.35GB no trend, queues +clean), but its ws accept loop went deaf mid-run — that diagnosis became +fix 8, so v1 certifies nothing beyond ingest. + +**Soak v2** (build `c39adfc`, all 8 fixes, consumer attached throughout): +**~75 minutes, every gauge green at every sample** — 1,477–1,858 conns, +RSS 1.23–1.35GB oscillating with no growth, metrics 0.027–0.57s +(1.4s once, during ramp), seq 154k→1.21M continuous, broadcast queue never +full, db queue always 0, consumer delivered ~890k frames / 4.8GB at +~270 frames/s and successfully RECONNECTED at the hour mark (re-proving +the fix-8 path under long-run conditions). The relay also sailed through +two disk-100% episodes without dropping connections. The window ended at +~75 minutes because **the fly machine's own 24-hour keepalive expired** +(clean exit_code=0 of the VM — see machine notes), not due to any relay +behavior. + +**Soak v3** (same build, reprovisioned machine): the formal uninterrupted +90-minute certification. + +- [ ] 90-minute verdict: +- [ ] RSS start/end: +- [ ] metrics latency range: +- [ ] consumer frames delivered: +- [ ] recommended next: overnight soak, then operator conversation about a + canary -- 2.51.2