From bf7bda5e5579983deb2b755694c6529a82026cc4 Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Sat, 8 Aug 2026 12:04:40 -0500 Subject: [PATCH] docs: routine deploy runbook + upstream head-to-head bench comparison deploying.md: the everyday ship-a-commit flow (admit run ~40min, publish, .env cutover, ~5min restart post-fba93d0), verification checklist, failure table. distinct from deployment-runbook.md (the whole-network experiment). benchmarks.md: same-machine ratios vs upstream f29815c with parity caveats. Co-Authored-By: Claude Fable 5 --- docs/benchmarks.md | 31 ++++++++++++++++++ docs/deploying.md | 78 ++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 109 insertions(+) create mode 100644 docs/deploying.md diff --git a/docs/benchmarks.md b/docs/benchmarks.md index ba0f224..3e733a9 100644 --- a/docs/benchmarks.md +++ b/docs/benchmarks.md @@ -94,3 +94,34 @@ The harness is deliberately simple — a fixed iteration count, no warmup discipline beyond a short prime, no statistical treatment. It catches order-of-magnitude regressions, not differences of a few percent. For precision, use a dedicated measurement. + +## Head-to-head vs upstream (2026-08-08, Apple M5 Pro, one run each) + +Same machine, same afternoon: `zig build bench -Doptimize=ReleaseSafe` here, +`go test -bench` at upstream `f29815c`. Single runs — treat as ratios with +~30-40% run variance, not as absolute truth. + +| path | stream | upstream (Go) | ratio | parity | +|---|---|---|---|---| +| writer, live shape | 3.20M rows/s | 2.13M events/s (sync) | ~1.5x | clean | +| writer, backfill shape | 5.30M rows/s | 4.43M events/s (async8, its best) | ~1.2x | clean | +| sealed block decode | 15.0M rows/s | 6.25M rows/s (654µs / 4096-event block) | ~2.4x | clean | +| block bloom lookup | 35 ns | 630 ns | ~18x | clean; includes our fixed-4096 bloom, see memory-ceiling note | +| cold fan-out, 1 consumer | 849K rows/s | 281K frames/s | ~3x | shapes differ slightly (their bench spans consumer counts) | +| seal | 3.04M rows/s (seal step only) | 1.22M events/s (append+flush+seal per op) | directionally ahead | NOT clean — different op boundaries | +| reader open | 123 ns | 21.2 µs | — | NOT comparable: ours parses resident bytes, theirs opens and reads a file | + +Context that keeps these honest: + +- Upstream jetstream V2 vendors its own decode — no indigo dependency — so + the atproto-bench indigo numbers say nothing about it. The zat-vs-their-Go + comparison on *format* work is the block-decode row (2.4x) via the same + sealed-segment shape (`atproto-bench/go-jss` uses their `segment` package). +- The live network runs ~200-450 events/s. Every row above is in the + millions/s: CPU is not the live tail's bottleneck for either + implementation. Where the margin actually pays: bootstrap decode+verify + (SDK-level: zat 249K frames/s CID-verified vs 91K for the fastest Go SDK + measured in atproto-bench), replay serving, compaction/rebuild scans, and + box size for the same workload. +- Nothing here is a promotion gate; parity rows marked NOT clean need a + matched-shape bench before being quoted as a ratio. diff --git a/docs/deploying.md b/docs/deploying.md new file mode 100644 index 0000000..a07749b --- /dev/null +++ b/docs/deploying.md @@ -0,0 +1,78 @@ +# deploying — routine code changes to the live box + +How to ship a commit to the running instance at stream.waow.tech, what each +stage costs, and what to expect while it happens. This is the everyday flow; +the whole-network *experiment* procedure (provisioning, watchdogs, cost +envelopes) is `deployment-runbook.md`. + +The instance is one Hetzner box (stream-cx43, 89.167.122.160, hel1) running +docker compose at `/opt/stream-experiment/`. The stream image is pinned in +`/opt/stream-experiment/.env` as `STREAM_IMAGE=atcr.io/zat.dev/stream:`. + +## the pipeline, in order + +| stage | command | measured cost | +|---|---|---| +| 1. land the change | commit + push to main (CI: fmt, debug+releasesafe tests) | suite floor ~32 min in CI, but admission below is the deploy gate | +| 2. admission | `STREAM_BUILD_HOST=ssh://root@89.167.122.160 ./scripts/admit run` | ~40 min total: ~5–8 min native amd64 build on the box + ~32 min gate suite | +| 3. publish | `./scripts/admit publish [sha]` | ~1–2 min (push + pin registry digest into `receipts/.json`) | +| 4. cut over | on the box: edit `STREAM_IMAGE` in `/opt/stream-experiment/.env`, then `docker compose up -d stream` | seconds to swap | +| 5. restart recovery | (automatic) | ~4–5 min to serving + dialing upstream; see below | + +Total wall clock from "fix committed" to "new build serving": **~45 minutes.** + +Notes on the stages: + +- **Admission is the gate, not CI.** A receipt (`receipts/.json`) binds + one full-suite run to one image digest; `scripts/admit verify ` + refuses anything without a fully-passing receipt. Never point `.env` at an + image that `admit verify` rejects. +- **Always build with `STREAM_BUILD_HOST`** (native amd64 on the box). + Local emulated builds on an arm64 workstation take 40+ minutes and produce + a cross-emulated approximation. The build ships the context over ssh, builds + remotely, and streams the image back so registry credentials never leave + the workstation. +- **Uncommitted trees are refused** (`require_clean_tree`) — the image tag is + the commit, by construction. + +## what to expect at restart + +- The process replays its cursor from the durable watermark; with + delete-compaction enabled, startup first rebuilds the live tombstone set. + Since `fba93d0` (header-level watermark skip, upstream spec §3.4) that + rebuild reads only segment headers below the watermark and costs seconds to + minutes, not hours. Before that fix a restart paid a ~3h full-archive scan + with the tail deaf throughout — if a restart ever takes that long again, + the skip has regressed; check the rebuild progress lines in the logs. +- **Subscribers are dropped at the restart.** zat-based consumers rotate to + their next configured host (official jetstreams) and come back on their own + schedule — they only return to stream when their current host drops them. + Do not conclude the deploy broke a consumer because it is camping on a + fallback host. +- The live tail replays the full gap on reconnect; the archive serves + throughout (the ~30s of docker swap excepted). + +## post-deploy verification (do all four) + +1. `curl -s https://stream.waow.tech/metrics | grep stream_build_info` — + the sha must be the one you shipped. +2. `upstream seq` advancing across two reads ~10s apart (an honest live-tail + check; `state live` alone only means the process is serving). +3. Delivery cadence: `tools/loadgen` one connection with `--report-s=1` for + ~60s — per-second deltas, no silent seconds. Run it once live-attach and + once with `--rewind-s=10` (cursor resume — the path that broke 2026-08-08). +4. Grafana (`grafana.stream.waow.tech`): rate, RSS, and dropped-events panels + at their usual shapes. + +Then commit the ops note: what shipped, why, what was verified. The operator +journal is `docs/ops-changelog.md` in the relay repo. + +## when a stage goes wrong + +| symptom | do | +|---|---| +| admission suite fails | fix it; there is no skip lever — a `skipped` verdict blocks admission like a failure | +| `admit verify` refuses at cutover | you are deploying an image with no passing receipt; go back to stage 2 | +| restart takes hours, tail deaf | the tombstone-rebuild header skip regressed (see above); the archive still serves — decide rollback vs. wait using the rebuild progress logs | +| consumers stay on fallback hosts | expected; bounce a consumer if you need it back on stream to verify a fix | +| box unreachable mid-deploy | the old container keeps running unless compose already swapped; verify with `docker ps` before assuming an outage | -- 2.51.2