jetstream v2 in zig stream.waow.tech
stream docs deployment-runbook.md
16 kB
Markdown

Deployment runbook — next whole-network experiment #

Written against the record of the three previous attempts and their failure modes. docs/full-network-experiment-2026-07.md is the durable field log; docs/semantic-parity.md is the gate; this file is what to actually do, in order, and what to do when a stage goes wrong.

Three terms used throughout, since they are load-bearing here and opaque cold:

  • the gate — semantic-parity.md, the row-by-row record of which behaviours are proven equivalent to upstream Jetstream V2. Nothing deploys on the strength of a passing test alone; the gate is what says a surface is actually settled.
  • a receipt — a JSON file under receipts/ binding one run of the full suite to one image digest (receipts/<sha>.json). It records each suite's verdict, and a skipped verdict blocks admission exactly like a failure. An admitted artifact is an image whose digest has such a receipt with everything passing.
  • cutover — the moment the lifecycle leaves merge for steady state and serving stops returning 503. Before it, the archive is still being assembled; after it, clients are being answered.

Every stage must be measurable. All three previous runs ended because of a condition the running system could not report.

What happened the last three times #

1 — 14 July 2026, terminated the same day #

CPX62, 4 TB volume, ~$18. Killed by a hard-coded 512 MiB getRepo CAR limit with no upstream equivalent: real repositories exceed it, so the crawl excluded them and the archive was incomplete. The archive was discarded.

Commissioning also found a per-worker memory ceiling too small for a real CAR (OOM treated as lifecycle-fatal), and pooled redirected connections producing dozens of false HttpConnectionClosing failures per minute that direct curl disproved.

Carried forward: no arbitrary size exclusion; incremental CAR consumption; exhaustion is a repository failure, not a process failure.

2 — 22 July 2026, destroyed early #

Isolated CCX43, independent watchdog, off-host observer, 2.5 TB XFS, ~$2.53. 26,099 repositories crossed durable completion; ~73.9 GB disk, 12.3 GB RSS.

Killed by a silent live-path death: the live scheduler stopped at 399,362 archived events while its WebSocket reader kept running and advanced its reconnect cursor 459,475 sequences past the last accepted frame. The process stayed up, bootstrap kept crawling, and ordinary health and backfill counters concealed it. The original error was not even retained — there was no terminal pipeline metric and no fail-loud capture supervision.

Carried forward: frames are acknowledged only after Stream accepts them; pipeline tickets, stage gaps and terminal failures must be scrapeable throughout bootstrap, not only in steady state.

3 — 23–24 July 2026, image e1926f3 #

Same isolated shape. The final receipt:

jetstream_backfill_discovered_total   100000
jetstream_backfill_completed_total         0
jetstream_backfill_progress_completed      0
process_resident_memory_bytes     24,263,409,664   (24.3 GB)
stream_events_total                   618,558
/data/stream                             26 GB

100,000 repositories were discovered and zero were durably complete. Completion was counted at the 100,000-entry dispatch boundary, so every restart — and the log shows four process starts — replayed nearly the whole batch while the dashboard's "durably committed" series climbed. That series was reading a process-local counter.

The run also crashed three separate ways (borrowed CAR/MST data outliving its owner, an UnsupportedFloat row escaping as process-fatal, a shutdown exceeding 45 s and being killed), took two unplanned restarts from ordinary firehose disconnects, and got slower when scaled from 100 to 200 workers (≈11 → ≈9 repos/s) with the in-flight limit saturated.

The deployed image was never admitted: e1926f3 included source and dependency changes made after the receipt that covered it, and the pipeline did not detect this.

What is different now #

Each failure mode above has a specific mitigation:

Then Now
512 MiB CAR exclusion removed; incremental consumption
Live path died silently terminal pipeline signals; frames acked only on accept
Completion at dispatch boundary per-repository, coupled to archive-writer durability
Progress panel read a process-local counter stream_backfill_repos_durable, read from the store; survives restart (measured 22 → 49)
Firehose disconnect killed the process websocket.zig setsockopt unreachable fixed; shipped via zat v0.3.18
Unencodable record killed ingest isolated per row, per wire
Retry failures hidden behind green health terminal failure stops the service
Unadmitted image deployed admit verify gates deploy.sh; it refuses sha256:adc276…
OOM retired a good repo as a malformed CAR exhaustion stays OutOfMemory, pinned by an allocation-failure sweep
Stale operator CIDR could lock you out provision.sh checks it against the live egress IP and refuses
Batch dispatch deadlocked before committing anything submit and consume interleaved; verified against bsky.network
Global parse mutex serialized 100 workers behind one core removed; upstream parses concurrently
In-flight byte budget blocked workers and could deadlock removed; upstream has no such bound
Scoped to a target nobody had measured Stage 0b feasibility check, with the head-of-list trap named
Experiment 3's tuning deployed unnoticed assert_deploy_args.py refuses backfill flags no receipt covers

Admitted artifact actually deployed for experiment 6: sha256:0f45468c2fdba928fb3a9830f79d4693960a9c0ae2174ae042ecb01dfd33e0cd (receipts/29705cb.json, 20/20 suites all pass, receipt cut 2026-07-28 04:36 UTC against an 05:01 UTC start, tag atcr.io/zat.dev/stream:29705cb).

This superseded eb95c14 (sha256:4300104135c6fd35bd13a264b744c5f842c41ee9a51e0cf1c57c9a59640551a2), which earlier revisions of this file named as the artifact for the run. Both were properly admitted; the runbook was not updated when the image moved. Verify the running digest against the receipt: docker inspect the workload container and compare to receipts/<sha>.json's image.registry_digest.

The eb95c14 fix below is carried forward into 29705cb: three sites reported OutOfMemory as a structural verdict about a repository's CAR, and the engine retires a repository permanently on that verdict. Under whole-network memory pressure (experiment 3 reached 24.3 GB) Stream would have discarded good repositories and recorded their PDS as the cause. See the whole-network bootstrap row in semantic-parity.md.

The plan #

Cost ceiling and deadline are set once, in provision.sh, and never silently extended. Previous envelope: 24 h, $50, ~$2.53–$18 actual.

Stage 0 — before spending anything (~30 min) #

  • python3 assert_empty_project.py — the project must be empty.
  • Recompute the operator /32; never reuse the one in .experiment.env. It is dynamic and has changed mid-run, which silently firewalls you out of your own workload while it keeps running fine. provision.sh checks it; nothing rechecks it during the run, so if probes go quiet, check this before concluding the workload is sick.
  • ./scripts/admit verify <digest> must admit the artifact you intend to run.
  • Confirm the artifact matches the commit you think it does.
  • python3 assert_deploy_args.py — the flags must match what was admitted. Admission covers the artifact, not its invocation.

Stage 0b — feasibility (~20 min, €0.03) #

The gate's twenty correctness suites do not measure duration. An admitted artifact can still be far from finishing inside the deadline; correctness evidence is not feasibility evidence.

Measured 2026-07-28:

Quantity Value How
repositories ~19.4M cursor-space binary search (22.7M) × measured density (0.854)
mean repo 0.45 MB 400 independent uniformly-random cursors
median repo 0.01 MB same
p99 / max 10.5 MB / 69 MB same
full pass ~8.7 TB product
our sustained rate ~460 Mbps CCX43, 100 workers
projection ~42 h at 460 Mbps, ~19 h at 1 Gbps arithmetic

Upstream documents ~16 hours, which needs ~1.2 Gbps — a single server with a few TB of disk, matching their design doc's stated profile.

Repository size is heavy-tailed and listRepos is ordered largest-first. The median repo is 10 KB; the first 100,000 entries average ~10 MB. Sampling consecutive entries, or extrapolating early throughput, overstates the network by ~21× and turns a 42-hour job into an apparent 40-day one. Sample independent random cursors, never consecutive ones — consecutive entries share a PDS and creation cohort, so forty of them is one sample, not forty.

Corollary for Stage 3: early throughput will look low and that is expected. The first hour crawls the heaviest accounts in the network. Judge progress by bytes moved, not repos completed, until the crawl is past the head.

If the projection exceeds the deadline, do not provision. Scale the goal to a bounded slice, or change the deadline.

Stage 1 — provision and arm (~20 min) #

CCX43 workload, 2.5 TB XFS, CPX12 watchdog, CPX12 observer, HEL1.

The watchdog is armed before the workload is trusted. It destroys billable resources on a hard deadline; "powered off" is not teardown.

If it fails: ./destroy.sh, verify the project is empty, then diagnose.

Stage 2 — deploy and first light (~15 min) #

deploy.sh refuses any digest admit verify does not admit — this is the e1926f3 guard. Then:

  • /healthz, /readyz respond; serving is gated (503) until steady state.
  • /metrics scrapes and pipeline tickets are already moving — experiment 2 died because these were only checked after steady state.
  • The observer, off-host, sees the public surface.

If it fails: capture a receipt first, then roll back. Experiment 2's root cause was lost because no receipt was taken.

Stage 3 — first hour, the one that matters #

Watch these four, in this order of importance:

  1. stream_backfill_repos_durable{status="complete"} climbing, and jetstream_backfill_completion_queued_total climbing with it. Neither experiment 3 nor 4 ever produced these. If discovery climbs and they stay at zero for more than ~15 minutes of active crawling, stop — continuing only builds an archive a restart discards. The two zeros distinguish the causes: handled climbing with queued at zero is the batch deadlock (check stream_backfill_queue_depth pinned at workers × 2); queued climbing with durable flat is a durability-boundary problem instead.

    Judge the rate by bytes, not repos. Sustained ~460 Mbps means the crawl is healthy even at 5 repos/s, because the head of listRepos is ~10 MB per repository against a 0.45 MB network mean. Repos/s should climb by more than an order of magnitude once past the head. A falling byte rate is the real alarm.

  2. Pipeline tickets submitted/claimed/emitted advancing together. A growing submitted-to-emitted gap is the experiment-2 signature: the reader is alive and the scheduler is not.

  3. RSS against the container limit. The in-flight byte budget was removed (upstream has none), so worker count is the only bound on concurrent CAR ownership. Measured: 100 workers sits near 6 GB; 200 and 400 workers OOM-killed a 15 GB box. On a 61 GB workload box 100 workers has wide headroom, but this is now an unguarded edge: if RSS climbs toward the limit, reduce workers, never raise the limit. Experiment 3 reached 24.3 GB. Bootstrap in-flight is capped at 8 GiB by default; the limit must not equal the budget.

  4. Restart count. Any unplanned restart is a finding, not noise. Capture a receipt before investigating.

Deliberate restart test, once, early. Restart the pod on purpose and confirm the durable progress series continues without resetting. This confirms the July durable-progress defect is absent, at the lowest cost early in the run.

Stage 4 — steady observation to the deadline #

Capture a receipt before every intentional change and before teardown. Do not scale workers to chase throughput: experiment 3 went 100 → 200 and got slower. If throughput disappoints, record it and analyse afterwards.

Stage 4b — if the run succeeds and should keep running #

The watchdog is a cost guard, not a judge: at the deadline it deletes every server, volume, firewall, network and IP in the project regardless of whether the experiment worked. A successful whole-network archive is deleted on schedule unless someone intervenes. Decide before the deadline, not after.

Protect the archive volume as soon as the run looks like succeeding.

hcloud volume enable-protection <volume-id> delete

Hetzner Cloud has no volume snapshots — protection is the mechanism. The watchdog's DELETE /volumes/<id> then fails, the servers are still destroyed so compute billing stops, and the archive survives to be reattached to a fresh box. Two consequences to expect:

  • the watchdog's drain_kind volumes loop retries until it gives up, so its systemd unit exits non-zero. That is the protection working, not a fault.
  • destroy.sh will also fail on a protected volume. Disabling protection is a deliberate step in tearing the archive down; that is the point.

Retention is not free: a 2.5 TB volume bills whether attached or not. Keeping the archive is a standing cost, so it is a decision, not a default.

Promoting the experiment to a standing instance is a separate piece of work from keeping its data. The experiment shape is deliberately temporary — sslip.io hostnames derived from the workload IP, a public Grafana share, a watchdog holding a hard deadline. A standing instance wants real DNS, a deadline that is not hours away, and a decision about whether the operator surfaces stay public. None of that is required to preserve the archive, and none of it should be done in a hurry at hour 71.

Stage 5 — teardown #

Server, volume, primary IPs, firewall, observer and watchdog all deleted, then assert_empty_project.py. Then write the field log entry while it is fresh — experiment 3's record existed only in a temporary handoff and was nearly lost.

If things go poorly #

Symptom Most likely Do
Discovery climbs, durable completions stay 0 dispatch-boundary or batch-deadlock regression Stop. Receipt, teardown. Reproduce on a small box against the real network first (€0.03, ~90 seconds).
queued_for_completion 0 while repos are handled the experiment-4 deadlock Stop. Check queue_depth against workers × 2; a pinned queue confirms it.
Submitted-to-emitted gap grows live scheduler dead behind a live reader Receipt including stream.log. Do not restart first — the state is the evidence.
Process exits 1 after firehose connection closed reconnect regression Check the websocket pin actually shipped in this image.
RSS approaching the limit concurrent CAR ownership; there is no byte budget any more Receipt, then reduce workers. Never raise the limit. Note that early RSS growth is a working set, not a leak — measured 8.8 → 12.5 → 6.8 GB in the first half hour.
Repeated GetRepoFailed for the same DIDs remote hosts, expected Not fatal. Confirm host parking and backoff engage.
Everything green but nothing progressing process-local counters masking a stall Trust the durable series over the process-local ones.

On any failure: capture a receipt before restarting, and record what the run showed before terminating it.