Deployment runbook — next whole-network experiment #
Written against the record of the three previous attempts and their failure
modes. docs/full-network-experiment-2026-07.md is the durable field log;
docs/semantic-parity.md is the gate; this file is what to actually do, in
order, and what to do when a stage goes wrong.
Three terms used throughout, since they are load-bearing here and opaque cold:
- the gate —
semantic-parity.md, the row-by-row record of which behaviours are proven equivalent to upstream Jetstream V2. Nothing deploys on the strength of a passing test alone; the gate is what says a surface is actually settled. - a receipt — a JSON file under
receipts/binding one run of the full suite to one image digest (receipts/<sha>.json). It records each suite's verdict, and askippedverdict blocks admission exactly like a failure. An admitted artifact is an image whose digest has such a receipt with everything passing. - cutover — the moment the lifecycle leaves merge for steady state and serving stops returning 503. Before it, the archive is still being assembled; after it, clients are being answered.
Every stage must be measurable. All three previous runs ended because of a condition the running system could not report.
What happened the last three times #
1 — 14 July 2026, terminated the same day #
CPX62, 4 TB volume, ~$18. Killed by a hard-coded 512 MiB getRepo CAR
limit with no upstream equivalent: real repositories exceed it, so the crawl
excluded them and the archive was incomplete. The archive was discarded.
Commissioning also found a per-worker memory ceiling too small for a real CAR
(OOM treated as lifecycle-fatal), and pooled redirected connections producing
dozens of false HttpConnectionClosing failures per minute that direct curl
disproved.
Carried forward: no arbitrary size exclusion; incremental CAR consumption; exhaustion is a repository failure, not a process failure.
2 — 22 July 2026, destroyed early #
Isolated CCX43, independent watchdog, off-host observer, 2.5 TB XFS, ~$2.53. 26,099 repositories crossed durable completion; ~73.9 GB disk, 12.3 GB RSS.
Killed by a silent live-path death: the live scheduler stopped at 399,362 archived events while its WebSocket reader kept running and advanced its reconnect cursor 459,475 sequences past the last accepted frame. The process stayed up, bootstrap kept crawling, and ordinary health and backfill counters concealed it. The original error was not even retained — there was no terminal pipeline metric and no fail-loud capture supervision.
Carried forward: frames are acknowledged only after Stream accepts them; pipeline tickets, stage gaps and terminal failures must be scrapeable throughout bootstrap, not only in steady state.
3 — 23–24 July 2026, image e1926f3 #
Same isolated shape. The final receipt:
jetstream_backfill_discovered_total 100000
jetstream_backfill_completed_total 0
jetstream_backfill_progress_completed 0
process_resident_memory_bytes 24,263,409,664 (24.3 GB)
stream_events_total 618,558
/data/stream 26 GB
100,000 repositories were discovered and zero were durably complete. Completion was counted at the 100,000-entry dispatch boundary, so every restart — and the log shows four process starts — replayed nearly the whole batch while the dashboard's "durably committed" series climbed. That series was reading a process-local counter.
The run also crashed three separate ways (borrowed CAR/MST data outliving its
owner, an UnsupportedFloat row escaping as process-fatal, a shutdown
exceeding 45 s and being killed), took two unplanned restarts from ordinary
firehose disconnects, and got slower when scaled from 100 to 200 workers
(≈11 → ≈9 repos/s) with the in-flight limit saturated.
The deployed image was never admitted: e1926f3 included source and
dependency changes made after the receipt that covered it, and the pipeline
did not detect this.
What is different now #
Each failure mode above has a specific mitigation:
| Then | Now |
|---|---|
| 512 MiB CAR exclusion | removed; incremental consumption |
| Live path died silently | terminal pipeline signals; frames acked only on accept |
| Completion at dispatch boundary | per-repository, coupled to archive-writer durability |
| Progress panel read a process-local counter | stream_backfill_repos_durable, read from the store; survives restart (measured 22 → 49) |
| Firehose disconnect killed the process | websocket.zig setsockopt unreachable fixed; shipped via zat v0.3.18 |
| Unencodable record killed ingest | isolated per row, per wire |
| Retry failures hidden behind green health | terminal failure stops the service |
| Unadmitted image deployed | admit verify gates deploy.sh; it refuses sha256:adc276… |
| OOM retired a good repo as a malformed CAR | exhaustion stays OutOfMemory, pinned by an allocation-failure sweep |
| Stale operator CIDR could lock you out | provision.sh checks it against the live egress IP and refuses |
| Batch dispatch deadlocked before committing anything | submit and consume interleaved; verified against bsky.network |
| Global parse mutex serialized 100 workers behind one core | removed; upstream parses concurrently |
| In-flight byte budget blocked workers and could deadlock | removed; upstream has no such bound |
| Scoped to a target nobody had measured | Stage 0b feasibility check, with the head-of-list trap named |
| Experiment 3's tuning deployed unnoticed | assert_deploy_args.py refuses backfill flags no receipt covers |
Admitted artifact actually deployed for experiment 6:
sha256:0f45468c2fdba928fb3a9830f79d4693960a9c0ae2174ae042ecb01dfd33e0cd
(receipts/29705cb.json, 20/20 suites all pass, receipt cut 2026-07-28
04:36 UTC against an 05:01 UTC start, tag atcr.io/zat.dev/stream:29705cb).
This superseded eb95c14
(sha256:4300104135c6fd35bd13a264b744c5f842c41ee9a51e0cf1c57c9a59640551a2),
which earlier revisions of this file named as the artifact for the run. Both
were properly admitted; the runbook was not updated when the image moved.
Verify the running digest against the receipt: docker inspect the workload
container and compare to receipts/<sha>.json's image.registry_digest.
The eb95c14 fix below is carried forward into 29705cb: three sites
reported OutOfMemory as a structural verdict
about a repository's CAR, and the engine retires a repository permanently on
that verdict. Under whole-network memory pressure (experiment 3 reached 24.3 GB)
Stream would have discarded good repositories and recorded their PDS
as the cause. See the whole-network bootstrap row in
semantic-parity.md.
The plan #
Cost ceiling and deadline are set once, in provision.sh, and never silently
extended. Previous envelope: 24 h, $50, ~$2.53–$18 actual.
Stage 0 — before spending anything (~30 min) #
python3 assert_empty_project.py— the project must be empty.- Recompute the operator
/32; never reuse the one in.experiment.env. It is dynamic and has changed mid-run, which silently firewalls you out of your own workload while it keeps running fine.provision.shchecks it; nothing rechecks it during the run, so if probes go quiet, check this before concluding the workload is sick. ./scripts/admit verify <digest>must admit the artifact you intend to run.- Confirm the artifact matches the commit you think it does.
python3 assert_deploy_args.py— the flags must match what was admitted. Admission covers the artifact, not its invocation.
Stage 0b — feasibility (~20 min, €0.03) #
The gate's twenty correctness suites do not measure duration. An admitted artifact can still be far from finishing inside the deadline; correctness evidence is not feasibility evidence.
Measured 2026-07-28:
| Quantity | Value | How |
|---|---|---|
| repositories | ~19.4M | cursor-space binary search (22.7M) × measured density (0.854) |
| mean repo | 0.45 MB | 400 independent uniformly-random cursors |
| median repo | 0.01 MB | same |
| p99 / max | 10.5 MB / 69 MB | same |
| full pass | ~8.7 TB | product |
| our sustained rate | ~460 Mbps | CCX43, 100 workers |
| projection | ~42 h at 460 Mbps, ~19 h at 1 Gbps | arithmetic |
Upstream documents ~16 hours, which needs ~1.2 Gbps — a single server with a few TB of disk, matching their design doc's stated profile.
Repository size is heavy-tailed and listRepos is ordered
largest-first. The median repo is 10 KB; the first
100,000 entries average ~10 MB. Sampling consecutive entries, or extrapolating
early throughput, overstates the network by ~21× and turns a 42-hour job into
an apparent 40-day one. Sample independent random cursors, never consecutive
ones — consecutive entries share a PDS and creation cohort, so forty of them
is one sample, not forty.
Corollary for Stage 3: early throughput will look low and that is expected. The first hour crawls the heaviest accounts in the network. Judge progress by bytes moved, not repos completed, until the crawl is past the head.
If the projection exceeds the deadline, do not provision. Scale the goal to a bounded slice, or change the deadline.
Stage 1 — provision and arm (~20 min) #
CCX43 workload, 2.5 TB XFS, CPX12 watchdog, CPX12 observer, HEL1.
The watchdog is armed before the workload is trusted. It destroys billable resources on a hard deadline; "powered off" is not teardown.
If it fails: ./destroy.sh, verify the project is empty, then diagnose.
Stage 2 — deploy and first light (~15 min) #
deploy.sh refuses any digest admit verify does not admit — this is the
e1926f3 guard. Then:
/healthz,/readyzrespond; serving is gated (503) until steady state./metricsscrapes and pipeline tickets are already moving — experiment 2 died because these were only checked after steady state.- The observer, off-host, sees the public surface.
If it fails: capture a receipt first, then roll back. Experiment 2's root cause was lost because no receipt was taken.
Stage 3 — first hour, the one that matters #
Watch these four, in this order of importance:
-
stream_backfill_repos_durable{status="complete"}climbing, andjetstream_backfill_completion_queued_totalclimbing with it. Neither experiment 3 nor 4 ever produced these. If discovery climbs and they stay at zero for more than ~15 minutes of active crawling, stop — continuing only builds an archive a restart discards. The two zeros distinguish the causes: handled climbing withqueuedat zero is the batch deadlock (checkstream_backfill_queue_depthpinned atworkers × 2);queuedclimbing withdurableflat is a durability-boundary problem instead.Judge the rate by bytes, not repos. Sustained ~460 Mbps means the crawl is healthy even at 5 repos/s, because the head of
listReposis ~10 MB per repository against a 0.45 MB network mean. Repos/s should climb by more than an order of magnitude once past the head. A falling byte rate is the real alarm. -
Pipeline tickets submitted/claimed/emitted advancing together. A growing submitted-to-emitted gap is the experiment-2 signature: the reader is alive and the scheduler is not.
-
RSS against the container limit. The in-flight byte budget was removed (upstream has none), so worker count is the only bound on concurrent CAR ownership. Measured: 100 workers sits near 6 GB; 200 and 400 workers OOM-killed a 15 GB box. On a 61 GB workload box 100 workers has wide headroom, but this is now an unguarded edge: if RSS climbs toward the limit, reduce workers, never raise the limit. Experiment 3 reached 24.3 GB. Bootstrap in-flight is capped at 8 GiB by default; the limit must not equal the budget.
-
Restart count. Any unplanned restart is a finding, not noise. Capture a receipt before investigating.
Deliberate restart test, once, early. Restart the pod on purpose and confirm the durable progress series continues without resetting. This confirms the July durable-progress defect is absent, at the lowest cost early in the run.
Stage 4 — steady observation to the deadline #
Capture a receipt before every intentional change and before teardown. Do not scale workers to chase throughput: experiment 3 went 100 → 200 and got slower. If throughput disappoints, record it and analyse afterwards.
Stage 4b — if the run succeeds and should keep running #
The watchdog is a cost guard, not a judge: at the deadline it deletes every server, volume, firewall, network and IP in the project regardless of whether the experiment worked. A successful whole-network archive is deleted on schedule unless someone intervenes. Decide before the deadline, not after.
Protect the archive volume as soon as the run looks like succeeding.
hcloud volume enable-protection <volume-id> delete
Hetzner Cloud has no volume snapshots — protection is the mechanism. The
watchdog's DELETE /volumes/<id> then fails, the servers are still destroyed
so compute billing stops, and the archive survives to be reattached to a fresh
box. Two consequences to expect:
- the watchdog's
drain_kind volumesloop retries until it gives up, so its systemd unit exits non-zero. That is the protection working, not a fault. destroy.shwill also fail on a protected volume. Disabling protection is a deliberate step in tearing the archive down; that is the point.
Retention is not free: a 2.5 TB volume bills whether attached or not. Keeping the archive is a standing cost, so it is a decision, not a default.
Promoting the experiment to a standing instance is a separate piece of work from keeping its data. The experiment shape is deliberately temporary — sslip.io hostnames derived from the workload IP, a public Grafana share, a watchdog holding a hard deadline. A standing instance wants real DNS, a deadline that is not hours away, and a decision about whether the operator surfaces stay public. None of that is required to preserve the archive, and none of it should be done in a hurry at hour 71.
Stage 5 — teardown #
Server, volume, primary IPs, firewall, observer and watchdog all deleted, then
assert_empty_project.py. Then write the field log entry while it is fresh
— experiment 3's record existed only in a temporary handoff and was nearly lost.
If things go poorly #
| Symptom | Most likely | Do |
|---|---|---|
| Discovery climbs, durable completions stay 0 | dispatch-boundary or batch-deadlock regression | Stop. Receipt, teardown. Reproduce on a small box against the real network first (€0.03, ~90 seconds). |
queued_for_completion 0 while repos are handled |
the experiment-4 deadlock | Stop. Check queue_depth against workers × 2; a pinned queue confirms it. |
| Submitted-to-emitted gap grows | live scheduler dead behind a live reader | Receipt including stream.log. Do not restart first — the state is the evidence. |
Process exits 1 after firehose connection closed |
reconnect regression | Check the websocket pin actually shipped in this image. |
| RSS approaching the limit | concurrent CAR ownership; there is no byte budget any more | Receipt, then reduce workers. Never raise the limit. Note that early RSS growth is a working set, not a leak — measured 8.8 → 12.5 → 6.8 GB in the first half hour. |
Repeated GetRepoFailed for the same DIDs |
remote hosts, expected | Not fatal. Confirm host parking and backoff engage. |
| Everything green but nothing progressing | process-local counters masking a stall | Trust the durable series over the process-local ones. |
On any failure: capture a receipt before restarting, and record what the run showed before terminating it.