jetstream v2 in zig stream.waow.tech
stream docs experiment-6-handoff.md
21 kB
Markdown

Experiment 6 — operational handoff #

Written 2026-08-04 ~01:40Z, mid-merging. Everything here was verified against the running system, not recalled. Times are given in UTC with CDT (UTC-5) in parentheses — the operator is US Central; earlier logs in this repo that say "EDT" are an hour off and mislabeled.

Companion reading, in priority order:

  • ~/.claude/projects/-Users-nate-tangled-org-zzstoatzz-io-relay/memory/ — the behavioural memories, especially feedback_monitoring_window_discipline, stream-accountcount-not-fetchable, stream-backfill-progress-metrics.
  • ~/tangled.org/zzstoatzz.io/notes/protocols/atproto/network-backfill.md — why the crawl behaves the way it does.
  • docs/full-network-experiment-2026-07.md — the running field log.
  • docs/gotchas.md — traps, including the gzip/transfer-buffer one.

1. What is running, in one paragraph #

A whole-network atproto backfill built into a Jetstream V2 archive. Bootstrap crawled bsky.network for ~145h and completed 16.26M repos (~1.7 TB, ~20B events). On 2026-08-03 07:04:58Z bootstrap was deliberately cut over to merging because the remaining tail was unfetchable, not slow (see §7). Merge drained its live-capture archive cleanly. It is now in the pending-repo pass, after which come discovery, rotate, compaction, and the steady_state commit that ungates serving. Nothing else is required of an operator until that happens.

The single question still open: can this archive actually serve? That is unproven until listSegments returns 200 and segments read back and decode.


2. Access #

SSH key ~/.ssh/waow_ed25519, user root. Always -o IdentitiesOnly=yes (the agent offers the wrong key otherwise).

host IP role
stream-20260728-0501 37.27.43.39 workload — the archive lives here
stream-20260728-0501-watchdog 65.109.0.222 kill-switch, holds the Hetzner API token
stream-20260728-0501-observer 37.27.41.161 idle this run

Hetzner firewalls: SSH (22) only from 99.64.254.139/32 and 73.50.144.240/32. These are dynamic operator IPs and have changed mid-run before — if SSH hangs, recheck your public IP before concluding the host is sick. Hetzner also recycles IPs between experiments, so a host-key mismatch on a new run is expected: ssh-keygen -R <ip>. Ports 80/443 on the workload are open to 0.0.0.0/0.

Public endpoints (Caddy, automatic TLS, no DNS needed — sslip.io encodes the IP):

  • https://stream.37-27-43-39.sslip.io/status — human-readable field report
  • https://dash.37-27-43-39.sslip.io — the live dashboard (§6)
  • https://grafana.37-27-43-39.sslip.io — password in /opt/stream-experiment/grafana-password.txt

There is no hcloud token configured locally. The only Hetzner API token is /etc/stream-watchdog/token on the watchdog. It has destroy permissions — use it read-only unless you intend to destroy something.

Reaching metrics #

The stream container has no curl or wget. Go through the prometheus container:

# PromQL
docker compose exec -T prometheus wget -qO- \
  'http://127.0.0.1:9090/api/v1/query?query=<urlencoded>'

# raw stream metrics / serving surface (container IP, DNS from busybox is flaky)
docker compose exec -T prometheus wget -qO- http://172.18.0.2:6060/metrics
docker compose exec -T prometheus wget -S -qO- http://172.18.0.2:8080/status

Ports: 6060 = metrics/debug, 8080 = serving. 172.18.0.2 is the stream container; re-read it with docker inspect if the container is recreated.

Local vs remote tooling — do not mix these up #

On the operator's machine: never invoke python3 or pip. Use uv / uvx / uv run, or jq for JSON. This is a standing rule in the operator's global instructions.

On the Hetzner hosts, uv is simply not installed — system python3 is what is there, and jq is present only on the watchdog. So the scripts in deploy/ops/ that run on the box (wsprobe.py, dash-snapshot) target system python3. That is a fact about how these hosts were provisioned, not a constraint: uv is a single static binary (curl -LsSf https://astral.sh/uv/install.sh | sh) and installing it on the hosts would let both sides use the same tooling. Worth doing on the next provision; not worth changing under a live run.


3. Deadlines — the replay window binds before the watchdog #

Two independent clocks. The earlier one is not the watchdog.

clock when what happens
firehose replay window 2026-08-06 07:04:58Z (Thu 2:04 AM CDT) the archive acquires a permanent hole
watchdog 2026-08-06 17:01:47Z (Thu 12:01 PM CDT) servers and the volume are deleted

Replay window. Live capture was torn down at phase -> merging (1785740698, 2026-08-03 07:04:58Z). Since then nothing is being captured. At steady_state the consumer resumes from the shared relay cursor and replays forward. bsky.network retains 72h (operator-confirmed). Miss it and events from 2026-08-03 07:04:58Z onward are gone for good — worse than finishing late.

Probe whether replay still works (returns 101 + binary commit frames if healthy):

ssh -i ~/.ssh/waow_ed25519 -o IdentitiesOnly=yes root@37.27.43.39 \
  '/usr/local/bin/wsprobe 32408688748'

32408688748 is the seq the live capture stopped at. Source is deploy/ops/wsprobe.py; installed to /usr/local/bin/wsprobe on the workload (it was in /tmp initially, which does not survive a reboot).

Watchdog. /etc/stream-watchdog/ on 65.109.0.222: armed (flag file), deadline (epoch), token. stream-experiment-expire.timer fires /usr/local/sbin/stream-experiment-expire every 60s; when now >= deadline it deletes every server, then drain_kind volumes, then itself.

It deletes the volume. The archive and the pre-merge snapshot are both on /dev/sdb. When it fires, everything is gone.

Current value 1786035707. Previous value backed up at /etc/stream-watchdog/deadline.bak-20260803T145236Z.


4. Current state and how to read it #

bash deploy/ops/check6.sh                  # from the stream repo; state in ~/.stream-exp6
curl -s https://dash.37-27-43-39.sslip.io/data.json | jq .

The ops scripts were originally written into a session scratchpad that does not survive. They now live in deploy/ops/ in this repo — check6.sh, wsprobe.py, dash-snapshot, dash-index.html. See deploy/ops/README.md.

As of 2026-08-04 01:33Z: phase 2 (merging), complete 16,260,762, pending 25,002 draining ~4,400/hr, listSegments 503 (correct), 0 stale .tmp, backfill/ 40 GB awaiting cleanup, restarts 1 (disk-full, 08-04 13:26Z, self-healed), disk 1.3 TB free, RSS ~56 GB, €145 spent.

Phase encoding: 1 bootstrap, 2 merging, 3 steady_state.

Metric literacy — these have all burned me #

  • phase = NONE means mid-restart, not a fault. Every other value is untrustworthy in that moment. Startup takes ~5 min (bloom load); wait it out.
  • Per-process counters reset on every restart. jetstream_ingest_events_appended_total sawtooths — it read 16.5M, 14.5M, 6.6M on successive reads. Only increase(...[24h]) is meaningful. The cumulative truth is the archive's next seq in the startup log (~19.97B).
  • stream_backfill_repos_durable{status="complete"} is the only monotonic bucket. The bucket sum is enumeration progress, not network size.
  • Check sample spacing before calling a trend. I twice read near-simultaneous samples as a time series and nearly escalated. Nested windows (1h/6h/12h) must move together, and a rate sampled while not_started ≈ 0 is a batch trough.
  • RSS ~46-56 GB is normal — resident bloom filters, not a leak.

5. Runbooks #

A0. Restarts (updated 2026-08-08) #

A restart costs ~4 min of archive recovery, then — with compaction enabled, which production runs since 2026-08-08 03:51Z — a synchronous tombstone rebuild before ingest dials. The rebuild logs a start line (with the loaded watermark and segment count) and progress every 200 segments, but it visits every segment file even below the committed watermark: ~3 h on the cx43 until the header-level skip fix lands. Plan restarts accordingly: the tail is deaf for the scan, the archive serves throughout, consumers ride their fallback hosts. The live instance is stream-cx43 (89.167.122.160); the watermark is committed (23,135,908,535 as of enablement) and advances with each compaction pass (4 h cadence from compactor start).

A. Routine check (nothing wrong) #

Read phase first. Confirm: listSegments 503 while phase < 3, stale_tmp 0, disk free » 0.4 TB, oom-kills still 1 (the one on 2026-07-29 06:06). Report two lines. Do not deploy, do not tear down.

B. Phase flips 2 → 3 (the win) — how to actually prove it #

A 200 is not proof. Do all of this:

# 1. gate opened
docker compose exec -T prometheus wget -S -qO- http://172.18.0.2:8080/xrpc/jetstream.listSegments
# 2. merge cleanup completed
du -sh /data/stream/backfill        # should be GONE, not 40 GB
find /data/stream -name '*.tmp' | wc -l   # 0
# 3. read data back out and decode it — the actual proof
#    pull a segment via getSegment and confirm it parses as jss
# 4. subscribe surface serves live events after replay catches up
# 5. capture a receipt: phase, counters, segment count, disk, timestamps

Then re-run the cursor probe: if replay was refused, say so plainly — a serving instance with a hole is not a success.

C. Compaction running long / replay window closing #

How to tell it started and is progressing. There is no "compaction started" log line to wait for. Watch stream_compaction_watermark_seq — it is 0 until the pass begins and then advances (compact/pass.zig:131). Pair it with stream_compaction_watermark_lag_us. CPU stays high throughout; that alone tells you nothing about progress.

docker compose exec -T prometheus wget -qO- \
  'http://127.0.0.1:9090/api/v1/query?query=stream_compaction_watermark_seq'

Compaction over 1.6 TB has never been measured. If steady_state has not committed and the replay window has < ~6h left, that is the decision point: finishing late is recoverable, a hole is not. Surface it immediately rather than letting it lapse quietly.

D. Process crashes during merging #

Expected to self-heal. restart: unless-stopped, and the crash matrix at lifecycle.zig:13 covers merging re-entry at-least-once (seal-guard + per-source cursor). Confirm recovery: resumed active segment appears after the restart. Escalate only if recovery stops resuming.

E. Roll back to the pre-merge snapshot — GONE as of 2026-08-04 13:4xZ #

The snapshot was deleted during the disk-full incident (see below). There is no rollback-to-premerge path anymore; the live archive is the only copy.

What happened, kept as the cautionary tale: the reflink snapshot cost ~0 space at creation, but merge + the pending pass rewrote enough segments that its unique blocks silently consumed the entire 1.3 TB of headroom. /data hit 100% (263 MB free) at 13:26Z on 08-04, the process crashed once on NoSpaceLeft, and recovery resumed cleanly on restart. Deleting the snapshot reclaimed 1.3 TB. df cannot warn you about reflink divergence — the space is consumed by writes to the live tree breaking sharing, and free space slides without any directory growing. If you take a CoW snapshot before a phase that rewrites data, budget for full divergence or monitor free space directly.

F. Move or remove the watchdog #

# extend (preferred — keeps the ceiling, just later)
ssh root@65.109.0.222 'cp -a /etc/stream-watchdog/deadline{,.bak-$(date -u +%Y%m%dT%H%M%SZ)}
  printf "%s\n" <new_epoch> > /etc/stream-watchdog/deadline
  d=$(cat /etc/stream-watchdog/deadline); now=$(date +%s); echo "left_h=$(( (d-now)/3600 ))"'

# disarm entirely (removes the cost ceiling — get explicit sign-off)
ssh root@65.109.0.222 'rm /etc/stream-watchdog/armed'

Always read back through the same path the script uses. Cost of keeping it alive is €0.84/h (~€20/day).

G. Deliberate teardown #

Let the watchdog do it — that is what it is for. Set deadline to a near-future epoch and confirm the project empties. Note: docs/deployment-runbook.md cites assert_empty_project.py, but that script is not in this repo — it lived in provisioning tooling that was never committed. Verify emptiness directly against the Hetzner API with the watchdog's token instead:

ssh root@65.109.0.222 't=$(cat /etc/stream-watchdog/token)
  for k in servers volumes floating_ips primary_ips; do
    printf "%s: " $k
    curl -fsS -H "Authorization: Bearer $t" "https://api.hetzner.cloud/v1/$k?per_page=50" | jq ".$k | length"
  done'

Do not hand-delete resources; the script's ordering (servers → volumes → IPs → self) exists because Hetzner refuses to delete an attached volume.

I. The replay window lapsed (worst case) #

wsprobe returns a non-101 handshake or an #info/OutdatedCursor frame. This means events from 2026-08-03 07:04:58Z to whenever ingest resumes are not recoverable from bsky.network.

Do not quietly continue. State it plainly, then:

  1. Confirm it. One failed probe is not proof — re-run, and try a cursor a few thousand seqs later to distinguish "this cursor aged out" from "the relay is refusing connections."
  2. Let it finish anyway. A partial archive with a known, dated hole is still a usable artifact and still answers the serving question. Do not tear down.
  3. Record the hole precisely — first and last missing seq — in the field log and in any receipt. An undocumented gap is far worse than a documented one.
  4. Repair is possible but is a separate project: re-running backfill for affected repos picks up their current state, since getRepo returns the repo as it is now. It does not recover the intermediate events.

H. Disk pressure #

Floor is 0.4 TB free. The snapshot lever was used on 2026-08-04 (see E) — it no longer exists. If disk tightens again the consumers are compaction rewrites and the archive itself, and the remaining lever is growing the volume via the Hetzner API (token on the watchdog) + xfs_growfs /data, as was done 2.5→3.0 TB on 07-30. The session monitor now warns below 700 GB free and alarms below the 400 GB floor.


6. The dashboard #

https://dash.37-27-43-39.sslip.io — free, rides the existing Caddy.

  • /usr/local/bin/dash-snapshot (python) writes /srv/dash/data.json
  • dash-snapshot.timer runs it every 60s
  • /srv/dash/index.html fetches data.json every 30s, same-origin
  • Caddy block appended to /opt/stream-experiment/Caddyfile; /srv/dash mounted read-only into the caddy container

It dies with the host on Thursday. Nothing outlives the volume. Backups: Caddyfile.bak-20260803T153044Z, compose.yaml.bak-pre-merge-20260803.


7. What happened, and the one config change #

when (UTC) event
07-28 05:01:47 bootstrap begins — ccx43, 2.5 TB, 100 workers, image 29705cb
07-29 06:06 one oom-kill (the only one all run)
07-30 ~19:00 deadline 120h→180h; volume 2.5→3.0 TB online via xfs_growfs
08-03 07:04:58 bootstrap cut over to merging (see below)
08-03 08:05:54 merge drain: 255,799,323 kept / 126,488,978 dropped across 137 sources
08-03 14:52 deadline 180h→228h (1786035707)
08-04 01:33 pending pass draining, 25,002 left

The cutover, and the flag #

The crawl fell to ~7.5 repos/s and stayed there. It was not degradation: 15 failing DIDs sampled through plc.directory showed 10 of 15 on atproto.brid.gy, which 301s to bridgy-hubble.microcosm.blue and answers {"error":"RepoNotFound","message":"repo not synchronized yet"}. Following redirects, 4 of 15 fetchable. Those accounts are counted by listHosts accountCount but their repos are not served by anyone — so the 19.4M target was never reachable.

Worse, merging commits only when the engine returns, which requires listRepos enumeration to reach its final page. A slow crawl also slows that approach, so merge would never have fired before the deadline.

Fix: added --backfill-max-repos=1 to the compose command and restarted. The engine breaks on max_reached or final_page (engine.zig:849) — both reach the same commit point.

Trap, verified by reading the code first: max_repos > 0 sets full_network = false (engine.zig:731), which skips loading the durable listRepos cursor and skips saveListReposCheckpoint. It does not delete the cursor (only selected_repos does that, line 730). Safe here because we were done crawling. Do not set this flag on a run you intend to resume.

To resume crawling later, remove the flag; the cursor at ~16.3M is intact.


8. Cost #

Verified against Hetzner's pricing API, not recalled.

resource €/h
ccx43 workload 0.528
cpx12 × 2 (watchdog, observer) 0.0288 each
volume 0.0767 /GB-month → 3 TB ≈ 0.315/h

Running: €0.84/h ≈ €20/day. Spent on experiment 6: ~€145 at 154h. Prior exploratory runs (experiments 1-5) were ~$25-30 total — cheap because they were stopped early and deliberately.

Steady-state, once right-sized (the ccx43 is sized for 100 crawl workers and is overkill for ingest+serve):

shape €/month
ccx43 + 3 TB (now) ~616
ccx33 + 2 TB ~348
ccx23 + 2 TB ~274

Right-size from a measurement after cutover, not from this table. Steady-state CPU has never been observed. Storage dominates and grows.


9. Invariants #

  1. Never let the replay window lapse before steady_state. A hole is permanent; lateness is not.
  2. The watchdog deletes the volume. Not just servers.
  3. Never conclude "absent" or "new" from your own grep or metric selector. I claimed RepairRateLimited was new when it was 38h old and merely missing from my pattern; I claimed the watchdog was unverifiable when it was two lookups away under a different unit name.
  4. Verify a flag's semantics in the source before setting it on a live run. --backfill-max-repos would have silently disabled cursor checkpointing.
  5. Do not deploy new code mid-run. zat 0.3.23 (which fixes the ~40-minute restart loop and gzip) is landed and pinned but not deployed; the running image is 29705cb.
  6. A 200 from listSegments is not proof of a working archive. Read data back.

10. When this run ends — carry these forward #

  • Deploy zat 0.3.23. Already pinned in build.zig.zon and pushed, never deployed. It fixes the ~40-minute restart loop (a pooled connection the relay reaped, surfacing as HttpConnectionClosing), makes gzip actually work on the getRepo hot path (~2.5x less data pulled from other operators' PDSes), and bounds stalled transfers. This run restarted 163 times because of that bug; each restart cost ~5 minutes of bloom-loading startup, roughly 12% of wall clock.
  • Right-size the box from a measurement, not from the table in §8.
  • Re-derive the denominator honestly. accountCount is an upper bound; see stream-accountcount-not-fetchable in memory.
  • The observer host was idle all run (load 0.24). If it is not earning its keep next time, do not provision it.

11. Open questions #

  • Can it serve? Unproven. This is the whole remaining point.
  • How long does compaction take on 1.6 TB? No measurement exists.
  • Is the kept/dropped ratio right? Final was 2.02:1 kept-over-dropped, the inverse of the brief's dropped >> kept expectation. The filter is matching (126M dropped ≠ 0). My reasoning — that repos backfilled early accumulate ~145h of correctly-kept live events, and largest-first ordering front-loads exactly the busiest repos — is consistent but unverified against segment timestamps. The failure it could indicate is duplicate events, not loss.
  • Actual coverage. ~16.26M against a 22.85M accountCount; fig and calabro both report 18-20M for relay-known repos on a ~41M reachable network. Where we really landed is worth measuring properly, not asserting.