Experiment 6 — operational handoff #
Written 2026-08-04 ~01:40Z, mid-merging. Everything here was verified against
the running system, not recalled. Times are given in UTC with CDT (UTC-5) in
parentheses — the operator is US Central; earlier logs in this repo that say
"EDT" are an hour off and mislabeled.
Companion reading, in priority order:
~/.claude/projects/-Users-nate-tangled-org-zzstoatzz-io-relay/memory/— the behavioural memories, especiallyfeedback_monitoring_window_discipline,stream-accountcount-not-fetchable,stream-backfill-progress-metrics.~/tangled.org/zzstoatzz.io/notes/protocols/atproto/network-backfill.md— why the crawl behaves the way it does.docs/full-network-experiment-2026-07.md— the running field log.docs/gotchas.md— traps, including the gzip/transfer-buffer one.
1. What is running, in one paragraph #
A whole-network atproto backfill built into a Jetstream V2 archive. Bootstrap
crawled bsky.network for ~145h and completed 16.26M repos (~1.7 TB, ~20B
events). On 2026-08-03 07:04:58Z bootstrap was deliberately cut over to merging
because the remaining tail was unfetchable, not slow (see §7). Merge drained
its live-capture archive cleanly. It is now in the pending-repo pass, after
which come discovery, rotate, compaction, and the steady_state commit that
ungates serving. Nothing else is required of an operator until that happens.
The single question still open: can this archive actually serve? That is
unproven until listSegments returns 200 and segments read back and decode.
2. Access #
SSH key ~/.ssh/waow_ed25519, user root. Always
-o IdentitiesOnly=yes (the agent offers the wrong key otherwise).
| host | IP | role |
|---|---|---|
stream-20260728-0501 |
37.27.43.39 |
workload — the archive lives here |
stream-20260728-0501-watchdog |
65.109.0.222 |
kill-switch, holds the Hetzner API token |
stream-20260728-0501-observer |
37.27.41.161 |
idle this run |
Hetzner firewalls: SSH (22) only from 99.64.254.139/32 and 73.50.144.240/32.
These are dynamic operator IPs and have changed mid-run before — if SSH hangs,
recheck your public IP before concluding the host is sick. Hetzner also recycles
IPs between experiments, so a host-key mismatch on a new run is expected:
ssh-keygen -R <ip>. Ports 80/443 on the
workload are open to 0.0.0.0/0.
Public endpoints (Caddy, automatic TLS, no DNS needed — sslip.io encodes the IP):
https://stream.37-27-43-39.sslip.io/status— human-readable field reporthttps://dash.37-27-43-39.sslip.io— the live dashboard (§6)https://grafana.37-27-43-39.sslip.io— password in/opt/stream-experiment/grafana-password.txt
There is no hcloud token configured locally. The only Hetzner API token is
/etc/stream-watchdog/token on the watchdog. It has destroy permissions — use it
read-only unless you intend to destroy something.
Reaching metrics #
The stream container has no curl or wget. Go through the prometheus
container:
# PromQL
docker compose exec -T prometheus wget -qO- \
'http://127.0.0.1:9090/api/v1/query?query=<urlencoded>'
# raw stream metrics / serving surface (container IP, DNS from busybox is flaky)
docker compose exec -T prometheus wget -qO- http://172.18.0.2:6060/metrics
docker compose exec -T prometheus wget -S -qO- http://172.18.0.2:8080/status
Ports: 6060 = metrics/debug, 8080 = serving. 172.18.0.2 is the stream
container; re-read it with docker inspect if the container is recreated.
Local vs remote tooling — do not mix these up #
On the operator's machine: never invoke python3 or pip. Use uv /
uvx / uv run, or jq for JSON. This is a standing rule in the operator's
global instructions.
On the Hetzner hosts, uv is simply not installed — system python3 is what
is there, and jq is present only on the watchdog. So the scripts in
deploy/ops/ that run on the box (wsprobe.py, dash-snapshot) target system
python3. That is a fact about how these hosts were provisioned, not a
constraint: uv is a single static binary
(curl -LsSf https://astral.sh/uv/install.sh | sh) and installing it on the
hosts would let both sides use the same tooling. Worth doing on the next
provision; not worth changing under a live run.
3. Deadlines — the replay window binds before the watchdog #
Two independent clocks. The earlier one is not the watchdog.
| clock | when | what happens |
|---|---|---|
| firehose replay window | 2026-08-06 07:04:58Z (Thu 2:04 AM CDT) | the archive acquires a permanent hole |
| watchdog | 2026-08-06 17:01:47Z (Thu 12:01 PM CDT) | servers and the volume are deleted |
Replay window. Live capture was torn down at phase -> merging
(1785740698, 2026-08-03 07:04:58Z). Since then nothing is being captured.
At steady_state the consumer resumes from the shared relay cursor and replays
forward. bsky.network retains 72h (operator-confirmed). Miss it and events
from 2026-08-03 07:04:58Z onward are gone for good — worse than finishing late.
Probe whether replay still works (returns 101 + binary commit frames if healthy):
ssh -i ~/.ssh/waow_ed25519 -o IdentitiesOnly=yes root@37.27.43.39 \
'/usr/local/bin/wsprobe 32408688748'
32408688748 is the seq the live capture stopped at. Source is
deploy/ops/wsprobe.py; installed to /usr/local/bin/wsprobe on the workload
(it was in /tmp initially, which does not survive a reboot).
Watchdog. /etc/stream-watchdog/ on 65.109.0.222: armed (flag file),
deadline (epoch), token. stream-experiment-expire.timer fires
/usr/local/sbin/stream-experiment-expire every 60s; when
now >= deadline it deletes every server, then drain_kind volumes, then
itself.
It deletes the volume. The archive and the pre-merge snapshot are both on
/dev/sdb. When it fires, everything is gone.
Current value 1786035707. Previous value backed up at
/etc/stream-watchdog/deadline.bak-20260803T145236Z.
4. Current state and how to read it #
bash deploy/ops/check6.sh # from the stream repo; state in ~/.stream-exp6
curl -s https://dash.37-27-43-39.sslip.io/data.json | jq .
The ops scripts were originally written into a session scratchpad that does not
survive. They now live in deploy/ops/ in this repo — check6.sh,
wsprobe.py, dash-snapshot, dash-index.html. See deploy/ops/README.md.
As of 2026-08-04 01:33Z: phase 2 (merging), complete 16,260,762,
pending 25,002 draining ~4,400/hr, listSegments 503 (correct),
0 stale .tmp, backfill/ 40 GB awaiting cleanup, restarts 1 (disk-full, 08-04 13:26Z, self-healed), disk 1.3 TB
free, RSS ~56 GB, €145 spent.
Phase encoding: 1 bootstrap, 2 merging, 3 steady_state.
Metric literacy — these have all burned me #
phase = NONEmeans mid-restart, not a fault. Every other value is untrustworthy in that moment. Startup takes ~5 min (bloom load); wait it out.- Per-process counters reset on every restart.
jetstream_ingest_events_appended_totalsawtooths — it read 16.5M, 14.5M, 6.6M on successive reads. Onlyincrease(...[24h])is meaningful. The cumulative truth is the archive'snext seqin the startup log (~19.97B). stream_backfill_repos_durable{status="complete"}is the only monotonic bucket. The bucket sum is enumeration progress, not network size.- Check sample spacing before calling a trend. I twice read near-simultaneous
samples as a time series and nearly escalated. Nested windows (1h/6h/12h) must
move together, and a rate sampled while
not_started≈ 0 is a batch trough. - RSS ~46-56 GB is normal — resident bloom filters, not a leak.
5. Runbooks #
A0. Restarts (updated 2026-08-08) #
A restart costs ~4 min of archive recovery, then — with compaction enabled, which production runs since 2026-08-08 03:51Z — a synchronous tombstone rebuild before ingest dials. The rebuild logs a start line (with the loaded watermark and segment count) and progress every 200 segments, but it visits every segment file even below the committed watermark: ~3 h on the cx43 until the header-level skip fix lands. Plan restarts accordingly: the tail is deaf for the scan, the archive serves throughout, consumers ride their fallback hosts. The live instance is stream-cx43 (89.167.122.160); the watermark is committed (23,135,908,535 as of enablement) and advances with each compaction pass (4 h cadence from compactor start).
A. Routine check (nothing wrong) #
Read phase first. Confirm: listSegments 503 while phase < 3, stale_tmp 0,
disk free » 0.4 TB, oom-kills still 1 (the one on 2026-07-29 06:06). Report
two lines. Do not deploy, do not tear down.
B. Phase flips 2 → 3 (the win) — how to actually prove it #
A 200 is not proof. Do all of this:
# 1. gate opened
docker compose exec -T prometheus wget -S -qO- http://172.18.0.2:8080/xrpc/jetstream.listSegments
# 2. merge cleanup completed
du -sh /data/stream/backfill # should be GONE, not 40 GB
find /data/stream -name '*.tmp' | wc -l # 0
# 3. read data back out and decode it — the actual proof
# pull a segment via getSegment and confirm it parses as jss
# 4. subscribe surface serves live events after replay catches up
# 5. capture a receipt: phase, counters, segment count, disk, timestamps
Then re-run the cursor probe: if replay was refused, say so plainly — a serving instance with a hole is not a success.
C. Compaction running long / replay window closing #
How to tell it started and is progressing. There is no "compaction started"
log line to wait for. Watch stream_compaction_watermark_seq — it is 0 until
the pass begins and then advances (compact/pass.zig:131). Pair it with
stream_compaction_watermark_lag_us. CPU stays high throughout; that alone tells
you nothing about progress.
docker compose exec -T prometheus wget -qO- \
'http://127.0.0.1:9090/api/v1/query?query=stream_compaction_watermark_seq'
Compaction over 1.6 TB has never been measured. If steady_state has not
committed and the replay window has < ~6h left, that is the decision point:
finishing late is recoverable, a hole is not. Surface it immediately rather than
letting it lapse quietly.
D. Process crashes during merging #
Expected to self-heal. restart: unless-stopped, and the crash matrix at
lifecycle.zig:13 covers merging re-entry at-least-once (seal-guard +
per-source cursor). Confirm recovery: resumed active segment appears after the
restart. Escalate only if recovery stops resuming.
E. Roll back to the pre-merge snapshot — GONE as of 2026-08-04 13:4xZ #
The snapshot was deleted during the disk-full incident (see below). There is no rollback-to-premerge path anymore; the live archive is the only copy.
What happened, kept as the cautionary tale: the reflink snapshot cost ~0 space
at creation, but merge + the pending pass rewrote enough segments that its
unique blocks silently consumed the entire 1.3 TB of headroom. /data hit 100%
(263 MB free) at 13:26Z on 08-04, the process crashed once on NoSpaceLeft, and
recovery resumed cleanly on restart. Deleting the snapshot reclaimed 1.3 TB.
df cannot warn you about reflink divergence — the space is consumed by
writes to the live tree breaking sharing, and free space slides without any
directory growing. If you take a CoW snapshot before a phase that rewrites data,
budget for full divergence or monitor free space directly.
F. Move or remove the watchdog #
# extend (preferred — keeps the ceiling, just later)
ssh root@65.109.0.222 'cp -a /etc/stream-watchdog/deadline{,.bak-$(date -u +%Y%m%dT%H%M%SZ)}
printf "%s\n" <new_epoch> > /etc/stream-watchdog/deadline
d=$(cat /etc/stream-watchdog/deadline); now=$(date +%s); echo "left_h=$(( (d-now)/3600 ))"'
# disarm entirely (removes the cost ceiling — get explicit sign-off)
ssh root@65.109.0.222 'rm /etc/stream-watchdog/armed'
Always read back through the same path the script uses. Cost of keeping it alive is €0.84/h (~€20/day).
G. Deliberate teardown #
Let the watchdog do it — that is what it is for. Set deadline to a near-future
epoch and confirm the project empties. Note: docs/deployment-runbook.md
cites assert_empty_project.py, but that script is not in this repo — it
lived in provisioning tooling that was never committed. Verify emptiness directly
against the Hetzner API with the watchdog's token instead:
ssh root@65.109.0.222 't=$(cat /etc/stream-watchdog/token)
for k in servers volumes floating_ips primary_ips; do
printf "%s: " $k
curl -fsS -H "Authorization: Bearer $t" "https://api.hetzner.cloud/v1/$k?per_page=50" | jq ".$k | length"
done'
Do not hand-delete resources; the script's ordering (servers → volumes → IPs → self) exists because Hetzner refuses to delete an attached volume.
I. The replay window lapsed (worst case) #
wsprobe returns a non-101 handshake or an #info/OutdatedCursor frame. This
means events from 2026-08-03 07:04:58Z to whenever ingest resumes are not
recoverable from bsky.network.
Do not quietly continue. State it plainly, then:
- Confirm it. One failed probe is not proof — re-run, and try a cursor a few thousand seqs later to distinguish "this cursor aged out" from "the relay is refusing connections."
- Let it finish anyway. A partial archive with a known, dated hole is still a usable artifact and still answers the serving question. Do not tear down.
- Record the hole precisely — first and last missing seq — in the field log and in any receipt. An undocumented gap is far worse than a documented one.
- Repair is possible but is a separate project: re-running backfill for
affected repos picks up their current state, since
getReporeturns the repo as it is now. It does not recover the intermediate events.
H. Disk pressure #
Floor is 0.4 TB free. The snapshot lever was used on 2026-08-04 (see E)
— it no longer exists. If disk tightens again the consumers are compaction
rewrites and the archive itself, and the remaining lever is growing the volume
via the Hetzner API (token on the watchdog) + xfs_growfs /data, as was done
2.5→3.0 TB on 07-30. The session monitor now warns below 700 GB free and
alarms below the 400 GB floor.
6. The dashboard #
https://dash.37-27-43-39.sslip.io — free, rides the existing Caddy.
/usr/local/bin/dash-snapshot(python) writes/srv/dash/data.jsondash-snapshot.timerruns it every 60s/srv/dash/index.htmlfetchesdata.jsonevery 30s, same-origin- Caddy block appended to
/opt/stream-experiment/Caddyfile;/srv/dashmounted read-only into the caddy container
It dies with the host on Thursday. Nothing outlives the volume. Backups:
Caddyfile.bak-20260803T153044Z, compose.yaml.bak-pre-merge-20260803.
7. What happened, and the one config change #
| when (UTC) | event |
|---|---|
| 07-28 05:01:47 | bootstrap begins — ccx43, 2.5 TB, 100 workers, image 29705cb |
| 07-29 06:06 | one oom-kill (the only one all run) |
| 07-30 ~19:00 | deadline 120h→180h; volume 2.5→3.0 TB online via xfs_growfs |
| 08-03 07:04:58 | bootstrap cut over to merging (see below) |
| 08-03 08:05:54 | merge drain: 255,799,323 kept / 126,488,978 dropped across 137 sources |
| 08-03 14:52 | deadline 180h→228h (1786035707) |
| 08-04 01:33 | pending pass draining, 25,002 left |
The cutover, and the flag #
The crawl fell to ~7.5 repos/s and stayed there. It was not degradation: 15
failing DIDs sampled through plc.directory showed 10 of 15 on atproto.brid.gy,
which 301s to bridgy-hubble.microcosm.blue and answers
{"error":"RepoNotFound","message":"repo not synchronized yet"}. Following
redirects, 4 of 15 fetchable. Those accounts are counted by listHosts accountCount but their repos are not served by anyone — so the 19.4M target was
never reachable.
Worse, merging commits only when the engine returns, which requires
listRepos enumeration to reach its final page. A slow crawl also slows that
approach, so merge would never have fired before the deadline.
Fix: added --backfill-max-repos=1 to the compose command and restarted. The
engine breaks on max_reached or final_page (engine.zig:849) — both reach the
same commit point.
Trap, verified by reading the code first:
max_repos > 0setsfull_network = false(engine.zig:731), which skips loading the durable listRepos cursor and skipssaveListReposCheckpoint. It does not delete the cursor (onlyselected_reposdoes that, line 730). Safe here because we were done crawling. Do not set this flag on a run you intend to resume.
To resume crawling later, remove the flag; the cursor at ~16.3M is intact.
8. Cost #
Verified against Hetzner's pricing API, not recalled.
| resource | €/h |
|---|---|
| ccx43 workload | 0.528 |
| cpx12 × 2 (watchdog, observer) | 0.0288 each |
| volume | 0.0767 /GB-month → 3 TB ≈ 0.315/h |
Running: €0.84/h ≈ €20/day. Spent on experiment 6: ~€145 at 154h. Prior exploratory runs (experiments 1-5) were ~$25-30 total — cheap because they were stopped early and deliberately.
Steady-state, once right-sized (the ccx43 is sized for 100 crawl workers and is overkill for ingest+serve):
| shape | €/month |
|---|---|
| ccx43 + 3 TB (now) | ~616 |
| ccx33 + 2 TB | ~348 |
| ccx23 + 2 TB | ~274 |
Right-size from a measurement after cutover, not from this table. Steady-state CPU has never been observed. Storage dominates and grows.
9. Invariants #
- Never let the replay window lapse before
steady_state. A hole is permanent; lateness is not. - The watchdog deletes the volume. Not just servers.
- Never conclude "absent" or "new" from your own grep or metric selector. I
claimed
RepairRateLimitedwas new when it was 38h old and merely missing from my pattern; I claimed the watchdog was unverifiable when it was two lookups away under a different unit name. - Verify a flag's semantics in the source before setting it on a live run.
--backfill-max-reposwould have silently disabled cursor checkpointing. - Do not deploy new code mid-run. zat
0.3.23(which fixes the ~40-minute restart loop and gzip) is landed and pinned but not deployed; the running image is29705cb. - A 200 from
listSegmentsis not proof of a working archive. Read data back.
10. When this run ends — carry these forward #
- Deploy zat
0.3.23. Already pinned inbuild.zig.zonand pushed, never deployed. It fixes the ~40-minute restart loop (a pooled connection the relay reaped, surfacing asHttpConnectionClosing), makes gzip actually work on thegetRepohot path (~2.5x less data pulled from other operators' PDSes), and bounds stalled transfers. This run restarted 163 times because of that bug; each restart cost ~5 minutes of bloom-loading startup, roughly 12% of wall clock. - Right-size the box from a measurement, not from the table in §8.
- Re-derive the denominator honestly.
accountCountis an upper bound; seestream-accountcount-not-fetchablein memory. - The observer host was idle all run (load 0.24). If it is not earning its keep next time, do not provision it.
11. Open questions #
- Can it serve? Unproven. This is the whole remaining point.
- How long does compaction take on 1.6 TB? No measurement exists.
- Is the kept/dropped ratio right? Final was 2.02:1 kept-over-dropped, the
inverse of the brief's
dropped >> keptexpectation. The filter is matching (126M dropped ≠ 0). My reasoning — that repos backfilled early accumulate ~145h of correctly-kept live events, and largest-first ordering front-loads exactly the busiest repos — is consistent but unverified against segment timestamps. The failure it could indicate is duplicate events, not loss. - Actual coverage. ~16.26M against a 22.85M
accountCount; fig and calabro both report 18-20M for relay-known repos on a ~41M reachable network. Where we really landed is worth measuring properly, not asserting.