# Experiment 6 — operational handoff Written 2026-08-04 ~01:40Z, mid-`merging`. Everything here was verified against the running system, not recalled. Times are given in UTC with CDT (UTC-5) in parentheses — **the operator is US Central**; earlier logs in this repo that say "EDT" are an hour off and mislabeled. Companion reading, in priority order: - `~/.claude/projects/-Users-nate-tangled-org-zzstoatzz-io-relay/memory/` — the behavioural memories, especially `feedback_monitoring_window_discipline`, `stream-accountcount-not-fetchable`, `stream-backfill-progress-metrics`. - `~/tangled.org/zzstoatzz.io/notes/protocols/atproto/network-backfill.md` — why the crawl behaves the way it does. - `docs/full-network-experiment-2026-07.md` — the running field log. - `docs/gotchas.md` — traps, including the gzip/transfer-buffer one. --- ## 1. What is running, in one paragraph A whole-network atproto backfill built into a Jetstream V2 archive. Bootstrap crawled `bsky.network` for ~145h and completed **16.26M repos** (~1.7 TB, ~20B events). On 2026-08-03 07:04:58Z bootstrap was deliberately cut over to `merging` because the remaining tail was **unfetchable, not slow** (see §7). Merge drained its live-capture archive cleanly. It is now in the **pending-repo pass**, after which come discovery, rotate, **compaction**, and the `steady_state` commit that ungates serving. Nothing else is required of an operator until that happens. **The single question still open: can this archive actually serve?** That is unproven until `listSegments` returns 200 *and* segments read back and decode. --- ## 2. Access SSH key `~/.ssh/waow_ed25519`, user `root`. Always `-o IdentitiesOnly=yes` (the agent offers the wrong key otherwise). | host | IP | role | |---|---|---| | `stream-20260728-0501` | `37.27.43.39` | **workload** — the archive lives here | | `stream-20260728-0501-watchdog` | `65.109.0.222` | kill-switch, holds the Hetzner API token | | `stream-20260728-0501-observer` | `37.27.41.161` | idle this run | Hetzner firewalls: SSH (22) only from `99.64.254.139/32` and `73.50.144.240/32`. **These are dynamic operator IPs and have changed mid-run before** — if SSH hangs, recheck your public IP before concluding the host is sick. Hetzner also recycles IPs between experiments, so a host-key mismatch on a *new* run is expected: `ssh-keygen -R `. Ports 80/443 on the workload are open to `0.0.0.0/0`. Public endpoints (Caddy, automatic TLS, no DNS needed — sslip.io encodes the IP): - `https://stream.37-27-43-39.sslip.io/status` — human-readable field report - `https://dash.37-27-43-39.sslip.io` — the live dashboard (§6) - `https://grafana.37-27-43-39.sslip.io` — password in `/opt/stream-experiment/grafana-password.txt` **There is no `hcloud` token configured locally.** The only Hetzner API token is `/etc/stream-watchdog/token` on the watchdog. It has destroy permissions — use it read-only unless you intend to destroy something. ### Reaching metrics The stream container has **no `curl` or `wget`**. Go through the prometheus container: ```bash # PromQL docker compose exec -T prometheus wget -qO- \ 'http://127.0.0.1:9090/api/v1/query?query=' # raw stream metrics / serving surface (container IP, DNS from busybox is flaky) docker compose exec -T prometheus wget -qO- http://172.18.0.2:6060/metrics docker compose exec -T prometheus wget -S -qO- http://172.18.0.2:8080/status ``` Ports: **6060** = metrics/debug, **8080** = serving. `172.18.0.2` is the stream container; re-read it with `docker inspect` if the container is recreated. ### Local vs remote tooling — do not mix these up **On the operator's machine: never invoke `python3` or `pip`.** Use `uv` / `uvx` / `uv run`, or `jq` for JSON. This is a standing rule in the operator's global instructions. **On the Hetzner hosts, `uv` is simply not installed** — system `python3` is what is there, and `jq` is present only on the watchdog. So the scripts in `deploy/ops/` that run *on the box* (`wsprobe.py`, `dash-snapshot`) target system `python3`. That is a fact about how these hosts were provisioned, **not a constraint**: `uv` is a single static binary (`curl -LsSf https://astral.sh/uv/install.sh | sh`) and installing it on the hosts would let both sides use the same tooling. Worth doing on the next provision; not worth changing under a live run. --- ## 3. Deadlines — the replay window binds before the watchdog Two independent clocks. **The earlier one is not the watchdog.** | clock | when | what happens | |---|---|---| | **firehose replay window** | **2026-08-06 07:04:58Z (Thu 2:04 AM CDT)** | the archive acquires a **permanent hole** | | watchdog | 2026-08-06 17:01:47Z (Thu 12:01 PM CDT) | servers **and the volume** are deleted | **Replay window.** Live capture was torn down at `phase -> merging` (`1785740698`, 2026-08-03 07:04:58Z). Since then **nothing is being captured.** At `steady_state` the consumer resumes from the shared relay cursor and replays forward. `bsky.network` retains **72h** (operator-confirmed). Miss it and events from 2026-08-03 07:04:58Z onward are gone for good — worse than finishing late. Probe whether replay still works (returns `101` + binary commit frames if healthy): ```bash ssh -i ~/.ssh/waow_ed25519 -o IdentitiesOnly=yes root@37.27.43.39 \ '/usr/local/bin/wsprobe 32408688748' ``` `32408688748` is the seq the live capture stopped at. Source is `deploy/ops/wsprobe.py`; installed to `/usr/local/bin/wsprobe` on the workload (it was in `/tmp` initially, which does not survive a reboot). **Watchdog.** `/etc/stream-watchdog/` on `65.109.0.222`: `armed` (flag file), `deadline` (epoch), `token`. `stream-experiment-expire.timer` fires `/usr/local/sbin/stream-experiment-expire` **every 60s**; when `now >= deadline` it deletes every server, then **`drain_kind volumes`**, then itself. > It deletes the **volume**. The archive and the pre-merge snapshot are both on > `/dev/sdb`. When it fires, everything is gone. Current value `1786035707`. Previous value backed up at `/etc/stream-watchdog/deadline.bak-20260803T145236Z`. --- ## 4. Current state and how to read it ```bash bash deploy/ops/check6.sh # from the stream repo; state in ~/.stream-exp6 curl -s https://dash.37-27-43-39.sslip.io/data.json | jq . ``` The ops scripts were originally written into a session scratchpad that does not survive. They now live in **`deploy/ops/`** in this repo — `check6.sh`, `wsprobe.py`, `dash-snapshot`, `dash-index.html`. See `deploy/ops/README.md`. As of 2026-08-04 01:33Z: phase **2 (merging)**, `complete` **16,260,762**, `pending` **25,002** draining ~4,400/hr, `listSegments` **503** (correct), 0 stale `.tmp`, `backfill/` 40 GB awaiting cleanup, restarts 1 (disk-full, 08-04 13:26Z, self-healed), disk **1.3 TB free**, RSS ~56 GB, **€145** spent. Phase encoding: `1` bootstrap, `2` merging, `3` steady_state. ### Metric literacy — these have all burned me - **`phase = NONE` means mid-restart, not a fault.** Every other value is untrustworthy in that moment. Startup takes ~5 min (bloom load); wait it out. - **Per-process counters reset on every restart.** `jetstream_ingest_events_appended_total` sawtooths — it read 16.5M, 14.5M, 6.6M on successive reads. Only `increase(...[24h])` is meaningful. The cumulative truth is the archive's `next seq` in the startup log (~19.97B). - **`stream_backfill_repos_durable{status="complete"}` is the only monotonic bucket.** The bucket *sum* is enumeration progress, not network size. - **Check sample spacing before calling a trend.** I twice read near-simultaneous samples as a time series and nearly escalated. Nested windows (1h/6h/12h) must move *together*, and a rate sampled while `not_started` ≈ 0 is a batch trough. - **RSS ~46-56 GB is normal** — resident bloom filters, not a leak. --- ## 5. Runbooks ### A0. Restarts (updated 2026-08-08) A restart costs ~4 min of archive recovery, then — with compaction enabled, which production runs since 2026-08-08 03:51Z — a synchronous tombstone rebuild before ingest dials. The rebuild logs a start line (with the loaded watermark and segment count) and progress every 200 segments, but it visits every segment file even below the committed watermark: **~3 h on the cx43** until the header-level skip fix lands. Plan restarts accordingly: the tail is deaf for the scan, the archive serves throughout, consumers ride their fallback hosts. The live instance is stream-cx43 (89.167.122.160); the watermark is committed (23,135,908,535 as of enablement) and advances with each compaction pass (4 h cadence from compactor start). ### A. Routine check (nothing wrong) Read phase first. Confirm: `listSegments` 503 while phase < 3, `stale_tmp` 0, disk free » 0.4 TB, oom-kills still **1** (the one on 2026-07-29 06:06). Report two lines. Do not deploy, do not tear down. ### B. Phase flips 2 → 3 (the win) — how to actually prove it A 200 is **not** proof. Do all of this: ```bash # 1. gate opened docker compose exec -T prometheus wget -S -qO- http://172.18.0.2:8080/xrpc/jetstream.listSegments # 2. merge cleanup completed du -sh /data/stream/backfill # should be GONE, not 40 GB find /data/stream -name '*.tmp' | wc -l # 0 # 3. read data back out and decode it — the actual proof # pull a segment via getSegment and confirm it parses as jss # 4. subscribe surface serves live events after replay catches up # 5. capture a receipt: phase, counters, segment count, disk, timestamps ``` Then re-run the cursor probe: if replay was refused, say so plainly — a serving instance with a hole is not a success. ### C. Compaction running long / replay window closing **How to tell it started and is progressing.** There is no "compaction started" log line to wait for. Watch `stream_compaction_watermark_seq` — it is `0` until the pass begins and then advances (`compact/pass.zig:131`). Pair it with `stream_compaction_watermark_lag_us`. CPU stays high throughout; that alone tells you nothing about progress. ```bash docker compose exec -T prometheus wget -qO- \ 'http://127.0.0.1:9090/api/v1/query?query=stream_compaction_watermark_seq' ``` Compaction over 1.6 TB has **never been measured**. If `steady_state` has not committed and the replay window has < ~6h left, that is the decision point: finishing late is recoverable, a hole is not. Surface it immediately rather than letting it lapse quietly. ### D. Process crashes during `merging` Expected to self-heal. `restart: unless-stopped`, and the crash matrix at `lifecycle.zig:13` covers `merging` re-entry at-least-once (seal-guard + per-source cursor). Confirm `recovery: resumed active segment` appears after the restart. Escalate only if recovery stops resuming. ### E. Roll back to the pre-merge snapshot — **GONE as of 2026-08-04 13:4xZ** The snapshot was **deleted** during the disk-full incident (see below). There is no rollback-to-premerge path anymore; the live archive is the only copy. What happened, kept as the cautionary tale: the reflink snapshot cost ~0 space at creation, but merge + the pending pass rewrote enough segments that its unique blocks silently consumed the entire 1.3 TB of headroom. `/data` hit 100% (263 MB free) at 13:26Z on 08-04, the process crashed once on NoSpaceLeft, and recovery resumed cleanly on restart. Deleting the snapshot reclaimed 1.3 TB. **`df` cannot warn you about reflink divergence** — the space is consumed by writes to the *live* tree breaking sharing, and free space slides without any directory growing. If you take a CoW snapshot before a phase that rewrites data, budget for full divergence or monitor free space directly. ### F. Move or remove the watchdog ```bash # extend (preferred — keeps the ceiling, just later) ssh root@65.109.0.222 'cp -a /etc/stream-watchdog/deadline{,.bak-$(date -u +%Y%m%dT%H%M%SZ)} printf "%s\n" > /etc/stream-watchdog/deadline d=$(cat /etc/stream-watchdog/deadline); now=$(date +%s); echo "left_h=$(( (d-now)/3600 ))"' # disarm entirely (removes the cost ceiling — get explicit sign-off) ssh root@65.109.0.222 'rm /etc/stream-watchdog/armed' ``` Always read back through the same path the script uses. Cost of keeping it alive is **€0.84/h (~€20/day)**. ### G. Deliberate teardown Let the watchdog do it — that is what it is for. Set `deadline` to a near-future epoch and confirm the project empties. **Note:** `docs/deployment-runbook.md` cites `assert_empty_project.py`, but that script is **not in this repo** — it lived in provisioning tooling that was never committed. Verify emptiness directly against the Hetzner API with the watchdog's token instead: ```bash ssh root@65.109.0.222 't=$(cat /etc/stream-watchdog/token) for k in servers volumes floating_ips primary_ips; do printf "%s: " $k curl -fsS -H "Authorization: Bearer $t" "https://api.hetzner.cloud/v1/$k?per_page=50" | jq ".$k | length" done' ``` Do **not** hand-delete resources; the script's ordering (servers → volumes → IPs → self) exists because Hetzner refuses to delete an attached volume. ### I. The replay window lapsed (worst case) `wsprobe` returns a non-101 handshake or an `#info`/`OutdatedCursor` frame. This means events from **2026-08-03 07:04:58Z** to whenever ingest resumes are not recoverable from `bsky.network`. Do not quietly continue. State it plainly, then: 1. **Confirm it.** One failed probe is not proof — re-run, and try a cursor a few thousand seqs *later* to distinguish "this cursor aged out" from "the relay is refusing connections." 2. **Let it finish anyway.** A partial archive with a known, dated hole is still a usable artifact and still answers the serving question. Do not tear down. 3. **Record the hole precisely** — first and last missing seq — in the field log and in any receipt. An undocumented gap is far worse than a documented one. 4. **Repair is possible but is a separate project**: re-running backfill for affected repos picks up their current state, since `getRepo` returns the repo as it is now. It does not recover the intermediate *events*. ### H. Disk pressure Floor is **0.4 TB free**. The snapshot lever was **used on 2026-08-04** (see E) — it no longer exists. If disk tightens again the consumers are compaction rewrites and the archive itself, and the remaining lever is growing the volume via the Hetzner API (token on the watchdog) + `xfs_growfs /data`, as was done 2.5→3.0 TB on 07-30. The session monitor now warns below 700 GB free and alarms below the 400 GB floor. --- ## 6. The dashboard `https://dash.37-27-43-39.sslip.io` — free, rides the existing Caddy. - `/usr/local/bin/dash-snapshot` (python) writes `/srv/dash/data.json` - `dash-snapshot.timer` runs it every 60s - `/srv/dash/index.html` fetches `data.json` every 30s, same-origin - Caddy block appended to `/opt/stream-experiment/Caddyfile`; `/srv/dash` mounted read-only into the caddy container **It dies with the host on Thursday.** Nothing outlives the volume. Backups: `Caddyfile.bak-20260803T153044Z`, `compose.yaml.bak-pre-merge-20260803`. --- ## 7. What happened, and the one config change | when (UTC) | event | |---|---| | 07-28 05:01:47 | bootstrap begins — ccx43, 2.5 TB, 100 workers, image `29705cb` | | 07-29 06:06 | one oom-kill (the only one all run) | | 07-30 ~19:00 | deadline 120h→180h; volume 2.5→3.0 TB online via `xfs_growfs` | | 08-03 07:04:58 | **bootstrap cut over to `merging`** (see below) | | 08-03 08:05:54 | merge drain: **255,799,323 kept / 126,488,978 dropped across 137 sources** | | 08-03 14:52 | deadline 180h→228h (`1786035707`) | | 08-04 01:33 | pending pass draining, 25,002 left | ### The cutover, and the flag The crawl fell to ~7.5 repos/s and stayed there. It was **not** degradation: 15 failing DIDs sampled through `plc.directory` showed **10 of 15 on `atproto.brid.gy`**, which `301`s to `bridgy-hubble.microcosm.blue` and answers `{"error":"RepoNotFound","message":"repo not synchronized yet"}`. Following redirects, **4 of 15 fetchable**. Those accounts are counted by `listHosts accountCount` but their repos are not served by anyone — so the 19.4M target was never reachable. Worse, `merging` commits only when the engine returns, which requires `listRepos` enumeration to reach its final page. A slow crawl also slows that approach, so **merge would never have fired** before the deadline. Fix: added `--backfill-max-repos=1` to the compose command and restarted. The engine breaks on `max_reached or final_page` (`engine.zig:849`) — both reach the same commit point. > **Trap, verified by reading the code first:** `max_repos > 0` sets > `full_network = false` (`engine.zig:731`), which **skips loading the durable > listRepos cursor** and **skips `saveListReposCheckpoint`**. It does *not* delete > the cursor (only `selected_repos` does that, line 730). Safe here because we > were done crawling. **Do not set this flag on a run you intend to resume.** To resume crawling later, remove the flag; the cursor at ~16.3M is intact. --- ## 8. Cost Verified against Hetzner's pricing API, not recalled. | resource | €/h | |---|---| | ccx43 workload | 0.528 | | cpx12 × 2 (watchdog, observer) | 0.0288 each | | volume | 0.0767 /GB-month → 3 TB ≈ 0.315/h | **Running: €0.84/h ≈ €20/day. Spent on experiment 6: ~€145** at 154h. Prior exploratory runs (experiments 1-5) were ~$25-30 total — cheap because they were stopped early and deliberately. Steady-state, once right-sized (the ccx43 is sized for 100 crawl workers and is overkill for ingest+serve): | shape | €/month | |---|---| | ccx43 + 3 TB (now) | ~616 | | ccx33 + 2 TB | ~348 | | ccx23 + 2 TB | ~274 | **Right-size from a measurement after cutover, not from this table.** Steady-state CPU has never been observed. Storage dominates and grows. --- ## 9. Invariants 1. **Never let the replay window lapse before `steady_state`.** A hole is permanent; lateness is not. 2. **The watchdog deletes the volume.** Not just servers. 3. **Never conclude "absent" or "new" from your own grep or metric selector.** I claimed `RepairRateLimited` was new when it was 38h old and merely missing from my pattern; I claimed the watchdog was unverifiable when it was two lookups away under a different unit name. 4. **Verify a flag's semantics in the source before setting it on a live run.** `--backfill-max-repos` would have silently disabled cursor checkpointing. 5. **Do not deploy new code mid-run.** zat `0.3.23` (which fixes the ~40-minute restart loop and gzip) is landed and pinned but **not deployed**; the running image is `29705cb`. 6. **A 200 from `listSegments` is not proof of a working archive.** Read data back. --- ## 10. When this run ends — carry these forward - **Deploy zat `0.3.23`.** Already pinned in `build.zig.zon` and pushed, never deployed. It fixes the ~40-minute restart loop (a pooled connection the relay reaped, surfacing as `HttpConnectionClosing`), makes gzip actually work on the `getRepo` hot path (~2.5x less data pulled from other operators' PDSes), and bounds stalled transfers. This run restarted **163 times** because of that bug; each restart cost ~5 minutes of bloom-loading startup, roughly 12% of wall clock. - **Right-size the box from a measurement**, not from the table in §8. - **Re-derive the denominator honestly.** `accountCount` is an upper bound; see `stream-accountcount-not-fetchable` in memory. - **The observer host was idle all run** (load 0.24). If it is not earning its keep next time, do not provision it. --- ## 11. Open questions - **Can it serve?** Unproven. This is the whole remaining point. - **How long does compaction take on 1.6 TB?** No measurement exists. - **Is the kept/dropped ratio right?** Final was 2.02:1 kept-over-dropped, the inverse of the brief's `dropped >> kept` expectation. The filter *is* matching (126M dropped ≠ 0). My reasoning — that repos backfilled early accumulate ~145h of correctly-kept live events, and largest-first ordering front-loads exactly the busiest repos — is consistent but **unverified against segment timestamps**. The failure it could indicate is duplicate events, not loss. - **Actual coverage.** ~16.26M against a 22.85M `accountCount`; fig and calabro both report 18-20M for relay-known repos on a ~41M reachable network. Where we really landed is worth measuring properly, not asserting.