# archive backup and restore — plan Status: **shelved 2026-08-28, nothing implemented.** Kept as a record of the reasoning. The write-once upload idea is sound, but a cold copy does not work in perpetuity: restore is O(archive) (download + full tombstone fold + compactor rewrite) and the archive grows without bound, while the upstream relay replay window is fixed. At some size the restore cannot complete inside the window and the copy can no longer yield a contiguous archive. The un-compacted copy also retains deleted rows indefinitely. A durable answer has to change an axiom — bounded retention, non-O(archive) compaction, or meta snapshots — which is the same open problem as prod's own compaction cost (`compaction-cost.md`) and upstream's dropped §6 Replication. Original text follows unchanged. Originally: **plan, nothing implemented.** Written 2026-08-23 from code reading, box measurements, and the Hetzner API. Numbers are from that day; re-measure before acting on them. ## why The archive is one copy: `stream-20260728-0501-archive`, 3 TB, hel1, attached to `stream-cx43`, `protection.delete = true` (verified via `hcloud`). Hetzner Cloud has no volume snapshots (`deployment-runbook.md`, "Stage 4b"), server snapshots do not include volumes, and `hcloud image list` for the project is empty. Delete-protection guards against API deletion — an operator, the watchdog, `destroy.sh` — and nothing else: not Hetzner losing the volume, not a filesystem fault, not a bad rewrite. The archive cannot be rebuilt from the network. A re-bootstrap assigns new seqs (every consumer cursor breaks) and cannot recover records or accounts deleted since July 2026. Upstream offers no answer here either: `docs/README.md` §6 Replication is "DROPPED — to be redesigned", the 2026-06-27 HA note is "exploration only", and jetstream "today runs single-node". ## what is on the volume `/data/stream` (1.7 TB used of 3 TB): | path | role | | --- | --- | | `segments/seg_<10 chars base-36>.jss` | **authoritative** event bytes. 6,750 sealed + 1 active, ~256 MiB each, contiguous indices from 0. The highest index is the active (unsealed) file, under its final name. | | `meta.rocksdb/` | **authoritative** for what segments cannot give back: relay cursor, lifecycle phase, compaction watermark, per-DID backfill status, host aggregates. | | `import-rules/` | timestamp-import rules (only matters if imports are used). | | `repair-scratch/`, `timestamp-import-jobs/`, `imports/`, `*.jss.tmp` | scratch; reset or reclaimed at boot. | Tombstones, blooms, the manifest, and the block cache are not files; they are rebuilt from segment headers/footers at `Archive.init`. ## what a restore is (traced in `main.zig`, `archive.zig:329 recover()`) Boot a process against `segments/` with **no** `meta.rocksdb`: - `recover()` scans the directory, truncates a torn tail, seals/resumes the last file, derives `next_seq` from sealed headers. Manifest rebuilt. Serving opens. - Missing `phase` reads as steady (`main.zig:692`): **no re-bootstrap.** - Missing `relay/cursor` → the live consumer dials upstream with no cursor, i.e. from the head. **Everything between the last archived event and now is lost for good.** This is the real damage of losing meta. - Missing `compaction/seq` → watermark 0 → full-archive tombstone fold before ingest dials (the multi-hour restart class from `deploying.md`). - Per-DID status, host aggregates, counts: gone; status page shows zeros; failed-repo healing starts with an empty roster. The escape hatch already in the code: `meta_store.Store.migrateLegacy` (`meta_store.zig:122-172`) imports flat files on first boot when the RocksDB key is absent: - `/cursor` — decimal text, the relay cursor (what `/status` shows as `upstream seq`). - `/compaction_watermark` — 9 bytes: `0x01` then u64 LE seq. - `/phase` — text, e.g. `steady_state`. So a usable backup is **segments + a tiny sidecar** (cursor, watermark, phase), not a RocksDB dump. Per-DID status is not recoverable and we accept that. ## the sync problem, and the reframe Measured 2026-08-23 on the box: 6,700 of 6,747 segment files carried an mtime within the last 24 h. Compaction metrics since the last restart (48 h): 5 passes, 30,638 of 33,669 examined segments rewritten, 8.1 TB rewritten — about 1.6 TB per pass, via `.tmp` → rename over the **same filename** (`segment_rewrite.zig:195`; `docs/compaction-cost.md` explains why it is O(archive)). A naive `rclone sync` on size+mtime would push on the order of 100 TB/month out of a box with a 22 TB traffic allowance (€1.20/TB overage; 2.1 TB used this month so far). The reframe that makes this cheap: **compaction only drops rows.** A segment's pre-compaction bytes are a strict superset of every later version and remain valid under the at-least-once contract (upstream `docs/README.md:74`: delete markers are positive events retained forever; create rows for already-dead records are allowed). Therefore: > Upload each segment **once, when it seals, and never again.** Genuine growth is what matters: 25 seals in 48.1 h of uptime ≈ 12.5 segments/day ≈ **3.4 GB/day ≈ 100 GB/month** of egress. The initial 1.7 TB upload happens once, inside the allowance. Cost of the remote copy (Cloudflare R2): storage at $0.015/GB-month ≈ $26/month for 1.7 TB, growing ~$1.50/month per month; ingress free; egress free (a restore pulls 1.7 TB back at no R2 cost; Hetzner ingress is free). Class A write operations at ~13/day are negligible. Hetzner Object Storage or B2 are alternatives in the same band; R2's zero egress is what makes the dress rehearsal (below) repeatable. Consequence for restore: the restored archive is *un-compacted* relative to prod. Two ways to boot it: 1. **Correct, slow** — omit `compaction_watermark` from the sidecar (or write 0). Startup re-folds the whole archive, then the steady compactor rewrites it. Deleted rows disappear again. This is the default for a real restore. 2. **Fast, lossy-in-the-other-direction** — write the watermark as `next_seq − 1`. No fold, instant ingest; rows deleted before the outage are resurrected and stay until some later tombstone happens to rewrite their segment. Acceptable only for a rehearsal or a read-only stand-in. Write-once has one blind spot: **timestamp-import patches** (`segment_patch.zig`) and the optional rebloom sweep also rewrite old segments in place. The backup keeps the pre-patch bytes. Import runs are rare and operator-driven; the plan is to re-upload the patched set after each import (the patch job knows which segments it touched) and to run a periodic checksum reconciliation that flags any local/remote divergence instead of silently skipping it. ## plan ### phase 0 — decide and prepare - Create an R2 bucket (`stream-archive` or similar) in the existing Cloudflare account; an API token scoped to that bucket, write-only for the box, kept in the sops store with the other stream credentials. - Install `rclone` on the box; R2 remote config from the token. - Extend `COSTS.md` with the new line before it appears on the cost dashboard, and add a `stream` attribution for the bucket in my-prefect-server's project mapping. ### phase 1 — initial upload - `rclone copy /data/stream/segments r2:/segments --ignore-existing --exclude --transfers 4 --bwlimit ` — run in tmux, bandwidth-limited so serving is unaffected; watch `jetstream_*` latency and the box's network graph while it runs. - Exclude rule must be dynamic: the active segment is the max index; anything below it is sealed. Confirm each candidate's header reports sealed (`stream inspect-segment`) — cheap, and the only defence against uploading a file mid-seal. - Upload the first sidecar: `cursor` (from `/status` "upstream seq"), `compaction_watermark` (current watermark from `jetstream_compaction_watermark_seq`), `phase`. - Verify: object count == sealed count, total bytes match, spot-check 20 segments by sha256 against R2's stored hash. ### phase 2 — incremental sync - A systemd timer on the box (hourly; a segment seals roughly every two hours at current live volume) that: lists local sealed segments not yet in R2 (an `rclone lsf` diff, or `--ignore-existing` copy), uploads them, then rewrites the sidecar objects. - The sidecar being slightly stale is fine: restoring a cursor a few minutes old just replays a few minutes at-least-once. A sidecar *newer* than the last uploaded segment is the dangerous direction (cursor past the durable archive → gap). Order: segments first, sidecar second, always. - Alert rule: `stream_backup_last_success_timestamp_seconds` older than a few hours → warning (textfile collector or a pushgateway-free equivalent; the rules file is `deploy/prometheus/rules.yml`). - Reconciliation (daily or weekly): `rclone check --size-only` is *not* enough (footer-only rewrites keep length); compare a local xxhash/sha256 manifest of headers+sizes against the R2 listing and log divergences. This is what catches a missed upload or an import patch. ### phase 3 — prove restore (the part that earns the word) 1. **Slice rehearsal (cheap, repeatable).** On heavypad or a throwaway VM: pull a contiguous head slice (segments 0..N) into a fresh data dir, write the sidecar, boot with no `meta.rocksdb`. Assert: process reaches serving, `listSegments` covers the slice, `getSegment` bytes match prod for the same indices, upstream Go client replays from seq 0. Confirm first whether `recover()` tolerates a non-contiguous set; if not, head slices only. 2. **Full dress rehearsal (once, then on format changes).** Temporary Hetzner server + 2 TB volume in hel1, `rclone copy` from R2 (free egress), sidecar, boot in mode 1 (watermark 0). Measure: copy time, fold time, time to serving. Diff `listSegments` against prod over the shared range. Point a consumer (the zat.dev/jetstream archive examples, e.g. `streamplace_chat.zig`) at it. Destroy everything afterwards; `assert_empty_project` analogue. Record the receipt in this file under "rehearsals". 3. Write the operator runbook section in `deploying.md`: the exact commands, which sidecar mode to use when, and the expected durations from the rehearsal. ### phase 4 — close the loop - Reconcile `docs/invariants.md:16` ("never edit a sealed file in place"): true of the inode, misleading about the name; say "rename over the same name" explicitly since the backup design depends on it. - Consider making the process write the three legacy flat files itself on each durable cursor/watermark commit (they are already read by `migrateLegacy`), so the sidecar is a byproduct rather than a scrape. Small change; decide after the first rehearsal shows whether the scrape is good enough. ## open questions - Does `recover()` accept a segment directory with gaps? (Determines whether slice rehearsals can use tail slices.) - Exact R2 pricing tier for Class A/B ops at this volume — negligible by estimate; confirm on the first bill. - Whether to also copy `import-rules/` (tiny; probably yes, as a RocksDB checkpoint or by re-importing the rule CSVs). - Whether the compaction design change being discussed in `compaction-cost.md` (bounded rewrites) changes any of the above. It doesn't change write-once; it only shrinks the gap between backup and prod. ## rehearsals _none yet._