jetstream v2 in zig stream.waow.tech
stream docs backup-plan.md
12 kB
Markdown

archive backup and restore — plan #

Status: shelved 2026-08-28, nothing implemented. Kept as a record of the reasoning. The write-once upload idea is sound, but a cold copy does not work in perpetuity: restore is O(archive) (download + full tombstone fold + compactor rewrite) and the archive grows without bound, while the upstream relay replay window is fixed. At some size the restore cannot complete inside the window and the copy can no longer yield a contiguous archive. The un-compacted copy also retains deleted rows indefinitely. A durable answer has to change an axiom — bounded retention, non-O(archive) compaction, or meta snapshots — which is the same open problem as prod's own compaction cost (compaction-cost.md) and upstream's dropped §6 Replication.

Original text follows unchanged.

Originally: plan, nothing implemented. Written 2026-08-23 from code reading, box measurements, and the Hetzner API. Numbers are from that day; re-measure before acting on them.

why #

The archive is one copy: stream-20260728-0501-archive, 3 TB, hel1, attached to stream-cx43, protection.delete = true (verified via hcloud). Hetzner Cloud has no volume snapshots (deployment-runbook.md, "Stage 4b"), server snapshots do not include volumes, and hcloud image list for the project is empty. Delete-protection guards against API deletion — an operator, the watchdog, destroy.sh — and nothing else: not Hetzner losing the volume, not a filesystem fault, not a bad rewrite.

The archive cannot be rebuilt from the network. A re-bootstrap assigns new seqs (every consumer cursor breaks) and cannot recover records or accounts deleted since July 2026. Upstream offers no answer here either: docs/README.md §6 Replication is "DROPPED — to be redesigned", the 2026-06-27 HA note is "exploration only", and jetstream "today runs single-node".

what is on the volume #

/data/stream (1.7 TB used of 3 TB):

path role
segments/seg_<10 chars base-36>.jss authoritative event bytes. 6,750 sealed + 1 active, ~256 MiB each, contiguous indices from 0. The highest index is the active (unsealed) file, under its final name.
meta.rocksdb/ authoritative for what segments cannot give back: relay cursor, lifecycle phase, compaction watermark, per-DID backfill status, host aggregates.
import-rules/ timestamp-import rules (only matters if imports are used).
repair-scratch/, timestamp-import-jobs/, imports/, *.jss.tmp scratch; reset or reclaimed at boot.

Tombstones, blooms, the manifest, and the block cache are not files; they are rebuilt from segment headers/footers at Archive.init.

what a restore is (traced in main.zig, archive.zig:329 recover()) #

Boot a process against segments/ with no meta.rocksdb:

  • recover() scans the directory, truncates a torn tail, seals/resumes the last file, derives next_seq from sealed headers. Manifest rebuilt. Serving opens.
  • Missing phase reads as steady (main.zig:692): no re-bootstrap.
  • Missing relay/cursor → the live consumer dials upstream with no cursor, i.e. from the head. Everything between the last archived event and now is lost for good. This is the real damage of losing meta.
  • Missing compaction/seq → watermark 0 → full-archive tombstone fold before ingest dials (the multi-hour restart class from deploying.md).
  • Per-DID status, host aggregates, counts: gone; status page shows zeros; failed-repo healing starts with an empty roster.

The escape hatch already in the code: meta_store.Store.migrateLegacy (meta_store.zig:122-172) imports flat files on first boot when the RocksDB key is absent:

  • <data-dir>/cursor — decimal text, the relay cursor (what /status shows as upstream seq).
  • <data-dir>/compaction_watermark — 9 bytes: 0x01 then u64 LE seq.
  • <data-dir>/phase — text, e.g. steady_state.

So a usable backup is segments + a tiny sidecar (cursor, watermark, phase), not a RocksDB dump. Per-DID status is not recoverable and we accept that.

the sync problem, and the reframe #

Measured 2026-08-23 on the box: 6,700 of 6,747 segment files carried an mtime within the last 24 h. Compaction metrics since the last restart (48 h): 5 passes, 30,638 of 33,669 examined segments rewritten, 8.1 TB rewritten — about 1.6 TB per pass, via <name>.tmp → rename over the same filename (segment_rewrite.zig:195; docs/compaction-cost.md explains why it is O(archive)). A naive rclone sync on size+mtime would push on the order of 100 TB/month out of a box with a 22 TB traffic allowance (€1.20/TB overage; 2.1 TB used this month so far).

The reframe that makes this cheap: compaction only drops rows. A segment's pre-compaction bytes are a strict superset of every later version and remain valid under the at-least-once contract (upstream docs/README.md:74: delete markers are positive events retained forever; create rows for already-dead records are allowed). Therefore:

Upload each segment once, when it seals, and never again.

Genuine growth is what matters: 25 seals in 48.1 h of uptime ≈ 12.5 segments/day ≈ 3.4 GB/day ≈ 100 GB/month of egress. The initial 1.7 TB upload happens once, inside the allowance.

Cost of the remote copy (Cloudflare R2): storage at $0.015/GB-month ≈ $26/month for 1.7 TB, growing ~$1.50/month per month; ingress free; egress free (a restore pulls 1.7 TB back at no R2 cost; Hetzner ingress is free). Class A write operations at ~13/day are negligible. Hetzner Object Storage or B2 are alternatives in the same band; R2's zero egress is what makes the dress rehearsal (below) repeatable.

Consequence for restore: the restored archive is un-compacted relative to prod. Two ways to boot it:

  1. Correct, slow — omit compaction_watermark from the sidecar (or write 0). Startup re-folds the whole archive, then the steady compactor rewrites it. Deleted rows disappear again. This is the default for a real restore.
  2. Fast, lossy-in-the-other-direction — write the watermark as next_seq − 1. No fold, instant ingest; rows deleted before the outage are resurrected and stay until some later tombstone happens to rewrite their segment. Acceptable only for a rehearsal or a read-only stand-in.

Write-once has one blind spot: timestamp-import patches (segment_patch.zig) and the optional rebloom sweep also rewrite old segments in place. The backup keeps the pre-patch bytes. Import runs are rare and operator-driven; the plan is to re-upload the patched set after each import (the patch job knows which segments it touched) and to run a periodic checksum reconciliation that flags any local/remote divergence instead of silently skipping it.

plan #

phase 0 — decide and prepare #

  • Create an R2 bucket (stream-archive or similar) in the existing Cloudflare account; an API token scoped to that bucket, write-only for the box, kept in the sops store with the other stream credentials.
  • Install rclone on the box; R2 remote config from the token.
  • Extend COSTS.md with the new line before it appears on the cost dashboard, and add a stream attribution for the bucket in my-prefect-server's project mapping.

phase 1 — initial upload #

  • rclone copy /data/stream/segments r2:<bucket>/segments --ignore-existing --exclude <active-segment-name> --transfers 4 --bwlimit <modest> — run in tmux, bandwidth-limited so serving is unaffected; watch jetstream_* latency and the box's network graph while it runs.
  • Exclude rule must be dynamic: the active segment is the max index; anything below it is sealed. Confirm each candidate's header reports sealed (stream inspect-segment) — cheap, and the only defence against uploading a file mid-seal.
  • Upload the first sidecar: cursor (from /status "upstream seq"), compaction_watermark (current watermark from jetstream_compaction_watermark_seq), phase.
  • Verify: object count == sealed count, total bytes match, spot-check 20 segments by sha256 against R2's stored hash.

phase 2 — incremental sync #

  • A systemd timer on the box (hourly; a segment seals roughly every two hours at current live volume) that: lists local sealed segments not yet in R2 (an rclone lsf diff, or --ignore-existing copy), uploads them, then rewrites the sidecar objects.
  • The sidecar being slightly stale is fine: restoring a cursor a few minutes old just replays a few minutes at-least-once. A sidecar newer than the last uploaded segment is the dangerous direction (cursor past the durable archive → gap). Order: segments first, sidecar second, always.
  • Alert rule: stream_backup_last_success_timestamp_seconds older than a few hours → warning (textfile collector or a pushgateway-free equivalent; the rules file is deploy/prometheus/rules.yml).
  • Reconciliation (daily or weekly): rclone check --size-only is not enough (footer-only rewrites keep length); compare a local xxhash/sha256 manifest of headers+sizes against the R2 listing and log divergences. This is what catches a missed upload or an import patch.

phase 3 — prove restore (the part that earns the word) #

  1. Slice rehearsal (cheap, repeatable). On heavypad or a throwaway VM: pull a contiguous head slice (segments 0..N) into a fresh data dir, write the sidecar, boot with no meta.rocksdb. Assert: process reaches serving, listSegments covers the slice, getSegment bytes match prod for the same indices, upstream Go client replays from seq 0. Confirm first whether recover() tolerates a non-contiguous set; if not, head slices only.
  2. Full dress rehearsal (once, then on format changes). Temporary Hetzner server + 2 TB volume in hel1, rclone copy from R2 (free egress), sidecar, boot in mode 1 (watermark 0). Measure: copy time, fold time, time to serving. Diff listSegments against prod over the shared range. Point a consumer (the zat.dev/jetstream archive examples, e.g. streamplace_chat.zig) at it. Destroy everything afterwards; assert_empty_project analogue. Record the receipt in this file under "rehearsals".
  3. Write the operator runbook section in deploying.md: the exact commands, which sidecar mode to use when, and the expected durations from the rehearsal.

phase 4 — close the loop #

  • Reconcile docs/invariants.md:16 ("never edit a sealed file in place"): true of the inode, misleading about the name; say "rename over the same name" explicitly since the backup design depends on it.
  • Consider making the process write the three legacy flat files itself on each durable cursor/watermark commit (they are already read by migrateLegacy), so the sidecar is a byproduct rather than a scrape. Small change; decide after the first rehearsal shows whether the scrape is good enough.

open questions #

  • Does recover() accept a segment directory with gaps? (Determines whether slice rehearsals can use tail slices.)
  • Exact R2 pricing tier for Class A/B ops at this volume — negligible by estimate; confirm on the first bill.
  • Whether to also copy import-rules/ (tiny; probably yes, as a RocksDB checkpoint or by re-importing the rule CSVs).
  • Whether the compaction design change being discussed in compaction-cost.md (bounded rewrites) changes any of the above. It doesn't change write-once; it only shrinks the gap between backup and prod.

rehearsals #

none yet.