archive backup and restore — plan #
Status: shelved 2026-08-28, nothing implemented. Kept as a record of the
reasoning. The write-once upload idea is sound, but a cold copy does not work
in perpetuity: restore is O(archive) (download + full tombstone fold +
compactor rewrite) and the archive grows without bound, while the upstream
relay replay window is fixed. At some size the restore cannot complete inside
the window and the copy can no longer yield a contiguous archive. The
un-compacted copy also retains deleted rows indefinitely. A durable answer has
to change an axiom — bounded retention, non-O(archive) compaction, or meta
snapshots — which is the same open problem as prod's own compaction cost
(compaction-cost.md) and upstream's dropped §6 Replication.
Original text follows unchanged.
Originally: plan, nothing implemented. Written 2026-08-23 from code reading, box measurements, and the Hetzner API. Numbers are from that day; re-measure before acting on them.
why #
The archive is one copy: stream-20260728-0501-archive, 3 TB, hel1, attached
to stream-cx43, protection.delete = true (verified via hcloud). Hetzner
Cloud has no volume snapshots (deployment-runbook.md, "Stage 4b"), server
snapshots do not include volumes, and hcloud image list for the project is
empty. Delete-protection guards against API deletion — an operator, the
watchdog, destroy.sh — and nothing else: not Hetzner losing the volume, not a
filesystem fault, not a bad rewrite.
The archive cannot be rebuilt from the network. A re-bootstrap assigns new seqs
(every consumer cursor breaks) and cannot recover records or accounts deleted
since July 2026. Upstream offers no answer here either: docs/README.md §6
Replication is "DROPPED — to be redesigned", the 2026-06-27 HA note is
"exploration only", and jetstream "today runs single-node".
what is on the volume #
/data/stream (1.7 TB used of 3 TB):
| path | role |
|---|---|
segments/seg_<10 chars base-36>.jss |
authoritative event bytes. 6,750 sealed + 1 active, ~256 MiB each, contiguous indices from 0. The highest index is the active (unsealed) file, under its final name. |
meta.rocksdb/ |
authoritative for what segments cannot give back: relay cursor, lifecycle phase, compaction watermark, per-DID backfill status, host aggregates. |
import-rules/ |
timestamp-import rules (only matters if imports are used). |
repair-scratch/, timestamp-import-jobs/, imports/, *.jss.tmp |
scratch; reset or reclaimed at boot. |
Tombstones, blooms, the manifest, and the block cache are not files; they are
rebuilt from segment headers/footers at Archive.init.
what a restore is (traced in main.zig, archive.zig:329 recover()) #
Boot a process against segments/ with no meta.rocksdb:
recover()scans the directory, truncates a torn tail, seals/resumes the last file, derivesnext_seqfrom sealed headers. Manifest rebuilt. Serving opens.- Missing
phasereads as steady (main.zig:692): no re-bootstrap. - Missing
relay/cursor→ the live consumer dials upstream with no cursor, i.e. from the head. Everything between the last archived event and now is lost for good. This is the real damage of losing meta. - Missing
compaction/seq→ watermark 0 → full-archive tombstone fold before ingest dials (the multi-hour restart class fromdeploying.md). - Per-DID status, host aggregates, counts: gone; status page shows zeros; failed-repo healing starts with an empty roster.
The escape hatch already in the code: meta_store.Store.migrateLegacy
(meta_store.zig:122-172) imports flat files on first boot when the RocksDB
key is absent:
<data-dir>/cursor— decimal text, the relay cursor (what/statusshows asupstream seq).<data-dir>/compaction_watermark— 9 bytes:0x01then u64 LE seq.<data-dir>/phase— text, e.g.steady_state.
So a usable backup is segments + a tiny sidecar (cursor, watermark, phase), not a RocksDB dump. Per-DID status is not recoverable and we accept that.
the sync problem, and the reframe #
Measured 2026-08-23 on the box: 6,700 of 6,747 segment files carried an mtime
within the last 24 h. Compaction metrics since the last restart (48 h): 5
passes, 30,638 of 33,669 examined segments rewritten, 8.1 TB rewritten — about
1.6 TB per pass, via <name>.tmp → rename over the same filename
(segment_rewrite.zig:195; docs/compaction-cost.md explains why it is
O(archive)). A naive rclone sync on size+mtime would push on the order of
100 TB/month out of a box with a 22 TB traffic allowance (€1.20/TB overage;
2.1 TB used this month so far).
The reframe that makes this cheap: compaction only drops rows. A segment's
pre-compaction bytes are a strict superset of every later version and remain
valid under the at-least-once contract (upstream docs/README.md:74: delete
markers are positive events retained forever; create rows for already-dead
records are allowed). Therefore:
Upload each segment once, when it seals, and never again.
Genuine growth is what matters: 25 seals in 48.1 h of uptime ≈ 12.5 segments/day ≈ 3.4 GB/day ≈ 100 GB/month of egress. The initial 1.7 TB upload happens once, inside the allowance.
Cost of the remote copy (Cloudflare R2): storage at $0.015/GB-month ≈ $26/month for 1.7 TB, growing ~$1.50/month per month; ingress free; egress free (a restore pulls 1.7 TB back at no R2 cost; Hetzner ingress is free). Class A write operations at ~13/day are negligible. Hetzner Object Storage or B2 are alternatives in the same band; R2's zero egress is what makes the dress rehearsal (below) repeatable.
Consequence for restore: the restored archive is un-compacted relative to prod. Two ways to boot it:
- Correct, slow — omit
compaction_watermarkfrom the sidecar (or write 0). Startup re-folds the whole archive, then the steady compactor rewrites it. Deleted rows disappear again. This is the default for a real restore. - Fast, lossy-in-the-other-direction — write the watermark as
next_seq − 1. No fold, instant ingest; rows deleted before the outage are resurrected and stay until some later tombstone happens to rewrite their segment. Acceptable only for a rehearsal or a read-only stand-in.
Write-once has one blind spot: timestamp-import patches (segment_patch.zig)
and the optional rebloom sweep also rewrite old segments in place. The backup
keeps the pre-patch bytes. Import runs are rare and operator-driven; the plan
is to re-upload the patched set after each import (the patch job knows which
segments it touched) and to run a periodic checksum reconciliation that flags
any local/remote divergence instead of silently skipping it.
plan #
phase 0 — decide and prepare #
- Create an R2 bucket (
stream-archiveor similar) in the existing Cloudflare account; an API token scoped to that bucket, write-only for the box, kept in the sops store with the other stream credentials. - Install
rcloneon the box; R2 remote config from the token. - Extend
COSTS.mdwith the new line before it appears on the cost dashboard, and add astreamattribution for the bucket in my-prefect-server's project mapping.
phase 1 — initial upload #
rclone copy /data/stream/segments r2:<bucket>/segments --ignore-existing --exclude <active-segment-name> --transfers 4 --bwlimit <modest>— run in tmux, bandwidth-limited so serving is unaffected; watchjetstream_*latency and the box's network graph while it runs.- Exclude rule must be dynamic: the active segment is the max index; anything
below it is sealed. Confirm each candidate's header reports sealed
(
stream inspect-segment) — cheap, and the only defence against uploading a file mid-seal. - Upload the first sidecar:
cursor(from/status"upstream seq"),compaction_watermark(current watermark fromjetstream_compaction_watermark_seq),phase. - Verify: object count == sealed count, total bytes match, spot-check 20 segments by sha256 against R2's stored hash.
phase 2 — incremental sync #
- A systemd timer on the box (hourly; a segment seals roughly every two hours at current live volume)
that: lists local sealed segments not yet in R2 (an
rclone lsfdiff, or--ignore-existingcopy), uploads them, then rewrites the sidecar objects. - The sidecar being slightly stale is fine: restoring a cursor a few minutes old just replays a few minutes at-least-once. A sidecar newer than the last uploaded segment is the dangerous direction (cursor past the durable archive → gap). Order: segments first, sidecar second, always.
- Alert rule:
stream_backup_last_success_timestamp_secondsolder than a few hours → warning (textfile collector or a pushgateway-free equivalent; the rules file isdeploy/prometheus/rules.yml). - Reconciliation (daily or weekly):
rclone check --size-onlyis not enough (footer-only rewrites keep length); compare a local xxhash/sha256 manifest of headers+sizes against the R2 listing and log divergences. This is what catches a missed upload or an import patch.
phase 3 — prove restore (the part that earns the word) #
- Slice rehearsal (cheap, repeatable). On heavypad or a throwaway VM:
pull a contiguous head slice (segments 0..N) into a fresh data dir, write
the sidecar, boot with no
meta.rocksdb. Assert: process reaches serving,listSegmentscovers the slice,getSegmentbytes match prod for the same indices, upstream Go client replays from seq 0. Confirm first whetherrecover()tolerates a non-contiguous set; if not, head slices only. - Full dress rehearsal (once, then on format changes). Temporary Hetzner
server + 2 TB volume in hel1,
rclone copyfrom R2 (free egress), sidecar, boot in mode 1 (watermark 0). Measure: copy time, fold time, time to serving. DifflistSegmentsagainst prod over the shared range. Point a consumer (the zat.dev/jetstream archive examples, e.g.streamplace_chat.zig) at it. Destroy everything afterwards;assert_empty_projectanalogue. Record the receipt in this file under "rehearsals". - Write the operator runbook section in
deploying.md: the exact commands, which sidecar mode to use when, and the expected durations from the rehearsal.
phase 4 — close the loop #
- Reconcile
docs/invariants.md:16("never edit a sealed file in place"): true of the inode, misleading about the name; say "rename over the same name" explicitly since the backup design depends on it. - Consider making the process write the three legacy flat files itself on each
durable cursor/watermark commit (they are already read by
migrateLegacy), so the sidecar is a byproduct rather than a scrape. Small change; decide after the first rehearsal shows whether the scrape is good enough.
open questions #
- Does
recover()accept a segment directory with gaps? (Determines whether slice rehearsals can use tail slices.) - Exact R2 pricing tier for Class A/B ops at this volume — negligible by estimate; confirm on the first bill.
- Whether to also copy
import-rules/(tiny; probably yes, as a RocksDB checkpoint or by re-importing the rule CSVs). - Whether the compaction design change being discussed in
compaction-cost.md(bounded rewrites) changes any of the above. It doesn't change write-once; it only shrinks the gap between backup and prod.
rehearsals #
none yet.