diff --git a/docs/experiment-6-handoff.md b/docs/experiment-6-handoff.md index 6829371..cca6b0e 100644 --- a/docs/experiment-6-handoff.md +++ b/docs/experiment-6-handoff.md @@ -173,16 +173,18 @@ Phase encoding: `1` bootstrap, `2` merging, `3` steady_state. ## 5. Runbooks -### A0. Restarts (updated 2026-08-07) - -A restart costs ~5 min of archive recovery before listeners bind, then the -firehose dial. With delete compaction ENABLED and a compaction watermark of 0, -startup additionally runs a synchronous full-archive tombstone rebuild -(`compactor.rebuild()`, before ingest spawns, no log output until done) — at -current archive size that is hours of deaf-but-serving. Production therefore -runs `--compaction-interval=0` (set in compose.yaml on the box, 2026-08-07): -rebuild skipped, deletes accumulate uncompacted. Do not re-enable compaction -until the rebuild is async or a watermark has been committed. +### A0. Restarts (updated 2026-08-08) + +A restart costs ~4 min of archive recovery, then — with compaction enabled, +which production runs since 2026-08-08 03:51Z — a synchronous tombstone +rebuild before ingest dials. The rebuild logs a start line (with the loaded +watermark and segment count) and progress every 200 segments, but it visits +every segment file even below the committed watermark: **~3 h on the cx43** +until the header-level skip fix lands. Plan restarts accordingly: the tail +is deaf for the scan, the archive serves throughout, consumers ride their +fallback hosts. The live instance is stream-cx43 (89.167.122.160); the +watermark is committed (23,135,908,535 as of enablement) and advances with +each compaction pass (4 h cadence from compactor start). ### A. Routine check (nothing wrong)