diff --git a/plan/harness.md b/plan/harness.md index d7816ba..5279c79 100644 --- a/plan/harness.md +++ b/plan/harness.md @@ -263,6 +263,45 @@ Neither of those makes it deliberate. - [x] `CORPUS_GLOBS` carries the compressed suffix, so `corpus save` keeps it - [ ] Compress a corpus when the run that produced it finishes, rather than by hand when the disk fills + +### "When the disk fills" happened, 2026-08-29 + +The item above predicted this exactly, and it went unclosed long enough to fire. + +A 240-match run reached **36 GB at 198 matches** - about 180 MB a match - with +the disk at **96% and 17 GB free**. `corpus save` copies, so the store would +have needed another 44 GB. The bench would have completed all 240, the save +would have failed on space, and a four-hour corpus would have been left in a +worktree's `runs/`, which is the state that has already destroyed two. + +**Nothing regressed, and that is the point.** Compression was never the writer's +job - this section already says so: *"Nothing writes it - `gzip` does, after a +match is finished"*. The corpus recorded three days earlier is `.gz` because +somebody ran `gzip` by hand. So there is no change to revert and nobody to +attribute it to: the cost has always been there, and it is paid by whoever +happens to record the next corpus. + +Compressing in flight, skipping any log held by a running container, took the +run directory from 36 GB to 2.5 GB and the disk from 17 GB to 50 GB free. Every +reader already accepts the compressed form, so nothing else changed. + +**Why nothing saw it coming.** Every metric being watched was healthy - match +count, timeouts, defaults, load, rate. It was found in a *readiness* check +rather than a progress check, and the difference generalises: progress checks +answer "is this going well", readiness checks answer "will the next step +succeed". The second is rarely asked until the next step has already failed, +and here the next step was the only one with no fallback, at the end, after all +the expensive work was done. + +- [ ] **Size the run before it starts, not at the save.** `sds bench` knows the + match count and can measure a log; a corpus at 180 MB a match against + available space is one arithmetic check at the point where it is cheap. + Warn, or refuse, rather than failing four hours later in the one step that + cannot be retried. Same shape as every other fix made that night: check + where it is cheap, not where it is fatal +- [ ] Until the writer compresses, `sds bench` should compress finished logs as + it goes. A run that needs a babysitter script to fit on disk is not + finished work - [ ] `phi` as an array against the row's `learnable` order would cut the log at the source rather than after the fact. It is a protocol change and a breaking one, so it needs its own argument