diff --git a/wiki/engineering/performance.md b/wiki/engineering/performance.md index fea15044..d24e15d9 100644 --- a/wiki/engineering/performance.md +++ b/wiki/engineering/performance.md @@ -104,9 +104,10 @@ The first observed default run on 2026-07-30 used Linux 6.17.9, an Intel i9-9900K, and rustc 1.93.0. It measured 2,000 release steps at 19 us average, 273 us p99, and 508 us maximum. On the same host and toolchain, the final feature-gated artifact measured generated B1 at 16 us average, 159 us -p99, and 305 us maximum, and saturated B1 at 196 us average, 399 us p99, and -643 us maximum. An earlier pre-gating saturated run measured 192 us average, -383 us p99, and 551 us maximum. These measurements establish wide headroom; +p99, and 305 us maximum. After adding post-measurement load and legality gates, +saturated B1 measured 227 us average, 551 us p99, and 863 us maximum. Earlier +saturated runs measured 192-196 us average, 383-399 us p99, and 551-643 us +maximum. These measurements establish wide headroom; the executable thresholds, not those historical numbers, are the regression boundary. @@ -133,14 +134,22 @@ owned; WORK, THINK, and LIE are concurrent; Filing, Network, Paper, Financial, JobAnomaly, Power, and Thermal custody are all nonempty; 30 real sensor-upkeep Thought sinks are active; one person carries work; the 24-item information buffer is full; and a resident procedure is installed. It passes strict -current-save parse/validation/apply gates before warmup and again at the -measurement boundary, publishes the live count of every required dimension, -and uses the same 10 ms p99 / 20 ms maximum core-step budgets. The constructor +current-save parse/validation/apply gates before warmup, at the measurement +boundary, and after the timed run. It publishes both boundary inventories and +fails unless rack, fleet, sensor, tap, mode, all-seven-evidence, standing-sink, +resident-procedure, and full-buffer load survives the complete sample. Fresh +person-carried work is consumable by design: it must exist at tick 800, may +complete through the ordinary schedule during measurement, and is reported at +both boundaries rather than falsely required to remain unfinished. The fixture +uses the same 10 ms p99 / 20 ms maximum core-step budgets. The constructor and its save-boundary helpers compile only for tests or the explicit `performance-fixture` Cargo feature required by `sim_perf`; normal library and frontend builds cannot call benchmark-only state construction. -Those legality roundtrips are fixture gates, not the still-pending timed durable -save/load evidence below. +Those legality roundtrips are fixture gates outside the timed sample, not the +still-pending timed durable save/load evidence below. Observer reporting is held +Silent only after all seven exact custody kinds exist, preventing the fixture +from ending in containment; criterion 4 therefore measures saturated B1 state, +routing, and standing-load cost, not near-containment escalation pressure. ### Growth beyond B1 — design boundary @@ -179,7 +188,8 @@ whole player surface stays responsive. native duration rather than the truncated microsecond display value. 3. ✅ Percentile calculation and exact requested tick progression have executable coverage. -4. ✅ A legal saturated-B1 fixture publishes all scale counts and passes the same +4. ✅ A legal saturated-B1 fixture publishes start/end scale counts, rejects + standing-load erosion across the timed window, and passes the same release-profile core-step budgets. 5. Generated and saturated current saves have separate write and load evidence that includes durability, validation, restore, and reconciliation work. diff --git a/wiki/log/2026-07-30-saturated-b1-performance.md b/wiki/log/2026-07-30-saturated-b1-performance.md index 2bb8cd89..86a601a5 100644 --- a/wiki/log/2026-07-30-saturated-b1-performance.md +++ b/wiki/log/2026-07-30-saturated-b1-performance.md @@ -12,13 +12,15 @@ Criterion 4 now has a second fail-closed release scenario rather than borrowing the ordinary generated-B1 result. `MISALIGNED_PERF_SCENARIO=saturated-b1` builds legal current-save state, validates and restores it before warmup, refreshes consumable custody at tick 800, validates and restores that exact measurement -state, then times only the same 2,000 calls to `Sim::advance`. +state, then times only the same 2,000 calls to `Sim::advance`. After timing, it +requires every standing load dimension to remain saturated and runs the strict +current-save roundtrip a third time. The fixture reaches the current B1 ceiling without impossible benchmark-only records: - all 240 authored Foundation rack sites are present; -- all 22 legally claimable machines are owned, with 6 WORK, 4 THINK, and 5 LIE +- all 22 legally claimable machines are owned, with 5 WORK, 5 THINK, and 5 LIE bodies online at the measurement boundary; - all 30 reachable sensor devices are retained, and every one has exactly one live `MaintainDeviceTap` Thought sink; @@ -31,9 +33,23 @@ The fixture uses ordinary current simulation paths for machine acquisition, network scan and compromise, sensor taps, procedure installation, carried work, and evidence authorship. Two fixture-only preparation seams set reachable relationship/funding prerequisites and fill the bounded information buffer; the -strict current-save roundtrip rejects malformed custody both before warmup and -at the timed boundary. Every exact reservoir helper also requires one and only -one matching open sink before it can fire. +strict current-save roundtrip rejects malformed custody before warmup, at the +timed boundary, and after measurement. Every exact reservoir helper also +requires one and only one matching open sink before it can fire. + +The post-measurement gate closes a false-evidence gap: the earlier fixture proved +its load only at tick 800. At tick 2,800 the run still had all 240 rack sites, 22 +owned machines, all 30 retained reachable sensors and tap sinks, 6 WORK / 6 THINK +/ 5 LIE machines, all seven custody kinds (6 Filing, 233 Network, 4 Paper, 1 +Financial, 1 JobAnomaly, 65 Power, 49 Thermal), 30 open Thought sinks, one +resident procedure, and the full 24/24 information buffer. The fresh carried +packet had correctly completed through Dana's ordinary schedule and is reported +as zero rather than kept alive by an impossible fixture-only delay. + +After exact routed records exist, observer report policies are held Silent so a +multi-day performance sample cannot terminate in containment. This deliberately +isolates saturated B1 state, routing, and standing-load cost; it is not evidence +for near-containment escalation performance. The fixture is not production simulation machinery. Its module compiles only for core tests or the explicit `performance-fixture` Cargo feature; the @@ -46,9 +62,9 @@ Observed on Linux 6.17.9, Intel i9-9900K, rustc 1.93.0, release profile: - 800 untimed warmup ticks; - 2,000 measured ticks, ending exactly at tick 2,800; -- average 196 us; -- p99 399 us against 10,000 us; -- maximum 643 us against 20,000 us; +- average 227 us; +- p99 551 us against 10,000 us; +- maximum 863 us against 20,000 us; - result: PASS. These host numbers are evidence, not the regression boundary. The executable