From d22888171978d92e3b43933346fa73ba4d2fd99e Mon Sep 17 00:00:00 2001 From: zzstoatzz Date: Sat, 25 Jul 2026 04:30:22 -0500 Subject: [PATCH] oracle: stop sampling a transient, and record what this harness cannot prove Two changes, one to a fixture and one to the audit. The bootstrap-pipeline receipt polled for phase==bootstrap live. Whether that gauge is observable at all depends on how long bootstrap happens to take: on a small simulator the whole lifecycle finishes in under four seconds and the window closes before the first scrape lands, while a large one holds it open for minutes. The fixture was a function of simulator size rather than of Stream's behavior. It now samples whenever the pipeline has demonstrably moved and proves bootstrap really happened from the cumulative transition counter, which cannot be missed by arriving late. A dispatch-scale kill test was written and then removed rather than weakened. The invariant it chased -- per-repository versus per-batch checkpointing -- is not observable in this harness. Repository completion is coupled to archive-writer durability, and in a 100-repo simulator world the archive flushes once covering the whole corpus, so after_repo_complete at any ordinal finds every repository already durable (checked at both one and two backfill workers), while mid_backfill_download fires before anything is durable. No cut leaves a durable prefix and remaining work at the same time. That is the coupling behaving correctly, not a defect: at production scale, with frequent in-flight flushes, the prefix is partial and the distinction becomes observable. The audit row now says exactly that instead of implying a test could be written here. zig build test (306), ReleaseSafe, differential-oracle, lifecycle oracle 9/9, power-loss oracle 20/20. Co-Authored-By: Claude Opus 5 (1M context) --- docs/semantic-parity.md | 2 +- tests/oracle.py | 30 ++++++++++++++++++++++++++---- 2 files changed, 27 insertions(+), 5 deletions(-) diff --git a/docs/semantic-parity.md b/docs/semantic-parity.md index 9ef71a4..16d09fe 100644 --- a/docs/semantic-parity.md +++ b/docs/semantic-parity.md @@ -74,7 +74,7 @@ blockers remain. The detailed bootstrap audit is in | Logging | **verified** | Real processes exercise JSON/text selection, level filtering, default values, and invalid configuration. | Semantic parity is limited to the public controls and rendered fields, not byte-identical Go logging internals. | | OpenTelemetry | **partial** | Provider configuration, propagation, sampling, batching, TLS/mTLS, and representative production spans have real collector tests. | “Every pinned production span” has not been independently re-audited in this pass. Treat the enumerated tested spans as evidence, not a blanket closure. | | Inspect/version command surfaces | **partial** | `version`, sealed `inspect-segment`, active inspection, and representative `inspect-all` reports have fixtures and golden comparisons. | Inspection does not compensate for missing online status tabs, and the current audit did not rerun every golden against the current commit. | -| Differential oracle | **blocked as an admission proof** | It provides useful event-log, final-state, restart, repair, and public-client coverage against a pinned upstream simulator. Focused production-boundary tests now prove per-repository writer-gated completion and same-disk reopen with a sibling held incomplete. | The oracle still needs an invariant-based process-kill world that scales this beyond a two-repository fixture; its old four-account `after_repo_complete` schedule could not detect a dispatch-scale replay defect. | +| Differential oracle | **blocked as an admission proof** | It provides useful event-log, final-state, restart, repair, and public-client coverage against a pinned upstream simulator. Focused production-boundary tests prove per-repository writer-gated completion and same-disk reopen with a sibling held incomplete. | Still not a substitute for scale. An attempt to add a dispatch-scale kill to `tests/oracle.py` was removed rather than weakened: the invariant is not observable in this harness. Repository completion is coupled to archive-writer durability, and in a 100-repo simulator world the archive flushes once covering the whole corpus, so `after_repo_complete` at any ordinal finds every repository already durable (verified at both 1 and 2 backfill workers), while `mid_backfill_download` fires before anything is durable. No cut leaves a durable prefix and remaining work at the same time, so per-repository and per-batch checkpointing are indistinguishable here. That is the coupling behaving correctly, not a defect — at production scale, with frequent in-flight flushes, the prefix is partial and the distinction becomes observable. This row is evidence that the experiment is where this gets settled, not a reason to withhold it. | | Strict power-loss oracle | **partial** | It exercises acknowledged-write reconstruction at named write/fsync/rename boundaries. Focused tests separately inject Store read failures for lifecycle phase, relay cursor, compaction watermark, and merge cursor; the merge restart guard also distinguishes source absence from filesystem failure. The completion reopen test proves completed-versus-interrupted repository classification around an unchanged listRepos cursor. | It does not yet inject missing manifest files during replay, long listRepos cursors, or storage loss specifically between archive fsync and the completion metadata commit. | | Exact artifact admission | **verified** | `scripts/admit` implements the procedure end to end and has now been run for real. Commit `b04bb34` passed all 20 suites — both oracles included — and the receipt `receipts/b04bb34.json` binds those results to `sha256:2f2bda74d78bd7c5596acc368a60946020ed1a6b2ab4759566d2196f2e7a71ec`, which is published. `admit verify` admits that digest and refuses `sha256:adc276…`, the `e1926f3` image that failed the July experiment; the experiment harness `deploy.sh` calls it before rollout. A receipt covers a digest, never a tag; a `skipped` suite blocks admission exactly like a failure. The artifact was built on a native linux/amd64 host rather than under emulation, so it is the same architecture as the deploy target rather than a cross-emulated approximation. `tests/admission_contract.py` pins the refusals. | The receipt is bound to `b04bb34`; any later commit needs its own `admit run` and `admit publish`. | diff --git a/tests/oracle.py b/tests/oracle.py index 3ce5fe6..31fa803 100644 --- a/tests/oracle.py +++ b/tests/oracle.py @@ -489,7 +489,17 @@ def run_bootstrap_pipeline_metrics_receipt(): "--backfill-workers=1", ], stdout=log, stderr=log) try: - deadline = time.time() + 60 + # Do not wait to catch phase==bootstrap live. Whether that gauge is + # observable at all depends on how long bootstrap happens to take: on a + # small simulator the whole lifecycle finishes in under four seconds and + # the window closes before the first scrape lands, while a large one + # holds it open for minutes. Sampling a transient made this fixture a + # function of simulator size rather than of Stream's behavior. + # + # Sample whenever the pipeline has demonstrably moved, then prove + # bootstrap really happened from the cumulative transition counter, + # which cannot be missed by arriving late. + deadline = time.time() + 120 receipt = None while time.time() < deadline: if p.poll() is not None: @@ -499,12 +509,23 @@ def run_bootstrap_pipeline_metrics_receipt(): except AssertionError: time.sleep(0.02) continue - submitted = metric(samples, "stream_pipeline_ticket", stage="submitted") - if metric(samples, "jetstream_orchestrator_phase") == 1 and submitted > 0: + if metric(samples, "stream_pipeline_ticket", stage="submitted") > 0: receipt = samples break time.sleep(0.02) - assert receipt is not None, "never observed an active bootstrap pipeline" + assert receipt is not None, "pipeline never scheduled any work" + + # Anti-vacuity: the run really did pass through bootstrap, and reached + # merging, before any of the counters below were read. + assert ( + metric( + receipt, + "jetstream_orchestrator_phase_transitions_total", + **{"from": "bootstrap", "to": "merging"}, + ) + > 0 + or metric(receipt, "jetstream_orchestrator_phase") == 1 + ), "run never entered or left bootstrap" submitted = metric(receipt, "stream_pipeline_ticket", stage="submitted") claimed = metric(receipt, "stream_pipeline_ticket", stage="claimed") @@ -538,6 +559,7 @@ def run_bootstrap_pipeline_metrics_receipt(): shutil.rmtree(DATA, ignore_errors=True) + def main(): try: with urllib.request.urlopen(SIM + "/xrpc/com.atproto.sync.listRepos?limit=1", timeout=5) as r: -- 2.51.2