diff --git a/experiments/README.md b/experiments/README.md new file mode 100644 index 0000000..f56e504 --- /dev/null +++ b/experiments/README.md @@ -0,0 +1,23 @@ +# Experiments + +This directory is the checked-in record of performance and representation experiments. +It exists so the repo keeps the paper trail next to the code instead of scattering +methods, results, and conclusions across chat logs or temporary files. + +Each study keeps the same minimum set of material: + +- a study-level README that explains the problem, the decision rule, and the current status +- a protocol document that explains how to rerun the study deterministically +- one subdirectory per approach or control condition +- machine-readable artifacts under each approach's `artifacts/` directory +- plain-English notes that explain the hypothesis, method, observations, and conclusion + +The current study is [event-shape-study/README.md](event-shape-study/README.md). + +Use these rules when adding new work here: + +1. Record the question before collecting new data. +2. Keep the collection commands deterministic and write them down. +3. Store raw machine-readable artifacts before writing summaries. +4. Separate observations from conclusions so later readers can challenge the interpretation. +5. Leave enough context that another maintainer could rerun the same study without guessing. \ No newline at end of file diff --git a/experiments/event-shape-study/methods.md b/experiments/event-shape-study/methods.md new file mode 100644 index 0000000..3b4ab46 --- /dev/null +++ b/experiments/event-shape-study/methods.md @@ -0,0 +1,143 @@ +# Event Shape Methods + +This study keeps each candidate self-contained so the code under test does not move around +between runs. Every approach directory can include its own `code/` snapshot, its own +artifacts, and its own notes. That keeps the benchmark input, the benchmark harness, and +the parser code aligned with the exact candidate being measured. + +## What gets measured + +Each recorded report captures two things: + +- timing samples from the approach-local `event_shape_bench.ts` +- retained-memory samples from the approach-local `event_shape_memory.ts` + +Both commands run in fresh processes. The report preserves the raw samples so later +bootstrap analysis does not need to reconstruct data from summary text. + +## Snapshot-local recording + +Record one approach with the study-local recorder: + +```bash +mise x deno@latest -- deno run --allow-run=mise,git --allow-read --allow-write \ + --no-lock \ + experiments/event-shape-study/tools/record_approach_artifacts.ts \ + --approach-dir=current-baseline \ + --variant=current-baseline \ + --runs=10 \ + --memory-repeats=5 +``` + +That command writes `artifacts/report.json`, `artifacts/recording.json`, and +`artifacts/commands.txt` inside the approach directory. + +The lower-level collector still exists when you only want a report file: + +```bash +mise x deno@latest -- deno run --allow-run=mise,git --allow-read --allow-write \ + --no-lock \ + experiments/event-shape-study/tools/collect_approach_report.ts \ + --approach-dir=current-baseline \ + --variant=current-baseline \ + --runs=10 \ + --memory-repeats=5 \ + --out=experiments/event-shape-study/current-baseline/artifacts/report.json +``` + +Record and compare a candidate against the baseline: + +```bash +mise x deno@latest -- deno run --allow-run=mise,git --allow-read --allow-write \ + --no-lock \ + experiments/event-shape-study/tools/record_approach_artifacts.ts \ + --approach-dir=planned-flat-eager-event-shape \ + --variant=planned-flat-eager-event-shape \ + --runs=10 \ + --memory-repeats=5 \ + --baseline-report=experiments/event-shape-study/current-baseline/artifacts/report.json +``` + +The recorder writes `comparison.json` and `comparison.txt` automatically when a baseline +report is supplied. + +You can still run the comparator directly when you want to inspect two existing reports: + +```bash +mise x deno@latest -- deno run --allow-read \ + --no-lock \ + experiments/event-shape-study/tools/compare_reports.ts \ + experiments/event-shape-study/current-baseline/artifacts/report.json \ + experiments/event-shape-study/planned-flat-eager-event-shape/artifacts/report.json +``` + +## Deterministic cross-candidate run + +The study also records an explicit round-robin schedule so baseline and candidate reports +are collected in a fixed order with the exact commands written to disk. + +```bash +mise x deno@latest -- deno run --allow-run=mise,git --allow-read --allow-write \ + --no-lock \ + experiments/event-shape-study/tools/run_cross_candidate_schedule.ts \ + --baseline=current-baseline \ + --approaches=current-baseline,planned-flat-eager-event-shape \ + --rounds=1 \ + --runs=10 \ + --memory-repeats=5 \ + --run-label=2026-05-16-baseline-vs-flat-eager-round-robin-01 +``` + +That command writes: + +- `cross-candidate-runs/