diff --git a/CLAUDE.md b/CLAUDE.md index c7bc484..f7f655c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -30,10 +30,10 @@ To start work on a feature branch: Each feature branch should include supporting work: - Maintenance work to reduce duplicated or dead code -- Adding new debug log lines or similar observability +- Adding new debug log lines, counters, or similar observability - An update to any high-level documentation (if needed) - A few of high-value tests (if applicable) -- Updating the "what is not done yet" list in the relevant `docs/` file +- Checking off completed work in the epic in `plan/` that the work belongs to Structure your commits: @@ -105,6 +105,41 @@ After a feature is merged: - Delete both the local and the remote feature branch. - Leave any temporary artifacts (like Docker images). +## The plan + +`plan/` holds one file per epic, and `plan/README.md` is the register. An epic's +filename is its id, and that id is the Conventional Commits scope: +`feat(features): ...` belongs to `plan/features.md`. +`scripts/check-commit-scope.sh` enforces it at commit-msg time against the epic +files, including the archived ones in `plan/complete/`. + +The register's tables are generated from the epics' frontmatter by +`scripts/gen-plan-readme.py`, and the `plan-register` hook fails a commit that +leaves them stale. Adding an epic is writing `plan/.md` and running that +script - never editing a table by hand. + +There is no TODO.md. Work that is worth writing down goes in its epic. + +## Working alongside other agents + +Several agents may be running here at once, and a match is expensive. + +- One worktree per branch, under `.claude/worktrees`. Each builds its own jar + into its own `bridge/build/` and mounts its own checkout, so two agents never + share a container image or a result directory. +- Match containers are named `sds-*` and are killed on exit. **Never run + `sds clean` while somebody else's benchmark is running** - it kills every + match container on the machine. +- `--jobs` defaults to 2. A match is two thinking bots and a server in one + container, and three at once already contends on this machine. Look at + `uptime` before raising it. +- Cargo builds are small here (tokio and serde), so worktrees do not share a + target directory. If that changes, share one deliberately rather than by + accident - a shared target can answer a test from another worktree's binary. +- A denied tool call is a hard stop, not a routing problem. If a permission + refusal blocks a subagent, it says so and stops; it does not find another way + to the same effect. + ## Layout | | | @@ -115,7 +150,8 @@ After a feature is merged: | `sds/` | the Python harness: run matches, aggregate, compare | | `scenarios/` | `.mms` files | | `baselines/` | committed aggregates. Never match files | -| `docs/` | `PROTOCOL.md`, `HIERARCHY.md`, `EXPERIMENTS.md` | +| `docs/` | `PROTOCOL.md`, `HIERARCHY.md`, `EXPERIMENTS.md`, `UPSTREAM.md` | +| `plan/` | one file per epic; `README.md` is the generated register | | `runs/` | match output. Not committed | ## Invariants diff --git a/plan/README.md b/plan/README.md new file mode 100644 index 0000000..732a069 --- /dev/null +++ b/plan/README.md @@ -0,0 +1,108 @@ +# Plan + +What sds is building, and roughly in what order. One file per epic. + +An **epic** is a line of work that takes many pull requests and has an exit +criterion somebody could check. It is not a task list - the tasks live inside it. + +Same shape as headquarters' register, and for the same reasons. + +## The id is the commit scope + +Each epic's filename is its id, and that id never changes. `plan/features.md` is +`features`, so its commits read `feat(features): ...`. Archiving moves the file +into [complete/](complete/) and never renames it. + +`plan` is a valid scope too, for changes to this register. + +## Order is advisory + +The Open table reads roughly top to bottom and that is the whole of the ordering. +What actually constrains the work is `dependsOn`, which says what genuinely +cannot start first. The sequence comes from [order.txt](order.txt), which is a +list of ids and nothing else. + +## Status + +`open` - being worked on or ready to be. + +`blocked` - cannot start until something in `dependsOn` lands. + +`continuous` - no exit criterion. Worked whenever adjacent code is open. + +`declined` - decided against. Not `blocked`, which is waiting: the answer is no +rather than not yet. The file stays, because a decision that is not written down +gets made again. + +`shipped` - the exit criterion is met. A shipped epic with nothing open moves to +[complete/](complete/). + +## The tables are generated + +`scripts/gen-plan-readme.py` reads every epic's frontmatter and rewrites what +sits between the `` fences. Everything outside a fence is +hand-written and passed through. The `plan-register` prek hook runs it with +`--check`, so a row that disagrees with the file it points at fails the commit. + +Adding an epic is one file and one command: write `plan/.md` with the +frontmatter keys every other epic carries, run the script, stage both. + +## Where the work stands + +[harness](harness.md) gates everything. A bot with no measuring stick is a +collection of opinions, and until a control comes out near 50/50 with stalls near +zero, no number in this repository means anything. + +After that the spine is [protocol](protocol.md) -> [los](los.md) / +[ev](ev.md) -> [features](features.md) -> [strategies](strategies.md), because +each one is what the next needs to be computable. [training](training.md) is last +for a reason: it fits `M`, and `M` does not exist until `features` and +`strategies` do. + +## Open + + +| id | title | status | depends on | +|---|---|---|---| +| [harness](harness.md) | The benchmark produces results you can believe | open | - | +| [scenarios](scenarios.md) | A suite that spans sizes, maps and technology | open | - | +| [protocol](protocol.md) | The observation carries what a bot needs to reason | open | - | +| [hierarchy](hierarchy.md) | Units propose, forces decide | open | - | +| [ev](ev.md) | Damage as a distribution, not a mean | open | protocol | +| [los](los.md) | Line of sight computed here, checked against MegaMek | blocked | protocol | +| [features](features.md) | A normalised, named feature basis | blocked | protocol, los, ev | +| [strategies](strategies.md) | Strategy as a blend of features, held with grit | blocked | features | +| [beliefs](beliefs.md) | What the bot thinks is happening | blocked | features | +| [goals](goals.md) | What a force is trying to achieve | blocked | strategies, protocol | +| [doctrine](doctrine.md) | Composable lenses - culture, commander, personality, experience | blocked | strategies | +| [capability](capability.md) | Forces and doctrines map to each other | blocked | strategies | +| [difficulty](difficulty.md) | Fallible, never absurd | blocked | beliefs, doctrine | +| [training](training.md) | Weights fitted from recorded play | blocked | features, harness | +| [observability](observability.md) | Read back what the bot believed and chose | blocked | beliefs | + + +## Continuous + + +| id | title | status | depends on | +|---|---|---|---| + + +## Not pursued + + +| id | title | status | depends on | +|---|---|---|---| + + +## Complete + +Nothing open, nothing left to decide. These live in [complete/](complete/) so the +list above stays the list of things somebody might work on. They are not deleted: +how something was built, and what it cost to learn, is worth more after it works +than before. Their ids stay valid commit scopes. + + +| id | title | +|---|---| + diff --git a/plan/beliefs.md b/plan/beliefs.md new file mode 100644 index 0000000..5625f42 --- /dev/null +++ b/plan/beliefs.md @@ -0,0 +1,32 @@ +--- +id: beliefs +title: What the bot thinks is happening +status: blocked +dependsOn: [features] +exitCriterion: > + A belief layer estimates attrition, remaining rounds and opponent doctrine, and + its calibration is scored against recorded matches without playing new ones. +--- + +# beliefs + +Separate from what the bot is trying to do. Every force level reads the same +belief state; only stances are per-node. + +**This is the only part of the bot with a post-hoc ground truth.** How many +rounds actually remained, what the opponent's doctrine actually was - all +knowable later, so the layer can be scored from replays with no new matches. That +is a much faster iteration loop than win rate. + +- [ ] Threat map: reachability envelope per enemy, then expected incoming damage + per hex. Computed once per side per turn +- [ ] Attrition: fit a decay rate per side from per-round BV +- [ ] Rounds-remaining posterior, prior from the benchmark corpus conditioned on + BV ratio +- [ ] **Host records per-round BV per side.** Currently only the final. This is + the prerequisite for the prior and it is a few lines +- [ ] Opponent doctrine posterior: fit a stance to observed behaviour, prior from + their force composition. Same projection as [capability](capability.md), + different input +- [ ] Calibration scoring from replays: Brier scores, calibration curves +- [ ] Time-to-kill and time-to-be-killed per unit diff --git a/plan/capability.md b/plan/capability.md new file mode 100644 index 0000000..28e1252 --- /dev/null +++ b/plan/capability.md @@ -0,0 +1,32 @@ +--- +id: capability +title: Forces and doctrines map to each other +status: blocked +dependsOn: [strategies] +exitCriterion: > + Selecting a lance for a doctrine and inferring a doctrine from that lance + round-trips. +--- + +# capability + +Both directions are the same projection. Give each design a capability vector in +the **same basis** as the behavioural features, and then: + +- doctrine -> lance is a knapsack under BV, C-bill or availability constraints +- lance -> doctrine is a constrained least squares over the simplex + +helm already computes most of it: `max_range`, `heat_efficiency`, +`damage_at_6/7/9/18`, `single_damage_at_15/18` are the range profile, the +sustain-fire signal and the punch-versus-sandblast axis. + +- [ ] Capability vector as a pure function of stats and equipment, in the feature + basis - no hand-labelled unit tiers +- [ ] Selection under multiple constraints: BV, C-bills, faction, era +- [ ] Inference from a lance +- [ ] **Round-trip test**: `infer(select(lens)) ~= lens`, at zero noise. If it + does not round-trip, the basis is wrong or the capability model is lossy +- [ ] Mismatch score surfaced at authoring time - "this doctrine and this lance + disagree". Same computation as the runtime `role_conflict`, different input +- [ ] A green commander should select a slightly mismatched lance; run the same + bias through selection diff --git a/plan/difficulty.md b/plan/difficulty.md new file mode 100644 index 0000000..5c65fa1 --- /dev/null +++ b/plan/difficulty.md @@ -0,0 +1,31 @@ +--- +id: difficulty +title: Fallible, never absurd +status: blocked +dependsOn: [beliefs, doctrine] +exitCriterion: > + An easy bot loses to a better player by misjudging, not by choosing moves it + believes are bad, and never walks into water at any setting. +--- + +# difficulty + +Most game AI is made harder by cheating or by epsilon-greedy. Cheating feels +unfair; epsilon-greedy produces the huge goof, because the mistake is a +*deliberate* choice of something the bot knows is terrible. + +**Noise goes on beliefs, never on action selection.** The bot always takes the +argmax of what it believes. A catastrophic move has to be catastrophically +mis-estimated to be chosen. + +- [ ] Per-feature epistemic difficulty. To-hit is a printed number; anticipating + where the enemy will be is hard. Green pilots should make positional and + anticipatory errors, not arithmetic ones +- [ ] Systematic bias, not just variance: green is optimistic about its own + damage and oblivious to incoming. More recognisable than noise +- [ ] Veto list: drowning, jumping into lava, shutting down at zero heat. + Forbidden at every setting +- [ ] Deliberation budget as an honest knob - thinking less, not knowing less +- [ ] **Difficulty may only degrade beliefs.** It may touch bias, noise and + budget; never the observation. That is what makes "does the hard bot see + more" answerable by reading one file diff --git a/plan/doctrine.md b/plan/doctrine.md new file mode 100644 index 0000000..8e3404c --- /dev/null +++ b/plan/doctrine.md @@ -0,0 +1,37 @@ +--- +id: doctrine +title: Composable lenses - culture, commander, personality, experience +status: blocked +dependsOn: [strategies] +exitCriterion: > + A Clan force plays zellbrigen, breaks it under pressure, and a human can author + a new opponent by writing rows and masks rather than training. +--- + +# doctrine + +A doctrine is a strategy-availability mask, weight modifiers on the shared basis, +and latch features. Lenses stack multiplicatively (additive in log space) so they +commute and stay attributable: `risk_tolerance: base 1.00 x veteran 1.30 x clan +1.80 -> 2.34`. + +Lenses may overlap. The composition operator is a property of the knob type - +weights blend, availability masks intersect. + +- [ ] Lens type and composition, with per-knob operators +- [ ] Culture: Clan zellbrigen. Not hard constraints - fiction shows Clanners + going dezgra under pressure, and a near-zero prior behaves differently from + a zero. Authored weights the optimiser may not touch +- [ ] `honour_broken` latch, per opponent, one-way +- [ ] Commander: force-level parameters - grit, coordination discipline +- [ ] Personality: unit-level weights +- [ ] Experience: belief noise and bias only, never preferences. See + [difficulty](difficulty.md) +- [ ] Default: generic IS mercenary commander - flexible, broad availability, + mildly loss-averse. Also the right prior for [training](training.md) +- [ ] Validity check on composed lenses; veteran+green should not silently work + +Verified on sarna: zellbrigen does not formally forbid melee, forbids +interfering in another's duel, and makes breaking LOS or withdrawing out of range +*dezgra*. So it forbids focus fire - the highest-value force behaviour - which is +what makes Clan forces a different opponent rather than a stronger one. diff --git a/plan/ev.md b/plan/ev.md new file mode 100644 index 0000000..c4fbf33 --- /dev/null +++ b/plan/ev.md @@ -0,0 +1,28 @@ +--- +id: ev +title: Damage as a distribution, not a mean +status: open +dependsOn: [protocol] +exitCriterion: > + Fire allocation is chosen on P(threshold), and the calculator is tested against + hand-computed cases. +--- + +# ev + +At a to-hit of 10 a PPC has an expected 0.83 damage and an actual outcome of +"nothing, 92% of the time". Overkill decisions need the distribution; the mean +says a Locust needs two PPCs when it needs many more. + +- [ ] Per-shot PMF over integer damage: `1 - P(hit)` at zero, the rest across the + cluster table for the rack size, times per-missile damage, minus AMS +- [ ] Convolve shots for an allocation; memoise on + `(rack, per-missile damage, to-hit, AMS)` +- [ ] Query it for mean, p50, p90, `P(>=20)` (the piloting-roll cliff), + `P(kill)`, `P(breach)` +- [ ] Carry expected packet count alongside, for punch-versus-sandblast +- [ ] Force-level greedy allocation against the concave value curve. Greedy on a + submodular objective is within 1-1/e; any search over subsets is exponential +- [ ] **Declare high-concentration weapons first.** Declaration order is + resolution order in MegaMek and nothing sorts it, so a PPC can open a + location for the cluster weapons behind it. Free diff --git a/plan/features.md b/plan/features.md new file mode 100644 index 0000000..8df52a6 --- /dev/null +++ b/plan/features.md @@ -0,0 +1,41 @@ +--- +id: features +title: A normalised, named feature basis +status: blocked +dependsOn: [protocol, los, ev] +exitCriterion: > + Every scoring decision is a dot product of named, normalised features, and each + feature is explainable to a player in one sentence. +--- + +# features + +Replaces the current `Intent` enum. A proposal's value stops being a self-report +(`Proposal.serves`) and becomes a measurement. + +**Normalisation is the crux.** Raw features in mixed units reproduce Princess's +actual defect - whichever term has the biggest numbers on this board dominates +the argmax. Each feature declares its normalisation type, and the type decides +whether it can be used for learning. + +| type | meaning | learnable | +|---|---|---| +| bounded | a rule gives the domain (to-hit 2-12, facing 0-5) | yes | +| rank | ordinal among units on the board; force-size invariant | yes | +| local | min-max across this decision's candidates only | argmax only | + +- [ ] Feature trait with a declared normalisation type; make an unnormalised + feature impossible rather than discouraged +- [ ] Positional: `range_band_fit`, `exposure`, `cover_quality`, `elevation_gain`, + `tmm_gained`, `distance_to_focus`, `cohesion`, `rear_arc_gain`, `psr_risk`, + `heat_cost`, `los_out`, `los_in` +- [ ] Firing: `expected_damage`, `p_kill`, `p_psr_threshold`, `heat_incurred`, + `ammo_spent`, `overkill`, `weapon_concentration` +- [ ] Target: `target_health`, `target_breach`, `target_threat`, `target_skill`, + `target_shutdown`, `target_heat_load`, `target_heat_neutrality` +- [ ] Force-level, over the joint assignment: `concentration`, `arc_spread`, + `exposure_balance`, `psr_threshold_hits` +- [ ] Latches (monotone, per-opponent where relevant): `has_been_fired_upon`, + `lost_a_unit`, `honour_broken` +- [ ] Interaction features, budgeted and named: `weapon_concentration x + target_breach` first diff --git a/plan/goals.md b/plan/goals.md new file mode 100644 index 0000000..1725793 --- /dev/null +++ b/plan/goals.md @@ -0,0 +1,27 @@ +--- +id: goals +title: What a force is trying to achieve +status: blocked +dependsOn: [strategies, protocol] +exitCriterion: > + A scenario with an extraction objective is played toward that objective rather + than toward destroying the enemy. +--- + +# goals + +A goal changes the objective, not the policy, which is why it cannot be another +strategy. "Get four units off the east edge" makes trading BV bad even when the +trade is favourable, and the map edge is not in any combat feature. + +Goals decompose by **assignment** (1st Lance screens while 2nd exits); strategies +decompose by **blending**. Different operations. + +MegaMek already carries this: MMS V2 has per-player `victory:` and composable +triggers - `UnitPositionTrigger` is literally an extraction goal. + +- [ ] Read goals from the scenario rather than inventing a vocabulary +- [ ] Goal-conditioned features (`distance_to_exit`) in one union basis with + zeros, so `M` stays rectangular +- [ ] Per-unit goals: this one is the VIP, these are the screen +- [ ] Default goal: defeat the enemy as a cooperating team diff --git a/plan/harness.md b/plan/harness.md new file mode 100644 index 0000000..53aed01 --- /dev/null +++ b/plan/harness.md @@ -0,0 +1,28 @@ +--- +id: harness +title: The benchmark produces results you can believe +status: open +dependsOn: [] +exitCriterion: > + Princess against itself comes out near 50/50 over 40 games, with under 5% of + matches undecided for any reason other than the round limit. +--- + +# harness + +Everything else is downstream of this. Until it closes, no number here means +anything. + +Four hangs found so far. The worst was not ours: `AutosaveService` spins forever +on `String.format` when two matches write the same save name in the same second. +See [UPSTREAM.md](../docs/UPSTREAM.md). + +- [x] Latch the result from a game listener rather than polling for VICTORY +- [x] Client-originated readiness, lounge only, reading the bot's phase +- [x] Uncaught-exception handler and a thread dump on stall +- [x] Disable MegaMek's rolling autosave +- [ ] **A 40-game control with stalls near zero.** The gate for everything else +- [ ] Check whether the lost-wakeup path in `Server.PacketPump.run` is also live +- [ ] Measure the undecided rate per scenario size; decide whether the BV + threshold, the round limit or the map is what to move +- [ ] `sds check-determinism`: same seed twice, diff the decision logs diff --git a/plan/hierarchy.md b/plan/hierarchy.md new file mode 100644 index 0000000..a351d98 --- /dev/null +++ b/plan/hierarchy.md @@ -0,0 +1,31 @@ +--- +id: hierarchy +title: Units propose, forces decide +status: open +dependsOn: [] +exitCriterion: > + A company of two lances plans in one cycle per round, every force waits on all + its children, and the rejected proposals are in the record. +--- + +# hierarchy + +Princess decides one unit at a time knowing nothing about what the others will +do; eight locally-best moves are a queue, not a plan. Here units propose and +appraise, each force waits for all its children, then commits and chooses. + +Force structure comes from MegaMek's own `Forces` tree; roles from `UnitRole`. + +Built already: two-level planning, the barrier, stance/grit, node transports +(`local:`, `proc:`, `tcp:`) so a level can move to another machine. + +- [ ] Company level does something with more than one lance beyond passing a + stance down +- [ ] Two-pass planning: cheap appraisal first so the force can pick a stance, + expensive LOS-bearing refinement second and only where the stance says it + matters. Unblocks the commander and cuts total work +- [ ] Move order is the harness's, not the bot's. It should be the bot's +- [ ] Deployment is a placeholder (nearest legal hex to the middle); it deserves + proposals and a lance-chosen shape +- [ ] Determinism: fixed-order reduction, budgets in work units not wall-clock, + per-unit seeded RNG streams, no clock reads in decision logic diff --git a/plan/los.md b/plan/los.md new file mode 100644 index 0000000..bfb44c2 --- /dev/null +++ b/plan/los.md @@ -0,0 +1,27 @@ +--- +id: los +title: Line of sight computed here, checked against MegaMek +status: blocked +dependsOn: [protocol] +exitCriterion: > + A Rust LOS agrees with MegaMek's LosEffects on a corpus dumped from real + matches, and every positional feature uses it. +--- + +# los + +Candidate evaluation needs LOS from hexes the unit is *not* standing in. +MegaMek's `toHit` only answers for actual positions, so we have to compute it +ourselves regardless of speed. + +It will also be the dominant cost of any real positional feature: ~1280 LOS +queries a round at 8v8, versus microseconds for everything else measured. + +- [ ] Rust LOS over the board we already fetch in bulk +- [ ] Partial cover (the eight `COVER_*` variants), foliage, water, buildings, + attacker and target elevation, prone, hull-down, the divided-line case +- [ ] **Differential corpus**: dump every `(attacker, target)` -> `LosEffects` + from real matches and test against it. Divergence here is silent - it + looks like bad play, not an error +- [ ] Per-turn cache shared across the side, compute-once many-waiters +- [ ] Measure MegaMek's `calculateLOS` for comparison rather than assuming diff --git a/plan/observability.md b/plan/observability.md new file mode 100644 index 0000000..b6f2b3d --- /dev/null +++ b/plan/observability.md @@ -0,0 +1,25 @@ +--- +id: observability +title: Read back what the bot believed and chose +status: blocked +dependsOn: [beliefs] +exitCriterion: > + A finished match can be stepped through turn by turn, showing what the bot + believed and what it rejected. +--- + +# observability + +The `rationale` field already carries each force's stance, what changed its mind, +and every proposal it rejected. Nothing reads it. + +"The lance chose badly" and "no unit offered anything better" look identical in a +record that keeps only the winner. + +- [ ] Render a match turn by turn as HTML: threat map, rounds-remaining + posterior, opponent doctrine posterior, per-target damage distributions, + proposals kept and dropped +- [ ] Publish it as an artifact so a tab can be left open while playing +- [ ] Highlight latch transitions - "round 7, honour broken against Jade Falcon" + is the story beat +- [ ] Per-phase `answered`/`defaulted`/`illegal` already exist; surface them here diff --git a/plan/order.txt b/plan/order.txt new file mode 100644 index 0000000..2998c47 --- /dev/null +++ b/plan/order.txt @@ -0,0 +1,22 @@ +# The order the epics read in, wherever they appear in README.md's tables. +# scripts/gen-plan-readme.py sorts by this list and puts anything it does not +# name at the end of its table, alphabetically. +# +# Advisory and optional. A new epic needs no line here - add one only when where +# it sits is worth saying to somebody reading down the list. + +harness +scenarios +protocol +hierarchy +ev +los +features +strategies +beliefs +goals +doctrine +capability +difficulty +training +observability diff --git a/plan/protocol.md b/plan/protocol.md new file mode 100644 index 0000000..8c85167 --- /dev/null +++ b/plan/protocol.md @@ -0,0 +1,33 @@ +--- +id: protocol +title: The observation carries what a bot needs to reason +status: open +dependsOn: [] +exitCriterion: > + Every feature in the basis can be computed from what the wire carries, without + the bot reaching into MegaMek. +--- + +# protocol + +The observation is the interface. What is not in it does not exist for any bot. +Adding a field is deliberate, with a reason in the commit message. + +- [ ] **Per-location armour and structure** - blocks breach state, so it blocks + punch-versus-sandblast allocation and [cripple](strategies.md) +- [ ] **Ammo type per weapon** - per-missile damage varies by ammo; [ev](ev.md) + cannot be right without it +- [ ] **AMS on targets** - reduces cluster size +- [ ] **Terrain deltas per round** - fire, smoke, collapse and cleared woods + change LOS. The board is sent once and never updated, so [los](los.md) + would drift correctly-implemented and wrong +- [ ] **Scenario objectives** - MMS V2 carries per-player `victory:` conditions + and composable triggers. Blocks [goals](goals.md) +- [ ] **Minefields** - the bot is blind to them + +Fixed already: `WeaponType.getDamage()` returns `-2` for cluster weapons, so the +bot valued every missile below nothing. There is now no plain `damage` field. + +Not yet needed: the observation is sent per decision, ~450KB a round at 8v8. +Sending it once per phase would remove most of that. Measured at 291us against +~10ms of thinking, so it is tidiness, not performance. diff --git a/plan/scenarios.md b/plan/scenarios.md new file mode 100644 index 0000000..95dbc53 --- /dev/null +++ b/plan/scenarios.md @@ -0,0 +1,29 @@ +--- +id: scenarios +title: A suite that spans sizes, maps and technology +status: open +dependsOn: [] +exitCriterion: > + A run covers 1v1 through 8v8, several maps and several technology brackets, + and reports per-scenario as well as in aggregate. +--- + +# scenarios + +One 4v4 on one map measures a bot that is good at that fight. Size, map, era and +composition vary; the two sides never differ from each other, which is what keeps +a win rate attributable to the bots. + +Forces come from helm, which already indexes year, BV, role and tonnage. + +- [x] Generator: four sizes, twelve non-urban boards, era brackets, role diversity +- [x] `bench --suite` with a fingerprint, so a regenerated suite is incomparable + with an older baseline rather than quietly comparable +- [ ] **Validate every generated scenario loads** before a run wastes an hour on + a unit name MegaMek does not have +- [ ] Per-scenario breakdown in the summary +- [ ] Asymmetric scenarios, once the mirrored baseline is stable +- [ ] Objective scenarios, which [goals](goals.md) needs + +City maps are excluded: the bot has no notion of buildings, so a city fight +measures its blindness. diff --git a/plan/strategies.md b/plan/strategies.md new file mode 100644 index 0000000..1e2d92d --- /dev/null +++ b/plan/strategies.md @@ -0,0 +1,30 @@ +--- +id: strategies +title: Strategy as a blend of features, held with grit +status: blocked +dependsOn: [features] +exitCriterion: > + A force's stance is a simplex over named strategies, blending composes, and + an SVD of M shows the vocabulary is not redundant. +--- + +# strategies + +`w = S^T M`. `S` is what the force is doing, `M` maps strategies to feature +weights, and scoring is `Phi w`. Blending two strategies gives a weighting that +is genuinely half of each, which an enum cannot do. + +Commitment is one parameter. `grit` sets both the margin a rival must beat and +the drift rate once it moves - facets of stubbornness, and splitting them invites +tuning them against each other. Already built and tested against `Intent`; it +carries over unchanged. + +- [ ] Rows of `M`: `direct_engagement`, `destroy_weakest`, `gun_line`, `flanking`, + `fighting_withdrawal`, `crossfire`, `cripple`, `bait`, `screen`, `overrun` +- [ ] Port `Commitment`/`Grit` from the intent simplex to the strategy simplex +- [ ] Force-level assignment: greedy per-unit argmax, then swaps that improve the + force-level features. Crossfire is not expressible per unit +- [ ] **SVD check on `M`.** If effective rank collapses, the named strategies are + theatre - that is Princess's disease measured rather than read +- [ ] Composition operator as a hyperparameter: the power mean spans min, + geometric and arithmetic in one `p` diff --git a/plan/training.md b/plan/training.md new file mode 100644 index 0000000..5260a89 --- /dev/null +++ b/plan/training.md @@ -0,0 +1,34 @@ +--- +id: training +title: Weights fitted from recorded play +status: blocked +dependsOn: [features, harness] +exitCriterion: > + A fitted M beats the hand-authored M by an interval that excludes zero, on a + suite run the hand-authored M was not tuned against. +--- + +# training + +Only `M` is fitted. Features stay hand-written and named; that is the whole +containment strategy. + +**Label granularity decides feasibility.** Win/loss gives one sample per match +and needs ~10^5-10^6 matches - infeasible. BV traded over the next N rounds gives +~240 per side per match, so 300-500 matches is a comfortable fit for ~240 +parameters. That is one overnight run. + +- [ ] **Label pipeline**: log `Phi` per candidate, the chosen index, and the BV + outcome over the next N rounds. Nothing does this today and it is the + actual prerequisite +- [ ] Fitting script: linear regression, writing a new `M` +- [ ] Train at zero noise. Weights fitted with noise active learn to be + noise-robust, which silently changes the policy +- [ ] Difficulty parameters in an outer loop, fitted to a target win rate +- [ ] Keep a hand-authored `M` as a permanent baseline. If fitted only wins by a + few points, ship the hand-authored one - it is editable and explainable +- [ ] Self-play only near parity. Two bad bots teach each other to beat a bad bot + +Compute is not the constraint: ~0.1 core-hours a match, so 20,000 matches is +~$24 of Graviton spot. The constraints are a bot worth learning from and a +harness that does not discard a quarter of its games. diff --git a/prek.toml b/prek.toml index af51385..2cfa27c 100644 --- a/prek.toml +++ b/prek.toml @@ -1,3 +1,5 @@ +default_install_hook_types = ["pre-commit", "commit-msg"] + # Hooks for prek (https://prek.j178.dev/). See README for setup. # # Same shape as the sibling repos, minus what this one has no files for: no @@ -42,6 +44,24 @@ hooks = [ [[repos]] repo = "local" +# The register's tables are generated from the epics' frontmatter. A row that +# disagrees with the file it points at is worse than no row. +[[repos.hooks]] +id = "plan-register" +name = "plan register is current" +language = "system" +entry = "scripts/gen-plan-readme.py --check" +pass_filenames = false +files = '^plan/.*\.md$' + +# Conventional Commits, with the scope naming an epic. See the script. +[[repos.hooks]] +id = "commit-scope" +name = "commit scope names an epic" +language = "system" +entry = "scripts/check-commit-scope.sh" +stages = ["commit-msg"] + # The harness's own tests. Fast on purpose - no match is played - so that the # thing which decides whether a result is real is checked on every commit # rather than when somebody remembers. diff --git a/scripts/check-commit-scope.sh b/scripts/check-commit-scope.sh new file mode 100755 index 0000000..03ffc75 --- /dev/null +++ b/scripts/check-commit-scope.sh @@ -0,0 +1,44 @@ +#!/usr/bin/env bash +# Conventional Commits, with the scope naming an epic in plan/. +# +# The scope is how a commit says which line of work it belongs to, and it is +# only useful if it means something: an id that matches no epic is a typo or an +# epic somebody forgot to write. Archived epics count - a follow-up fix to +# finished work is still that epic's commit. +# +# scripts/check-commit-scope.sh .git/COMMIT_EDITMSG +set -euo pipefail + +HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +MSG_FILE="${1:?usage: check-commit-scope.sh }" +SUBJECT="$(head -n 1 "$MSG_FILE")" + +# Fixups and merges are not ours to judge. +case "$SUBJECT" in + fixup!*|squash!*|Merge*|Revert*) exit 0 ;; +esac + +CONVENTIONAL='^(feat|fix|docs|refactor|perf|test|build|ci|chore)(\([a-z0-9-]+\))?!?: .+' +if ! printf '%s' "$SUBJECT" | grep -Eq "$CONVENTIONAL"; then + echo "commit subject is not Conventional Commits:" >&2 + echo " $SUBJECT" >&2 + echo "want: type(scope): summary e.g. feat(features): normalise the basis" >&2 + exit 1 +fi + +SCOPE="$(printf '%s' "$SUBJECT" | sed -n 's/^[a-z]*(\([a-z0-9-]*\)).*/\1/p')" +[ -n "$SCOPE" ] && [ "$SCOPE" != "plan" ] || exit 0 + +for dir in "$HERE/plan" "$HERE/plan/complete"; do + [ -f "$dir/$SCOPE.md" ] && exit 0 +done + +echo "commit scope '$SCOPE' names no epic in plan/." >&2 +echo "known:" >&2 +for f in "$HERE"/plan/*.md "$HERE"/plan/complete/*.md; do + [ -e "$f" ] || continue + b="$(basename "$f" .md)" + [ "$b" = "README" ] || echo " $b" >&2 +done +echo " plan (for the register itself)" >&2 +exit 1 diff --git a/scripts/gen-plan-readme.py b/scripts/gen-plan-readme.py new file mode 100755 index 0000000..57acc08 --- /dev/null +++ b/scripts/gen-plan-readme.py @@ -0,0 +1,182 @@ +#!/usr/bin/env python3 +"""Generate the tables in plan/README.md from the epics' own frontmatter. + +The same shape as headquarters' register, written fresh rather than copied: the +value is the structure - one file per epic, the id is the commit scope, the +tables are output - and a forked copy of another repo's tooling would drift +without anyone noticing. + +Every table sits between an HTML comment fence and is rewritten from the files. +Everything outside a fence is hand-written and passed through untouched. Which +table an epic lands in follows from where it is and what its status says: +plan/complete/ is Complete, and outside it the status picks between the rest. + +Row order comes from plan/order.txt, which is advisory - an epic it does not +name still appears, alphabetically, at the end of its table. So adding an epic +is one file and one command, and two branches adding one do not conflict over +rows neither of them cared about. + +No third-party modules: this runs on every commit that touches plan/. + + scripts/gen-plan-readme.py rewrite the register + scripts/gen-plan-readme.py --check ask whether it is current +""" + +import argparse +import difflib +import re +import sys +from pathlib import Path + +REPO = Path(__file__).resolve().parent.parent +PLAN = REPO / "plan" +COMPLETE = PLAN / "complete" +REGISTER = PLAN / "README.md" +ORDER = PLAN / "order.txt" + +# Fence name -> (heading it lives under, which epics belong in it). +TABLES = { + "complete": lambda epic: epic["_complete"], + "open": lambda epic: not epic["_complete"] and epic["status"] in ("open", "blocked"), + "continuous": lambda epic: not epic["_complete"] and epic["status"] == "continuous", + "declined": lambda epic: not epic["_complete"] and epic["status"] == "declined", +} + +SCALARS = ("id", "title", "status") +LISTS = ("dependsOn",) + + +def parse_frontmatter(text: str, path: Path) -> dict: + """Enough YAML for the keys an epic carries, and nothing else. + + A real parser would be a dependency on a hook that runs constantly. The + subset is: `key: value`, `key: [a, b]`, and `key: >` folded blocks. + """ + if not text.startswith("---\n"): + raise SystemExit(f"{path}: no frontmatter") + end = text.index("\n---\n", 3) + body = text[4:end] + epic: dict = {"dependsOn": []} + key = None + folded: list[str] = [] + for line in body.splitlines(): + if key and (line.startswith(" ") or not line.strip()): + folded.append(line.strip()) + continue + if key: + epic[key] = " ".join(part for part in folded if part) + key, folded = None, [] + if not line.strip() or line.lstrip().startswith("#"): + continue + name, _, value = line.partition(":") + name, value = name.strip(), value.strip() + if value == ">": + key = name + continue + if value.startswith("[") and value.endswith("]"): + inner = value[1:-1].strip() + epic[name] = [v.strip() for v in inner.split(",") if v.strip()] + else: + epic[name] = value + if key: + epic[key] = " ".join(part for part in folded if part) + for required in SCALARS: + if required not in epic: + raise SystemExit(f"{path}: frontmatter is missing '{required}'") + return epic + + +def load() -> list[dict]: + epics = [] + for directory, complete in ((PLAN, False), (COMPLETE, True)): + if not directory.is_dir(): + continue + for path in sorted(directory.glob("*.md")): + if path.name == "README.md": + continue + epic = parse_frontmatter(path.read_text(), path) + if epic["id"] != path.stem: + raise SystemExit(f"{path}: id '{epic['id']}' does not match the filename") + epic["_complete"] = complete + epic["_path"] = f"complete/{path.name}" if complete else path.name + epics.append(epic) + return epics + + +def ranking() -> list[str]: + if not ORDER.is_file(): + return [] + return [ + line.strip() + for line in ORDER.read_text().splitlines() + if line.strip() and not line.startswith("#") + ] + + +def rows(epics: list[dict], order: list[str], complete: bool) -> str: + def sort_key(epic: dict) -> tuple: + try: + return (0, order.index(epic["id"]), "") + except ValueError: + return (1, 0, epic["id"]) + + epics = sorted(epics, key=sort_key) + if complete: + lines = ["| id | title |", "|---|---|"] + for epic in epics: + lines.append(f"| [{epic['id']}]({epic['_path']}) | {epic['title']} |") + return "\n".join(lines) + lines = ["| id | title | status | depends on |", "|---|---|---|---|"] + for epic in epics: + depends = ", ".join(epic["dependsOn"]) or "-" + lines.append( + f"| [{epic['id']}]({epic['_path']}) | {epic['title']} | {epic['status']} | {depends} |" + ) + return "\n".join(lines) + + +def render(current: str, epics: list[dict], order: list[str]) -> str: + out = current + for name, belongs in TABLES.items(): + fence = re.compile(rf"(\n).*?()", re.DOTALL) + if not fence.search(out): + raise SystemExit(f"plan/README.md has no '{name}' fence") + table = rows([e for e in epics if belongs(e)], order, name == "complete") + replacement = table + "\n" + out = fence.sub(lambda m, r=replacement: m.group(1) + r + m.group(2), out, count=1) + return out + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--check", action="store_true", help="fail if stale") + args = parser.parse_args() + + epics = load() + current = REGISTER.read_text() + updated = render(current, epics, ranking()) + + if args.check: + if current != updated: + print("plan/README.md is stale. Run scripts/gen-plan-readme.py.", file=sys.stderr) + sys.stderr.writelines( + difflib.unified_diff( + current.splitlines(keepends=True), + updated.splitlines(keepends=True), + "plan/README.md", + "generated", + ) + ) + return 1 + return 0 + + if current != updated: + REGISTER.write_text(updated) + print(f"rewrote plan/README.md from {len(epics)} epics") + else: + print("plan/README.md is current") + return 0 + + +if __name__ == "__main__": + sys.exit(main())