From ad1db0ae0b42ca4505329d00bf8cd97f2fb60589 Mon Sep 17 00:00:00 2001 From: "@permadeath.com" Date: Thu, 20 Aug 2026 11:12:18 -0400 Subject: [PATCH] docs(features): write down that damage lands somewhere Records the per-location model in the features, ev, protocol and training epics, the measured cost in docs/PERFORMANCE.md, and sign priors for p_breach, value_destroyed and p_mission_kill. Also records two harness hazards found on the way: --bot silently defaults to the random bot, and selecting work by pattern rather than identity takes other people's. --- docs/PERFORMANCE.md | 17 +++++++++++++++++ plan/ev.md | 13 +++++++++---- plan/features.md | 44 +++++++++++++++++++++++++++++++++++++++++++ plan/harness.md | 46 +++++++++++++++++++++++++++++++++++++++++++++ plan/protocol.md | 7 +++++++ plan/training.md | 26 +++++++++++++++++++++++++ sds/train.py | 14 ++++++++++++++ 7 files changed, 163 insertions(+), 4 deletions(-) diff --git a/docs/PERFORMANCE.md b/docs/PERFORMANCE.md index f4ef806..e825696 100644 --- a/docs/PERFORMANCE.md +++ b/docs/PERFORMANCE.md @@ -17,10 +17,27 @@ needs sixteen cores is not a number. | fire allocation, 48 shots x 8 targets, greedy PMF convolution | 162-543 us | | one shot's PMF, built cold | 0.2 us | | parse one 28KB observation | 291 us | +| per-location damage, one volley, cold cache | 19 us | +| per-location damage, one volley, warm cache | 0.6 us | +| 20 candidates, per-location | 33 us | The 3x spread on fire allocation between runs is contention, not variance in the work. Take the high numbers. +Per-location damage costs 33 us for a whole unit's 20 candidates, against 66 us +for the fire allocation beside it. That is affordable, and the reason is the +share. Damage arriving at a location is memoised on the shots and on that +location's chances in 36, and a pristine front table has only four distinct +shares - 7, 5, 4 and 1 - so eight locations cost four convolutions rather than +eight, and a second candidate with the same volley costs 0.6 us instead of 19. +A whittled-down target has fewer distinct shares still, because the destroyed +locations' rolls fold into the ones behind them. + +Nothing here enumerates critical slots or recomputes battle value per location. +That would be roughly 20 candidates x 8 locations x a battle value pass, and it +is what the capped-candidate invariant exists to prevent; what each location +holds is one scalar off the observation instead. + Fire allocation measures `sds_core::ev`, not a stand-in. It is 48 shots against 8 targets but only ~10 distinct `(rack, packet damage, to-hit)` triples, so the memo answers 4838 of 4848 asks and the convolution is sparse-into-dense: 12 diff --git a/plan/ev.md b/plan/ev.md index 8612486..c5970d0 100644 --- a/plan/ev.md +++ b/plan/ev.md @@ -24,15 +24,20 @@ The calculator lives in `crates/sds-core/src/ev.rs`. Nothing calls it yet. `(rack, per-missile damage, to-hit, cluster modifier)` - [x] Query it for mean, p50, p90, `P(>=20)` (the piloting-roll cliff) - [x] Carry expected packet count alongside, for punch-versus-sandblast -- [ ] `P(kill)` and `P(breach)`. `at_least(threshold)` is the whole of the - arithmetic; the thresholds wait on per-location armour, which the - observation does not carry +- [x] `P(kill)` and `P(breach)`, against the location the damage reaches rather + than the target's total. `at_least(threshold)` is still the whole of the + arithmetic; what was missing was not the threshold but the distribution to + apply it to. `EvCache::location_allocation` carries the per-location + compound distribution, and `hitloc` supplies the shares. See + `plan/features.md` for the size of the error this removed - [ ] Feed AMS into the cluster modifier. The key has the field and the roll applies it; nothing on the wire says a target mounts AMS - [ ] Force-level greedy allocation against the concave value curve. Waits on force-level assignment machinery, which does not exist yet. Greedy on a submodular objective is within 1-1/e; any search over subsets is exponential -- [ ] **Declare high-concentration weapons first.** Declaration order is +- [ ] **Declare high-concentration weapons first.** Now measurable: `p_breach` + says whether a location is expected to open, which is the thing a later + cluster weapon would be firing through. Declaration order is resolution order in MegaMek and nothing sorts it, so a PPC can open a location for the cluster weapons behind it. Free, once there is an allocation to order diff --git a/plan/features.md b/plan/features.md index 69059f6..e7eab9a 100644 --- a/plan/features.md +++ b/plan/features.md @@ -49,6 +49,15 @@ whether it can be used for learning. weight fitted against it came from eight rows; `waste_above` against the mission-kill threshold is the fixed form, and `thin_columns` in `sds/train.py` is the guard that would have caught it +- [x] Firing, per location: `p_kill`, `p_mission_kill` and `overkill` measured + against the location the damage reaches rather than against the target's + total, plus `p_breach` and `value_destroyed`. See `hitloc` below +- [ ] Retire `weapon_concentration` once `p_breach` has a fitted weight. It + compares the average packet against the biggest packet in the volley, + which is a statement about the guns; `p_breach` asks the same question of + the target and answers it from the distribution. Both are kept for now + because redefining a feature under its own name would silently change what + every weight already recorded for it means - [x] Target: `target_health`, `target_breach`, `target_threat`, `target_skill` - [x] Target, worth: `target_original_bv` and `target_tonnage` as `rank`, because raw BV and raw tonnage have no domain a rule gives and the question is @@ -74,6 +83,41 @@ whether it can be used for learning. - [ ] Interaction features, budgeted and named: `weapon_concentration x target_breach` first - now measurable against a basis the bot scores with +## Damage lands somewhere + +Three features compared the *whole volley's* distribution against a *single +location's* health, as if every point arrived in the same place. The centre +torso takes 7 rolls in 36 and a leg takes 4, so `p_kill` was overconfident - and +not by a constant. A threshold needing `k` packets in one place falls off like +`share^k`, so measured against a 20-point centre torso two AC/20s were 3.5x +overconfident, three PPCs 11.7x and four LRM-20s over a thousand times. No +fitted weight can absorb an error that swings with the shape of the volley, and +a fit correcting for one pushes the feature negative - which is what `p_kill` +drew across 17 overnight runs. + +`crates/sds-core/src/hitloc.rs` holds the model. Tables and the transfer chain +come out of MegaMek through `bridge/sds/SdsHitLocations.java`, never a rulebook. + +- [x] Hit tables for front, left, right and rear, read out of + `Mek.rollHitLocation` with the dice pinned through `Compute.setRNG` +- [x] Damage that would land on a destroyed location transfers inward, from + `Mek.getTransferLocation` and `Mek.getDependentLocation`. The derived + table is a function of `(side, destroyed set)` and there are 36 reachable + keys, so it is memoised rather than rebuilt per candidate +- [x] Per-location damage as a compound distribution: the packets arriving at a + location are binomial in the packets that landed, memoised on the shots + and the location's share in thirty-sixths +- [x] One scalar per location for what is inside it, from the contents summary + the wire now carries. Not a critical model: no slot enumeration, no + per-crit battle value, which would be a battle value pass per location per + candidate and is what the capped-candidate invariant exists to prevent +- [ ] Punch, kick and the quadruped front table. Different tables, and the bot + fires none of them yet +- [ ] A joint over two locations, if "either leg" ever needs to be exact. The + marginals are independent today, which is a lower bound on the union + because the locations are negatively correlated - a packet on the left leg + is not on the right. The correction is small and points the safe way + ## What the migration cost Movement candidates are measured against a *projected* volley from the hex they diff --git a/plan/harness.md b/plan/harness.md index 102f19c..fa0dfd9 100644 --- a/plan/harness.md +++ b/plan/harness.md @@ -115,6 +115,52 @@ one identified, benign cause rather than that the rate is measured. The round-limit games are excluded by the criterion by design, and both are legitimate. Neither is the pilot fault and neither is River Delta. +## `--bot` silently substitutes a different bot + +`sds one` and `sds bench` default `--bot` to `python3 /work/bots/random_bot.py`. +Omit the flag and the run measures the random bot and reports the result in the +same shape as any other, with the sds seat named `sds`. + +The worked example: a 24-game bench meant to measure a feature change came back +`princess 17, sds 0, 100.0%`. Nothing errored, nothing warned, and the number is +a perfectly plausible loss for a bot that had just been changed - which is the +worst possible shape for a wrong answer. It took two people and several hours to +catch, and only because a second, unrelated number looked wrong. The same +default also explains a phantom "a quarter of movement proposals are illegal", +which was the random bot proposing at random. + +`baseline.json` records the bot faithfully, so the evidence was there to read. +That is not enough: a default nobody meant to use should not be quotable. + +- [ ] Make the bot explicit. Any of: require `--bot`; print the bot, the seat + and the scenario at the top of every run; or refuse to write a baseline + whose bot is the random one without a flag saying so. Printing it is the + cheapest and would have caught this in the first line of output + +## Selecting work by pattern takes other people's work + +Three instances of one mistake, all of them on this machine and all of them +inside a week: + +- `sds clean` kills every match container named `sds-*`, including a benchmark + somebody else started. CLAUDE.md names it because it has already cost one. +- `pkill -f "sds.cli bench"` stops a bench, and stops every other agent's bench + at the same time. Used to stop a run that was measuring the wrong thing; + nothing of anyone else's was running, which was luck rather than care. +- `git worktree remove --force` over a hardcoded list of names took a 360-match + corpus and an agent's uncommitted branch with it, because the list was written + before the work started and not checked against `git worktree list` after. + +The generalisation: **anything that selects work by pattern rather than by +identity can take someone else's.** Container names are already branch-tagged +for exactly this reason, and that tagging only helps if the thing doing the +killing matches on the tag. + +- [ ] Give the destructive commands a way to name only their own work: a + `--mine` that matches the current branch's tag, or a required scope + argument, so that the safe form is also the short one. A rule that lives + only in a document is followed until somebody is in a hurry + ## A seed does not reproduce a match `2v2-clan-invasion-s101020`, seed 31, same binary, six runs: victories at rounds diff --git a/plan/protocol.md b/plan/protocol.md index 196efed..879c2b2 100644 --- a/plan/protocol.md +++ b/plan/protocol.md @@ -16,6 +16,13 @@ Adding a field is deliberate, with a reason in the commit message. - [x] **Per-location armour and structure** - blocks breach state, so it blocks punch-versus-sandblast allocation and [cripple](strategies.md). On the wire as `locations`, with rear armour on the three locations that have it +- [x] **Per-location contents** - armour says how hard a location is to break, + not what breaking it costs the target. A weapon count and their damage, + live ammunition and whether CASE contains it, engine, gyro and cockpit + slots, and an actuator count. One shared addition: [cripple](strategies.md) + needs it for arc-aligned targeting, and a mission kill needs it before + "all its guns are gone" can mean anything. Deliberately a summary and not + a critical slot listing - [x] **Ammo type per weapon** - per-missile damage varies by ammo; [ev](ev.md) cannot be right without it. On the wire as a weapon's `ammo`. `damagePerPacket` is still computed from the launcher alone; making it diff --git a/plan/training.md b/plan/training.md index ae4b1f6..bb398f2 100644 --- a/plan/training.md +++ b/plan/training.md @@ -508,6 +508,32 @@ weapon count is not the answer: Read MegaMek for the bin-to-weapon mapping rather than inferring it. +## A name is not a meaning + +The manifest records the sorted list of every feature name in a corpus, and +refuses a fit whose names do not line up. That catches a feature added or +removed. It does not catch the case this change is. + +`p_kill` kept its name and changed its quantity. It used to ask whether the +whole volley cleared the centre torso's health; it now asks whether the centre +torso or the head is actually destroyed, and the two differ by a factor of +between 3.5 and over a thousand depending on the shape of the volley. +`p_mission_kill` and `overkill` moved the same way. A manifest comparing name +lists sees `p_kill` on both sides and says the corpus is fine. + +**Every weight fitted against the old `p_kill` is meaningless**, including the +label comparison's `p_kill +0.225` at 91% agreement. Nothing in a corpus +recorded before this change says which `p_kill` wrote it, and nothing would +object. + +- [ ] Give the feature basis a version the manifest records, bumped when a + feature's meaning changes rather than when the set does - the same rule + `sds/epoch.py` applies to a match, applied to the basis. A name list + answers "are these the same columns"; this answers "do the columns mean + the same thing", and only the second question would have caught this + change. The epoch itself is the wrong place: it versions what a match is, + and the bot's own scoring is not that + ## Leaving the map: not yet, and why Fighting to destruction is wrong in the long run - a scenario can be won on diff --git a/sds/train.py b/sds/train.py index 8e015a2..aeccce6 100644 --- a/sds/train.py +++ b/sds/train.py @@ -81,11 +81,25 @@ WEIGHTS = REPO / "weights" # it crosses the aim penalty or the shutdown roll, ammo matters when the bins # are nearly out. Those features have clean signs. These do not, so they are # not claimed. See `plan/features.md` on heat breakpoints. +# +# `p_breach` and `value_destroyed` are claimed, and the test that admits them is +# the same one that threw the other two out: neither carries a cost to the +# shooter. Stripping a location to bare structure cannot leave us worse off - +# past it every further point rolls for a critical - and breaking what the +# target fights with is the objective rather than a proxy for it. Heat is a +# price paid, and the sign of a price depends on what it bought. +# +# `p_mission_kill` was an omission rather than a decision: it arrived with the +# `overkill` fix and this list was not revisited. Its sentence forces a sign as +# plainly as `p_kill`'s does. EXPECTED_SIGNS: dict[str, int] = { "expected_damage": +1, "p_kill": +1, + "p_mission_kill": +1, "p_psr_threshold": +1, "target_breach": +1, + "p_breach": +1, + "value_destroyed": +1, "overkill": -1, } -- 2.51.2