--- id: unit-types title: Vehicles and infantry play, not just Meks status: open dependsOn: [protocol, features] exitCriterion: > A force of Meks, vehicles and infantry is played without a unit type taking an inert default, and `sds spread` shows no feature reading a constant for one type and a range for another. --- # unit-types The Core Rules cover Meks, combat vehicles and infantry. Aerospace, dropships and anything that leaves the ground are out of scope and stay out. Today the basis is written for Meks and it shows in the vocabulary: `canStand` asks a `Mek` about its gyro, `level_tmm` reads a movement modifier from hexes moved and a jump, `heat_*` is four features about a scale a vehicle does not have, and `target_tonnage` now spans to 200 tons because the unit files say a combat vehicle reaches it - so the *denominator* already admits vehicles while the features around it do not. Nothing refuses a vehicle. That is the problem: a feature that is meaningless for a type reads a constant for it, and a constant within a decision teaches a difference-framed fit nothing. `sds spread` is the instrument - a column that moves for Meks and is flat for tanks is the signature, and it is invisible in a win rate. ## What it needs - A census first: what does MegaMek actually hand us for a vehicle and for infantry, and which of the named features return something meaningful. Probe it rather than reason about it. - A decision on features that cannot apply. A heat feature for a vehicle is not zero, it is undefined, and those are different things to a fit. - Movement: vehicles read a terrain column Meks do not, and infantry move by different rules again. `MoveHex` and the pathfinder have to answer for all three or say which they cannot. - Damage: infantry take damage by a different table entirely, so `hitloc` and the volley path need to know what they are shooting at. ## The two sweeps, and what the second one adds The zero-as-absence sweep above found **three** columns gated on `has_locations()`. A later sweep for columns that read *differently* against a vehicle - not just zero - found **five**, and the two extra are the interesting half: - **`p_psr_threshold`** reads a mean of **0.39** and a p95 of **0.99** against a vehicle, for the twenty-point piloting cliff. `Entity.canFall()` returns `false` for a combat vehicle, probed directly. It is a Mek rule applied to a machine MegaMek says it does not govern, and it errs **upward**. - **`overkill`** correlates **+0.9925** with `p_kill` against a vehicle, against +0.789 against a Mek: on a body where every location is lethal, "took more than it can absorb" and "was destroyed" are the same event one point apart. **Neither could have been found by the first sweep, and that is not a criticism of it.** A sweep for zeros cannot find a column that reads 0.39, and a sweep for absences cannot find one that has become redundant. The generalisation both support is that the pattern is not "zero means absence" - zero was the costume. It is **a value arrived at without the thing it claims to measure**, and it can err upward as easily as down. The full family, the discriminator that separates the real asymmetries from the artefacts, and the evidence for what a vehicle location is worth are below. ## What MegaMek answers for a combat vehicle Probed with `sds.SdsVehicleHits`, the vehicle companion to `SdsHitLocations`: the dice are pinned through `Compute.setRNG` and the tables are read out of `Tank.rollHitLocation`, the same code that resolves the shots. - **Seven locations, five of them reachable.** `BD FR RS LS RR TU FT`. No normal attack on any of the four arcs ever rolls the body or the second turret, so the model carries five and ignores two rather than refusing a target that reports them. - **No transfer chain.** `Tank.getTransferLocation` answers `LOC_DESTROYED` for every location - the same sentinel a Mek's head and centre torso give. So every location a vehicle has is a lethal one, and `hitloc::lethal` now derives both bodies' sets from that map instead of from a hand-written constant. - **No dependents and no rear armour pool.** A vehicle's four facings *are* four of its locations, so the arc picks the location rather than a second armour pool on the same one. - **Motive damage is a flag on the roll, not on the location.** Rolls 3, 4, 5 and 9 arrive with `EFFECT_VEHICLE_MOVE_DAMAGED` on every arc - 13 ways of 36 - while 6, 7 and 8 land on the same armour and do not. It cannot be modelled as a property of a location and is carried as its own share. - **Criticals are on the roll too:** 2 and 12 on every arc, and 8 from the left and right. - **`getAllowedPhysicalAttacks()` returns 1**, the same as a Mek's. - **A turretless design rolls a different table**, with 10-12 going to the arc's own location. Dumped rather than derived: MegaMek takes a different branch, and a redirect applied on our side would have been a guess that happened to agree. ### A defect in MegaMek's vehicle front table Front 5 and front 9 both return the **left** side. Every other arc separates the pair - `left` gives FR and RR, `rear` gives LS and RS - and the printed front table separates them as well. The consequence is that a vehicle shot from in front never takes a right-side hit and takes left-side hits at twice the rate. It was checked rather than assumed, because an answer this clean is usually the instrument: both rows consume exactly two d6 through the pinned source, and forcing a third die to each of its six values moves neither. `SIDE_LOC_MAPPING` is `{FRONT->FR, LEFT->LS, RIGHT->RS, REAR->RR}`, so the mapping is not the cause. The bot models what resolves the shots, so it models this, and `a_vehicles_front_table_hits_one_side_twice_and_that_is_megameks` is the test that says so the day it changes. ## The motive damage table, dumped `sds motive-damage` reads it out of `TWGameManager.vehicleMotiveDamage`, which mutates a `Tank` inside a game rather than being a pure function of a pinned die - so it needs a real scenario, a crewed and deployed entity and a manager. The scenario is reloaded for every cell, because motive damage accumulates and immobilisation is sticky. Recorded in `docs/livefire/motive-damage.txt`. **The outcome is a pure function of `roll + modifier`**, verified across all 44 cells of a four-modifier sweep: | 2d6 + modifier | result | |---|---| | 2-5 | no effect | | 6-7 | minor: +1 driving | | 8-9 | moderate: +2 driving, -1 MP | | 10-11 | heavy: +3 driving, half MP | | **12 or more** | **major: the vehicle is immobile** | And the modifier is the arc, from `Entity.getMotiveSideMod`: **front +0, rear +1, either side +2.** So the number `p_mission_kill` should carry, per motive-flagged hit: | arc | needs | P(immobilised) | |---|---|---| | front | 2d6 >= 12 | 1/36 = 0.028 | | rear | 2d6 >= 11 | 3/36 = 0.083 | | left or right | 2d6 >= 10 | 6/36 = 0.167 | ### What this says about the feature as it currently stands `p_mission_kill` for a vehicle currently reads `P(a motive-flagged hit lands)`, which measured a mean of 0.688 over 9,249 real firing candidates. The honest figure is that probability multiplied by the table above, so **the crippling end is overstated by roughly a factor of 36 from the front**. **The size of the error is not the worst part.** At 0.688 the column is non-zero on 99.9% of vehicle candidates and near-constant across them, so within a decision it barely discriminates between one tank and another while dominating every comparison against a Mek. Flat within a decision teaches a difference-framed fit nothing - the `sds spread` failure this epic exists to catch - so it degrades play *and* hides that it is doing so, in a direction no win rate would reveal. **Do not fit a weight on this column until it is calibrated.** The calibrated form is a genuine improvement rather than just a smaller number: it is six times more likely to cripple from the flank than from the front, so for the first time the feature says something about *where to shoot a tank from*, which the flat version could not express at all. ## Answered by jmm - **Model everything.** "We should just work on modeling for anything that isn't modeled yet." There is no tactical preference to preserve: a basis that favoured Meks because three columns collapsed was not a bot anybody chose, and once a vehicle scores honestly the preference is a decision somebody can make on evidence. - **`p_mission_kill` spans unit types on purpose.** "For ex this is one reason it's 'mission killed' and not 'shot leg off': tanks can be mobility-damaged in the same way." A Mek losing a leg and a vehicle taking motive damage are the same event at the level the bot reasons about, so motive damage belongs *in* that feature rather than beside it. That is what [`hitloc::LocationDamage::motive`] feeds, and it is why a vehicle's crippling set is empty rather than being a Mek's with different names. - **Two movement points per level is a hovercraft rule**, not a ground vehicle one - "level changes are harder (they want flat terrain)". So the tank-in-woods discrepancy `claude/move-legality` measured, 3 against 4, has some other cause and is still unexplained. The measurement was real; the rule attributed to it was not. - **A unit may make a move it cannot afford** if the destination is legal and it is the unit's whole turn. Given about infantry and generalised deliberately; it comes up for a Mek down to 1 walk MP as well. Not a vehicle rule and not this branch's code, recorded here so it is not lost. ## What an analysis pass actually costs **A pairwise sweep over `tactics-hub-mirrored` - 300 matches of decision logs, 3.6 GB on disk - runs at 147 MB resident.** Measured, not estimated, at 1m43s into a steady-state run over 80 logs. That is the opposite of what everyone assumed, including the person who warned about it and the person who took the warning. The reasoning that produced the wrong answer was "300 matches of JSON is a lot of memory", which is an inference from a data size to a memory cost with the shape of the program left out. **The general form, because it is not about memory.** The inference ran from a *data size* to a *resource cost* with the shape of the program left out. Any estimate of that form is a guess wearing arithmetic: the same input at the same size costs three orders of magnitude more or less depending on whether the program **holds** it or **walks past** it. Ask which before quoting a number, and if the answer is not known, measure rather than cite a precedent from a program of the other shape. **The shape is the whole thing.** `sds train` holds a corpus in memory and its OOM on this machine is real and recorded. A sweep *walks past* the corpus: `json.loads` per line, never a whole file, and accumulators that are about 2,600 floats per phase. Same input, same size, three orders of magnitude apart. The `train` precedent does not transfer, and citing it felt rigorous while being an assumption. So: the binding constraint on this machine during a run is **the match containers and their JVMs**, not analysis. Two containers plus their JVMs is what 28 of 30 GB looks like, and that is simply what a run costs. **Cap it anyway.** `CAP_MEM=2G capped python3 ...` is thirteen times headroom on a 147 MB job and buys nothing in expectation. It is still right: `MemoryMax` with `MemorySwapMax=0` makes the analysis structurally incapable of being the kernel's victim-pick, whatever the figure says. A guard that is unnecessary in expectation is exactly what belongs around somebody else's two hours of matches. ## Observability: to-hit could be logged, and is dropped one step before `expected_damage` reads 1.4x higher against a vehicle and the likeliest explanation - vehicles are easier to hit, so more of the volley lands - could not be established, because the decision log carries no to-hit. **It could.** The number exists all the way to the point candidates are built. `Shot.to_hit` is on the wire (`wire.rs`, emitted by `Observation.java` from `WeaponAttackAction.toHit`), and `ShotKey.to_hit` is what `ev::ShotPmf` uses to compute every damage distribution. What drops it is the menu: candidates are reduced to `Vec<(String, FeatureVector)>` - a label and a phi - before anything is written, so a quantity that is not a feature cannot survive to the log. So carrying it is widening that tuple, not plumbing a new signal. A mean or minimum to-hit per candidate would answer this question and is a local change. **This is the second question tonight a corpus should have been able to answer and could not** - the first being which arc a shot was taken from, which needed `examples/sidetable.rs` to reconstruct from positions that *were* logged. The pattern is the same: the log records what the fit consumes, so any question about *why* a candidate scored as it did has to be reconstructed or cannot be asked. ### For jmm: is the decision log a fit input or a diagnostic record? Nobody has decided, and **it is being decided by default every time somebody cannot answer a question.** Twice tonight, on this branch alone: - **Which arc a shot was taken from.** Reconstructible, but only because positions happened to be logged - it needed `examples/sidetable.rs` to dump `los::side_table` and an analysis that reads the table back. - **What the to-hit was.** Not reconstructible. It exists on the wire and drives every damage distribution, and the menu drops it one step before writing. The log currently records exactly what the fit consumes, which is a coherent design. The cost is that *why* a candidate scored as it did is not answerable from a corpus, and that question keeps arising - three of tonight's findings needed it and one could not be closed. **Proposed with the question rather than done quietly:** carry a per-candidate to-hit summary, which is widening `Vec<(String, FeatureVector)>` and nothing more. It is a local change, and it is the first instance of whatever the general answer turns out to be - so it is worth deciding the general question first rather than patching the one case and leaving the design unmade. ## A refused move is re-proposed identically, forever Found while measuring the illegal-move rate on the first pure-vehicle suite anyone has played. It is **not a vehicle defect**; vehicles only make it visible. **The pathology.** In `1v1-succession-wars-s4010001-vehicle`, one unit produced **38 refusals over rounds 3 to 40, and every one was the same step list** - `TURN_RIGHT, FORWARDS, FORWARDS, FORWARDS`. MegaMek refuses it, the unit takes the inert default and stands still, and next round it proposes exactly the same move again. For thirty-seven rounds. **The mechanism, confirmed in code by the movement agent.** `SdsClient.buildPath` returns `null` when `isMoveLegal` fails, and `askMovement` increments a tally and returns an empty path - **the refusal never crosses back to the Rust bot.** So the bot observes a world in which the unit simply did not move, and being deterministic given state and seed it proposes the same thing again. It is not that the bot ignores the refusal; it is never told. That also means the fix is protocol work rather than a patch: a client that substituted its own move would break "every order has exactly one possible author", which is the invariant the whole design rests on. It is with jmm. **And it is worse than "re-proposes identically".** In `2v2-jihad-s102033` a Bishamon was refused in **36 of 40 rounds and never moved at all** - ten distinct step lists, all landing in the same hexes, seventeen of them identical. So the failure is not one stuck proposal but a **whole candidate neighbourhood being refused**, with the bot cycling among refused variants for an entire match. In that corpus 29 of 52 refused units were refused more than once, so the concentration caution applies and does not explain it away. **It is general.** The mixed arms, which are mostly Meks, show the same shape at lower amplitude - worst repeat **11** on `origin/main` and **4** on this stack, against 63 and 67 refusals. So a Mek can livelock too; it is just rarer because a Mek is refused less often. **Vehicles amplify it** because their refusal rate is far higher. Over the whole 24-match arm, 301 movement decisions: | | answered | illegal | defaulted | |---|---|---|---| | tracked (n=175) | 47% | 28% | 25% | | hover (n=89) | 44% | **36%** | 20% | | wheeled (n=37) | **78%** | 16% | 5% | | **all** (n=301) | **50%** | **29%** | **21%** | **Half of all vehicle movement decisions are answered.** The other half split between a move MegaMek refused and no move proposed at all, and the 21% defaulted is its own number rather than part of the refusals. ### The comparison has to be within one suite An earlier revision compared this against "7-8% on the mostly-Mek mixed arms", which was the *blended* rate of a suite containing both. Split by the moving unit's own type, the mixed arms say: | | Mek | vehicle | |---|---|---| | this stack | 7.6% illegal (n=367) | **21.3%** illegal (n=183) | | `origin/main` | 5.0% illegal (n=403) | **28.5%** illegal (n=151) | So **within a single suite the gap is about four to six times**, not the thirty the blended figure suggested. **And the suite matters more than the bot.** The movement agent measured Meks near `main` at **0.22%** refused over `tactics-hub-mirrored` (20 of 9,109), against **5.0%** for `origin/main`'s own bot on the mixed suite here - a twenty-fold difference with the same bot and different scenarios. Whatever drives that is a property of the boards or forces these generators draw, and it means a refusal rate quoted without its suite is not a number. The vehicle figures above are only comparable to the Mek figures **beside** them. That elevated baseline is consistent with `claude/move-legality` + PR #313, which measured that a Mek pays one movement point per level and everything else pays two, and which is not on `main` - so this bot under-prices every vehicle level change. **Wheeled being the least affected** at 16% is worth a look by whoever picks that up: #313 added a wheeled column nobody had, and wheeled is the motion doing best here. ### The two defects, separated by the repeat histogram 87 refusals over 42 distinct `(unit, step-list)` pairs: | repeats | distinct proposals | |---|---| | 38 | 1 | | 4 | 1 | | 2 | 5 | | 1 | 35 | **One proposal accounts for 38 of the 87 refusals - 44% of them.** The other 49 are spread over 41 distinct proposals, 35 of which were refused exactly once. That is the clean separation: a long tail of one-off refusals is what a pricing error looks like, and a single proposal repeated 38 times is the livelock. Both are present, and neither figure would be visible in the other's aggregate. ### Correcting a number this file carried An earlier revision of this section said refusals ran "62% ... 38% for tracked, 38% for hover and 14% for wheeled". Those came from the first four matches of the arm, before it finished. **The full 24-match figures are the table above** - 29% overall rather than 62%. The partial sample was dominated by the livelocked scenario, which is exactly the hazard of reading an arm while it is still running. The finding survived; the magnitudes did not. **The two are separate defects and want separate fixes.** #313 makes the bot propose fewer illegal moves; it does nothing about proposing the same illegal move forever once it does. ### A wrong turn worth recording The first cut of this analysis grouped refusals by board and read the result as confirmation of #313: Desert Mountain 95%, Rolling Hills 75%, City Hills 33%, flat boards 0%. Mountains and hills at the top of a list about level changes is a very persuasive picture. **It was pattern-matching on map names.** Parsing the actual `.board` files and correlating measured elevation range against refusal rate gives **+0.031 over 13 boards** - no relationship at all. Badlands 1 has the *largest* relief on the suite and refused nothing; City Hills Residential has a relief of 1 and refused 46%. The board grouping was a proxy, and the thing it proxied for turned out to be one scenario with a livelocked unit contributing 38 of about 100 decisions. Nearly reported as independent confirmation of another agent's PR. The check that caught it was measuring the quantity the names stood for. ## Rule: a synthetic object gets a synthetic answer **Four times in one night, across the agents working here.** Written as a rule because it has stopped being an anecdote: - A bare `new Tank()` reported every location as not-destroyed immediately after being destroyed - an uninitialised entity, not a result. It is why this epic answers lethality from the transfer map instead. - An uncrewed entity answered zero hits where a crewed one answered four. - `TWGameManager.vehicleMotiveDamage` needed a real game, a real crewed entity and a real manager before it would answer at all. - `applyCriticalHit` threw on a `CriticalSlot` built by hand for `WEAPON_DESTROYED`, `FUEL_TANK` and `TURRET_DESTROYED`, because the slot names no real mount. **The rule.** When probing MegaMek, start from a real entity in a real game - loaded from a scenario, crewed, deployed - and only fall back to a constructed object once the real one has answered. The reverse order costs a probe, a plausible wrong answer, and the time spent believing it. The tell is an answer that is uniformly empty: all-`false`, all-zero, or a throw naming something the object never had. **And the corollary, which cost the most time tonight:** an all-empty answer from a synthetic object is *not evidence of absence*. It is evidence the question was not asked. ## Recorded negatives: why two pulls on this branch have no number Both pulls from this work - `3mubgtasnuy2z` and `3mubj6zsl4n2a`, created ninety minutes apart - show `?` where tangled.org's number goes, while other pulls number immediately. Two hypotheses were tested and **both are refuted**, written down so nobody spends time on them a third time: - **Not the missing `Change-Id` trailer.** These commits carry none, but `claude/test-threads-doc` carries none and is #306, and `claude/move-legality` carries none across six commits and is #287. `claude/kill-my-containers` does carry one and is #310, so both combinations of (has trailer, is numbered) occur. - **Not patch size.** `claude/expected-crits` is 25 commits and is #270; these are 16 and 26. - **Not indexer backlog.** #310 was created *after* the first of these and was numbered immediately. The records are valid and readable by record key, and **neither pull should be re-created**: a retried `pr create` can leave a duplicate record on the PDS, and then one branch has two pulls. This is jmm's tooling and his account, and it is with him. ## `pgrep` matches the shell that launched it Four independent rediscoveries of this in one night, across three agents. A watcher looking for its own job finds its own `bash -c` wrapper, because the pattern is in that wrapper's command line, and reports the job still running after it has finished - or, worse, `pkill -f` kills the shell that issued it. The fix each of them arrived at separately: **match the process, not the command line.** `ps -C python3 -o cmd` lists only real Python processes and cannot include the shell; `pgrep -f pairsweep` cannot tell them apart. Anything that polls for its own completion, or kills by pattern, wants the first form. ## Measuring an arm Two rules, both bought rather than reasoned: - **Smoke-test one game before committing an arm to 24.** Run one, read the decision log, and confirm the seat answered with candidates - `outcomes {'answered': 32}` and 9,617 candidates is a working bot; `{'defaulted': 162}` and zero candidates is a seat that never started. A whole baseline arm was lost to the second, and it printed weapon tables and a win rate while being a benchmark of the inert default. - **Report the per-scenario breakdown even when the pooled result is clean.** At the sample sizes a suite of this size gives, one scenario can carry an entire difference, and a single pooled share hides it. A summary standing in for the thing it summarises is the same failure as every other one in this file. ## For jmm: what "actual tanks, real combat vehicles" should mean The pool filter is `unit_type = 'Tank'` plus canon, valid, 20-100 tons, BV above zero, at least two weapons, and a motion type from a list. **The motion list is the question**, and it is currently `Tracked, Wheeled, Hover`. Here is what each choice actually admits, in the succession-wars era, so the decision is over a list rather than an abstraction. ### In the pool today: 256 designs | motion | designs | heaviest examples | |---|---|---| | Tracked | 139 | Mars Assault Vehicle (100t), Behemoth Heavy Tank (100t), Alacorn Mk II (95t) | | Hover | 70 | Condor Heavy Hover Tank (50t), Drillson Heavy Hover Tank (50t), Harrier (50t) | | Wheeled | 47 | Giant Guardian (80t), Gallant Urban Assault Tank (70t), Ishtar Fire Support (65t) | 20 to 100 tons, mean 52. Roles: Striker 61, Missile Boat 52, Scout 38, Juggernaut 32, Brawler 29, Sniper 24, Ambusher 10, Skirmisher 4, none 6. ### Excluded today, and worth confirming - **Naval (4)** and **Hydrofoil (3)** - Mauna Kea Command Vessel, Monitor Naval Vessel, Sea Skimmer. They need water to exist on and the boards are mostly land. Excluded, and this seems clearly right. - **WiGE** and **Submarine** - excluded for the same reason plus rules nothing here implements. - **VTOL** is a different `unit_type` and never enters. `hitloc` refuses one explicitly: it reports every abbreviation a turreted tank has plus a rotor, and reading one as a tank is the partial-match hazard that made a quadruped look like a biped. ### Decided and built, subject to jmm overriding any of it Built on the recommendations below rather than blocking on an answer; each is a filter change and a regenerate, not rework. **What was actually excluded is seven named utility designs** - three Coolant Trucks, the R10 coolant ICV, two MASH Trucks and the Engineering Vehicle - and nothing else. Pools are now succession-wars 218, clan-invasion 60, jihad 176. **The APC exclusion was not built, and that is a change from the recommendation.** Every way of expressing "it is a transport" turned out to be wrong when the list it returned was read: - **Not a troop bay.** 58 designs carry `troopspace`, and they include the Goblin Medium Tank, the Turhan Urban Combat Vehicle, the Prowler and the Skulker Scout Tank - combat vehicles that happen to carry a squad. - **Not firepower.** The plain **Vedette Medium Tank**, the generic Inner Sphere line tank, reads 7.0 - below the Engineering Vehicle's neighbours. Any threshold that drops a MASH Truck drops the Vedette. - **Not role.** The Coolant Truck 135-K is filed as `Ambusher`. So the utility vehicles are named instead, which is small and reviewable. The toothless end of the APC question is bounded and can be dropped in one word if jmm wants it: the twelve 20-ton `Heavy Hover/Tracked/Wheeled APC` variants at firepower 4 to 12, and the ten Badger Tracked Transports at 30 tons. They are still in. No committed scenario contains an excluded design, so the suites and the measurement in `comparison.md` are unaffected by this filter. ### The three questions, and the recommendations they were built on Read against "actual tanks mainly, like real combat vehicles" and "skip odd or rare equipment types", which points at the plain cases throughout. 1. **Hover: keep them in?** *Recommend: in.* 70 of 256. A Condor and a Drillson are plain combat vehicles that anybody would call tanks, not odd equipment. The one caveat is that their movement is the least tank-like - the level-change charge `claude/move-legality` measured turned out to be a hovercraft rule rather than a ground-vehicle one, and that is still not fully settled - so if the answer is "in", movement is where it will show. 2. **Transports and APCs: drop them?** *Recommend: out.* Twelve are in the pool on the two-weapon minimum - the Badger Tracked Transport family, the Heavy Hover APC - and several carry exactly the two the filter demands. A Badger with two machine guns in a line battle measures the harness rather than the bot, which is the argument that already excludes IndustrialMeks from the Mek pool. Raising the weapon minimum for vehicles drops them. 3. **Is 20 tons the right floor?** *Recommend: keep it.* Inherited from the Mek pool. Combat vehicles go lighter, and a very light one dies to a single volley - a real machine, but a thin benchmark row, and not worth widening the pool for while the question is whether tanks work at all. Nothing here is changed until answered: both pool defects on this branch so far came from trusting a filter rather than reading what it returned. ## A prediction for the calibrated motive column, registered before it is built The calibrated `p_mission_kill` is six times more likely to cripple from the flank than head-on - 6/36 against 1/36 - where the flat version was constant across arcs. So: 1. **The share of shots taken at a vehicle's flank or rear rises relative to head-on**, on vehicle targets. 2. **The effect is larger on vehicle targets than on Mek targets.** A Mek's crippling locations are its legs and the hit tables do not favour a leg from the flank the way motive damage does, so the Mek arm is close to a control. 3. **Magnitude: small.** `p_mission_kill` is one weighted column among fifty-one and the arcs already differ for other reasons - rear armour is thinner, and `rear_arc_gain` and the side tables already push that way. I predict the flank-and-rear share of shots at vehicles moves by **under 10 points**, and I would not be surprised by 3. **If it does not move at all, that is a real finding** and not a failed change: it would say a single column cannot steer target selection against the rest of the basis, which is worth knowing before anybody fits a weight expecting it to. ### And a second prediction, about where the bot stands The same asymmetry should move *movement*, not only target selection, because the firing basis is measured from every candidate stand: a stand that puts the shooter on a tank's flank scores a `p_mission_kill` six times the one head-on. 4. **A unit closing on a vehicle arrives at its flank or rear more often than one closing on a Mek.** Predicted to be the weaker of the two effects and possibly invisible: the flank is already preferred for reasons that have nothing to do with motive damage - thinner rear armour, `rear_arc_gain`, the side hit tables - so this column is adding to a push that already exists rather than creating one. I predict **under 5 points** and would not be surprised by nothing measurable. ## The charge, probed: what a vehicle may do and what it would cost to wire Both halves of the rule confirmed against MegaMek, with a 100-ton Gürteltier and a 100-ton Atlas placed adjacent and both done moving: - **A vehicle may charge a Mek.** `ChargeAttackAction.toHit` answers 7 after a one-hex run, 8 at three hexes, 9 at six - `5 (base) + 2 (attacker ran) + target movement`. Damage scales with the run: 0, 20 and 50 points at one, three and six hexes, against 10 taken by the vehicle each time. A six-hex charge is a fifty-point attack, which is more than most vehicles' guns. - **A Mek may not charge a vehicle.** `IMPOSSIBLE: Target is not a 'Mek`, at every distance. MegaMek's rule, not ours. - The first refusal a vehicle gets is `Target must be done with movement`, which is a state precondition rather than a unit-type rule - so a charge is declared against a machine that has already committed. ### Why this is not a gate to flip `Observation.physicalsAt` gates on `shooter instanceof Mek` and emits punches and kicks as pseudo-weapons in the firing menu. **A charge cannot be carried that way**, and `SdsBattleArmor.rush` on `main` already says why: a punch costs a limb and is declared in the physical phase, while a charge costs a *move* - it is declared during movement, takes a `MovePath` ending in the target's hex, its damage scales with distance travelled, and it hurts the attacker. That is why neither charge nor death-from-above is on the wire today, for any unit type. So wiring it is a movement-space feature: the candidate generator has to offer "run into that hex and declare a charge", and price it as damage dealt against damage taken. It is the same work for a Mek charging a Mek, and it overlaps the branch that owns physicals. **It is also not on the acceptance test's path.** jmm's test is a Mek kicking a combat vehicle to death; a kick is a physical-phase attack, it already reaches the wire, there is no target-side unit-type gate, and `a_kick_against_a_vehicle_prices_through_the_same_table` pins that it prices correctly. The vehicle charge adds a capability vehicles are entitled to and blocks nothing. ## For the Advance tactic, on another branch A tank's vulnerability is **directional in a way a Mek's is not**. Motive damage is +0 from the front, +1 from the rear and +2 from either side, so the same flagged hit is six times more likely to immobilise a vehicle from its flank than head-on - and a vehicle's four armour facings are four separate locations, so there is no transfer to spread a concentrated flank attack. `Advance` is a closing tactic. **A bot that closes head-on at a tank arrives at its least vulnerable facing.** Not this epic's to act on, and the interaction is recorded here because the branch building `Advance` has no way to know it otherwise. ## The flank baseline, registered before the calibration runs Measured on the two arms already recorded, **neither of which had the calibration** - so this is the "before" the flank predictions are against. The arc is not recomputed by the analysis: `crates/sds-core/examples/sidetable.rs` dumps `los::side_table`, the function the bot itself used, and the analysis reads that table. A port would be a second instrument, and two instruments agreeing proves nothing about either. | run | target | n | front | left | right | rear | flank+rear share | |---|---|---|---|---|---|---|---| | after (this stack, flat motive) | Mek | 345 | 264 | 28 | 14 | 39 | 0.2348 | | after (this stack, flat motive) | vehicle | 94 | 63 | 14 | 4 | 13 | **0.3298** | | before (`origin/main`) | Mek | 382 | 287 | 19 | 24 | 52 | 0.2487 | | before (`origin/main`) | vehicle | 78 | 64 | 4 | 2 | 8 | **0.1795** | **The move from 0.1795 to 0.3298 is not what the prediction is about.** Neither arm had the calibration: that gap is honest scoring and the pools, which changed *whether a vehicle is worth shooting at all* and therefore which shots got taken. It is context for reading the next run, not a preview of the arc modifier's effect, and it must not be quoted as one. **Read it cautiously in its own right, too.** At 78 to 94 vehicle shots the standard error on a share near 0.3 is about 0.048, so even that gap is roughly two standard errors - suggestive, not established. What it does establish is that **the instrument works and the sample is thin**. 78 to 94 shots per arm is not enough to see a 5-point move - the prediction is one standard error at this sample size, so the run could not tell it being right from its being wrong. That is a design fault found before launching rather than after reading an uninformative null, which is what registering a magnitude is for. **So the run is two arms, in this order:** 1. **The vehicle-only suite first.** 18 scenarios where every target is a tank, which buys several times the vehicle shots for the same match budget. This is the arm that can answer the registered flank prediction. 2. **The mixed suite second**, because that is the question actually asked - tanks and Meks together against Princess, the realistic case. Vehicle-only answers the science; mixed answers "are we getting better". ## Who each arm plays, and what will be said about the result **All three arms are sds against Princess.** None is self-play - `sds bench` seats Princess opposite by default and the script passes no `--self-play`, which the seat files confirm (`sds_North` against `princess_South`, alternating). This is stated because a flank share measured against our own bot and one measured against Princess are different quantities, and a reader with only the suite and the bot in front of them would have no way to know which they had. ### The win rate already recorded, disclosed before the run The two mixed arms already on disk, computed from `results.json` by mapping each seat to its team rather than assuming the order: | arm | decided | sds | princess | win rate | |---|---|---|---|---| | this stack, mixed | 24/24 | 11 | 13 | **45.8%** | | `origin/main`, mixed | 24/24 | 14 | 10 | **58.3%** | **This branch's arm came out lower than the baseline's, and that is being said here rather than left for someone to find.** It is also uninformative: at 24 games the standard error on a win rate is about 10 points and on the *difference* about 14, so a 12.5-point gap is under one standard error. The harness prints the same conclusion in its own words - about 97 decided games to separate 10 points, 385 to separate 5. Neither number is evidence about this change in either direction. ### What will be said if arm 3 comes out badly, decided now - **"Poor" means below 50%**, and poor is the honest expectation: the bot is being asked to fight Princess with a unit type it could not score at all a day ago, on a suite drawn to include machines nobody has tuned against. - **Arm 3 will be reported as it comes out**, in the same table whether it is 40% or 60%, with the sample size beside it. - **No claim of improvement or regression will be made from arm 3 alone.** At 24 games nothing between roughly 30% and 70% is distinguishable from 50%, so the only honest reading of a single arm is "consistent with parity". - **No weight will be tuned to improve it.** Nothing on this branch or the one below it fits a weight, and a number that moved because a weight was nudged after seeing it would be worthless - especially with a `p_mission_kill` column whose calibration is one commit old. - If arm 3 is genuinely bad in a way the sample *can* carry - a win rate near zero, or a collapse in `answered` - that is a defect to find, not a figure to report, and the run stops there. ## What the Princess benchmark can and cannot attribute The run that measures this stack will be **a state-of-the-bot measurement, not an A/B of one change.** By the time it runs, three things will have landed together: vehicles scoring honestly at all, the pools that put them on the board, and the calibrated motive column. A win rate against Princess measures the three as a set. **So it cannot separate the calibration's effect from the pools'.** Do not report it as though it could, and do not let a later reader infer it. If the calibration needs attributing on its own, that is a paired run with the column switched and nothing else moved, and it is not this one. **The flank predictions are the part that is attributable**, and that is why they were registered rather than a win rate. Nothing else in the stack pushes on approach direction: the pools change *which* machines are on the board and the honest scoring changes *whether a vehicle is worth shooting*, but only the arc modifier says a tank is easier to cripple from the side. If the flank share moves, this column moved it; if it does not, that is a real finding about how much one column can steer against fifty others. ### What is already checked on the acceptance-test path jmm's test for the whole effort is a win where a Mek kicks a combat vehicle to death. Three things on that path were checked here rather than assumed: - **A vehicle can be a physical target.** `Observation.physicalsAt` gates on the *shooter* being a Mek and puts no unit-type gate on the target; legality comes from `KickAttackAction.toHit`, which accepts a tank. There is no hidden `instanceof Mek` on the target, which is the failure that would have been invisible until exactly this test. - **A kick against a vehicle prices exactly, not approximately.** `Tank.rollHitLocation` returns identical rows for `HIT_NORMAL`, `HIT_PUNCH` and `HIT_KICK` on all four arcs, so this repository using the normal table for a physical - a documented approximation for a Mek - is exact for a vehicle. - **The figures are sane.** A kick is one large packet, which is the shape a vehicle is least able to absorb: no transfer chain, and every location lethal, so a single packet through anywhere ends it. `a_kick_against_a_vehicle_prices_through_the_same_table` pins that a kick finishes a worn tank and does not finish a healthy one. # Next epic: vehicle columns need a scale of their own, not the Mek's ## The test, before any of the instances **A large ratio is not itself a defect.** The question is never "does this column read differently against a vehicle" - it should. The question is whether it reads differently **for a reason about the machine or a reason about our reader**. `p_kill` reads 25.9 times higher against a vehicle because a vehicle *is* that much easier to kill: every one of its locations is lethal, and that comes from MegaMek's own transfer map. A bot that prefers shooting vehicles for that reason is making a defensible tactical judgement, and "fix everything that differs" would break the one column telling the truth. **And the pattern behind all of it is not "zero means absence".** That was the costume, not the shape. Every other defect found on this branch was a zero standing in for something unmeasured - three scoring columns, five movement outputs, a bot binary the container could not see. Then `p_psr_threshold` turned up reading a **mean of 0.39 and a p95 of 0.99 for an event that cannot happen**. The real pattern is **a value arrived at without the thing it claims to measure**, and it can err upward just as easily as down. Looking only for zeros would have missed it. ## The family **The basis was built for a machine whose worth lives in its location contents, and a vehicle's does not.** That single fact has now produced five separate findings, and they are one family rather than five bugs. Measured over 10,533 real firing candidates from the recorded run, by target type: | column | Mek mean | vehicle mean | ratio | verdict | |---|---|---|---|---| | `p_kill` | 0.0048 | 0.1232 | **25.9x** | **real** | | `value_destroyed` | 0.0048 | 0.0797 | **17x** | **artefact** | | `overkill` | 0.0238 | 0.1117 | 4.7x | **real ratio, but the column stops being its own** | | `p_psr_threshold` | 0.1693 | 0.3878 | 2.3x | **artefact** | | `expected_damage` | 0.2248 | 0.3219 | 1.4x | probably real, mechanism not established | | `p_mission_kill` | 0.0004 | 0.688 | ~1700x | **artefact, fixed on `claude/vehicle-motive`** | | the four `heat_*` | varies | ~0, flat | - | **sentinel** | **Telling the real ones from the artefacts is the whole job.** A large ratio is not itself a defect: `p_kill` reads 25.9 times higher against a vehicle because a vehicle *is* that much easier to kill - every one of its locations is lethal, and that comes from MegaMek's own transfer map rather than from an omission. A bot that prefers shooting vehicles for that reason is making a defensible tactical judgement. The artefacts are the ones where the number is high or low for a reason that is about our reader rather than about the machine. ## The artefacts, each verified against MegaMek **`value_destroyed`, 17x.** Its denominator is `Location::value()` - guns plus actuators plus 8 engine plus 8 gyro plus 20 cockpit - and those flags come from `Observation.contents` reading `Mek.SYSTEM_*` and `Mek.ACTUATOR_*` out of a location's critical slots. **A combat vehicle has no critical slots**: probed, `systemSlots=0` against an Atlas's 31, because MegaMek models a vehicle's criticals as unit-level results rather than as contents of a location. So vehicle value-at-stake is guns and ammunition alone, the denominator is a fraction of a Mek's, and because every location is lethal the ratio saturates - p95 0.92 against 0.03. **`p_psr_threshold`, 2.3x.** It is "the chance we hit the target hard enough to make it roll not to fall over", measured as the Mek twenty-point piloting cliff. **A vehicle cannot fall**: `Entity.canFall()` returns `false` for one and `true` for a Mek, probed directly. So this column reads a mean of 0.39 and a p95 of 0.99 for an event that cannot happen. It is the sharpest of the family because it is not an omission producing a wrong scale - it is a Mek rule applied to a machine MegaMek says it does not govern. **`overkill`, 4.7x - the ratio is real and the *column* is the problem.** A vehicle genuinely overflows more: fewer locations, smaller pools, so a volley spills past one sooner. But `overkill` is `P(some location takes more than it can absorb)` and `p_kill` is `P(some location is destroyed)`, and on a body where **every location is lethal those are the same event one point apart**. Measured: `corr(overkill, p_kill)` is **+0.9925 against a vehicle** and +0.789 against a Mek. So for a vehicle the column carries almost nothing `p_kill` does not already say, and a difference-framed fit would be putting two weights on one signal. That is not a wrong number - it is a column that stops being an independent one. This module has the precedent written down: *"If it ever finds them equal everywhere, one of them is redundant and should be deleted rather than fitted - which is what happened to `los_in`/`los_out` and to `overkill`."* It happened to `overkill` once already, for Meks. **`expected_damage`, 1.4x - probably real, and left honest.** The likeliest mechanism is that vehicles are easier to hit: no jumping, and many are slow, so more of the volley lands. That is a property of the machine rather than of the reader, which would make it real. It is **not established** - the to-hit is not carried per candidate in the decision log - and it is recorded as unexplained rather than waved through, because a 1.4x that turned out to be structural would be the same family again. **The `heat_*` family, flat.** `getHeatCapacity()` returns **999** for a combat vehicle - checked across 41 distinct designs, energy-armed and ballistic alike - which is MegaMek saying "no heat scale", not a capacity. Five of our columns divide by it, so they read about zero for every vehicle in every decision. The values are right and that is why nobody found it; what is wrong is that nothing distinguishes "no heat scale" from "a Mek running cold". Same shape as `UPSTREAM.md` #6, where `WeaponType.getDamage()`'s sentinel once reached the bot as real damage - a recurrence rather than a new bug. ## Evidence for the scale decision, probed `sds.SdsVehicleCriticals` applies each of MegaMek's fifteen vehicle criticals to a fresh crewed, deployed Gürteltier MBT through `TWGameManager.applyCriticalHit`, and reads what it cost. A fresh entity per critical, because they accumulate. Baseline: walk 3, run 5, six usable weapons, **battle value 2120**, 100 tons. | critical | walk | run | weapons | **battle value** | flags set | |---|---|---|---|---|---| | ENGINE | 3->0 | 5->0 | 6->5 | **2120->1264** (-40%) | engineHit, turretLocked | | TURRET_DESTROYED | - | - | - | **2120->1169** (-45%) | (threw part-way) | | AMMO | - | - | - | **2120->1840** (-13%) | - | | SENSOR | - | - | - | 2120 | sensors=1 | | CREW_KILLED | - | - | - | 2120 | crewHits=6, crewDead **false** | | DRIVER, COMMANDER, CREW_STUNNED, STABILIZER, WEAPON_JAM, TURRET_JAM, TURRET_LOCK, CARGO | - | - | - | 2120 | none observed | **Three of fifteen need more context than a synthetic critical slot carries.** `WEAPON_DESTROYED` threw `NoSuchElementException` and `FUEL_TANK` and `TURRET_DESTROYED` threw `NullPointerException` - they want a real slot naming a real mount, not one built by hand. That is the same answer as the motive table and the munition handlers, for the third time: **a synthetic object gets a synthetic answer, and the real path needs the real object.** Next time it should be the first thing tried. **Two results are anomalies rather than findings**, and are recorded as such rather than used: `CREW_KILLED` sets `crewHits=6` but leaves `crewDead` false and the unit alive, and `TURRET_LOCK` does not set `isTurretLocked` although an engine hit does. Both are probably the same missing-context problem. Neither should be leaned on until re-probed through a path that resolves damage properly. ### What this says about the shape **A vehicle's worth is unit-level with a location-level trigger, not location-level.** That is not a rescaling of the Mek model, it is a different shape. Every one of these criticals is a property of the machine - its engine, its crew, its sensors - reached *through* a location but not stored in one. So "what is this location worth" has no answer for a vehicle in the terms `Location::value` uses, and any number invented for it would be a number about our reader again. **`calculateBattleValue()` is the unit that does work.** It is MegaMek's own scalar for what a machine is worth, it responds to vehicle criticals in sensible proportions - engine 40% of the unit, turret 45%, ammunition 13% - and it applies to both body plans without translation. It is also **already on the wire**: `target_current_bv` and `target_original_bv` are existing features, so the observation carries it and nothing new has to be plumbed to try it. ### The ruling **Use MegaMek's per-component battle value**: the sum of `EquipmentType.getBV` over the equipment mounted in a location, carried as `component_bv`. The name is deliberate - it is the offensive half only and is not a location's battle value. Pricing a location with mobility in it needs a battle value calculator that can be asked about a hypothetical machine, which is Helm's job. It re-prices every Mek location as a side effect, so it is measured on Meks and not only on vehicles. ## What the epic has to decide Not "translate the Mek reader". There is nothing to translate: a vehicle's engine, crew and motive system do not live in a location, so there is no slot to read. **What a vehicle location is worth is a modelling decision**, and it needs the same treatment the motive column just got - probe what MegaMek does to a vehicle on a critical, then choose a scale and say so. **Until then, neither `value_destroyed` nor `p_mission_kill` should carry a fitted weight over a mixed suite**, and `p_psr_threshold` should not be read as meaningful against a vehicle at all. The column means two different things depending on what is being shot at, and a fit cannot see the difference. ## The first of the family, in detail: `value_destroyed` Measured over **436,416 real firing candidates** from the 48-match mixed suite (`corpora/20260831T070313Z-bench`, commit `dc336d5`), which carries the `component_bv` pricing above. Targets split by `movementMode`: 371,238 candidates at a biped, 65,178 at a tracked, wheeled or hover vehicle. | target | non-zero | mean | p95 | max | at the clamp | |---|---|---|---|---|---| | Mek | 93.6% | 0.0283 | 0.1584 | 1.0000 | 0.05% | | vehicle | 45.9% | 0.1965 | 0.9509 | 1.0000 | 3.58% | A vehicle reads **6.9 times** a Mek on the same column. Pricing a location by battle value took that from 17 times; what is left is a different thing and has a different cause. ### A unit dies once, and the column charged for it per location `worth_hundredths` prices a location whose loss ends the unit at the *whole unit*, and `expected_value_destroyed` summed `P(destroyed) x value` over every location. On a biped that is two terms. **A combat vehicle answers `lethal` on every location it has**, so five terms each worth the whole machine were added together for one death. The size of it, from the same corpus - post-fix a vehicle's reading is exactly its `p_kill`, so `value_destroyed / p_kill` is the double count: | target | median | mean | p95 | max | |---|---|---|---|---| | vehicle | 1.001 | 1.018 | 1.073 | 1.359 | Small in the middle and up to 36% at the top, and it pushed **3.58% of vehicle candidates past 1.0**, where `Score::fraction` clamps. A clamped column does not read as wrong, it reads as certain, and every volley past the line ordered identically - among exactly the volleys that kill. `LocationDamage::any_lethal` is the union and is what `p_kill` already reads, so the two now come from one place. The remaining locations are counted only in the worlds where the machine is not already dead, which bounds the figure by the unit's battle value: `component_bv` is the offensive half of a location and its sum is strictly below the whole. Vehicle mean falls from 0.1965 to 0.1911, Mek from 0.0283 to a little under it. ### What the column has left to say about a vehicle **Nothing of its own.** Every vehicle location is lethal and worth the whole unit, so the non-lethal term is empty and `value_destroyed / at_stake` is identically `p_kill`. The corpus says so before the fix does: the median ratio is 1.001. That is not a defect in the reader. It is what a location's worth *is* when the location has no separable worth: a vehicle's engine, crew and motive system are reached through a location and not stored in one. The figure that would give the column something independent is `BV(intact) - BV(without this location)`, with the mobility term in it, and it needs a battle value calculator that can be asked about a hypothetical machine. That is Helm's, and the wire stays a plain per-location number so the producer can be swapped with no change on this side. **The hold on fitting `value_destroyed` over a mixed suite stands**, for a narrower reason than before: not that the column is inflated on vehicles, but that it is a copy of `p_kill` on them and an independent measurement on Meks, so one weight means two things. `p_mission_kill` is a separate case and no longer this one: a vehicle's reading comes from `motive_immobilised` rather than from a location's worth. ### No decision moves today `value_destroyed` carries no weight in `weights/hand-authored.json`, and `sds counterfactual value_destroyed --phase firing` reports 162 of 162 live menus already scored with the column at nought. `Engage` prices the column, so `retune` scales it, and scaling nought is nought. **A correction to a column at zero weight cannot move an argmax**, and this one does not claim to. What it changes is what the column will read when something does weigh it. ## Also swept: three things, none urgent **1. A vehicle's heat capacity is a sentinel, and five columns divide by it.** `Entity.getHeatCapacity()` returns **999** for a combat vehicle - checked across 41 distinct designs from the vehicle suite, energy-armed and ballistic alike, and it is 999 for every one. It is MegaMek saying "this machine has no heat scale", not a capacity. The bot divides by it. `Unit::heat_load` is `heat / capacity`, `situation.rs` takes `heat / capacity.max(1)`, and `end_of_turn_heat` subtracts it, so every heat column reads about zero for a vehicle and `heat_shutdown_risk` and friends read exactly zero. **The values are right** - a vehicle cannot overheat, so its heat risk genuinely is nought - which is exactly why nobody would notice. What is wrong is that they are **flat within every decision for a vehicle**, and nothing distinguishes "this machine has no heat scale" from "this Mek is running cold". That is this epic's own signature and the motive column's failure one level over. It is also the same shape as `UPSTREAM.md` #6, where `WeaponType.getDamage()` returns a sentinel that once reached the bot as real damage. Here the arithmetic happens to come out right; that is luck, not design. **2. The vehicle suite over-draws hover, within noise.** 84 units drawn against the pool's proportions: | motion | pool | drawn | |---|---|---| | Tracked | 52.8% | 52.4% | | Hover | 27.3% | **35.7%** | | Wheeled | 19.9% | **11.9%** | At 84 units the standard error on a share near 0.27 is about 0.048, so both gaps are under two standard errors and consistent with sampling. Weight classes spread sensibly - 26 light, 18 medium, 18 heavy, 22 assault, mean 56 tons. So the draw is not broken. It is worth knowing anyway, because hover is the motion whose movement is least tank-like and the one still carrying an open rules question: a suite that is a third hover leans on exactly that. **3. One of this branch's own tests was true of its fixture, not of the rule.** `a_vehicles_mission_kill_is_motive_and_a_meks_is_not` asserted `immobilised < entry / 10`. From the front a major result needs 12 on 2d6 and the two figures are about 36 apart, so it passed; from a flank it needs 10 and they are about 6 apart, so the same assertion would have failed on a side shot for no reason a reader could see. Now checked against the arc's own odds. The one remaining fixed bound is labelled in the test as a regression bound rather than a rule. ## Open questions for jmm - **A destroyed turret still takes rolls 10-12.** The probe says MegaMek does not redirect them, so the model parks that share on a dead location where it does nothing. That was probed on a bare `Tank` and the same probe's lethality half returned all-`false`, which is too clean to trust - so this one is *recorded, not settled*. It wants a real design in a real game. - **Does a hovercraft pay the level charge on a drop as well as a climb?** The ground-vehicle half of this is answered above; the climb-versus-drop half is not. It wants a probe of what MegaMek does rather than a rule read off the answer that was given. - **Why does a tank in woods cost us 3 and MegaMek 4?** The hovercraft rule was the explanation offered for it and is not the explanation, so the measurement is now unattached. `claude/move-legality`'s, not this branch's. - Infantry. Nothing here touches them, and they take damage by a different table again - `hitloc` now has somewhere to put that, which it did not before. - ProtoMeks: the unit files carry 86 of them, between 2 and 15 tons. Core Rules or not? ## What the basis does to a target it cannot model Swept the whole feature basis for figures that can read zero, and asked of each whether zero is a measurement or an absence. **29 sites, and all but three are measurements.** The three that are not share one cause, and it is this epic's. The discriminator that matters is whether the guard's input varies *within* one decision. A guard on "no enemies on the board", "no friends", "this machine carries no ammunition" or "this machine cannot move" reads the same for every candidate on the menu, so it cannot move an argmax - it shifts the whole vector and nothing else. Most of the basis is that kind: - `los_in`, `los_out`, `rear_arc_exposure`, `rear_arc_gain`, `cover_quality`, `range_spread`, `range_band_fit`, `overlook` - no enemy is a real zero share of the enemy force. - `friend_support`, `cohesion` - the same for friends. - `ammo_spent`, `heat_ammo_explosion_risk`, `heat_mp_penalty` - properties of the shooter, identical across its own candidates. - `concentration`, `arc_spread` - no assignment to be concentrated. - `elevation_gain` already gets this right and says so: no enemy returns the *middle* of its domain, not zero, "the value a flat board would give". So does `Options::UNSEARCHED`, which reads a hex's mobility as 1.0 rather than 0.0 when no search was run. **Three vary within a decision, and all three are gated on `has_locations()`** - whether the target is a biped Mek the hit tables describe. The same volley at a Mek and at a tank, measured: | | expected damage | `p_breach` | `expected_criticals` | `value_destroyed` | |---|---|---|---|---| | Mek | 13.8 | 0.258 | 0.018 | 0.004 | | tank | 13.8 | **0.000** | **0.000** | **0.000** | The shot is identical and every column that comes off a location collapses. All three carry positive weights, so **the basis prefers shooting the Mek, for a reason that is a modelling gap rather than a tactical fact.** `p_kill` is the exception and shows what the fix looks like: it falls back to a whole-unit threshold, which is wrong in the way `hitloc` documents but is *honestly* wrong rather than silently zero. ### Why the two that got it right, got it right `elevation_gain` and `Options::UNSEARCHED` both map an absence to a **benign value** rather than to zero, and both could only do that because somebody asked what absence meant *at the point of writing the guard*. Every one of the sites that got it wrong had a plausible local reason to return zero - no enemies is genuinely no share of the enemy force, no bins genuinely spends no rounds - and zero is a legal value in all of these domains, so an absence can wear it without anything failing. That is the whole difficulty: nothing breaks, and the figure is consumed as a preference. ## How often the preference is actually exercised Sized on the recorded corpora rather than argued about. The trigger is narrower than "not a Mek": a **quadruped** is a `Mek` with eight locations, and it reports `HD, CT, RT, LT, FRL, FLL, RRL, RLL` - so `hitloc::index_of` recognises four of the eight and `has_locations` refuses it. `hitloc`'s own module note predicted exactly this case. Measured over `tactics-hub-mirrored`, 300 matches and **8,699 firing decisions**, by reading the menus out of the decision logs and asking MegaMek what each of the 380 chassis the bot shot at actually is: | | | |---|---| | firing decisions | 8,699 | | with a quadruped on the menu | **211 (2.43%)** | | offering a quadruped *and* a biped | **178 (2.05%)** | | matches involving one at all | 14 of 300 | | chassis that are quadrupeds | 7 of 380 | The other two corpora with data, `corrected-basis-explore` and `hitloc-handauthored`, are **0.00%**: their suites are biped-only, and the two scenarios that carry a quadruped today were recorded before it was in them. **2% is a correctness note rather than a live bias**, which downgrades the question rather than closing it. For what it is worth and not as a verdict: when a quadruped was on the menu the bot chose it 128 times and chose a biped 77, so the three collapsed columns are a thumb on the scale rather than a veto - but whether those choices were right is not something this count can say, and it is not what it was measured to answer. - [x] Size the exposure. 2.05% of firing decisions in the largest corpus, 0% in the others - [x] Sweep the basis and classify. Pinned by `a_target_the_tables_do_not_describe_reads_zero_and_should_not`, which fails the day vehicles get a hit table - which is the point - [ ] Until they do, decide whether the absence should read as the benign value the way `Options::UNSEARCHED` does, or stay at zero. **It is not obviously wrong to prefer the target we can reason about**, and that is why this is a question rather than a fix: a bot that shoots what it understands is defensible, a bot that does so because three columns silently collapsed is not. Worth deciding with jmm rather than patching - [ ] `target_breach` was checked and is fine: `breach_threshold` walks whatever locations the target has, so a tank gets a real answer. Its zero guard is for an observation older than per-location armour, which no live match produces ## `expected_criticals` against a vehicle: the mechanism question > **Both parts are implemented in this stack.** Part one - > `Projected::expected_criticals: Option`, gated on the body having a > critical model rather than on having locations - is in `(1/3)`. Part two, > `target_described`, is in `(3/3)` below. The analysis is kept because it is > the argument for the shape, and because it names the two options that were > ruled out and why. The vehicles work replaces `has_locations()` with `Body::classify()` and an optional `LocationDamage`, which means a naive rebase switches this column on for vehicles and computes them from the **Mek** critical table. A combat vehicle has `systemSlots = 0` against an Atlas's 31 - MegaMek models its criticals as unit-level results reached *through* a location rather than stored in one - so there is nothing per-location for that table to be about. ### What the code actually looks like, because it changes the answer The five columns do not fail the same way, and only two of them are mine to gate: | column | today, against a target the tables do not describe | |---|---| | `p_kill` | falls back to a whole-unit threshold - wrong, but *answers* | | `p_mission_kill` | explicit `0.0`, with a comment | | `expected_criticals` | explicit `0.0`, with a comment | | `value_destroyed` | **no gate at all** - zero because every `live[]` is false | | `at_stake` | the same, implicitly | That is the fact that decides the mechanism. **Making `Projected::expected_criticals` an `Option` is cheap** - three real construction sites and one test - **and it fixes one fifth of a shared defect**, while `value_destroyed` keeps returning a silent zero from arithmetic that has no gate to change. Five optional fields would be five independent decisions about what `None` means at a boundary that must yield a number anyway. ### The mechanism, if one is wanted A feature must produce a number for the fit, so "optional" cannot mean "absent" at the scoring boundary; it can only mean "imputed, and said so". The standard way to carry that into a *linear* model is the pair: impute a fixed value and add an **indicator column** recording that you imputed. One feature - `target_described`, 1.0 when the hit tables describe the target and 0.0 when they do not - rides on `Projected` like the rest. One entry in `for_each_feature!`, one vignette, one field. **No existing column changes signature**, older weight files lack it and contribute zero, and it covers all five at once rather than the one that happens to be in the rebase. Its honest limit: an indicator lets a linear fit apply a constant *offset* for those rows. It cannot apply a different *slope*, so it is better bookkeeping and not a better model. The real answer is columns that describe a vehicle's worth, which is this epic. ### The recommendation **(a) is out, and my own sweep is what rules it out.** That sweep lists `expected_criticals` at 0.018 for a Mek and 0.000 for a tank as one of exactly three sites where zero is an *absence* rather than a measurement. Recommending "leave vehicles excluded" would be recommending the defect I catalogued. Struck. **(b) is out on the measurement.** A Mek critical table against `systemSlots = 0` is a confident wrong number, and confidently wrong is worse than absent. **So (c), and the mechanism is two parts, because one of them alone is cosmetic.** **Part one: `Projected::expected_criticals: Option`.** Three real construction sites and one test. This is where "the producer could not compute it" gets stated, and it is what makes the case greppable rather than a comment. **Part two is the one that matters.** A feature must yield a number for the fit, so `Option` at the producer does not change a single figure the fit sees - it still imputes at the boundary, and imputing zero is numerically identical to today. To make the absence *distinguishable*, which is the standard this branch spent its time on, the vector has to carry the distinction: an indicator column, `target_described`, 1.0 when the hit tables describe the target and 0.0 when they do not. That pair is the textbook handling of a quantity that is undefined for some rows: impute a fixed value, and record that you imputed. The fit learns the average correction for those rows; a log reader can see why four columns are zero at once. **Build part two for the family, not for this column.** `value_destroyed` and `at_stake` have no gate to change - they are zero because every `live[]` is false - so an `Option` on my field fixes one fifth and leaves the other four exactly as silent. One indicator covers all five. **Cost.** Part one is small. Part two is one `for_each_feature!` entry, one field on `Projected`, wiring in both producers, and a **compulsory vignette**, which is the real expense - the docs build refuses a feature without one. No weight file changes: an older set lacks the column and contributes zero, so behaviour is identical until something is fitted against it. Neither part needs matches. **Its honest limit.** An indicator lets a linear fit apply a constant *offset* for those rows. It cannot apply a different *slope*, so it is better bookkeeping rather than a better model. Columns that describe a vehicle's worth are the real answer and they are this epic. - [ ] What is missing today is not a basis field but a **log line**: nothing in a decision says "these columns are zero because the tables do not describe this target". That is `observability`, costs no fit change, and would have made all five of these findable without a rebase - [ ] Then `target_described` as one column for the family, decided rather than merged, if the fit still wants it ## What the zero sweep could not have found Worth recording against the sweep rather than leaving it to be rediscovered. It swept 29 zero-reading sites and found 3 absences; the vehicles work has since found 5. The two it missed could not have appeared in it: - **`p_psr_threshold`** reads mean 0.39, p95 0.99, for a fall `Entity.canFall()` says a vehicle cannot take. It errs **upward**. Fixed: the wire carries `canFall`, and the column is `None` for a machine no piloting roll applies to. - **`heat_*`** reads a 999 sentinel as a capacity, which comes out near zero and correct **by luck**. Fixed: the wire carries `tracksHeat`, compared against `Entity.DOES_NOT_TRACK_HEAT` in `Observation.java` so no sentinel value is written on our side. The predicate was "reads zero", and that is the costume. The shape is **a value arrived at without the thing it claims to measure**, which can err upward as easily as down, and can be accidentally right. The discriminator survives - does the guard's input vary within one decision? - because it is about *which* absences distort a ranking rather than about how to find them. What needs replacing is the search. Three distinguishable shapes: 1. **An input is absent** and something imputes. Greppable, and what the sweep found. 2. **The quantity is undefined for this actor** and nothing checks. `canFall` is this: the input is present and fine, the *claim* is inapplicable. 3. **A sentinel is read as a value.** 999 capacity, `DAMAGE_BY_CLUSTER_TABLE`, `IArmorState::ARMOR_NA`. Only the first is a grep. The second needs each feature's claim tested against a taxonomy of actor kinds, which is an audit rather than a sweep - and is what the vehicles work is doing by hand. The third is greppable if the sentinels are enumerated, and MegaMek's are: they are the same family this branch met in `getDamage`, and they are worth listing once somewhere. ## `target_described`, and MegaMek's sentinels enumerated The indicator is built. `hitloc::describes` is the one predicate both producers ask - it was duplicated in `Volley::has_locations` before - and `firing::TargetDescribed` carries the answer into the vector as a flag. **It is not "is this a Mek", and it stopped being that under it.** The column was written when a combat vehicle was undescribed, and the body-plan work beneath it gave vehicles their own tables. `describes` is now `Body::classify(..).is_some()`, so a vehicle reads **1.00** and the columns this flags *measure* against one. What it excludes is a body plan nothing here rolls for: a quadruped, whose `FRL/FLL/RRL/RLL` no plan knows; a tripod, whose extra `CL` has no dump; and an observation carrying no locations. The worked figure, its note, the docs and the test that used a tank as the undescribed case were all rewritten to a quadruped for that reason. **And it does not explain every zero it sits beside.** `expected_criticals` is `None` for a vehicle *and* for an undescribed target, because a vehicle is described by a *hit* table and still has no *critical* model - MegaMek resolves its criticals as unit-level results and it reports `systemSlots = 0`. So a vehicle reads `target_described` 1.00 and `expected_criticals` zero, and nothing in the basis yet says why. **What it covers, named in the column's own docs so its purpose is legible:** `expected_criticals`, which is an `Option` and imputes zero; and `p_mission_kill`, `value_destroyed`, `at_stake` and `p_breach`, which have **no gate at all** and are zero because every `live` flag in the profile is false. That asymmetry is the argument for one indicator over four `Option` fields. `p_kill` is the exception and falls back to a whole-unit threshold, which is wrong in a documented way rather than silent. **What it does not cover, also in the docs, because someone will assume it did:** `p_psr_threshold` prices a fall `Entity.canFall()` says a vehicle cannot take and errs *upward* - an indicator cannot subtract a claim that should never have been made - and the `heat_*` group reads a 999 sentinel as a capacity, which comes out near zero and correct by luck. Both want the quantity not computed, not a flag saying it was. **Both are now done that way**, each carrying the answer on the wire rather than deriving it: `canFall` gates the cliff to `None`, and `tracksHeat` gates the heat scale. **The honest limit is in the column's doc comment**: a constant offset for those rows, not a different slope. Better bookkeeping, not a better model. The real answer is columns that describe a vehicle's worth. ### A file with a regeneration command is not a file to edit That is the whole rule, it costs nothing to apply, and it was learned the expensive way twice in one night. **When behaviour changes, the committed renders of that behaviour change with it, and nothing in the reasoning about the code prompts you to notice.** This stack changed what `hitloc::describes` answers. Five things said the old answer, and only one of them was code: - the column's own doc comment - the worked figure's per-case note - `docs/features/target_described.svg`, a checked-in render - `docs/FEATURES.md`, generated prose - a **passing test** asserting a tank reads `0.00` The test is the worst of them and it is worth saying why. It did not merely fail to catch the change - it existed to *defend* the old belief, so the branch carried the wrong claim pinned in place by something green. **A passing assertion reads as evidence rather than as a record of what somebody believed at the time**, so the next person trusts it harder than an untested assumption. And the render caught the author of this section out while writing it. The `FEATURES.md` paragraph was edited by hand, the generator was then re-run for the SVG, and it **overwrote the correction and restored the false text**. The real source was a `reading:` string in `examples/vignettes/firing.rs`. Editing a generated file is not a small mistake that gets corrected later; it is a correction that gets silently reverted by the next build, which is worse than not making it. The tell was free and sitting in the failure message the whole time: `tests/vignettes.rs` prints `cargo run -j 2 -p sds-core --example featuredoc` when the docs go stale. Before editing a docs file, check whether anything writes it. ### The sentinel list, read out of the jar The third shape from the sweep write-up - a sentinel read as a value - is greppable *only if the sentinels are enumerated*, so here they are: | constant | value | met as | |---|---|---| | `WeaponType.DAMAGE_BY_CLUSTER_TABLE` | **-2** | a rack's damage; reached the bot as negative expected damage | | `WeaponType.DAMAGE_VARIABLE` | **-3** | a Snub-Nose PPC and a VSP laser | | `WeaponType.DAMAGE_SPECIAL` | **-4** | not yet met | | `WeaponType.DAMAGE_ARTILLERY` | **-5** | an Arrow IV, once minus five points | | `Entity.DOES_NOT_TRACK_HEAT` | **999** | a vehicle's heat capacity, the `heat_*` group | | `Entity.LOC_NONE` | **-1** | a location a unit does not have | | `Entity.LOC_DESTROYED` | **-2** | transfer, where damage goes when a location is gone | | `Entity.UNLIMITED_JUMP_DOWN` | **999** | not yet met | | armour states | **-1, -2, -3** | `ARMOR_NA`, `ARMOR_DOOMED`, `ARMOR_DESTROYED`, already handled by `points()` | Every one is a small negative or 999, and every one is a legal-looking number in its own field - which is why `getDamage()` returning -2 became a negative expectation and 999 became a heat capacity nothing could exceed. **A reader that takes a raw MegaMek integer without naming which sentinels it can be is the third shape**, and this table is what makes that greppable rather than an audit. - [ ] Sweep the bridge for raw MegaMek integers reaching the wire without a sentinel check. `Observation.points()` is the pattern that does it right; the question is how many readers there are and how many name their sentinels ## `destroyLocation` cannot produce a BV differential Outside a `Game`, `Entity.destroyLocation` marks a location doomed and stops: armour and internal go to `IArmorState.ARMOR_DOOMED` (-2) rather than `ARMOR_DESTROYED` (-3), `isLocationBad` stays false, no criticals are destroyed, and walk and run MP do not change. Criticals and MP recalculation happen later, in the game manager's damage resolution. So `BV(intact) - BV(intact, location destroyed)` cannot be computed this way: it measures armour and structure pools being zeroed, and never reaches the movement term. It throws nothing and returns well-ordered numbers. `new Game()` throws `ExceptionInInitializerError`, so attaching one is not the easy fix it looks like.