bayes for days
sds plan unit-types.md
78 kB
Markdown
at main


id: unit-types title: Vehicles and infantry play, not just Meks status: open dependsOn: [protocol, features] exitCriterion: > A force of Meks, vehicles and infantry is played without a unit type taking an inert default, and sds spread shows no feature reading a constant for one type and a range for another. #

unit-types #

The Core Rules cover Meks, combat vehicles and infantry. Aerospace, dropships and anything that leaves the ground are out of scope and stay out.

Today the basis is written for Meks and it shows in the vocabulary: canStand asks a Mek about its gyro, level_tmm reads a movement modifier from hexes moved and a jump, heat_* is four features about a scale a vehicle does not have, and target_tonnage now spans to 200 tons because the unit files say a combat vehicle reaches it - so the denominator already admits vehicles while the features around it do not.

Nothing refuses a vehicle. That is the problem: a feature that is meaningless for a type reads a constant for it, and a constant within a decision teaches a difference-framed fit nothing. sds spread is the instrument - a column that moves for Meks and is flat for tanks is the signature, and it is invisible in a win rate.

What it needs #

  • A census first: what does MegaMek actually hand us for a vehicle and for infantry, and which of the named features return something meaningful. Probe it rather than reason about it.
  • A decision on features that cannot apply. A heat feature for a vehicle is not zero, it is undefined, and those are different things to a fit.
  • Movement: vehicles read a terrain column Meks do not, and infantry move by different rules again. MoveHex and the pathfinder have to answer for all three or say which they cannot.
  • Damage: infantry take damage by a different table entirely, so hitloc and the volley path need to know what they are shooting at.

The two sweeps, and what the second one adds #

The zero-as-absence sweep above found three columns gated on has_locations(). A later sweep for columns that read differently against a vehicle - not just zero - found five, and the two extra are the interesting half:

  • p_psr_threshold reads a mean of 0.39 and a p95 of 0.99 against a vehicle, for the twenty-point piloting cliff. Entity.canFall() returns false for a combat vehicle, probed directly. It is a Mek rule applied to a machine MegaMek says it does not govern, and it errs upward.
  • overkill correlates +0.9925 with p_kill against a vehicle, against +0.789 against a Mek: on a body where every location is lethal, "took more than it can absorb" and "was destroyed" are the same event one point apart.

Neither could have been found by the first sweep, and that is not a criticism of it. A sweep for zeros cannot find a column that reads 0.39, and a sweep for absences cannot find one that has become redundant. The generalisation both support is that the pattern is not "zero means absence" - zero was the costume. It is a value arrived at without the thing it claims to measure, and it can err upward as easily as down.

The full family, the discriminator that separates the real asymmetries from the artefacts, and the evidence for what a vehicle location is worth are below.

What MegaMek answers for a combat vehicle #

Probed with sds.SdsVehicleHits, the vehicle companion to SdsHitLocations: the dice are pinned through Compute.setRNG and the tables are read out of Tank.rollHitLocation, the same code that resolves the shots.

  • Seven locations, five of them reachable. BD FR RS LS RR TU FT. No normal attack on any of the four arcs ever rolls the body or the second turret, so the model carries five and ignores two rather than refusing a target that reports them.
  • No transfer chain. Tank.getTransferLocation answers LOC_DESTROYED for every location - the same sentinel a Mek's head and centre torso give. So every location a vehicle has is a lethal one, and hitloc::lethal now derives both bodies' sets from that map instead of from a hand-written constant.
  • No dependents and no rear armour pool. A vehicle's four facings are four of its locations, so the arc picks the location rather than a second armour pool on the same one.
  • Motive damage is a flag on the roll, not on the location. Rolls 3, 4, 5 and 9 arrive with EFFECT_VEHICLE_MOVE_DAMAGED on every arc - 13 ways of 36 - while 6, 7 and 8 land on the same armour and do not. It cannot be modelled as a property of a location and is carried as its own share.
  • Criticals are on the roll too: 2 and 12 on every arc, and 8 from the left and right.
  • getAllowedPhysicalAttacks() returns 1, the same as a Mek's.
  • A turretless design rolls a different table, with 10-12 going to the arc's own location. Dumped rather than derived: MegaMek takes a different branch, and a redirect applied on our side would have been a guess that happened to agree.

A defect in MegaMek's vehicle front table #

Front 5 and front 9 both return the left side. Every other arc separates the pair - left gives FR and RR, rear gives LS and RS - and the printed front table separates them as well. The consequence is that a vehicle shot from in front never takes a right-side hit and takes left-side hits at twice the rate.

It was checked rather than assumed, because an answer this clean is usually the instrument: both rows consume exactly two d6 through the pinned source, and forcing a third die to each of its six values moves neither. SIDE_LOC_MAPPING is {FRONT->FR, LEFT->LS, RIGHT->RS, REAR->RR}, so the mapping is not the cause.

The bot models what resolves the shots, so it models this, and a_vehicles_front_table_hits_one_side_twice_and_that_is_megameks is the test that says so the day it changes.

The motive damage table, dumped #

sds motive-damage reads it out of TWGameManager.vehicleMotiveDamage, which mutates a Tank inside a game rather than being a pure function of a pinned die - so it needs a real scenario, a crewed and deployed entity and a manager. The scenario is reloaded for every cell, because motive damage accumulates and immobilisation is sticky. Recorded in docs/livefire/motive-damage.txt.

The outcome is a pure function of roll + modifier, verified across all 44 cells of a four-modifier sweep:

2d6 + modifier result
2-5 no effect
6-7 minor: +1 driving
8-9 moderate: +2 driving, -1 MP
10-11 heavy: +3 driving, half MP
12 or more major: the vehicle is immobile

And the modifier is the arc, from Entity.getMotiveSideMod: front +0, rear +1, either side +2.

So the number p_mission_kill should carry, per motive-flagged hit:

arc needs P(immobilised)
front 2d6 >= 12 1/36 = 0.028
rear 2d6 >= 11 3/36 = 0.083
left or right 2d6 >= 10 6/36 = 0.167

What this says about the feature as it currently stands #

p_mission_kill for a vehicle currently reads P(a motive-flagged hit lands), which measured a mean of 0.688 over 9,249 real firing candidates. The honest figure is that probability multiplied by the table above, so the crippling end is overstated by roughly a factor of 36 from the front.

The size of the error is not the worst part. At 0.688 the column is non-zero on 99.9% of vehicle candidates and near-constant across them, so within a decision it barely discriminates between one tank and another while dominating every comparison against a Mek. Flat within a decision teaches a difference-framed fit nothing - the sds spread failure this epic exists to catch - so it degrades play and hides that it is doing so, in a direction no win rate would reveal. Do not fit a weight on this column until it is calibrated.

The calibrated form is a genuine improvement rather than just a smaller number: it is six times more likely to cripple from the flank than from the front, so for the first time the feature says something about where to shoot a tank from, which the flat version could not express at all.

Answered by jmm #

  • Model everything. "We should just work on modeling for anything that isn't modeled yet." There is no tactical preference to preserve: a basis that favoured Meks because three columns collapsed was not a bot anybody chose, and once a vehicle scores honestly the preference is a decision somebody can make on evidence.
  • p_mission_kill spans unit types on purpose. "For ex this is one reason it's 'mission killed' and not 'shot leg off': tanks can be mobility-damaged in the same way." A Mek losing a leg and a vehicle taking motive damage are the same event at the level the bot reasons about, so motive damage belongs in that feature rather than beside it. That is what [hitloc::LocationDamage::motive] feeds, and it is why a vehicle's crippling set is empty rather than being a Mek's with different names.
  • Two movement points per level is a hovercraft rule, not a ground vehicle one - "level changes are harder (they want flat terrain)". So the tank-in-woods discrepancy claude/move-legality measured, 3 against 4, has some other cause and is still unexplained. The measurement was real; the rule attributed to it was not.
  • A unit may make a move it cannot afford if the destination is legal and it is the unit's whole turn. Given about infantry and generalised deliberately; it comes up for a Mek down to 1 walk MP as well. Not a vehicle rule and not this branch's code, recorded here so it is not lost.

What an analysis pass actually costs #

A pairwise sweep over tactics-hub-mirrored - 300 matches of decision logs, 3.6 GB on disk - runs at 147 MB resident. Measured, not estimated, at 1m43s into a steady-state run over 80 logs.

That is the opposite of what everyone assumed, including the person who warned about it and the person who took the warning. The reasoning that produced the wrong answer was "300 matches of JSON is a lot of memory", which is an inference from a data size to a memory cost with the shape of the program left out.

The general form, because it is not about memory. The inference ran from a data size to a resource cost with the shape of the program left out. Any estimate of that form is a guess wearing arithmetic: the same input at the same size costs three orders of magnitude more or less depending on whether the program holds it or walks past it. Ask which before quoting a number, and if the answer is not known, measure rather than cite a precedent from a program of the other shape.

The shape is the whole thing. sds train holds a corpus in memory and its OOM on this machine is real and recorded. A sweep walks past the corpus: json.loads per line, never a whole file, and accumulators that are about 2,600 floats per phase. Same input, same size, three orders of magnitude apart. The train precedent does not transfer, and citing it felt rigorous while being an assumption.

So: the binding constraint on this machine during a run is the match containers and their JVMs, not analysis. Two containers plus their JVMs is what 28 of 30 GB looks like, and that is simply what a run costs.

Cap it anyway. CAP_MEM=2G capped python3 ... is thirteen times headroom on a 147 MB job and buys nothing in expectation. It is still right: MemoryMax with MemorySwapMax=0 makes the analysis structurally incapable of being the kernel's victim-pick, whatever the figure says. A guard that is unnecessary in expectation is exactly what belongs around somebody else's two hours of matches.

Observability: to-hit could be logged, and is dropped one step before #

expected_damage reads 1.4x higher against a vehicle and the likeliest explanation - vehicles are easier to hit, so more of the volley lands - could not be established, because the decision log carries no to-hit. It could.

The number exists all the way to the point candidates are built. Shot.to_hit is on the wire (wire.rs, emitted by Observation.java from WeaponAttackAction.toHit), and ShotKey.to_hit is what ev::ShotPmf uses to compute every damage distribution. What drops it is the menu: candidates are reduced to Vec<(String, FeatureVector)> - a label and a phi - before anything is written, so a quantity that is not a feature cannot survive to the log.

So carrying it is widening that tuple, not plumbing a new signal. A mean or minimum to-hit per candidate would answer this question and is a local change.

This is the second question tonight a corpus should have been able to answer and could not - the first being which arc a shot was taken from, which needed examples/sidetable.rs to reconstruct from positions that were logged. The pattern is the same: the log records what the fit consumes, so any question about why a candidate scored as it did has to be reconstructed or cannot be asked. ### For jmm: is the decision log a fit input or a diagnostic record?

Nobody has decided, and it is being decided by default every time somebody cannot answer a question. Twice tonight, on this branch alone:

  • Which arc a shot was taken from. Reconstructible, but only because positions happened to be logged - it needed examples/sidetable.rs to dump los::side_table and an analysis that reads the table back.
  • What the to-hit was. Not reconstructible. It exists on the wire and drives every damage distribution, and the menu drops it one step before writing.

The log currently records exactly what the fit consumes, which is a coherent design. The cost is that why a candidate scored as it did is not answerable from a corpus, and that question keeps arising - three of tonight's findings needed it and one could not be closed.

Proposed with the question rather than done quietly: carry a per-candidate to-hit summary, which is widening Vec<(String, FeatureVector)> and nothing more. It is a local change, and it is the first instance of whatever the general answer turns out to be - so it is worth deciding the general question first rather than patching the one case and leaving the design unmade.

A refused move is re-proposed identically, forever #

Found while measuring the illegal-move rate on the first pure-vehicle suite anyone has played. It is not a vehicle defect; vehicles only make it visible.

The pathology. In 1v1-succession-wars-s4010001-vehicle, one unit produced 38 refusals over rounds 3 to 40, and every one was the same step list - TURN_RIGHT, FORWARDS, FORWARDS, FORWARDS. MegaMek refuses it, the unit takes the inert default and stands still, and next round it proposes exactly the same move again. For thirty-seven rounds.

The mechanism, confirmed in code by the movement agent. SdsClient.buildPath returns null when isMoveLegal fails, and askMovement increments a tally and returns an empty path - the refusal never crosses back to the Rust bot. So the bot observes a world in which the unit simply did not move, and being deterministic given state and seed it proposes the same thing again. It is not that the bot ignores the refusal; it is never told.

That also means the fix is protocol work rather than a patch: a client that substituted its own move would break "every order has exactly one possible author", which is the invariant the whole design rests on. It is with jmm.

And it is worse than "re-proposes identically". In 2v2-jihad-s102033 a Bishamon was refused in 36 of 40 rounds and never moved at all - ten distinct step lists, all landing in the same hexes, seventeen of them identical. So the failure is not one stuck proposal but a whole candidate neighbourhood being refused, with the bot cycling among refused variants for an entire match. In that corpus 29 of 52 refused units were refused more than once, so the concentration caution applies and does not explain it away.

It is general. The mixed arms, which are mostly Meks, show the same shape at lower amplitude - worst repeat 11 on origin/main and 4 on this stack, against 63 and 67 refusals. So a Mek can livelock too; it is just rarer because a Mek is refused less often.

Vehicles amplify it because their refusal rate is far higher. Over the whole 24-match arm, 301 movement decisions:

answered illegal defaulted
tracked (n=175) 47% 28% 25%
hover (n=89) 44% 36% 20%
wheeled (n=37) 78% 16% 5%
all (n=301) 50% 29% 21%

Half of all vehicle movement decisions are answered. The other half split between a move MegaMek refused and no move proposed at all, and the 21% defaulted is its own number rather than part of the refusals.

The comparison has to be within one suite #

An earlier revision compared this against "7-8% on the mostly-Mek mixed arms", which was the blended rate of a suite containing both. Split by the moving unit's own type, the mixed arms say:

Mek vehicle
this stack 7.6% illegal (n=367) 21.3% illegal (n=183)
origin/main 5.0% illegal (n=403) 28.5% illegal (n=151)

So within a single suite the gap is about four to six times, not the thirty the blended figure suggested.

And the suite matters more than the bot. The movement agent measured Meks near main at 0.22% refused over tactics-hub-mirrored (20 of 9,109), against 5.0% for origin/main's own bot on the mixed suite here - a twenty-fold difference with the same bot and different scenarios. Whatever drives that is a property of the boards or forces these generators draw, and it means a refusal rate quoted without its suite is not a number. The vehicle figures above are only comparable to the Mek figures beside them.

That elevated baseline is consistent with claude/move-legality + PR #313, which measured that a Mek pays one movement point per level and everything else pays two, and which is not on main - so this bot under-prices every vehicle level change. Wheeled being the least affected at 16% is worth a look by whoever picks that up: #313 added a wheeled column nobody had, and wheeled is the motion doing best here.

The two defects, separated by the repeat histogram #

87 refusals over 42 distinct (unit, step-list) pairs:

repeats distinct proposals
38 1
4 1
2 5
1 35

One proposal accounts for 38 of the 87 refusals - 44% of them. The other 49 are spread over 41 distinct proposals, 35 of which were refused exactly once. That is the clean separation: a long tail of one-off refusals is what a pricing error looks like, and a single proposal repeated 38 times is the livelock. Both are present, and neither figure would be visible in the other's aggregate.

Correcting a number this file carried #

An earlier revision of this section said refusals ran "62% ... 38% for tracked, 38% for hover and 14% for wheeled". Those came from the first four matches of the arm, before it finished. The full 24-match figures are the table above - 29% overall rather than 62%. The partial sample was dominated by the livelocked scenario, which is exactly the hazard of reading an arm while it is still running. The finding survived; the magnitudes did not.

The two are separate defects and want separate fixes. #313 makes the bot propose fewer illegal moves; it does nothing about proposing the same illegal move forever once it does.

A wrong turn worth recording #

The first cut of this analysis grouped refusals by board and read the result as confirmation of #313: Desert Mountain 95%, Rolling Hills 75%, City Hills 33%, flat boards 0%. Mountains and hills at the top of a list about level changes is a very persuasive picture.

It was pattern-matching on map names. Parsing the actual .board files and correlating measured elevation range against refusal rate gives +0.031 over 13 boards - no relationship at all. Badlands 1 has the largest relief on the suite and refused nothing; City Hills Residential has a relief of 1 and refused 46%. The board grouping was a proxy, and the thing it proxied for turned out to be one scenario with a livelocked unit contributing 38 of about 100 decisions.

Nearly reported as independent confirmation of another agent's PR. The check that caught it was measuring the quantity the names stood for.

Rule: a synthetic object gets a synthetic answer #

Four times in one night, across the agents working here. Written as a rule because it has stopped being an anecdote:

  • A bare new Tank() reported every location as not-destroyed immediately after being destroyed - an uninitialised entity, not a result. It is why this epic answers lethality from the transfer map instead.
  • An uncrewed entity answered zero hits where a crewed one answered four.
  • TWGameManager.vehicleMotiveDamage needed a real game, a real crewed entity and a real manager before it would answer at all.
  • applyCriticalHit threw on a CriticalSlot built by hand for WEAPON_DESTROYED, FUEL_TANK and TURRET_DESTROYED, because the slot names no real mount.

The rule. When probing MegaMek, start from a real entity in a real game - loaded from a scenario, crewed, deployed - and only fall back to a constructed object once the real one has answered. The reverse order costs a probe, a plausible wrong answer, and the time spent believing it. The tell is an answer that is uniformly empty: all-false, all-zero, or a throw naming something the object never had.

And the corollary, which cost the most time tonight: an all-empty answer from a synthetic object is not evidence of absence. It is evidence the question was not asked.

Recorded negatives: why two pulls on this branch have no number #

Both pulls from this work - 3mubgtasnuy2z and 3mubj6zsl4n2a, created ninety minutes apart - show ? where tangled.org's number goes, while other pulls number immediately. Two hypotheses were tested and both are refuted, written down so nobody spends time on them a third time:

  • Not the missing Change-Id trailer. These commits carry none, but claude/test-threads-doc carries none and is #306, and claude/move-legality carries none across six commits and is #287. claude/kill-my-containers does carry one and is #310, so both combinations of (has trailer, is numbered) occur.
  • Not patch size. claude/expected-crits is 25 commits and is #270; these are 16 and 26.
  • Not indexer backlog. #310 was created after the first of these and was numbered immediately.

The records are valid and readable by record key, and neither pull should be re-created: a retried pr create can leave a duplicate record on the PDS, and then one branch has two pulls. This is jmm's tooling and his account, and it is with him.

pgrep matches the shell that launched it #

Four independent rediscoveries of this in one night, across three agents. A watcher looking for its own job finds its own bash -c wrapper, because the pattern is in that wrapper's command line, and reports the job still running after it has finished - or, worse, pkill -f kills the shell that issued it.

The fix each of them arrived at separately: match the process, not the command line. ps -C python3 -o cmd lists only real Python processes and cannot include the shell; pgrep -f pairsweep cannot tell them apart. Anything that polls for its own completion, or kills by pattern, wants the first form.

Measuring an arm #

Two rules, both bought rather than reasoned:

  • Smoke-test one game before committing an arm to 24. Run one, read the decision log, and confirm the seat answered with candidates - outcomes {'answered': 32} and 9,617 candidates is a working bot; {'defaulted': 162} and zero candidates is a seat that never started. A whole baseline arm was lost to the second, and it printed weapon tables and a win rate while being a benchmark of the inert default.
  • Report the per-scenario breakdown even when the pooled result is clean. At the sample sizes a suite of this size gives, one scenario can carry an entire difference, and a single pooled share hides it. A summary standing in for the thing it summarises is the same failure as every other one in this file.

For jmm: what "actual tanks, real combat vehicles" should mean #

The pool filter is unit_type = 'Tank' plus canon, valid, 20-100 tons, BV above zero, at least two weapons, and a motion type from a list. The motion list is the question, and it is currently Tracked, Wheeled, Hover. Here is what each choice actually admits, in the succession-wars era, so the decision is over a list rather than an abstraction.

In the pool today: 256 designs #

motion designs heaviest examples
Tracked 139 Mars Assault Vehicle (100t), Behemoth Heavy Tank (100t), Alacorn Mk II (95t)
Hover 70 Condor Heavy Hover Tank (50t), Drillson Heavy Hover Tank (50t), Harrier (50t)
Wheeled 47 Giant Guardian (80t), Gallant Urban Assault Tank (70t), Ishtar Fire Support (65t)

20 to 100 tons, mean 52. Roles: Striker 61, Missile Boat 52, Scout 38, Juggernaut 32, Brawler 29, Sniper 24, Ambusher 10, Skirmisher 4, none 6.

Excluded today, and worth confirming #

  • Naval (4) and Hydrofoil (3) - Mauna Kea Command Vessel, Monitor Naval Vessel, Sea Skimmer. They need water to exist on and the boards are mostly land. Excluded, and this seems clearly right.
  • WiGE and Submarine - excluded for the same reason plus rules nothing here implements.
  • VTOL is a different unit_type and never enters. hitloc refuses one explicitly: it reports every abbreviation a turreted tank has plus a rotor, and reading one as a tank is the partial-match hazard that made a quadruped look like a biped.

Decided and built, subject to jmm overriding any of it #

Built on the recommendations below rather than blocking on an answer; each is a filter change and a regenerate, not rework. What was actually excluded is seven named utility designs - three Coolant Trucks, the R10 coolant ICV, two MASH Trucks and the Engineering Vehicle - and nothing else. Pools are now succession-wars 218, clan-invasion 60, jihad 176.

The APC exclusion was not built, and that is a change from the recommendation. Every way of expressing "it is a transport" turned out to be wrong when the list it returned was read:

  • Not a troop bay. 58 designs carry troopspace, and they include the Goblin Medium Tank, the Turhan Urban Combat Vehicle, the Prowler and the Skulker Scout Tank - combat vehicles that happen to carry a squad.
  • Not firepower. The plain Vedette Medium Tank, the generic Inner Sphere line tank, reads 7.0 - below the Engineering Vehicle's neighbours. Any threshold that drops a MASH Truck drops the Vedette.
  • Not role. The Coolant Truck 135-K is filed as Ambusher.

So the utility vehicles are named instead, which is small and reviewable. The toothless end of the APC question is bounded and can be dropped in one word if jmm wants it: the twelve 20-ton Heavy Hover/Tracked/Wheeled APC variants at firepower 4 to 12, and the ten Badger Tracked Transports at 30 tons. They are still in.

No committed scenario contains an excluded design, so the suites and the measurement in comparison.md are unaffected by this filter.

The three questions, and the recommendations they were built on #

Read against "actual tanks mainly, like real combat vehicles" and "skip odd or rare equipment types", which points at the plain cases throughout.

  1. Hover: keep them in? Recommend: in. 70 of 256. A Condor and a Drillson are plain combat vehicles that anybody would call tanks, not odd equipment. The one caveat is that their movement is the least tank-like - the level-change charge claude/move-legality measured turned out to be a hovercraft rule rather than a ground-vehicle one, and that is still not fully settled - so if the answer is "in", movement is where it will show.
  2. Transports and APCs: drop them? Recommend: out. Twelve are in the pool on the two-weapon minimum - the Badger Tracked Transport family, the Heavy Hover APC - and several carry exactly the two the filter demands. A Badger with two machine guns in a line battle measures the harness rather than the bot, which is the argument that already excludes IndustrialMeks from the Mek pool. Raising the weapon minimum for vehicles drops them.
  3. Is 20 tons the right floor? Recommend: keep it. Inherited from the Mek pool. Combat vehicles go lighter, and a very light one dies to a single volley - a real machine, but a thin benchmark row, and not worth widening the pool for while the question is whether tanks work at all.

Nothing here is changed until answered: both pool defects on this branch so far came from trusting a filter rather than reading what it returned.

A prediction for the calibrated motive column, registered before it is built #

The calibrated p_mission_kill is six times more likely to cripple from the flank than head-on - 6/36 against 1/36 - where the flat version was constant across arcs. So:

  1. The share of shots taken at a vehicle's flank or rear rises relative to head-on, on vehicle targets.
  2. The effect is larger on vehicle targets than on Mek targets. A Mek's crippling locations are its legs and the hit tables do not favour a leg from the flank the way motive damage does, so the Mek arm is close to a control.
  3. Magnitude: small. p_mission_kill is one weighted column among fifty-one and the arcs already differ for other reasons - rear armour is thinner, and rear_arc_gain and the side tables already push that way. I predict the flank-and-rear share of shots at vehicles moves by under 10 points, and I would not be surprised by 3.

If it does not move at all, that is a real finding and not a failed change: it would say a single column cannot steer target selection against the rest of the basis, which is worth knowing before anybody fits a weight expecting it to.

And a second prediction, about where the bot stands #

The same asymmetry should move movement, not only target selection, because the firing basis is measured from every candidate stand: a stand that puts the shooter on a tank's flank scores a p_mission_kill six times the one head-on.

  1. A unit closing on a vehicle arrives at its flank or rear more often than one closing on a Mek. Predicted to be the weaker of the two effects and possibly invisible: the flank is already preferred for reasons that have nothing to do with motive damage - thinner rear armour, rear_arc_gain, the side hit tables - so this column is adding to a push that already exists rather than creating one. I predict under 5 points and would not be surprised by nothing measurable.

The charge, probed: what a vehicle may do and what it would cost to wire #

Both halves of the rule confirmed against MegaMek, with a 100-ton Gürteltier and a 100-ton Atlas placed adjacent and both done moving:

  • A vehicle may charge a Mek. ChargeAttackAction.toHit answers 7 after a one-hex run, 8 at three hexes, 9 at six - 5 (base) + 2 (attacker ran) + target movement. Damage scales with the run: 0, 20 and 50 points at one, three and six hexes, against 10 taken by the vehicle each time. A six-hex charge is a fifty-point attack, which is more than most vehicles' guns.
  • A Mek may not charge a vehicle. IMPOSSIBLE: Target is not a 'Mek, at every distance. MegaMek's rule, not ours.
  • The first refusal a vehicle gets is Target must be done with movement, which is a state precondition rather than a unit-type rule - so a charge is declared against a machine that has already committed.

Why this is not a gate to flip #

Observation.physicalsAt gates on shooter instanceof Mek and emits punches and kicks as pseudo-weapons in the firing menu. A charge cannot be carried that way, and SdsBattleArmor.rush on main already says why: a punch costs a limb and is declared in the physical phase, while a charge costs a move - it is declared during movement, takes a MovePath ending in the target's hex, its damage scales with distance travelled, and it hurts the attacker. That is why neither charge nor death-from-above is on the wire today, for any unit type.

So wiring it is a movement-space feature: the candidate generator has to offer "run into that hex and declare a charge", and price it as damage dealt against damage taken. It is the same work for a Mek charging a Mek, and it overlaps the branch that owns physicals.

It is also not on the acceptance test's path. jmm's test is a Mek kicking a combat vehicle to death; a kick is a physical-phase attack, it already reaches the wire, there is no target-side unit-type gate, and a_kick_against_a_vehicle_prices_through_the_same_table pins that it prices correctly. The vehicle charge adds a capability vehicles are entitled to and blocks nothing.

For the Advance tactic, on another branch #

A tank's vulnerability is directional in a way a Mek's is not. Motive damage is +0 from the front, +1 from the rear and +2 from either side, so the same flagged hit is six times more likely to immobilise a vehicle from its flank than head-on - and a vehicle's four armour facings are four separate locations, so there is no transfer to spread a concentrated flank attack.

Advance is a closing tactic. A bot that closes head-on at a tank arrives at its least vulnerable facing. Not this epic's to act on, and the interaction is recorded here because the branch building Advance has no way to know it otherwise.

The flank baseline, registered before the calibration runs #

Measured on the two arms already recorded, neither of which had the calibration - so this is the "before" the flank predictions are against. The arc is not recomputed by the analysis: crates/sds-core/examples/sidetable.rs dumps los::side_table, the function the bot itself used, and the analysis reads that table. A port would be a second instrument, and two instruments agreeing proves nothing about either.

run target n front left right rear flank+rear share
after (this stack, flat motive) Mek 345 264 28 14 39 0.2348
after (this stack, flat motive) vehicle 94 63 14 4 13 0.3298
before (origin/main) Mek 382 287 19 24 52 0.2487
before (origin/main) vehicle 78 64 4 2 8 0.1795

The move from 0.1795 to 0.3298 is not what the prediction is about. Neither arm had the calibration: that gap is honest scoring and the pools, which changed whether a vehicle is worth shooting at all and therefore which shots got taken. It is context for reading the next run, not a preview of the arc modifier's effect, and it must not be quoted as one.

Read it cautiously in its own right, too. At 78 to 94 vehicle shots the standard error on a share near 0.3 is about 0.048, so even that gap is roughly two standard errors - suggestive, not established.

What it does establish is that the instrument works and the sample is thin. 78 to 94 shots per arm is not enough to see a 5-point move - the prediction is one standard error at this sample size, so the run could not tell it being right from its being wrong. That is a design fault found before launching rather than after reading an uninformative null, which is what registering a magnitude is for.

So the run is two arms, in this order:

  1. The vehicle-only suite first. 18 scenarios where every target is a tank, which buys several times the vehicle shots for the same match budget. This is the arm that can answer the registered flank prediction.
  2. The mixed suite second, because that is the question actually asked - tanks and Meks together against Princess, the realistic case. Vehicle-only answers the science; mixed answers "are we getting better".

Who each arm plays, and what will be said about the result #

All three arms are sds against Princess. None is self-play - sds bench seats Princess opposite by default and the script passes no --self-play, which the seat files confirm (sds_North against princess_South, alternating). This is stated because a flank share measured against our own bot and one measured against Princess are different quantities, and a reader with only the suite and the bot in front of them would have no way to know which they had.

The win rate already recorded, disclosed before the run #

The two mixed arms already on disk, computed from results.json by mapping each seat to its team rather than assuming the order:

arm decided sds princess win rate
this stack, mixed 24/24 11 13 45.8%
origin/main, mixed 24/24 14 10 58.3%

This branch's arm came out lower than the baseline's, and that is being said here rather than left for someone to find. It is also uninformative: at 24 games the standard error on a win rate is about 10 points and on the difference about 14, so a 12.5-point gap is under one standard error. The harness prints the same conclusion in its own words - about 97 decided games to separate 10 points, 385 to separate 5. Neither number is evidence about this change in either direction.

What will be said if arm 3 comes out badly, decided now #

  • "Poor" means below 50%, and poor is the honest expectation: the bot is being asked to fight Princess with a unit type it could not score at all a day ago, on a suite drawn to include machines nobody has tuned against.
  • Arm 3 will be reported as it comes out, in the same table whether it is 40% or 60%, with the sample size beside it.
  • No claim of improvement or regression will be made from arm 3 alone. At 24 games nothing between roughly 30% and 70% is distinguishable from 50%, so the only honest reading of a single arm is "consistent with parity".
  • No weight will be tuned to improve it. Nothing on this branch or the one below it fits a weight, and a number that moved because a weight was nudged after seeing it would be worthless - especially with a p_mission_kill column whose calibration is one commit old.
  • If arm 3 is genuinely bad in a way the sample can carry - a win rate near zero, or a collapse in answered - that is a defect to find, not a figure to report, and the run stops there.

What the Princess benchmark can and cannot attribute #

The run that measures this stack will be a state-of-the-bot measurement, not an A/B of one change. By the time it runs, three things will have landed together: vehicles scoring honestly at all, the pools that put them on the board, and the calibrated motive column. A win rate against Princess measures the three as a set.

So it cannot separate the calibration's effect from the pools'. Do not report it as though it could, and do not let a later reader infer it. If the calibration needs attributing on its own, that is a paired run with the column switched and nothing else moved, and it is not this one.

The flank predictions are the part that is attributable, and that is why they were registered rather than a win rate. Nothing else in the stack pushes on approach direction: the pools change which machines are on the board and the honest scoring changes whether a vehicle is worth shooting, but only the arc modifier says a tank is easier to cripple from the side. If the flank share moves, this column moved it; if it does not, that is a real finding about how much one column can steer against fifty others.

What is already checked on the acceptance-test path #

jmm's test for the whole effort is a win where a Mek kicks a combat vehicle to death. Three things on that path were checked here rather than assumed:

  • A vehicle can be a physical target. Observation.physicalsAt gates on the shooter being a Mek and puts no unit-type gate on the target; legality comes from KickAttackAction.toHit, which accepts a tank. There is no hidden instanceof Mek on the target, which is the failure that would have been invisible until exactly this test.
  • A kick against a vehicle prices exactly, not approximately. Tank.rollHitLocation returns identical rows for HIT_NORMAL, HIT_PUNCH and HIT_KICK on all four arcs, so this repository using the normal table for a physical - a documented approximation for a Mek - is exact for a vehicle.
  • The figures are sane. A kick is one large packet, which is the shape a vehicle is least able to absorb: no transfer chain, and every location lethal, so a single packet through anywhere ends it. a_kick_against_a_vehicle_prices_through_the_same_table pins that a kick finishes a worn tank and does not finish a healthy one.

Next epic: vehicle columns need a scale of their own, not the Mek's #

The test, before any of the instances #

A large ratio is not itself a defect. The question is never "does this column read differently against a vehicle" - it should. The question is whether it reads differently for a reason about the machine or a reason about our reader. p_kill reads 25.9 times higher against a vehicle because a vehicle is that much easier to kill: every one of its locations is lethal, and that comes from MegaMek's own transfer map. A bot that prefers shooting vehicles for that reason is making a defensible tactical judgement, and "fix everything that differs" would break the one column telling the truth.

And the pattern behind all of it is not "zero means absence". That was the costume, not the shape. Every other defect found on this branch was a zero standing in for something unmeasured - three scoring columns, five movement outputs, a bot binary the container could not see. Then p_psr_threshold turned up reading a mean of 0.39 and a p95 of 0.99 for an event that cannot happen. The real pattern is a value arrived at without the thing it claims to measure, and it can err upward just as easily as down. Looking only for zeros would have missed it.

The family #

The basis was built for a machine whose worth lives in its location contents, and a vehicle's does not. That single fact has now produced five separate findings, and they are one family rather than five bugs. Measured over 10,533 real firing candidates from the recorded run, by target type:

column Mek mean vehicle mean ratio verdict
p_kill 0.0048 0.1232 25.9x real
value_destroyed 0.0048 0.0797 17x artefact
overkill 0.0238 0.1117 4.7x real ratio, but the column stops being its own
p_psr_threshold 0.1693 0.3878 2.3x artefact
expected_damage 0.2248 0.3219 1.4x probably real, mechanism not established
p_mission_kill 0.0004 0.688 ~1700x artefact, fixed on claude/vehicle-motive
the four heat_* varies ~0, flat - sentinel

Telling the real ones from the artefacts is the whole job. A large ratio is not itself a defect: p_kill reads 25.9 times higher against a vehicle because a vehicle is that much easier to kill - every one of its locations is lethal, and that comes from MegaMek's own transfer map rather than from an omission. A bot that prefers shooting vehicles for that reason is making a defensible tactical judgement. The artefacts are the ones where the number is high or low for a reason that is about our reader rather than about the machine.

The artefacts, each verified against MegaMek #

value_destroyed, 17x. Its denominator is Location::value() - guns plus actuators plus 8 engine plus 8 gyro plus 20 cockpit - and those flags come from Observation.contents reading Mek.SYSTEM_* and Mek.ACTUATOR_* out of a location's critical slots. A combat vehicle has no critical slots: probed, systemSlots=0 against an Atlas's 31, because MegaMek models a vehicle's criticals as unit-level results rather than as contents of a location. So vehicle value-at-stake is guns and ammunition alone, the denominator is a fraction of a Mek's, and because every location is lethal the ratio saturates - p95 0.92 against 0.03.

p_psr_threshold, 2.3x. It is "the chance we hit the target hard enough to make it roll not to fall over", measured as the Mek twenty-point piloting cliff. A vehicle cannot fall: Entity.canFall() returns false for one and true for a Mek, probed directly. So this column reads a mean of 0.39 and a p95 of 0.99 for an event that cannot happen. It is the sharpest of the family because it is not an omission producing a wrong scale - it is a Mek rule applied to a machine MegaMek says it does not govern.

overkill, 4.7x - the ratio is real and the column is the problem. A vehicle genuinely overflows more: fewer locations, smaller pools, so a volley spills past one sooner. But overkill is P(some location takes more than it can absorb) and p_kill is P(some location is destroyed), and on a body where every location is lethal those are the same event one point apart. Measured: corr(overkill, p_kill) is +0.9925 against a vehicle and +0.789 against a Mek. So for a vehicle the column carries almost nothing p_kill does not already say, and a difference-framed fit would be putting two weights on one signal.

That is not a wrong number - it is a column that stops being an independent one. This module has the precedent written down: "If it ever finds them equal everywhere, one of them is redundant and should be deleted rather than fitted - which is what happened to los_in/los_out and to overkill." It happened to overkill once already, for Meks.

expected_damage, 1.4x - probably real, and left honest. The likeliest mechanism is that vehicles are easier to hit: no jumping, and many are slow, so more of the volley lands. That is a property of the machine rather than of the reader, which would make it real. It is not established - the to-hit is not carried per candidate in the decision log - and it is recorded as unexplained rather than waved through, because a 1.4x that turned out to be structural would be the same family again.

The heat_* family, flat. getHeatCapacity() returns 999 for a combat vehicle - checked across 41 distinct designs, energy-armed and ballistic alike - which is MegaMek saying "no heat scale", not a capacity. Five of our columns divide by it, so they read about zero for every vehicle in every decision. The values are right and that is why nobody found it; what is wrong is that nothing distinguishes "no heat scale" from "a Mek running cold". Same shape as UPSTREAM.md #6, where WeaponType.getDamage()'s sentinel once reached the bot as real damage - a recurrence rather than a new bug.

Evidence for the scale decision, probed #

sds.SdsVehicleCriticals applies each of MegaMek's fifteen vehicle criticals to a fresh crewed, deployed Gürteltier MBT through TWGameManager.applyCriticalHit, and reads what it cost. A fresh entity per critical, because they accumulate. Baseline: walk 3, run 5, six usable weapons, battle value 2120, 100 tons.

critical walk run weapons battle value flags set
ENGINE 3->0 5->0 6->5 2120->1264 (-40%) engineHit, turretLocked
TURRET_DESTROYED - - - 2120->1169 (-45%) (threw part-way)
AMMO - - - 2120->1840 (-13%) -
SENSOR - - - 2120 sensors=1
CREW_KILLED - - - 2120 crewHits=6, crewDead false
DRIVER, COMMANDER, CREW_STUNNED, STABILIZER, WEAPON_JAM, TURRET_JAM, TURRET_LOCK, CARGO - - - 2120 none observed

Three of fifteen need more context than a synthetic critical slot carries. WEAPON_DESTROYED threw NoSuchElementException and FUEL_TANK and TURRET_DESTROYED threw NullPointerException - they want a real slot naming a real mount, not one built by hand. That is the same answer as the motive table and the munition handlers, for the third time: a synthetic object gets a synthetic answer, and the real path needs the real object. Next time it should be the first thing tried.

Two results are anomalies rather than findings, and are recorded as such rather than used: CREW_KILLED sets crewHits=6 but leaves crewDead false and the unit alive, and TURRET_LOCK does not set isTurretLocked although an engine hit does. Both are probably the same missing-context problem. Neither should be leaned on until re-probed through a path that resolves damage properly.

What this says about the shape #

A vehicle's worth is unit-level with a location-level trigger, not location-level. That is not a rescaling of the Mek model, it is a different shape. Every one of these criticals is a property of the machine - its engine, its crew, its sensors - reached through a location but not stored in one. So "what is this location worth" has no answer for a vehicle in the terms Location::value uses, and any number invented for it would be a number about our reader again.

calculateBattleValue() is the unit that does work. It is MegaMek's own scalar for what a machine is worth, it responds to vehicle criticals in sensible proportions - engine 40% of the unit, turret 45%, ammunition 13% - and it applies to both body plans without translation. It is also already on the wire: target_current_bv and target_original_bv are existing features, so the observation carries it and nothing new has to be plumbed to try it.

The ruling #

Use MegaMek's per-component battle value: the sum of EquipmentType.getBV over the equipment mounted in a location, carried as component_bv. The name is deliberate - it is the offensive half only and is not a location's battle value. Pricing a location with mobility in it needs a battle value calculator that can be asked about a hypothetical machine, which is Helm's job.

It re-prices every Mek location as a side effect, so it is measured on Meks and not only on vehicles.

What the epic has to decide #

Not "translate the Mek reader". There is nothing to translate: a vehicle's engine, crew and motive system do not live in a location, so there is no slot to read. What a vehicle location is worth is a modelling decision, and it needs the same treatment the motive column just got - probe what MegaMek does to a vehicle on a critical, then choose a scale and say so.

Until then, neither value_destroyed nor p_mission_kill should carry a fitted weight over a mixed suite, and p_psr_threshold should not be read as meaningful against a vehicle at all. The column means two different things depending on what is being shot at, and a fit cannot see the difference.

The first of the family, in detail: value_destroyed #

Measured over 436,416 real firing candidates from the 48-match mixed suite (corpora/20260831T070313Z-bench, commit dc336d5), which carries the component_bv pricing above. Targets split by movementMode: 371,238 candidates at a biped, 65,178 at a tracked, wheeled or hover vehicle.

target non-zero mean p95 max at the clamp
Mek 93.6% 0.0283 0.1584 1.0000 0.05%
vehicle 45.9% 0.1965 0.9509 1.0000 3.58%

A vehicle reads 6.9 times a Mek on the same column. Pricing a location by battle value took that from 17 times; what is left is a different thing and has a different cause.

A unit dies once, and the column charged for it per location #

worth_hundredths prices a location whose loss ends the unit at the whole unit, and expected_value_destroyed summed P(destroyed) x value over every location. On a biped that is two terms. A combat vehicle answers lethal on every location it has, so five terms each worth the whole machine were added together for one death.

The size of it, from the same corpus - post-fix a vehicle's reading is exactly its p_kill, so value_destroyed / p_kill is the double count:

target median mean p95 max
vehicle 1.001 1.018 1.073 1.359

Small in the middle and up to 36% at the top, and it pushed 3.58% of vehicle candidates past 1.0, where Score::fraction clamps. A clamped column does not read as wrong, it reads as certain, and every volley past the line ordered identically - among exactly the volleys that kill.

LocationDamage::any_lethal is the union and is what p_kill already reads, so the two now come from one place. The remaining locations are counted only in the worlds where the machine is not already dead, which bounds the figure by the unit's battle value: component_bv is the offensive half of a location and its sum is strictly below the whole.

Vehicle mean falls from 0.1965 to 0.1911, Mek from 0.0283 to a little under it.

What the column has left to say about a vehicle #

Nothing of its own. Every vehicle location is lethal and worth the whole unit, so the non-lethal term is empty and value_destroyed / at_stake is identically p_kill. The corpus says so before the fix does: the median ratio is 1.001.

That is not a defect in the reader. It is what a location's worth is when the location has no separable worth: a vehicle's engine, crew and motive system are reached through a location and not stored in one. The figure that would give the column something independent is BV(intact) - BV(without this location), with the mobility term in it, and it needs a battle value calculator that can be asked about a hypothetical machine. That is Helm's, and the wire stays a plain per-location number so the producer can be swapped with no change on this side.

The hold on fitting value_destroyed over a mixed suite stands, for a narrower reason than before: not that the column is inflated on vehicles, but that it is a copy of p_kill on them and an independent measurement on Meks, so one weight means two things.

p_mission_kill is a separate case and no longer this one: a vehicle's reading comes from motive_immobilised rather than from a location's worth.

No decision moves today #

value_destroyed carries no weight in weights/hand-authored.json, and sds counterfactual <corpus> value_destroyed --phase firing reports 162 of 162 live menus already scored with the column at nought. Engage prices the column, so retune scales it, and scaling nought is nought. A correction to a column at zero weight cannot move an argmax, and this one does not claim to. What it changes is what the column will read when something does weigh it.

Also swept: three things, none urgent #

1. A vehicle's heat capacity is a sentinel, and five columns divide by it. Entity.getHeatCapacity() returns 999 for a combat vehicle - checked across 41 distinct designs from the vehicle suite, energy-armed and ballistic alike, and it is 999 for every one. It is MegaMek saying "this machine has no heat scale", not a capacity.

The bot divides by it. Unit::heat_load is heat / capacity, situation.rs takes heat / capacity.max(1), and end_of_turn_heat subtracts it, so every heat column reads about zero for a vehicle and heat_shutdown_risk and friends read exactly zero. The values are right - a vehicle cannot overheat, so its heat risk genuinely is nought - which is exactly why nobody would notice.

What is wrong is that they are flat within every decision for a vehicle, and nothing distinguishes "this machine has no heat scale" from "this Mek is running cold". That is this epic's own signature and the motive column's failure one level over. It is also the same shape as UPSTREAM.md #6, where WeaponType.getDamage() returns a sentinel that once reached the bot as real damage. Here the arithmetic happens to come out right; that is luck, not design.

2. The vehicle suite over-draws hover, within noise. 84 units drawn against the pool's proportions:

motion pool drawn
Tracked 52.8% 52.4%
Hover 27.3% 35.7%
Wheeled 19.9% 11.9%

At 84 units the standard error on a share near 0.27 is about 0.048, so both gaps are under two standard errors and consistent with sampling. Weight classes spread sensibly - 26 light, 18 medium, 18 heavy, 22 assault, mean 56 tons. So the draw is not broken. It is worth knowing anyway, because hover is the motion whose movement is least tank-like and the one still carrying an open rules question: a suite that is a third hover leans on exactly that.

3. One of this branch's own tests was true of its fixture, not of the rule. a_vehicles_mission_kill_is_motive_and_a_meks_is_not asserted immobilised < entry / 10. From the front a major result needs 12 on 2d6 and the two figures are about 36 apart, so it passed; from a flank it needs 10 and they are about 6 apart, so the same assertion would have failed on a side shot for no reason a reader could see. Now checked against the arc's own odds. The one remaining fixed bound is labelled in the test as a regression bound rather than a rule.

Open questions for jmm #

  • A destroyed turret still takes rolls 10-12. The probe says MegaMek does not redirect them, so the model parks that share on a dead location where it does nothing. That was probed on a bare Tank and the same probe's lethality half returned all-false, which is too clean to trust - so this one is recorded, not settled. It wants a real design in a real game.
  • Does a hovercraft pay the level charge on a drop as well as a climb? The ground-vehicle half of this is answered above; the climb-versus-drop half is not. It wants a probe of what MegaMek does rather than a rule read off the answer that was given.
  • Why does a tank in woods cost us 3 and MegaMek 4? The hovercraft rule was the explanation offered for it and is not the explanation, so the measurement is now unattached. claude/move-legality's, not this branch's.
  • Infantry. Nothing here touches them, and they take damage by a different table again - hitloc now has somewhere to put that, which it did not before.
  • ProtoMeks: the unit files carry 86 of them, between 2 and 15 tons. Core Rules or not?

What the basis does to a target it cannot model #

Swept the whole feature basis for figures that can read zero, and asked of each whether zero is a measurement or an absence. 29 sites, and all but three are measurements. The three that are not share one cause, and it is this epic's.

The discriminator that matters is whether the guard's input varies within one decision. A guard on "no enemies on the board", "no friends", "this machine carries no ammunition" or "this machine cannot move" reads the same for every candidate on the menu, so it cannot move an argmax - it shifts the whole vector and nothing else. Most of the basis is that kind:

  • los_in, los_out, rear_arc_exposure, rear_arc_gain, cover_quality, range_spread, range_band_fit, overlook - no enemy is a real zero share of the enemy force.
  • friend_support, cohesion - the same for friends.
  • ammo_spent, heat_ammo_explosion_risk, heat_mp_penalty - properties of the shooter, identical across its own candidates.
  • concentration, arc_spread - no assignment to be concentrated.
  • elevation_gain already gets this right and says so: no enemy returns the middle of its domain, not zero, "the value a flat board would give". So does Options::UNSEARCHED, which reads a hex's mobility as 1.0 rather than 0.0 when no search was run.

Three vary within a decision, and all three are gated on has_locations() - whether the target is a biped Mek the hit tables describe. The same volley at a Mek and at a tank, measured:

expected damage p_breach expected_criticals value_destroyed
Mek 13.8 0.258 0.018 0.004
tank 13.8 0.000 0.000 0.000

The shot is identical and every column that comes off a location collapses. All three carry positive weights, so the basis prefers shooting the Mek, for a reason that is a modelling gap rather than a tactical fact.

p_kill is the exception and shows what the fix looks like: it falls back to a whole-unit threshold, which is wrong in the way hitloc documents but is honestly wrong rather than silently zero.

Why the two that got it right, got it right #

elevation_gain and Options::UNSEARCHED both map an absence to a benign value rather than to zero, and both could only do that because somebody asked what absence meant at the point of writing the guard. Every one of the sites that got it wrong had a plausible local reason to return zero - no enemies is genuinely no share of the enemy force, no bins genuinely spends no rounds - and zero is a legal value in all of these domains, so an absence can wear it without anything failing. That is the whole difficulty: nothing breaks, and the figure is consumed as a preference.

How often the preference is actually exercised #

Sized on the recorded corpora rather than argued about. The trigger is narrower than "not a Mek": a quadruped is a Mek with eight locations, and it reports HD, CT, RT, LT, FRL, FLL, RRL, RLL - so hitloc::index_of recognises four of the eight and has_locations refuses it. hitloc's own module note predicted exactly this case.

Measured over tactics-hub-mirrored, 300 matches and 8,699 firing decisions, by reading the menus out of the decision logs and asking MegaMek what each of the 380 chassis the bot shot at actually is:

firing decisions 8,699
with a quadruped on the menu 211 (2.43%)
offering a quadruped and a biped 178 (2.05%)
matches involving one at all 14 of 300
chassis that are quadrupeds 7 of 380

The other two corpora with data, corrected-basis-explore and hitloc-handauthored, are 0.00%: their suites are biped-only, and the two scenarios that carry a quadruped today were recorded before it was in them.

2% is a correctness note rather than a live bias, which downgrades the question rather than closing it. For what it is worth and not as a verdict: when a quadruped was on the menu the bot chose it 128 times and chose a biped 77, so the three collapsed columns are a thumb on the scale rather than a veto - but whether those choices were right is not something this count can say, and it is not what it was measured to answer.

expected_criticals against a vehicle: the mechanism question #

Both parts are implemented in this stack. Part one - Projected::expected_criticals: Option<f32>, gated on the body having a critical model rather than on having locations - is in (1/3). Part two, target_described, is in (3/3) below. The analysis is kept because it is the argument for the shape, and because it names the two options that were ruled out and why.

The vehicles work replaces has_locations() with Body::classify() and an optional LocationDamage, which means a naive rebase switches this column on for vehicles and computes them from the Mek critical table. A combat vehicle has systemSlots = 0 against an Atlas's 31 - MegaMek models its criticals as unit-level results reached through a location rather than stored in one - so there is nothing per-location for that table to be about.

What the code actually looks like, because it changes the answer #

The five columns do not fail the same way, and only two of them are mine to gate:

column today, against a target the tables do not describe
p_kill falls back to a whole-unit threshold - wrong, but answers
p_mission_kill explicit 0.0, with a comment
expected_criticals explicit 0.0, with a comment
value_destroyed no gate at all - zero because every live[] is false
at_stake the same, implicitly

That is the fact that decides the mechanism. Making Projected::expected_criticals an Option<f32> is cheap - three real construction sites and one test - and it fixes one fifth of a shared defect, while value_destroyed keeps returning a silent zero from arithmetic that has no gate to change. Five optional fields would be five independent decisions about what None means at a boundary that must yield a number anyway.

The mechanism, if one is wanted #

A feature must produce a number for the fit, so "optional" cannot mean "absent" at the scoring boundary; it can only mean "imputed, and said so". The standard way to carry that into a linear model is the pair: impute a fixed value and add an indicator column recording that you imputed.

One feature - target_described, 1.0 when the hit tables describe the target and 0.0 when they do not - rides on Projected like the rest. One entry in for_each_feature!, one vignette, one field. No existing column changes signature, older weight files lack it and contribute zero, and it covers all five at once rather than the one that happens to be in the rebase.

Its honest limit: an indicator lets a linear fit apply a constant offset for those rows. It cannot apply a different slope, so it is better bookkeeping and not a better model. The real answer is columns that describe a vehicle's worth, which is this epic.

The recommendation #

(a) is out, and my own sweep is what rules it out. That sweep lists expected_criticals at 0.018 for a Mek and 0.000 for a tank as one of exactly three sites where zero is an absence rather than a measurement. Recommending "leave vehicles excluded" would be recommending the defect I catalogued. Struck.

(b) is out on the measurement. A Mek critical table against systemSlots = 0 is a confident wrong number, and confidently wrong is worse than absent.

So (c), and the mechanism is two parts, because one of them alone is cosmetic.

Part one: Projected::expected_criticals: Option<f32>. Three real construction sites and one test. This is where "the producer could not compute it" gets stated, and it is what makes the case greppable rather than a comment.

Part two is the one that matters. A feature must yield a number for the fit, so Option at the producer does not change a single figure the fit sees - it still imputes at the boundary, and imputing zero is numerically identical to today. To make the absence distinguishable, which is the standard this branch spent its time on, the vector has to carry the distinction: an indicator column, target_described, 1.0 when the hit tables describe the target and 0.0 when they do not.

That pair is the textbook handling of a quantity that is undefined for some rows: impute a fixed value, and record that you imputed. The fit learns the average correction for those rows; a log reader can see why four columns are zero at once.

Build part two for the family, not for this column. value_destroyed and at_stake have no gate to change - they are zero because every live[] is false - so an Option on my field fixes one fifth and leaves the other four exactly as silent. One indicator covers all five.

Cost. Part one is small. Part two is one for_each_feature! entry, one field on Projected, wiring in both producers, and a compulsory vignette, which is the real expense - the docs build refuses a feature without one. No weight file changes: an older set lacks the column and contributes zero, so behaviour is identical until something is fitted against it. Neither part needs matches.

Its honest limit. An indicator lets a linear fit apply a constant offset for those rows. It cannot apply a different slope, so it is better bookkeeping rather than a better model. Columns that describe a vehicle's worth are the real answer and they are this epic.

What the zero sweep could not have found #

Worth recording against the sweep rather than leaving it to be rediscovered. It swept 29 zero-reading sites and found 3 absences; the vehicles work has since found 5. The two it missed could not have appeared in it:

  • p_psr_threshold reads mean 0.39, p95 0.99, for a fall Entity.canFall() says a vehicle cannot take. It errs upward. Fixed: the wire carries canFall, and the column is None for a machine no piloting roll applies to.
  • heat_* reads a 999 sentinel as a capacity, which comes out near zero and correct by luck. Fixed: the wire carries tracksHeat, compared against Entity.DOES_NOT_TRACK_HEAT in Observation.java so no sentinel value is written on our side.

The predicate was "reads zero", and that is the costume. The shape is a value arrived at without the thing it claims to measure, which can err upward as easily as down, and can be accidentally right.

The discriminator survives - does the guard's input vary within one decision? - because it is about which absences distort a ranking rather than about how to find them. What needs replacing is the search. Three distinguishable shapes:

  1. An input is absent and something imputes. Greppable, and what the sweep found.
  2. The quantity is undefined for this actor and nothing checks. canFall is this: the input is present and fine, the claim is inapplicable.
  3. A sentinel is read as a value. 999 capacity, DAMAGE_BY_CLUSTER_TABLE, IArmorState::ARMOR_NA.

Only the first is a grep. The second needs each feature's claim tested against a taxonomy of actor kinds, which is an audit rather than a sweep - and is what the vehicles work is doing by hand. The third is greppable if the sentinels are enumerated, and MegaMek's are: they are the same family this branch met in getDamage, and they are worth listing once somewhere.

target_described, and MegaMek's sentinels enumerated #

The indicator is built. hitloc::describes is the one predicate both producers ask - it was duplicated in Volley::has_locations before - and firing::TargetDescribed carries the answer into the vector as a flag.

It is not "is this a Mek", and it stopped being that under it. The column was written when a combat vehicle was undescribed, and the body-plan work beneath it gave vehicles their own tables. describes is now Body::classify(..).is_some(), so a vehicle reads 1.00 and the columns this flags measure against one. What it excludes is a body plan nothing here rolls for: a quadruped, whose FRL/FLL/RRL/RLL no plan knows; a tripod, whose extra CL has no dump; and an observation carrying no locations. The worked figure, its note, the docs and the test that used a tank as the undescribed case were all rewritten to a quadruped for that reason.

And it does not explain every zero it sits beside. expected_criticals is None for a vehicle and for an undescribed target, because a vehicle is described by a hit table and still has no critical model - MegaMek resolves its criticals as unit-level results and it reports systemSlots = 0. So a vehicle reads target_described 1.00 and expected_criticals zero, and nothing in the basis yet says why.

What it covers, named in the column's own docs so its purpose is legible: expected_criticals, which is an Option and imputes zero; and p_mission_kill, value_destroyed, at_stake and p_breach, which have no gate at all and are zero because every live flag in the profile is false. That asymmetry is the argument for one indicator over four Option fields. p_kill is the exception and falls back to a whole-unit threshold, which is wrong in a documented way rather than silent.

What it does not cover, also in the docs, because someone will assume it did: p_psr_threshold prices a fall Entity.canFall() says a vehicle cannot take and errs upward - an indicator cannot subtract a claim that should never have been made - and the heat_* group reads a 999 sentinel as a capacity, which comes out near zero and correct by luck. Both want the quantity not computed, not a flag saying it was. Both are now done that way, each carrying the answer on the wire rather than deriving it: canFall gates the cliff to None, and tracksHeat gates the heat scale.

The honest limit is in the column's doc comment: a constant offset for those rows, not a different slope. Better bookkeeping, not a better model. The real answer is columns that describe a vehicle's worth.

A file with a regeneration command is not a file to edit #

That is the whole rule, it costs nothing to apply, and it was learned the expensive way twice in one night.

When behaviour changes, the committed renders of that behaviour change with it, and nothing in the reasoning about the code prompts you to notice.

This stack changed what hitloc::describes answers. Five things said the old answer, and only one of them was code:

  • the column's own doc comment
  • the worked figure's per-case note
  • docs/features/target_described.svg, a checked-in render
  • docs/FEATURES.md, generated prose
  • a passing test asserting a tank reads 0.00

The test is the worst of them and it is worth saying why. It did not merely fail to catch the change - it existed to defend the old belief, so the branch carried the wrong claim pinned in place by something green. A passing assertion reads as evidence rather than as a record of what somebody believed at the time, so the next person trusts it harder than an untested assumption.

And the render caught the author of this section out while writing it. The FEATURES.md paragraph was edited by hand, the generator was then re-run for the SVG, and it overwrote the correction and restored the false text. The real source was a reading: string in examples/vignettes/firing.rs. Editing a generated file is not a small mistake that gets corrected later; it is a correction that gets silently reverted by the next build, which is worse than not making it.

The tell was free and sitting in the failure message the whole time: tests/vignettes.rs prints cargo run -j 2 -p sds-core --example featuredoc when the docs go stale. Before editing a docs file, check whether anything writes it.

The sentinel list, read out of the jar #

The third shape from the sweep write-up - a sentinel read as a value - is greppable only if the sentinels are enumerated, so here they are:

constant value met as
WeaponType.DAMAGE_BY_CLUSTER_TABLE -2 a rack's damage; reached the bot as negative expected damage
WeaponType.DAMAGE_VARIABLE -3 a Snub-Nose PPC and a VSP laser
WeaponType.DAMAGE_SPECIAL -4 not yet met
WeaponType.DAMAGE_ARTILLERY -5 an Arrow IV, once minus five points
Entity.DOES_NOT_TRACK_HEAT 999 a vehicle's heat capacity, the heat_* group
Entity.LOC_NONE -1 a location a unit does not have
Entity.LOC_DESTROYED -2 transfer, where damage goes when a location is gone
Entity.UNLIMITED_JUMP_DOWN 999 not yet met
armour states -1, -2, -3 ARMOR_NA, ARMOR_DOOMED, ARMOR_DESTROYED, already handled by points()

Every one is a small negative or 999, and every one is a legal-looking number in its own field - which is why getDamage() returning -2 became a negative expectation and 999 became a heat capacity nothing could exceed. A reader that takes a raw MegaMek integer without naming which sentinels it can be is the third shape, and this table is what makes that greppable rather than an audit.

destroyLocation cannot produce a BV differential #

Outside a Game, Entity.destroyLocation marks a location doomed and stops: armour and internal go to IArmorState.ARMOR_DOOMED (-2) rather than ARMOR_DESTROYED (-3), isLocationBad stays false, no criticals are destroyed, and walk and run MP do not change. Criticals and MP recalculation happen later, in the game manager's damage resolution.

So BV(intact) - BV(intact, location destroyed) cannot be computed this way: it measures armour and structure pools being zeroed, and never reaches the movement term. It throws nothing and returns well-ordered numbers.

new Game() throws ExceptionInInitializerError, so attaching one is not the easy fix it looks like.