Prediction, written before the arms were run #
Recorded 2026-08-30, before either arm of the berserker-lance benchmark had produced a number.
Design. Six mirrored 4v4 scenarios at BV 9000, every one of whose twelve
forces classifies Berserker/Close and opens on advance. Self-play, so all
eight machines in a match are under the arm being tested. Two arms over the
same (scenario, seed) pairs: --tactics engage pins every force on engage and
is the control; --tactics live lets the formation's opening stand. Melee rate
is physical attacks MegaMek resolved, from its own game log, per match.
What I expect.
- The live arm lands more physical attacks per match than the control. It
has to, or
Advanceis not doing the one thing it names. I do not have a prior for the size, because nothing in this repository has ever measured a melee rate - the bot could not throw a punch until this branch. - The control arm is not zero. Engage still prices a fist now that
ceilingandprojected_shotsagree, so a machine that happens to be in contact will swing. If the control is exactly zero I should suspect the arm switch rather than celebrate the contrast. - The live arm's matches run longer in rounds. Closing to contact takes turns, and a force that spends them walking is a force not shooting.
- Both arms answer close to 100% of decisions. If the live arm's
answereddrops, the tactic is producing menus the bot cannot resolve and any melee difference is an artefact of that instead.
What would falsify the whole thing. A difference smaller than the spread between seeds within one arm. That spread is the noise floor and is computed from the same 24 matches, not asserted - if the paired difference does not clear it, the honest report is that these arms cannot be separated at this sample size.
What the arms came out at #
Everything below this line was produced after the prediction above was written.
Three arms of 24 matches on scenarios/berserker-lance - six mirrored 4v4s at
BV 9000 whose twelve forces all classify Berserker/Close - self-play, so all
eight machines in a match are on the arm under test, --explore 0, paired on
identical (scenario, seed) pairs. The jar and the bot binary were checksummed
before the batch and verified unchanged after it.
The rate #
| arm | matches | physicals | per match | median | sd |
|---|---|---|---|---|---|
engage (control) |
24 | 161 | 6.71 | 4.0 | 6.94 |
live (advance) |
24 | 252 | 10.50 | 10.0 | 5.88 |
Paired difference +3.79 physicals a match, bootstrap 95% CI [+1.67, +5.83], 18 pairs up against 5 down and 1 tie (sign test p ~ 0.011). Dropping the one stalled match strengthens it: +4.04, CI [+1.87, +6.13].
answered was 3646/3732 (97.7%) under engage and 2643/2679 (98.7%) under
advance. Both bots played the matches they are credited with.
The noise floor, measured rather than assumed #
plan/harness.md records that a seed does not reproduce a match, so within-arm
spread is the wrong instrument: the floor is what the same configuration run
twice produces. The third arm is the control arm again, same seeds, same jar.
| comparison | paired difference | 95% CI | up/down/tie |
|---|---|---|---|
| floor: engage vs engage | +0.17 | [-1.17, +1.29] | 5 / 3 / 16 |
| signal: advance vs engage | +3.79 | [+1.67, +5.83] | 18 / 5 / 1 |
The signal is 22 times the floor and the two intervals do not overlap - the floor's upper bound is +1.29 and the signal's lower bound is +1.67. The difference is larger than what repeating the same configuration produces, which is the thing that had to be shown.
The prediction that was wrong #
I predicted advance would make matches run longer in rounds, because closing takes turns. They run shorter: 7.4 rounds against 11.5. The mechanism is backwards from how I stated it - closing to contact does not lengthen a fight, it ends one, because melee is decisive once you arrive.
That makes the headline an understatement. Per round rather than per match:
| arm | physicals per round |
|---|---|
| engage | 0.585 |
| advance | 1.416 (+142%) |
Advance throws 56% more physicals a match while playing 36% fewer rounds to do it.
Self-play is not reproducible either, and that is new #
plan/harness.md blames the divergence on Princess's precognition thread
drawing from Compute's RNG out of order. There is no Princess in a self-play
run, so the third arm answers a question nobody had asked: same binary, same
seeds, same scenarios, run twice.
16 of 24 matches came back identical on every field. The other 8 differed in end round, physicals, decisions and answered, and one differed in outcome. So the divergence is not entirely Princess's, and something in our own stack or in MegaMek's server threading is nondeterministic. Two thirds of matches reproducing exactly is also more structure than "a seed does not reproduce a match" suggests, and it is worth someone establishing which third is which.
This is reported for plan/harness.md rather than as part of the melee result:
the comparison above does not depend on it, because the floor was measured under
exactly the nondeterminism the signal was measured under.
What this does not say #
- Nothing about win rate. Advance closes to contact and throws punches; that it wins more is unmeasured and not claimed.
- Nothing about any suite but this one. Berserker lances are rare at 4v4 - 5.6% of forces, against 29% at 8v8 - and this suite was built by filtering a BV 9000 pool for scenarios whose forces both classify. It is the population the tactic is for, not a general one.
- The bot cannot decline a physical. Every physical-phase decision in both arms declared one (verified 17/17 in the logs, and it is why the decision count and the resolved count are equal). A kick that misses can force a piloting roll, and nothing weighs that.