bayes for days
sds docs PRINCESS.md
4.0 kB

Princess #

Why this repository exists. MegaMek's own bot has been hand-maintained for over a decade and still plays badly in ways that are structural rather than careless. Read from the 0.51.0 source; megamek/client/bot is ~30k lines, of which Princess.java is 4,306, FireControl.java 3,766 and BasicPathRanker.java 2,061.

Most of that bulk is rules coverage, not intelligence. The intelligence is maybe 2k lines.

How it decides #

  1. getEntityToMove() picks a unit by a scalar calculateMoveIndex.
  2. PathEnumerator generates every legal path.
  3. BasicPathRanker.rankPath() scores each and takes the argmax:
utility = -fallMod + braveryMod - aggressionMod - herdingMod
        + movementMod - crowdingTolerance - facingMod - selfPreservationMod
utility -= utility * offBoardMod
  1. FireControl separately scores firing plans by damage / crits / kill / heat.

The four structural faults #

Greedy, one unit and one turn at a time. No joint plan. Coordination is approximated by herdingMod, a penalty on distance from the friendly centroid — which is exactly the conga line. Nothing in the model can express "you hold that ridge while I flank".

Damage dealt is a max; damage taken is a sum. In rankPath:

if (damageEstimate.firingDamage < eval.getMyEstimatedDamage()) {
    damageEstimate.firingDamage = eval.getMyEstimatedDamage();
}
...
expectedDamageTaken += eval.getEstimatedEnemyDamage();

A hex that can shoot three enemies scores the same as one that can shoot the best of them, while every extra arc it enters is charged in full. Princess cannot see the value of a crossfire, and is systematically pessimistic about advancing.

The terms are not commensurable. fallMod is a probability times a slider, aggressionMod is hexes times a slider, facingMod is a facing difference times 50, braveryMod is in damage points. Nothing is normalised, so which term dominates depends on map size, unit count and unit type. That is why tuning a slider that fixes one map breaks another, and the 1–10 sliders in BehaviorSettings are the user-facing surface of the same defect.

No enemy model. evaluateUnmovedEnemy guesses; nothing projects where a threat will be next turn. Princess cannot reason "I will be shot from that hill".

Corroborating symptoms #

  • Cost is O(paths × enemies × weapons). Princess keeps a moveEvaluationTimeEstimate and chats an ETA to the players before thinking.
  • Four path rankers, one (UtilityPathRanker) an unfinished rewrite.
  • Physicals still delegate to the pre-Princess PhysicalCalculator, with the comment "the original bot's physical options seem superior".
  • ~30 chat-command classes, so live tuning by a human is the expected workflow.
  • megamek/ai/dataset/ exists solely to serialise game states — the maintainers are collecting training data, which is agreement that hand-tuning has topped out.

What this repository does differently #

Princess sds
one unit at a time units propose, forces decide — HIERARCHY.md
max dealt, sum taken features summed over targets, normalised — plan/features.md
incommensurable terms every feature declares a normalisation type
no enemy model plan/beliefs.md
every legal path ~20 curated candidates; the cap is the performance design
reaches into Entity one documented observation — PROTOCOL.md

Learning from it anyway #

crates/sds-bot/src/imitate.rs reconstructs Princess's movement from two consecutive observations and records it as training rows, to bootstrap the weight vector: a label per decision instead of one per match.

It inherits the faults above. A weight fitted from Princess's choices has learned Princess's opinion of a crossfire, which is the second fault on this page. It is a starting position and has to be re-measured against Princess before it means anything - see plan/training.md.