bayes for days
Rust 57%
Python 24%
Java 17%
Module Management System 2%
Shell <1%
CSS <1%
Dockerfile <1%

README.md

sds #

SDS is a bot for playing BattleTech matches via MegaMek, implemented as a Rust-based hierarchical decision-making system with an interpretable linear utility model.

Project goals #

SDS is being developed as the bot opponent for matches on lance.blue. It is intended to serve as a complete replacement for the existing MegaMek bot, "Princess". The project goals attempt to address several current limitations:

SDS is a project for humans. #

SDS is developed using the latest AI agents, but it must remain accessible to human contributors. AI agents are not used during gameplay. Any AI-generated code must remain human-readable, well-structured, and testable. The evaluation harness and training capabilities must remain runnable by humans with reasonable compute resources.

SDS is fun to play against. #

This is very subjective, of course. Some examples, though:

  • Plays "like a human", neither hyper-optimized nor clueless.
  • Uses unit strengths effectively, so that matches feel different.
  • Decisions should be legible to the human player, but occasionally challenging or surprising.
  • No stalling, kiting, or otherwise drawing out gameplay.
  • No "weak points" that trivialize beating the bot.

SDS capabilities strictly exceed those of the current MegaMek bot. #

This means that SDS should:

  • have a >50% win rate against Princess
  • calculate turns faster (on average)
  • calculate turns more reliably (worst-case)
  • support more unit types
  • support more interactions
  • support more types of gameplay or scenarios

SDS makes decisions at multiple levels. #

Units and forces are both represented as domain objects by SDS, and there is a collective planning layer. This enables advanced capabilities like spotting for artillery, flanking opponents, or mounting transports.

SDS decisionmaking is highly adjustable. #

SDS can be customized by applying a "lens" that adjusts features. Lenses are composable, allowing the AI to express concepts like "Green, Mercenary, Tank Commander" or "Aggressive, Veteran, Clan Star Captain". This allows scenarios to adjust either raw difficulty, or overall tendency to pick certain tactics/behaviors, in a way that doesn't make the bot feel "stupid".

SDS decisionmaking must remain explainable to humans. #

SDS makes decisions at both unit and force levels using machine learning. However, unlike "AI agents", all features are repeatable, interpretable, and explainable: the features are hand-written and named, and only the weights that combine them are fitted. Every feature must be extensively documented, including with diagrams generated from the code that measures it.

SDS also ships with a live viewer for its decisions, which can be run alongside a match to visualize how the bot understands the current match and what decision it plans to make next.

SDS is deterministic and reproducible. #

Two instances of SDS playing each other on the same scenario with the same seed should be guaranteed to replay to the same outcome, regardless of factors like parallelism or system load.

Design #

The learning is offline: the features are hand-written and named, and only the weights that combine them are fitted from recorded play. tokio provides safe parallelism, and a small Java bridge reads and writes MegaMek. The codebase is almost entirely agent-authored, though the design has had extensive and ongoing human input.

More details can be found at docs/ (WORK IN PROGRESS)

CLAUDE DOCS BELOW: WATCH OUT!! #

What works today #

A lance-vs-lance match, headless, with any mix of Princess and external bot seats, and the statistics to say whether a difference between two bots is real or is dice.

./scripts/build.sh                                  # once
./sds.sh one                                        # Princess vs Princess
./sds.sh one --sds-seat South \
    --bot "python3 /work/bots/random_bot.py"        # bot vs Princess
./sds.sh control --games 40                         # the harness self-test
./sds.sh bench --games 40 \
    --bot "python3 /work/bots/random_bot.py"        # the measurement
./sds.sh play                                       # you, at MegaMek, against the bot
./sds.sh view runs/<run-dir>                        # what the bot thought
./sds.sh one --replay --sds-seat South \
    --bot /work/target/release/sds-bot              # + a GIF of the match
./sds.sh los-dump scenarios/suite/*.mms \
    --out crates/sds-core/tests/corpus/los.jsonl    # regenerate the LOS corpus
./sds.sh pathfind-dump                              # regenerate the movement corpus
./sds.sh arc-dump                                   # the arc corpus, which is NOT committed

A 4v4 on one map sheet takes about 90 seconds.

view turns the decision log an SDS seat writes next to its result into one self-contained HTML file: each force's stance and how long it has held it, what last changed its mind, every proposal a unit offered with the one that was taken marked, and the per-phase counters. Units are named the way MegaMek names them - Turkina C (#4) - and --watch keeps the page current while a match is still being played.

one --replay writes <match-tag>.gif beside the result: MegaMek's minimap, one frame per round. It is the game's own view of the match rather than ours.

Why the harness came first #

MegaMek's own bot, Princess, has been hand-tuned for over a decade and still walks into water. The reason is structural rather than careless — it scores every legal path for one unit with a weighted sum of terms measured in incommensurable units, takes the argmax, and repeats for the next unit. docs/PRINCESS.md has the specifics, including the max-dealt sum-taken asymmetry that makes it blind to a crossfire.

None of that is fixable without a way to tell whether a change helped. So: harness first, bot second.

Running an experiment #

One idea, one branch, one comparison in the pull request:

git checkout -b claude/flank-weighting origin/main
# change the bot
./scripts/build.sh
./sds.sh bench --games 60 --against main

That writes comparison.md, which is the PR body. See docs/EXPERIMENTS.md — and note that the comparison refuses a baseline measured under different rules, and says in words when a difference is not supported by the sample.

Fitting M #

The bot scores a candidate as a dot product: a weight per named feature. The features are hand-written and stay that way; only the weight vector M is fitted.

Each one is explainable in a sentence, and docs/FEATURES.md is that sentence plus a small situation the feature was measured in and the number it came out at. The page and its figures are generated, and a test fails if either has gone stale:

cargo run -j 2 -p sds-core --example featuredoc            # the markdown and the figures
cargo run -j 2 -p sds-core --example featuredoc -- --html  # and docs/features.html
cargo run -j 2 -p sds-core --example featuredoc -- --check # are they current?

--html writes the whole catalogue as one browsable page - grouped by family, figures inlined, with a light/dark switch that reaches inside them. Open docs/features.html in a browser; it needs nothing beside it. That file is gitignored, because every byte of it is derived from the markdown and the SVGs, which are not.

A new feature lands with its worked example or it does not compile. There is no list to add a name to: Documented::example has no default, and the docs build expands the same for_each_feature! that builds the catalogue.

./sds.sh bench --games 300
./sds.sh corpus save runs/<the run> overnight
./sds.sh train overnight --out weights/fitted.json

train plays nothing. It reads the *.decisions.jsonl a run left behind — each candidate's feature vector, which one was taken — joins each log to its match result on the match tag, and fits one weight per feature by least squares on the difference between the chosen candidate and the ones it beat.

bench --imitate records the opposing seat's moves as well, reconstructed from consecutive observations, and train --imitation fits them: a label per decision instead of one per match. That clones the opposing bot's opinions along with its moves — plan/training.md says what it inherits, and why it is a bootstrap rather than the goal.

Keeping a corpus #

runs/ is gitignored and sits inside a worktree, so a corpus there is one git worktree remove --force from gone — that is how 360 matches and 21937 decisions were lost, with nothing left on disk that even named them.

corpus save copies a run — copies, never links — into corpora/<name>/data/ beside the main checkout, where removing a worktree cannot reach it, and writes corpora/<name>/manifest.json next to it. The data is ignored and the manifest is committed, so a corpus that is gone is still known to have existed and is describable: the commit, the epoch, the suite fingerprint, the bot command, the weights it played, whether exploration was on and at what rate, the counts, and the sorted list of every feature name in it.

./sds.sh corpus save runs/20260819T192254Z-bench overnight --notes "why"
./sds.sh corpus list
./sds.sh train overnight --out weights/fitted.json

train takes a corpus name, a run directory or a single log. Given a manifest it checks it first, on the rule sds/baseline.py already applies to a benchmark: what changes the meaning of a number travels with the number. A different epoch, a --feature the corpus does not have, or data the manifest does not describe stops the fit and says which; --stale-ok fits it anyway. A feature the baseline M uses that the corpus never recorded is a warning printed before and after the report rather than a refusal — every corpus recorded today is missing the positional features, and a guard that fires every time is one nobody reads.

What the fit refuses #

Two things it refuses to do. A local feature — one min-maxed across a single decision's own candidates — never gets a weight, because a weight fitted against one has learned a board rather than the game; the Rust side makes that a compile error and train enforces the same rule from the local list each log record carries. And it prints, loudly, any weight whose sign contradicts the feature's own one-sentence description: overkill coming out positive is a broken label, not a discovery.

Then play the fitted set. The bot takes --weights, and the harness passes a bot command through whole, so it rides along in --bot:

./sds.sh bench --games 60 --against main \
  --bot "/work/target/release/sds-bot --weights /work/weights/fitted.json"

Without the flag the bot plays weights/hand-authored.json, which is the permanent baseline: a fitted set that only beats an arbitrary one has proven nothing. That file and Weights::hand_authored are checked against each other by a test, and sds-bot --print-weights writes it out again.

A weights file that cannot be read is fatal and says why — a missing file, a name no feature answers to, a local feature that may not carry a weight. It also says which features the file leaves at zero that the baseline weights. Silently ignoring one would make a training run look like it worked and change nothing.

The label is the match's final BV differential, given to every decision in that match, so a good move in a lost match is labelled bad. It averages out over matches and does not over decisions. Whether a fitted M is actually better is a bench of the bot carrying it, --against the same baseline as the hand-authored one — the fit itself proves nothing.

The commands, and why they exist #

  • control plays Princess against Princess. It must come out near 50/50. If it does not, the harness is biased — seat order, deployment edge, the RNG, bridge latency — and every number it has ever printed is suspect. Run it after touching the harness, before believing anything else.
  • bench plays a bot against Princess, alternating which faction the bot takes, and reports a win rate with a Wilson interval.
  • one plays a single match and prints everything, for looking at. --replay also writes the match as an animated GIF next to the result, one frame per round, drawn by MegaMek's own minimap. It costs an X server inside the container and is off everywhere else: a benchmark does not want a GIF per match.
  • play puts you in one of the seats, at MegaMek's own client, and leaves the same result, round report and decision log behind. For the failures a win rate cannot describe. docs/PLAYING.md.

baselines lists what has been recorded, corpus stores and lists training corpora, and clean kills match containers left by a harness that was killed rather than interrupted.

bench and one play /work/target/release/sds-bot unless --bot says otherwise, and print which bot they are running in their first line. The floor bot is still there and still worth running — --bot "python3 /work/bots/random_bot.py" — but you have to ask for it. It used to be the default, and twice now a run of it has been read as a run of the change under test; see plan/harness.md for both.

Reading one decision #

./sds.sh explain runs/<the run> --phase FIRING

Every decision line in a match's .decisions.jsonl already carries the whole menu the bot ranked: each candidate's label, its measured features and the value the argmax used. sds-bot writes the weight vector beside the log, and explain prints the two together — one table per decision, a column per candidate, a row per feature, showing the raw value, the weight and the product. The default the bot could have taken instead (hold fire, stand still) is a column like any other, so "it chose badly" and "everything on the menu was worse than doing nothing" are different pictures.

sds view renders a whole match as HTML; explain is for the one decision that made no sense.

Watching a match as it is played #

./sds.sh watch runs/<the run>          # or one .decisions.jsonl
# then open http://127.0.0.1:8737/

watch tails the decision log by byte offset and serves it to a page meant to sit beside the MegaMek window. The bot embeds no server and publishes no port: SdsClient flushes every decision as it writes it, and the run directory is inside the repository the match container mounts, so the file on disk is already live. The same command replays a finished run — a log nobody is appending to is a log whose tail is empty.

The page is a hex map of the board the scenario named, read from $MM_HOME (nothing is vendored), with:

  • every proposal's end hex coloured by rank, not by absolute value. What is ranked is a dropdown: the proposal's value, the damage it deals, the damage it takes, its risk, or any single feature it measured.
  • terrain drawn over the colouring, one glyph per kind, and cliff edges drawn on the hex edge they belong to.
  • each unit's footprint outlined along real hex edges, a chevron where it stands, a ring on the hex it took and a dashed line between the two.
  • beside the map: which unit acted and which were eligible, why the force replanned, the stance it was holding, and the chosen candidate against its runner-up broken into terms — the same arithmetic explain prints.

The two failures view exists to separate are counted in the panel: "the lance chose badly" (a better proposal was offered and passed on) against "nobody offered anything" (every proposal worth about nothing).

--port moves it — several agents share this machine — and --max-proposals sets how many of a unit's proposals reach the page, best first, with the chosen one always kept. The bot still reports every candidate it scored; the cut is the reader's, and the page says how many it dropped.

How the numbers avoid lying #

  • Mirrored forces. scenarios/mirror-lance.mms gives both sides the same four machines and the same pilots. A 200 BV edge swamps any tactical difference a new bot is likely to make.
  • Sides alternate. No two deployment edges are equally good, so the bot under test plays each of them half the time, and the tally follows the bot rather than the corner of the map.
  • Seeded dice. Compute.setRNG takes a seeded generator, so two runs draw the same numbers. This does not make a match deterministic — see bridge/sds/SeededRandom.java for what it does and does not buy.
  • Undecided games are not half a win. A match that hits the round limit has no winner and is excluded from the rate and reported separately. Counting it as a draw would make a bot that refuses to engage look average.
  • The winner is MegaMek's, never inferred. Not "the side with more units left" — that turns a benchmark of who wins into a benchmark of who hides.
  • Confidence intervals, always. 12-8 is a 60% win rate whose interval runs from 39% to 78%. Separating a five-point difference from noise takes about 385 decided games; the tool prints that reminder next to every result.

The bot #

crates/ holds the real bot: units propose, forces wait for all of their units and then commit to a stance, and a node is an address so any level can be moved to another machine. docs/HIERARCHY.md has the design and, more usefully, the list of what it does not do yet.

How a bot plugs in #

A bot is a process. It reads newline-delimited JSON on stdin and writes it on stdout — see docs/PROTOCOL.md. It answers a decision, or it passes and takes the harness's default.

No SDS decision is ever played by Princess. SdsClient extends BotClient, shares no code with Princess, and never falls back to it — not on a pass, a timeout, or a crash. The reason is debugging: if a seat sometimes plays Princess's move, no line in a match log tells you whose decision you are looking at, and every investigation starts by working out whether the thing you are staring at is even yours.

The defaults are ours and deliberately inert — stand still, hold fire, deploy in the first legal hex. A bot that answers nothing therefore stands where it landed and is shot to pieces, which is the correct and legible outcome rather than a respectable opponent wearing our name.

The line between what may be borrowed and what may not: MegaMek's rules yes, MegaMek's bot no. WeaponAttackAction.toHit is rules — every human player has that number on screen before choosing — so the observation carries it. The hex ranking in BotClient.getStartingCoordsArray is tactics, and is not called even though SdsClient inherits it.

The reference bots in bots/ are floors rather than opponents; none is meant to be good:

  • bots/passthrough_bot.py passes on everything, so its lance deploys, stands still and never fires. That is the floor beneath the floor, and the fastest check that the bridge is carrying decisions at all: if a real bot scores the same as this, its answers are not arriving.
  • bots/random_bot.py walks somewhere legal without looking at the enemy, shoots everything at whatever it is most likely to hit, and never thinks about heat. A bot that cannot beat it is not an improvement whatever it scores against Princess.

Requirements #

Docker. Nothing else — there is no host JDK and none is wanted, following the pattern in helm/bridge and ~/.cache/mul-build.

MegaMek is mounted at run time from an extracted release, never vendored: MM_HOME defaults to ~/.cache/mul-build/megamek. No MegaMek data lives in this repository.

Hooks #

uv tool install prek
prek install

Licensing #

The bridge is compiled against a stock MegaMek release and links its GPL-3.0 code. MegaMek's data is CC BY-NC-SA 4.0 and is not redistributed here.