sds #
SDS is a bot for playing BattleTech matches via MegaMek, implemented as a Rust-based hierarchical decision-making system with an interpretable linear utility model.
Project goals #
SDS is being developed as the bot opponent for matches on lance.blue. It is intended to serve as a complete replacement for the existing MegaMek bot, "Princess". The project goals attempt to address several current limitations:
SDS is a project for humans. #
SDS is developed using the latest AI agents, but it must remain accessible to human contributors. AI agents are not used during gameplay. Any AI-generated code must remain human-readable, well-structured, and testable. The evaluation harness and training capabilities must remain runnable by humans with reasonable compute resources.
SDS is fun to play against. #
This is very subjective, of course. Some examples, though:
- Plays "like a human", neither hyper-optimized nor clueless.
- Uses unit strengths effectively, so that matches feel different.
- Decisions should be legible to the human player, but occasionally challenging or surprising.
- No stalling, kiting, or otherwise drawing out gameplay.
- No "weak points" that trivialize beating the bot.
SDS capabilities strictly exceed those of the current MegaMek bot. #
This means that SDS should:
- have a >50% win rate against Princess
- calculate turns faster (on average)
- calculate turns more reliably (worst-case)
- support more unit types
- support more interactions
- support more types of gameplay or scenarios
SDS makes decisions at multiple levels. #
Units and forces are both represented as domain objects by SDS, and there is a collective planning layer. This enables advanced capabilities like spotting for artillery, flanking opponents, or mounting transports.
SDS decisionmaking is highly adjustable. #
SDS can be customized by applying a "lens" that adjusts features. Lenses are composable, allowing the AI to express concepts like "Green, Mercenary, Tank Commander" or "Aggressive, Veteran, Clan Star Captain". This allows scenarios to adjust either raw difficulty, or overall tendency to pick certain tactics/behaviors, in a way that doesn't make the bot feel "stupid".
SDS decisionmaking must remain explainable to humans. #
SDS makes decisions at both unit and force levels using machine learning. However, unlike "AI agents", all features are repeatable, interpretable, and explainable: the features are hand-written and named, and only the weights that combine them are fitted. Every feature must be extensively documented, including with diagrams generated from the code that measures it.
SDS also ships with a live viewer for its decisions, which can be run alongside a match to visualize how the bot understands the current match and what decision it plans to make next.
SDS is deterministic and reproducible. #
Two instances of SDS playing each other on the same scenario with the same seed should be guaranteed to replay to the same outcome, regardless of factors like parallelism or system load.
Design #
The learning is offline: the features are hand-written and named, and only the weights that combine them are fitted from recorded play. tokio provides safe parallelism, and a small Java bridge reads and writes MegaMek. The codebase is almost entirely agent-authored, though the design has had extensive and ongoing human input.
More details can be found at docs/ (WORK IN PROGRESS)
CLAUDE DOCS BELOW: WATCH OUT!! #
What works today #
A lance-vs-lance match, headless, with any mix of Princess and external bot seats, and the statistics to say whether a difference between two bots is real or is dice.
./scripts/build.sh # once
./sds.sh one # Princess vs Princess
./sds.sh one --sds-seat South \
--bot "python3 /work/bots/random_bot.py" # bot vs Princess
./sds.sh control --games 40 # the harness self-test
./sds.sh bench --games 40 \
--bot "python3 /work/bots/random_bot.py" # the measurement
./sds.sh play # you, at MegaMek, against the bot
./sds.sh view runs/<run-dir> # what the bot thought
./sds.sh one --replay --sds-seat South \
--bot /work/target/release/sds-bot # + a GIF of the match
./sds.sh los-dump scenarios/suite/*.mms \
--out crates/sds-core/tests/corpus/los.jsonl # regenerate the LOS corpus
./sds.sh pathfind-dump # regenerate the movement corpus
./sds.sh arc-dump # the arc corpus, which is NOT committed
A 4v4 on one map sheet takes about 90 seconds.
view turns the decision log an SDS seat writes next to its result into one
self-contained HTML file: each force's stance and how long it has held it, what
last changed its mind, every proposal a unit offered with the one that was
taken marked, and the per-phase counters. Units are named the way MegaMek names
them - Turkina C (#4) - and --watch keeps the page current while a match is
still being played.
one --replay writes <match-tag>.gif beside the result: MegaMek's minimap,
one frame per round. It is the game's own view of the match rather than ours.
Why the harness came first #
MegaMek's own bot, Princess, has been hand-tuned for over a decade and still walks into water. The reason is structural rather than careless — it scores every legal path for one unit with a weighted sum of terms measured in incommensurable units, takes the argmax, and repeats for the next unit. docs/PRINCESS.md has the specifics, including the max-dealt sum-taken asymmetry that makes it blind to a crossfire.
None of that is fixable without a way to tell whether a change helped. So: harness first, bot second.
Running an experiment #
One idea, one branch, one comparison in the pull request:
git checkout -b claude/flank-weighting origin/main
# change the bot
./scripts/build.sh
./sds.sh bench --games 60 --against main
That writes comparison.md, which is the PR body. See
docs/EXPERIMENTS.md — and note that the comparison
refuses a baseline measured under different rules, and says in words when a
difference is not supported by the sample.
Fitting M #
The bot scores a candidate as a dot product: a weight per named feature. The
features are hand-written and stay that way; only the weight vector M is
fitted.
Each one is explainable in a sentence, and docs/FEATURES.md is that sentence plus a small situation the feature was measured in and the number it came out at. The page and its figures are generated, and a test fails if either has gone stale:
cargo run -j 2 -p sds-core --example featuredoc # the markdown and the figures
cargo run -j 2 -p sds-core --example featuredoc -- --html # and docs/features.html
cargo run -j 2 -p sds-core --example featuredoc -- --check # are they current?
--html writes the whole catalogue as one browsable page - grouped by family,
figures inlined, with a light/dark switch that reaches inside them. Open
docs/features.html in a browser; it needs nothing beside it. That file is
gitignored, because every byte of it is derived from the markdown and the SVGs,
which are not.
A new feature lands with its worked example or it does not compile. There is no
list to add a name to: Documented::example has no default, and the docs build
expands the same for_each_feature! that builds the catalogue.
./sds.sh bench --games 300
./sds.sh corpus save runs/<the run> overnight
./sds.sh train overnight --out weights/fitted.json
train plays nothing. It reads the *.decisions.jsonl a run left behind — each
candidate's feature vector, which one was taken — joins each log to its match
result on the match tag, and fits one weight per feature by least squares on the
difference between the chosen candidate and the ones it beat.
bench --imitate records the opposing seat's moves as well, reconstructed from
consecutive observations, and train --imitation fits them: a label per decision
instead of one per match. That clones the opposing bot's opinions along with its
moves — plan/training.md says what it inherits, and why it is a bootstrap
rather than the goal.
Keeping a corpus #
runs/ is gitignored and sits inside a worktree, so a corpus there is one
git worktree remove --force from gone — that is how 360 matches and 21937
decisions were lost, with nothing left on disk that even named them.
corpus save copies a run — copies, never links — into corpora/<name>/data/
beside the main checkout, where removing a worktree cannot reach it, and
writes corpora/<name>/manifest.json next to it. The data is ignored and the
manifest is committed, so a corpus that is gone is still known to have existed
and is describable: the commit, the epoch, the suite fingerprint, the bot
command, the weights it played, whether exploration was on and at what rate, the
counts, and the sorted list of every feature name in it.
./sds.sh corpus save runs/20260819T192254Z-bench overnight --notes "why"
./sds.sh corpus list
./sds.sh train overnight --out weights/fitted.json
train takes a corpus name, a run directory or a single log. Given a manifest it
checks it first, on the rule sds/baseline.py already applies to a benchmark:
what changes the meaning of a number travels with the number. A different epoch,
a --feature the corpus does not have, or data the manifest does not describe
stops the fit and says which; --stale-ok fits it anyway. A feature the baseline
M uses that the corpus never recorded is a warning printed before and after the
report rather than a refusal — every corpus recorded today is missing the
positional features, and a guard that fires every time is one nobody reads.
What the fit refuses #
Two things it refuses to do. A local feature — one min-maxed across a single
decision's own candidates — never gets a weight, because a weight fitted against
one has learned a board rather than the game; the Rust side makes that a compile
error and train enforces the same rule from the local list each log record
carries. And it prints, loudly, any weight whose sign contradicts the feature's
own one-sentence description: overkill coming out positive is a broken label,
not a discovery.
Then play the fitted set. The bot takes --weights, and the harness passes a
bot command through whole, so it rides along in --bot:
./sds.sh bench --games 60 --against main \
--bot "/work/target/release/sds-bot --weights /work/weights/fitted.json"
Without the flag the bot plays weights/hand-authored.json, which is the
permanent baseline: a fitted set that only beats an arbitrary one has proven
nothing. That file and Weights::hand_authored are checked against each other
by a test, and sds-bot --print-weights writes it out again.
A weights file that cannot be read is fatal and says why — a missing file, a
name no feature answers to, a local feature that may not carry a weight. It
also says which features the file leaves at zero that the baseline weights.
Silently ignoring one would make a training run look like it worked and change
nothing.
The label is the match's final BV differential, given to every decision in that
match, so a good move in a lost match is labelled bad. It averages out over
matches and does not over decisions. Whether a fitted M is actually better is
a bench of the bot carrying it, --against the same baseline as the
hand-authored one — the fit itself proves nothing.
The commands, and why they exist #
controlplays Princess against Princess. It must come out near 50/50. If it does not, the harness is biased — seat order, deployment edge, the RNG, bridge latency — and every number it has ever printed is suspect. Run it after touching the harness, before believing anything else.benchplays a bot against Princess, alternating which faction the bot takes, and reports a win rate with a Wilson interval.oneplays a single match and prints everything, for looking at.--replayalso writes the match as an animated GIF next to the result, one frame per round, drawn by MegaMek's own minimap. It costs an X server inside the container and is off everywhere else: a benchmark does not want a GIF per match.playputs you in one of the seats, at MegaMek's own client, and leaves the same result, round report and decision log behind. For the failures a win rate cannot describe. docs/PLAYING.md.
baselines lists what has been recorded, corpus stores and lists training
corpora, and clean kills match containers left by a harness that was killed
rather than interrupted.
bench and one play /work/target/release/sds-bot unless --bot says
otherwise, and print which bot they are running in their first line. The floor
bot is still there and still worth running — --bot "python3 /work/bots/random_bot.py" — but you have to ask for it. It used to be the
default, and twice now a run of it has been read as a run of the change under
test; see plan/harness.md for both.
Reading one decision #
./sds.sh explain runs/<the run> --phase FIRING
Every decision line in a match's .decisions.jsonl already carries the whole
menu the bot ranked: each candidate's label, its measured features and the
value the argmax used. sds-bot writes the weight vector beside the log, and
explain prints the two together — one table per decision, a column per
candidate, a row per feature, showing the raw value, the weight and the
product. The default the bot could have taken instead (hold fire, stand still)
is a column like any other, so "it chose badly" and "everything on the menu was
worse than doing nothing" are different pictures.
sds view renders a whole match as HTML; explain is for the one decision
that made no sense.
Watching a match as it is played #
./sds.sh watch runs/<the run> # or one .decisions.jsonl
# then open http://127.0.0.1:8737/
watch tails the decision log by byte offset and serves it to a page meant to
sit beside the MegaMek window. The bot embeds no server and publishes no port:
SdsClient flushes every decision as it writes it, and the run directory is
inside the repository the match container mounts, so the file on disk is
already live. The same command replays a finished run — a log nobody is
appending to is a log whose tail is empty.
The page is a hex map of the board the scenario named, read from $MM_HOME
(nothing is vendored), with:
- every proposal's end hex coloured by rank, not by absolute value. What is ranked is a dropdown: the proposal's value, the damage it deals, the damage it takes, its risk, or any single feature it measured.
- terrain drawn over the colouring, one glyph per kind, and cliff edges drawn on the hex edge they belong to.
- each unit's footprint outlined along real hex edges, a chevron where it stands, a ring on the hex it took and a dashed line between the two.
- beside the map: which unit acted and which were eligible, why the force
replanned, the stance it was holding, and the chosen candidate against its
runner-up broken into terms — the same arithmetic
explainprints.
The two failures view exists to separate are counted in the panel: "the lance
chose badly" (a better proposal was offered and passed on) against "nobody
offered anything" (every proposal worth about nothing).
--port moves it — several agents share this machine — and --max-proposals
sets how many of a unit's proposals reach the page, best first, with the chosen
one always kept. The bot still reports every candidate it scored; the cut is
the reader's, and the page says how many it dropped.
How the numbers avoid lying #
- Mirrored forces.
scenarios/mirror-lance.mmsgives both sides the same four machines and the same pilots. A 200 BV edge swamps any tactical difference a new bot is likely to make. - Sides alternate. No two deployment edges are equally good, so the bot under test plays each of them half the time, and the tally follows the bot rather than the corner of the map.
- Seeded dice.
Compute.setRNGtakes a seeded generator, so two runs draw the same numbers. This does not make a match deterministic — seebridge/sds/SeededRandom.javafor what it does and does not buy. - Undecided games are not half a win. A match that hits the round limit has no winner and is excluded from the rate and reported separately. Counting it as a draw would make a bot that refuses to engage look average.
- The winner is MegaMek's, never inferred. Not "the side with more units left" — that turns a benchmark of who wins into a benchmark of who hides.
- Confidence intervals, always. 12-8 is a 60% win rate whose interval runs from 39% to 78%. Separating a five-point difference from noise takes about 385 decided games; the tool prints that reminder next to every result.
The bot #
crates/ holds the real bot: units propose, forces wait for all of their units
and then commit to a stance, and a node is an address so any level can be moved
to another machine. docs/HIERARCHY.md has the design and,
more usefully, the list of what it does not do yet.
How a bot plugs in #
A bot is a process. It reads newline-delimited JSON on stdin and writes it on stdout — see docs/PROTOCOL.md. It answers a decision, or it passes and takes the harness's default.
No SDS decision is ever played by Princess. SdsClient extends BotClient,
shares no code with Princess, and never falls back to it — not on a pass, a
timeout, or a crash. The reason is debugging: if a seat sometimes plays
Princess's move, no line in a match log tells you whose decision you are looking
at, and every investigation starts by working out whether the thing you are
staring at is even yours.
The defaults are ours and deliberately inert — stand still, hold fire, deploy in the first legal hex. A bot that answers nothing therefore stands where it landed and is shot to pieces, which is the correct and legible outcome rather than a respectable opponent wearing our name.
The line between what may be borrowed and what may not: MegaMek's rules yes,
MegaMek's bot no. WeaponAttackAction.toHit is rules — every human player has
that number on screen before choosing — so the observation carries it. The hex
ranking in BotClient.getStartingCoordsArray is tactics, and is not called even
though SdsClient inherits it.
The reference bots in bots/ are floors rather than opponents; none is meant
to be good:
bots/passthrough_bot.pypasses on everything, so its lance deploys, stands still and never fires. That is the floor beneath the floor, and the fastest check that the bridge is carrying decisions at all: if a real bot scores the same as this, its answers are not arriving.bots/random_bot.pywalks somewhere legal without looking at the enemy, shoots everything at whatever it is most likely to hit, and never thinks about heat. A bot that cannot beat it is not an improvement whatever it scores against Princess.
Requirements #
Docker. Nothing else — there is no host JDK and none is wanted, following the
pattern in helm/bridge and ~/.cache/mul-build.
MegaMek is mounted at run time from an extracted release, never vendored:
MM_HOME defaults to ~/.cache/mul-build/megamek. No MegaMek data lives in this
repository.
Hooks #
uv tool install prek
prek install
Licensing #
The bridge is compiled against a stock MegaMek release and links its GPL-3.0 code. MegaMek's data is CC BY-NC-SA 4.0 and is not redistributed here.