Performance #
The budget the design is built against, and what actually costs anything.
Measured with cargo run --release -p sds-core --example perf, single-threaded,
on an i7-1270P under load — deployed runs get one to two CPUs, so a number that
needs sixteen cores is not a number.
Measured #
Four runs, uptime load average 8-18: several agents building and a 40-game
benchmark, which is this machine's normal state. Each figure is the example's
own answer, which is the fastest of five batches of CPU time — see How these
are measured for why it is neither a mean nor a wall
clock. The range across the three runs is what is left after both.
| operation | cost |
|---|---|
| threat map, 8 enemies x 544 hexes | 5.2-11.9 us |
| 160 LOS queries, no cache | 175-458 us (1.1-2.9 us each) |
| 1280 LOS queries, cold cache (160 traces + 1120 hits) | 364-675 us |
| 1280 LOS queries, warm cache | 66-69 us (52-54 ns each) |
| one warm LOS ask over the sweep's own key set | 72-78 ns |
stance blend w = S^T M |
0.1 us |
| score 160 candidates x 30 features | 2.0-2.4 us |
| fire allocation, 48 shots x 8 targets, greedy PMF convolution | 159-403 us |
| one shot's PMF, built cold | 0.2 us |
| parse one 28KB observation | 96-123 us |
| per-location damage, one volley, cold cache | 44-101 us |
| per-location damage, one volley, warm cache | 1.4-2.7 us |
| 20 candidates, per-location | 66-178 us |
volley::gather, one (stand, position) pair |
615-1083 ns |
volley::best_volley, warm memo |
768-1736 ns |
stands::score_stands, one 8v8 unit-decision |
2.0-4.0 s |
| — per exchange | 1451-2870 ns |
stands::rank over that sweep |
84-215 ms (197-503 ns/entry) |
Everything above the bold rows is a hot step. The bold rows are the hot loop, and they are where the runtime is: see The sweep.
The 2-3x spread between runs is contention, not variance in the work. Take the high numbers.
Per-location damage costs 33 us for a whole unit's 20 candidates, against 66 us for the fire allocation beside it. That is affordable, and the reason is the share. Damage arriving at a location is memoised on the shots and on that location's chances in 36, and a pristine front table has only four distinct shares - 7, 5, 4 and 1 - so eight locations cost four convolutions rather than eight, and a second candidate with the same volley costs 0.6 us instead of 19. A whittled-down target has fewer distinct shares still, because the destroyed locations' rolls fold into the ones behind them.
Nothing here enumerates critical slots or recomputes battle value per location. That would be roughly 20 candidates x 8 locations x a battle value pass, and it is what the capped-candidate invariant exists to prevent; what each location holds is one scalar off the observation instead.
Fire allocation measures sds_core::ev, not a stand-in. It is 48 shots against
8 targets but only ~10 distinct (rack, packet damage, to-hit) triples, so the
memo answers 4838 of 4848 asks and the convolution is sparse-into-dense: 12
outcomes into 128 buckets, not 128 x 128. Without the memo it is the same work
485 times over.
The sweep #
stands::score_stands is 93% of the runtime, and nothing measured it until
now. Counters over 229 unit-decisions and 1805 s of thinking put 96.9% of it
inside the sweep and 93.2% inside the volley estimate the sweep calls twice a
pair — 76% of that in volley::gather and 24% in the memo lookup it feeds.
Paths were 0.0%, envelopes 0.0%, rank 2.8%, surface 0.3%.
The benchmark's fixture is two real MegaMek sheets side by side — Map Set 5/16x17 Open Terrain 1 and Map Set 4/16x17 Heavy Forest 1, out of
tests/corpus/pathfind.jsonl — with eight 5/8 Meks a side carrying six guns
across five mounts, both forces at their run allowance because nobody has moved.
That is 327 stands against 8 enemies over 2133 positions: 697,491 pairs and
1,394,982 exchanges in one unit-decision, of which 38.8% store nothing and the
rest occupy 25 MiB. The volley memo answers 99.5% of its asks.
At 2.0-4.0 s a sweep it lands close to the 7.6 s a real unit-decision took in the match logs, which is the check that says the fixture is the right shape and not a toy.
The per-item figures are the ones a change has to move. The sweep is not
expensive because any step in it is slow — one exchange is well under a
microsecond and one line of sight ask is ~72 ns. It is expensive because there
are millions of them. A fixture half this size should report the same
ns/exchange, and a change that improves the model should be read there rather
than in the total, which moves whenever the fixture does.
Three things inside it are worth naming, because they are where the
ns/exchange goes:
- Line of sight is ~10-13% of the sweep, at ~72 ns an ask and one ask per exchange. The share has moved twice while the ask itself barely has: it was ~18% at 265 ns before the memo was made cheap, ~5% at 67 ns straight after, and back to ~10-13% once the volley key stopped being built — the numerator fell by a quarter and then the denominator fell further. Read the share against the build that produced it, and see what a warm ask costs before sizing work against any of them.
gatherruns before the memo, because it is what builds the key. A 99% hit rate does not make key construction free — a hit pays for it in full, which is why thegatherrow and thebest_volleyrow are close together.LocationProfile::ofis ~7.5%, built once an exchange, and its cost isUnit::locationdoing a linear scan ofStringcomparisons eight times a profile.
What a warm ask costs #
98-99% of line of sight asks are hits, so "what does an ask cost" means "what does a hit cost", and for a long time this file answered with a row that was not measuring one.
1280 LOS queries, cached built a fresh LosCache inside each timed batch.
160 of its 1280 asks were therefore real traces at ~1.3 us apiece, and they were
most of the batch: the row reported ~300 ns an ask and was read as the price of
a hit. It is the price of a round's fill. The row is now
1280 LOS queries, cold cache, and a warm cache row beside it fills the memo
before the clock starts.
That row has a fault of its own in the other direction: 160 distinct keys is a
working set that fits anywhere, and it reports 52-54 ns. The figure to quote is
sweep-wide LOS ask (warm), which builds the AttackInfo and asks the memo
over the sweep's own pairs — 34,180 distinct keys in one unit-decision, which is
the working set the memo really faces. 72-78 ns, minimum of three batches
over five runs, against a hex::distance control in the same loop shape at
3-9 ns.
The older figures in this file are not wrong, they are stale, and the order they were taken in matters:
| build | warm ask | share of the sweep |
|---|---|---|
| before the memo was made cheap | 265 ns | ~18% |
| after (by-value entries, multiply-rotate hash) | 67 ns | ~5% |
| after the volley key stopped being built | 72 ns | ~10-13% |
265 ns is what justified making a hit cheap, and 213 -> 76 ns is what that
change moved a get by in isolation. The share then went up again without the
ask changing, because the rest of the exchange got cheaper around it. A share is
a ratio and both halves move.
How these are measured #
examples/perf.rs reports the minimum of five batches of CPU time, and both
halves of that are load-bearing on a shared machine.
Not a mean. Several agents build here and a benchmark of 40 games runs beside them. The same binary has been timed at 59.5 ms and 129.7 ms an hour apart, and an experiment that deliberately doubled the work once came back faster. A mean over that spread is a measurement of the other agents' work and it moves when they finish; the minimum is the closest available reading of what the code costs with nothing in the way. It is a lower bound and it is meant to be — nothing here is a service-level figure, and every number exists to compare one build against another.
Not wall clock. A run descheduled for 40 ms because somebody else's build
got the core has a wall time 40 ms longer and has done identical work. The
kernel already counts what we were given: /proc/self/schedstat's first field
is nanoseconds on CPU. That leaves the contention that is real — cache and
memory bandwidth are shared whether or not we are running — and removes the
waiting, which is the part that moves by a factor of two.
It ticks once a millisecond, so the harness grows each batch until it has run for at least 100 ms and divides by the count that actually ran. Without that the cheap rows quantise: they came out as 5.0, 15.0 and 130.0 us, which are the clock's numbers and not the code's.
Where /proc/self/schedstat is not readable everything falls back to wall
clock, and the header line the example prints says which of the two produced the
table.
A/B-ing a change #
The minimum-of-five CPU figure above makes one run readable. It does not make two runs comparable, and every performance number in this repository used to come from a shell loop that assumed it did. In one afternoon on this machine those loops produced:
- a first reading of 45% for a change that eight properly controlled pairs put at 32-39%;
- a 14% for a change that twenty-five controlled runs put at 5-7%, which is what its mechanism predicted;
- an A/B in which the control - literally unchanged code compiled into both binaries - came out at 1.43 and then 1.99;
- a doubling experiment that came back negative, which is not a thing that can happen;
- and two statistics disagreeing on the same seven samples - raw medians said 7%, within-run normalisation said 24%. Neither was bad arithmetic. The gap between them was the only thing in the room saying the sample was too small to carry either.
None of those are variance in the code. They are the machine: several agents build here, the load average has ranged from 4 to 37 in a day, and the same unchanged binary has been timed at 59.5 ms and 129.7 ms an hour apart. A figure read off build A at noon and build B at half past is a measurement of what happened in between.
scripts/ab.py is the method that has actually worked, made repeatable so that
nobody has to remember it and so that a number in a PR body can be reproduced.
scripts/ab.py --before /tmp/perf-main --after /tmp/perf-branch --pairs 15 \
--metric 'stands::score_stands' \
--control 'parse one observation'
It does four things:
- Runs the two binaries alternately, swapping which goes first each pair, so neither build owns the quiet half of the run.
- Reads two figures out of each run: the
--metricunder test, and a--controlthe change does not touch. - Normalises within the run - metric divided by control - before anything crosses between builds. Whatever the machine was doing during that run was done to both figures, so it cancels.
- Compares the medians of those per-run ratios, with a bootstrap interval, and prints the load average at the start and the end.
It reports both statistics - the raw before/after and the normalised one - side by side, always, and says so when they disagree by more than five points. That disagreement is a reading in its own right: the 7%-and-24% pair above was one sample looked at two ways, and the only thing that made it obviously unquotable was seeing the two numbers next to each other.
A selector is a regular expression. Capture a group and that group is the
number; capture nothing and the last number on a matching line is taken, which
is what examples/perf.rs puts there - so volley::gather reads its
ns/gather and score 160 candidates reads its us rather than the 160 in its
name. A selector that matches two rows is an error, not a coin toss.
The refusal #
When the control's own before/after ratio is more than 10% from 1.0, the tool prints no result at all. The control is code the change does not touch, so its ratio has one honest value and that value is 1.0. Anything else means the machine moved under the measurement in a way the within-run normalisation did not absorb, and no figure from that run is readable — including the one you wanted. It exits 2 and says so.
That refusal is the most useful part of the tool. Two of the four failures above would have been caught by it and nothing else; the 1.43 control was sitting in the output the whole time and was read past.
Below the refusal there are three warnings, which do not gate: fewer than
--min-pairs pairs (8); a control whose own spread within a build is over a
quarter of its median; and the raw and normalised readings disagreeing by more
than --max-divergence (0.05). All three mean the number is worth less than it
looks.
What it looks like #
Driven by a stand-in binary that prints canned perf output under a machine
factor swinging 1.0 to 3.0 - which is this machine - with a true improvement of
20% in the first and third columns of the fixture:
=== a real 20% improvement, control steady
medians before after change
metric 3344 2785 -16.7% (uncontrolled - do not quote this)
control 212.2 220.8 +4.1% (should be ~0%)
normalised 15.7617 12.6096 -20.0%
RESULT -20.0% normalised, -16.7% raw (95% interval -20.0% to -20.0%, 8 pairs)
=== the control itself moved 1.5x between the builds
metric 3228 2486 -23.0%
control 204.8 295.6 +44.4% (should be ~0%)
normalised 15.7623 8.4065 -46.7%
WARNING the two statistics disagree by 23.7 points (raw -23.0%, normalised -46.7%)
REFUSED TO REPORT
the control moved 1.444x between builds ... No figure from this run is readable.
=== identical binaries on both sides - a null A/B
normalised 15.7629 15.7611 -0.0%
RESULT -0.0% normalised, +0.7% raw (95% interval -0.0% to +0.0%, 8 pairs)
The first column is the whole argument: the naive reading is -16.7% and the true answer is -20.0%, and the difference is entirely the 4.1% the machine happened to drift. The second exits 2 and prints no result. The third is a null A/B - the same binary on both sides - and comes out at -0.0% normalised while the uncontrolled reading still wanders 0.7%.
tests/test_ab.py drives those three cases from a fake binary rather than a
real build, which is the only way to check that the refusal fires when it should:
a live run cannot be made to move its control on demand.
Running one #
Build a release binary per branch and keep them side by side. A worktree per branch is already the workflow, so this is two builds and two copies:
cargo build -j 2 --release -p sds-core --example perf
cp target/release/examples/perf /tmp/perf-<branch>
Then pick a control the change genuinely cannot reach. parse one observation
is serde and nothing else; threat map (8 enemies) is arithmetic over a slice.
A control inside the same module as the change is not a control.
examples/perf takes about 40 seconds a run under load, so 15 pairs is roughly
20 minutes of one core. --pairs 8 is the floor for a figure worth quoting.
--json writes every run, its whole output and the summary. That is what to
attach when a PR's number is questioned, and it is also re-readable:
scripts/ab.py --from-json run.json \
--metric 'volley::gather' --control 'threat map'
runs nothing and re-scrapes the measurement already taken. The runs are the expensive part and the machine they happened on is gone; asking a different row of them is free, and it is the only honest way to answer "what did that change do to the row I did not look at".
Per round, 8v8, 16 decisions #
Bot: ~6-10 ms typical, ~30 ms pessimistic. Per unit decision, 1-3 ms.
That was before the sweep existed and it is now wrong by three orders of magnitude — a unit-decision is 2-4 seconds, and The sweep is why. The rest of this section is still true of everything outside it.
Dominated today by 16 observation parses (~4.7 ms). That is a statement about how little thinking exists yet, not about JSON being slow: scoring 160 candidates takes 3 us because the scoring is one dot product.
The JVM side is now measured rather than estimated. sds los-dump times
MegaMek's own LosEffects.calculateLOS over 2383 calls on four boards and gets
95-107 us mean across runs. That is a cold JVM with no warm-up pass, so
treat it as an upper bound - but it is the right order, and it lands at the top of the 20-200
us range this file used to guess. WeaponAttackAction.toHit runs 8 shooters x 6
weapons x 8 targets = 384 of those per firing phase, so ~40 ms/round on the
host: still several times the bot's whole budget.
What will dominate once the thinking lands #
Line of sight, and it turned out cheaper than feared. exposure,
cover_quality, los_in, los_out, rear_arc_gain and a terrain-aware threat
map all need it, and it is O(hexes along the line) per pair: ~1280 queries a
round at 8v8. The nine positional features that shipped ask through one
Posture::survey per candidate, so the count is two traces per enemy per
candidate however many features read the result.
This file predicted 5-20 us each and 6-25 ms/round. Measured, one query is
0.7-1.5 us and a whole round is 0.3-0.7 ms — better than a fortieth of
the guess even taking the high numbers, and about a tenth of what the host
spends answering the same question.
The per-turn LOS cache shared across the side still earns its place: at 8v8
it answers seven queries in eight, which is where 1.5 us falls to 0.6 us.
MegaMek does the same internally — the server passes a losCache into
filterEntities.
What blows up #
- Candidate explosion. Everything here assumes ~20 curated candidates per unit. Enumerating paths the way Princess does is 100-1000x, and puts a "thinking..." message back in front of the player. The cap is the design.
- EV over subsets. Greedy with an incremental PMF is O(shots x targets). Any search over allocation subsets is exponential.
toHitscaling as shooters x weapons x targets. A 12v12 of missile boats is ~4x worse — still only ~300 ms/round.- Observation per decision — 28KB x 16 = ~450KB a round encoded and decoded. Sending it once per phase would remove most of that. Tidiness, not performance, until the thinking lands.
Board size is linear and cheap: four sheets is ~34 us for the threat map. Memory
is a non-issue — a PMF is 128 floats, the threat map 544, M is 8x30, all
L2-resident.
Headroom #
On one CPU with a 100 ms/unit budget we use ~3 ms. Roughly 30x headroom. That is the answer to whether the EV calculator, per-location armour, opponent modelling and multi-turn projection are affordable: yes, provided candidates stay capped.
Concurrency #
The tokio fan-out buys nothing on one core and is not there for speed — a force must have heard from every unit before it decides. It degrades to sequential cleanly and must not be removed as a pointless optimisation.
Where concurrency does pay, even on one CPU, is a compute-once many-waiters LOS
cache: eight units will ask overlapping questions. sds_core::los::LosCache is
that cache - a Mutex around a memo, with a Condvar per entry so the second
asker waits rather than computing the same line again. Nothing iterates the map,
so no result depends on its order.
What a hit costs is its own question. At 98-99% hits the memo is working and
the bill is the asking: a warm LosCache::get measured at 213ns minimum over
a 2,176-key working set, of which SipHash over the 48-byte AttackInfo was
~110ns and the rest was the Arc<Slot> clone and the second lock a hit took to
read the answer out of the slot. Two changes, neither of which touches what is
computed: a resolved line lives in the map by value, so a hit copies it out
under the lock it already holds, and the map hashes with a seedless multiply-
rotate rather than SipHash. 76ns minimum after, over eight alternating runs
of each build.
End to end, examples/stands spends 5-7% less user CPU - min ratio 0.960,
median 0.927, mean 0.932 over twenty-five alternating runs of each build - and
prints a byte-identical board. That figure is worth stating carefully, because a
first measurement of it came out at 14% and did not survive being taken again:
the arithmetic says 7%. Line of sight is ~18% of the sweep, the sweep is ~60% of
this binary and a get lost 64% of its cost, so 0.18 x 0.60 x 0.64 is where the
win has to land. A whole-binary number taken in a lucky window on a shared
machine will beat its own mechanism, and when it does, the mechanism is right.
It lives for the match, and deployment shares it. Measured over eight rounds
of scenarios/mirror-lance.mms, one seat, with scripts/los-growth.sh:
| round | phase | lines held | new | MB | asked | hit% |
|---|---|---|---|---|---|---|
| 0 | deployment | 37,656 | 37,656 | 4.74 | 10,054,462 | 99.63 |
| 1 | movement | 38,674 | 1,018 | 4.87 | 20,563,140 | 99.995 |
| 2 | movement | 58,278 | 19,604 | 7.34 | 21,108,220 | 99.91 |
| 4 | movement | 61,486 | 710 | 7.74 | 15,626,016 | 99.995 |
| 6 | movement | 61,900 | 0 | 7.79 | 3,213,544 | 100.00 |
| 8 | movement | 64,020 | 1,600 | 8.06 | 3,416,936 | 99.95 |
It flattens: the first two rounds put down 94% of what eight rounds hold, and rounds five to eight add 0.05 MB each. 90.4M lines were asked for and 64,020 computed. There is deliberately no bound on it - a cap chosen before the growth was measured would have been a guess, and 8 MB does not need one.
The same eight rounds with the old per-round clear recompute 20-45k lines every
round instead. At the 0.7-1.5 us a line costs that is tens of milliseconds
against a round that thinks for ten seconds, so the clear was costing almost
nothing: the reason to drop it is that its stated reason - "units move" - was
never true of a cache keyed on AttackInfo.
That argument is silent about the thing that does invalidate a line. Units
moving does not; terrain changing does, and terrain changes in a real fight -
buildings collapse, fires start, smoke drifts and thins. The lifetime is safe
today only because the board document is sent once per match and
docs/PROTOCOL.md says outright that mid-fight terrain is not tracked, so the
bot cannot learn that a hex changed. Whichever change first puts smoke or
buildings into the observation has to decide this cache's lifetime in the same
breath, or a line traced in round 1 will answer a round 6 question through smoke
that was not there - silently, with nothing failing. plan/los.md has the note;
CLAUDE.md's invariant already says "per turn", which is the form that survives.
Results must not depend on completion order. Reduce in fixed order, budget in work units rather than wall-clock, give each unit its own seeded RNG stream, and never read a clock inside decision logic.
Comparing against Princess #
The turn clock measures wall time between turn changes, which is the same quantity for both sides and says nothing about how much of the machine each side used to fill it. That gap is not small here, and it runs one way:
| threads doing the work | |
|---|---|
| Princess | one decision thread, plus one Precognition thread |
| sds | one tokio runtime, units fanned out across its workers |
MegaMek contains no parallelStream, no ForkJoinPool and no ExecutorService
outside two non-bot classes; megamek/client/bot has exactly one thread of its
own, Princess.precognitionThread. BasicPathRanker.rankPath over every legal
path is strictly serial. So an unpinned wall-clock comparison between these two
is a comparison of two different amounts of hardware, and no sample size fixes
it.
So a run whose timings are going to be quoted uses --pin. Each seat gets
one physical core and nothing else:
- Seats are separate processes - see
bridge/sds/SdsSeat.java- so onetasksetper seat covers its JVM, the per-turn threads MegaMek creates inside it,Precognition, JIT, GC, and the bot subprocess with every tokio worker in it. A mask is inherited acrosscloneandexec, which is what makes one call enough. - The host JVM, and so the server and the clock, is started with a mask that excludes the seats' cores, so nothing of MegaMek's own drifts onto one.
SDS_WORKER_THREADS=1, because four workers time-slicing one core is context switching and nothing else.- Cores are chosen by
sds/cores.py, never by CPU number. Two traps it exists for: on an SMT part-c 0and-c 1are usually the same physical core, and this is a hybrid part where-c 0is a 4.8GHz P core and-c 8a 3.5GHz E one. Seats only ever get distinct physical cores of one speed class, and the SMT siblings of those cores are kept out of the container's cpuset so a seat's core is a whole core.
A pinned figure and an unpinned one are two different measurements. The result
document records which it was in pinning, and sds stats prints it beside the
thinking-time table.
What the split costs. A match is three JVMs rather than one. Measured over a
mirror-lance match: a seat peaks at 668MB resident, the host at 865MB, so
~2.2GB a match against roughly a third of that before. Seats are capped at
-Xmx2g because the default maximum heap is a quarter of the machine's RAM and
three JVMs each claiming 7.5GB is how --jobs 2 ran a 30GB machine down to 3GB
free. Memory is now what bounds --jobs, not CPU.
Match throughput #
~200 s/match at --jobs 3, so ~54 matches/hour; solo is ~88 s, so three at
once buys only ~20% — one match already saturates ~2 cores. At ~0.1 core-hours a
match, 20,000 matches is ~$24 of Graviton spot. Compute is not the constraint;
see plan/training.md.