diff --git a/plan/README.md b/plan/README.md index 46bbc1a..6dc1f6c 100644 --- a/plan/README.md +++ b/plan/README.md @@ -70,6 +70,7 @@ frontmatter keys every other epic carries, run the script, stage both. | [fire-modes](fire-modes.md) | Ultra and Rotary autocannon rates are a choice | open | ev | | [heat-damage](heat-damage.md) | Making the enemy hot is a way to win | open | ev, features | | [infantry](infantry.md) | A platoon is a headcount, not a body plan | open | unit-types, features | +| [latency](latency.md) | A unit-decision has a budget, and the worst case honours it | open | candidates | | [matchup](matchup.md) | The bot can see what it is up against | open | features, scenarios | | [melee](melee.md) | The bot throws a punch | open | features, ev | | [mines](mines.md) | Ground that hurts to walk on | open | protocol, beliefs | diff --git a/plan/latency.md b/plan/latency.md new file mode 100644 index 0000000..a6abc43 --- /dev/null +++ b/plan/latency.md @@ -0,0 +1,119 @@ +--- +id: latency +title: A unit-decision has a budget, and the worst case honours it +status: open +dependsOn: [candidates] +exitCriterion: > + On the 8v8 perf fixture a unit-decision never exceeds a stated work budget, + the decision is a function of the inputs and the budget alone, and the + median is under a fifth of a second. +--- + +# latency + +README promises a bot that calculates turns faster than Princess on average +and more reliably in the worst case. `docs/PERFORMANCE.md` measures a +unit-decision at **2 to 4 seconds** on the 8v8 fixture, against a design +budget of 100 ms and Princess's tens of milliseconds, and nothing bounds the +worst case at all. Nothing in the sweep is slow: one exchange is well under a +microsecond and one line-of-sight ask is tens of nanoseconds. There are 1.4 +million of them, and there is a multiplier the document does not count: +`Bot::plan` rebuilds whenever an enemy settles or one of ours is displaced, +and every rebuild sweeps every unit (`main.rs:955-1030`, `1105`), so a side +can pay N rebuilds of N sweeps in one round. `carry.rs` recovers the pairs +that did not change and nothing else. + +The exit criterion's numbers are a proposal. A fifth of a second is where a +player stops noticing; the budget constant is the part that matters, because +it is what makes the worst case a number. + +## Where the time is + +The loop is stand group by (hex, TMM) → stand → foe → position in the enemy +envelope → two halves (`stands.rs:1731-2055`). Per half: a `SideMemo` probe; +on a miss a `LineTable` bearing and line; then `best_volley_with_line` +(`volley.rs:1706`), which recomputes the side table with an `atan2` +(`1736`), packs a `FiresKey`, probes the level-1 memo, packs a `LandKey`, +probes level-2. A level-2 miss builds a `LocationProfile` (`hitloc.rs:1206-1224`: +one allocation and about 130 `String` compares, on the order of 120k times in +a cold sweep) and solves. The result is a 68-byte `Exchange` (`stands.rs:509`) +reserved at the envelope's full size per (stand, foe): 47 MB reserved, 28 MB +written on the fixture. `Carried::of` then clones the whole sweep +(`carry.rs:154`) and `Memos` keeps one per unit for the round, so a side at +8v8 holds a few hundred megabytes of exchanges beside a JVM capped at 2 GB. +Memory, not CPU, is what bounds `--jobs` today. + +Two rows in `PERFORMANCE.md` are stale and would misdirect work: line of sight +is no longer one ask per exchange since `LineTable` landed (about 38k asks a +sweep, milliseconds), and the host's `toHit` count is not 384 a phase. +`Observation.firing()` prices every eligible shooter at every twist on every +decision (`Observation.java:461-477`, `540-574`), which is nearer 8,600 calls +a phase, close to a second of JVM time at the measured 100 µs each. + +## The order to do it in + +Every item below is exact - the same decision, cheaper - except the two that +say otherwise. + +1. **A work-unit budget over a bound-first order.** Pass one prices every + (stand hex, foe hex) cell from the `LineTable` and `ReachTable::best_at` - + an upper bound on outgoing, a lower bound on incoming, side- and + gait-agnostic. Stands are swept in order of that bound until + `stats.exchanges` reaches the budget; the rest keep their bounds and are + reported with their precision. Exact for a budget the sweep finishes + inside, approximate past it, and a function of the inputs and the budget + alone, so it keeps the completion-order invariant. This is the only item + that meets the worst-case promise, and "units report every scored + candidate" becomes "every candidate is reported, with how precisely". +2. **Compact the exchange record and stop cloning the sweep.** Outcomes are + already interned in the level-2 memo; an `Exchange` can be an index and + two `u32`s, about 12 bytes, five times smaller, and `present` can be + reserved at the measured density. `Carried` can own the sweep, or an + `Arc` of it, instead of a copy. +3. **Hoist the profile, flatten the memo, reuse the bearing.** One + `LocationProfile` per (target, side) per turn - 32 a sweep - instead of + one per miss; a dense `Vec` indexed by (fires id, foe, + side) in place of the `LandKey` hash and the duplicate `outcomes` + `BTreeMap` with its cloned keys; the bearing the sweep already holds + passed in rather than recomputed. Together roughly a halving of the hit + path. +4. **Store the enemy envelope per position group with multiplicity.** + `PositionGroups` already shares the compute (1.39x measured); the store + and the collapse still run per position. Exact. +5. **Two-pass refinement.** Sweep exactly only the stands whose pass-one + bounds can reach the frontier or the weights' top K. Approximate for the + unrefined stands, deterministic for a fixed K and a total ordering key, + and the second thing that bounds work. `plan/hierarchy.md` has asked for + this since the start. +6. **The host's share.** A per-phase cache of `toHit` keyed on (shooter, + twist, weapon, mode, target), invalidated for a shooter when it declares; + `terrain{mp, barred, stands}` computed once per (entity, board) rather + than per decision; the observation sent once per phase with per-turn + deltas, which `plan/protocol.md` already lists. + +What not to spend on: pruning the envelope by dominance (the collapse is a +mean over all of it, so dropping members changes every score - a real cut +needs a belief over enemy intent, which is [beliefs](beliefs.md)); SIMD or +bitset line of sight; the observation parse, which is two milliseconds a +round. + +## Determinism + +One real hazard was found and is fixed: `by_lance` in `main.rs:1113` was a +`HashMap` iterated at `1146` and `1172`, so a company's `f32` appraisal sum +and the order decisions were recorded followed the process's hash seed. Every +other `HashMap` in the two crates is looked up, never iterated to a result, +and every `Instant` read lands in a counter rather than a decision. The +`explore::Stream` key omits the phase, so both phases of a round share one +"explore?" draw - by design, and worth knowing when reading a corpus. + +- [ ] Correct the two stale rows in `PERFORMANCE.md` and count the rebuild multiplier +- [ ] `Exchange` at ~12 bytes; `present` reserved at density; the sweep owned, not cloned +- [ ] `LocationProfile` once per (target, side) per turn; `Location.name` interned at parse +- [ ] Dense level-2 memo; drop the `outcomes` map on the fast path; pass the bearing in +- [ ] Position groups stored and collapsed with multiplicity +- [ ] Pass-one bounds from the `LineTable` and `ReachTable` +- [ ] The work-unit budget, its constant, and the precision field on a reported candidate +- [ ] Two-pass refinement to the frontier and the top K +- [ ] A per-phase `toHit` cache and once-per-board terrain in `Observation.java` +- [ ] The p99 of a unit-decision printed by `sds stats`, beside the mean