diff --git a/README.md b/README.md index 3508604..dda637c 100644 --- a/README.md +++ b/README.md @@ -70,6 +70,12 @@ candidate's feature vector, which one was taken — joins each log to its match result on the match tag, and fits one weight per feature by least squares on the difference between the chosen candidate and the ones it beat. +`bench --imitate` records the opposing seat's moves as well, reconstructed from +consecutive observations, and `train --imitation` fits them: a label per decision +instead of one per match. That clones the opposing bot's opinions along with its +moves — `plan/training.md` says what it inherits, and why it is a bootstrap +rather than the goal. + Two things it refuses to do. A `local` feature — one min-maxed across a single decision's own candidates — never gets a weight, because a weight fitted against one has learned a board rather than the game; the Rust side makes that a compile diff --git a/docs/PRINCESS.md b/docs/PRINCESS.md index 7e82e7d..db9c343 100644 --- a/docs/PRINCESS.md +++ b/docs/PRINCESS.md @@ -76,3 +76,14 @@ threat will be next turn. Princess cannot reason "I will be shot from that hill" | no enemy model | `plan/beliefs.md` | | every legal path | ~20 curated candidates; the cap is the performance design | | reaches into `Entity` | one documented observation — [PROTOCOL.md](PROTOCOL.md) | + +## Learning from it anyway + +`crates/sds-bot/src/imitate.rs` reconstructs Princess's movement from two +consecutive observations and records it as training rows, to bootstrap the +weight vector: a label per decision instead of one per match. + +It inherits the faults above. A weight fitted from Princess's choices has +learned Princess's opinion of a crossfire, which is the second fault on this +page. It is a starting position and has to be re-measured against Princess +before it means anything - see `plan/training.md`. diff --git a/docs/PROTOCOL.md b/docs/PROTOCOL.md index 7b12614..0e651fa 100644 --- a/docs/PROTOCOL.md +++ b/docs/PROTOCOL.md @@ -193,7 +193,50 @@ Written from `SdsClient` and flushed per line, so the file can be read while the match is still being played. `sds view ` renders one as HTML. The directory is the one named by `--log-dir`, defaulting to wherever `--out` -writes the result. It is also what the bot process gets as `SDS_LOG_DIR`. +writes the result. The bot process is told about it, and about which match it is +playing, through the environment: + +| variable | set by | meaning | +|---|---|---| +| `SDS_LOG_DIR` | `SdsClient` | where this run's logs go | +| `SDS_MATCH_TAG` | `SdsClient` | the match tag, so a file the bot writes itself can be joined to the result | +| `SDS_SEAT` | `SdsClient` | this seat's name | +| `SDS_SEED` | the harness | the seed a stochastic bot must be repeatable from | +| `SDS_IMITATE` | the harness, on `bench --imitate` | write the imitation corpus below | + +### The imitation corpus + +`/-.imitation.jsonl`, written by the **bot** rather +than by the host, and only when `SDS_IMITATE` is set. One line per opposing move +the bot could reconstruct from two consecutive observations of one movement +phase: + +```json +{"round": 3, "phase": "MOVEMENT", "owner": 1, "observer": "North", + "unit": 7, + "candidates": [ + {"label": "close on Griffin GRF-1N", "phi": { ... }, "value": 0.63}, + {"label": "observed move", "phi": { ... }, "value": 0.11}], + "chosen": 1, + "learnable": [ ... ], "local": [ ... ], + "policy": "observed"} +``` + +The training row is the same schema, so a fit reads it with no new parsing. Two +fields are not on a decision-log line: + +- `owner` — the id of the player whose move this was. The observation carries + player ids and no names; the match result carries both, and `sds train` + resolves the one to the other so the row is labelled with **that** player's + outcome. `observer` is the seat that watched, and is not the author. +- `policy` — `argmax` on the bot's own rows, `observed` here. On an `observed` + row `chosen` is somebody else's move and `value` is what *our* weights thought + of it, so `chosen` is deliberately not the argmax. Nothing downstream can tell + a corpus check from the point of the row without being told. + +Separate file and separate flag on both sides (`bench --imitate`, +`train --imitation`): what it teaches is another bot's opinion, faults included. +`docs/PRINCESS.md` and `plan/training.md` say which faults. ### The defaults diff --git a/plan/training.md b/plan/training.md index 868f600..2bcbf42 100644 --- a/plan/training.md +++ b/plan/training.md @@ -41,6 +41,13 @@ parameters. That is one overnight run. file a bench can be pointed at. If a fitted set only wins by a few points, ship this one - it is editable and explainable - [ ] Self-play only near parity. Two bad bots teach each other to beat a bad bot +- [x] **Imitation, to bootstrap**: reconstruct the opposing seat's movement from + two consecutive observations and record it as a training row whose + `chosen` is what they did. `sds bench --imitate` writes + `-.imitation.jsonl`; `sds train --imitation` fits it. See + `crates/sds-bot/src/imitate.rs`, and the section below for what it costs +- [ ] Re-measure any imitation-fitted `M` against the bot it was cloned from. + Until that number exists, an imitation fit has proved nothing Compute is not the constraint: ~0.1 core-hours a match, so 20,000 matches is ~$24 of Graviton spot. The constraints are a bot worth learning from and a @@ -61,3 +68,84 @@ it needs one thing nothing writes yet: BV per side per round in the match result or the decision log. When that exists, the change is `label_of` in `sds/train.py` and nothing else - the fit already takes one number per decision rather than one per match. + +## Imitation, and what it inherits + +The label above is one number per match. Imitation is one label per *decision*: +the enemy moved, and the hex it moved to is the answer. That is thousands of +rows a night instead of a few hundred, with no credit assignment problem at all, +which is why it is worth doing even though what it teaches is somebody else's +opinion. + +### How it works + +No Princess code runs, and nothing is asked of the opponent. The observation +already carries every visible unit's position, facing and `done` flag, so two +consecutive observations of one movement phase bracket a move: + +1. a unit that was `done: false` and is now `done: true` took its turn in + between, and where it is now is where it chose to be; +2. our own candidate menu is generated for it from its state in the **previous** + observation, with the observation mirrored so its side reads as friendly; +3. the observed hex goes into that menu as one more candidate, before anything + is measured - `damage_lead` is min-maxed across a decision's own candidates, + so a candidate added afterwards would not share the row's spread; +4. every candidate is measured with the same features the bot measures its own + with, and the observed one is recorded as `chosen`. + +Princess's own candidate set is not needed and could not be reproduced: it +enumerates every legal path and this bot caps at ~20 curated ones. What is +needed is only what it chose. + +The row is the same schema the bot's own decisions use, so `difference_rows` and +`fit` need no branch. `policy` is what tells the two apart, and the rows live in +a different file so a corpus is not silently part self-play and part clone. + +Moves that cannot be reconstructed are dropped and counted, never guessed at: a +unit not in the previous observation, a unit destroyed, a unit whose turn the +pair of observations does not bracket, and a displacement further than its own +movement points can explain. + +### The label + +The row is labelled by the outcome of the **player who made the move**, not the +bot that watched it - `label_of` is already per-seat, and the imitation row +carries the observed player's id for exactly this. So a match Princess won is a +positive label on its moves and a match it lost is a negative one, which is +"move the way Princess moves when Princess wins" and its anti-imitation half in +the same expression. It also degrades correctly as this bot improves: once +Princess starts losing, imitating it stops being rewarded. + +The alternative - a constant +1 on the winner's moves only - is a special case +of this with the magnitude thrown away, and there is no reason to throw it away. + +Worth being precise about what the fit then is. The design matrix is +`phi_chosen - phi_rejected` and there is no intercept, so with a label that is +constant within a match the objective is asking for a **fixed margin**: it wants +`w . (phi_chosen - phi_rejected)` to equal the label on every comparison. That is +a ranking objective, and it is the right shape. It is not a hinge, though - it +penalises a margin that is too *large* as well as one that is too small, so a +candidate our features already rank far above the rest still contributes error. +That is a defect worth knowing about before reading much into a coefficient. + +### The cost, which is the point of writing this down + +`docs/PRINCESS.md` is the argument for this repository existing. +`BasicPathRanker.rankPath` takes the **maximum** damage a hex can deal and the +**sum** of the damage it can take, so Princess cannot see the value of a +crossfire and is systematically pessimistic about advancing. **Cloning its moves +clones that blindness.** The same goes for `herdingMod`, which is the conga line. + +This does not break "MegaMek's rules yes, MegaMek's bot no": no Princess code is +linked, imported or invoked, and every order the bot gives still has exactly one +possible author. But that invariant's own wording is that inheriting a tactical +opinion **silently** is the thing the design exists to avoid - and this +deliberately inherits one. So it is recorded here, in `imitate.rs`, and in the +fit's own report, which names how many of its rows are imitation and says what +they were cloned from. + +Imitation is a **bootstrap to competence, not the goal.** Weights fitted this +way must be re-measured against the bot they were cloned from before they are +believed, and the long-term path is outcome-based training that exceeds it. A +future reader must not be able to conclude this bot is independent of Princess +when part of it was copied from Princess.