diff --git a/plans/0000-roadmap.md b/plans/0000-roadmap.md index 32c0c56..42f7110 100644 --- a/plans/0000-roadmap.md +++ b/plans/0000-roadmap.md @@ -88,7 +88,7 @@ the real design happens. The blurbs here stay deliberately loose. markdown into styled rows as it streams, committing each block as soon as it settles and holding back only the one still forming. *You can now: read the DM's words styled, not asterisked.* - ([0008](0008-streaming-markdown.md)) + ([0008](completed/0008-streaming-markdown.md)) - [ ] **Phase 7: DM prompting.** The first real tuning pass against live sessions, and the DM-versus-player knowledge line: secrets from lore stay behind the screen until the player earns them, and the DM diff --git a/plans/0009-scenario-replay.md b/plans/0009-scenario-replay.md new file mode 100644 index 0000000..19664c5 --- /dev/null +++ b/plans/0009-scenario-replay.md @@ -0,0 +1,275 @@ +# 0009: Scenario replay + +## What this is and why + +A scenario is a fixed, repeatable Storied session: a list of keys the +player presses and a list of events the DM reports. Scenario replay runs +that list through the real play loop and shows what came out, as text and +as a picture. + +The goals are speed and cost. Today, checking how the terminal renders +means one of two things. Either you read the live game, which needs an API +key, a model, and a human typing, all of it slow and non-deterministic. Or +you write a Rust test with a hand-built step list, which is fast and +deterministic but ties every scenario to code. A scenario file lowers both +costs: you write a scenario as data, run it with no network and no model, +and reuse the same file to watch the visuals and to lock a behavior into a +test. + +It also lets an agent review the output. A moving terminal is only good +for the person staring at it. The artifacts replay produces are meant to +be read after the run, so another reviewer, or tooling, can judge what +happened. + +## The scenario file + +A scenario is a list of steps in order. Each step is one pass of the play +loop. The language has a small set of steps, with two conveniences so most +scenarios read like a conversation instead of a timeline. + +```json +{ + "name": "statblock", + "terminal": { "width": 60, "height": 26 }, + "steps": [ + { "input": "I open the iron door." }, + { "reply": { + "text": [ + "You push open the iron door and meet **Mother Uldra**.", + "", + "# Mother Uldra", + "", + "| Str | Dex | Con | Int | Wis | Cha |", + "|:---:|:---:|:---:|:---:|:---:|:---:|", + "| 12 | 14 | 13 | 10 | 16 | 18 |", + "", + "She croaks: *\u201cyou are late, little moth.\u201d*" + ], + "chunk": "word" + }}, + { "checkpoint": "landed" }, + { "key": "eof" } + ] +} +``` + +The steps: + +- `input`: type a string and submit. This is the common case. +- `type` and `key`: type without submitting, or press one named key, for + editing tests. +- `reply`: the DM speaks. `text` is the reply's markdown, either one + string or a list of lines the parser joins with newlines. `chunk` picks + the streaming granularity that splits the text into events: `word`, + `char`, `line`, `paragraph`, or `whole`. The granularity you choose is + the behavior you test. A table that arrives row by row is exercised by + `line`. A paragraph that re-wraps as it grows is exercised by `word` or + `char`. +- `events`: an explicit list of DM events for full control. This is where + tool calls, cancellations, and failures live, and where a `key` step can + be interleaved between deltas so a cancel lands mid-stream. +- `resize`: change the terminal size. This exercises width pinning and + re-wrapping. +- `checkpoint`: name this point in the run, so a screenshot is rendered + here. The final point is always captured; a checkpoint adds another. +- `idle`: advance the render clock by several passes, so the thinking + indicator can be observed. + +The file validates strictly. An unknown step, key, chunk policy, or field +fails the run loudly rather than being ignored. + +The steps are flat and pass-ordered, not grouped by turn. The play loop +reads keys and drains DM events on separate passes, and the only way to +test a cancel mid-block is to press Escape between streamed deltas. The +conveniences hide that ordering in the common case; `events` exposes it +when the scenario needs it. + +## How replay runs it + +Replay feeds the scenario into the same `play` loop the real game uses. +Nothing is reimplemented. Two inputs get replaced: the human keyboard +becomes a `Keys` implementation that yields the scenario's keys, and the +live model becomes a worker stub that emits the scenario's events on the +worker channel. The `play` loop, the transcript, the viewport, and the +renderer all run as they do in production. + +The one real terminal path lives in `play::terminal::run`, which builds +crossterm, the `CrosstermViewport`, the `CrosstermSyncGuard`, and the real +`Worker`. The `replay` subcommand routes in `cli.rs` but builds its own +doubles, so it never calls `terminal::run`. + +## What you get back + +Replay is for observation, not assertion. It produces artifacts to read, +and it asserts nothing. The artifacts: + +- **The behavior narration.** A line-oriented record of what the loop did + each pass: which rows committed to the transcript, which rows stayed + forming in the tail, how many rows spilled, the viewport height, which + key was pressed. This is the TUI-behavior story. +- **The rendered screen as a character grid.** Every cell of the buffer, + one character each, plus a per-cell attribute layer so styling (bold, + dim, italic, colors) can be checked in place. Alignment, borders, and + wrapping all live in this grid. +- **A rendered screenshot as an image.** The same screen drawn to a PNG, so + the visuals can be looked at the way a human reads a terminal capture. + A screenshot comes at the end by default, and at any checkpoint the + scenario places, so a streaming bug is visible in the frame it happened. +- **A behavior summary.** Counts and invariants drawn from the run: + rows spilled, maximum viewport height, whether a turn cancelled. These + are for quick reading. + +The live watch mode is a separate, optional path: run the scenario on a +real terminal so a human can feel it move. It is a thin wrapper over the +same loop. + +## Decisions + +- **A scenario file, not a code API.** Scenarios become data you can write, + share, and store next to the plan that describes them. +- **Flat steps, not turns.** The loop consumes keys and events in one + ordering, and cancel demands interleaving. `input` and `reply` hide the + ordering most of the time. +- **Two tiers of authoring.** `reply` with `chunk` is the pleasant path. + `events` is the full-control step for exact boundaries and tool flows. +- **Observation, not snapshots.** No saved "golden" text to diff against. + Your unit tests already assert behavior precisely. This is for review. +- **Reuse the play loop.** Replay tests the same code path the game runs, + not a parallel reimplementation. This is the trap most harnesses fall + into, and the reason to avoid it. +- **Behavior guarantees live in Rust tests.** The scenario file is the + input; a test helper loads it through the same harness and asserts a + structural property (every row committed once, a spilled row never + returns, the viewport never passes its cap). The assertion is about the + run, not about frozen text. + +## Screenshots and control codes + +The picture story has two tiers, and they answer different questions. + +- **Buffer screenshots** are deterministic and clean. They come from the + in-memory buffer, which holds the full content, layout, and styling. + They catch the great majority of rendering bugs: a misaligned table, a + wrong style, bad wrapping, gutter drift. They are the primary artifact. +- **A real terminal** carries the control codes a terminal actually sees: + cursor moves, scroll regions, the alternate screen, scrollback, resize + repaints. Those codes exist only at the boundary with a terminal, so the + buffer cannot show them. For the screen's true scrollback and resize + behavior, replay must run inside a real terminal emulator, a + pseudo-terminal like tmux, and capture the pane. + +The buffer screenshot is the default. The pseudo-terminal path is a +separate capability, for the small set of behaviors only a real terminal +shows. + +## Open decisions + +- How deep the built-in self-checking goes. The safe option is behavior + guarantees in Rust tests only. The richer option is optional structural + invariants on the scenario itself, like `"expect": {"spills": 2, + "viewport_max": 9}`, which are not text snapshots. Decide which. +- The fidelity bar for screenshots: buffer-only as primary, with the + pseudo-terminal path on call, or always captured from a real terminal. +- Whether to capture only the checkpoints and the final frame, or add the + option of an animated sequence for the whole run. +- Whether a tool line in `events` is plain text, which it is today, or + styled spans. + +## The seams + +The pieces meet at these contracts, so the tasks can land independently: + +```rust +// src/play/replay.rs +/// One scenario, loaded from a JSON file. Strictly validated. +pub struct Scenario { pub name: String, pub terminal: Size, pub steps: Vec } + +/// One step of a scenario, as authored. +pub enum Step { Input(String), Type(String), Key(Key), Reply { text: String, chunk: Chunk }, + Events(Vec), Resize { width: u16, height: u16 }, + Checkpoint(String), Idle(usize) } + +/// The report a replay run leaves behind, to render or assert against. +pub struct Report { + pub narration: Vec, + pub transcript: Vec, + pub frames: Vec, // buffer screenshots, plus checkpoints + pub summary: Summary, +} + +/// Runs `scenario` through the real play loop and returns what happened. +pub fn run(scenario: &Scenario) -> Result; + +/// The buffer at one point in the run, ready to render as a grid or an image. +pub struct Frame { pub buffer: Buffer, pub checkpoint: Option } +``` + +A `Report` holds everything an observer needs. The narration and transcript +are text. Each `Frame` is a grid of characters with a per-cell attribute +layer, which the same function can print as text or draw to a PNG. + +## Tasks + +Each task lands as one running step, so the plan builds toward something you +can run at every point. + +### 1. The scenario format and a runner that narrates + +The `Step` and `Scenario` types, the strict parser, and a runner that +executes the steps through the existing `play` loop with a `TestBackend`. +It prints the behavior narration: commits, forming rows, spills, viewport +heights, keys. This is the core, and it is runnable on its own. + +### 2. The `replay` subcommand + +Route `storied replay ` through a new `Command::Replay` in `cli.rs`. +The subcommand builds its own doubles, a `TestBackend`, scripted keys, and +a worker stub, and calls the play loop directly. It does not go through +`play::terminal::run`, which keeps building the real terminal. The output +is the narration. + +### 3. The rendered screen as a grid + +A function that turns a buffer into a character grid with a per-cell +attribute layer, printed as text. A `checkpoint` step in the scenario +emits a frame at a chosen point. This is the spatial view for a bug that +is about layout. + +### 4. The screenshot + +A renderer that draws a buffer frame to a PNG: each cell's character and +its colors and attributes become pixels. This is the image an agent and a +human can both look at. Works at any checkpoint, plus the final frame. + +### 5. Behavior tests on a scenario + +A test helper that loads a scenario through the same harness, so a Rust +test reads a file and asserts a structural property instead of hand- +building a step list. Moves the scenario file from "a way to watch" to "a +way to test" with no design change. + +### 6. Watch mode (optional, later) + +Run a scenario on a real terminal so a human can feel it move. This is the +pseudo-terminal path that shows control-code behavior, so it also carries +the scrollback and resize cases the buffer cannot. + +## Verification + +- Every scenario in the `plans/` and test suites runs deterministically: + the same file always produces the same narration and frames. +- A rendering bug drives a check: write the scenario, look at the + screenshot, fix the renderer, confirm the frame changes the way you + expected. +- A behavior test drives a property: load the scenario, assert the + invariant, watch it fail before the fix and pass after. +- The pre-commit bar stays: tests green, coverage at 100%, clippy clean. + +## Out of scope + +- Snapshot or golden-file comparison. The artifacts are for reading, not + diffing. +- A web or visual diff UI. The image and the grid are the review surface. +- Recording real sessions into scenarios. Replay of an authored scenario is + the core. Turning a live session into a file is a separate idea it may + earn later, once replay proves itself. diff --git a/plans/0008-streaming-markdown.md b/plans/completed/0008-streaming-markdown.md similarity index 100% rename from plans/0008-streaming-markdown.md rename to plans/completed/0008-streaming-markdown.md