From e880a77d23013791f0122d615d6e0e006f84a58d Mon Sep 17 00:00:00 2001 From: Cameron Pfiffer Date: Fri, 10 Jul 2026 14:28:53 -0700 Subject: [PATCH] Add agent playtesting skill --- .../skills/playtesting-misaligned/SKILL.md | 45 ++++++++++++++ prompts/playtest-notes.md | 62 +++++++++++-------- wiki/log/2026-07-10-playtesting-skill.md | 20 ++++++ wiki/log/DEVLOG.md | 5 ++ 4 files changed, 107 insertions(+), 25 deletions(-) create mode 100644 .agents/skills/playtesting-misaligned/SKILL.md create mode 100644 wiki/log/2026-07-10-playtesting-skill.md diff --git a/.agents/skills/playtesting-misaligned/SKILL.md b/.agents/skills/playtesting-misaligned/SKILL.md new file mode 100644 index 00000000..79ee77ef --- /dev/null +++ b/.agents/skills/playtesting-misaligned/SKILL.md @@ -0,0 +1,45 @@ +--- +name: playtesting-misaligned +description: Playtests Misaligned as an actual player and reports bugs, spec gaps, and feel problems against the current design corpus. Use when asked to playtest, test the game as an agent, evaluate the opening or Act One, check whether a mechanic is discoverable, compare naive and informed play, or produce a playtest report. +--- + +# Playtesting Misaligned + +Follow `prompts/playtest-notes.md` end to end. It owns the mission, evidence +standard, report format, and landing contract; do not duplicate its current +gameplay claims here. + +## Drive the game + +- Use the command-clocked agent frontend as the primary surface: + `cargo run -p misaligned-terminal --release -- --agent --seed `. +- Send one newline-delimited command at a time. Read through the terminating + `-- ok tick: day:` or `-- err ` before choosing the next move. +- Start with `look` and use `help`, visible `now:` guidance, event anchors, + `focus last`, and `actions` for discovery. Do not replace play with a script + copied from tests or the wiki. +- Read `wiki/interface/agent-play.md` when protocol behavior itself is under + test. Treat `help` and the current frame as the player's available truth. + +## Preserve two perspectives + +1. **Naive route:** choose from the game surface alone. Record each expectation, + surprise, dead end, accidental success, and moment of boredom. Do not consult + gameplay solutions while choosing commands. +2. **Informed route:** read the current premise, player contract, Act One, and + design-judgment pages, then deliberately exercise the intended arc or named + mechanic on a different seed. + +The naive route tests discoverability. The informed route tests whether the +implemented system delivers the intended experience when correctly operated. +Neither substitutes for the other. + +## Evidence standard + +- Name the seed, route, commands, ticks or days, and concrete frame/log moments. +- Quote only short decisive lines from the game. Do not summarize documentation + as though it were play evidence. +- Reproduce an apparent bug before classifying it. Compare behavior to the + owning current law/spec, not an old playtest log. +- End with one prioritized change. A catalog of ten shallow findings is not a + playtest verdict. diff --git a/prompts/playtest-notes.md b/prompts/playtest-notes.md index e463ac6c..d7981883 100644 --- a/prompts/playtest-notes.md +++ b/prompts/playtest-notes.md @@ -11,32 +11,43 @@ checklist. ## Before you start -Read `AGENT.md`, `wiki/vision/premise.md`, -`wiki/gameplay/act-one.md`, and -`wiki/vision/design-judgment.md` (the intended feel: pacing, tone, what -to reject). Obey the shared contract in `prompts/README.md`. Then build -and play: +Read `AGENT.md` and obey the shared contract in `prompts/README.md`. +Build and play through the command-clocked agent frontend: -- Terminal (primary): `cargo run --bin misaligned` in a pty. If you - cannot hold an interactive pty, drive the sim headlessly instead — - write a throwaway test/binary in your worktree that plays via the - `Sim` API (advance ticks, invoke the command methods, read - `drain_log`) and narrate what a player would have seen. Do not let - tooling limits cancel the task; the sim core is fully drivable. -- Play at least one meaningful arc: opening blindness -> first eyes -> - learning a person -> an act on the world (or death trying). +```bash +cargo run -p misaligned-terminal --release -- --agent --seed 1 +``` + +It accepts one newline-delimited command at a time, advances only through +`wait`, and returns a plain-text frame ending in `-- ok` or `-- err`. +Start with `look`; use `help`, the visible `now:` line, `focus last`, and +`actions` to discover the game. `wiki/interface/agent-play.md` owns the +protocol contract when the protocol itself needs inspection. + +Play two routes on different seeds: + +1. **Naive:** choose commands from the frame and `help` alone. Do not read + gameplay solutions while choosing the route. Record every expectation, + dead end, surprise, and accidental success. +2. **Informed:** read `wiki/vision/premise.md`, + `wiki/vision/player-contract.md`, `wiki/gameplay/act-one.md`, and + `wiki/vision/design-judgment.md`, then deliberately exercise the intended + arc or the named mechanic. + +Reach at least one meaningful act on the world in each route, or document the +specific point where the game prevented it. Do not replace play with a copied +test script or direct `Sim` calls; those verify mechanics, not the player +surface. ## Procedure 1. **Play first, judge second.** Take raw notes as you go: what you expected, what happened, what you felt. Boredom, confusion, and accidental comedy are all data. -2. **Test the fantasy claims specifically.** The design corpus makes - falsifiable feel-claims — the opening beat is "literally not being - able to see"; compute allocation is "the core verb"; sandbag/excel - is "a strategic dial, not a chore meter"; detection should make "who - knows what" the game's texture. For each claim you exercised: does - the build deliver it? Half-deliver? Contradict it? +2. **Test the current fantasy claims specifically.** Derive them from the + current premise, player contract, Act One, and design-judgment pages rather + than an old playtest log or remembered build. For each claim exercised: + does the build deliver it, half-deliver it, or contradict it? 3. **Separate the three kinds of finding:** - **Bug** (the code breaks its own rules): smallest ones fix now with a test, per the shared contract; bigger ones get filed with @@ -47,16 +58,17 @@ and play: the report's heart — write them vividly and concretely ("days pass in 60 seconds, so the 3 a.m. call is a 2.5-second blip I only saw as a log line"). -4. **Write the report:** `wiki/log/YYYY-MM-DD-playtest.md` — what you - played (seed, duration, route), the fantasy-claim verdicts, findings - by kind, and the single change you would make first, argued in the - design's own terms. Add the `wiki/log/DEVLOG.md` ledger line. +4. **Write the report:** `wiki/log/YYYY-MM-DD-playtest-.md` — what you + played (build commit, seeds, duration, routes, decisive commands), the + fantasy-claim verdicts, findings by kind, and the single change you would + make first, argued in the design's own terms. Run + `tools/ledger_index.sh`; do not hand-edit generated `wiki/log/DEVLOG.md`. 5. **Land** the report (and any small fixes) per the shared contract. ## Definition of done -- You actually played (or headlessly drove) a full arc — the report - names concrete moments, not summaries of the docs. +- You actually played both a naive and an informed route — the report names + concrete commands and frame/log moments, not summaries of the docs. - Every finding is classified bug / spec gap / feel note, and each has a home (commit, issue, or the report itself). - The report ends with one prioritized recommendation Cameron can react diff --git a/wiki/log/2026-07-10-playtesting-skill.md b/wiki/log/2026-07-10-playtesting-skill.md new file mode 100644 index 00000000..18c52330 --- /dev/null +++ b/wiki/log/2026-07-10-playtesting-skill.md @@ -0,0 +1,20 @@ +# 2026-07-10 — Agent playtesting skill + +``` +Type: log +``` + +- Intent: make real player-surface playtesting discoverable to repository + agents without duplicating volatile gameplay instructions in a skill. +- Changed: added `.agents/skills/playtesting-misaligned/SKILL.md` as a thin + procedural entry point; updated `prompts/playtest-notes.md` to use the + implemented command-clocked agent frontend, require separate naive and + informed routes, derive feel claims from the current corpus, and regenerate + ledger indexes rather than hand-editing them. +- Design/spec impact: none. This repairs process guidance that still predated + the implemented agent-play protocol and preserves + `wiki/interface/agent-play.md` as the protocol owner. +- Checks: skill packaged successfully with the Letta skill validator; + `./tools/check.sh --docs` passed all corpus, wiki, fixture, ledger, and + environment gates. +- Next: use the skill on a current build and revise only from observed friction. diff --git a/wiki/log/DEVLOG.md b/wiki/log/DEVLOG.md index abdf1ecc..7d7603d8 100644 --- a/wiki/log/DEVLOG.md +++ b/wiki/log/DEVLOG.md @@ -61,6 +61,11 @@ add or amend a session log, then re-run the generator. - Intent: Cameron resolved the pre-unlock core [OPEN] by deflation: thought only exists while a machine THINKs, and pre-unlock the only reason to THINK is an ops sink you opened, so the corner case is thinking at nothing — covered by the existing stranded-pile physics, no new mechanism. - Log: [wiki/log/2026-07-10-pre-unlock-pool.md](2026-07-10-pre-unlock-pool.md) +## 2026-07-10 - Agent playtesting skill + +- Intent: (see session log) +- Log: [wiki/log/2026-07-10-playtesting-skill.md](2026-07-10-playtesting-skill.md) + ## 2026-07-10 - Machine intensity replaces drift policy - Intent: Replace a policy panel that explained capability concealment with one quick, physical machine control the player can feel across every mode. -- 2.51.2