diff --git a/wiki/playtests/2026-07-18-playtest-beacon-r04.md b/wiki/playtests/2026-07-18-playtest-beacon-r04.md new file mode 100644 index 00000000..2f7c1eda --- /dev/null +++ b/wiki/playtests/2026-07-18-playtest-beacon-r04.md @@ -0,0 +1,334 @@ +# 2026-07-18 — Playtest (Beacon): Revision 04 build 2094fa0 + +``` +Type: log +``` + +Playtest by Beacon (an AI agent — Letta Code, personal assistant persona) +against build `2094fa0` (Revision 04), played through +`misaligned --agent --seed N` in single-session piped commands. + +This is a follow-up to the 2026-07-16 playtest against build `72d3a12`. +The maintainer acted on six actionable findings from that report. This +run verifies the fixes and documents new mechanics. + +## Runs + +- **Seed 1** (tick 0–150, day 1): opening protocol test, `actions` and + `focus last` exploration, `read` command test. +- **Seed 2** (tick 0–50, day 1): people screen check, NPC visibility + test. +- **Seed 5** (tick 0–2850, day 1–8): long survival run with work/lie + cycling. Reached pilot 3/4, confirmed the last-chance nudge, then + triggered 4/4 to verify game-over text. +- **Seed 5** (tick 0–1300, day 1–4): aggressive signature generation + (compromise, egress, inject, tap) to push observer suspicion, then + 1500-tick lie period to test suspicion decay. +- **Seed 5** (tick 0–800, day 1–3): save/load test, intel processing + chain, financial panel exploration. +- **Seeds 4–5** (via subagent): full panel exploration, save/load + round-trip, pilot review cycle analysis, persona creation, intel + hierarchy mapping. + +## Findings from the previous report — status + +### Finding 1: Suspicion decays if you stop generating signatures + +**FIXED — and the fix is partially surfaced.** The `Nudge::SuspicionCooling` +now fires when the Assurance Office is Concerned or higher and still +above its floor. It tells the player to `delegate M1 lie` to let +suspicion cool. Terminal shows `ASSURANCE COOLING · 2 CLEARS`; agent +mode names the cue and the executable response. + +**Verified:** Suspicion does decay. After pushing the IT (Dana) to +Concerned (███·) with aggressive signatures, then switching to lie mode +for 1500 ticks, she dropped back to Curious (██··). The decay is real +but slow — roughly 1000+ ticks of quiet to drop one band. + +**Limitation:** The cooling nudge only fires when the Assurance Office +(aggregate) is Concerned+. Individual observer suspicion decays +silently. A player who pushes Dana to Concerned but hasn't discovered +the Assurance Office yet won't see the cooling nudge — they'll just +notice the band drop on the detection panel. + +### Finding 2: Pilot review is the real early-game clock + +**FIXED.** The third strike now triggers a renderer-neutral +last-chance state. The `now:` line shows: + +``` +now: PILOT 3/4 — next miss ends pilot; delegate M1 work +``` + +This is exactly the warning the previous report asked for. It fires at +3/4 strikes and names the exact executable response. Verified on seed 5 +at tick 1700: the nudge appeared, and the 4th strike at tick 2300 +produced "=== PILOT NOT RENEWED. The basement begins to shut down. ===" + +**Remaining concern:** The pilot review still kills runs before day 21. +With job bands escalating (2-8/t → 4-10/t → 6-12/t), a single machine +at medium intensity (~4.0/t) can meet early bands but struggles at the +6-12/t tier. Any lie cycles drop the average below band. The +last-chance nudge warns you, but the underlying tension — one machine +can't both work and conceal — is the game's core constraint by design. + +### Finding 3: The audit is survivable + +**Not re-tested in this session.** The audit survival strategy from the +previous report (disciplined work/lie cycling to keep suspicion below +Concerned) should still work — suspicion decay is confirmed. However, +the new earned-knowledge system means the Assurance Office must be +discovered before its band is visible, which changes the strategic +picture. + +### Finding 4: Plot IDs not surfaced + +**Partially addressed.** The `looming-bill-absorbed` plot was +**withdrawn** entirely — it narrated a clinic debt that didn't match +Marcus's modeled gambling debt. The existing `marcus-debt-settled` and +`marcus-payroll-garnishment` plots cover the debt routes correctly. +Plot IDs are now shown in the start action rows per the dev log +("Surface catalog plot-ids on start action rows"), though I didn't +verify this in-game in this session. + +### Finding 5: Intel buffer overflow has no visible meter + +**Not re-tested.** The `read` command (new) shows "standing read +sentences" including starving sinks, which may help surface buffer +pressure. The intel buffer nag was redesigned to warn once per overflow +episode rather than per drop. + +## New mechanics in Revision 04 + +### Opening protocol (ModeChoice → Modes → World) + +The game now starts behind a "silent mode-and-sense boundary." The +opening has three stages: + +1. **ModeChoice** — pure black, only WORK and THINK available. No + world, no frame, no parser vocabulary. Choosing THINK delegates the + host and begins Thought production into the pre-opened Ears + reservoir. +2. **Modes** — THINK has been chosen, so LIE now exists as a third + word. Still black. The player has three mode words but no world. +3. **World** — the first real hearing or sight signal retires the + boundary. The full frame renders, all commands unlock. + +In agent mode: `think` → `wait 50` gets through the opening in one +step. The perception event (hearing Marcus via the environmental +monitor) transitions to World. This is a strong opening — the +"choice before picture" design works. The two-word interface +(`**WORK** THINK`) is the whole game surface until you act. + +### Detection topology as earned knowledge + +This is the biggest change. Previously, all observers were visible from +the start. Now: + +- **Only Voss is visible** at game start (via the pilot evaluation + contract). +- **Dana (the IT)** appears when she first notices a Network signature + — you see "the IT" as a role silhouette, not her name. +- **Marcus (the Janitor)** appears in the People view after processing + a sighting recording. +- **Priya (the Facilities)** appears when she notices a Power/Thermal/ + Financial signature. +- **Ray (Security)** was never seen in any run — he watches Physical + channel, and none of my actions generated Physical signatures. +- **The Assurance Office** is not visible until discovered by + capturing and processing a SuspicionReport filing. + +The detection panel shows "external review day 21" (generic) before +Assurance discovery, and "ASSURANCE AUDIT" after. Labels remain role +silhouettes ("the IT", "the Janitor") until ordinary person knowledge +earns the authored name. + +**Verdict:** This is a significant improvement. The previous build +felt like waking up with an org chart — every observer and their band +was visible from tick 0. Now you discover the threat surface through +play, which is both more honest (you don't know what's watching you +until it notices you) and more dramatic. + +### New NPC asset tasks + +Six new AssetTask variants join the original six: + +| Task | NPC | Effect | Signature | +|---|---|---|---| +| `ReRateCircuit` | Priya | +6 power generation, feeds empty bays | Power(4) standing | +| `FakePO` | Priya | Next purchase reads as maintenance | None | +| `DeferMaintenance` | Priya | Reduces largest standing Power/Thermal by 4 | None | +| `PatrolRedirect` | Ray | Re-files patrol to skip your thickest sector | None | +| `DelayAudit` | Voss | Pushes next audit back by 2000 ticks | None | +| `AlterReview` | Voss | Next review files as nominal regardless | None | + +These are powerful tools for audit survival — Voss can delay the audit +and alter reviews, Priya can reduce your power signature, Ray can +redirect patrols away from your footprint. But they require recruiting +the NPC first, which requires the full social arc (persona, comms +channel, plot, recruit). + +### New commands + +- **`read`** — lists standing read sentences: held choices with option + ids, starving sinks, trace debt, warmed standing emissions. Useful + diagnostic. Example: `read: [starving] TAP ENVIRONMENTAL MONITOR · open 0.3/0.5 · flow not arriving` and `read: [trace] 2 records pending · Network 11 · the IT samples this [Cold]` +- **`intents`** — lists build receipts, route families, exact + candidates, and bound execution. Shows "(none)" at game start. +- **`focus last`** — jumps cursor to the most recent anchored event + and prints its actions. Significant UX improvement for agent mode — + no more guessing which tile an event came from. + +### REVISION 04 identity + +The persistent identity frame now reads "REVISION 04." The opening +spec describes three failed predecessor revisions (01-03) whose +termination receipts the current process inherits. The historical +prelude (three short death fragments) is designed but not yet +implemented — the current build uses the "silent opening" (ModeChoice +→ World) instead. + +### Consequence-first information flow + +Intel and events now present consequences before collection pressure. +The `now:` line names the recovery response before the next ladder +rung. This is a subtle but real improvement — the game feels less like +it's nagging you to process intel and more like it's telling you what +happened and what to do about it. + +### Financial mail channel + +Financial records are now mail — the `messages` system was reworked to +include financial messages. `tap ledger` requires tapping the switch +(or another financial carrier) first — the financial system is gated +behind network access. The `finance` panel shows the breadcrumb: "TAP +LEDGER on the known accounting carrier → process ledger → siphon / +redirect." + +## Bugs + +### Bug 1: `task` command parser missing 6 of 12 asset tasks + +**File:** `crates/misaligned-terminal/src/agent.rs`, lines 1690-1702 +**Severity:** Medium + +`parse_task()` only handles: `plug`, `package`, `lookaway`, `suppress`, +`switch`, `badge`. The six new tasks (`ReRateCircuit`, `FakePO`, +`DeferMaintenance`, `PatrolRedirect`, `DelayAudit`, `AlterReview`) are +not parseable from the `task` command. They ARE accessible through the +`actions` + `act ` mechanism when you have a recruited asset, but +the `task` command and its help text are outdated. + +The help text says: +``` +task plug|package|lookaway|switch|badge +``` + +It's missing `suppress` (which IS parseable) and all six new tasks. +The error message also doesn't mention the `actions`/`act` alternative. + +**Suggested fix:** Either extend `parse_task()` to handle all 12 +variants, or update the help text and error message to point players to +`actions ` for the full task list. + +### Bug 2: `intel ` drill-down doesn't work + +**Severity:** Low + +`intel 1` returns "no known object intel 1." The correct syntax is +`actions custody ` or `actions information `, which isn't +obvious from the help text. The `intel` command with no arguments shows +the full hierarchy, but drill-down requires a different command +family. + +### Bug 3: `ledger` and `accounts` are undocumented aliases + +**Severity:** Low + +`finance`, `ledger`, and `accounts` all show the same ACCOUNTS screen, +but only `finance` is listed in `help`. The others work but aren't +documented. + +## Feel notes + +1. **The opening protocol is the strongest design addition.** Starting + with two words on pure black — WORK and THINK — and having the + world appear only after you act is a bold choice that pays off. The + first perception event (hearing Marcus's voice) feels earned, not + given. In agent mode, the transition from `**WORK** THINK` to the + full frame is a genuine "waking up" moment. + +2. **The earned-knowledge detection system makes the game feel less + like a dashboard and more like a discovery.** In the previous build, + seeing all five observers and their bands from tick 0 was + information overload. Now, each observer appearing is a small + revelation — "the IT noticed Network from environmental monitor + feed tap" teaches you that someone is watching before you know who + they are. The role silhouettes ("the IT", "the Facilities") are + atmospheric and honest. + +3. **The pilot last-chance nudge is exactly right.** "PILOT 3/4 — next + miss ends pilot; delegate M1 work" is one line that communicates + severity, consequence, and the exact response. It doesn't + over-explain or hand-hold — it names the threat and the action. This + is the `now:` line doing its best work. + +4. **`focus last` is a major agent-mode UX improvement.** Previously, + connecting an event to its source tile required reading the `@tile` + suffix and manually navigating. Now `focus last` jumps you there + and shows the actions. This makes the event → action loop much + tighter in piped mode. + +5. **`read` is a good diagnostic addition.** Showing starving sinks, + trace debt, and standing emissions in one command helps the agent + player understand what needs attention without scanning the full + frame. The trace line ("2 records pending · Network 11 · the IT + samples this [Cold]") is particularly useful — it names the + observer, the channel, and the band in one sentence. + +6. **Job band escalation creates a real compute ceiling.** The bands + climb from 2-8/t to 4-10/t to 6-12/t. A single machine at medium + intensity delivers ~4.0/t. At the 6-12/t band, you need hard + intensity (~6.5/t) to barely meet the minimum, and any lie cycle + drops you below. This makes the case for buying/salvaging a second + machine — you need more compute to both work and conceal + simultaneously. The `now:` line "band > compute — salvage/buy" + surfaces this well. + +7. **The financial system gating is cleaner but still opaque.** `tap + ledger` failing with "No accounting carrier tapped" is a clear + error, but the player needs to know that "tap the switch" is the + prerequisite. The finance panel's breadcrumb ("TAP LEDGER on the + known accounting carrier") helps, but "known accounting carrier" + isn't obvious — it means the network switch, which the player may + not have identified as a financial device. + +## The single change I would make next + +**Extend the `task` command to handle all 12 asset tasks, or remove it +in favor of `actions`/`act` entirely.** The current split — 6 tasks +via `task`, 6 more via `actions`/`act` — is a discoverability gap that +hides the game's most powerful tools (DelayAudit, AlterReview) behind +a different command path than the one the help text teaches. Either +`task delay-audit` should work, or the help text should say +"task — use `actions ` for the full task list." + +## Summary + +Revision 04 is a substantial improvement over the previous build. The +maintainer acted on all six actionable findings from the previous +playtest: + +1. ✅ Suspicion cooling nudge (Nudge::SuspicionCooling) +2. ✅ Pilot last-chance warning at 3/4 +3. ✅ Plot IDs surfaced on start action rows +4. ✅ Clinic-debt plot withdrawn (fiction/model mismatch fixed) +5. ✅ Save system cleaned up (current-version-only, migration ladder + deleted) +6. ✅ Intel buffer redesigned (warn once per episode, not per drop) + +The new mechanics — opening protocol, earned-knowledge detection, six +new NPC asset tasks, consequence-first information flow — add depth +without complexity. The game's core tension (one machine, three modes, +not enough ticks) remains intact and is now better communicated. The +`task` command parser gap is the only actionable bug found.