diff --git a/docs-ai/064-agent-completion-signals/004-s2-work-note.md b/docs-ai/064-agent-completion-signals/004-s2-work-note.md index de713c37..b463c3b7 100644 --- a/docs-ai/064-agent-completion-signals/004-s2-work-note.md +++ b/docs-ai/064-agent-completion-signals/004-s2-work-note.md @@ -56,12 +56,10 @@ recorded in the existing 064 plan/action documents. resolution also closes the exact just-launched tab or pane before returning failure. - Focused tests for the final audit passed 22/22; the final full gate passed 2453 reported / 2455 verified tests with zero failures. -- Durable skill governance forbids editing skill files directly in this environment. Skill - Workshop proposal `prowl-cli-20260823-a59b46f453` was applied after explicit owner - authorization and installed the S2 command/rubric update as the managed workspace - `prowl-cli` skill with a clean scan. Workshop has no repository-target selector, so the - tracked `skills/prowl-cli/SKILL.md` copy remains unchanged; repository bundling is the one - remaining delivery gap rather than an approval wait. +- Skill Workshop proposal `prowl-cli-20260823-a59b46f453` was applied to the managed workspace + during the original implementation. Review follow-up then synchronized the approved dispatch + wait flow, heuristic rubric, and structured errors into tracked `skills/prowl-cli/SKILL.md`, + with a repository contract test preventing the bundled copy from drifting back to polling. ## Progress log @@ -98,6 +96,9 @@ recorded in the existing 064 plan/action documents. unexpected ancestry generation. The resolver test now injects a nil start-date provider for its fake PID graph. Focused tests passed 3/3, `make check` passed, and the full gate again passed 2453 reported / 2455 verified tests with zero failures. +- 2026-08-23 — Review follow-up synchronized the approved S2 flow and heuristic rubric into the + tracked bundled skill. Repository tests now require strict dispatch waiting, structured error + details, timeout re-arming, and explicit rejection of heuristic task completion. ## Fresh Debug E2E diff --git a/scripts/test_prowl_cli_skill.py b/scripts/test_prowl_cli_skill.py new file mode 100644 index 00000000..15925257 --- /dev/null +++ b/scripts/test_prowl_cli_skill.py @@ -0,0 +1,29 @@ +import pathlib +import unittest + + +class ProwlCLISkillTests(unittest.TestCase): + @classmethod + def setUpClass(cls): + root = pathlib.Path(__file__).resolve().parents[1] + cls.skill = (root / "skills" / "prowl-cli" / "SKILL.md").read_text() + + def test_documents_shipped_dispatch_wait_flow(self): + self.assertIn('dispatch="$(printf', self.skill) + self.assertIn('prowl agents wait --dispatch "$dispatch"', self.skill) + self.assertIn("DISPATCH_INCOMPLETE", self.skill) + self.assertNotIn("When\n`agents wait` ships", self.skill) + + def test_labels_heuristic_results_and_requires_screen_review(self): + self.assertIn('confidence == "heuristic"', self.skill) + self.assertIn("never treat heuristic evidence as task completion", self.skill) + self.assertIn("re-arm the wait", self.skill) + + def test_documents_structured_wait_errors(self): + self.assertIn(".error.details", self.skill) + self.assertIn("DISPATCH_NEEDS_INPUT", self.skill) + self.assertIn("WAIT_TIMEOUT", self.skill) + + +if __name__ == "__main__": + unittest.main() diff --git a/skills/prowl-cli/SKILL.md b/skills/prowl-cli/SKILL.md index 02854d0d..a6de7447 100644 --- a/skills/prowl-cli/SKILL.md +++ b/skills/prowl-cli/SKILL.md @@ -91,15 +91,24 @@ Review the current branch against its base. Report only actionable findings with EOF )" pane="$(printf '%s\n' "$launch" | jq -r '.data.target.pane.id')" -prowl read --pane "$pane" --last 200 --wait-stable --json +dispatch="$(printf '%s\n' "$launch" | jq -r '.data.dispatch.id')" +if result="$(prowl agents wait --dispatch "$dispatch" --include-screen 40 --json)"; then + printf '%s\n' "$result" | jq -r '.data.receipt.summary, .data.target.pane.id' +else + printf '%s\n' "$result" | jq '.error.code, .error.details' +fi ``` -The returned pane is the launched agent; `.data.launch` records the resolved Profile. `--prompt -` -requires a pipe or heredoc (never interactive stdin); Prowl carries up to 256 KiB outside initial -PTY input through a command portable across zsh, bash, and fish. Put larger requirement sets in -a repository file and prompt the Profile to read it. Use `read --wait-stable` today. When -`agents wait` ships, prefer it for deterministic completion before the final read. Add -`--background` when the split must not change focus or select a hidden anchor's tab/worktree. +The returned pane is the launched agent; `.data.launch` records the resolved Profile and +`.data.dispatch` is the exact assignment receipt. `--prompt -` requires a pipe or heredoc +(never interactive stdin); Prowl carries up to 256 KiB outside initial PTY input through a +command portable across zsh, bash, and fish. Put larger requirement sets in a repository file +and prompt the Profile to read it. Add `--background` when the split must not change focus or +select a hidden anchor's tab/worktree. + +Only a succeeded dispatch receipt proves that prompted assignment completed. The receipt may +arrive before the TUI paints its final response; if the next action sends another prompt to the +same pane, wait for an idle condition or read a stable screen first. Create a fresh tab in a listed worktree: @@ -140,7 +149,7 @@ prowl close --tab "$tab" --force --json # --force skips the GUI confirmation f ## Parsing JSON Output -Every `--json` response is `{ "ok", "command", "schema_version", "data": {...} }`; failures are `{ "ok": false, "command", "schema_version", "error": { "code", "message" } }`. Parser errors (bad flags) print plain text even with `--json`, so check the exit code before piping into `jq`. When JSON sits in a shell variable, use `printf '%s\n' "$json" | jq …` — zsh `echo` can turn `\u001B` escapes back into control characters. Pass shell values into `jq` with `--arg`. +Every `--json` response is `{ "ok", "command", "schema_version", "data": {...} }`; failures are `{ "ok": false, "command", "schema_version", "error": { "code", "message", "details"? } }`. Wait failures use governed `.error.details` for the retained dispatch record or last condition evidence. Parser errors (bad flags) print plain text even with `--json`, so check the exit code before piping into `jq`. When JSON sits in a shell variable, use `printf '%s\n' "$json" | jq …` — zsh `echo` can turn `\u001B` escapes back into control characters. Pass shell values into `jq` with `--arg`. Key fields by command: @@ -148,9 +157,11 @@ Key fields by command: - `agents` → `.data.agents[]` with `.status`, `.raw_state`, `.detection_reason`, `.type`, `.name`, `.pane.{id,focused,cwd}`, `.tab`, `.worktree`, `.project.{name,branch,path}`. - `agents read` → `.data.agent`, `.data.blocker.text`, `.data.result.{state,text}` — `pending`, `unavailable`, `missing`, `incomplete`, `too_large` carry no partial text. - `agents signal` → `.data.pane.{id,worktree_id}`, `.data.signal.{event,source,confidence,at,session_id,detail,claimed_origin}`; optional fields are omitted. +- `agents wait --dispatch` → `.data.receipt`, immutable `.data.target`, `.data.signals`, optional `.data.screen`; nonzero results retain the record and evidence under `.error.details`. +- `agents wait --until …` → `.data.observation.{status,raw_state,source,confidence,at,revision}`, `.data.signals`, and optional `.data.screen`. - `read` → `.data.text`, `.data.line_count`, `.data.truncated`, `.data.mode`, `.data.source`; `.data.stabilized` / `.data.waited_ms` with `--wait-stable`. - `send` → `.data.input`, `.data.wait.{exit_code,duration_ms}` when waiting, `.data.capture.{text,line_count,truncated}` with `--capture`. -- `create tab` / `open` → `.data.target.{pane,tab,worktree}`; `create pane` → `.data.anchor`, `.data.direction`, `.data.target`; Profile launches also include `.data.launch.{profile_id,profile_name,agent}`. +- `create tab` / `open` → `.data.target.{pane,tab,worktree}`; `create pane` → `.data.anchor`, `.data.direction`, `.data.target`; Profile launches also include `.data.launch.{profile_id,profile_name,agent}`, and prompted launches require `.data.dispatch.{id,state,created_at}`. - `profiles list` → `.data.profiles[]` with `.id`, `.name`, `.enabled`, `.runtime`, `.availability.{status,reason}`. Terminal text is `.data.text` (read) and `.data.capture.text` (send) — never `.content`, `.output`, or `.stdout`. @@ -158,16 +169,21 @@ Terminal text is `.data.text` (read) and `.data.capture.text` (send) — never ` ## Reading Agent Output - For Codex/Claude Code, `prowl agents read` beats scraping: check `.data.agent.status`, inspect `.data.blocker.text` before answering a prompt with `send`/`key` (read and write are not atomic), and only trust `.data.result.text` when `state == "complete"`. `--result-only` prints the raw trusted result and fails otherwise; it cannot combine with `--json`. -- For everything else, `prowl read --pane "$pane" --last 200 --wait-stable --json` blocks until the screen stops changing. `task.status` flips to `idle` before a TUI finishes painting, so poll it only to wait for a `working` agent, then still read with `--wait-stable`: +- For an unpaired or manually launched agent, use one condition wait instead of a polling loop: ```bash -for i in 1 2 3 4 5 6; do - task_state="$(prowl list --json | jq -r --arg p "$pane" '.data.items[] | select(.pane.id == $p) | .task.status')" - [ "$task_state" = idle ] && break - sleep 1 -done +result="$(prowl agents wait "$pane" --until idle --include-screen 40 --timeout 600 --json)" +printf '%s\n' "$result" | jq '.data.observation, .data.screen' ``` + Exact/high evidence can establish the requested observable condition. If + `jq -e '.data.observation.confidence == "heuristic"'` matches, inspect the included stable + screen and, when needed, `prowl agents read "$pane" --json`. A finished answer with an empty + prompt is positive evidence; a spinner/tool footer means working; a permission dialog or + explicit question means blocked. Always use task context, and never treat heuristic evidence as task completion or perform destructive follow-up from it alone. A timeout leaves the task unresolved: inspect the pane, then re-arm the wait rather than assuming completion. +- `DISPATCH_NEEDS_INPUT` means the exact worker needs intervention. `DISPATCH_INCOMPLETE` means + its turn ended without the required receipt. Both retain a pending receipt, so inspect + `.error.details`, respond or nudge as appropriate, and wait again. - Rendered screens can truncate or fold content. When you need an agent's complete output, have the command write a file (`… > /tmp/out.txt`) and read that; shell redirection avoids the agent's own sandbox prompts. - `read` returning fewer lines than `--last` with `truncated: false` means the pane simply has less history — do not retry. `--source detection` returns the exact detector input instead of the viewport; it exists for diagnosing agent-state detection (see `components/agent-detection.md` in the docs folder), not for everyday reading. @@ -197,7 +213,10 @@ done - `PROFILE_NOT_FOUND` / `PROFILE_NOT_UNIQUE`: re-run `prowl profiles list --json`; choose an enabled Profile UUID. - `NO_ACTIVE_PANE`: focused-pane targeting found nothing — pass `--pane`. `SOURCE_REQUIRED`: a caller-owned command (`agents signal`, selector-free `handoff`) could not map process ancestry to a Prowl pane. `AGENT_GONE`: the signal's caller pane closed before recording. - `EMPTY_INPUT`, `INVALID_ARGUMENT`, `UNSUPPORTED_KEY`, `INVALID_REPEAT`: fix the arguments (`prowl --help`). -- `CAPTURE_UNSUPPORTED`: drop `--capture` and use `read --wait-stable` or file redirection. `WAIT_TIMEOUT`: raise `--timeout` or use `--no-wait`. +- `CAPTURE_UNSUPPORTED`: drop `--capture` and use `read --wait-stable` or file redirection. +- `WAIT_TIMEOUT`: inspect `.error.details`, then re-arm the wait if the task remains active. +- `DISPATCH_FAILED` / `DISPATCH_ABANDONED` / `AGENT_GONE`: the exact dispatch is terminal; inspect its retained record and immutable target in `.error.details`. +- `DISPATCH_NEEDS_INPUT` / `DISPATCH_INCOMPLETE`: the dispatch remains pending; inspect, intervene if appropriate, and wait again. - `PATH_NOT_FOUND` / `PATH_NOT_DIRECTORY` / `PATH_NOT_ALLOWED`: fix the path given to `open` or `create tab --path`. ## Handing Off Your Task @@ -220,4 +239,4 @@ Required sections are `## Objective`, `## Current State`, and `## Next Steps`; o ## Command Set -`list`, `agents`, `agents read`, `agents signal`, `profiles list`, `read`, `send`, `key`, `focus`, `create tab`, `create pane`, `close`, `handoff to`, `handoff save`, and `open` (default). There is no CLI `quit`; close temporary tabs or panes with an explicit `close`. `tab create`, `tab close`, and `pane close` remain deprecated aliases for one release. +`list`, `agents`, `agents read`, `agents signal`, `agents dispatch-complete`, `agents dispatch-abandon`, `agents wait`, `profiles list`, `read`, `send`, `key`, `focus`, `create tab`, `create pane`, `close`, `handoff to`, `handoff save`, and `open` (default). There is no CLI `quit`; close temporary tabs or panes with an explicit `close`. `tab create`, `tab close`, and `pane close` remain deprecated aliases for one release.