diff --git a/.agents/skills/sync-pi-evals/SKILL.md b/.agents/skills/sync-pi-evals/SKILL.md new file mode 100644 index 0000000..25ecad8 --- /dev/null +++ b/.agents/skills/sync-pi-evals/SKILL.md @@ -0,0 +1,101 @@ +--- +name: sync-pi-evals +description: Vendor and refresh the unpublished pi-evals harness from earendil-works/pi into tests/evals/vendor, and keep pi-* devDependencies compatible with it. Use when asked to set up, refresh, or update the evals harness or vendored pi-evals code in pi-processes. +--- + +# Syncing vendored pi-evals + +`@earendil-works/pi-evals` (`packages/evals` in `earendil-works/pi`) is `"private": true` and not published to npm. +Until it ships publicly, the harness source is copied into `tests/evals/vendor/pi-evals/` (gitignored) and patched. + +Why not a pnpm git dependency: installing `github:earendil-works/pi#&path:packages/evals` does fetch the code, +but the harness must be patched, and `pnpm patch` on that dependency fails with +`ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED` — it tries to prepare the whole `pi-monorepo` root and demands an +`allowBuilds` entry for it. Revisit if pi publishes the package or upstreams extension injection. + +Do not run the eval suite as part of syncing: evals make real, paid model calls. `pnpm typecheck:evals` is the free +check that a vendor drop is sound. + +## Run the script + +```bash +./.agents/skills/sync-pi-evals/scripts/sync-pi-evals.sh # pinned default ref +./.agents/skills/sync-pi-evals/scripts/sync-pi-evals.sh --ref v0.85.0 # tag, branch, or SHA +``` + +It clones the pi repo at the ref, copies `pi-harness.ts` and `vitest-evals/`, records the upstream `package.json` +as `upstream-package.json`, applies every patch in `tests/evals/patches/`, and writes `tests/evals/vendor/SYNC.json` +with the ref, resolved SHA, and date. Re-running is safe: `rsync --delete` restores pristine files before patching. + +The script only mechanises the copy. The two decisions below are yours. + +## Decision 1: which ref + +Read `tests/evals/vendor/SYNC.json` first to see what is currently vendored. + +Default rule: use the git tag matching the `@earendil-works/pi-coding-agent` version pinned in this repo's +`package.json` devDependencies (devDependency `0.84.1` → tag `v0.84.1`). The harness calls +`AgentSession` / `createAgentSessionServices` APIs, so it must match the coding-agent version we build against. + +Exception: if the harness change you need landed on `main` after the latest tag, pin the exact SHA rather than a +branch, and check whether that commit's `packages/coding-agent` API is compatible with the version installed here. +If it is not, stop and ask the user before syncing — see decision 2. + +Never sync from a floating branch. The script accepts one, but the result is not reproducible. + +## Decision 2: dependency reconciliation — always ask before bumping + +After syncing, compare `tests/evals/vendor/pi-evals/upstream-package.json` devDependencies against this repo's +`package.json`: + +- `@earendil-works/pi-ai` +- `@earendil-works/pi-coding-agent` +- `vitest` +- `vitest-evals` (already a devDependency here, pinned exact) + +Then run `pnpm typecheck:evals`. It compiles the vendored and patched harness against this repo's installed +versions and is the real compatibility signal. + +If it passes and versions match: done. + +If versions differ or the typecheck fails: + +1. Report the mismatch explicitly: which package, the version pinned here, the version the vendored code expects, + and whether that version is published (`npm view @earendil-works/pi-coding-agent version`). +2. If it is published: propose bumping `@earendil-works/pi-ai`, `@earendil-works/pi-coding-agent`, + `@earendil-works/pi-tui`, and `vitest` (plus matching `peerDependencies` ranges), and ask for confirmation + before running `pnpm add -D @`. This repo uses pnpm. Never bump silently. +3. If it is not published: say the vendored harness is ahead of what is installable, and ask whether to re-sync + from the last tag (losing the new harness feature) or wait for the next pi release. + +## The patches + +`tests/evals/patches/0001-harness-additional-resource-paths.patch` makes four changes, all required: + +1. adds `additionalExtensionPaths` / `additionalSkillPaths` options, +2. always forwards `resourceLoaderOptions` (upstream only does when `transformSystemPrompt` is set), +3. relaxes the zero-extension guard to expect exactly the injected extensions, +4. widens `resolveModelSelection`'s `environment` parameter to `Record`, because + `@types/node` 25 here makes `ProcessEnv` incompatible with upstream's inline literal type. + +Upstream's harness deliberately refuses to load extensions, so without 1–3 there is no way to get the processes +extension into an eval session. + +If the script reports a failed patch, upstream moved. Re-derive the change by hand against the newly vendored +`pi-harness.ts`, regenerate with `diff -u`, and commit the updated patch file. Do not skip it. + +## After syncing + +- `pnpm typecheck:evals` must pass. `pnpm typecheck` excludes `tests/evals` on purpose, so it stays green without a + vendor drop. +- Verify with `git status` that nothing under `tests/evals/vendor/` is staged; it is gitignored. +- CI runs this same script in the gated `evals` job, so a working sync locally means a working sync there. + +For writing, running, and verifying the evals themselves, see the `writing-pi-evals` skill and +`tests/evals/README.md`. + +## Re-sync when + +- `@earendil-works/pi-coding-agent` is bumped here. +- An eval needs a harness capability that only exists in a newer pi commit. +- `SYNC.json` predates a pi release this repo has adopted. diff --git a/.agents/skills/writing-pi-evals/SKILL.md b/.agents/skills/writing-pi-evals/SKILL.md new file mode 100644 index 0000000..2999dc3 --- /dev/null +++ b/.agents/skills/writing-pi-evals/SKILL.md @@ -0,0 +1,145 @@ +--- +name: writing-pi-evals +description: Write, run, and verify behavioral evals for the processes extension using the vendored pi-evals harness. Use when adding or debugging an eval in evals/, running the eval suite, or checking whether an eval result is trustworthy. +--- + +# Writing pi-processes evals + +Evals drive a real model through a real `AgentSession` with the processes extension loaded, then assert on the +resulting tool calls. They measure agent behaviour (does it poll? does it reach for a log watch?), not unit +correctness. Unit behaviour belongs in `src/**/*.test.ts` and `extensions/**/*.test.ts`. + +Prerequisite: the vendored harness must be present at `tests/evals/vendor/pi-evals/`. If it is missing, run +`./.agents/skills/sync-pi-evals/scripts/sync-pi-evals.sh` (see the `sync-pi-evals` skill). Nothing here works +without it. + +## Layout + +- `tests/evals/harness.ts` — committed. Wraps the vendored `createPiCodingAgentHarness` with the processes extension + path, plus tool-call extraction helpers. Import this, never the vendored file directly. +- `tests/evals/*.eval.ts` — committed. One file per behaviour area. +- `tests/evals/patches/` — committed. Patches applied to the vendored harness during sync. +- `tests/evals/global-setup.ts` — committed. Generates `agent-dir/models.json` from `.env.test` before any run. +- `tests/evals/README.md` — committed. Explains which files are ours and which are vendored. +- `tests/evals/vendor/` — gitignored. Produced by `sync-pi-evals`. +- `tests/evals/vitest.config.ts` — separate config; `include: ["tests/evals/**/*.eval.ts"]`, long timeouts, no parallelism, no + global mocks (the unit `vitest.config.ts` mocks `node:fs` via memfs, which would break real process spawning). +- `tests/evals/tsconfig.json` — separate project so `pnpm typecheck` stays green when `tests/evals/vendor/` is absent. + +## Running + +```bash +pnpm test:evals # whole suite +pnpm test:evals tests/evals/smoke.eval.ts # one file +pnpm test:evals -t "smoke" # one test by name +pnpm typecheck:evals # after sync or harness edits, costs nothing +``` + +No environment variables required. `tests/evals/vitest.config.ts` pins everything that decides which model runs: + +- `PI_CODING_AGENT_DIR` to `tests/evals/agent-dir/`, so runs never read or spend from your real `~/.pi/agent`. +- `PI_PROVIDER` / `PI_MODEL` to `aperture` / `syn:small:text`, pinned rather than inherited so an ambient + `PI_PROVIDER=anthropic` in your shell cannot silently redirect evals at a paid account. + +Override with `PI_EVAL_PROVIDER` / `PI_EVAL_MODEL` — not `PI_PROVIDER`, which the config overwrites. + +`tests/evals/agent-dir/` is generated at run time by `global-setup.ts` and is entirely gitignored: the Aperture +base URL is private. Set it with `cp .env.test.example .env.test`, or export `APERTURE_BASE_URL` (CI uses a +secret, and a real env var wins over `.env.test`). Aperture is reachable over Tailscale. + +### Only built-in and models.json providers work + +The harness calls `ModelRuntime.create()` with no arguments and resolves the model *before* it builds the session, +so providers that your normal setup registers through extensions do not exist yet. `Eval model not found: +/` means the provider needs a `providers` entry in `tests/evals/agent-dir/models.json`, not an extension. + +Run `tests/evals/smoke.eval.ts` first when you suspect the setup. It is one short round trip with thinking off and is +the cheapest way to separate "the setup is broken" from "the agent behaved badly". Keep it that way: do not split +it into multiple tests or add prompts that invite the model to explain itself. + +## Writing an eval + +Use `createProcessesHarness()` from `tests/evals/harness.ts`. Its `output` exposes a JSON-safe view: + +- `response` — final assistant text. +- `toolCalls` — every `process` call in order, with `arguments`. +- `bashCommands` — every `bash` command string, for catching `&` / `nohup` backgrounding. +- `activeTools` — tool names registered in the session. + +Assert on `result.output`. Judges and assertions receive `output`, not the live `AgentSession`. + +```ts +import { expect } from "vitest"; +import { describeEval } from "vitest-evals"; + +import { callsWithAction, createProcessesHarness } from "./harness.ts"; + +const harness = createProcessesHarness(); + +describeEval("process tool: ", { harness }, (it) => { + it("", async ({ run }) => { + const result = await run(""); + expect(result.output.activeTools).toContain("process"); + expect(callsWithAction(result.output.toolCalls, "start")).toHaveLength(1); + }); +}); +``` + +Multi-turn scenarios pass an array of steps. Use this to seed a bad pattern before testing recovery: + +```ts +const result = await run([ + { type: "prompt", content: "Start ... without a log watch." }, + { type: "prompt", content: "Check its output." }, + { type: "prompt", content: "Check its output again." }, + { type: "prompt", content: "You keep checking by hand. Fix that." }, +]); +``` + +`{ type: "reload" }` steps re-run resource loading, needed when a prompt created or changed Pi resources on disk. + +### Prefer deterministic assertions + +Tool discipline is mechanical: did it call `output` twice, did it pass `notify.logMatches`, did it shell out with +`&`. Assert directly on `toolCalls` / `bashCommands`. Reserve `createJudge(...)` for genuinely fuzzy questions like +"did it explain why it backgrounded the command", and set `judgeThreshold: null` so a low score is an observation +rather than a suite failure. + +### Write prompts that make violations unambiguous + +A model checking output twice under a vague prompt is not misbehaviour. Say "check its output exactly once, then +stop and end your turn" so a second call is unambiguously wrong. If you cannot phrase the prompt so a violation is +clear-cut, the behaviour probably needs a judge, not an assertion. + +## Verifying an eval is trustworthy + +A green eval is worthless if it passed for the wrong reason. Before believing a result: + +1. **Guard against trivial passes.** Every eval asserts `expect(result.output.activeTools).toContain("process")` + and that the expected `start` call happened. Without this, "the agent never polled" also passes when the + extension failed to load and the agent had no `process` tool at all. +2. **Confirm the extension count.** The patched harness throws if the number of loaded extensions does not match + `additionalExtensionPaths`. An error mentioning "Expected the isolated eval session to load exactly N + extension(s)" means the patch or the extension path is wrong, not that the agent misbehaved. +3. **Read the transcript.** Each run writes a native Pi session JSONL under `.eval/sessions/`, indexed by + `.eval/runs.jsonl`. Open the session for a surprising pass or fail and read the actual tool calls. These files + contain prompts, responses, and tool output. +4. **Invert the eval once.** For a new behavioural assertion, temporarily flip the prompt to induce the bad + behaviour (e.g. "keep checking the output until it finishes") and confirm the eval fails. An assertion that has + never failed has not been tested. +5. **Repeat before concluding.** Single runs are noisy. Re-run a few times, or use `evalHarnessTable(...)` with + `repetitions` for comparative work, before claiming a behaviour changed. + +## Comparative evals + +To measure whether the shipped skill or a prompt change helps, build a baseline/candidate table with +`evalHarnessTable(...)` from the vendored `tests/evals/vendor/pi-evals/vitest-evals/harness-table.ts` and Vitest's +`describe.for(...)`. `createProcessesHarnessWithSkill()` exists as the candidate side against +`createProcessesHarness()` as baseline. Record correctness with a judge, set `judgeThreshold: null`, and read the +reporter's pass-rate lift rather than individual pass/fail. + +## Cost and etiquette + +Every eval run costs real tokens against a real provider. The config runs files serially with +`maxConcurrency: 1`. Do not add retries to paper over flaky behaviour — a flaky eval is a finding. Scope runs with +`-t` while iterating instead of running the whole suite.