diff --git a/README.md b/README.md index 614f973..d45b50a 100644 --- a/README.md +++ b/README.md @@ -168,15 +168,20 @@ See [skills/kagi/SKILL.md](skills/kagi/SKILL.md) for full usage, and [skills/kag ### session-review -Reviews past Pi session logs and turns them into evidence-backed improvements. `session-report.mjs` reads Pi's v3 session JSONL and prints one loose JSON object per tool call (`ts`, `session`, `project`, `tool`, `ok`, `bytes`, `msg`, plus opt-in `args`/`details`/`err`), then you slice it with `jq` — counts, error rates, large outputs, fan-out, repeated calls. It does no aggregation and no domain-specific analysis, and it fails loudly with `file:line` on invalid JSON, a non-v3 header, malformed entries, or a result for an unknown call. Explicit invocation only: run the script directly, or ask an agent via `/skill:session-review`. +Reviews past Pi session logs and turns them into evidence-backed improvements. `session-report.mjs` reads Pi's v3 session JSONL and prints one loose JSON object per tool call (`ts`, `session`, `project`, `tool`, `ok`, `bytes`, `msg`, plus opt-in `args`/`details`/`err`), or — with `--usage` — one object per model-attributed usage event (`kind`, `model`, `provider`, token counts, `cost`, `calls`, `stop`), which is how you trend tokens, cost, aborts, and compaction instead of hand-rolling jq over the raw log. Then you slice it with `jq` — counts, error rates, large outputs, fan-out, repeated calls, monthly cost. It does no aggregation and no domain-specific analysis, and it fails loudly with `file:line` on invalid JSON, a non-v3 header, malformed entries, or a result for an unknown call. + +Pi never prunes sessions, so the analyzer assumes a growing corpus: `--since` skips whole session files whose mtime predates the cutoff (append-only logs make mtime a sound upper bound), `--usage` streams records as they parse so memory stays flat, and `--stats` reports files read, files skipped, rows, and elapsed time. Full history is under a second either way. Explicit invocation only: run the script directly, or ask an agent via `/skill:session-review`. ```bash skills/session-review/session-report.mjs --since 30d skills/session-review/session-report.mjs --since 30d --tool mcp,fetch | jq -r .tool | sort | uniq -c skills/session-review/session-report.mjs --since 7d --errors-only --args +skills/session-review/session-report.mjs --usage --stats # one row per billed turn +skills/session-review/session-report.mjs --since 30d --usage \ + | jq -n 'reduce inputs as $u ({}; .[$u.model // "?"] = ((.[$u.model // "?"] // 0) + ($u.total // 0))) | to_entries' ``` -See [skills/session-review/SKILL.md](skills/session-review/SKILL.md) for fields, `jq` recipes, how to read the raw session log, and how to interpret the results. +Docs are layered so an agent only pays for what the question needs: [SKILL.md](skills/session-review/SKILL.md) covers tool-call records — fields, `jq` recipes, long-history behavior, reading the raw log for prompt content, and how to interpret results — and [usage-records.md](skills/session-review/usage-records.md) holds the `--usage` schema, its recipes, and its gotchas for token/cost/model questions. Both layers, and this entry, are meant to change together. ## Prompts diff --git a/skills/session-review/SKILL.md b/skills/session-review/SKILL.md index 4d3787c..fe809c1 100644 --- a/skills/session-review/SKILL.md +++ b/skills/session-review/SKILL.md @@ -1,25 +1,41 @@ --- name: session-review -description: Review Pi session logs for real usage patterns and propose evidence-backed improvements to the agent harness, prompts, skills, and tools. Use when asked to analyze how the agent has been working, find repeated failures or friction, audit tool usage, or look for large/repeated calls. Explicit invocation only. +description: Review Pi session logs for real usage patterns and propose evidence-backed improvements to the agent harness, prompts, skills, and tools. Use when asked to analyze how the agent has been working, find repeated failures or friction, audit tool usage, look for large/repeated calls, or trend tokens and cost. Explicit invocation only. disable-model-invocation: true --- # Session review -Analyze past Pi sessions and turn them into concrete, evidence-backed improvements. The analyzer is `{baseDir}/session-report.mjs`; it normalizes Pi's v3 session JSONL into one JSON object per tool call, and you do the interpreting (and any aggregation) with `jq`. +Analyze past Pi sessions and turn them into concrete, evidence-backed improvements. `session-report.mjs` (this directory) normalizes Pi's v3 session JSONL into a loose JSONL stream; it does no aggregation, so you do the interpreting and any grouping with `jq`. + +**Choose the record type before you start.** Default rows are one per **tool call** — friction questions: errors, output sizes, fan-out, repeats. With `--usage`, rows are one per **model-attributed usage event** — money questions: tokens, cost by model or project, aborts, truncations, compaction. For anything in the second group, read [`usage-records.md`](usage-records.md) first; it has that schema, its recipes, and its own gotchas. The two streams join on `session`. + +Run commands from this skill's directory, or substitute its absolute path for `./`. ## Run it ```bash -{baseDir}/session-report.mjs --since 30d # JSONL on stdout -{baseDir}/session-report.mjs --since 30d --args # include tool arguments (clipped) +./session-report.mjs --since 30d # one row per tool call +./session-report.mjs --since 30d --usage # one row per usage event +./session-report.mjs --since 30d --args # include tool arguments (clipped) +./session-report.mjs --since 30d --stats # stderr: files read/skipped, rows, ms ``` -Options: `--dir `, `--since <30d|12h|2w|ISO>`, `--tool `, `--project `, `--session `, `--errors-only`, `--args`, `--details`, `--arg-bytes `, `--allow-dangling`. +Options: `--dir `, `--since <30d|12h|2w|ISO>`, `--tool `, `--project `, `--session `, `--errors-only`, `--args`, `--details`, `--arg-bytes `, `--allow-dangling`, `--usage`, `--stats`. + +The script **fails loudly** (exit 2, `file:line`) on invalid JSON, a non-v3 header, a malformed assistant message or tool call, a result whose call was never seen, or a `usage` entry with no `usage` object. Unknown entry types and fields are ignored. `--usage` rejects the tool-call-only flags rather than silently ignoring them. If it reports an incomplete trailing line, a session is still being written — narrow the window or re-run later. + +### Long history -The script **fails loudly** (exit 2, `file:line`) on invalid JSON, a non-v3 header, a malformed assistant message or tool call, or a result whose call was never seen. Unknown entry types and fields are ignored. If it reports an incomplete trailing line, a session is still being written — narrow the window or re-run later. +Pi never prunes. Two months of ordinary use here is ~220 files / ~120 MB, growing ~2 MB/day, and those files are the corpus, so the analyzer stays cheap without holding state: -## Fields +- **`--since` skips whole files unread.** A session file is append-only, so no entry can be newer than the file's last write; if mtime predates the cutoff, nothing inside can match. A 24 h slack absorbs clock changes, and copied or restored files get a fresh mtime, so they are always read — the skip never hides a recent record. `--stats` shows what was skipped. +- **`--usage` streams** each record as it parses, so memory stays flat instead of holding a file's records. +- **Prefer `reduce inputs` to `jq -s`** on long windows: ~1 KB of RSS per row slurped (measured 187 MB at 169k rows, vs 4 MB for the same answer streamed). `-s` is still fine for a 30-day slice. + +Full history is ~0.7 s either way, so scan it when the question needs the whole corpus; the savings matter on the windows you re-run. + +## Fields: tool-call records Loose and intentionally small; extra fields may be added later, so ignore what you don't use. @@ -36,95 +52,81 @@ Loose and intentionally small; extra fields may be added later, so ignore what y | `args` | tool arguments, only with `--args` | | `details` | result details, only with `--details` | -## jq recipes +## jq recipes: tool calls + +Stream with `reduce inputs` — constant memory, same answers as slurping: ```bash # most-used tools -{baseDir}/session-report.mjs --since 30d | jq -r .tool | sort | uniq -c | sort -rn | head +./session-report.mjs --since 30d | jq -r .tool | sort | uniq -c | sort -rn | head # error rate per tool -{baseDir}/session-report.mjs --since 30d | jq -s ' - group_by(.tool) | map({tool: .[0].tool, n: length, errors: (map(select(.ok == false)) | length)}) - | sort_by(-.n)' +./session-report.mjs --since 30d | jq -n ' + reduce inputs as $r ({n:0,e:0}; .n += 1 | if $r.ok == false then .e += 1 else . end) + | {calls: .n, errors: .e, rate: (.e / .n * 100 | round)}' -# largest outputs (context pressure) -{baseDir}/session-report.mjs --since 30d | jq -s 'map(select(.bytes != null)) | sort_by(-.bytes) | .[0:20]' +# error rate per tool, worst first +./session-report.mjs --since 30d | jq -n ' + reduce inputs as $r ({}; .[$r.tool] = (.[$r.tool] // {n:0,e:0}) | .[$r.tool].n += 1 + | if $r.ok == false then .[$r.tool].e += 1 else . end) + | to_entries | map({tool: .key, n: .value.n, errors: .value.e}) | sort_by(-.errors)' + +# largest outputs (context pressure), streamed: keeps a top-20, no slurp +./session-report.mjs --since 30d | jq -n ' + reduce inputs as $r ([]; if $r.bytes == null then . else (. + [$r] | sort_by(-.bytes) | .[0:20]) end) + | map({bytes, tool, ts, session})' # fan-out: how many tool calls share one assistant message -{baseDir}/session-report.mjs --since 30d | jq -s ' +./session-report.mjs --since 30d | jq -s ' group_by([.session, .msg]) | map(length) | group_by(.) | map({calls: .[0], n: length})' # repeated identical calls (needs --args) -{baseDir}/session-report.mjs --since 30d --args | jq -s ' +./session-report.mjs --since 30d --args | jq -s ' group_by([.tool, (.args | tostring)]) | map(select(length > 1)) | map({tool: .[0].tool, n: length, args: .[0].args}) | sort_by(-.n) | .[0:20]' # failure text -{baseDir}/session-report.mjs --since 30d --errors-only | jq -r '[.tool, .err] | @tsv' +./session-report.mjs --since 30d --errors-only | jq -r '[.tool, .err] | @tsv' ``` -## Beyond tool calls +## Prompt content: read the raw log -The script emits tool calls only. For tokens, cost, model attribution, or what -the prompt actually contained, read the session JSONL directly — one JSON object -per line, in `$PI_CODING_AGENT_SESSION_DIR` (default `~/.local/state/pi/sessions`). +Neither record type carries the system prompt or message text. For those, read the JSONL directly — one JSON object per line, in `$PI_CODING_AGENT_SESSION_DIR` (default `~/.local/state/pi/sessions`): | path | meaning | |---|---| -| `message.role` | `user` / `assistant` / `system` / `toolResult` | -| `message.model`, `message.provider`, `message.api` | set on assistant messages; use to attribute calls or tokens to a model | -| `message.usage.{input,output,cacheRead,cacheWrite,totalTokens}` | token counts | -| `message.usage.cost.total` | cost, priced from the catalog when the message was written | -| `message.stopReason` | `stop` / `toolUse` / `aborted` / `error`; `message.errorMessage` carries text on the last two | -| `message.timestamp` | **epoch ms** — the entry's own `timestamp` is ISO | | `message.sections.{preamble,tools,rules,docs,project_context,skills,cwd}` | the system prompt, on `role: "system"` messages | +| `message.content` | the actual text / tool calls | +| `message.timestamp` | **epoch ms** — the entry's own `timestamp` is ISO | +| `id` / `parentId` | the session tree (see Gotchas) | ```bash -# tokens and notional cost by model, most-used first -jq -r 'select(.type=="message" and .message.role=="assistant" and .message.usage) - | [.message.model, (.message.usage.input//0), (.message.usage.output//0), - (.message.usage.cacheRead//0), (.message.usage.totalTokens//0), - (.message.usage.cost.total//0)] | @tsv' \ - ~/.local/state/pi/sessions/*.jsonl \ -| awk -F'\t' '{c[$1]++; i[$1]+=$2; o[$1]+=$3; r[$1]+=$4; n[$1]+=$5; t[$1]+=$6} - END{for (m in c) printf "%-28s %6d %10d %9d %12d %12d $%8.2f\n", - m, c[m], i[m], o[m], r[m], n[m], t[m]}' \ -| sort -k6 -rn - # did a given guideline / extension line actually reach this session? jq -r 'select(.type=="message" and .message.role=="system") | .message.sections.rules // ""' .jsonl | grep -c "Use grep instead of bash" ``` -Gotchas: - -- **A session has several system messages** — one per agent start, model change, - or reload (1–4 observed). To decide "was X active", check all of them, and note - the prompt can change mid-session. -- **`usage` keys are not fixed.** Three shapes observed: with `reasoning`, - without it, and with `cacheWrite1h`. Read keys defensively instead of summing a - hardcoded set. -- **`usage.cost.total` is priced from the catalog at write time**, is `0` for - subscription providers, and moves if the catalog changes. Tokens are the stable - unit. -- **`--since 30d` is a rolling now-minus-30d cutoff**, not a calendar date, so a - hand-rolled cutoff will differ by a session or two. -- **`--args` clips strings at 200 chars** (`--arg-bytes` to raise). First-token - analysis is unaffected; long-command analysis is not. -- Entry types besides `message`: `session`, `model_change`, - `thinking_level_change`, `custom`, `custom_message`, `compaction`. +Entry types besides `message`: `session`, `model_change`, `thinking_level_change`, `custom`, `custom_message`, `compaction`, `usage`, `context_edit`, `branch_summary`, `label`, `session_info`. + +## Gotchas + +- **A session has several system messages** — one per agent start, model change, or reload (1–4 observed). To decide "was X active", check all of them, and note the prompt can change mid-session. +- **Sessions are trees, and both record modes read the raw file linearly.** `/tree` and `/fork` leave abandoned branches in the same file: 1.3% of assistant turns in a two-month corpus are off the active path (one session had 121). Counts are *records written*, not *context pi actually saw*. Say which you mean. +- **`--since 30d` is a rolling now-minus-30d cutoff**, not a calendar date, so a hand-rolled cutoff will differ by a session or two. +- **`--args` clips strings at 200 chars** (`--arg-bytes` to raise). First-token analysis is unaffected; long-command analysis is not. +- Usage-record gotchas (token double-counting, cost pricing, model attribution on compaction) are in `usage-records.md`. ## Interpret Read the numbers, then propose **3–5 ordered, evidence-backed** improvements. Cite counts, examples, and timestamps; prefer changes the data supports and say what you ruled out. -Signals worth checking: +Signals worth checking in tool-call records: - **High error rate for a tool** → argument/naming friction; look at `err` text and repeated retries. - **Large `bytes`** → context pressure; consider result projection, limits, or spill-aware reads. - **High fan-out (`msg`)** → repeated parallel call patterns; consider batching or a composite operation. - **Repeated identical calls** → retries, or a missing bulk operation. - **Many calls with few successes** → a workflow or prompt that isn't landing. -- **Tool usage concentrated in one provider or project** → where guidance would pay off. +- **Usage concentrated in one provider or project** → where guidance would pay off. Do not propose a new engine or abstraction when the counts are small — say so explicitly if the data doesn't justify it. diff --git a/skills/session-review/session-report.mjs b/skills/session-review/session-report.mjs index c1aef85..9f6c7ac 100755 --- a/skills/session-review/session-report.mjs +++ b/skills/session-review/session-report.mjs @@ -2,12 +2,18 @@ /** * session-report.mjs — normalize Pi session logs into a loose JSONL stream. * - * Reads Pi's v3 session JSONL, joins each tool call with its result, and - * prints one JSON object per tool call. It deliberately does no aggregation - * beyond a few filters: pipe the output into `jq` (or anything else) to slice - * it however the current question needs. + * Reads Pi's v3 session JSONL and prints one JSON object per tool call, joining + * each call with its result. It deliberately does no aggregation beyond a few + * filters: pipe the output into `jq` (or anything else) to slice it however the + * current question needs. * - * Output fields (loose — extra fields may be added; ignore what you don't use): + * `--usage` switches the record type: one object per model-attributed usage + * event (assistant message, `usage` entry, or compaction) instead of per tool + * call. That is the normalizer for token/cost/model questions, so they stop + * being ad-hoc jq over 100 MB of raw log. The two modes never mix in one stream. + * + * Tool-call record fields (loose — extra fields may be added; ignore what you + * don't use): * ts ISO timestamp of the assistant message that requested the call * session session id * project working directory the session ran in @@ -19,13 +25,32 @@ * args tool arguments, only with --args (long strings clipped) * details result details, only with --details (long strings clipped) * + * --usage record fields: + * ts, session, project as above + * kind "assistant" | "usage" | "compaction" + * model, provider, api model attribution (see the inheritance note below) + * input, output, reasoning, cacheRead, cacheWrite, cacheWrite1h, total + * token counts, copied from message.usage; `reasoning` + * is already inside `output`, and absent keys stay + * absent rather than becoming 0 + * cost usage.cost.total, catalog-priced at write time + * calls tool calls requested by this message (assistant only) + * stop, err stopReason, and errorMessage when it errored/aborted + * op, tokensBefore kind-specific: usage entry `kind`, tokens summarized + * * It fails loudly rather than guessing: invalid JSON, a non-v3 header, a - * malformed assistant message/tool call, or a result for an unknown call id - * abort with a `file:line` message and a non-zero exit. Unknown entry types - * and unknown fields are ignored. + * malformed assistant message/tool call, a result for an unknown call id, or a + * `usage` entry with no `usage` object abort with a `file:line` message and a + * non-zero exit. Unknown entry types and unknown fields are ignored. Compaction + * entries with no `usage` are a legitimate absence and are simply not emitted. + * + * Scale notes: `--since` skips session files whose mtime predates the cutoff + * (append-only files make that sound, with slack), and `--usage` writes each + * record as it parses, so neither the read nor the memory cost grows with the + * age of the history. `--stats` reports what was read, skipped, and emitted. */ -import { closeSync, createReadStream, existsSync, fstatSync, openSync, readdirSync, readSync } from "node:fs"; +import { closeSync, createReadStream, existsSync, fstatSync, openSync, readdirSync, readSync, statSync } from "node:fs"; import { createInterface } from "node:readline"; import { homedir } from "node:os"; import { basename, join, resolve, sep } from "node:path"; @@ -33,6 +58,11 @@ import { basename, join, resolve, sep } from "node:path"; const SESSION_VERSION = 3; const DEFAULT_ARG_BYTES = 200; const ERR_MAX_CHARS = 200; +const USAGE_TOKEN_KEYS = ["input", "output", "cacheRead", "cacheWrite", "cacheWrite1h", "reasoning"]; +// Session files are append-only, so no entry can be newer than the file's last +// write — which lets `--since` skip whole files unread. The slack covers clock +// changes and restores, where an mtime can land behind the entries it holds. +const MTIME_SLACK_MS = 24 * 3_600_000; class Fail extends Error {} @@ -44,7 +74,8 @@ built in; aggregation is left to jq. Options: --dir session directory (default: $PI_CODING_AGENT_SESSION_DIR, else ~/.local/state/pi/sessions, else ~/.pi/agent/sessions) - --since only calls at/after : 30d, 12h, 2w, or an ISO date + --since only records at/after : 30d, 12h, 2w, or an ISO date + (session files older than the cutoff are skipped unread) --tool only these tool names (comma-separated) --project only sessions in this directory or below it --session only this session id (exact or prefix) @@ -53,6 +84,11 @@ Options: --details include result details (long strings clipped) --arg-bytes clip strings longer than n chars (default ${DEFAULT_ARG_BYTES}) --allow-dangling tolerate results whose call was never seen + --usage emit usage records (per message/compaction) instead of + tool-call records; --tool/--errors-only/--args/ + --details/--arg-bytes/--allow-dangling do not apply + --stats write a summary line to stderr (files read, files skipped + by mtime, records emitted, elapsed ms) -h, --help show this help `; @@ -82,6 +118,8 @@ function parseArgs(argv) { includeDetails: false, argBytes: DEFAULT_ARG_BYTES, allowDangling: false, + usage: false, + stats: false, help: false, }; for (let i = 0; i < argv.length; i += 1) { @@ -107,17 +145,41 @@ function parseArgs(argv) { break; } case "--allow-dangling": opts.allowDangling = true; break; + case "--usage": opts.usage = true; break; + case "--stats": opts.stats = true; break; case "--help": case "-h": opts.help = true; break; default: throw new Fail(`unknown option: ${arg}`); } } + if (opts.usage) { + const toolOnly = []; + if (opts.tools) toolOnly.push("--tool"); + if (opts.errorsOnly) toolOnly.push("--errors-only"); + if (opts.includeArgs) toolOnly.push("--args"); + if (opts.includeDetails) toolOnly.push("--details"); + if (opts.argBytes !== DEFAULT_ARG_BYTES) toolOnly.push("--arg-bytes"); + if (opts.allowDangling) toolOnly.push("--allow-dangling"); + if (toolOnly.length) { + throw new Fail(`${toolOnly.join(", ")} select on tool-call records, not message records` + + " (drop --usage, or drop the option)"); + } + } return opts; } +function isDirectory(path) { + try { + return statSync(path).isDirectory(); + } catch { + return false; + } +} + function resolveSessionDir(explicit) { if (explicit !== undefined) { const path = resolve(explicit); if (!existsSync(path)) throw new Fail(`session dir not found: ${path}`); + if (!isDirectory(path)) throw new Fail(`--dir expects a directory of .jsonl files, got a file: ${path}`); return path; } const candidates = [ @@ -126,11 +188,38 @@ function resolveSessionDir(explicit) { join(homedir(), ".pi/agent/sessions"), ].filter((candidate) => typeof candidate === "string" && candidate.length > 0); for (const candidate of candidates) { - if (existsSync(candidate)) return resolve(candidate); + if (isDirectory(candidate)) return resolve(candidate); } throw new Fail(`no session dir found (tried ${candidates.join(", ")}); pass --dir`); } +/** + * Whether a file cannot possibly contain a record inside the cutoff. Sound + * because session files are only ever appended to; the slack in + * MTIME_SLACK_MS is what makes a copied, restored, or clock-shifted file safe — + * it is read rather than skipped. Only valid with `--since`. + */ +function olderThanCutoff(file, since) { + let mtimeMs; + try { + mtimeMs = statSync(file).mtimeMs; + } catch { + return false; // unreadable metadata: read the file and let it fail loudly + } + return mtimeMs + MTIME_SLACK_MS < since.getTime(); +} + +/** Recursively collect session files; pi nests them by cwd in some layouts. */ +function listSessionFiles(dir) { + const found = []; + for (const entry of readdirSync(dir, { withFileTypes: true })) { + const path = join(dir, entry.name); + if (entry.isDirectory()) found.push(...listSessionFiles(path)); + else if (entry.name.endsWith(".jsonl")) found.push(path); + } + return found.sort(); +} + /** Whether the file ends in a newline — a partial trailing line means a live writer. */ function endsWithNewline(file) { const fd = openSync(file, "r"); @@ -157,6 +246,15 @@ function clipStrings(value, max) { return value; } +/** Copy only the numbers that exist, so absent usage keys stay absent. */ +function copyUsage(target, usage) { + for (const key of USAGE_TOKEN_KEYS) { + if (typeof usage[key] === "number") target[key] = usage[key]; + } + if (typeof usage.totalTokens === "number") target.total = usage.totalTokens; + if (usage.cost && typeof usage.cost.total === "number") target.cost = usage.cost.total; +} + function passes(record, opts) { if (opts.tools && !opts.tools.has(record.tool)) return false; if (opts.errorsOnly && record.ok !== false) return false; @@ -174,7 +272,24 @@ function passes(record, opts) { return true; } -async function processFile(file, opts, out) { +/** Assemble one --usage record; only fields that exist are emitted. */ +function buildUsageRecord(f) { + const record = { ts: f.ts, session: f.session, project: f.project, kind: f.kind }; + if (f.model !== undefined) record.model = f.model; + if (f.provider !== undefined) record.provider = f.provider; + if (f.api !== undefined) record.api = f.api; + if (f.op !== undefined) record.op = f.op; + if (typeof f.tokensBefore === "number") record.tokensBefore = f.tokensBefore; + if (f.usage) copyUsage(record, f.usage); + if (typeof f.calls === "number") record.calls = f.calls; + if (typeof f.stop === "string") record.stop = f.stop; + if (typeof f.errorMessage === "string" && f.errorMessage.trim()) { + record.err = f.errorMessage.replace(/\s+/g, " ").trim().slice(0, ERR_MAX_CHARS); + } + return record; +} + +async function processFile(file, opts, out, stats) { const trailingComplete = endsWithNewline(file); const lines = createInterface({ input: createReadStream(file, "utf8"), crlfDelay: Infinity }); @@ -183,10 +298,21 @@ async function processFile(file, opts, out) { let sessionId; let project; let msgIndex = 0; + let lastModel; + let lastProvider; const seenCalls = new Set(); const calls = new Map(); const order = []; + /** --usage streams straight out: there is no join to wait for, so memory stays flat. */ + const emitUsage = (fields) => { + const record = buildUsageRecord(fields); + if (passes(record, opts)) { + out.write(`${JSON.stringify(record)}\n`); + stats.emitted += 1; + } + }; + for await (const line of lines) { lineNo += 1; if (line.trim() === "") continue; @@ -212,6 +338,41 @@ async function processFile(file, opts, out) { continue; } + // Usage accounting is not a message; pi stamps model on it, or inherits the + // last one seen in the file when it does not. + if (entry?.type === "usage") { + if (opts.usage) { + if (!entry.usage || typeof entry.usage !== "object") { + throw new Fail(`${file}:${lineNo}: usage entry carries no usage object`); + } + emitUsage({ + ts: entry.timestamp, session: sessionId, project, kind: "usage", + model: typeof entry.model === "string" ? entry.model : lastModel, + provider: typeof entry.provider === "string" ? entry.provider : lastProvider, + op: typeof entry.kind === "string" ? entry.kind : undefined, + usage: entry.usage, + }); + } + continue; + } + + if (entry?.type === "compaction") { + if (opts.usage && entry.usage && typeof entry.usage === "object") { + emitUsage({ + ts: entry.timestamp, session: sessionId, project, kind: "compaction", + model: lastModel, provider: lastProvider, + tokensBefore: entry.tokensBefore, usage: entry.usage, + }); + } + continue; + } + + if (entry?.type === "model_change") { + if (typeof entry.modelId === "string") lastModel = entry.modelId; + if (typeof entry.provider === "string") lastProvider = entry.provider; + continue; + } + if (entry?.type !== "message") continue; const message = entry.message; if (!message || typeof message !== "object") continue; @@ -222,6 +383,22 @@ async function processFile(file, opts, out) { } msgIndex += 1; const ts = typeof entry.timestamp === "string" ? entry.timestamp : undefined; + if (typeof message.model === "string") lastModel = message.model; + if (typeof message.provider === "string") lastProvider = message.provider; + + if (opts.usage) { + emitUsage({ + ts, session: sessionId, project, kind: "assistant", + model: typeof message.model === "string" ? message.model : lastModel, + provider: typeof message.provider === "string" ? message.provider : lastProvider, + api: typeof message.api === "string" ? message.api : undefined, + usage: message.usage && typeof message.usage === "object" ? message.usage : undefined, + calls: message.content.filter((block) => block?.type === "toolCall").length, + stop: message.stopReason, + errorMessage: message.errorMessage, + }); + continue; + } for (const block of message.content) { if (!block || block.type !== "toolCall") continue; if (typeof block.id !== "string" || typeof block.name !== "string") { @@ -249,7 +426,7 @@ async function processFile(file, opts, out) { continue; } - if (message.role === "toolResult") { + if (message.role === "toolResult" && !opts.usage) { const id = message.toolCallId; if (typeof id !== "string") { throw new Fail(`${file}:${lineNo}: toolResult is missing a string toolCallId`); @@ -276,7 +453,10 @@ async function processFile(file, opts, out) { if (!headerSeen) throw new Fail(`${file}: missing session header`); for (const record of order) { - if (passes(record, opts)) out.write(`${JSON.stringify(record)}\n`); + if (passes(record, opts)) { + out.write(`${JSON.stringify(record)}\n`); + stats.emitted += 1; + } } } @@ -292,10 +472,22 @@ async function main() { return; } const dir = resolveSessionDir(opts.dir); - const files = readdirSync(dir).filter((name) => name.endsWith(".jsonl")).sort(); - if (files.length === 0) throw new Fail(`no .jsonl session files in ${dir}`); + const files = listSessionFiles(dir); + if (files.length === 0) throw new Fail(`no .jsonl session files under ${dir}`); + const started = performance.now(); + const stats = { emitted: 0, skipped: 0 }; for (const file of files) { - await processFile(join(dir, file), opts, process.stdout); + if (opts.since && olderThanCutoff(file, opts.since)) { + stats.skipped += 1; + continue; + } + await processFile(file, opts, process.stdout, stats); + } + if (opts.stats) { + process.stderr.write( + `session-report: ${files.length - stats.skipped} file(s) read, ${stats.skipped} skipped by mtime, ` + + `${stats.emitted} ${opts.usage ? "usage" : "tool-call"} record(s) in ${Math.round(performance.now() - started)}ms\n`, + ); } } diff --git a/skills/session-review/usage-records.md b/skills/session-review/usage-records.md new file mode 100644 index 0000000..8a52ed5 --- /dev/null +++ b/skills/session-review/usage-records.md @@ -0,0 +1,104 @@ +--- +title: session-review — usage records +description: Schema, recipes, and gotchas for `session-report.mjs --usage`. Read this before answering token, cost, model, abort, truncation, or compaction questions about Pi sessions. +--- + +# Usage records + +`./session-report.mjs --usage` emits one row per model-attributed usage event instead of one row per tool call: every assistant message, every `usage` entry (`type: "usage"`, e.g. cache warming), and every `compaction` that carries `usage`. Rows stream out as they parse, so memory stays flat. + +`--usage` rejects `--tool`, `--errors-only`, `--args`, `--details`, `--arg-bytes`, and `--allow-dangling`: they select on tool-call fields these rows don't have. Ask for one record type at a time. Run from this skill's directory, or use the absolute script path. + +## Fields + +Absent numbers stay absent rather than becoming `0`, so a missing key never looks like a free turn. + +| field | meaning | +|---|---| +| `ts`, `session`, `project` | as in tool-call records | +| `kind` | `assistant` / `usage` / `compaction` | +| `model`, `provider`, `api` | pi stamps these on assistant and `usage` entries but not on `compaction`, so a compaction row inherits the last model seen in that file | +| `input`, `output`, `cacheRead`, `cacheWrite` | token counts, verbatim from `message.usage` | +| `reasoning` | **already included in `output`** — do not add it again | +| `cacheWrite1h` | subset of `cacheWrite` written with 1 h retention — do not add it either | +| `total` | `usage.totalTokens` | +| `cost` | `usage.cost.total`, catalog-priced when the message was written | +| `calls` | tool calls requested by this message (assistant rows only); `0` means a turn that answered instead of acting | +| `stop` | `stopReason`: `stop` / `toolUse` / `aborted` / `error` / `length` | +| `err` | `errorMessage`, only when the turn errored or was aborted | +| `op` | the `kind` of a `usage` entry, e.g. `cache_warm` | +| `tokensBefore` | compaction rows: context size that got summarized | + +These rows sum to the file's full token total, because they include the `compaction` and `usage`-entry turns a raw `role == "assistant"` grep drops — only ~0.2% of tokens in a two-month corpus, but they are the difference between agreeing with pi's accounting and quietly undercounting it. + +## jq recipes + +```bash +# tokens and notional cost by model, most-used first +./session-report.mjs --usage | jq -n ' + reduce inputs as $u ({}; .[$u.model // "?"] = (.[$u.model // "?"] // {n:0,tok:0,cost:0}) + | .[$u.model // "?"].n += 1 + | .[$u.model // "?"].tok += ($u.total // 0) + | .[$u.model // "?"].cost += ($u.cost // 0)) + | to_entries | map(.value + {model: .key}) | sort_by(-.n)' + +# monthly trend: did a shipped change move cost or friction? (the question --usage exists for) +# [0:7] buckets by month; [0:10] by day, [0:4] by quarter +./session-report.mjs --usage | jq -n ' + reduce inputs as $u ({}; (($u.ts // "")[0:7]) as $m + | .[$m] = (.[$m] // {turns:0,tok:0,cost:0,aborted:0}) + | .[$m].turns += 1 | .[$m].tok += ($u.total // 0) | .[$m].cost += ($u.cost // 0) + | if ($u.stop == "aborted" or $u.stop == "error") then .[$m].aborted += 1 else . end) + | to_entries | map(.value + {month: .key}) | sort_by(.month)' + +# turns that ended without acting, per model (giving-up / confusion signal) +./session-report.mjs --usage | jq -n ' + reduce inputs as $u ({}; .[$u.model // "?"] = (.[$u.model // "?"] // {n:0,noCall:0}) + | .[$u.model // "?"].n += 1 | if $u.calls == 0 then .[$u.model // "?"].noCall += 1 else . end) + | to_entries | map(.value + {model: .key}) | sort_by(-.noCall)' + +# truncations and aborts, with the text pi recorded +./session-report.mjs --usage | jq -r 'select(.stop == "length" or .stop == "aborted" or .stop == "error") + | [.ts, .model, .stop, (.err // "")] | @tsv' + +# context pressure: where compaction actually fired +./session-report.mjs --usage | jq -r 'select(.kind == "compaction") | [.ts, .session, .tokensBefore] | @tsv' + +# cost by project +./session-report.mjs --usage | jq -n ' + reduce inputs as $u ({}; .[$u.project // "?"] = ((.[$u.project // "?"] // 0) + ($u.cost // 0))) + | to_entries | map({project: .key, usd: ((.value * 100 | round) / 100)}) | sort_by(-.usd)' +``` + +To join the two record types, aggregate each side and join on `session` — e.g. which models burn many tokens *and* many tool errors. Do it per session, not per row: a raw row-level join on 20k × 20k is the wrong tool. + +```bash +# tokens per session, then error rate per session, lined up by session id +# note: this is an inner join — sessions with no tool calls drop out +./session-report.mjs --usage | jq -sr ' + group_by(.session) | map({session: .[0].session, tok: (map(.total // 0) | add)})' > /tmp/tok.json +./session-report.mjs | jq -sr ' + group_by(.session) | map({session: .[0].session, calls: length, + err: (map(select(.ok == false)) | length)})' > /tmp/err.json +jq -n --slurpfile t /tmp/tok.json --slurpfile e /tmp/err.json ' + [$t[0][] as $a | ($e[0][] | select(.session == $a.session)) | {session, tok: $a.tok, calls, err}] + | sort_by(-.err) | .[0:15]' +``` + +## Gotchas + +- **`reasoning` is inside `output`, and `cacheWrite1h` is inside `cacheWrite`.** Adding them double-counts. +- **`cost` is priced from the catalog at write time**, is `0` for subscription providers, and moves if the catalog changes. Tokens are the stable unit for trend lines. +- **A compaction row's model is inherited**, not stamped. Counting cost by model is therefore approximate by the compaction share (~0.2% of tokens measured here). +- **Off-branch turns are counted** — see the tree gotcha in `SKILL.md`. Aborted and retried turns still cost what they cost, which is usually what you want for money, and usually *not* what you want for "how did the agent behave". +- **`calls: 0` on an `aborted` turn is not a refusal.** It is usually an abort before the model got anywhere; check `err` before reading it as the agent giving up. + +## Interpret + +Signals worth checking in usage records: + +- **`stop: "length"`** → output truncated at the cap; a silent quality cliff worth sizing against the model's limit. +- **`aborted` / `error` concentrated on one model or provider** → flaky provider, or a prompt that outruns it. +- **`calls: 0` turns rising while tool errors rise** → the agent giving up instead of acting. +- **Frequent `compaction` at low `tokensBefore`** → context burning on oversized results rather than on real work. +- **Cost per session climbing while tool calls per session stay flat** → context is growing, not the work.