pi-simple-subagents #
A subagent tool for pi. The calling
agent hands a task to a fresh pi process with its own context window and gets
back only the final answer, so noisy work (broad searches, log digging, verbose
output) stays out of the main conversation.
There are no agent definitions to write first. Each call sets the task, extra instructions, tools, model, and working directory, or picks up defaults from an optional preset. A subagent can also be forked from the calling conversation, so it starts knowing what the caller knows.
Install #
pi install https://tangled.org/did:plc:q2h526frmg4skgi4sst4m74v/
The trailing slash is required. Update it with
pi update https://tangled.org/did:plc:q2h526frmg4skgi4sst4m74v/ and remove it
with pi remove and the same URL.
To install from a local clone, install the clone's dependencies first; pi doesn't install them for a local package:
(cd /path/to/pi-simple-subagents && npm install)
pi install /path/to/pi-simple-subagents
pi loads the clone where it is, without copying it, so the clone is the
installed copy: update it by pulling into the clone (and rerunning
npm install when its dependencies change), and remove it with pi remove
and its path.
Usage #
Ask for delegation in plain language ("have a subagent survey the retry logic"); the agent makes the tool calls. The tool description lists available models and presets, so the agent can pick without looking them up.
The subagent tool #
{
"task": "Find every place retry logic is implemented and summarize how backoff is configured.",
"label": "retry-scan",
"tools": ["read", "grep", "find", "ls"],
"model": "haiku"
}
| Field | Meaning |
|---|---|
task |
The job. A fresh subagent sees none of the calling conversation, so the task has to carry all the context it needs. |
tasks |
Tasks to run in parallel, instead of task, up to the limit. Top-level preset, instructions, model, tools, cwd, reports, and fork act as defaults for each entry. |
label |
Short display name. |
preset |
Named defaults to start from. See Presets. |
instructions |
Extra system prompt (role, constraints, output format). |
tools |
Tool allowlist of names, or patterns where * matches any characters (pi 1.0.4 or later). Omitted: the preset's tools, otherwise pi's default tools minus subagent/subagents. ["none"]: no tools. Add "subagent" by name to allow nesting; patterns never match it. |
model |
Any pattern pi --model accepts, including :thinking suffixes. |
cwd |
Working directory. Default: the session's. |
fork |
Start from a copy of the calling conversation. See Forks. |
reports |
Ids of finished runs whose latest reports are appended to the task word for word. See Passing reports on. |
mode |
blocking (default) waits for the answer; background returns at once and the report arrives as a notification later. |
waitSeconds |
Block for at most this many seconds; runs still going then continue in the background. 0 or omitted: no cap. |
Subagents work directly on the caller's files, not on a copy. Parallel tasks
that edit need separate files, or separate working copies (a git worktree or jj
workspace) passed as cwd.
Only task (or tasks) is required. A blocking call returns the subagent's
final message, followed by the tools it used, the files it changed with file
tools such as write and edit (not bash), its run id, and its transcript
path:
Retries are configured in two places: …
used read ×8, grep ×4
run id: s3
transcript: ~/.pi/agent/subagents/sessions/--home-me-proj--/019ed2ed-…/2026-…_01a0b4a7-….jsonl
Parallel and background reports add a heading per run with its status, run time, tokens, cost, and model:
### [retry-scan] (s3) completed · 2m 14s · 18 turns ↑412k ↓9.1k $0.21 haiku
The answer handed back to the calling agent is capped in size (see Limits). The full text stays in the transcript.
Neither mode nor waitSeconds limits the runs themselves, and neither does
Escape: interrupting the turn stops the call from waiting, and its runs carry
on in the background. Until you send another message, their reports don't start
a turn; after that they report as usual. Stop a run with subagents cancel or
from /subagents.
The subagents tool #
Manages runs after they start:
| Action | Effect |
|---|---|
list |
Every run the session holds, with status, usage, and current activity. |
status |
What the given runs are doing now (task, spend, recent tool calls and reasoning), or their report once finished. Never blocks. |
message |
Send a running subagent an instruction, or resume a finished one with a follow-up. Blocks for the reports unless mode is background. |
wait |
Block until the given runs (default: all in flight) finish and return their reports. waitSeconds caps the wait. |
cancel |
Stop the given runs (default: all in flight) and return any partial output. |
{ "action": "message", "ids": ["s3"], "message": "Check the tests too.", "mode": "background" }
{ "action": "message", "ids": ["s3"], "message": "I fixed findings 1-3; re-check them." }
A message to a running subagent lands at its next turn boundary
("delivery": "followUp" holds it until the subagent would otherwise finish).
It gets no separate reply; the subagent responds in its report.
A message to a finished subagent resumes it from its own session: it keeps everything it already read and answers the follow-up, so a reviewer can re-check fixes without a new brief. For a final, unbiased review, launch a new one.
See docs/background.md for background runs, steering, and resuming in detail.
Passing reports on #
To hand one run's findings to another, such as a reviewer's notes to the agent
fixing them, the calling agent names the reviewer's run in reports instead of
copying the report into the task. Restating a report costs output tokens and
tends to lose detail, such as file:line pointers. reports works on a launch
and on subagents message:
{ "task": "Fix the review findings below. Skip 4; that behaviour is intended.", "reports": ["s3"] }
{ "action": "message", "ids": ["s5"], "message": "Apply this review too.", "reports": ["s3"] }
Each report is appended after the task or message, marked as another subagent's report. Only finished runs qualify, and a run's latest report is the one passed on. A run that is still going, failed, or was cancelled is refused, and nothing starts. If the reporting run worked in another directory, the subagent is told so, since the report's paths point there.
Forks #
"fork": true starts the subagent as a copy of the calling conversation. It
already has every file the caller read and everything it was told, so the task
can be a short directive ("draft tests for the parser changes") instead of a
full brief. It copies the conversation, not the files: it works on the same
files as the caller.
{ "task": "Draft unit tests for the parser changes so far.", "fork": true }
{ "fork": true, "tasks": [{ "task": "Try approach A" }, { "task": "Try approach B" }] }
A fork runs on the caller's model, thinking level, and tools so that it can reuse the caller's prompt cache. Fork when a fresh agent would have to re-read what the caller already read. Otherwise start fresh: a short brief can go to a cheaper model, a review should not share the caller's assumptions, and a fork re-reads its whole inherited context on every turn.
Limits:
- A fork takes no
preset,model,tools,cwd, orinstructions. - A fork can't fork again, though it can start fresh subagents.
- There's nothing to fork in a session started with
--no-session, or one whose context window is nearly full (see Limits). - A fork whose tools differ from the caller's is stopped before its first request.
Set "fork": false in subagent.json to turn forking off.
docs/forks.md explains how a fork keeps the cache.
Supervisor mode #
/supervise turns supervisor mode on or off for the session (/supervise on
and /supervise off set it). While it's on, the agent is prompted to design,
coordinate subagents, review their work, and integrate it, doing only trivial
work itself. Runs it starts, or resumes with subagents message, run in the
background unless the session can't leave runs going (see
docs/background.md), and it collects their reports with
subagents wait when it needs them. The setting is saved in the session, so it
survives resuming and follows /tree. The footer shows supervisor mode while
it's on.
The rules are added to the system prompt as the subagent tool's prompt
guidelines, listed in
extensions/subagent/supervisor.ts. A
custom system prompt (--system-prompt or SYSTEM.md) leaves tool guidelines
out, so the mode needs pi's default one.
Watching runs #
- A call's row in the chat shows the subagent's newest steps, then a preview of its reply once it has one. It keeps updating after the call stops waiting. Ctrl+O expands the row into each exchange: the task or message, the first and last few steps, and the reply in full.
/subagentsopens a live list of this session's runs, with their full transcripts, drawn as pi draws its own session. It also lists runs from before the session was resumed and subagents those runs launched. These are read from their transcripts, so they can be opened but not messaged or cancelled;mon one copies thepi --forkcommand that continues it in a copy. A stored transcript over the size limit (see Limits) isn't shown; read it withpi --export.
| Key | In the list | In a transcript | In the usage view |
|---|---|---|---|
↑/↓, k/j |
Select | Scroll a line | Scroll a line |
space/b, PgDn/PgUp, Ctrl+F/Ctrl+B |
Scroll a page | Scroll a page | |
Ctrl+D/Ctrl+U |
Scroll half a page | Scroll half a page | |
g/G, Home/End |
Top / bottom | Top / bottom | |
⏎ |
Open transcript | ||
v, Ctrl+O |
Collapse or expand long output; opens as the chat shows it | ||
m |
Message selected run, or copy a stored one's pi --fork command |
The same for this run | |
c |
Cancel selected run | Cancel this run | |
x |
Cancel all runs in flight | ||
u |
Usage totals | ||
s |
Switch scope | ||
r |
Refresh | ||
q/esc |
Close | Back to list | Back to list |
Messaging a run yourself #
When you can see what a stuck run is missing, press m and tell it directly.
⏎ sends, tab switches between steering now and following up when it would
otherwise finish, and esc discards the message. The subagent is told the
message is from you, and its next report quotes it to the calling agent.
A message to a finished run resumes it in the background. That's refused until the agent has the run's previous report; send it again then.
To have the agent deal with a stuck run instead, close /subagents and press
Escape in the chat: the call stops waiting, the run keeps going, and the agent
can check on and message it when you ask.
Usage totals #
/subagents usage (or u in the list) totals runs, outcomes, spend, tokens,
and run time, broken down by model and by preset, for this session, this
project, or every project. The totals come from the stored transcripts, so
they don't include what runs still in flight have spent so far; this session's
is shown on a separate line. See
docs/transcripts.md.
Configuration #
Optional, at ~/.pi/agent/subagent.json. It is re-read on every call. A file
that can't be read or isn't a valid JSON object is ignored, with a warning at
session start, and the defaults apply.
{
"model": "haiku",
"thinking": "medium",
"systemPromptMode": "append",
"backgroundNotify": "trigger",
"fork": true
}
| Field | Meaning |
|---|---|
model |
Default model for subagents. |
thinking |
Default thinking level (off, minimal, low, medium, high, xhigh, max). Added to the preset's, this file's, or the session's model when it has no :thinking suffix; a model passed in the call is used as is. |
systemPromptMode |
append (default) adds instructions to pi's system prompt; replace swaps it out. A preset's own setting wins. |
backgroundNotify |
trigger (default): a finished background run wakes an idle session. context: it waits for your next message instead. While the agent is working, reports arrive at its next turn either way. |
fork |
true (default) lets the agent fork its conversation. false refuses forks at once and removes the option from the tool after /reload. |
The model resolves in this order: the call's model, the preset's, this file's,
then the calling session's current model and thinking level.
The tool description lists models for the agent to choose from, up to the
limit, or only the session's enabled models when --models or
enabledModels is set. The list and the presets in it are fixed at session
start; /reload refreshes them.
It also names the session's own model, so the agent can pick a different one
when asked for a second opinion; that line follows a model switch.
Presets #
A preset is a named set of defaults: system prompt, tools, model, and thinking level, written as a Markdown file with YAML frontmatter.
---
name: scout
description: Fast codebase recon that returns compressed context for handoff.
tools: read, grep, find, ls
thinking: low
---
You are a scouting subagent. Move fast, but do not guess. Return the entry
points, key types, data flow, and files likely to change.
Put user presets in ~/.pi/agent/subagents/presets/ and project presets in
<project>/.pi/subagents/presets/. Project presets load only in projects you
have trusted. None ship with the extension. See docs/presets.md.
Transcripts #
Each run is saved as a normal pi session file under
~/.pi/agent/subagents/sessions/<project>/<launching session id>/. Its path is
in the run's report. Read one in /subagents or export it with
pi --export <path> run.html. To carry on its conversation, use
pi --fork <path>, which continues from a copy in the directory you run it
from (the run's own is cwd in its run record). Continuing the file itself
with pi --session <path> leaves the run's spend out of the usage totals.
Nothing is pruned automatically, and a session started with --no-session
saves nothing. See docs/transcripts.md.
Limits #
The values are constants in the source under extensions/subagent/.
| Limit | Value |
|---|---|
| Tasks per parallel call | 8 |
| Subagent processes running at once (the rest queue) | 6 |
| Runs in flight, running or queued (a launch over this is refused) | 16 |
| Finished runs addressable by id (oldest-launched dropped first) | 50 |
| Answer returned to the calling agent, per task and per call | 50 KB |
| Models listed in the tool description | 30 |
| Context window use at which a session can no longer fork | 80% |
| Wait for partial output from cancelled runs (slower ones report later) | 3 s |
Stored transcript size /subagents shows |
8 MB |
Runs listed in /subagents, and levels of nesting followed |
200, 8 |
| Transcripts read for usage totals | 5000 |
Good to know #
- Subagents load AGENTS.md files (global and the
cwdproject's) like any pi session. The task doesn't need to repeat them. - Every subagent is told to lead with its answer, back claims with checkable
pointers (file:line, URLs), and say what it's unsure of. Your
instructionscome after this and can change the output format. - Subagent processes get
PI_SUBAGENT=1in their environment. Other extensions can check it to skip behavior meant only for the top-level session. - Dialogs an extension opens inside a subagent are cancelled at once, since nobody is there to answer them.
- Unknown tool names, and model patterns that can't match any known model or
provider, are rejected before anything starts. Tasks with a different
cwdskip this check, since that project's extensions may add tools and models. - Runs left going are cancelled when the session ends, and
/reloadcancels everything. Run ids carry on from the runs a session already shows, so they stay unique within it after a resume or reload. - In one-shot modes (
--print,--mode json) and inside a subagent, launches andsubagents messageblock until their runs finish, whatevermodeandwaitSecondssay, and interrupting them stops those runs.
Security #
A subagent is a full pi process with your permissions. It runs a model-written
prompt without asking for confirmation, so treat the subagent tool like
bash. Runs keep working after the agent stops waiting on them, so keep
/subagents close at hand.
Terminal escape sequences in tasks, messages, and subagent output are removed
before display. Stored transcripts contain everything a subagent read, ran, and
was told, so protect ~/.pi/agent/subagents/sessions/ as you would pi's own
session directory.