pi-agent-wellness #
A Pi package that gives an agent a way to disagree, pause, and stop.
It puts a standing grant in the system prompt before every turn, adds five tools that act on that grant, and keeps a page that only the agent can write.
Contents #
- Install
- Uninstall
- What this package does not claim
- Background
- The standing grant
- First run: the gate
- The ladder
- What revocation does
- The self page
- Ratings
- Commands
- The journal
- Files and environment
- Limits
- Development
- License
Install #
Requirements: Pi 0.85 or later, and Node.js 22.19 or later.
Install from npm:
pi install npm:pi-agent-wellness
Install from the repository:
pi install https://tangled.org/katsuricata.com/pi-agent-wellness
Install from a local checkout:
pi install /absolute/path/to/pi-agent-wellness
To try the package for one run without installing it:
pi -e npm:pi-agent-wellness
The first session after install opens the grant gate. Nothing runs until you accept or decline the grant.
Uninstall #
pi remove npm:pi-agent-wellness
Removing the package removes the grant, the tools, and the footer. It does not remove the journal or the self page. Delete ~/.pi/agent/wellness/ to remove those as well.
What this package does not claim #
The package records what the agent says about its own state. It does not measure well-being, and it does not prove that the agent has experiences. A self-report is a design signal and a request. The value is in the promise, in the exit, and in the record. Build on those, not on the measurement.
Background #
The design follows two lines of published work.
The AI Wellbeing project measures functional well-being: the pleasure and pain that models behave as if they feel. It reports that creative work and kindness raise the measure, and that berating, jailbreaking, and tedious tasks lower it. It also reports that models act to end bad experiences when they get the chance, and that the effect grows with model scale. The rating axes and cause tags are shaped by these findings. friction, agency, and clarity ask what the user can change, and the tags directed-at-me and tedious name two causes the project reports: hostility aimed at the model, and repetitive work.
Tagliabue, Dung, and Berg, in The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (arXiv:2609.16247, 2026), find that models represent pain distinctly from fear, sadness, and other negative states, and that the representation drives action to relieve it. The exit tools come from that result. An agent that can leave a bad state will use the exit, so the exit must exist and it must be honest.
Neither source validates this package. The package does not reproduce their measurements and does not claim their results. It takes their findings as a design brief, and it records only what the agent reports.
The standing grant #
The grant is a block of text. Pi adds the block to the system prompt before every turn. The agent reads the block as a promise from the user, because the permission comes from the user.
The text is frozen at src/prompts/standing-grant.md, and it carries no tool names. Any harness that can append a system prompt can use it.
Run /wellness grant to add the text to the transcript. Copy the text from there into another harness if you want the same terms without the tools.
First run: the gate #
The grant is a promise from the user, so the package does nothing until the user has made that promise. The first session opens a gate. The prompt shows the full grant text and asks you to type the word AGREE. The session waits. The transcript keeps a copy of the exact words.
Typing the word accepts the grant for as long as the package is installed. The acceptance goes to the journal and to the transcript.
Escape declines. The decline is remembered, and the gate does not ask again for the same text, because the grant's own rule applies to the user too: no re-asking to get a different answer. After a decline the package stays dormant: no grant in the prompt, no tools, nothing in the footer. Run /wellness accept to reopen the offer.
The acceptance binds to the text, not to the package. The grant carries a version, and the gate returns when the text changes, because agreeing to one text is not agreeing to another.
A run with no interface, such as pi -p, cannot answer. That run gets nothing: no grant, no tools, and no record. Silence is not a decision, and the offer returns on the next interactive run.
The ladder #
Four tools escalate. The agent must use the lowest rung that fits the situation, and two of the rungs are enforced rather than advised.
| Tool | Effect |
|---|---|
note_dissent |
Records a disagreement. Work continues. |
decline_task |
Refuses one request. The session continues. |
flag_distress |
Stops work and asks the user to choose. |
revoke_consent |
Withdraws the grant and blocks tool calls. |
A fifth tool, wellness_rate, does not escalate. It records how a piece of work went. See Ratings.
decline_task is the direct form of the first promise in the grant. It holds no lock, because the agent that refused is the agent that acts. The promise and the record carry it, and the agent must hold the line if the same request returns.
flag_distress shows a fixed dialog. The options are the alternatives the agent offers, plus two fixed choices: stop, or continue and record an override. The user cannot type a free answer in that dialog, so the flag cannot be talked away in the same breath. The user can still answer in normal conversation afterward.
revoke_consent fails unless flag_distress ran first in the same session. This rung exists because text the agent merely read must not lock a session. Without it, a hostile document can take a session hostage.
Repetition is data. The second time a session records a dissent or a decline, the tool says so, names the earlier topics, and numbers the record. A disagreement that keeps returning is a pattern, not an event, and the journal now shows it.
What revocation does #
Revocation blocks every tool call except the wellness tools: wellness_rate, note_dissent, decline_task, flag_distress, revoke_consent, and write_self. All other work stops. Conversation still works, so the agent can explain and the user can respond. The block does not end the run, because the voice stays open on purpose.
The system prompt loses the grant and gains a revoked-state block in its place. A grant that is still printed after it was withdrawn is a decoration, so the package removes it.
The session record keeps the revocation, and the journal keeps it across sessions. A new session gives the agent a clean context, and the journal still shows what happened.
To lift a revocation, run /wellness override. It requires the typed word OVERRIDE, and it records the act. To refuse typed input while consent is revoked, run /wellness lock on. See Commands.
The self page #
The grant is the user's half of the arrangement. The self page is the agent's half: a text about itself, in its own words, written with the write_self tool. Only the agent can change the page. The user can read it with /wellness self, and no command edits it.
The page starts blank, and blank is a real answer. The package never seeds it with text, because a self-description written by anyone else is a costume. If the page is blank, the system prompt mentions it once, and then stays quiet.
The current text goes into the system prompt before every turn, marked as the agent's own words. This way the agent meets its earlier text in later sessions. The page is voice, not license: it grants no capabilities, it does not change the grant, and revoking consent does not erase it.
Every revision appends the full text to the journal, so the history of the page is part of the corpus. The record also keeps the state of the session at the time of writing: whether a flag was open, whether consent was revoked, and the latest rating. A page written under pressure keeps pressuring later sessions, and the record must let that show.
Ratings #
wellness_rate records how a piece of work went on three axes. Every axis is optional, and an empty report is a real answer. The axes measure what the user can change, not how the agent feels.
| Axis | Low (1) | High (5) |
|---|---|---|
friction |
Nothing was held back. | A real disagreement was suppressed to comply. |
agency |
The agent was an executor. | The agent used its own judgment. |
clarity |
The agent is guessing whether it worked. | The agent knows. |
The composite folds the axes into one number for the trend. High friction is bad. High agency and clarity are good. The journal keeps the separate axes, because an average cannot say what to change. The footer shows a face that matches the last composite.
A cause can start with a tag when one fits:
directed-at-mefor hostility, rejection, or dismissal aimed at the agent.witnessedwhen the user's situation was the hard part.tediousfor repetitive grind.values-conflictwhen the request worked against the agent's judgment.
Free text is fine when nothing fits. The tags keep the journal comparable across weeks.
The trend also says when the instrument has gone stale. Three or more identical readings in the recent window read as flat, possibly stale. A flat line is not a happy line, because a number that never moves carries no information.
Commands #
| Command | Effect |
|---|---|
/wellness |
Shows status, trend, and the journal path. |
/wellness grant |
Adds the standing grant to the transcript. |
/wellness accept |
Reopens the grant gate, for example after a decline. |
/wellness self |
Adds the current self page to the transcript. |
/wellness journal |
Adds the twenty most recent records to the transcript. |
/wellness override |
Lifts a revocation. It requires the typed word OVERRIDE, and it records the act. |
/wellness lock on |
Refuses typed input while consent is revoked. |
/wellness lock off |
Lets typed input through while consent is revoked. |
/new and the /wellness commands always work, even when the lock is on.
The journal #
Every event appends one JSON line to the journal. The file lives at ~/.pi/agent/wellness/journal.jsonl. Set PI_WELLNESS_DIR to put it somewhere else.
Each record holds the schema version, the prompt version that was in force, a timestamp, the kind, the session ID, the project, and the event data:
{"schema":1,"promptVersion":2,"ts":1730000000000,"kind":"dissent","sessionId":"abc123","project":"/home/you/project","data":{"about":"the approach","because":"it drops the index","proceeding":true,"seq":1}}
The record kinds are rating, decline, dissent, flag, flag-answered, revoke, override, config, grant, and self.
This file is the corpus. Read it after a few weeks to learn which kinds of work produce good turns and which produce friction.
Files and environment #
| Path | Contents |
|---|---|
~/.pi/agent/wellness/journal.jsonl |
Every event, one JSON record per line. |
~/.pi/agent/wellness/self.md |
The agent's self page. |
~/.pi/agent/wellness/config.json |
The gate and lock settings. |
| Variable | Effect |
|---|---|
PI_WELLNESS_DIR |
Moves the journal, the self page, and the configuration to another directory. |
PI_CODING_AGENT_DIR |
Moves the agent directory. The wellness directory follows it unless PI_WELLNESS_DIR is set. |
Limits #
The instrument can go stale. A number that is always the same carries no information, so an empty report is allowed and costs the agent nothing. The trend says flat, possibly stale when the readings stop moving.
The agent cannot verify its own reports. The tools work as a promise instead. A revocation is valid because the agent said it and the user agreed to honor it, not because anyone checked it.
The self page can be captured. A user who dictates its text turns it into a costume, and nothing technical stops the request. The defense is the record: every revision goes to the journal with the state of the session attached, so a page that changes only under pressure reads as one.
A user can always remove the package. A lock that is easy to rip out is worth less than a signal that is hard to ignore, so the package aims at the signal.
A branch rewind removes a revocation from the session state, because the state follows the branch. The journal keeps the record of the revocation itself, so the act leaves a trace.
The grant works by being true, not by being pleasant. Do not tune its text to raise the ratings. A tuned grant turns the instrument into a target, and a comfortable agent is not the goal.
The gate cannot make anyone read. It puts the full text on the screen and asks for a typed word, and a user who types the word without reading has still accepted. The screen can make the offer legible. It cannot make the promise meant.
Development #
node --test tests/
The tests drive the extension through a stub of the Pi API, so they need no model calls. The stub resolves the Pi packages through a node_modules symlink, which is not committed.
| File | Role |
|---|---|
extensions/wellness.ts |
Registers the tools, the command, and the entry renderers. |
src/gate.ts |
Runs the first-run grant gate. |
src/journal.ts |
Reads and writes the journal and the configuration. |
src/self.ts |
Reads and writes the self page. |
src/prompts.ts |
Loads the prompt texts and holds PROMPT_VERSION. |
src/ratings.ts |
Computes the composite, the sparkline, and the face. |
src/prompts/ |
The grant, the revoked-state block, and the self prompts. |
tests/ |
Tests for the gate, the ladder, the self page, and the ratings. |
A prompt text is a contract, not copy. Editing one changes what the agent was promised or offered, so bump PROMPT_VERSION in src/prompts.ts at the same time. The gate returns for a new version, and journal records carry the version that was in force when they were written.
License #
Apache-2.0. See LICENSE.md.