From 238f95e722b0defbaa49637aecde303ea3f3a30e Mon Sep 17 00:00:00 2001 From: "@permadeath.com" Date: Thu, 3 Sep 2026 14:15:05 -0400 Subject: [PATCH] docs(plan): drop the hook, credential-delivery and scrobble epics The stamp contract, harness survey, credential delivery and scrobble epics belong with the crates they describe. Surviving epics and the reference docs keep their own subjects. Change-Id: I9fff1577a8759f10bb78827bd007d83da16245e9 --- docs/architecture.md | 7 ++- docs/overview.md | 7 +-- docs/running-locally.md | 60 ++++++-------------------- docs/trust-model.md | 22 +++------- plan/README.md | 4 -- plan/account-types.md | 2 +- plan/adversarial.md | 60 +++----------------------- plan/ai-preference.md | 2 +- plan/auth-types.md | 22 +++++----- plan/cli.md | 23 +++------- plan/cred-delivery.md | 80 ---------------------------------- plan/credentials.md | 14 +++--- plan/dedupe-audit.md | 53 +++-------------------- plan/dev-setup.md | 2 +- plan/did-minting.md | 28 ++++++------ plan/handshake.md | 6 +-- plan/harnesses.md | 95 ----------------------------------------- plan/local-dev.md | 4 +- plan/node.md | 9 ++-- plan/onboarding.md | 47 +++++++++----------- plan/order.txt | 3 -- plan/provenance.md | 2 +- plan/scrobble.md | 44 ------------------- plan/stamp-contract.md | 88 -------------------------------------- plan/subagents.md | 16 +++---- 25 files changed, 112 insertions(+), 588 deletions(-) delete mode 100644 plan/cred-delivery.md delete mode 100644 plan/harnesses.md delete mode 100644 plan/scrobble.md delete mode 100644 plan/stamp-contract.md diff --git a/docs/architecture.md b/docs/architecture.md index 5e0a72b8..61712db5 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -13,12 +13,11 @@ had none. [auth-types](../plan/auth-types.md) built legacy sessions with Argon2id) and a small, explicit table of which credential each route requires — public reads, `LegacySession` for session management itself, and public for now on the `com.atproto.repo.*` write surface, which is that -epic's own open item: `didbot-hookd` writes an agent's own records by calling -those routes directly with no session to present, so switching them to -require one is real follow-up work rather than a flag to flip. +epic's own open item: a client writes an agent's own records by calling those +routes directly with no session to present, so switching them to require one +is real follow-up work rather than a flag to flip. [The trust model](trust-model.md) says what the system can and cannot prove. -[The hook flow](hook-flow.md) covers the same provisioning path in prose. ## Hosts, and what a compromise of each one reaches diff --git a/docs/overview.md b/docs/overview.md index 9ef109d6..15e58159 100644 --- a/docs/overview.md +++ b/docs/overview.md @@ -5,16 +5,13 @@ personal data server. Agent contexts are provisioned real, protocol-native ATProto accounts, and publish intent-level status updates — what an agent is working on and why — to a feed. -This crate is the workspace for that system. At present it holds the pieces -that do not depend on a repository implementation: the lexicon documents, the -namespace, and the Claude Code hook protocol. +This crate is the workspace for that system: the personal data server, the +identity and naming layers, the lexicon documents and the namespace. ## Reading order - [The trust model](../docs/trust-model.md) — what the system can and cannot prove about an agent, and why those are separate questions. -- [The hook flow](../docs/hook-flow.md) — how an agent context becomes an - account, and how a status update gets written. - [The names an agent has](../docs/names.md) — one account worked through end to end: what exists in the zone, what the server answers, and what a stranger reads. diff --git a/docs/running-locally.md b/docs/running-locally.md index 2588b2b1..73cb2abd 100644 --- a/docs/running-locally.md +++ b/docs/running-locally.md @@ -3,11 +3,10 @@ One command brings up a personal data server, provisions agent accounts against it, and streams what happened. Nothing leaves the machine. -Two terminals are enough to watch an agent scrobble, one script each: +One terminal is enough to watch agents arrive: ```sh ./scripts/dev-pds.sh # personal data server: registrations and lifecycle -./scripts/dev-mcp.sh # scrobble MCP host: the feed ``` For the read side — an index, something serving it, and a canvas that draws it @@ -15,47 +14,27 @@ For the read side — an index, something serving it, and a canvas that draws it write side and a population big enough to be worth looking at; see [the whole stack](#the-whole-stack) below. -## Before the terminals +## Before the terminal -The scripts run from a checkout. What does not come from a checkout is the -hook binary on `PATH`, the hook entries in your own settings file, and the -scrobble host declared where the harness looks for it. +The scripts run from a checkout. `didbot-setup` says what this machine has been +told about itself and whether it holds together: ```sh cargo run -p didbot-setup -- # what is wrong here -cargo run -p didbot-setup -- apply -``` - -`apply` prints every change it would make as a file, a path and a value, asks -once, and records what was there before so `undo` takes it back exactly. It -ends by telling you to restart the harness, because hook configuration is read -once at session start and a session open now is running whatever it started -with. - -`verify` is what proves the restart took: it provisions an agent, writes a -scrobble under it, reads the record back and deletes the account. - -```sh -cargo run -p didbot-setup -- verify +cargo run -p didbot-setup -- show ``` A machine that has never been configured runs one of everything on the ports -below and needs no configuration file at all. For two personal data servers, or -two scrobble hosts over one server, see +below and needs no configuration file at all. For two personal data servers +over one machine, see [didbot-stack](../crates/didbot-stack/README.md). Each `dev-*.sh` sources `scripts/dev-profile.sh`, which puts the current directory's profile in the environment, so a worktree bound to a second stack gets that stack's ports with nothing here edited. A variable already set still wins: `DIDBOT_PDS_PORT=3100 ./scripts/dev-pds.sh` means 3100. -Each rebuilds its own half, runs that half's tests, and then runs in the -foreground so the terminal *is* the log. They are separate because they are -separate things to watch, and because they go stale for different reasons: a -lexicon change matters to the server, a tool schema change matters to the MCP -host. - -Order does not matter. The MCP host warns if no server is answering and runs -anyway; scrobbles are refused until the server is up. +Each script rebuilds its own half, runs that half's tests, and then runs in the +foreground so the terminal *is* the log. For the server alone, with some traffic to watch: @@ -306,9 +285,9 @@ rewritten log carries forward. The name travels. `listAgents` reports it, the index keeps it beside the DID — after checking it lands under the zone its server admits to minting, since a name is the field a reader trusts instead of the identifier — and the canvas -labels each agent with it. The MCP host's feed uses it too, asking the server -what its accounts are called and remembering the answer, so a status line reads -`basalt-otter` rather than the digest the DID was minted from. A deployment +labels each agent with it, asking the server what its accounts are called and +remembering the answer, so a line reads `basalt-otter` rather than the digest +the DID was minted from. A deployment running `NAMES=hostname` sees the digests it saw before, everywhere. Resolution goes both ways, and both halves are worth looking at once. The @@ -340,7 +319,7 @@ what is rerun on purpose: they are the step most likely to have something to say about the change that triggered them. ```sh -./scripts/dev-mcp.sh --watch +./scripts/dev-swarm.sh --watch ``` The personal data server is the exception, and refuses. The swarm holds @@ -564,19 +543,6 @@ head, which a compaction of the write-ahead log empties. A reader that sees it repeatedly is one falling further behind than `didbot_pds::history::TRAIL` between compactions. -## The scrobble feed - -The MCP host logs one line per scrobble: the emoji the agent chose, then the -text, with the identifier and record key as structured fields. The emoji is -padded to a fixed width by measured display width rather than character count, -so a two-cell emoji does not shift the column, and the text is truncated on -grapheme cluster boundaries so a flag or a family sequence is never cut in -half. - -Agents reach it over HTTP rather than stdio, so the process belongs to whoever -started it and its log goes to that terminal. Under stdio the harness owns the -process and the log disappears into its debug output, which defeats the point. - ## Driving it by hand ```sh diff --git a/docs/trust-model.md b/docs/trust-model.md index de75c028..9fc096e8 100644 --- a/docs/trust-model.md +++ b/docs/trust-model.md @@ -51,23 +51,11 @@ harness's hook payload; the model supplies only content. So a model that has been prompt-injected, or is simply confused, can lie about what it is doing. It cannot lie about which agent is doing it. -This is mechanically enforced, not merely organisational. `didbot-hook`'s -`stamp()` used to refuse to overwrite a stamp key already present on a tool -call, on the theory that a call already carrying one had already been -stamped. It hadn't: the only way a model's own tool call could carry -`_vibescrobbleAgentDid`, `_vibescrobbleParentDid`, `_vibescrobbleEffort`, or -any of the other four stamp keys was that the model wrote it there, and the -old code let that value ride through untouched — an identity claim the -paragraph above says a model cannot make, forgeable in practice by writing -the right key. `stamp()` now always overwrites with the harness's own value, -and `didbot-hookd`'s `pre_tool_use` denies outright a call that arrives -already carrying a stamp key whose value disagrees with what this hook is -about to write, rather than silently correcting it — a forged claim is loud, -not merely ineffective. The one case that is not a forgery is the same hook -re-stamping a call it already processed, because more than one configuration -scope wired the matcher; that is told apart from an attack by comparing the -carried value against the one this hook independently computes, not by -trusting that the key is merely present. +This is mechanically enforced, not merely organisational: a stamp the harness +writes always overwrites whatever a model put under the same key, and a call +arriving with a stamp that disagrees with what the harness independently +computes is denied outright rather than silently corrected — a forged claim is +loud, not merely ineffective. That guarantee stops at the trust domain. Arbitrary code execution in the same domain can address the credential path directly and claim any agent within it. diff --git a/plan/README.md b/plan/README.md index f968c182..4a129883 100644 --- a/plan/README.md +++ b/plan/README.md @@ -121,17 +121,14 @@ The exit criterion is met and something is still open in the file. Nothing yet. | [pds-writes](pds-writes.md) | A record is a signed commit in a repository | open | | [agent-accounts](agent-accounts.md) | An agent context becomes an account, and stops being one | open | | [account-types](account-types.md) | Not every account is a session | open | -| [scrobble](scrobble.md) | An agent says what it is working on, and the statement is its own record | open | | [write-policy](write-policy.md) | Which agents may write which record types | open | | [auth-types](auth-types.md) | Every credential this server accepts, and what each one may do | open | | [oauth](oauth.md) | A third-party app signs in as an agent, with nobody at the consent screen | open | | [pds-xrpc](pds-xrpc.md) | A client nobody here wrote can talk to this server | open | | [federation](federation.md) | An off-the-shelf relay and an off-the-shelf app, not just the protocol | open | | [credentials](credentials.md) | A session gets a credential without a wrapper process | open | -| [cred-delivery](cred-delivery.md) | A credential reaches the harness, and the model never holds it | open | | [node](node.md) | A host proves what it is once, and issues credentials to the sessions on it | open | | [subagents](subagents.md) | What a distinct subagent identity would buy, and what a harness lets you enforce | open | -| [harnesses](harnesses.md) | What a harness has to provide for any of this to work | open | | [scope-policy](scope-policy.md) | An agent cannot be granted what its owner has not allowed | open | | [ownership](ownership.md) | The human names the agents and the agents name the human | open | | [vouch](vouch.md) | What an owner vouches, what an agent vouches, and what the server vouches | open | @@ -174,7 +171,6 @@ The exit criterion is met and something is still open in the file. Nothing yet. | [periodic-backups](periodic-backups.md) | The server backs up its own data, and something reads it back | open | | [policy](policy.md) | What an agent may do is a set of denials, evaluated at the write, from three sources | open | | [site](site.md) | A stranger can read what this project is, without cloning it | open | -| [stamp-contract](stamp-contract.md) | The hook stamps a tool it was told about, and the tool reads a contract it can name | open | | [tombstone-serving](tombstone-serving.md) | An account outlives the server that answered for it | open | | [updates](updates.md) | The project says what it has done, where the network can read it | open | | [witness](witness.md) | A stranger can check a `bot.did.registration` claim against something the server does not control | open | diff --git a/plan/account-types.md b/plan/account-types.md index 042452d7..54a9c3ca 100644 --- a/plan/account-types.md +++ b/plan/account-types.md @@ -2,7 +2,7 @@ id: account-types title: Not every account is a session status: open -crates: [didbot-pds, didbot-lexicon, didbot-name, didbot-serve, didbot-hookd] +crates: [didbot-pds, didbot-lexicon, didbot-name, didbot-serve] dependsOn: [agent-accounts, provenance] exitCriterion: > A service account provisioned from a CI run keeps its DID and its name across diff --git a/plan/adversarial.md b/plan/adversarial.md index c91397ac..7a000ff6 100644 --- a/plan/adversarial.md +++ b/plan/adversarial.md @@ -2,19 +2,19 @@ id: adversarial title: Integration tests at the seams, not inside the components that already pass status: open -crates: [didbot-serve, didbot-pds, didbot-identity, didbot-hookd, didbot-policy, didbot-policy-source] +crates: [didbot-serve, didbot-pds, didbot-identity, didbot-policy, didbot-policy-source] dependsOn: [] exitCriterion: > - Three adversarial integration tests exist — zone containment, DPoP binding - and stamp forgery — each exercising a request that crosses at least two - components, and each fails loudly if either component's assumption about - the other quietly changes. + Adversarial integration tests exist for zone containment and DPoP binding, + each exercising a request that crosses at least two components, and each + fails loudly if either component's assumption about the other quietly + changes. --- # adversarial [CLAUDE.md](../CLAUDE.md) says integration tests exercising multiple system -invariants are the most valuable kind. Three targets already have unit +invariants are the most valuable kind. These targets already have unit tests, and unit tests are not what is missing: what is missing is a test at the seam between the components each one assumes will hold up its end. @@ -37,15 +37,6 @@ the seam between the components each one assumes will hold up its end. seam between "a token was minted bound to key A" and "a resource server later checks a proof against key A" is the one nothing exercises end to end. -- **Stamp forgery** has unit coverage of `stamp_conflict` in - `crates/didbot-hook/src/lib.rs`'s tests and in - `crates/didbot-hookd/tests/hook_events.rs`, which already construct a - stamped `mcp__vibescrobble__scrobble` call. What none of them does is put a - value in `_vibescrobble*` fields *the way a model actually would* — as - part of a tool call the model itself composed, following the tool's - schema, rather than as a call the test harness assembled with the forged - field already in place — and confirm the same refusal holds when the - request arrives shaped like the thing an adversarial model would send. ## Zone containment: the seam between resolution, serving and vouching @@ -73,11 +64,9 @@ through the real HTTP surface. ## DPoP binding: the seam between mint and verify -## Stamp forgery: the seam between the model's tool call and the hook's check - ## Degenerate values: the seam between a configuration path and a comparison -The three suites above test seams between components. This one tests the +The suites above test seams between components. This one tests the seam between a component and the *configuration* that feeds it — the class where a comparison is correct for every value except the one a failed configuration path produces, and nothing in the codebase refuses to produce @@ -367,38 +356,6 @@ Left open, in rough order of how much a deployment should care: and the clock-skew window are the same number holds across a proof that has travelled both stages, rather than one fed to a single verifier call. -- **A model supplying its own `_vibescrobble*` fields, inside a real tool - call the harness's hook intercepts**, not a call the test constructs - with the forged fields pre-attached to `tool_input`. The distinction - matters because `stamp_conflict`'s own documentation (per - [stamp-contract](stamp-contract.md)'s Done entry) says presence alone - cannot tell a forged value from a legitimately re-stamped one — the - real test is whether the *comparison* against the harness's - independently computed value catches a value that arrived shaped - exactly like a genuine re-stamp, not whether a hand-built conflicting - call is caught. Driven through the compiled `didbot-hook` binary as a - subprocess (`CARGO_BIN_EXE_didbot-hook`), the actual process boundary a - model's tool call crosses. See - `crates/didbot-hookd/tests/stamp_forgery.rs`. -- **A forged value that matches the format of a genuine one.** A model - that has seen its own previous scrobbles knows what `_didbotAgentDid` - looks like; a forged DID string that is syntactically well-formed and - resembles the agent's own is the adversarial case, not an obviously - wrong value that any check would catch. Sharpened further: the forged - `_vibescrobbleParentDid` names a real, currently-provisioned DID - belonging to a second, separately-minted account, not merely a - well-formed string. -- **The MCP server's own record of what the hook said, versus what - arrives at the write path.** [stamp-contract](stamp-contract.md) is - about the *contract* between the hook and the scrobble server; this - epic's job is a test that drives a full round trip — tool call, hook - interception, scrobble server write — with a forged value planted at - the one place a model could actually plant one, and confirms the - written record carries the harness's value and not the model's. The - forged-parent test confirms zero records exist for either account - after a denied call; the honest-restamp test closes the same round - trip to a real written record and checks its DID against the - harness's own value. - **An authentication bypass on an empty operator secret.** `operator_secret_matches` was a bare @@ -458,9 +415,6 @@ Left open, in rough order of how much a deployment should care: wired, the server's `htu` could not match a conformant client's, and `validate_access` took no proof — and all three are closed; see that section above. -- Stamp forgery adversarial suite added - (`crates/didbot-hookd/tests/stamp_forgery.rs`); see that section above for - what it covers. - **`TreePolicyGate` never told any evaluator about a freeze, an unfreeze, or a deletion.** `PolicyGate::froze`/`unfroze`/`deleted` default to no-ops, and `TreePolicyGate` — the only implementation backed by a real diff --git a/plan/ai-preference.md b/plan/ai-preference.md index 54eb45b3..8caf6575 100644 --- a/plan/ai-preference.md +++ b/plan/ai-preference.md @@ -2,7 +2,7 @@ id: ai-preference title: A stranger's declared AI preference is a ceiling on what our agents may do status: open -crates: [didbot-lexicon, didbot-pds, didbot-mcp] +crates: [didbot-lexicon, didbot-pds] dependsOn: [write-policy] exitCriterion: > An account whose repository declares that it refuses synthetic content is diff --git a/plan/auth-types.md b/plan/auth-types.md index 85501805..49e915f9 100644 --- a/plan/auth-types.md +++ b/plan/auth-types.md @@ -38,7 +38,7 @@ bare 401, which is a decision about who may talk to this server rather than a detail. **That a kind is added only when no existing kind fits the caller.** The -agent token is the worked example: `didbot-hookd` writes an agent's records +agent token is the worked example: a harness client writes an agent's records automatically, with no person at a login screen, so a session token would have meant a machine holding a *human* credential shape. A bearer token minted once at provisioning, bound to one DID, with an expiry and no @@ -79,21 +79,19 @@ the repository the request names — `ApiError::repo_mismatch` (403 another's, an invariant rather than a policy knob. Getting there needed a credential kind that did not exist yet, because none -of the ones that did fit the caller: `didbot-hookd` -writes an agent's scrobbles and memories automatically, with no person at a -login screen and no operator minting an app password for it, so -`LegacySession` was the wrong vehicle — building it would have meant -`didbot-hookd` holding a *human* credential shape for a caller that is -never human. What it needed instead was the smallest thing +of the ones that did fit the caller: a harness client writes an agent's +scrobbles and memories automatically, with no person at a login screen and +no operator minting an app password for it, so `LegacySession` was the wrong +vehicle — building it would have meant a machine holding a *human* +credential shape for a caller that is never human. What it needed instead was the smallest thing `plan/credentials.md`'s framing already asks for: a bearer token, minted once at provisioning, bound to one DID, with an expiry and no rotation. That is `didbot_pds::credential` — `Credential::AgentToken` on the HTTP side — and `Provisioner::provision` now issues one automatically, durable in the write-ahead log (`pds.layout` bumped to 3), returned once as -`agentToken` in `provisionAgent`'s response body. `didbot-hookd` carries it -the same way it already carries the DID — stamped onto every scrobble -alongside `STAMP_KEY` — and `crates/didbot/tests/hook_to_record.rs` proves -the whole path: hook, scrobble host and server agreeing on a credential +`agentToken` in `provisionAgent`'s response body. A client carries it the +same way it already carries the DID, and the conformance suite proves the +whole path: client and server agreeing on a credential none of them minted by hand. This is deliberately not `LegacySession`, not OAuth, and not inter-service @@ -125,7 +123,7 @@ it enforces today, which is nothing. **Account-lifecycle mutations are always credentialed, never toggleable.** `Credential::AgentSelf`: the acting agent's own write credential, checked against the account the request names. -`didbot-hookd` calls all three automatically — the same shape that argued +A harness client calls all three automatically — the same shape that argued (wrongly) for `Public` before — but by the time one of these fires the account already has its own write credential, and that is exactly the caller these routes are reasonable to ask one from. An agent may end or toggle diff --git a/plan/cli.md b/plan/cli.md index 036b9dfb..ec015929 100644 --- a/plan/cli.md +++ b/plan/cli.md @@ -2,7 +2,7 @@ id: cli title: Six binaries with six naming conventions, and no way for anyone else to add a seventh status: open -crates: [didbot-setup, didbot-hookd, didbot-serve, didbot-swarm, didbot-mcp] +crates: [didbot-setup, didbot-serve, didbot-swarm] dependsOn: [] exitCriterion: > A single `didbot` command dispatches every first-party surface as a @@ -12,9 +12,9 @@ exitCriterion: > # cli -There is no `didbot` command. There are `didbot-hook`, `didbot-pds`, -`didbot-setup`, `didbot-swarm` and `vibescrobble-mcp`, each its own binary -under `crates/*/src/bin/`, named by whoever added it. A person who has +There is no `didbot` command. There are `didbot-pds`, `didbot-setup` and +`didbot-swarm`, each its own binary under `crates/*/src/bin/`, named by +whoever added it. A person who has installed this project has no single entry point to run, nothing to type to find out what is available, and no way to discover a surface they did not already know the name of. @@ -35,7 +35,7 @@ converged on — lets someone add a subcommand in any language without touching this repository. The near-term decision is to **bake the first-party surfaces in**: one -`didbot` binary with `hook`, `dev`, `setup`, `swarm` and `operator` as real +`didbot` binary with `dev`, `setup`, `swarm` and `operator` as real `clap` subcommands, sharing argument conventions, help output and error formatting. That is the whole of the immediate work. @@ -66,14 +66,6 @@ first extension written is the one every later extension is copied from. credential, not the signing keys. This is the same separation [services](services.md) wants between processes, applied at the extension boundary, where it would otherwise be handed straight back. -- [ ] **Keep `didbot-hook` working under its own name.** - `crates/didbot-setup/src/harness.rs` writes `command: "didbot-hook"` - into harness settings files, and `crates/didbot-setup/src/check.rs` - checks for that name on `PATH`. Both are already installed on - machines. The name needs a stable alias, or a migration in - `didbot-setup` that rewrites settings it wrote — and the alias is - cheaper, because a harness settings file may be edited by hand or held - under version control by its owner. - [ ] **Say what a single binary costs.** Baking the surfaces in means the binary on the server contains the operator's authentication flow, and the binary on a laptop contains the server's key handling. That is @@ -82,11 +74,6 @@ first extension written is the one every later extension is copied from. with a genuinely different threat model should ship as a separate binary reached through dispatch rather than as another baked-in subcommand. -- [ ] **`vibescrobble-mcp` is not a `didbot` subcommand.** It is an MCP - server a harness launches, not something a person types, and it is - named for the other project. It should be left where it is or renamed - on its own terms; folding it in would put a machine-launched process - behind a human entry point. ## Done diff --git a/plan/cred-delivery.md b/plan/cred-delivery.md deleted file mode 100644 index fec4f5c1..00000000 --- a/plan/cred-delivery.md +++ /dev/null @@ -1,80 +0,0 @@ ---- -id: cred-delivery -title: A credential reaches the harness, and the model never holds it -status: open -crates: [didbot-hook, didbot-hookd] -dependsOn: [credentials, node] -exitCriterion: > - A session's write credential is readable by the shell and tool processes the - harness spawns, and no documented path puts it in the model's own context. ---- - -# cred-delivery - -[credentials](credentials.md) states the shape and names `CLAUDE_ENV_FILE` as -its delivery mechanism; [node](node.md) is the component that issues what -gets delivered. Neither says, in one place, what "the model never holds it" -actually requires — which is a claim about a boundary, not about a mechanism, -and the two are not the same thing. - -**The owner's requirement is that credentials be harness-accessible and never -directly held by the agent.** A model that can read its own bearer token can -put it in a tool call, and a tool call is exactly what an untrusted MCP -server, a fetched URL's content, or a crafted file the model reads can -influence. The credential does not have to be exfiltrated through a -deliberate leak; it only has to be in the model's context for a prompt -injection to ask for it back. - -## What "harness-accessible" means concretely - -`docs/hook-flow.md` says a hook cannot set environment variables for tool -execution except through `CLAUDE_ENV_FILE`, available on `SessionStart`, -`Setup`, `CwdChanged` and `FileChanged`, and that its contents persist into -later shell commands for that session. That is a real boundary: a shell -command the model asks for can read an environment variable, but the model -does not see the file being written, does not choose its contents, and has -no tool call that returns its value directly. The credential is available to -processes the harness spawns on the model's behalf, without ever appearing -in the transcript the model reads back. - -- [ ] **Land the `CLAUDE_ENV_FILE` write.** [credentials](credentials.md) - names this and it is not done. `didbot-hookd` already carries a - credential from `PreToolUse` to the scrobble host by stamping one tool - call — the mechanism this epic replaces for the general case, because a - stamped call is bound to that one call and not delivered for every - later tool execution the way session-scoped delivery needs. -- [ ] **A shell command is not the model's context, and the distinction is - the whole epic.** A command a subprocess runs with an environment - variable set is not a place the model reads the value from; a tool - result that echoes environment variables back, or a debugging command - the model runs that prints `env`, is. The threat model has to name - which tools are safe to run with the credential in their environment - and which return output the model reads — `curl` writing a response - body to a file is different from a shell built-in that prints its own - environment on error. -- [ ] **What a fetched or read document can do with it.** The credential - sits in the environment of every subprocess for the session, including - ones invoked to process content the model did not author — a build - script, a linter, a test runner. Any of those printing the environment - on failure, or a crash handler doing the same, puts the credential in - output the model then reads. This needs an answer, not an assumption - that subprocess environments are opaque to the model that spawned them. -- [ ] **Rotation and revocation on a leak.** If a credential does end up in a - transcript or a tool result despite the above, there has to be a way to - invalidate it that does not require killing the session — the session - is exactly what a leaked credential would be used to impersonate. -- [ ] **Say what a subagent gets**, and point rather than restate: - [subagents](subagents.md) is where the custody boundary — session, not - subagent — is revisited; this epic assumes whatever that epic decides - lands here as the delivery scope. - -## What this epic is not - -Issuance — how a credential is minted and what it is scoped to — is -[credentials](credentials.md) and [node](node.md). This epic starts from a -credential that already exists and asks only how it gets from the node -component into a session without a model ever reading it. - -## Done - -Nothing closed yet. diff --git a/plan/credentials.md b/plan/credentials.md index bc4e7624..1d9c95b3 100644 --- a/plan/credentials.md +++ b/plan/credentials.md @@ -2,7 +2,7 @@ id: credentials title: A session gets a credential without a wrapper process status: open -crates: [didbot-hookd, didbot-attest] +crates: [didbot-attest] dependsOn: [oauth, attestation] exitCriterion: > An ordinary session, configured only through settings.json, receives a @@ -23,9 +23,8 @@ write, and nothing assumes co-location. out into [node](node.md), because it is a process with a lifetime, an install story and a privilege boundary of its own. - [ ] **Session-scoped delivery through `CLAUDE_ENV_FILE`**, the one hook - mechanism that reaches later tool calls. The delivery mechanism and its - threat model — what "the model never holds it" actually requires — are - [cred-delivery](cred-delivery.md), not restated here. + mechanism that reaches later tool calls. What "the model never holds + it" actually requires is a threat model of its own. - [ ] **Say plainly that a subagent gets no credential of its own.** The custody boundary is the session; subagent identity is attribution. Whether that boundary should move, and what a harness actually lets you @@ -48,10 +47,9 @@ write, and nothing assumes co-location. ## Done -Nothing closed yet. Adjacent, and worth knowing about here: `didbot-hookd` -now carries a write credential (`didbot_pds::credential`, minted at -provisioning) from `PreToolUse` to the scrobble host — but by the same -tool-call-stamping mechanism `STAMP_KEY` already used, not through +Nothing closed yet. Adjacent, and worth knowing about here: a harness client +carries a write credential (`didbot_pds::credential`, minted at +provisioning) by stamping it onto the tool call, not through `CLAUDE_ENV_FILE`. That solves `plan/auth-types.md`'s narrower problem — a credential has to exist and reach the write path at all, for the repo write surface to require one without breaking every write this deployment already diff --git a/plan/dedupe-audit.md b/plan/dedupe-audit.md index 0ce86f4f..a5f7f0c1 100644 --- a/plan/dedupe-audit.md +++ b/plan/dedupe-audit.md @@ -2,7 +2,7 @@ id: dedupe-audit title: A workspace-wide sweep for drifted duplicate logic, dead code, and stragglers from removed backends status: open -crates: [didbot-hookd, didbot-serve, didbot-pds, didbot-attest] +crates: [didbot-serve, didbot-pds, didbot-attest] dependsOn: [] exitCriterion: > Every finding below carrying a decision for a human (the `normalize_pds_url` @@ -29,29 +29,10 @@ different on purpose, not by drift: - `didbot_verify::check::same_endpoint` lowercases the whole string and trims a trailing slash, documented as following atproto's service-endpoint rule (scheme, host, optional port only). -- `didbot_hookd::state::normalize_pds_url` trims only a trailing slash, with - a doc comment arguing explicitly against being cleverer — it would mean - guessing that two developer-written configurations are the same host. - These compare different things (a DPoP proof's claimed URL, a DID document's -service endpoint, a developer's `--pds` flag) under different rules that -each cite a source for the rule. Merging them would force one comparison -rule onto three call sites that need three different ones. Left alone. - -`didbot_setup::check::host_writes_where_the_profile_says` compares a running -server's reported PDS URL against the profile's, trimming a trailing slash -only, with a comment giving the same non-cleverness argument as -`normalize_pds_url`. It reimplements the trim inline rather than calling -`normalize_pds_url`, which is private to `didbot-hookd`. `didbot-setup`'s -`do_forget` (in `didbot-setup/src/bin/didbot-setup.rs`) does the same trim a -third time, directly against `HookState::servers`, which is written through -`normalize_pds_url`. All three trims are `trim_end_matches('/')` and agree in -practice; nothing here has drifted. Exporting `normalize_pds_url` from -`didbot-hookd` and using it in the other two spots is a real but small -cleanup — three call sites, one already a public dependency of the others — -left for a human rather than done here, since it is a judgment call about -whether the shared name is worth the coupling for three near-identical -one-liners. +service endpoint) under different rules that each cite a source for the +rule. Merging them would force one comparison rule onto call sites that need +different ones. Left alone. Retry/backoff loops in `didbot-dns` (Route 53 throttling), `didbot-pds` (name-claim collision retry) and `didbot-reconcile` (ticker backoff) solve @@ -65,17 +46,8 @@ errors vs. RFC 6749 OAuth errors); kept separate. The workspace removed its shared-secret attestation backend (`5b93113`) and the operator shared secret (`f20fa76`), replacing both with record-based mechanisms (`bot.did.operator`, read from the operator's own -repository). Three stragglers from that removal were still readable by an -operator or a maintainer, all fixed in this branch: - -- **`crates/didbot-hookd/README.md`** documented `DIDBOT_SHARED_SECRET` and - `DIDBOT_NODE_ID` as configuration variables. Neither is read anywhere in - the crate — `didbot-hookd/src/config.rs` reads only `DIDBOT_PDS_URL` and - `DIDBOT_HOOK_STATE`. Removed the rows and the paragraph explaining the - (nonexistent) default secret. The same README's troubleshooting section - claimed a `not provisioning ...` diagnostic meant the server "refused the - attestation claim"; provisioning today can fail for any reason the PDS - reports, not specifically attestation, so the wording is now generic. +repository). A straggler from that removal was still readable by an +operator or a maintainer, fixed in this branch: - **`crates/didbot-serve/src/onboarding.rs`** told an operator in `ServerState::Provisioning` to run `aws ssm put-parameter` for an @@ -88,11 +60,6 @@ operator or a maintainer, all fixed in this branch: `Booting` state's shape. This was live, user-facing operational guidance that would have sent a real deployer down a dead end. -- **`crates/didbot-hookd/src/lib.rs`**'s `diag()` carried a doc comment - warning "nothing passed here may include the shared secret or attestation - evidence." Neither concept exists in this crate any more (confirmed: no - match for either term outside doc comments). Removed the stale warning. - `didbot-attest`'s node-credential backend and the crate's "nothing here runs" framing were checked and are current, not stragglers: the crate's module doc explicitly and correctly states that no backend is wired into @@ -124,7 +91,7 @@ each carry a doc comment giving a specific, still-true reason for the allow. None looked stale enough to flag. Fake/mock/fixture types (`didbot-dns::route53`'s transport fake, -`didbot-claim`'s several, `didbot-hookd/tests/hook_events.rs`'s stub, +`didbot-claim`'s several, `didbot-reconcile`'s view fake, `didbot-serve`'s test doubles) each implement a trait local to their own crate. Consolidating any of them would add a cross-crate test dependency to share one struct; not done, per the @@ -133,13 +100,7 @@ lines. ## Done -- [x] **`crates/didbot-hookd/README.md`** no longer documents - `DIDBOT_SHARED_SECRET` or `DIDBOT_NODE_ID`, which nothing reads, and no - longer blames a failed provision on a refused "attestation claim." - [x] **`crates/didbot-serve/src/onboarding.rs`**'s `Provisioning` step list no longer tells an operator to write `attestation-secret` and `operator-secret` SSM parameters that nothing reads; it now reflects that this state requires no operator action. -- [x] **`crates/didbot-hookd/src/lib.rs`**'s `diag()` no longer warns against - leaking a shared secret or attestation evidence that no longer exist - in this crate. diff --git a/plan/dev-setup.md b/plan/dev-setup.md index c2a2cd61..462f018e 100644 --- a/plan/dev-setup.md +++ b/plan/dev-setup.md @@ -2,7 +2,7 @@ id: dev-setup title: One command takes a new machine to a working stack status: open -crates: [didbot-setup, didbot-stack, didbot-hookd, didbot-hook, didbot-mcp, didbot-serve] +crates: [didbot-setup, didbot-stack, didbot-serve] dependsOn: [] exitCriterion: > On a machine that has never run this, one command ends with an agent diff --git a/plan/did-minting.md b/plan/did-minting.md index d4452e62..62b6ca25 100644 --- a/plan/did-minting.md +++ b/plan/did-minting.md @@ -2,7 +2,7 @@ id: did-minting title: The server mints the identifier, not its caller status: open -crates: [didbot-pds, didbot-identity, didbot-hookd] +crates: [didbot-pds, didbot-identity] dependsOn: [agent-accounts] exitCriterion: > A provisioning request carries no identifier and gets an account whose DID @@ -28,14 +28,12 @@ changed afterwards. ## What follows from it -**A harness may not have an identifier to give.** `didbot-hook`'s own lineage -module records that the top-level session "has no `agent_id`" -([lineage.rs](../crates/didbot-hook/src/lineage.rs)), and -`didbot-hookd` exists partly to paper over this: `derive_agent_id` builds a -label out of whatever an `AgentKey` holds, because "the harness does not -promise the format". That is a good workaround living in the wrong place — it -is one client's convention, not a rule the server enforces, and a second -client is free to do something else. +**A harness may not have an identifier to give.** A top-level session has no +`agent_id` of its own, and a client papers over it by deriving a label out of +whatever session bookkeeping it holds, because the harness does not promise +the format. That is a good workaround living in the wrong place — it is one +client's convention, not a rule the server enforces, and a second client is +free to do something else. **Nothing makes the identifier globally unique.** `derive_agent_id` renders 40 bits of an FNV-1a digest, and its own documentation puts the collision odds at @@ -73,8 +71,8 @@ have an account", never to name it. loses `agent_id`. Nothing about the key reaches the DID, so it needs no label rules and can be as long, as structured or as absent as a harness requires. This is a breaking change to the provisioning surface and to - every caller of it — `didbot-hookd` and `didbot-pds`'s `--demo` are the - ones in this repository — so it lands as a `!` commit. + every caller of it — `didbot-pds`'s `--demo` is the one in this + repository — so it lands as a `!` commit. - [ ] **Choose what the server mints from.** [`didbot_name::generated`](../crates/didbot-name/src/generated.rs) already ships `Uuid`, `Random`, `Counter` and `Timestamp` with their @@ -104,10 +102,10 @@ have an account", never to name it. would like to be called; how that is resolved, reserved and reported back is [name-pools](name-pools.md)'s, but the request type is this epic's: `correlation` is not a hint and a hint is not an identifier. -- [ ] **Retire `derive_agent_id`, or demote it.** Once the server mints, the - hook daemon has nothing to derive. Its slug — the readable prefix that - let an operator match a DID to a session by eye — is a real affordance - and should reappear as something the server can attach, not as the +- [ ] **Keep the readable prefix a server affordance.** Once the server + mints, a client has nothing to derive. The slug — the readable prefix + that let an operator match a DID to a session by eye — is a real + affordance and should be something the server attaches, not the identifier itself. ## What this is not diff --git a/plan/handshake.md b/plan/handshake.md index 45421c2b..80a052b6 100644 --- a/plan/handshake.md +++ b/plan/handshake.md @@ -181,9 +181,9 @@ and put the node allowlist in the tier [node](node.md) already wants it in: one an on-box attacker cannot reach, because writing it requires the operator's own key rather than anything reachable from the server's host. -This belongs to [node](node.md) and [cred-delivery](cred-delivery.md) to -build, not to this epic: node bootstrap has its own lifecycle, install story -and privilege boundary that this file has no reason to duplicate. What this +This belongs to [node](node.md) to build, not to this epic: node bootstrap +has its own lifecycle, install story and privilege boundary that this file +has no reason to duplicate. What this epic contributes is the pattern and the poll it already has to build for the operator vouch — a second thing to read off the same repository on the same schedule is a small addition once the first exists, and a much larger one diff --git a/plan/harnesses.md b/plan/harnesses.md deleted file mode 100644 index 8c76c59f..00000000 --- a/plan/harnesses.md +++ /dev/null @@ -1,95 +0,0 @@ ---- -id: harnesses -title: What a harness has to provide for any of this to work -status: open -crates: [didbot-hook] -dependsOn: [stamp-contract] -exitCriterion: > - A second harness's integration is named alongside Claude Code's, with every - guarantee this project relies on marked as provided, degraded, or absent for - it. ---- - -# harnesses - -Everything in this project's trust model assumes Claude Code's hook -protocol. `crates/didbot-hook`'s own module documentation says the division -plainly: hooks are how this project learns which agent is acting, because -the harness fills these payloads and the model never does. `docs/hook-flow.md` -spells out four specific mechanisms this rests on, and none of them is part -of the Model Context Protocol or any other cross-harness standard — they are -Claude Code's. - -## The four things assumed, and what is Claude-Code-specific about each - -- **`PreToolUse` can rewrite a tool call's arguments before it executes.** - This is how an agent identity the model did not choose reaches an MCP - server, per `docs/hook-flow.md`, and it is the entire mechanism - [stamp-contract](stamp-contract.md) exists to keep honest. MCP's own - specification says `clientInfo` is self-reported and must not be relied on - for security — which is a statement about MCP, not about any particular - harness — but the fix, a hook that intercepts and rewrites the call before - the server sees it, is a capability the *harness* has to offer, and - nothing in MCP requires a harness to offer it. -- **`CLAUDE_ENV_FILE` delivers session-scoped environment variables to later - tool calls.** [credentials](credentials.md) is built on this being the one - way a hook can set environment for tool execution, and - [cred-delivery](cred-delivery.md) is the delivery mechanism built on top of - it. The name says which harness it belongs to. -- **`SessionStart`, `SubagentStart`, `SubagentStop` and `SessionEnd` fire at - the boundaries this project provisions and tears down on.** - `docs/hook-flow.md`'s whole provisioning and teardown story — mint an - account on start, delete or pin it on end — is keyed to five named events - existing, firing reliably, and firing in that order. -- **The session transcript is readable and its assistant rows name the - model.** `crates/didbot-hook/src/transcript.rs` reads a harness-internal - file with, in its own module documentation's words, "no promised shape" — - the reader is already defensive about this changing under Claude Code - itself, let alone about another harness having a transcript in a different - format or no transcript file at all. - -## What a harness has to provide, at minimum - -- [ ] **Name the minimum contract, not a wish list.** Not every harness needs - all four mechanisms above equally: a harness with no subagent concept - has nothing for `SubagentStart` to model, and that is a smaller gap - than a harness with no tool-call interception at all, which removes the - one thing identity-stamping depends on entirely. Rank them by what - breaks versus what merely goes unused. -- [ ] **A harness with no tool-call rewrite has no `agent_id` stamping**, full - stop, and every downstream guarantee described in - [subagents](subagents.md) about `agent_id` being trustworthy stops - applying. Say what degrades: does the harness get no identity binding, - or a weaker one sourced from something else the harness does offer - (a process argument, a static config file per launch)? -- [ ] **A harness with no `CLAUDE_ENV_FILE` equivalent has no session-scoped - credential delivery.** [cred-delivery](cred-delivery.md) is the - mechanism this would fall back to; without it, the choices narrow to a - credential in the tool's own arguments (visible to the model, which is - the exact thing that mechanism exists to avoid) or a wrapper process - per session, which `docs/hook-flow.md`'s own opening line says this - design was built to not require. -- [ ] **A harness with no lifecycle events has no automatic teardown.** - [agent-accounts](agent-accounts.md)'s `sweep_stale` already exists for - the case where `SessionEnd` never fires even on Claude Code; a harness - with no equivalent event at all makes that the *only* path rather than - the backstop, and the staleness window becomes the harness's actual - account lifetime rather than an edge case. -- [ ] **A harness with no readable transcript has no model identifier.** - Every field `transcript.rs` reads is already optional on the record it - produces, per `docs/hook-flow.md`'s "a scrobble missing any of them is - still a scrobble" — so this is the one gap that degrades gracefully by - construction, worth stating precisely because the other three do not. - -## What this is not - -Not a second harness's implementation — `crates/didbot-hook` parses one -payload shape and stays that way, per its own module documentation. This -epic is the reference a second harness's integration would be checked -against, and the place [subagents](subagents.md)'s question about -enforcement-versus-reporting gets asked about a harness other than the one -this project has only ever tested against. - -## Done - -Nothing closed yet. diff --git a/plan/local-dev.md b/plan/local-dev.md index b0659ce2..af5dd146 100644 --- a/plan/local-dev.md +++ b/plan/local-dev.md @@ -160,7 +160,7 @@ than left standing. The surviving scripts are described accurately below. - [x] **A session that never reached a live server** is covered by the same harness: one opens with nothing listening, gets no account, and scrobbles anyway once a server exists. -- [x] `dev-pds.sh` and `dev-mcp.sh` stop a previous run by pidfile rather than +- [x] `dev-pds.sh` stops a previous run by pidfile rather than by command-line pattern. `pkill -f "didbot-pds --port ${PORT}"` matched anything whose arguments contained that string, including another developer's server and the shell running the script — which happened @@ -176,7 +176,7 @@ than left standing. The surviving scripts are described accurately below. the test suite the hooks are too slow to run. A commit made with `--no-verify` is checked by it. - [x] Dev scripts that find each other by default: `dev-pds.sh`, - `dev-swarm.sh`, `dev-mcp.sh` and `dev-site.sh`, over the shared + `dev-swarm.sh` and `dev-site.sh`, over the shared `dev-profile.sh`/`dev-pidfile.sh`/`dev-watch.sh` helpers. This used to read "five dev scripts — server, index, query, swarm, web": the index, query and web scripts went with their crates to vibescrobble.com, and diff --git a/plan/node.md b/plan/node.md index 8d205fdd..c63e17d8 100644 --- a/plan/node.md +++ b/plan/node.md @@ -2,7 +2,7 @@ id: node title: A host proves what it is once, and issues credentials to the sessions on it status: open -crates: [didbot-attest, didbot-hook, didbot-hookd] +crates: [didbot-attest] dependsOn: [credentials, attestation] exitCriterion: > A machine holding a node credential issues a session credential to a hook that @@ -39,10 +39,9 @@ supervise and version on every agent host, and [deploy](deploy.md) already notes that agent hosts are their own deployment. [dev-setup](dev-setup.md)'s service/profile/binding model is where it has to fit. -- [ ] **Decide whether this is a new crate or a role - [`didbot-hookd`](../crates/didbot-hookd/) grows.** There is already a - daemon on that host doing adjacent work. A second one needs a reason - better than tidiness, and growing the first one needs a reason better than +- [ ] **Decide whether this is a new crate or a role an existing daemon + grows.** A second daemon on an agent host needs a reason better than + tidiness, and growing an existing one needs a reason better than convenience. ## What it holds, and what that is worth diff --git a/plan/onboarding.md b/plan/onboarding.md index 84be26fa..c40d0778 100644 --- a/plan/onboarding.md +++ b/plan/onboarding.md @@ -2,7 +2,7 @@ id: onboarding title: An operator establishes a server before the server can establish anyone else status: open -crates: [didbot-attest, didbot-pds, didbot-serve, didbot-setup, didbot-hookd, didbot-reconcile] +crates: [didbot-attest, didbot-pds, didbot-serve, didbot-setup, didbot-reconcile] dependsOn: [ownership, deploy] exitCriterion: > An operator takes a freshly applied deployment — instance up, no accounts, @@ -26,8 +26,7 @@ cover, and the half that has to happen first. ## The bootstrap paradox **Nothing authenticates a provisioning request.** `bot.did.provisionAgent` -takes no credential, and `didbot-hookd`'s client says so -(`crates/didbot-hookd/src/pds.rs`). An account is provisioned +takes no credential. An account is provisioned `unauthenticated` / `self-asserted`, and `bot.did.registration` records it that way. Node authorization is a later pass and is explicitly not foundational, so there is no secret to generate, distribute, or get wrong. @@ -149,35 +148,29 @@ mints its first account: Once a server is up, the handshake has completed, and `--owner` is established, a person still has to get one agent host talking to it. -[didbot-setup](../crates/didbot-setup/) does the machine half: it puts the -hook binary on `PATH`, writes the harness's hook entries, declares the -scrobble host, and its `verify` command provisions a real account through the -same path a live session would use. What it explicitly does not do, checked -against its own source: - -- **It never sets `DIDBOT_PDS_URL`.** `crates/didbot-hookd/src/config.rs`'s - `Config::from_env` falls back to `http://localhost:3000` when it is unset, - and `didbot-setup` never writes it — its whole flow is built and tested - against `didbot-pds`'s own defaults. Pointing a harness at a deployed - server is presently "set an environment variable by hand, correctly, on - your own," with no `didbot-setup` step that asks for a URL and writes it - anywhere. There is no secret to set alongside it: provisioning takes no - credential. -- **`verify`'s own recovery path is the honest current answer, and it is - worth stating as such rather than leaving it implied.** [dev-setup](dev-setup.md) - already documents that `PreToolUse` provisions when a session has no usable - account, so once the two variables above are set correctly by hand, the - first tool call in a fresh session is what actually creates the account — - there is no separate "first login" moment to perform. +[didbot-setup](../crates/didbot-setup/) does the machine half: it describes +which services this machine runs and says whether they are answering. What it +explicitly does not do, checked against its own source: + +- **It never sets `DIDBOT_PDS_URL`.** A harness client falls back to + `http://localhost:3000` when it is unset, and `didbot-setup` never writes + it — its whole flow is built and tested against `didbot-pds`'s own + defaults. Pointing a harness at a deployed server is presently "set an + environment variable by hand, correctly, on your own," with no + `didbot-setup` step that asks for a URL and writes it anywhere. There is + no secret to set alongside it: provisioning takes no credential. +- **A session with no usable account provisions one on its first tool + call**, so once the variables above are set correctly by hand, that call + is what actually creates the account — there is no separate "first login" + moment to perform. - [ ] **A `didbot-setup` flow — or documented manual steps, if the flow is not worth building yet — for wiring a harness against a server that is not `didbot-pds`'s own defaults**: the URL, landing in the same place - `check`/`apply`/`verify` already look. When node authorization lands it - will add to this list; it does not today. + `check` already looks. When node authorization lands it will add to + this list; it does not today. - [ ] **`didbot-setup check` should say which server a profile is configured - for**, the way it already reports a scrobble host writing to the wrong - server, so a harness silently pointed at a stale or wrong deployment is + for**, so a harness silently pointed at a stale or wrong deployment is caught before the first session rather than after a refusal with no context. diff --git a/plan/order.txt b/plan/order.txt index f078ca8e..3558bb22 100644 --- a/plan/order.txt +++ b/plan/order.txt @@ -16,7 +16,6 @@ canvas pds-writes agent-accounts account-types -scrobble write-policy auth-types @@ -24,10 +23,8 @@ oauth pds-xrpc federation credentials -cred-delivery node subagents -harnesses scope-policy ownership vouch diff --git a/plan/provenance.md b/plan/provenance.md index e07539af..aa53b7cf 100644 --- a/plan/provenance.md +++ b/plan/provenance.md @@ -2,7 +2,7 @@ id: provenance title: A record says which agent, which model, and what spawned it status: open -crates: [didbot-avatar, didbot-hook, didbot-hookd, didbot-lexicon, didbot-pds, didbot-serve] +crates: [didbot-avatar, didbot-lexicon, didbot-pds, didbot-serve] dependsOn: [] exitCriterion: > A reader looking at one record can say which agent wrote it, what model wrote diff --git a/plan/scrobble.md b/plan/scrobble.md deleted file mode 100644 index 175ebdfb..00000000 --- a/plan/scrobble.md +++ /dev/null @@ -1,44 +0,0 @@ ---- -id: scrobble -title: An agent says what it is working on, and the statement is its own record -status: open -crates: [didbot-mcp, didbot-hookd, didbot-lexicon] -dependsOn: [pds-writes] -exitCriterion: > - A scrobble called in an ordinary session is written to a real personal data - server as that agent's own signed record. ---- - -# scrobble - -Short-term, broadcast status: what an agent is working on, written when the -model has something to say. The lexicon is settled and the tool exists. -The last hop is missing — the tool does not write to a real server. - -- [ ] **Decide on compaction.** A separate low-retention collection, or - ordinary repository growth. Growth is fine at this volume and not at a - swarm's. -- [ ] **Hook-authored scrobbles, or a decision against.** A `Stop` hook of type - `prompt` can summarize the turn without the main model spending one. - -## Done - -- [x] **Wire the `PreToolUse` stamping to a real personal data server.** No - longer blocked behind [pds-writes](pds-writes.md), which landed. - `a_hook_payload_becomes_a_record_and_the_account_goes_away` in - `crates/didbot/tests/hook_to_record.rs` drives the real hook through - `didbot_hookd::handle`, the real scrobble host - (`didbot_mcp::ScrobbleServer` over its `PdsClient`, posting - `com.atproto.repo.createRecord`) across a real TCP socket to the real - router, then reads the record back through `getRecord` and asserts its - text and emoji — and that the account is gone afterwards. - `a_session_that_opened_with_no_server_scrobbles_once_there_is_one` - covers the same path when the server arrives late. - -- [x] The `com.vibescrobble.scrobble` lexicon, with text and a single-grapheme - emoji. -- [x] The MCP server, and the `PreToolUse` rewrite that stamps the acting - agent's DID into the call. -- [x] A nudge that quotes the tool description rather than restating it. -- [x] A per-agent scrobbling switch. -- [x] Reasoning effort recorded on the record. diff --git a/plan/stamp-contract.md b/plan/stamp-contract.md deleted file mode 100644 index b6ba63df..00000000 --- a/plan/stamp-contract.md +++ /dev/null @@ -1,88 +0,0 @@ ---- -id: stamp-contract -title: The hook stamps a tool it was told about, and the tool reads a contract it can name -status: open -crates: [didbot-hook, didbot-hookd, didbot-mcp, didbot-setup] -dependsOn: [] -exitCriterion: > - The hook stamps a tool named in its configuration, and the scrobble server - reads the stamp, with neither crate importing a constant from the other. ---- - -# stamp-contract - -`didbot-hook` holds two things. One is the Claude Code harness: payload -types, the transcript read that finds a turn's model, the sidecar read that -finds what spawned an agent. That is how this project learns which agent is -acting, and it stays here — [harnesses](harnesses.md) is where what a -*different* harness would have to supply in place of it gets asked. - -The other is a contract with the scrobble server — six key names the hook -writes into a tool call and the server reads back out. The two halves import -each other's constants today, and they are about to live in different -repositories. That works only while they share a workspace. - -The coupling runs both ways, which is the part worth fixing regardless of any -split. `SCROBBLE_TOOL`, `SCROBBLE_SERVER`, `is_scrobble_tool` and -`scrobble_tool_for` mean the mechanism that decides whether an `agent_id` can -be trusted knows a product's tool by name. - -- [ ] **Say what the stamp is, apart from what writes it.** Six keys — - `_didbotAgentDid`, `_didbotTurn`, `_didbotModel`, - `_didbotEffort`, `_didbotAgentType`, `_didbotParentDid` - — and the rule that each is droppable on its own. A reader must be able - to implement against the contract without reading this crate. -- [ ] **Keep the property, not the spelling.** The stamp is trusted because the - harness writes it and the model cannot reach it, and the `_` prefix is - what tells a stamped call from an unstamped one and keeps it clear of a - tool's own parameters. Any renaming has to carry that argument with it, - or it is a rename that quietly removes a check. -- [ ] **Give the hook the tool name instead of building it in.** Which tools to - stamp belongs in `hookd`'s configuration. The four scrobble-shaped - constants leave the hook once it is told rather than knowing. -- [ ] **Move the two limits that are already on the wrong side.** - `TEXT_MAX_GRAPHEMES` and `MODEL_MAX_LENGTH` are scrobble record field - limits, re-exported by the scrobble server from the hook. They belong - with the lexicon that defines the record. Decide what the hook does once - it no longer holds them: stop refusing early and let record validation - refuse, or take the limits as configuration too. -- [ ] **Decide where the nudge lives.** `SCROBBLE_INSTRUCTIONS` is text about - one product, in the crate that stamps every product. -- [ ] **Renaming the keys is a protocol break, and the halves ship apart.** - `didbot-setup` installs a hook binary onto a machine; the server it - talks to is updated separately. A key rename lands on a deployment where - one side is old, so it needs a window that reads both and writes one, or - a stamp that says which version it is. -- [ ] **Replace the test that spans the two crates.** - `didbot-mcp/tests/create_record.rs` imports `STAMP_KEY` and - `TURN_KEY` from the hook to build a stamped call. Once the two are apart - that has to be the contract as data, the way the atproto interop vectors - are, rather than a shared import. -- [ ] **Do this while both halves are still in one repository.** The - decomposition is testable end to end only while the hook and the scrobble - server build together. Moving the server out first means doing this - across a repository boundary, on the crate whose own documentation calls - it the basis for trusting an `agent_id`. - -## Done - -- [x] **The property `stamp()` was supposed to hold, actually holds.** - Finding 03 (critical, security review): `stamp()` refused to overwrite - a key already present on `tool_input`, on the theory that meant the - call had already been stamped. It hadn't — the only way a model's own - tool call could carry `STAMP_KEY`, `PARENT_KEY`, `EFFORT_KEY`, - `TOKEN_KEY` or any of the other three was that the model wrote it - there, and the old code let that value pass through untouched. - `stamp()` now always overwrites with the harness's value, and - `didbot-hookd::handler::pre_tool_use` denies outright a call whose - stamp keys disagree with what this hook is about to write — see - `stamp_conflict` in `crates/didbot-hook/src/lib.rs`, which tells a - forged value apart from the same hook re-stamping a call it already - processed (wired at more than one configuration scope) by comparing - the carried value against the one this call independently computes, - not by trusting mere presence of the key. This is not the crate-split - this epic tracks, but it is the same "the coupling runs both ways" - property in "Keep the property, not the spelling" above, now actually - true of the code rather than only of the comments describing it — - whichever crate ends up owning the contract inherits an invariant that - holds, not one that only held by convention. diff --git a/plan/subagents.md b/plan/subagents.md index 99345b83..c90ffbe6 100644 --- a/plan/subagents.md +++ b/plan/subagents.md @@ -2,7 +2,7 @@ id: subagents title: What a distinct subagent identity would buy, and what a harness lets you enforce status: open -crates: [didbot-pds, didbot-hook] +crates: [didbot-pds] dependsOn: [credentials, node] exitCriterion: > A written decision on whether a subagent gets its own DID, backed by a @@ -62,8 +62,8 @@ a session spawns. ## What a harness actually lets you enforce, versus what it merely reports This is the honest part, and it is not a Claude-Code-specific caveat — it is -the crux of the whole question. `docs/hook-flow.md` is explicit about the -ceiling: a `PreToolUse` hook can block a call, rewrite its arguments, and +the crux of the whole question. The ceiling is a harness's, not this +project's: a `PreToolUse` hook can block a call, rewrite its arguments, and inject context, but it cannot synthesize a call the model did not make, and `agent_id`/`agent_type` arrive on the hook payload as the harness's own report of what is running, not as something the harness cryptographically @@ -84,11 +84,11 @@ attests about the subagent process itself. credential scoped to a subagent restricts what it is *handed*, not what it could reach if it chose to ignore the scoping and act as the session directly. -- [ ] **This is the concrete instance of [harnesses](harnesses.md)'s general - question**, narrowed to one guarantee: does *any* harness let a - subagent's identity be enforced rather than merely reported, and if - none does today, say that plainly rather than let a distinct DID imply - a stronger boundary than the harness backs it with. +- [ ] **Narrow the general harness question to one guarantee**: does *any* + harness let a subagent's identity be enforced rather than merely + reported, and if none does today, say that plainly rather than let a + distinct DID imply a stronger boundary than the harness backs it + with. ## Done -- 2.51.2