From ec779f805f178b052dfdafc881c08370158143fe Mon Sep 17 00:00:00 2001 From: "@permadeath.com" Date: Wed, 2 Sep 2026 15:37:00 -0400 Subject: [PATCH] docs(onboarding): say what the bootstrap actually is Provisioning takes no credential, so there is no attestation secret to generate, distribute or check. The paradox is now only that a fresh deployment must reach a state its operator can claim. Co-Authored-By: Claude Opus 5 (1M context) --- plan/onboarding.md | 153 +++++++++++++++++---------------------------- 1 file changed, 56 insertions(+), 97 deletions(-) diff --git a/plan/onboarding.md b/plan/onboarding.md index bf6d42f6..06806a12 100644 --- a/plan/onboarding.md +++ b/plan/onboarding.md @@ -25,79 +25,45 @@ cover, and the half that has to happen first. ## The bootstrap paradox -`didbot-attest` ships one backend, -[`NodeCredentialBackend`](../crates/didbot-attest/src/node_credential.rs). The -rest of this section is written against a shared-secret backend that no longer -exists, and does not describe the server as it stands — see the last item -below. - -One half of this is already enforced: `didbot-dev` refuses to start against a -`--zone` outside `.localhost` while `--secret` is still the published -default (`crates/didbot-serve/src/bin/didbot-dev.rs`, the check right before -`assemble`). That closes the loudest version of the mistake — a production -zone quietly attesting the whole internet — by making it a startup error -instead. - -It is not the whole paradox, because the check only catches the one field it -can see: - -- **Nothing generates the real secret or gets it onto the host.** - `infra/templates/user_data.sh.tftpl` fetches - `$SSM_PREFIX/attestation-secret` from SSM Parameter Store at boot and fails - the whole boot (`aws ssm get-parameter` with no `||` fallback) if that - parameter does not exist — which is the right failure, but nothing in - `infra/` creates the parameter. An operator who has never read this file - has no way to know that step exists, and no documented command for doing - it. -- **The attestation secret is the only secret this gap applies to.** There - is no operator secret: [auth-types](auth-types.md) records that a server - learns which DID operates it by reading `bot.did.operator` out of the - operator's own repository ([handshake](handshake.md)), and holds nothing - presentable back to it. So there is one parameter to get onto the host, - not two. `didbot-dev` - itself only `warn!`s once, to a log an operator may not be watching this - early. -- **The secret that gates provisioning has to reach every agent host too, and - nothing carries it there.** `SharedSecretBackend`'s claim is symmetric: the - same string the server checks is the string `didbot-hookd` signs with - (`crates/didbot-hookd/src/config.rs`'s `DIDBOT_SHARED_SECRET`, defaulting - to the identical placeholder). [didbot-setup](../crates/didbot-setup/) - wires the hook binary, the settings.json entries and the scrobble host - declaration onto a machine — it never touches this variable. A real - deployment therefore needs a second, undocumented distribution step: the - same value that came out of SSM for the server has to land in every agent - host's environment, by some channel nobody here specifies, or every - attestation claim that host produces is checked against a secret the - server does not hold and is refused. - -So: the startup check is sufficient for the one failure mode it targets — -a real zone holding the checked-in string — and insufficient for the -paradox as a whole. The actual answer has three parts, and none of them -exist as instructions today: - -- [ ] **Write the operator's first commands.** Generate the attestation - secret and write it to `$SSM_PREFIX/attestation-secret`, before the - instance's first boot, or before the next one if it already ran - without it. A script or a documented `aws ssm put-parameter` — either - is fine, but right now there is neither. -- [ ] **Write the operator's handshake command.** The bootstrap that used to - be "paste a secret the server compares" is now: write a - `bot.did.operator` record into the operator's own repository, at the - rkey the server is watching, and wait for `crate::ownership_poll` to - find it. `didbot-claim` has the local half; what onboarding still owes - is the sequence a first-time operator follows against a real - deployment, and how they tell whether the server has seen it yet. -- [ ] **Rewrite the bootstrap paradox against the credential backend.** The - section above, and the two items that follow it, are written against a - shared secret this workspace no longer has. Whether the paradox - survives the change is the question; the prose cannot answer it as it - stands. -- [ ] **Say how the shared secret reaches an agent host**, or replace the - question: this is the same gap [node](node.md) is designed to close - (a node credential issued out of band, rather than a secret copied by - hand), and until that lands, onboarding's honest answer is "copy the - value some way you control, into `DIDBOT_SHARED_SECRET`, and treat it - exactly as seriously as the server's own copy." +**Nothing authenticates a provisioning request.** `bot.did.provisionAgent` +takes no credential, and `didbot-hookd`'s client says so +(`crates/didbot-hookd/src/pds.rs`). An account is provisioned +`unauthenticated` / `self-asserted`, and `bot.did.registration` records it +that way. Node authorization is a later pass and is explicitly not +foundational, so there is no secret to generate, distribute, or get wrong. + +`didbot-attest` still ships two backends — +[`NodeCredentialBackend`](../crates/didbot-attest/src/node_credential.rs) and +[`AwsInstanceIdentityBackend`](../crates/didbot-attest/src/aws/mod.rs) — and +`didbot-dev` wires neither. They are the vocabulary a later pass will use, not +something a deployment configures today. + +So the paradox is not what it was. It was: the secret that lets an operator +act on a fresh deployment must reach the host before the deployment can be +acted on. There is no such secret now, and the question that replaces it is +narrower — a fresh deployment has to reach a state its operator can claim, +using only what the deployment already has. + +That question is answered elsewhere and this epic should not re-answer it. +The server mints its own identity through an administrative path that never +travels the xrpc write path, so its `did:json` and signing key exist before +it serves anything; the operator then claims it from their own machine with +`didbot-claim`, and the server observes that claim through its own poll. See +[handshake](handshake.md) and `docs/server-lifecycle.md`. + +What onboarding still owes is the sequence, not the mechanism: + +- [ ] **Write the operator's first commands**, in order, against a real + deployment: what to run, on which machine, and what each one is waiting + for. `didbot-claim` has the local half and `didbot-claim --check` has + the pre-flight; nothing walks a first-time operator through them. +- [ ] **Say how an operator tells whether the server has seen the claim yet.** + The poll is the only thing that moves the server to `claimed`, and its + interval is the delay an operator experiences as "nothing happened". +- [ ] **Say what a deployment does before it has any operator policy.** + The policy epic puts a pre-flight between provisioning and + `claimed`, so an instance is never online and unpolicied. The commands + for that half belong here too. ## What "first log in" means here @@ -111,8 +77,7 @@ frame and the epic should stop implying one. What exists instead: their *own* repository and this server reads it (`crates/didbot-serve/src/ownership_poll.rs`). Nothing is presented to the server and nothing is compared by it. There is no operator - credential: the shared secret this section used to describe is gone, and - with it every HTTP surface whose only gate it was — see + credential, and no HTTP surface is gated on one — see [auth-types](auth-types.md). - Local access to the box, which is what the e-stop admin socket's `0700`/`0600` permissions amount to and the whole of that socket's @@ -190,20 +155,14 @@ scrobble host, and its `verify` command provisions a real account through the same path a live session would use. What it explicitly does not do, checked against its own source: -- **It never sets `DIDBOT_SHARED_SECRET` or `DIDBOT_PDS_URL` to anything but - the development defaults.** `crates/didbot-hookd/src/config.rs`'s - `Config::from_env` falls back to `http://localhost:3000` and the checked-in - placeholder secret whenever those variables are unset, and `didbot-setup` - never writes them — its whole flow is built and tested against - `didbot-dev`'s own defaults. Pointing a harness at a deployed server is - presently "set two environment variables by hand, correctly, on your own," - with no `didbot-setup` step that asks for a URL and a secret and writes - them anywhere. -- **It never validates the node id it will attest as is one the deployment - allowlists.** `DIDBOT_NODE_ID` defaults to `dev-node`, which is the one id - `didbot-dev` itself allowlists; a real `SharedSecretBackend` is constructed - with whatever node ids the operator chose, and nothing checks the two - agree before the first provisioning attempt fails. +- **It never sets `DIDBOT_PDS_URL`.** `crates/didbot-hookd/src/config.rs`'s + `Config::from_env` falls back to `http://localhost:3000` when it is unset, + and `didbot-setup` never writes it — its whole flow is built and tested + against `didbot-dev`'s own defaults. Pointing a harness at a deployed + server is presently "set an environment variable by hand, correctly, on + your own," with no `didbot-setup` step that asks for a URL and writes it + anywhere. There is no secret to set alongside it: provisioning takes no + credential. - **`verify`'s own recovery path is the honest current answer, and it is worth stating as such rather than leaving it implied.** [dev-setup](dev-setup.md) already documents that `PreToolUse` provisions when a session has no usable @@ -213,14 +172,14 @@ against its own source: - [ ] **A `didbot-setup` flow — or documented manual steps, if the flow is not worth building yet — for wiring a harness against a server that is - not `didbot-dev`'s own defaults**: the URL, the shared secret, and a - node id the deployment actually allowlists, landing in the same place - `check`/`apply`/`verify` already look. -- [ ] **`didbot-setup check` should say which server and which secret a - profile is configured for**, the way it already reports a scrobble host - writing to the wrong server, so a harness silently pointed at a stale - or wrong deployment is caught before the first session rather than - after a provisioning refusal with no context. + not `didbot-dev`'s own defaults**: the URL, landing in the same place + `check`/`apply`/`verify` already look. When node authorization lands it + will add to this list; it does not today. +- [ ] **`didbot-setup check` should say which server a profile is configured + for**, the way it already reports a scrobble host writing to the wrong + server, so a harness silently pointed at a stale or wrong deployment is + caught before the first session rather than after a refusal with no + context. ## Failure modes worth naming -- 2.51.2