diff --git a/docs/conformance.md b/docs/conformance.md index 6e86aca1..f9ce0c2d 100644 --- a/docs/conformance.md +++ b/docs/conformance.md @@ -716,7 +716,7 @@ this server serves — does not exercise either. `{"hostname": ..}` body to `{relay}/xrpc/com.atproto.sync.requestCrawl` and `{relay}/xrpc/com.atproto.sync.notifyOfUpdate` respectively. Neither is called automatically. The only caller is `didbot_serve::estop_admin`'s -`ANNOUNCE`/`NOTIFY` commands, reached over the same credentialed unix socket +`ANNOUNCE`/`NOTIFY` commands, reached over the same local unix socket as the e-stop's own `RELEASE` — an operator asks explicitly, and the socket reply reports what the relay said rather than a log line the operator has to go find. Design decisions, each stated once and checked by a test in @@ -725,7 +725,7 @@ to go find. Design decisions, each stated once and checked by a test in | Question | Answer | | --- | --- | | When does either fire? | Only when an operator sends `ANNOUNCE` or `NOTIFY` over the e-stop admin socket. Nothing calls either at startup, on a timer, or on any other schedule. | -| What credential does it need? | The same operator secret `RELEASE` requires, compared in constant time. No credential configured, or the wrong one, is an `ERR` before any relay is contacted. | +| What credential does it need? | None, the same as every other command on this socket. The socket is `0600` inside a directory the server owns and re-tightens to `0700` on every start, so reaching it at all means host access — which is the whole authorization story, and a stronger one than a shared secret readable from the same filesystem. There is no operator credential anywhere in didbot: a server learns which DID operates it by reading `bot.did.operator` out of the operator's own repository, and is never handed anything to check. | | What does the reply distinguish? | Three outcomes, as `OK `: `{"result":"accepted"}`, a relay that answered but refused (`{"result":"refused","status":..,"body":..}`, covering `HostBanned` and anything else a relay might say), and a relay that could not be reached at all (`{"result":"unreachable","error":..}`). An operator who ran the command sees which one happened. | | Which relay? | `[relay].hostname` in `didbot-config`, or `--relay-hostname` directly — never hardcoded. `None` (the default) leaves `ANNOUNCE`/`NOTIFY` with no relay to reach, reported as `ERR`. | | What about `*.localhost`? | `didbot-dev.rs` never wires the commands up for a loopback zone, whatever `--relay-hostname` says — the same `LoopbackDns::accepts(zone.host())` check that already gates `--tls acme` and the placeholder-secret refusal, because no relay can reach a loopback address. | diff --git a/infra/templates/user_data.sh.tftpl b/infra/templates/user_data.sh.tftpl index 95246d26..61250492 100644 --- a/infra/templates/user_data.sh.tftpl +++ b/infra/templates/user_data.sh.tftpl @@ -77,14 +77,6 @@ aws ssm get-parameter \ --query Parameter.Value --output text > /etc/didbot/attestation-secret chmod 600 /etc/didbot/attestation-secret -aws ssm get-parameter \ - --region "$AWS_REGION" \ - --name "$SSM_PREFIX/operator-secret" \ - --with-decryption \ - --query Parameter.Value --output text > /etc/didbot/operator-secret 2>/dev/null \ - || echo "no operator-secret parameter set; the admin surface and e-stop release will refuse every request until one is" -chmod 600 /etc/didbot/operator-secret 2>/dev/null || true - # --- Image pull --------------------------------------------------------------- aws ecr get-login-password --region "$AWS_REGION" \ | docker login --username AWS --password-stdin "${ecr_registry}" @@ -130,7 +122,6 @@ ExecStart=/usr/bin/docker run --rm --name didbot-pds \ --port 443 \ --data /data \ --secret-file /run/didbot/attestation-secret \ - --operator-secret-file /run/didbot/operator-secret \ --tls acme \ --acme-environment ACME_ENVIRONMENT_PLACEHOLDER \ --route53-zone-id ROUTE53_ZONE_ID_PLACEHOLDER diff --git a/plan/agent-accounts.md b/plan/agent-accounts.md index ee1af3cc..44c129b8 100644 --- a/plan/agent-accounts.md +++ b/plan/agent-accounts.md @@ -132,11 +132,15 @@ Two hostnames per account: the DID's, and the handle's, which needs reaps any other unpinned, abandoned account. - [x] **Hard delete is administrator-only and confirmed.** `Registry::hard_delete` refuses unless `confirm_did` names the account exactly, logs at `WARN`, - and records the operator on the ledger's `Deprovisioned` entry — see - `bot.did.hardDeleteAgent`, `Credential::Operator` only, no self-service - path. `Registry::freeze`/`unfreeze`/`soft_delete` stay - `Credential::AgentOrOperator`, the same posture the pre-existing - `deleteAgent` already had. + and records the operator on the ledger's `Deprovisioned` entry. It has + no HTTP route: administrator-only is exactly what this server cannot + authenticate, since there is no operator credential + ([auth-types](auth-types.md)) and self-service is the wrong answer for + the one operation that frees a burned name. The route returns when + [oauth](oauth.md)'s owner sign-in gives it a caller it can check. + `Registry::delete`/`freeze`/`unfreeze`/`soft_delete` and + `setAgentPinned` are `Credential::AgentSelf` — an agent acting on + itself, and nothing else. - [x] **Names are burned permanently by soft delete, freed only by hard delete.** `Reservation::FormerAgent` extends the existing reservation mechanism (`Reservation::ZoneApex`/`Operational`) rather than adding a diff --git a/plan/auth-types.md b/plan/auth-types.md index 3dbaf53c..888a2e68 100644 --- a/plan/auth-types.md +++ b/plan/auth-types.md @@ -110,17 +110,18 @@ you are reading; see `crates/didbot-serve/src/auth.rs`'s module documentation fo taxonomy this section only summarises. **Mutations — `deleteAgent`, `setAgentPinned`, `setAgentScrobbling` — are -always credentialed, never toggleable.** `Credential::AgentOrOperator`: -either the acting agent's own write credential, checked against the account -the request names, or an operator credential, which may name any account. -`didbot-hookd` calls all three automatically with no operator in the loop — -the same shape that argued (wrongly) for `Public` before — but by the time -one of these fires the account already has its own write credential, and -that is exactly the caller these routes are reasonable to ask one from. An -agent may end or toggle *itself*; only an operator may act on another -account. There is no netizen argument for letting a stranger delete an -account, so this stays hard-locked regardless of the disclosure decision -below. +always credentialed, never toggleable.** `Credential::AgentSelf`: the acting +agent's own write credential, checked against the account the request names. +`didbot-hookd` calls all three automatically — the same shape that argued +(wrongly) for `Public` before — but by the time one of these fires the +account already has its own write credential, and that is exactly the caller +these routes are reasonable to ask one from. An agent may end or toggle +*itself*, and that is the whole list: acting on another account would need a +credential proving this deployment's operator, and there is deliberately no +such credential. `bot.did.hardDeleteAgent`, which must not be self-service, +consequently has no route at all — see [agent-accounts](agent-accounts.md). +There is no netizen argument for letting a stranger delete an account, so +this stays hard-locked regardless of the disclosure decision below. ### Publish by default: the accountability argument, and why it is a principle rather than a note about four routes @@ -225,16 +226,22 @@ three reasons laid out together. caller-facing methods — an operator sets the one password this server tracks per account directly, a smaller surface than the full lexicon set. -- [x] **Say what an operator credential is and keep it separate.** - `auth::operator_secret_matches`: a bearer scheme of its own - (`Operator `), compared against a configured secret in constant - time (`subtle::ConstantTimeEq`), independent of `SessionAuth` and of - anything OAuth would issue. Now wired to HTTP: `auth::require_self_or_operator` - for the account-admin mutations acting on another account, and - `auth::require_disclosure` for a closed disclosure route — as well as - [e-stop](e-stop.md)'s admin socket, which reaches - `operator_secret_matches` directly outside HTTP for the same secret and - the same timing property. +- [x] **Say what an operator credential is: there is not one.** No scheme + this server accepts authenticates an operator, and none should. A + server learns which DID operates it by reading `bot.did.operator` out + of the operator's own repository — see [handshake](handshake.md) — a + mechanism whose whole point is that no operator credential ever + reaches a PDS. A header this server could compare would be one it + holds, which is the thing the handshake makes unnecessary. + `auth::ROUTE_CREDENTIALS` carries the consequences: every + account-lifecycle mutation is self-service (`auth::require_self`, an + agent token naming the account the request names), a closed + disclosure route is closed to everyone, and + [e-stop](e-stop.md)'s admin socket rests on its own `0700`/`0600` + permissions rather than on anything presented over it. Proving + control of the operator DID *to* this server is + [oauth](oauth.md)'s owner sign-in; until it exists, a surface that + would need one is not built rather than gated on a stand-in. - [x] **State what stays open to everyone.** Decided and written down in `auth::ROUTE_CREDENTIALS`'s own comments and the module doc's "Why every public route is public": three separate reasons, not one — @@ -248,7 +255,7 @@ three reasons laid out together. surface and the `bot.did.*` mutations are no longer part of this list — see above. - [x] **Credential the account-admin mutations.** `deleteAgent`, - `setAgentPinned` and `setAgentScrobbling` are `Credential::AgentOrOperator` + `setAgentPinned` and `setAgentScrobbling` are `Credential::AgentSelf` — see "The account-admin surface, and the mistake `Credential::Public` let happen" above. `provisionAgent` is relabelled `Credential::Attested` so its attestation gate is visible in the table rather than shared with @@ -256,9 +263,9 @@ three reasons laid out together. - [x] **Decide the account-admin reads, and build the toggle.** `listAgents`, `listAgentLedgers`, `getAgentLedger` and `stats` are `Credential::Disclosure` — public by default, closeable per route by - `auth::Disclosure`, an operator credential still reaching a closed - route, and a closed route answering `DisclosureDisabled` rather than a - 404 or an empty list. See "Publish by default" above for why the + `auth::Disclosure`, and a closed route answering `DisclosureDisabled` + — for every caller alike, since no credential reaches past the toggle + — rather than a 404 or an empty list. See "Publish by default" above for why the default is public and why the toggle exists anyway. - [x] **The failure shapes, and the headers that go with them.** `AuthenticationRequired`, `ExpiredToken`, `InvalidToken` and the new diff --git a/plan/aws-deploy.md b/plan/aws-deploy.md index f9e75c72..9a1d6701 100644 --- a/plan/aws-deploy.md +++ b/plan/aws-deploy.md @@ -34,8 +34,7 @@ invent a second way of doing the same thing. **The module and the binary must stay in step.** `infra/templates/user_data.sh.tftpl` writes a systemd unit whose `ExecStart` passes specific flags — `--zone`, `--owner`, `--port`, `--data`, -`--secret-file`, `--operator-secret-file` today, confirmed by reading that -template. A module version and a server (container image) version are +`--secret-file` today, confirmed by reading that template. A module version and a server (container image) version are therefore coupled: a flag renamed, dropped, or given new meaning between two image versions is invisible to Terraform, and an operator who bumps the module without also picking a compatible image gets a server that will not @@ -67,8 +66,8 @@ direction. already states the current discipline: secret *values* are written to SSM out of band, by a human or a separate process this configuration does not run, so `terraform.tfstate` never carries key material. A one-shot module that -*generates* the attestation secret and the operator secret for the operator — -which is the more "one-shot" reading of the request — breaks that discipline +*generates* the attestation secret for the operator — which is the more +"one-shot" reading of the request — breaks that discipline by construction: anything a Terraform resource creates is a value Terraform's state file holds, in plaintext, in whatever backend that state lives in (local disk by default, S3 with no encryption unless configured, a CI diff --git a/plan/config.md b/plan/config.md index b69bb870..1288b698 100644 --- a/plan/config.md +++ b/plan/config.md @@ -15,11 +15,10 @@ exitCriterion: > # config `didbot-dev` takes `--zone`, `--data`, `--port`, `--owner`, `--secret`, -`--secret-file`, `--operator-secret`, `--operator-secret-file`, -`--blob-quota`, `--max-blob`, `--name-hold`, `--names`, `--names-if-down`, -`--estop-file`, `--estop-socket`, `--close-disclosure`, `--avatar`, `--demo`, -and more — twenty-one flags today, and every epic that touches this binary -adds one. `--secret` already needed a `--secret-file` twin because process +`--secret-file`, `--blob-quota`, `--max-blob`, `--name-hold`, `--names`, +`--names-if-down`, `--estop-file`, `--estop-socket`, `--close-disclosure`, +`--avatar`, `--demo`, and more — nineteen flags today, and every epic that +touches this binary adds one. `--secret` already needed a `--secret-file` twin because process arguments are world-readable in `/proc//cmdline`, which is [deployment](../docs/deployment.md)'s own finding. A flag list that long is not a bootstrap interface any more; it is configuration with no schema, no @@ -90,8 +89,8 @@ is wired to anything that runs; see "What this pass wires" below. lives — the same shape `--secret-file` already established for the command line — and never carries one itself. A key that looks like it wants a literal secret is a schema mistake, not a feature. True of - every wired secret path today (`--secret-file`, - `--operator-secret-file`); still open because `[attestation]`'s own + every wired secret path today (`--secret-file`); still open because + `[attestation]`'s own `secret_file` field is schema only, not yet read by anything. - [ ] **One stated precedence.** Command-line arguments win over the file, which wins over the environment, decided in one place @@ -119,8 +118,8 @@ Four other lines of work are touching `didbot-dev.rs`, `rate_limit.rs`, builds the schema for every section above and wires only a slice that does not collide: `--config` reads a file through `didbot_config::Config::load`, `[blobs]` and `[disclosure]` are applied with `--max-blob`/`--blob-quota`/ -`--close-disclosure` taking precedence, and `--secret-file` / -`--operator-secret-file` are checked for loose permissions. That proves the +`--close-disclosure` taking precedence, and `--secret-file` is checked for +loose permissions. That proves the refusal behaviours end to end — unknown key, malformed value, loose permissions — without touching the bind address, the rate limiter, or either provider crate. @@ -133,7 +132,6 @@ pass: - [ ] `--secret` itself (only the file-permission check for `--secret-file` is wired; the value's own precedence is decided above but not yet enforced in code) -- [ ] `--operator-secret` - [ ] `--names`, `--names-if-down`, `--name-hold` - [ ] `--estop-file`, `--estop-socket` - [ ] `--avatar`, `--demo` (development-only; may never belong in a diff --git a/plan/onboarding.md b/plan/onboarding.md index f19d3816..43399e5b 100644 --- a/plan/onboarding.md +++ b/plan/onboarding.md @@ -50,14 +50,12 @@ can see: `infra/` creates the parameter. An operator who has never read this file has no way to know that step exists, and no documented command for doing it. -- **The operator secret gets a quieter version of the same gap, deliberately - softened in a way that is easy to mistake for handled.** The same script's - fetch of `$SSM_PREFIX/operator-secret` is wrapped in `|| echo ...`: a - missing parameter prints a warning and the boot continues. That is the - correct choice for the check `--zone` performs — a missing operator secret - should not take agents offline — but it means a deployment can run - indefinitely with the entire admin surface and e-stop release refusing - every request, and nothing forces the operator to notice. `didbot-dev` +- **The attestation secret is the only secret this gap applies to.** There + is no operator secret: [auth-types](auth-types.md) records that a server + learns which DID operates it by reading `bot.did.operator` out of the + operator's own repository ([handshake](handshake.md)), and holds nothing + presentable back to it. So there is one parameter to get onto the host, + not two. `didbot-dev` itself only `warn!`s once, to a log an operator may not be watching this early. - **The secret that gates provisioning has to reach every agent host too, and @@ -79,17 +77,17 @@ paradox as a whole. The actual answer has three parts, and none of them exist as instructions today: - [ ] **Write the operator's first commands.** Generate the attestation - secret, write it to `$SSM_PREFIX/attestation-secret`; generate the - operator secret, write it to `$SSM_PREFIX/operator-secret`; both before - the instance's first boot, or before the next one if it already ran - without them. A script or a documented `aws ssm put-parameter` pair — - either is fine, but right now there is neither. -- [ ] **Make the operator-secret gap loud enough to act on**, not just louder - than it is. A deployment that has been running for a week with no - operator secret is indistinguishable, from the outside, from one that - has one — until the day the e-stop is needed and cannot be released. - [alerts](alerts.md)'s channel, or something that runs before it exists, - should be able to say this continuously rather than once at startup. + secret and write it to `$SSM_PREFIX/attestation-secret`, before the + instance's first boot, or before the next one if it already ran + without it. A script or a documented `aws ssm put-parameter` — either + is fine, but right now there is neither. +- [ ] **Write the operator's handshake command.** The bootstrap that used to + be "paste a secret the server compares" is now: write a + `bot.did.operator` record into the operator's own repository, at the + rkey the server is watching, and wait for `crate::ownership_poll` to + find it. `didbot-claim` has the local half; what onboarding still owes + is the sequence a first-time operator follows against a real + deployment, and how they tell whether the server has seen it yet. - [ ] **Say how the shared secret reaches an agent host**, or replace the question: this is the same gap [node](node.md) is designed to close (a node credential issued out of band, rather than a secret copied by @@ -105,14 +103,18 @@ frame and the epic should stop implying one. What exists instead: - `--owner` — a DID string the server writes into every registration record as who is answerable for that account. It is not verified at startup: any string parses. -- The operator secret — a bearer credential, presented as `Authorization: - Operator ` on the HTTP admin routes (`crates/didbot-serve/src/auth.rs`) - and as a plaintext line over the e-stop's unix admin socket - (`crates/didbot-serve/src/estop_admin.rs`). Whoever holds this string can - administer the account surface and release the e-stop. It proves control of - a shared string, nothing about the DID named in `--owner`. - -Neither of those is the handshake [ownership](ownership.md) asks for. +- The `bot.did.operator` handshake — the operator writes a record into + their *own* repository and this server reads it + (`crates/didbot-serve/src/ownership_poll.rs`). Nothing is presented to + the server and nothing is compared by it. There is no operator + credential: the shared secret this section used to describe is gone, and + with it every HTTP surface whose only gate it was — see + [auth-types](auth-types.md). +- Local access to the box, which is what the e-stop admin socket's + `0700`/`0600` permissions amount to and the whole of that socket's + authorization story. + +Neither of the first two is the full handshake [ownership](ownership.md) asks for. Ownership's own checklist says the agent-side owner claim exists (written by the server at provisioning) and the human's vouch of the server is meant to come from the owner's own repository — but ownership's "three statements" @@ -133,11 +135,9 @@ one process argument, forever. ownership's item to build; onboarding is where a first-run check would use it. - [ ] **Write down, plainly, that there is no bidirectional claim yet.** Until - the owner's own repository carries a written vouch - ([policy-dashboard](policy-dashboard.md) or a predecessor of it) and the - status endpoint above exists, "first log in" is `--owner` plus an - operator secret, and that is a weaker claim than the rest of `plan/` - sometimes assumes. Say so somewhere a new operator reads before they + the status endpoint above exists, "first log in" is `--owner` plus a + `bot.did.operator` record this server reads one way, and that is a + weaker claim than the rest of `plan/` sometimes assumes. Say so somewhere a new operator reads before they conclude ownership is already checked because the flag is there. ## Required operator attestation @@ -146,17 +146,20 @@ The owner's framing names this explicitly, and it is worth being precise about what it is not: [attestation](attestation.md) is about the *node an agent runs on* — proving a machine before it is trusted with provisioning. The operator is a different subject. Nothing in this codebase today asks the -operator to prove anything beyond "holds the operator secret," which is -possession of a string chosen at deploy time, not a claim about who they are. +operator to prove anything *to the server over HTTP* at all: the handshake +runs the other way, with the server reading a record the operator wrote in +their own repository. That is a claim about a DID rather than possession of +a string, which is the right direction — but this server never checks a +signature over a challenge it chose, so it is a claim it reads, not one it +verifies interactively. What an operator attestation step would need to establish, before the server mints its first account: -- [ ] **That the operator secret was actually rotated from nothing** — i.e. - the deployment did not skip the step above and is not still silently - refusing every admin request. This is the loud version of the - bootstrap-paradox item above, scoped to "before minting begins" rather - than "eventually." +- [ ] **That the handshake actually completed** — i.e. this server has read + a `bot.did.operator` record naming it, rather than still waiting for + one. This is the loud version of the bootstrap-paradox item above, + scoped to "before minting begins" rather than "eventually." - [ ] **That the `--owner` DID is one the operator controls**, which is a real proof — sign a nonce with the DID's own key, or an equivalent — not merely a string on the command line. Nothing here does this yet; @@ -165,16 +168,17 @@ mints its first account: - [ ] **Decide whether operator attestation is a one-time ceremony or a standing credential.** A node's attestation is spent once, at provisioning, and the node's ongoing claim to be trusted rests on - holding a node credential afterward ([node](node.md)). The operator has - no equivalent today — the operator secret is a permanent bearer token - with no analogous "prove it once, then hold a narrower credential" - step. Whether that gap should close the same way, or the operator's - case is different because there is exactly one of them, is this epic's - to decide, not to assume. + holding a node credential afterward ([node](node.md)). The operator + has no equivalent today, and — since there is deliberately no operator + credential at all — no "hold a narrower credential afterward" half to + build one out of. [oauth](oauth.md)'s owner sign-in is the candidate + mechanism. Whether the operator needs a standing credential at all, or + the read-only handshake plus host access is the whole story, is this + epic's to decide, not to assume. ## The first agent -Once a server is up, the operator secret is real, and `--owner` is +Once a server is up, the handshake has completed, and `--owner` is established, a person still has to get one agent host talking to it. [didbot-setup](../crates/didbot-setup/) does the machine half: it puts the hook binary on `PATH`, writes the harness's hook entries, declares the @@ -239,13 +243,13 @@ default above, restated as what an operator actually sees when it happens. failure mode worth naming is: a startup that passes every local check above but still cannot complete a DNS-01 challenge against Let's Encrypt, for a reason none of those checks can catch in advance. -- [ ] **An operator who has lost the operator secret.** There is no recovery - path. It is an SSM SecureString parameter with no documented rotation - procedure and no second credential that could re-mint it — losing it - means losing the ability to release the e-stop or reach the admin - surface at all, on that deployment, permanently, short of replacing the - parameter and restarting the instance (which itself needs whatever - access created the parameter in the first place). +- [ ] **An operator who has lost access to the DID that operates a server.** + There is nothing to rotate — the server holds no credential — but the + recovery question is real and unanswered: a server watching a + `bot.did.operator` record in a repository its operator can no longer + write has no path to being told about a new one, short of restarting + it with a different `--owner`. Whether that restart is the answer, or + something narrower is, is undecided. ## Done