--- id: capacity title: A deployment stops minting before somebody else's quota does status: open crates: [didbot-pds, didbot-dns, didbot-identity, didbot-serve] dependsOn: [zone-scale, account-types] exitCriterion: > A deployment configured with a maximum account count refuses provisioning once that many accounts hold records, names which zone is full, and an operator watching the dashboard saw the number climbing before it refused anything. --- # capacity Moving from a wildcard record to a real record per host, which [agent-accounts](agent-accounts.md) and [zone-scale](zone-scale.md) already did, is what created this ceiling. A wildcard record covers every hostname underneath it for the cost of one record; the PDS managing its own zone records per agent, so that per-agent DNS-01 issuance and precise withdrawal are possible, means every agent now has a cost measured in a hosted zone's own finite record count. This epic is that decision's bill, not a surprise found later — the tradeoff was worth making and the bill still has to be paid somewhere. ## What an account costs, verified [`docs/deployment.md`](../docs/deployment.md) states it plainly: two hostnames per account, one for the DID and one for the handle, and both must resolve before the account exists. `Route53Dns::publish` in [`crates/didbot-dns/src/route53.rs`](../crates/didbot-dns/src/route53.rs) writes one `ChangeResourceRecordSets` `CREATE` per host it is given, so an account is two record sets, not one. The certificate does not add to that per account: [agent-accounts](agent-accounts.md) commits to one wildcard certificate per zone over ACME DNS-01, so the `_acme-challenge` TXT records it needs are fixed overhead on the zone, not a cost that scales with how many agents the zone holds. Route53's own quotas, checked against AWS's current documentation rather than assumed: | Quota | Default | Adjustable | |---|---|---| | Records per hosted zone (`MAX_RRSETS_BY_ZONE`) | 10,000 | Yes, via Service Quotas or a support request; AWS charges extra above 10,000 | | Hosted zones per AWS account | 500 | Yes, via Service Quotas | | `ResourceRecord` elements per `ChangeResourceRecordSets` request | 1,000 (doubled for `UPSERT`) | No | Source: [Amazon Route 53 quotas](https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/DNSLimitations.html), "Quotas on records" and "Quotas on hosted zones". Both the per-zone record quota and the per-account zone quota are soft — raisable by a support request — which matters for what "the wall" means below: it moves, but not on this server's schedule, and not for free above 10,000 records. At two records per account and the unmodified 10,000-record default, one hosted zone holds **five thousand accounts** before Route53's own quota refuses a write mid-provisioning. That is the raw ceiling this epic exists to stay under, not the number a deployment should configure — see the default below. ## Historical accounts count, not just live ones [account-types](account-types.md) is relaxing the sweep-on-session-end default and making a name permanent once issued rather than freed at the end of a session. The record a DID publishes has to keep resolving for as long as a reader might hold a reference to it — an `at://` URI copied out of a log, a citation, a mention — which is longer than the session that minted the account. An account whose DID document must still resolve still holds its two records, whether or not anything is presently running as that account. So the population this quota is checked against is not "live sessions right now." It is every account whose records are still published: - **Live** (provisioned, in use or pinned): two records held. - **Retained past session end** (soft-deleted, inside its retention window, still resolving because something might still reference it): two records held. This is the state [account-types](account-types.md) is making the common case rather than the exception, and it is why "historical" is the right word for the population that matters here, not "active." - **Withdrawn** (retention window closed, `Registry::delete`'s DNS withdrawal has actually run): zero records held. The name itself moves to `NameRegistry`'s `Released` hold, which is a name-pool cost tracked in [zone-scale](zone-scale.md), not a DNS-record cost tracked here. - **Reserved** (a zone apex, or an operational block in `NameRegistry`): zero DNS records. A reservation is a name held out of the naming pool; it never had a DNS record to begin with. A cap that only ever counted the first bucket would let a deployment mint past what its own zone can hold, because the retained bucket is exactly the one growing while [account-types](account-types.md) relaxes the default that used to shrink it. ## The tasks - [ ] **The cap counts the historical population, not the live one.** Whatever counts against `max_accounts` has to include every account still holding records — live, pinned, and retained-past-session-end — and exclude only the ones actually withdrawn. Getting this wrong in either direction is a real failure: undercounting lets provisioning walk past the real ceiling, and overcounting (counting a `Released` name hold, which has no record) refuses provisioning that would have succeeded. - [ ] **Publish occupancy so it is watched, not discovered.** [zone-scale](zone-scale.md) already asks for `bot.did.stats` to carry how much of the zone's record budget is spoken for; this epic adds nothing new there beyond making sure the number it publishes is the historical count above, not a live-session count. [ops-dashboard](ops-dashboard.md) is where an operator watches it — a panel showing accounts held against the configured cap and against Route53's own quota, the same two numbers this epic's ceiling is built from. [alerts](alerts.md) is the channel for a warning ahead of the wall, on the same terms as every other quietly-failing condition it already carries: a threshold crossing, not a full-zone event, because thirty seconds' notice at 99.9% full is not what an operator needs. No fourth mechanism belongs here; all three already exist for exactly this purpose. - [ ] **Subzones are the real answer, and they only work if chosen early.** Each hosted zone in `didbot_identity::ZoneRegistry` carries its own Route53 record budget, so a deployment routing agents across several zones through `didbot-dns`'s `MultiZoneDns` multiplies its ceiling by the number of zones it configures. That only helps a deployment that adopts a naming scheme with subzones in it before it needs the capacity: a DID is permanent, so an agent already minted at `claudes.pds.did.bot` cannot be moved under a new subzone later without breaking every reference to it. A deployment expecting to grow past a few thousand accounts should choose its subzone layout — by team, by agent type, by cohort, whatever the deployment's own axis is — at first configuration, not when the occupancy panel turns amber. This is the single most useful thing an operator sizing a deployment can be told, and it belongs in the operator-facing documentation this epic's config work produces, not only in this file. - [ ] **Decide, and implement, what happens at the wall.** Three answers, not mutually exclusive with each other in principle but a deployment has to pick one to start: refuse provisioning outright; refuse and name which zone is full, so an operator knows which subzone to add capacity to or which to route around; or place the account into another configured zone automatically. The third interacts with the rule [zone-scale](zone-scale.md) already established for `ZoneRouter`: nothing a request asserts may choose its own zone, so automatic placement at the wall has to be the router's decision from the same admission facts it already uses (harness, agent type), not a fallback a caller can trigger by exhausting the zone it wanted. Refusing and naming the full zone is the smallest correct answer and the one to build first; automatic placement is worth having once there is more than one non-default zone actually configured to place into. - [ ] **Hosted zones per account is a second, higher ceiling worth stating once, not tracking continuously.** Five hundred zones (the default, also raisable) times a few thousand accounts per zone is a population no deployment this project is aware of is near. It is not worth a config setting or a dashboard panel; it is worth one sentence in the operator documentation this epic's config section produces, so a deployment that is somehow near it does not have to rediscover the number. ## Done - [x] **A configured cap on accounts, with a sane default.** `--max-accounts` or `[capacity] max_accounts` sets it, and `DEFAULT_MAX_ACCOUNTS` is 1000, under half the raw ceiling above. `provision_agent` refuses with `503 AccountCapReached`, naming the cap, before it parses the request or writes anything, so a caller never sees a `ChangeResourceRecordSets` failure after the keys and the first hostname exist. The cap is one number for the whole deployment. It is checked against `Registry::account_count`, the account store's size, which is cheap enough to read on every request.