--- id: alerts title: The server says when it is failing, somewhere its operator will see it status: open crates: [didbot-serve, didbot-pds, didbot-lexicon] dependsOn: [pds-writes, ownership] exitCriterion: > A certificate renewal that has been failing for a day is a record in this server's own repository naming its operator, a reader can tell it from a transient blip, and the operator can tell a healthy server from one that has stopped talking. --- # alerts This server holds a DID and a repository, so it can write records like anything else on the network. That makes the notification channel one it already has: when something is wrong, say so as a record, and name the operator. [deploy](deploy.md) lists what wants saying — certificate expiry, a policy poll that has not succeeded, clock skew against the attestation window, and disk — under "alert on the things that fail quietly". This epic is where that list gets a mechanism instead of an intention. The appeal is that it adds no dependency. There is no paging service to hold an account with, no SMTP credential on the box, and nothing new to compromise: the server signs an alert with the key it already signs commits with, and the operator reads it with any client. It is also public, which is the same property [ownership](ownership.md) leans on — an outsider can see that this deployment reported a problem, and when. ## The failure that breaks it, and what to do about it **The failures most worth hearing about are the ones that stop the alert being delivered.** Writing the record is a local operation and will succeed; being *read* is not. If the wildcard certificate expires, TLS is down, the firehose is unreachable, and the alert about the expired certificate is behind the thing it is about. A full disk stops the write itself. A partition stops both. So this channel is an early warning and must not be mistaken for an alarm. - [ ] **A heartbeat, so silence means something.** An alert is a positive signal: its presence says something is wrong, and its absence says nothing at all. A record written on a schedule inverts that — the operator watches for the gap, and the modes that suppress an alert suppress the heartbeat too. Without one, a server that has stopped talking is indistinguishable from a server with nothing to report, which is the same distinction the disclosure routes in [auth-types](auth-types.md) had to make and the opposite way round. - [ ] **Say plainly, in the epic and in the operator documentation, what this channel cannot tell you.** A mechanism believed to be an alarm when it is a warning is worse than no mechanism, because it is trusted. ## What an alert may contain - [ ] **No paths, no tokens, no internal state.** A record is public and permanent. The exposure review found filesystem paths reaching anonymous clients through 500 bodies; the same discipline applies harder here, because a leak in an error response is transient and a leak in a repository is not. An alert names *what* is failing and *since when*, and leaves the detail to a log nobody else can read. - [ ] **Deduplicate, and rate-limit.** A renewal retried hourly must not become twenty-four permanent records a day. One record per condition, updated or superseded rather than repeated, and a floor on how often a flapping condition may write at all. Repository growth is the cost of getting this wrong and it is not recoverable. - [ ] **Decide what clears an alert.** A condition that recovers should say so, or the operator is left reading a permanent record of a problem that ended. ## The shape - [ ] **A lexicon for the record.** There is no lexicon epic — a `bot.did.*` schema belongs to whichever epic needs it, and this is that epic. It wants a severity, a condition identifier stable enough to deduplicate on, a first-seen time, and the operator it names. - [ ] **Which DID writes it.** The service DID, the one [ownership](ownership.md) already has the operator vouching for, so a reader checking who is complaining lands on a chain that is already checkable. - [ ] **Tagging the operator is a mention**, and delivery is [mentions](mentions.md)'s question rather than this epic's. Until that lands, an alert is a record the operator has to look for. - [ ] **The conditions themselves**, each of which lives with the thing it watches rather than here: certificate renewal in [agent-accounts](agent-accounts.md), the policy poll in [policy-store](policy-store.md), clock skew in [attestation](attestation.md), and disk in [deploy](deploy.md). This epic owns the channel; it does not own the thresholds. ## Where it does not belong - [ ] **Not a replacement for [ops-dashboard](ops-dashboard.md).** That pane shows live state and can lie about it under compromise. This channel has the opposite property and the opposite weakness: it cannot be edited after the fact, and it can be silenced. They answer different questions and neither subsumes the other. - [ ] **Not the [e-stop](e-stop.md).** The stop path is local and immediate precisely because it must work when everything else does not. An alert crosses a repository and a reader; that is fine for a warning and useless for a halt. ## Done Nothing closed yet.