id: alerts title: The server says when it is failing, somewhere its operator will see it status: open crates: [didbot-serve, didbot-pds, didbot-lexicon] dependsOn: [pds-writes, ownership] exitCriterion: > A certificate renewal that has been failing for a day is a record in this server's own repository naming its operator, a reader can tell it from a transient blip, and the operator can tell a healthy server from one that has stopped talking. #
alerts #
This server holds a DID and a repository, so it can write records like anything else on the network. That makes the notification channel one it already has: when something is wrong, say so as a record, and name the operator.
deploy lists what wants saying — certificate expiry, a policy poll that has not succeeded, clock skew against the attestation window, and disk — under "alert on the things that fail quietly". This epic is where that list gets a mechanism instead of an intention.
The appeal is that it adds no dependency. There is no paging service to hold an account with, no SMTP credential on the box, and nothing new to compromise: the server signs an alert with the key it already signs commits with, and the operator reads it with any client. It is also public, which is the same property ownership leans on — an outsider can see that this deployment reported a problem, and when.
The failure that breaks it, and what to do about it #
The failures most worth hearing about are the ones that stop the alert being delivered. Writing the record is a local operation and will succeed; being read is not. If the wildcard certificate expires, TLS is down, the firehose is unreachable, and the alert about the expired certificate is behind the thing it is about. A full disk stops the write itself. A partition stops both.
So this channel is an early warning and must not be mistaken for an alarm.
What an alert may contain #
The shape #
Where it does not belong #
Done #
Nothing closed yet.