# Zone reconciliation `didbot_reconcile::ZoneState` is what one **DNS zone** stands at right now, as distinct from `didbot_pds::ServerState`, which is what this *deployment* may do. The two are separate machines on purpose: a hostname withdrawn from the zone by hand breaks agents that already exist whether or not the operator's claim currently stands, so repairing it cannot be a step in a bootstrap that finished months ago. It **acts**. `didbot-claim`'s checks observe a deployment from the operator's own machine and report; this reads the zone from inside, compares it against what this deployment published, and writes the difference back. ```mermaid stateDiagram-v2 [*] --> blocked blocked --> in_sync: a read succeeded and nothing is wrong blocked --> drifted: a read succeeded and something is in_sync --> drifted: a name is missing or points elsewhere drifted --> reconciling: a second read confirms the same drift drifted --> in_sync: it was propagation, not drift reconciling --> in_sync: the repair landed reconciling --> blocked: every outstanding repair is quarantined in_sync --> blocked: the zone could not be read ``` | state | repairs records | observation is evidence | deployment keeps serving | | --- | --- | --- | --- | | `blocked` | no | no | **yes** | | `drifted` | yes | yes | **yes** | | `reconciling` | yes | yes | **yes** | | `in-sync` | yes | yes | **yes** | The last column is constant, and that is the point: this machine is a repairer, not a gate. `serving_is_never_gated_on_this_machines_state` asserts it by destructuring `ZonePolicy` exhaustively over every state, so a capability that could gate a request stops the test compiling. ## What counts as wrong Decided per name, against the intent — the hostnames this deployment published and what each should point at. | observed | verdict | what happens | | --- | --- | --- | | intended name, nothing in the zone | **missing** | republished | | intended name, a different target | **pointing elsewhere** | replaced | | a name the intent does not know | **unattributed** | reported, never touched | | a different TTL | not drift | ignored | | `TXT` values | out of scope | ignored | | the `NS` delegation above the zone | not ours | ignored | A **TTL** is left alone structurally rather than by choice: `RecordTarget` carries none, so an intent cannot express one. That is also the right answer — lowering a TTL before a migration is ordinary practice, and a wrong TTL degrades how fresh an answer is, not whether there is one. **`TXT`** is out of scope because one name holds a set of values that the ACME renewal path writes and withdraws within the length of one challenge; a reconciler working from a five-minute snapshot would race it every run. ## What it never does - **Delete a record it cannot prove stale.** `Repair` has two variants, `Publish` and `Replace`, and no third. `Drift::Unattributed` yields no repair at all, so a name this deployment did not publish has nothing that could act on it. A read that fails produces `ViewError`, not an empty zone — which is why `ZoneView::observe` returns a `Result` where `DnsProvider::published` returns a bare `Vec`. - **Gate serving on its own state.** See the table above. - **Touch a zone it does not manage.** `ZoneAuthority::ReadOnly` blocks before the first read. It never creates a zone under any circumstances, so `dns.may_create_zones` is not a posture it can reach past: a zone that is not there is `ViewError::ZoneAbsent`, which blocks and says so. ## How often, and how hard | bound | default | why | | --- | --- | --- | | interval | 5 minutes | `Route53Dns`'s own default record TTL: never re-examine a name faster than a resolver could have seen the last change | | backoff | double to 1 hour | a tick that could not advance waits longer — a gate failed, or every drifted name is quarantined and the zone is `blocked`. One tick that makes progress clears it | | confirmations | 2 reads | drift must survive a second read before anything is written, so a change in flight is not fought | | repairs per tick | 25 | a wholesale-emptied zone costs a bounded number of writes per tick, not thousands at once | | accepted repairs per name | 5 | a name that keeps drifting back after a repair the provider *accepted* is quarantined and reported. The counter clears the moment the name reads correct, and a repair the provider refused never adds to it | Between the interval and the confirmation count, drift is repaired five to ten minutes after it happens. That is this epic's answer to `plan/onboarding.md`'s open question of how long drift is propagation and how long it is a fault: long enough that a propagating change is never fought, short enough that a withdrawn hostname is back before an agent's next session. The attempt budget is aimed at people rather than at providers, and only at people. A human deliberately changing a record by hand gets argued with five times and then reported, rather than forever. A repair the provider *refused* is not charged to it, and the two bounds are not interchangeable. Quarantine is permanent for the life of the process — its only exit is observing the name correct — so spending it on a provider fault strands the name: nothing but this machine would ever put the record back, and quarantine is what stopped it. A refused repair fails a gate instead, which is what the backoff is computed from, so a provider outage costs one attempt per tick at a rate falling to hourly and the name is repaired on the first tick after the provider recovers. Every one of these counters lives in memory, so a restart discards all of them — the backoff included. What keeps that safe rather than expensive is the confirmation count: a repair needs two consecutive reads from one process, so a process that crash-loops faster than its own interval reads the zone and writes nothing at all. ## The gates `didbot_fsm::Gates`, declared in causal order, so one failure reports as one cause rather than as every check below it: 1. `zone is managed` 2. `zone is readable` 3. `intent is known` 4. `intended records are present` 5. `intended records point here` The last two are the reason this machine brought three-valued gates back to `didbot-fsm`. A missing record awaiting its confirming read is `Pending`: it is wrong, and nothing has been attempted, and reporting that as a failure would make the first tick of a healthy process look like an outage — and the backoff is computed from exactly that signal.