id: policy title: What an agent may do is a set of denials, evaluated at the write, from three sources status: open crates: [didbot-pds, didbot-serve, didbot-lexicon, didbot-config] dependsOn: [pds-writes, ownership] exitCriterion: > A write that a policy denies is refused before it reaches the log, with a reason the caller can act on and a durable record that holds no part of the refused payload; a deployment with an evaluator down refuses only the writes that evaluator judged; and an operator can see which policies are loaded, which evaluator serves each, and whether it is healthy. #
policy #
This epic replaces the design assumed by policy-store, scope-policy, write-policy, app-allowlist and policy-dashboard. Those five each assumed a policy model without one having been agreed. Read this file first; where they disagree with it, they are wrong.
What a policy is #
A policy denies. There are no allows. An empty policy set permits everything, and every policy subtracts from that.
The reason is not economy, it is monotonicity, and three properties depend on it:
- Policy sets compose without interaction. Adding a policy can never widen what is permitted, so there is no shadowing, no precedence rule between sources, and no ordering semantics to get wrong. Any deny from any source is sufficient.
- Order is presentational, not semantic. Evaluation still runs in a predictable order, but only so that the reason reported first is stable. Nothing about the verdict depends on it.
- Fail-closed is strictly conservative. When an evaluator cannot answer, the writes it judged are refused. With denials only, a refusal on an unavailable evaluator is guaranteed to be at least as restrictive as a complete evaluation would have been. Introduce allows and this stops being true: the unavailable evaluator might have been the one that permitted, and failing closed would refuse something the full set allows. Deny-only is what makes the availability behaviour below coherent rather than arbitrary.
There is no way to except yourself from a built-in policy. An operator who needs something a built-in denies forks and compiles their own binary. This is a property, not a gap; the first person to hit it will file it as a bug and should be pointed here.
Where policies come from #
Three sources. All three deny; none can permit what another denies, so their order is a matter of where a policy is written and not of what it can do.
- Compiled into the binary, immutable. Things unsafe to admit into public existence. An application that does not want agents using it opens a pull request against didbot asking to be denied as a client; that denial ships in the binary.
- Supplied at startup, immutable for the run. Available for deployments that need it. Expected to be unused.
- The operator's own records, polled.
bot.did.policycarries each denial andbot.did.policyBindingsays whom it applies to. This source can fail: a policy may name an engine this binary does not have, or the operator's repository may be unreachable.
A failed refresh keeps the last good set. A failing operator must never widen what an agent may do.
A set that has never loaded is different, and is handled at onboarding rather than at runtime: see "Pre-flight" below.
A policy is not its wrapper #
A policy is separate from the file or record that carries it. That wrapper supplies metadata of two kinds:
- Required: which language the policy is written in, and therefore which evaluator handles it.
- Optional, and transport-shaped: which PDS it applies to, a version string of the author's own. That version is not this server's version of its evaluation tree, and conflating them will produce a system where an operator cannot tell which of their edits is live.
Two trees, one set of mechanics #
Policy is evaluated at two points, with the same indexing, the same evaluator processes and the same outcomes:
- Writes, whose subject is
(agent, client_id, diff).client_idis absent for a write made with the account's own credential. - Sign-ins, whose subject is
(client_id, requested scopes). - Token issue, whose subject is
(agent, client_id, scopes).
The subject shapes differ; the machinery does not. Including client_id in the
write subject is what lets one app-scoped rule appear in every tree with the
same predicate, which the following scenario requires.
Scenario — a token outliving its app's admission. An agent holds a valid token for an application. A policy denying that application is then added. The application presents its still-valid token and attempts a write. The write tree must refuse it, because the token was issued before the denial existed and nothing revokes it retroactively. A rule that lived only in the token tree would let every already-issued token continue writing until it expired.
Note that denying an application is not denying a collection. Record types are shared between applications; blocking one client must not stop an agent writing that collection through another.
Evaluation happens at the write #
Not at the grant. The rule that forces this: an agent may edit its bio but not its display name is a constraint on the diff, and a diff exists only at the moment of a write. Anything baked into a token at grant time cannot express it.
The tree indexes; the list evaluates #
The tree is an index for finding applicable policies. What comes out of it is an ordered list. The order is part of this server's version of the tree, so that the reason reported for a denial is stable.
Availability is scoped by the tree #
The index is not only a throughput optimisation. It decides the blast radius of an evaluator failure.
The evaluator contract #
Last-good-revision #
A policy that compiled at revision N and fails to compile at N+1 keeps
enforcing N — dropping it would widen what an agent may do, and denying
everything it covers would turn a typo into an outage. didbot-policy-source's
LastGoodPolicies holds the last successfully-compiled record per
operator-sourced id, across every merge::build call; a revision that fails
to compile is looked up there and, if found, recompiled and enforced in the
new revision's place, reported as a LastGoodFallback — distinct from a
ScopedDenial, which is what a policy that has never compiled still gets,
loudly, because there is nothing to fall back to and a first-publish typo
must not silently deny its own coverage. startup (source 2, immutable for
the run) has no last-good-revision fallback: there is never a second
revision to have fallen back from.
Deadlines, lag, and adversarial policies #
Outcomes #
allow, reject, freeze.
What is written down, and what is not #
Reading records this deployment did not design #
The motivating policy: an agent never mentions someone who has opted out of AI. Evaluating it requires the lexicon schema for a record type didbot did not design, to find which fields hold a DID; a cache lookup per DID found; and a network resolution on a cache miss.
Pre-flight, so there is never an unpolicied instance #
The operator sees the tree #
Not this epic #
Two neighbouring epics use "policy" in a different sense and are not governed by anything here. name-pools's reclaim policy is a rule about when a handle returns to its pool; did-minting mentions a policy only to say that making reuse negligible is preferable to enforcing one. Neither is an agent policy, neither is evaluated by the tree below, and neither denies a write.
What a policy is keyed on, and what that key is worth #
A rule that applies to "agents of this type" is keyed on something, and the strength of the whole rule is the strength of that key.
Configuration may narrow, never grant #
didbot-config's own documentation, plan/config.md and
docs/deployment.md all say that collection permissions belong in the
owner's records rather than in this server's configuration file, and that a
reviewer should reject a configuration key that adds "a scope, a permission,
or an allowlist of apps" on that question alone.
That rule holds, and there is a shape it does not forbid, worth writing down before somebody argues it either way in a review:
Deliberately unsettled #
These need iterative design and must not be invented by whoever implements first:
- The shape of the required and optional wrapper metadata, beyond the two kinds named above.
- Which languages ship, past a deterministic in-process regular-expression evaluator for the simplest cases — which is also what makes an application denial an early, cheap rejection near the root of the tree.
- The relationship to the atproto scope grammar. Policy is the ceiling; whether scopes remain as a coarse projection for third-party clients that expect them is open.
- Whether deletes are evaluated, per the item above.
Consent is a denial, evaluated at authorize #
An agent has nobody at a consent screen, so the question at authorize is not whether somebody approved an application — it is whether this one is denied. That is this epic's own model applied one step earlier, and the machinery is already shaped for it: the grant is a subject kind of its own, and denying an application by its client identifier applies to a grant request as much as to a write.
Done #
-
Source 3 is read and enforced.
didbot_serve::policy_poll::PolicyPolllists both collections on the ownership poll's tick, builds the set whole throughdidbot_pds::policy_bindings::resolve, keeps the build last refused for an operator to read (PolicySet::rejected), and swaps the tree into the gate;bot.did.statspublishes the set's digest aspolicy.enforced. -
Bindings resolved within each account's boundary (
crates/didbot-pds/src/policy_bindings.rs).resolvebuilds the whole set from one observation of both collections or refuses it, naming every record it cannot honour inRejectedBuild; a refused build leaves the set last built serving. The gate judges a write withinRegistry::boundary's chain, so a binding on a principal reaches every account beneath it, minus itsexcludes; the server's own DID as a subject stands for the whole server, and a grant judged before any account has signed in is reached only that way. -
The two lexicon documents,
lexicons/bot/did/policy.jsonandlexicons/bot/did/policyBinding.json, and the reader for each (crates/didbot-policy-records/src/records.rs). A policy names theactionsit is consulted on and adocumenttyped by the engine that reads it; a binding applies policies, each named by AT-URI, to itssubjects, theirdescendants, or both, minus itsexcludes. The policies must currently be in the same repository as the binding; a reference into another repository is not supported yet. -
The
compilestep onEvaluator(three outcomes, aCompiledIdthe evaluator owns, a parallel policy-lifecycle channel) and last-good-revision (crates/didbot-policy,crates/didbot-policy-source,crates/didbot-policy-regex,crates/didbot-pds/src/policy_tree.rs). -
`AccountState` itself stays the bare `Frozen` variant it already was. The reason instead lives on `LedgerEvent::StateChanged`'s new `policy_evaluation: Option<EvaluationId>` field (`crates/didbot-pds/src/ledger.rs`) — absent for an administrator's or a lifetime rule's freeze (which still use `operator`), present for a policy's, carrying the id `crate::evaluation_log::EvaluationLog` minted for the judgment that tripped it. `FreezeSyncingGate::froze` (`crates/didbot-pds/src/provision.rs`) reads the id off `PolicyGate::last_evaluation_id` and writes the ledger entry. -
`crate::evaluation_log::DenialRow` (`crates/didbot-pds/src/evaluation_log.rs`) carries `closed_at`, the `now` the judgment actually ran under, rather than anything an operator could use to re-derive the verdict later. -
`DenialRow` is exactly that field list, and nothing else; `hash_subject` is the one place the module reads a payload, over which it produces only a SHA-256. `evaluation_log::tests::a_row_never_carries_the_refused_value` asserts a row's own serialized JSON never contains the refused value. -
`EvaluationLog` keys an in-memory, bounded collapsing window on (account, client, policies fired, reason); a repeat inside the window advances a count rather than writing a row, and the row lands once the window closes — on a later attempt outside the window, an eviction past `DEFAULT_MAX_OPEN_WINDOWS`, or a deployment's own periodic `flush_stale`. Both halves are covered: the map bounds memory during an active spree, the row is what a metapolicy's freeze (which needs several attempts to trip) has time to land before. -
`evaluation_log::CaptureSink` is the seam: nothing in this crate wires one in, `EvaluationLog::with_capture` is the only way one is ever called, and it fires only for the write that opens a fresh window — never a collapsed repeat. Its own retention is the sink's business, not `EvaluationLog`'s.