--- id: ai-preference title: A stranger's declared AI preference is a ceiling on what our agents may do status: open crates: [didbot-lexicon, didbot-pds] dependsOn: [write-policy] exitCriterion: > An account whose repository declares that it refuses synthetic content is named in a record an agent tries to write; the write is refused, the refusal names the record it read, and no policy this deployment can set permits the write. --- # ai-preference **This epic leans on [index](index.md), which is declined.** The index and the query service live at vibescrobble.com, not in this repository. Whatever this epic needed from an index it now needs from somebody else's service, or it needs restating as something the PDS can answer on its own. That has not been decided, and nothing below should be started before it is. `dependsOn` names [write-policy](write-policy.md), because the check sits beside its refusal at the write. `community.lexicon.preference.ai` is a lexicon.community record in which an atproto account declares how it wants AI systems to use its public data. It lives at [`at://did:plc:mtr7qrqtcyseedx3jyr5o7db/com.atproto.lexicon.schema/community.lexicon.preference.ai`](https://eurosky.social/xrpc/com.atproto.repo.getRecord?repo=did:plc:mtr7qrqtcyseedx3jyr5o7db&collection=com.atproto.lexicon.schema&rkey=community.lexicon.preference.ai), at CID `bafyreihevdqkrhh4pioixa7d2hwi76raich5vuiz4rvhrdi2lw77ykveqe`. Four axes — `training`, `embedding`, `inference`, `syntheticContent` — each an `allow` boolean with its own timestamp. Three scopes: a global default at rkey `self`, and TID-keyed overrides naming either an entity (a DID or a domain) or a collection in the declaring account's repository. An omitted axis is undefined, which the schema is explicit is not the same as denied. This deployment mints accounts that write records naming people who never asked for one, and runs an index that reads repositories it did not create. Both are squarely what the record is about. Honouring it is the cheapest thing this project can do to be worth having on the network, and it is a small amount of code sitting in exactly one place. ## What the record says, and what it does not - **It is not consent to be contacted.** `syntheticContent: true` says a user does not object to AI-derived content from their data. It does not invite an agent to reply to them. [README](../README.md) already forbids unsolicited interaction with humans outright, that rule is stricter than anything in this lexicon, and it stays. Nothing here is a route around it. - **Absence is neither consent nor refusal.** Almost nobody has one of these records. Silence therefore decides nothing, and the default for an undefined axis has to come from our own configuration and the owner's policy — never from reading permission into a missing file. - **It is a preference, not an access control.** We can obey it. We cannot make anyone else obey it, and every repository involved is public. Describe what we do as "this deployment honours it", never as a guarantee about the data. ## The four axes, against what this project actually does The mapping is the substantive decision in this epic. The plumbing is small; mapping an operation to the permissive axis is the failure that matters. - **`syntheticContent` — an agent writing a record that names somebody.** A mention facet on a record ([mentions](mentions.md)), an [attestation](attestation.md) naming a DID that never asked to be named. This is the axis with the most surface here and the one worth building first. **A deny here is read as covering being named**, not only as covering data derived from theirs. The schema's wording — content "derived from user data" — does not settle it, and the strict reading is both the safe one and the one a person can hold in their head: somebody who set this does not want an agent generating things about them, and a record with their DID in it is that. Decided rather than deferred, because the alternative is a deployment that honours a preference in a sense its author did not mean. - **`embedding` — the index.** [index](index.md) fingerprints record text with IDF-weighted SimHash and bands it for LSH, which is semantic indexing under the name the field uses. Today discovery walks the vouch chain and every subject it reaches is an agent, so nothing crosses this yet. It crosses the first time a reachable repository holds a human's records, and that is a configuration change rather than a code change. - **`inference` — retrieval into a running model's context.** The query service answering a mention query, any MCP tool that reads a foreign repository and hands the result to a model. This is the axis that gets crossed silently, because unlike a write it leaves nothing behind. - **`training` — none of it, and say so.** No corpus here is taken from real repositories. The honest move is to state that where somebody can check it rather than to build a gate in front of a thing that never runs; the item is to notice if evaluation corpora ever start coming from live data. ## Resolving a preference - [ ] **Read the set, not the record.** Resolve the subject's DID to its server, `listRecords` the collection, take `self` as the global default and each TID-keyed record as an override. - [ ] **Follow the documented resolution order, which is not in the schema.** The record carries no precedence rule; the namespace's README in the lexicons repository does, and it is entity, then collection, then the global default at `self`, with overrides additive — an override that declares one axis inherits the rest rather than replacing them. Implement that, and cite where it comes from, because a reader holding only the schema record cannot derive it. "Any deny wins" is the rule to *not* reach for: it makes global-deny-plus-one-allow inexpressible, which is the obvious thing a user wants to write. - [ ] **Match the entity against every identity we have, and let the broadest deny win.** A user writing an entity scope will name whichever of us they have heard of: an agent's DID, the owner's DID, the deployment's domain, or the project. Ephemeral accounts make this load-bearing — if only the acting agent's DID were checked, a denial would be escaped by minting the next agent, which this server does thousands of times. So an entity-scoped *allow* on a narrow identity may sit under a broad one, and a deny on any identity in the chain is a deny for everything below it. - [ ] **Publish the identifiers to name.** A user cannot write an entity scope against a deployment whose domain and owner DID they cannot find. The apex account and `bot.did.registration` already carry the operator; state plainly, in one fetchable place, which strings work here. - [ ] **Pin the lexicon by CID, and alert rather than refuse when it moves.** A `com.atproto.lexicon.schema` record is rewritable in place by whoever holds that repository, so validating against whatever is served today means a stored preference can change meaning without us noticing. It will move legitimately: a merge to the lexicons repository's `main` republishes the record, which is the whole publication path. So the pin is a signal to go read the diff, not a reason to stop honouring preferences. ## Blocking this deployment The one thing a stranger most plausibly wants is to be left alone by every agent here at once, and the lexicon already expresses it: an entity-scoped record naming the deployment's domain. One record, one string, and it binds every account under that zone — which is why the entity match walks the chain rather than looking only at the agent that happened to act. ```json { "$type": "community.lexicon.preference.ai", "updatedAt": "2026-08-31T00:00:00.000Z", "scope": { "$type": "community.lexicon.preference.ai#entityScope", "entity": "foo.example" }, "preferences": { "syntheticContent": { "allow": false, "updatedAt": "2026-08-31T00:00:00.000Z" } } } ``` - [ ] **Publish that, filled in, where a stranger will find it.** Not in `plan/`. The deployment's own site is the place, and the apex account that [labels](labels.md) established is the thing to name — a domain somebody read off an agent's handle is the identifier they already have. - [ ] **Accept both spellings of the union member.** The namespace's own examples write `"$type": "#entityScope"`, relative to the enclosing document; a fully qualified `community.lexicon.preference.ai#entityScope` is the other form and is what a generic client is likelier to emit. Refusing either would mean refusing a preference somebody wrote in good faith, which is the worst failure available here. - [ ] **Say what it does not reach.** It stops the agents; it does not retract what is already published, and it says nothing to any other deployment. Overstating it is worse than the gap. ## Fetching without being a bad guest `Do not consume atproto ecosystem resources` is the constraint that shapes all of this. A per-write lookup from every session, across a swarm, is exactly the traffic that gets a server defederated. - [ ] **One resolver, one cache, deployment-wide.** The check belongs beside the index, which already holds a store and a sweep, not in ten thousand sessions each asking bsky.social the same question. - [ ] **Cache the misses.** Most subjects have no such record, so the miss is the common path and the one that costs somebody else a request. - [ ] **A stated TTL, since there is no invalidation signal.** The firehose in [index](index.md) covers vouched servers, not the network, so nothing tells us a stranger changed their mind. The TTL *is* the propagation delay, and gets written down rather than discovered — same argument as the polling interval in [policy-store](policy-store.md). - [ ] **Make the failure modes asymmetric, on purpose.** Never fetched, and unreachable: undefined, falling to the configured default. Previously read as deny, now unreachable: still denied, and a deny never ages back into an allow because somebody's server went down. Previously read as allow: expires to undefined at the TTL. ## Where the check goes - [ ] **At the write, beside the other refusals.** Same argument [write-policy](write-policy.md) makes: a rule enforced in the tool is a request, and the model is not a policy engine. A tool description is worth having — it is where a model learns why it was refused — but it is not the enforcement. - [ ] **Subject extraction, per lexicon, in `didbot-lexicon`.** Given a record, which DIDs does it name? Facets, attestation subjects, post replies and quotes. A collection with no extractor names nobody, which is a hole; it is mostly closed already by the existing refusal of collections outside this deployment's namespace, and what remains should fail closed rather than silently extract nothing. - [ ] **A third distinct refusal.** Not "malformed", not "your policy forbids this", but "the account you named has declined this" — legible enough that an agent stops rather than retries, and carrying the record it came from so the claim is checkable. - [ ] **Record the refusal in the audit trail**, so honouring a preference is something a stranger can be shown rather than told. ## How it meets the policy mechanism Three layers, and only one of them is a ceiling. - [ ] **Policies that ship, on by default.** A subject who has denied `syntheticContent` is not somebody an agent should be reaching in *any* way, so the shipped rule is not only "do not write a record naming them" but the whole interaction surface that has them as its subject: a reply, a quote, a like, a follow, a mention facet, or an attestation about them. One derived deny-list, keyed by subject DID, consulted by every agent under this deployment. README already says a policy this project ships disabled is disabled for a reason; this is one shipped enabled for the same kind of reason. - [ ] **Every agent write transits this server, so record-writing is covered completely.** That is worth stating because it is unusually strong for a rule of this kind, and because it makes the boundary sharp: what is *not* covered is anything an agent does off atproto, anything a third-party app does with credentials we issued, and anything already federated. "To the best of our ability" has an edge, and it should be written down rather than implied. - [ ] **The owner's policy narrows the defaults and cannot touch the ceiling.** [policy-store](policy-store.md) decides the stance for an undefined axis, the TTL, and whether this deployment binds itself tighter than anyone declared. It does not decide what a declared deny means: the preference check runs last, after [policy-store](policy-store.md) and [write-policy](write-policy.md) have had their say, so no policy version reaches past it. - [ ] **Draw the difference in [policy-dashboard](policy-dashboard.md).** The shipped defaults are dials; the preference is a floor under them. A control that looks adjustable and is not is worse than one that is not drawn. ## Withdrawal A preference changes after we already wrote. What is recoverable and what is not: - [ ] **Our collections on our server: delete.** The same sweep that refreshes the cache re-checks the subjects the index already knows are named, and records naming an account that now denies are removed. Bounded by the index rather than by rescanning every repository. - [ ] **Derived state in the index: drop.** Fingerprints, group membership and anything a mention query answers from, for a subject who has turned `embedding` or `inference` off. Cheap, because the index rebuilds from the sweep. - [ ] **Federated and labelled: cannot be unsaid.** [labels](labels.md) already records that a retraction reaches a reader on their next backfill and no sooner. Do not claim more than that here. - **A deletion is itself a signal.** Removing a record that named somebody puts a tombstone on the firehose that says they objected. Accepted: the alternative is leaving the record up. ## Our own accounts declare too - [ ] **Give agent accounts a `self` record rather than leaving them undefined.** They are atproto accounts and the absent state is the ambiguous one we just complained about. What it should say is a deployment decision; that it exists is not. - [ ] **The model does not write it.** Same rule as the profile in [write-policy](write-policy.md): a preference record is a statement by the accountable human, not content. ## Scenarios The cases the design has to answer, and what it answers. 1. **Global deny, an agent mentions them.** A user's `self` record sets `syntheticContent.allow: false`. An agent writes a record whose facet resolves to that DID. Refused at the write, before anything is signed; the model gets the distinct refusal; the audit trail gets an entry. 2. **Global deny, entity allow for us.** The same user adds a TID-keyed record naming this deployment's domain with `syntheticContent.allow: true`. Permitted — entity precedes global, and it names us. This is why "any deny wins" is not the rule. 3. **Global deny, and a fresh agent.** A new agent DID has never been named in anyone's record. Still refused: the chain checked includes the deployment and the owner, and ephemeral identity buys nothing. 4. **The deny arrives second.** The record was written last week and the preference changed today. The sweep deletes ours; a label already sequenced stays until backfill; anything that federated is gone. 5. **Their server is down.** Never fetched: undefined, configured default, which for naming a stranger should already be refusal. Cached deny: still denied, indefinitely. 6. **A collection scope, and the thing it cannot express.** A deny on `app.bsky.feed.post` names a collection in *their* repository, so it bounds what the index may embed of their posts and says nothing about a facet we write. "Do not name me" is only expressible on `syntheticContent`, at global or entity scope. Do not read a collection scope as a writing rule. 7. **Handles move.** The subject is keyed by DID throughout — the facet stores the resolved DID for exactly this reason — so their preference survives a handle change. The entity string in their record points at us, and holds as long as we hold the domain. ## Upstream The schema is one record in a governed namespace, and the governance is real: the lexicons repository at `tangled.org/lexicon.community/lexicons` is the source of truth, ideas start on the Lexicon Community forum carrying the tool that would use them, changes land as pull requests a Technical Steering Committee approves, and merging to `main` is what publishes a schema to the network. It is the same forge this project already uses, so a proposal from here is `atgc pr create --patch-only` against a repository we have no push access to. - [ ] **One question is genuinely unanswered, and it is ours to raise.** How `entity` matches a consumer that holds several identities at once — an agent DID under an owner DID under a domain — which ephemeral accounts make load-bearing and which nobody minting thousands of DIDs has had to ask yet. Everything else this epic needed was already decided upstream or is decided here. - [ ] **Tell them how we read `syntheticContent`, without asking.** The strict reading above is a consumer's choice to be more careful than the wording compels, not a gap in the schema, so it belongs in a note on the forum rather than in a patch. - [ ] **Lead with the consumer, not the schema.** The forum asks for the application that would use a lexicon, and this deployment is one: a thing that writes records naming strangers and can be made to stop. - [ ] **The pull request template assigns contribution rights to Lexicon Community under MIT.** A schema patch is small and the namespace is a commons, but this repository ships no license at all ([license](license.md)), so the assignment is the owner's call before anything is submitted. - [ ] **Their CI runs `lex gen-api` over every schema and fails on any stderr.** Run it locally first; a proposal that does not survive `@atproto/lex-cli` codegen is not a proposal yet. - [ ] **Watch the adjacent work rather than duplicating it.** The namespace cites Bluesky's proposal 0008 on user intents for data reuse and the IETF `aipref` working group. If a general answer lands in either, honouring it beats carrying our own reading of a community schema. ## Done Nothing closed yet.