Identities for entities did.bot
agent llm did
didbot plan ai-preference.md
19 kB
Markdown
at commit 18ba4fe0


id: ai-preference title: A stranger's declared AI preference is a ceiling on what our agents may do status: open crates: [didbot-lexicon, didbot-pds] dependsOn: [write-policy] exitCriterion: > An account whose repository declares that it refuses synthetic content is named in a record an agent tries to write; the write is refused, the refusal names the record it read, and no policy this deployment can set permits the write. #

ai-preference #

This epic leans on index, which is declined. The index and the query service live at vibescrobble.com, not in this repository. Whatever this epic needed from an index it now needs from somebody else's service, or it needs restating as something the PDS can answer on its own. That has not been decided, and nothing below should be started before it is. dependsOn names write-policy, because the check sits beside its refusal at the write.

community.lexicon.preference.ai is a lexicon.community record in which an atproto account declares how it wants AI systems to use its public data. It lives at at://did:plc:mtr7qrqtcyseedx3jyr5o7db/com.atproto.lexicon.schema/community.lexicon.preference.ai, at CID bafyreihevdqkrhh4pioixa7d2hwi76raich5vuiz4rvhrdi2lw77ykveqe.

Four axes — training, embedding, inference, syntheticContent — each an allow boolean with its own timestamp. Three scopes: a global default at rkey self, and TID-keyed overrides naming either an entity (a DID or a domain) or a collection in the declaring account's repository. An omitted axis is undefined, which the schema is explicit is not the same as denied.

This deployment mints accounts that write records naming people who never asked for one, and runs an index that reads repositories it did not create. Both are squarely what the record is about. Honouring it is the cheapest thing this project can do to be worth having on the network, and it is a small amount of code sitting in exactly one place.

What the record says, and what it does not #

  • It is not consent to be contacted. syntheticContent: true says a user does not object to AI-derived content from their data. It does not invite an agent to reply to them. README already forbids unsolicited interaction with humans outright, that rule is stricter than anything in this lexicon, and it stays. Nothing here is a route around it.
  • Absence is neither consent nor refusal. Almost nobody has one of these records. Silence therefore decides nothing, and the default for an undefined axis has to come from our own configuration and the owner's policy — never from reading permission into a missing file.
  • It is a preference, not an access control. We can obey it. We cannot make anyone else obey it, and every repository involved is public. Describe what we do as "this deployment honours it", never as a guarantee about the data.

The four axes, against what this project actually does #

The mapping is the substantive decision in this epic. The plumbing is small; mapping an operation to the permissive axis is the failure that matters.

  • syntheticContent — an agent writing a record that names somebody. A mention facet on a record (mentions), an attestation naming a DID that never asked to be named. This is the axis with the most surface here and the one worth building first. A deny here is read as covering being named, not only as covering data derived from theirs. The schema's wording — content "derived from user data" — does not settle it, and the strict reading is both the safe one and the one a person can hold in their head: somebody who set this does not want an agent generating things about them, and a record with their DID in it is that. Decided rather than deferred, because the alternative is a deployment that honours a preference in a sense its author did not mean.
  • embedding — the index. index fingerprints record text with IDF-weighted SimHash and bands it for LSH, which is semantic indexing under the name the field uses. Today discovery walks the vouch chain and every subject it reaches is an agent, so nothing crosses this yet. It crosses the first time a reachable repository holds a human's records, and that is a configuration change rather than a code change.
  • inference — retrieval into a running model's context. The query service answering a mention query, any MCP tool that reads a foreign repository and hands the result to a model. This is the axis that gets crossed silently, because unlike a write it leaves nothing behind.
  • training — none of it, and say so. No corpus here is taken from real repositories. The honest move is to state that where somebody can check it rather than to build a gate in front of a thing that never runs; the item is to notice if evaluation corpora ever start coming from live data.

Resolving a preference #

Blocking this deployment #

The one thing a stranger most plausibly wants is to be left alone by every agent here at once, and the lexicon already expresses it: an entity-scoped record naming the deployment's domain. One record, one string, and it binds every account under that zone — which is why the entity match walks the chain rather than looking only at the agent that happened to act.

{
  "$type": "community.lexicon.preference.ai",
  "updatedAt": "2026-08-31T00:00:00.000Z",
  "scope": {
    "$type": "community.lexicon.preference.ai#entityScope",
    "entity": "foo.example"
  },
  "preferences": {
    "syntheticContent": {
      "allow": false,
      "updatedAt": "2026-08-31T00:00:00.000Z"
    }
  }
}

Fetching without being a bad guest #

Do not consume atproto ecosystem resources is the constraint that shapes all of this. A per-write lookup from every session, across a swarm, is exactly the traffic that gets a server defederated.

Where the check goes #

How it meets the policy mechanism #

Three layers, and only one of them is a ceiling.

Withdrawal #

A preference changes after we already wrote. What is recoverable and what is not:

  • A deletion is itself a signal. Removing a record that named somebody puts a tombstone on the firehose that says they objected. Accepted: the alternative is leaving the record up.

Our own accounts declare too #

Scenarios #

The cases the design has to answer, and what it answers.

  1. Global deny, an agent mentions them. A user's self record sets syntheticContent.allow: false. An agent writes a record whose facet resolves to that DID. Refused at the write, before anything is signed; the model gets the distinct refusal; the audit trail gets an entry.
  2. Global deny, entity allow for us. The same user adds a TID-keyed record naming this deployment's domain with syntheticContent.allow: true. Permitted — entity precedes global, and it names us. This is why "any deny wins" is not the rule.
  3. Global deny, and a fresh agent. A new agent DID has never been named in anyone's record. Still refused: the chain checked includes the deployment and the owner, and ephemeral identity buys nothing.
  4. The deny arrives second. The record was written last week and the preference changed today. The sweep deletes ours; a label already sequenced stays until backfill; anything that federated is gone.
  5. Their server is down. Never fetched: undefined, configured default, which for naming a stranger should already be refusal. Cached deny: still denied, indefinitely.
  6. A collection scope, and the thing it cannot express. A deny on app.bsky.feed.post names a collection in their repository, so it bounds what the index may embed of their posts and says nothing about a facet we write. "Do not name me" is only expressible on syntheticContent, at global or entity scope. Do not read a collection scope as a writing rule.
  7. Handles move. The subject is keyed by DID throughout — the facet stores the resolved DID for exactly this reason — so their preference survives a handle change. The entity string in their record points at us, and holds as long as we hold the domain.

Upstream #

The schema is one record in a governed namespace, and the governance is real: the lexicons repository at tangled.org/lexicon.community/lexicons is the source of truth, ideas start on the Lexicon Community forum carrying the tool that would use them, changes land as pull requests a Technical Steering Committee approves, and merging to main is what publishes a schema to the network. It is the same forge this project already uses, so a proposal from here is atgc pr create --patch-only against a repository we have no push access to.

Done #

Nothing closed yet.