diff --git a/.env.prod.example b/.env.prod.example index f6ae7b0..27d8513 100644 --- a/.env.prod.example +++ b/.env.prod.example @@ -16,9 +16,29 @@ POSTGRES_PASSWORD=CHANGE_ME # proxying CDN edge breaks. BRIDGE_HOSTNAME=tdpl.io +# Origin the NATIVE Coves users' ActivityPub actors are served under — the +# second origin on the bridge's listener, split from BRIDGE_HOSTNAME by a Host +# router. REQUIRED IN PRODUCTION: unset, the process refuses to start with +# config: AP_USER_ORIGIN is required in production +# (this is why a HEAD deploy onto a pre-v2 .env fails at boot — fail closed). +# Must be https here, and must not be BRIDGE_HOSTNAME or a subdomain of it. +# +# Baked into every actor_id minted under it. Serving reads the stored actor_id, +# never this value, so changing it does NOT move existing actors — it only +# changes where new ones appear. Treat it as permanent. +# Caddy must proxy this origin's AP paths to the bridge; see DEPLOY.md. +# +# NO TRAILING SLASH. It must be a bare scheme://host — any path (including a +# lone "/"), query, fragment, or userinfo is refused at boot. +AP_USER_ORIGIN=https://coves.social + # 32-byte key-encryption key sealing per-actor signing keys at rest # (AES-256-GCM). Generate with: openssl rand -hex 32 # Losing it means losing every bridged repo's signing keys — back it up. +# THERE IS NO ROTATION PATH. Nothing in the codebase re-seals existing +# ciphertext under a new KEK; changing this value orphans every sealed key +# (every bridged identity and the PLC escrow rotation key) with no recovery. +# See DEPLOY.md "Not implemented" before you touch it. BRIDGE_KEK=CHANGE_ME # Bearer token for the /admin API. Generate with: openssl rand -hex 32 @@ -38,17 +58,56 @@ RELAY_HOSTS=https://bsky.network RELAY_ADMIN_PASSWORD=CHANGE_ME # Task-14 Jetstream consumer: native Coves users' opt-outs, profiles, posts, -# comments and votes flowing outward. DEFAULT OFF and left off until the -# outbound-delivery/acceptance/echo-moderation seams (tasks 15-17) land: with -# it on, the pipeline runs but delivery is a logging noop, postv2 is skipped -# (no acceptance engine), and opt-out deleteRemote + account deletions are only -# recorded, not acted on. Flip to true (and set JETSTREAM_URL) only once those -# seams are wired. +# comments and votes flowing outward. DEFAULT OFF. The seams it feeds are now +# all wired — outbound delivery (15), the acceptance engine (16), echo +# suppression and moderation (17) — so turning it on runs the REAL pipeline: +# postv2 is admitted and a community-signed acceptance written, opt-out +# deleteRemote and confirmed account deletions actually purge at peers. It is +# still default-off because it writes durable outbound state, and because +# nothing about being able to run it makes it safe to run it un-canaried. +# Read the staged rollout in DEPLOY.md before flipping this. CONSUMER_ENABLED=false # The self-hosted Jetstream the consumer subscribes to (ws:// or wss://). # REQUIRED when CONSUMER_ENABLED=true; validated at boot whenever set, so a bad -# value fails fast instead of becoming a silent reconnect loop. Points at the -# jetstream service in docker-compose.prod.yml (host-local). Safe to stage here -# ahead of flipping CONSUMER_ENABLED. -# JETSTREAM_URL=ws://jetstream:6008/subscribe +# value fails fast instead of becoming a silent reconnect loop. Safe to stage +# ahead of flipping CONSUMER_ENABLED — which is the point of staging it. +# +# Points at the jetstream service in docker-compose.prod.yml, which listens on +# :8080 (JETSTREAM_ADDR). :6060 is its debug/metrics listener and serves no +# /subscribe. Compose already defaults this; uncomment only to override. +# JETSTREAM_URL=ws://tidepool-prod-jetstream:8080/subscribe + +# Delivery workers. 0 = OFF, and AND-gated on CONSUMER_ENABLED — there is no +# OUTBOUND_ENABLED. With the consumer on and this at 0, outbound intent +# accumulates durably and nothing is POSTed; raising it starts draining that +# backlog immediately, so raise it deliberately. +OUTBOUND_WORKERS=0 + +# Kill switches. A blocked delivery is PARKED — still pending, resumed when the +# switch clears — never dropped. Read once at boot; changing one needs a +# service recreate, not a signal. +# OUTBOUND_DRY_RUN translate + log, POST nothing +# OUTBOUND_DISABLED global stop +# OUTBOUND_DISABLED_HOSTS comma-separated inbox hosts (lowercased) +# OUTBOUND_DISABLED_COMMUNITIES comma-separated community AP ids, EXACT +# match ("https://lemmy.world/c/comicstrips") +# OUTBOUND_DISABLED_ACTORS comma-separated actor DIDs, EXACT match +# There is no allowlist form: a one-community canary is spelled by disabling +# every OTHER community in the follow list. See DEPLOY.md. +# OUTBOUND_DRY_RUN=false +# OUTBOUND_DISABLED=false +# OUTBOUND_DISABLED_HOSTS= +# OUTBOUND_DISABLED_COMMUNITIES= +# OUTBOUND_DISABLED_ACTORS= + +# Acceptance-engine flood cap: how many posts one native author may have +# accepted into one bridged community. 0 = unlimited. Tidepool signs the +# community's acceptance, so this bounds what it vouches for. +# ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY=50 + +# Reconciliation sweep cadence. NOT an on/off switch: the sweep is wired +# unconditionally, runs once at startup, and 0 is refused. It is read-only +# (reports divergence, repairs nothing), so leaving it alone is correct. +# DIVERGENCE_INTERVAL=15m +# DIVERGENCE_ACCEPTANCE_STALE_AFTER=12h diff --git a/DEPLOY.md b/DEPLOY.md new file mode 100644 index 0000000..0e9b0e3 --- /dev/null +++ b/DEPLOY.md @@ -0,0 +1,701 @@ +# Deploying Tidepool v2 + +The production runbook for the v2 surface: native Coves users federating +outward (tasks 13–17), on top of the v1 inbound bridge that is already running. + +**This document describes what is in the code, not what was planned.** Where a +lever does not exist, it says so under [Not implemented](#not-implemented) +rather than describing a plausible one. A runbook that names a knob nobody +built is worse than one admitting the knob is missing, because it sends an +operator hunting for it during an incident. Every claim below cites the file it +came from; if you change the code, change the citation. + +Related: [`SELF_HOSTED_RELAY.md`](SELF_HOSTED_RELAY.md) (the relay + Jetstream +ingest path), [`README.md`](README.md#configuration) (the full config table and +admin route table). + +--- + +## 1. Read this before deploying HEAD + +**A HEAD deploy onto a pre-v2 `/opt/tidepool/.env` does not start.** It exits +during `config.Load` with: + +``` +config: AP_USER_ORIGIN is required in production +``` + +`AP_USER_ORIGIN` is read through `stringVar`, which falls back to a dev default +in development and returns an error in production +(`internal/config/config.go:797-806`, read at `:414`). The pre-v2 env file has +no such variable. + +**This is correct behaviour and must not be "fixed" by defaulting it away.** +The origin is baked into every `actor_id` this deployment mints; serving then +derives every URL from the *stored* `actor_id` and never from config +(`internal/personas/serving.go:20-23`). A wrong value is therefore not a +restart away from repaired — it is a set of federated identities pointing at +the wrong place, permanently. Two further boot-time guards exist for the same +reason: it must be `https` in production (`config.go:782-784`), and its host +must not be `BRIDGE_HOSTNAME` or a subdomain of it, which would shadow the +bridged-handle namespace (`config.go:786-790`). + +**It must also be a bare origin — no trailing slash.** +`personas.CanonicalizeOrigin` refuses any path, query, fragment, or userinfo +(`internal/personas/origin.go:42-46`), and `https://coves.social/` has a path +of `/`. That is a boot failure, not a normalization. + +`docker-compose.prod.yml` now supplies `AP_USER_ORIGIN` with a +`${AP_USER_ORIGIN:-https://coves.social}` default, so a `git pull` plus the +normal deploy command is sufficient. Set it explicitly in `.env` anyway if this +deployment is not tdpl.io/coves.social. + +### Minimum `.env` delta + +Nothing else is *required* — every other v2 variable has a real default in +every environment. See `.env.prod.example` for the full annotated set. + +```sh +# required; the rest of the v2 knobs default safely +AP_USER_ORIGIN=https://coves.social +``` + +### Deploy command + +Unchanged, and still targeted — never a bare `up -d`: + +```sh +cd /opt/tidepool && git pull +docker compose -f docker-compose.prod.yml up -d --build tidepool +``` + +The migrate one-shot and the server share one locally-built image and the +server only starts after `service_completed_successfully` on the migration, so +a failed migration holds the old container up rather than starting a new one +against an unmigrated schema. + +Confirm the boot actually got past config: + +```sh +docker logs --tail 50 tidepool-prod | grep -E 'listening|config:' +curl -sf localhost:8091/xrpc/_health +``` + +--- + +## 2. The v2 flag topology + +Everything v2 is off by default and turns on in one order. Nothing here is a +runtime toggle: **config is read exactly once**, at +`cmd/tidepool/main.go:109`, and there is no reload signal — changing any value +means editing `.env` and recreating the service. + +``` +CONSUMER_ENABLED=false ← nothing v2 runs. Today's state. + │ + │ ON: the Jetstream consumer dials JETSTREAM_URL. The REAL enqueuer + │ persists outbound intent; the acceptance engine admits postv2 + │ and writes community-signed acceptances; opt-out deleteRemote + │ and confirmed account deletions purge at peers. + │ /admin/admissions* comes into existence. + ▼ +OUTBOUND_WORKERS=0 ← intent accumulates, nothing is POSTed. + │ + │ >0: delivery workers drain the queue onto the wire. + ▼ +OUTBOUND_DISABLED / _HOSTS / _COMMUNITIES / _ACTORS / OUTBOUND_DRY_RUN + ← per-scope parking, applied per delivery. +``` + +Three properties of that diagram are load-bearing and are the ones operators +get wrong: + +**There is no `OUTBOUND_ENABLED`.** Delivery starts only when +`OUTBOUND_WORKERS > 0` **AND** `CONSUMER_ENABLED` +(`cmd/tidepool/main.go:716` and `:843`, inside `startConsumer`, which is only +called under the consumer flag at `:537`). `OUTBOUND_WORKERS` also defaults to +**0**, via `intVarNonNegative` rather than `intVar`, precisely so that 0 is a +legal value and not a config error (`config.go:501`, `:688-702`). + +**Consumer-on / workers-0 is a designed staging step, not a broken state.** +With the consumer running, the real persisting enqueuer is always wired — it +writes `outbound_activities`/`outbound_deliveries` inside the consumer's gate +transaction, so intent past the gate is never dropped +(`cmd/tidepool/main.go:695-713`). Raising `OUTBOUND_WORKERS` later drains +whatever accumulated, immediately. Budget for that. + +**`OUTBOUND_DISABLED` parks, it does not fail.** A blocked delivery stays +`pending` and resumes when the switch clears; it is never poisoned and never +cancelled (`internal/outbound/worker.go:235-242`, +`internal/outbound/switches.go:5-12`). Engaging a kill switch loses nothing. +`OUTBOUND_DRY_RUN` parks the same way, after translating and logging +(`worker.go:243-248`). + +Scope matching, from `config.go:519-525`: + +| Switch | Matching | +|---|---| +| `OUTBOUND_DISABLED_HOSTS` | inbox host, **lowercased** on load and compared case-insensitively — a kill switch must fail closed on case | +| `OUTBOUND_DISABLED_COMMUNITIES` | community **AP id**, e.g. `https://lemmy.world/c/comicstrips`, **exact and case-sensitive**, compared against the delivery's ordering key (`worker.go:237-240`) | +| `OUTBOUND_DISABLED_ACTORS` | actor **DID**, exact and case-sensitive | + +There is no allowlist form of any of these. See +[Staged rollout](#5-staged-rollout) for what that means for a canary. + +--- + +## 3. The admin surface at incident time + +Full route table with request bodies: [README](README.md#the-admin-api). All +routes are bearer-authenticated (`Authorization: Bearer $ADMIN_TOKEN`) behind +one middleware group (`internal/ingest/follow.go:148-162`, +`internal/accept/admin.go:65-71`). Production publishes `127.0.0.1:8091` on the +box so admin calls skip Caddy entirely. + +Four behaviours worth knowing *before* you need them, because each one reads +like a fault and is not: + +**`POST /admin/communities/reconcile` → 501 means `FOLLOW_LIST_PATH` is +unset.** The follow reconciler is only constructed when the path is configured +(`cmd/tidepool/main.go:468-483`), and the handler nil-checks it +(`internal/ingest/follow.go:575-579`). Production *does* set +`FOLLOW_LIST_PATH=/repo/communities.yaml`, so a 501 there means the env changed, +not that reconciliation broke. + +**`/admin/admissions` and `/admin/admissions/readmit` → 404 means the consumer +is off.** Those routes are registered inside the `if cfg.ConsumerEnabled` block +(`cmd/tidepool/main.go:537-556`) because a force re-admit needs the acceptance +engine, which only exists there. A 404 is "the consumer is not running", never +"the endpoint is broken". Every other `/admin` route answers regardless. + +**`POST /admin/outbound/redrive` refuses an unscoped redrive.** With no +`activity`, no `community`, and no explicit `{"all":true}` it returns 400 +(`internal/ingest/follow.go:210-213`) — an unscoped redrive would re-attempt +every poisoned delivery at once, and a malformed body must not become a silent +fleet-wide replay. `POST /admin/outbound/cancel` likewise demands exactly one +of `actor` or `community` (`follow.go:236-239`). + +**`GET /admin/outbound` answers even with the consumer off.** The deliveries +store is wired unconditionally (`cmd/tidepool/main.go:454`), so an empty +`by_state` map is the truth about the queue, not a symptom of misconfiguration. + +```sh +T="Authorization: Bearer $ADMIN_TOKEN" +curl -s -H "$T" localhost:8091/admin/outbound # queue depth by state +curl -s -H "$T" localhost:8091/admin/divergence # one sweep, synchronously +curl -s -H "$T" localhost:8091/admin/metrics # tidepool* expvars +curl -s -H "$T" 'localhost:8091/admin/admissions?status=rejected' +``` + +--- + +## 4. coves.social Caddy — a CROSS-REPO change + +**There is no Caddyfile in this repository.** TLS for both `tdpl.io` and +`coves.social` terminates in the **Coves** Caddy (`coves-prod-caddy`), which +reaches this stack over the shared external `coves-prod-network`. The work +below is an edit to `~/Code/coves/Caddyfile` — a Coves-repo change, deployed on +the Coves side. + +### Why Caddy has to do this at all + +`AP_USER_ORIGIN=https://coves.social` puts the native users' ActivityPub surface +on the **Coves** hostname, served by the Tidepool process. Tidepool's Host +router sends requests whose `Host` is `coves.social` to the persona surface, and +everything under `tdpl.io` to the bridge (`internal/personas/hostrouter.go:89-115`). +Caddy currently proxies `coves.social` entirely to the AppView, so **none of +those AP paths reach Tidepool today.** + +The apex needs a content-negotiated split, and Tidepool deliberately cannot do +it itself: `internal/personas/instance.go:32-34` states in its own comment that +content negotiation is absent *because Caddy owns it in production*, and that a +peer sending `ld+json`, `activity+json`, or no `Accept` at all must still get +the actor. Peers fetch `GET /` to find the instance actor; browsers fetch +`GET /` to find the web app. Only the edge can tell them apart. + +### ⚠️ The Caddyfile inode trap — read this first + +**Edit the Caddyfile in place. Never replace the file.** The production Caddy +mounts the Caddyfile as a single-file bind mount, and Docker pins a single-file +mount to the **inode present when the container started**. Anything that +replaces the file rather than writing through it — `git checkout`, `mv`, +`sed -i`, most editors' atomic-save, `scp` of a new file — leaves the container +serving the **old** contents forever, while `cat` on the host shows the new +ones. This has bitten this project before. It is the same trap that made +Tidepool's own `communities.yaml` a *directory* mount +(`docker-compose.prod.yml:18-24`). + +Safe edits: `nano`/`vim` with backupcopy=yes, `cat > file`, `tee`. After any +edit, confirm the container sees it: + +```sh +docker exec coves-prod-caddy cat /etc/caddy/Caddyfile | grep -c '/ap/' +docker exec coves-prod-caddy caddy validate --config /etc/caddy/Caddyfile \ + --adapter caddyfile +docker exec coves-prod-caddy caddy reload --config /etc/caddy/Caddyfile \ + --adapter caddyfile +``` + +If the container's copy does not show your change, the inode moved: recreate +the Caddy container (`docker compose -f docker-compose.prod.yml up -d +--force-recreate caddy` in `/opt/coves`). + +### The change + +Inside the **existing `coves.social { … }` site block** — do not add a second +block for the same hostname — add the AP routes. `handle` blocks in one site +are mutually exclusive and sorted by path specificity, so +`/.well-known/webfinger` wins over the existing `/.well-known/*` static block, +and `/ap/*` is disjoint from everything already there. + +```caddyfile +coves.social { + # ── Tidepool's native-user AP surface (AP_USER_ORIGIN) ────────────── + # These paths belong to the bridge, not the AppView. More specific than + # the /.well-known/* static block below, so they win the handle sort. + # + # NOTE: no `header_up Host` on these proxies. Caddy v2 forwards the + # original Host by default, and Tidepool's Host router keys on it to + # choose the persona surface over the bridge surface + # (internal/personas/hostrouter.go). Rewriting Host to the upstream + # address would 421 every one of these requests. + handle /.well-known/webfinger { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + handle /.well-known/nodeinfo { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + handle /nodeinfo/2.0 { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + # /ap/actor/{did}, /ap/actor/{did}/outbox, /ap/object/*, /ap/activity/*, + # and POST /ap/inbox — the shared inbox for this origin. + handle /ap/* { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + + # ── Apex: split by Accept ─────────────────────────────────────────── + # `handle /` matches the apex EXACTLY (a Caddy path matcher is exact + # unless it ends in *), so this replaces only the bare "/" case that + # the catch-all used to serve. Same-name directives run in Caddyfile + # order, so the matched reverse_proxy is tried before the fallback. + # + # Two Accept lines, not one substring: Lemmy and Mastodon send + # application/activity+json, but the AS2 spec form is + # application/ld+json;profile="…activitystreams", which shares no + # useful substring with the first. Values for the SAME header field + # are OR'ed. + handle / { + @ap { + header Accept *application/activity+json* + header Accept *application/ld+json* + } + reverse_proxy @ap tidepool:80 { + header_up X-Real-IP {remote_host} + } + # Fallback: the web app, with the SAME upstream options as the + # catch-all below — copy them verbatim, header_up Host included + # (DPoP htu matching depends on it). + reverse_proxy appview:8080 { + health_uri /xrpc/_health + health_interval 30s + health_timeout 5s + header_up Host {host} + header_up X-Real-IP {remote_host} + header_up X-Forwarded-For {remote_host} + header_up X-Forwarded-Proto {scheme} + header_up X-Forwarded-Host {host} + } + } + + # … existing handle /.well-known/*, /client-metadata.json, /img/*, + # catch-all handle, headers, CSP, encode — all unchanged … +} +``` + +### Verify from outside + +```sh +# instance actor (Application), not the web app +curl -s -H 'Accept: application/activity+json' https://coves.social/ | jq '.type,.id' +# expect: "Application" "https://coves.social/" + +# the AS2 spelling must reach the same document +curl -s -H 'Accept: application/ld+json; profile="https://www.w3.org/ns/activitystreams"' \ + https://coves.social/ | jq '.type' + +# a browser must still get the web app +curl -sI -H 'Accept: text/html' https://coves.social/ | head -1 + +curl -s 'https://coves.social/.well-known/webfinger?resource=acct:alice@coves.social' | jq . +curl -s https://coves.social/.well-known/nodeinfo | jq . +curl -s https://coves.social/nodeinfo/2.0 | jq '.software.name' # "tidepool" +``` + +A 421 from any of these means the `Host` header was rewritten on the way +through. A 404 on webfinger for a user who has never federated is correct — +actors are minted lazily on first federating interaction. + +### Not changing + +`tdpl.io`, the per-instance wildcard blocks, and the on-demand catch-all are +untouched by v2. The `on_demand_tls ask` gate still points at +`http://tidepool:80/.well-known/tidepool-tls-ask`, which is served on the +bridge Host (`cmd/tidepool/main.go:165`) and is unaffected by anything above. + +--- + +## 5. Staged rollout + +Decision 19's canary. Each step is a separate `.env` edit plus +`docker compose -f docker-compose.prod.yml up -d tidepool`, because config is +read once at boot. + +### Step 0 — Caddy first + +Do section 4 **before** enabling the consumer. A native actor that federates +outward triggers Lemmy to fetch its webfinger, its actor document, and the +origin apex, all at `coves.social`. If Caddy is not routing those to Tidepool +yet, Lemmy gets the web app or a 404 and caches the failure. + +### Step 1 — consumer on, delivery off + +``` +CONSUMER_ENABLED=true +OUTBOUND_WORKERS=0 +``` + +Nothing reaches any peer. Watch for a day: + +```sh +curl -s -H "$T" localhost:8091/admin/metrics | jq '{ + cursor_age: .tidepool_consumer_cursor_age_seconds, + last_event_age: .tidepool_consumer_last_event_age_seconds, + connected: .tidepool_consumer_connected, + dead_letters: .tidepool_consumer_dead_letters, + reconnects: .tidepool_consumer_reconnects +}' +curl -s -H "$T" localhost:8091/admin/outbound # intent accumulating +curl -s -H "$T" 'localhost:8091/admin/admissions?status=rejected' | jq +``` + +`tidepool_consumer_cursor_age_seconds` and +`tidepool_consumer_last_event_age_seconds` are what make a *stalled* consumer +visible: the process stays up and the healthcheck stays green while events +quietly stop arriving (`cmd/tidepool/main.go:827-830`). A climbing +`tidepool_consumer_dead_letters` means events are failing and being parked for +the redriver, not lost. + +Two reading rules for those keys: + +- **The whole `tidepool_consumer_*` family only exists while the consumer + runs** — `PublishMetrics` is called inside `startConsumer` + (`cmd/tidepool/main.go:830`). Absent keys mean the consumer is off, not that + it is broken and silent. The `tidepool_divergence_*` gauges are the opposite: + registered at package init, so they are always present. +- **`tidepool_consumer_dead_letters: -1` is not a count.** It is the sentinel + for "storage could not be read" (`internal/consume/metrics.go:26-30`), chosen + because a `0` would claim the backlog is empty at exactly the moment nobody + can tell. + +Check the rejections before letting anything out. A misconfigured +`ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY`, a stale community mapping, or a +consent state you did not expect all surface here as `decisionCode`s while the +blast radius is still zero. + +### Step 2 — one community + +There is **no allowlist**, so a one-community canary is spelled as a denylist +over the rest of `communities.yaml`. With today's four entries, canarying +`!comicstrips@lemmy.world` means: + +``` +OUTBOUND_WORKERS=1 +OUTBOUND_DISABLED_COMMUNITIES=https://lemmy.world/c/selfhosted,https://lemmy.world/c/fediverse,https://lemmy.ml/c/linux +``` + +AP ids, exact case — not the `!name@host` spelling `communities.yaml` uses. +Confirm the ids you are about to paste rather than constructing them; the +`community` field of the list response **is** the AP group id +(`internal/ingest/follow.go:283-300`): + +```sh +curl -s -H "$T" localhost:8091/admin/communities | jq -r '.communities[].community' +``` + +**Adding a community to `communities.yaml` later does not add it to this +denylist.** The reconciler will subscribe it and it will start federating on +the next sweep. Whenever the follow list grows during a canary, extend +`OUTBOUND_DISABLED_COMMUNITIES` in the same change. + +Optionally precede this with `OUTBOUND_DRY_RUN=true` for one cycle: every +delivery is translated and logged, nothing is POSTed, and the parked deliveries +resume when you clear it. That validates the translator against real records +without touching a peer. + +### On announcement throttling — what actually exists + +Decision 19 asks for "deliberate throttling of initial actor announcements". +**Read this carefully, because the obvious knob is the wrong one.** + +`MINT_RATE_PER_MINUTE` / `MINT_BURST` (production overrides them to **10/20**, +`docker-compose.prod.yml`, because the public `plc.directory` 429s mint bursts +during community backfill) gate **inbound** DID minting only. They are wired +into `ingest.NewMintGate`, whose sole consumer is the materializer's minter +(`cmd/tidepool/main.go:271-276`, `:314`) — the path where an unseen *Lemmy* +author gets an atproto DID. + +**There is no rate limiter on the outbound side.** `internal/outbound` contains +no `rate.Limiter`, no sleep, and no throttle of any kind; the persona actors +the consumer mints do not pass through `mintGate` (`main.go:539` hands +`personasService` directly to `startConsumer`). The only levers on outbound +volume are: + +- `OUTBOUND_WORKERS` — worker **concurrency**, not a rate. Workers poll with a + 1s idle interval (`main.go:53`); with a full queue they run flat out. +- the scoped kill switches — which communities/actors/hosts may deliver at all. + +So the canary *is* the throttle: `OUTBOUND_WORKERS=1` plus a one-community +scope. Do not go looking for an announcement rate knob; nobody built one, and +this is filed under [Not implemented](#not-implemented). + +### Step 3 — widen + +Remove entries from `OUTBOUND_DISABLED_COMMUNITIES` one at a time, raising +`OUTBOUND_WORKERS` as the queue justifies. Before each widening: + +```sh +# poison depth — the queue's own verdict +curl -s -H "$T" localhost:8091/admin/outbound | jq .by_state + +# echo suppression: bridge-origin content must never re-materialize +curl -s -H "$T" localhost:8091/admin/metrics | jq '{ + mapped_object: .tidepool_echo_drops_mapped_object, + local_activity: .tidepool_echo_drops_local_activity, + local_actor: .tidepool_echo_drops_local_actor, + ancestor: .tidepool_echo_drops_ancestor_short_circuit +}' + +# the reconciliation report +curl -s -H "$T" localhost:8091/admin/divergence | jq '{counts, truncated}' +``` + +What each one means: + +- **`by_state.poisoned` climbing** — deliveries exhausting their retries. + Inspect, fix, then `redrive` **scoped** to the affected community. +- **echo drop counters rising steadily** — expected and healthy: our own + content arriving back from Lemmy and being correctly refused. A counter at + **zero** while native content is flowing is the alarming case; it means + classification is not firing and re-materialization is possible. +- **`tidepool_divergence_*` gauges** — the read-only reconciliation sweep + (`GET /admin/divergence` runs one synchronously). Watch + `acceptance_undelivered_stale` (a delivery pending past + `DIVERGENCE_ACCEPTANCE_STALE_AFTER`, default 12h), + `acceptance_undelivered_poisoned`, and `vote_recast_undelivered`. Also watch + `tidepool_divergence_sweep_failures` and + `tidepool_divergence_sweep_age_seconds`: a failed sweep publishes **nothing** + and the gauges keep their previous values + (`internal/ingest/divergence.go:584-591`), so a stale age with quiet gauges + is a *silent* failure mode. The sweep never repairs anything — it reports. + +### Rollback + +Rollback is a **kill switch, not an un-deploy**. In escalation order, each +step being one `.env` edit plus `up -d tidepool`: + +1. **`OUTBOUND_DISABLED_COMMUNITIES=`** — park one community. + Everything else keeps flowing; the parked deliveries resume when you clear + it. +2. **`OUTBOUND_DISABLED=true`** — park everything outbound. The consumer keeps + running and keeps recording intent; nothing reaches any peer. This is the + big red button and it is **lossless**. +3. **`OUTBOUND_WORKERS=0`** — stop the workers entirely. Equivalent effect to + (2) for delivery; prefer (2), because a parked delivery carries a recorded + reason and a stopped worker does not. +4. **`CONSUMER_ENABLED=false`** — stop consuming. Intent stops being recorded. + The consumer resumes from its stored cursor when re-enabled, so this is + recoverable, but it is the only step that stops *observing*, and + `/admin/admissions` disappears with it — taking your triage view down at the + moment you most want it. Prefer 1–3. + +Do **not** roll back by cancelling deliveries. `POST /admin/outbound/cancel` +is for consent withdrawal and community removal; a cancelled delivery is not +resumable the way a parked one is. + +Only redeploy the previous image if the failure is a code defect rather than a +federation outcome. The kill switches address the latter faster and without a +schema-version question. + +--- + +## 6. Not implemented + +Everything in this section is a real operational gap. None of it has a +mechanism in the code today. It is written down because an operator who +assumes one of these exists will look for it during the exact incident where +looking costs the most. + +### Key rotation — `BRIDGE_KEK` and per-actor RSA keys + +**No rotation path exists. Not partial, not manual, not scripted.** + +Every mention of `BRIDGE_KEK` in this repository is a warning, never a +procedure. `internal/config/config.go:49-53` documents the value; the sealing +itself is AES-256-GCM in `internal/identity/keys.go`. There is nothing that +re-seals existing ciphertext under a new key: the binary has exactly two +subcommands, `tidepool` and `tidepool migrate` +(`cmd/tidepool/main.go:69-78`). + +Blast radius of losing or changing it: every per-actor RSA signing key for +every bridged identity is sealed under it, plus the PLC **escrow rotation +key** (`internal/identity/keys.go:140-144`). Change the KEK and every one of +those ciphertexts becomes undecryptable — no bridged actor can sign, and the +escrow key that could recover the DIDs is itself sealed under the key you just +replaced. Approximately 950 identities. There is no recovery. + +*Naming trap:* `LoadOrCreateRotationKey` is **not** KEK rotation. It loads or +generates the did:plc escrow/recovery key — an atproto identity concept — +which is itself sealed under the KEK. Do not read that symbol as evidence that +rotation is implemented. + +What a real rotation would require, none of which exists: a key-version +column or KEK-id alongside each sealed blob; a dual-read custodian that tries +the new KEK then the old; an online re-seal pass over `bridged_actors` and +`service_keys`; and a cutover that retires the old KEK only after the pass +completes. The ciphertext does carry a one-byte version prefix +(`internal/identity/keys.go:127`), which is a hook someone could build on, but +nothing reads it as a key selector today. + +Per-actor **RSA** rotation is equally undefined: rotating an actor's key means +republishing `publicKey` in its actor document and having every peer that +cached it re-fetch, with no grace-overlap mechanism in the code to publish two +keys at once. + +**Until this is built, treat `BRIDGE_KEK` as immutable, and back it up +somewhere that survives the loss of the server.** + +### Backup and restore + +**No procedure exists, and no tooling.** `docker-compose.prod.yml:46` mounts +`./backups:/backups` into the Postgres container. Nothing writes to it. There +is no cron, no `pg_dump` wrapper, no restore drill, and no documented RPO/RTO. + +What is at risk, in order of irreplaceability: + +1. **`BRIDGE_KEK`** — lives in `/opt/tidepool/.env`, not in Postgres, and is + not covered by any database backup. Losing it is unrecoverable (above). +2. **`bridged_actors` / `service_keys`** — the sealed signing keys. Losing + these loses the identities even if the KEK survives. +3. **Repo blocks and commits** — the atproto repos themselves. Re-derivable + from upstream only by re-bridging, which mints new DIDs; the old at-uris do + not come back. +4. **`outbound_*`, `admissions`** — in-flight federation state. Losing it + double-sends or drops deliveries. + +Anything actually built here should be a separate task with a **restore +drill**, since an unverified backup is a claim, not a capability. + +### A divergence off switch + +**There is none.** The reconciliation sweep is wired unconditionally — the +constructor and `go divergence.Run(ctx)` sit outside every flag at +`cmd/tidepool/main.go:485-506` — and `Run` performs one sweep **immediately**, +before its first tick (`internal/ingest/divergence.go:561-565`). Setting +`DIVERGENCE_INTERVAL=0` does not disable it: `durationVar` rejects zero and +negative values and the process refuses to start +(`internal/config/config.go:666-668`). + +The only available lever is a **large interval** — `DIVERGENCE_INTERVAL=8760h` +quiets the background pass. The startup sweep still runs, once, on every boot, +and `GET /admin/divergence` still works. + +This is defensible: the sweep writes nothing, to peers or to our own tables, +which is exactly what makes an always-on schedule safe. But if a sweep is ever +implicated in an incident (lock pressure, a long multi-table scan — two of its +legs scan `outbound_activities` in full, per `FOLLOWUPS.md`), **the runbook +answer is "raise the interval and restart", not "disable it", because disabling +is not possible.** + +### Periodic vote re-seed + +**Nothing re-seeds vote aggregates on a schedule.** `SeedPostCounts` has one +caller, on the community-backfill path. + +This makes one documented self-healing claim conditional in a way that matters: +baseline-only voters can drift when a later flip or clear has no per-voter +baseline row to retract, and the standing note is that "a re-seed heals the +aggregate" (`FOLLOWUPS.md`, *Votes*). For a **quiet community that is never +backfilled again, the next re-seed is never.** The drift is permanent, silently, +and nothing reports it — the divergence sweep compares atproto state against +outbound state, not vote aggregates against the origin instance. + +The manual lever is a backfill of the affected community +(`POST /admin/communities/backfill`), which re-seeds as a side effect. That is +a workaround, not a scheduled heal, and it does other work besides. + +--- + +## 7. Support matrix + +| Peer | Status | Notes | +|---|---|---| +| **Lemmy 0.19.x** | **Targeted.** Strictness ceiling and e2e target | See the version discrepancy below | +| **PieFed** | Best-effort | No pinned instance in the harness; conformance is held by captured-wire fixtures. Its votes arrive from anonymous per-user actors, which is fine for tallies but means no per-voter identity | +| **Lemmy 1.0-beta** | Tracked, **not targeted** | Vote `FederationMode`, inbox collapsing, and `NoteWrapper` all change behaviour we depend on. No harness coverage | +| **Mastodon** | Incidental | The `security/v1` context is published so its parser accepts our `publicKey`, and its hosts appear in production handle subdomains. Not a target; not tested | + +### ⚠️ Version discrepancy — unresolved, needs a decision + +The tree contradicts itself about which Lemmy version is pinned: + +- `e2e/lemmy/Dockerfile:35` pins **`ARG LEMMY_VERSION=0.19.19`**. This is what + `make e2e` actually builds and tests against. +- `PLAN.md:430` (decision 19) says "**Lemmy 0.19.20** is the pinned strictness + ceiling and e2e target". `tasks/18-e2e-deploy.md:21` repeats it, and the + 0.19.20 source is cited as the authority for specific verified behaviours + across `tasks/13`, `14`, `15`, `17` — the `Delete`-summary convention, the + `check_bot_account` rule, `Instance`-enum strictness. + +**The matrix above says "0.19.x" deliberately, because writing either number +alone would be a claim the tree does not support.** The behaviours we verified +were read from 0.19.20 source; the behaviours we *test* are 0.19.19's. + +This is not resolvable from the docs — it needs a call: + +1. bump `e2e/lemmy/Dockerfile` to `0.19.20` so the tested version matches the + decided one (preferred; the e2e stack is owned by a separate task and this + file is out of scope for this change), **or** +2. amend decision 19 to name 0.19.19 as the pin and re-verify the source + claims against that tag. + +Until one of those happens, do not cite a specific patch version as "the +supported one" in operator-facing material. + +--- + +## 8. Known operational gaps carried forward + +Not new, but they shape what the runbook above can promise: + +- **Per-IP rate limits degrade to global ones behind Caddy.** Every in-process + limiter keys on `RemoteAddr` and ignores `X-Forwarded-For`, so from behind + the proxy they see only Caddy's container IP. Rate limiting at the edge is + the fix; tracked in `FOLLOWUPS.md`. +- **`ENVIRONMENT=production` has never been exercised end to end.** The e2e + harness runs in development mode (migrations-on-start, HTTP, private fetch, + strict lexicon validation). The production-only refusals — `BRIDGE_SCHEME=http`, + `ALLOW_PRIVATE_FETCH`, `AP_HOST_FALLTHROUGH_DEV`, and the `AP_USER_ORIGIN` + https rule — are unit-covered, not harness-covered. Section 1 exists because + of this. +- **Production lexicon validation records and writes rather than failing.** + A strict-first rollout should wait until + `tidepool_lexicon_validation_failures` stays at zero in production. diff --git a/FOLLOWUPS.md b/FOLLOWUPS.md index 3127a00..0b361d6 100644 --- a/FOLLOWUPS.md +++ b/FOLLOWUPS.md @@ -279,17 +279,59 @@ task documents and git history rather than this list. ## Production rollout -- Bluesky's public relay accepts new PDS hosts, but its documented default - allowance is only 100 accounts, 50 repo-stream events/second, 2,600/hour, - and 21,000/day. Tidepool mints one repo DID per bridged actor/community and - will exceed the account cap quickly. Arrange a relay limit increase before - broad subscriptions, or operate a suitably bootstrapped relay. +RESOLVED — Bluesky's public-relay account cap (100 accounts, 50 ev/s) is no +longer load-bearing: the self-hosted relay + Jetstream in +`docker-compose.prod.yml` are the app's ingest path and carry only our two PDS +hosts with an effectively unlimited account limit. `bsky.network` remains the +wider-visibility path only. Runbook: `SELF_HOSTED_RELAY.md`. + +DOCUMENTED, not resolved — the v2 deploy gaps below now have a written home in +`DEPLOY.md` (§6 "Not implemented") with their blast radius. Writing them down +is not building them; they stay open here: + +- **No `BRIDGE_KEK` / per-actor RSA rotation path.** Nothing re-seals existing + ciphertext under a new KEK, and the binary's only subcommand is `migrate` + (no args = serve). Changing the KEK orphans every bridged identity's signing key + *and* the PLC escrow rotation key sealed under it (~950 identities, no + recovery). Would need a key-version selector on each sealed blob, a + dual-read custodian, an online re-seal pass, and a cutover. +- **No backup or restore procedure.** `docker-compose.prod.yml` mounts + `./backups` into the Postgres container and nothing writes to it. Note + `BRIDGE_KEK` lives in `.env` and is not covered by any database backup at + all. Whatever is built needs a restore *drill* — an unverified backup is a + claim, not a capability. +- **No divergence off switch.** The sweep is wired unconditionally, sweeps once + at startup, and `DIVERGENCE_INTERVAL=0` is refused by `durationVar`. The only + lever is a large interval. Defensible while the sweep stays read-only; revisit + if it is ever implicated in lock pressure. +- **No outbound announcement rate limiter.** Decision 19 asks for "deliberate + throttling of initial actor announcements"; `internal/outbound` has no rate + limiter, and `MINT_RATE_PER_MINUTE`/`MINT_BURST` gate the **inbound** mint + path only (`ingest.NewMintGate`, materializer-only consumer). Today the + throttle is the canary itself: `OUTBOUND_WORKERS=1` plus a one-community + scope. +- **Kill switches and every other knob are boot-time only.** Config is read once + at `cmd/tidepool/main.go:109` with no reload signal, so engaging a kill switch + during an incident requires a container recreate. A SIGHUP reload (or an + admin-write switch table) would cut that latency. +- **Scoped kill switches are denylist-only.** There is no allowlist form, so a + one-community canary must enumerate every *other* subscribed community — and + adding a community to `communities.yaml` silently escapes an existing canary. + +Still open, unchanged: + - `ENVIRONMENT=production` has not been exercised end-to-end. The harness uses development mode for migrations-on-start, HTTP/private fetching, and strict lexicon validation. - Production lexicon validation currently records a metric and writes the record instead of failing it. A strict-first rollout should happen only after `tidepool_lexicon_validation_failures` remains zero in production. +- **Lemmy pin contradiction.** `e2e/lemmy/Dockerfile:35` pins + `LEMMY_VERSION=0.19.19`; PLAN decision 19 (`PLAN.md:430`) and + `tasks/18-e2e-deploy.md:21` name **0.19.20** as the pinned strictness ceiling + and e2e target, and the 0.19.20 source is cited as the authority for verified + behaviours across tasks 13/14/15/17. Either bump the Dockerfile or amend the + decision; until then, operator-facing docs say `0.19.x`. ## Sync surface @@ -319,8 +361,11 @@ task documents and git history rather than this list. deadlock-avoidance property has no true concurrency test. - Vote subject resolution occurs outside the mutation transaction, leaving a narrow race with deletion. -- Baseline-only voters can temporarily drift: a later flip or clear lacks a - per-voter baseline row to retract. A re-seed heals the aggregate. +- Baseline-only voters can drift: a later flip or clear lacks a per-voter + baseline row to retract. A re-seed heals the aggregate — but nothing + re-seeds on a schedule (see "Nothing re-seeds periodically" above), so for a + quiet community that is never backfilled again the drift is permanent and + unreported. "Temporarily" was the wrong word. ## Materializer and storage diff --git a/README.md b/README.md index 2f024fb..554a135 100644 --- a/README.md +++ b/README.md @@ -242,12 +242,24 @@ relays, or public Lemmy instances. ## Configuration -Environment variables with logged dev defaults (see -`internal/config/config.go`); everything below is **required in -production**: +All configuration is environment variables, read **once at process start** +(`config.Load`) — there is no reload signal, so changing any value below means +recreating the container. + +Two classes, and the difference matters at boot: + +- **Required in production** — `DATABASE_URL`, `LISTEN_ADDR`, + `BRIDGE_HOSTNAME`, `PLC_DIRECTORY_URL`, `BRIDGE_KEK`, `ADMIN_TOKEN`, + `AP_USER_ORIGIN`. These have *dev defaults only*; unset with + `ENVIRONMENT=production` the process refuses to start (`config: is + required in production`). Fail-closed on purpose: every one of them is baked + into identities or authority, where a defaulted guess is worse than no boot. +- **Tuning knobs** — everything else. Real defaults in every environment, + logged when applied. | Variable | Dev default | Meaning | |---|---|---| +| `ENVIRONMENT` | `development` | `development` or `production`; any other value is refused at boot. Development enables migrations-on-start, dev defaults, and strict lexicon validation. Production additionally *refuses* `BRIDGE_SCHEME=http`, `ALLOW_PRIVATE_FETCH`, `ALLOW_DEV_REQUEST_CRAWL`, and `AP_HOST_FALLTHROUGH_DEV` | | `DATABASE_URL` | local dev postgres | bridge state | | `LISTEN_ADDR` | `:8091` | HTTP bind address | | `BRIDGE_HOSTNAME` | `localhost` | public domain of the bridge; anchors handles and the PDS endpoint in minted DID docs | @@ -258,6 +270,8 @@ production**: | `USER_AGENT` | derived | outbound HTTP user agent | | `ALLOW_PRIVATE_FETCH` | off | dev-only: disables the SSRF egress guard (AP fetches **and** PLC directory requests) so localhost targets work | | `FIREHOSE_RETENTION` | `72h` | how long `firehose_events` rows are kept for `subscribeRepos` cursor replay (Go duration; a background pruner trims older events hourly) | +| `MAX_BLOB_BYTES` | `5242880` (5 MiB) | outer transport budget for remote media (avatars, banners, post images) the materializer downloads per blob. Individual lexicon slots impose tighter caps (avatars 1 MB); this is the ceiling over all of them. Fails closed — oversized media is dropped, never truncated | +| `PROFILE_REFRESH_TTL` | `24h` | how stale a bridged actor's materialized profile may get before the materializer re-fetches it. `Update{Person\|Group}` refreshes immediately regardless — this covers what Lemmy never federates (bio edits, and so the `#nobridge` marker) | | `RELAY_HOSTS` | *(optional)* | comma-separated relays to send `com.atproto.sync.requestCrawl` to on startup (each retried on a bounded budget — the relay calls back into `describeServer` before subscribing, which can race process start); in development the request is logged, never sent, unless `ALLOW_DEV_REQUEST_CRAWL` opts in | | `ALLOW_DEV_REQUEST_CRAWL` | off | dev-only: actually SEND `requestCrawl` to `RELAY_HOSTS` in development (exists for the e2e stack's local BigSky); refused in production, where sending is already the behavior | | `ADMIN_TOKEN` | `dev-admin-token` | bearer token protecting the `/admin` API | @@ -280,13 +294,51 @@ production**: | `SYNC_MAX_SUBSCRIBERS` | `100` | concurrent `subscribeRepos` connection cap | | `FOLLOW_LIST_PATH` | *(optional)* | declarative follow list (see below); unset = the `/admin` API is the only subscription control | | `FOLLOW_LIST_INTERVAL` | `15m` | follow-list reconciler sweep cadence | -| `CONSUMER_ENABLED` | **off** | turns on the task-14 Jetstream consumer (native users' opt-outs, profiles, posts, comments, votes → durable outbound state). Default off: it writes durable state and hands work to delivery seams that are still stubbed until tasks 15–17 land, so a deployment that has not been wired end to end should not silently start accumulating it. Enabling it now runs the pipeline with a logging-noop enqueuer (nothing is delivered), a nil acceptance engine (postv2 skipped), and nil destructive/terminal tiers (opt-out `deleteRemote` and account deletions recorded, not acted on) | +| `DIVERGENCE_INTERVAL` | `15m` | cadence of the reconciliation sweep (task 17e) that compares atproto state against outbound state and publishes the `tidepool_divergence_*` gauges. **Not an on/off switch:** the sweep is wired unconditionally, runs once at startup before its first tick, and `0` is refused — it is read-only (it reports, never repairs), which is what makes an always-on schedule safe. `GET /admin/divergence` runs one on demand | +| `DIVERGENCE_ACCEPTANCE_STALE_AFTER` | `12h` | how long a pending delivery may sit before the sweep reports its acceptance as **stale**. The report's one crying-wolf knob — shorter and every in-flight post is a finding, longer and a queue that stopped this morning is not in tonight's report. 12h is derived from the retry schedule (~2–3h to poison) plus the causal wait budget (6h), not picked | +| `CONSUMER_ENABLED` | **off** | turns on the Jetstream consumer (task 14): native users' opt-outs, profiles, posts, comments and votes flowing outward. Default off because it writes durable outbound state, and because a deployment that has not been canaried should not start accumulating it — not because the seams behind it are stubbed. They are wired: with it on, the **real** enqueuer persists outbound intent, the acceptance engine admits postv2 and writes community-signed acceptances, and opt-out `deleteRemote` / confirmed account deletions actually purge at peers. It also gates two other things — the `OUTBOUND_WORKERS` AND, and whether `/admin/admissions*` exists at all | | `JETSTREAM_URL` | *(optional)* | the self-hosted Jetstream the consumer subscribes to (`ws://` or `wss://`); **required** when `CONSUMER_ENABLED`, and validated at boot whenever set so a typo fails fast instead of becoming a reconnect loop. May be staged ahead of the flag | - -## Subscribing to communities (admin API) - -Community subscriptions are operator-driven, over bearer-token-protected -endpoints (`Authorization: Bearer $ADMIN_TOKEN`): +| `OUTBOUND_WORKERS` | `0` (**off**) | how many delivery workers run. **There is no `OUTBOUND_ENABLED`:** delivery starts only when this is `>0` *and* `CONSUMER_ENABLED`. With the consumer on and this at `0`, outbound intent still accumulates durably and nothing is POSTed — which is the intended staging step, not a broken state. Raising it drains the accumulated backlog immediately | +| `OUTBOUND_DISABLED` | off | global delivery kill switch. A blocked delivery is **parked** — it stays `pending` and resumes when the switch clears — never poisoned, never cancelled. Engaging it loses nothing; it stops the wire | +| `OUTBOUND_DISABLED_HOSTS` | *(empty)* | comma-separated inbox **hosts** to park. Lowercased on load and compared case-insensitively — a kill switch must fail closed on case | +| `OUTBOUND_DISABLED_COMMUNITIES` | *(empty)* | comma-separated community **AP ids** to park (`https://lemmy.world/c/comicstrips`), matched **exactly and case-sensitively** against the delivery's ordering key. There is no allowlist form: a one-community canary is spelled by disabling every other community | +| `OUTBOUND_DISABLED_ACTORS` | *(empty)* | comma-separated actor **DIDs** to park, exact match | +| `OUTBOUND_DRY_RUN` | off | translate and log every delivery, POST nothing. Parks like the kill switches, so nothing is lost — the difference is that the translation ran and is in the log | +| `ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY` | `50` | acceptance-engine flood cap: how many posts one native author may have accepted into one bridged community. `0` = unlimited. Tidepool signs the community's acceptance, so this bounds what it vouches for | + +## The admin API + +Every route below is bearer-protected (`Authorization: Bearer $ADMIN_TOKEN`) +and mounted on the bridge's own `Host`. Production publishes port +`127.0.0.1:8091` on the box for exactly this — admin calls do not round-trip +through Caddy. + +| Route | Purpose | Answers 501/404 when | +|---|---|---| +| `POST /admin/communities` | subscribe (WebFinger → Group → materialize → signed `Follow`) | — | +| `DELETE /admin/communities` | unsubscribe (`Undo{Follow}`; records kept, content stops) | — | +| `GET /admin/communities` | list subscriptions and their state | — | +| `POST /admin/communities/backfill` | on-demand outbox backfill | backfill unconfigured (**501**) | +| `POST /admin/communities/reconcile` | force one follow-list sweep | **501** — `FOLLOW_LIST_PATH` is unset. The reconciler is only built when the path is set, so this is "no follow list configured", not a broken endpoint | +| `POST /admin/reemit` | re-emit a repo's records as delete+create pairs (relay cold-start gap) | repo manager unconfigured (**501**) | +| `POST /admin/objects/sweep-deleted` | origin-verified cleanup of missed deletes | sweeper unconfigured (**501**) | +| `GET /admin/divergence` | run one reconciliation sweep synchronously and return the report | always wired in a normal deployment | +| `GET /admin/outbound` | delivery queue depth by state (`pending`/`poisoned`/`cancelled`/…) | — the store is always wired, so this answers even with the consumer off and workers at 0. An empty queue then is the truth, not a misconfiguration | +| `POST /admin/outbound/redrive` | reset poisoned deliveries to pending | — | +| `POST /admin/outbound/cancel` | park an actor's or a community's pending deliveries as cancelled | — | +| `GET /admin/admissions` | list acceptance decisions with `status`, `decisionCode`, `evaluatedCid`; filter by `?status=`/`?community=` | **404** — these routes are registered **only** when `CONSUMER_ENABLED`. A 404 here means the consumer is off, not that the endpoint is broken | +| `POST /admin/admissions/readmit` | force re-admit a rejected post | **404**, same reason | +| `GET /admin/metrics` | expvar counters filtered to the `tidepool*` prefix (never Go's `cmdline`/`memstats`) | — | + +`POST /admin/outbound/redrive` **refuses an unscoped redrive**: send +`{"activity":"…"}`, `{"community":"…"}`, or an explicit `{"all":true}`. A +missing filter is a 400, never a silent fleet-wide replay of every poisoned +delivery. `POST /admin/outbound/cancel` requires exactly one of +`{"actor":"did:…"}` or `{"community":"https://…"}`. + +### Subscribing to communities + +Community subscriptions are operator-driven: ```sh # follow: WebFinger → fetch Group → materialize community → signed Follow @@ -622,6 +674,35 @@ Note TLS: a single wildcard certificate only covers one label level, while bridged handles sit two levels below `BRIDGE_HOSTNAME` — terminate TLS with on-demand certificate issuance (e.g. Caddy) or per-instance wildcard certs. +## Production deploy + +The runbook is **[`DEPLOY.md`](DEPLOY.md)**: the boot-time config gate, the v2 +flag topology (`CONSUMER_ENABLED` → `OUTBOUND_WORKERS` → kill switches), the +cross-repo Caddy change that puts the native-user AP surface on +`coves.social`, the staged canary and its rollback order, and — explicitly — +the things that have **no** mechanism today (KEK/RSA rotation, backup/restore, +a divergence off switch, a periodic vote re-seed). +[`SELF_HOSTED_RELAY.md`](SELF_HOSTED_RELAY.md) covers the relay + Jetstream +ingest path. + +One thing to know before deploying HEAD onto an existing box: `AP_USER_ORIGIN` +is required in production and was not in the pre-v2 env file, so the process +fails at startup rather than coming up half-configured. That is the intended +behaviour — see DEPLOY.md §1. + +### Support matrix + +| Peer | Status | +|---|---| +| **Lemmy 0.19.x** | targeted — the e2e target and strictness ceiling | +| **PieFed** | best-effort, behind captured-wire conformance (votes arrive from anonymous per-user actors: fine for tallies, no per-voter identity) | +| **Lemmy 1.0-beta** | tracked, not targeted (vote `FederationMode`, inbox collapsing, `NoteWrapper`) | +| **Mastodon** | incidental — the `security/v1` context is published so its parser accepts our `publicKey`; not a target, not tested | + +The patch version is deliberately written as `0.19.x`: `e2e/lemmy/Dockerfile` +pins `0.19.19` while decision 19 and the task docs name `0.19.20`. That +contradiction is unresolved — see DEPLOY.md §7. + ## License Tidepool is licensed under the [GNU Affero General Public License v3.0](LICENSE) diff --git a/SELF_HOSTED_RELAY.md b/SELF_HOSTED_RELAY.md index 9d273fb..e356854 100644 --- a/SELF_HOSTED_RELAY.md +++ b/SELF_HOSTED_RELAY.md @@ -30,7 +30,16 @@ at bsky.network, so native Coves signup #101 would silently vanish from the app exactly like the bridged commenters did. Design decisions and their reasons live as comments on the compose -services; this file is the runbook. +services; this file is the runbook for the **ingest** path only. The v2 +deploy — config gate, the `CONSUMER_ENABLED`/`OUTBOUND_WORKERS` staging +order, the coves.social Caddy change, rollout and rollback — is +[`DEPLOY.md`](DEPLOY.md). + +Note the direction: this file describes the relay and Jetstream that feed +**Coves**. Tidepool's own consumer reads the same Jetstream from the other +side, over `JETSTREAM_URL` +(`ws://tidepool-prod-jetstream:8080/subscribe`, :8080 = `JETSTREAM_ADDR`; +:6060 is the debug listener and serves no `/subscribe`). ## Not in scope here (Phase 2+, Coves repo) diff --git a/docker-compose.prod.yml b/docker-compose.prod.yml index 1249b33..9481667 100644 --- a/docker-compose.prod.yml +++ b/docker-compose.prod.yml @@ -23,6 +23,13 @@ # plain `git pull` is enough; the reconciler picks the new file up on its # next sweep (FOLLOW_LIST_INTERVAL, default 15m) — no container restart. # +# v2 env: the `tidepool` service below gained the task-13/14/15/16/17 knobs +# (AP_USER_ORIGIN, CONSUMER_ENABLED, JETSTREAM_URL, OUTBOUND_*, the admission +# quota, DIVERGENCE_*). AP_USER_ORIGIN is REQUIRED in production, so deploying +# HEAD onto a pre-v2 /opt/tidepool/.env fails at boot rather than coming up +# half-configured. The staged-rollout order, the kill switches and the +# coves.social Caddy changes are in DEPLOY.md. +# # Ops note (README "Sync surface"): the bridge's per-IP rate limiters key on # RemoteAddr and see only Caddy's container IP from behind the proxy — # per-IP limits degrade to global ones until rate limiting exists at the @@ -104,6 +111,84 @@ services: # instead (SELF_HOSTED_RELAY.md). RELAY_HOSTS: ${RELAY_HOSTS:-https://bsky.network} FOLLOW_LIST_PATH: /repo/communities.yaml + + # --- v2 (tasks 13-17): the native-user surface and the outbound half --- + # + # REQUIRED IN PRODUCTION. config.stringVar refuses an unset value outside + # development (internal/config/config.go:797-806), so a HEAD deploy onto + # an env file without this dies at boot with + # config: AP_USER_ORIGIN is required in production + # — fail-closed, and deliberately so: this origin is baked into every + # actor_id this deployment mints, so a wrong value is not a restart away + # from fixed, it is a set of federated identities pointing at the wrong + # place forever. Must be https in production (config.go:782-784) and must + # not be BRIDGE_HOSTNAME or a subdomain of it (config.go:786-790), which + # would shadow the bridged-handle namespace. + # + # Serving derives every URL from the stored actor_id, never from this + # value — so changing it later does NOT migrate existing actors, it only + # changes where NEW ones are minted. + AP_USER_ORIGIN: ${AP_USER_ORIGIN:-https://coves.social} + + # The Jetstream consumer: native Coves users' records flowing OUTWARD. + # Default off. Staged here ahead of the flag on purpose — JETSTREAM_URL is + # validated whenever it is set (config.go:481-489), not only when the + # consumer is on, so a typo is a boot failure with a clear message today + # rather than a silent reconnect loop on the day someone flips the switch. + # + # Container NAME, not the `jetstream` service alias, for the same reason + # DATABASE_URL uses tidepool-prod-postgres: this service also joins + # coves-prod-network, and a bare alias is only unambiguous until the Coves + # stack grows a service by the same name. :8080 is JETSTREAM_ADDR below; + # :6060 is its debug listener and serves no /subscribe. + CONSUMER_ENABLED: ${CONSUMER_ENABLED:-false} + JETSTREAM_URL: ${JETSTREAM_URL:-ws://tidepool-prod-jetstream:8080/subscribe} + + # Delivery workers. 0 = OFF, and it is AND-gated on CONSUMER_ENABLED + # (cmd/tidepool/main.go:716, 843) — there is no OUTBOUND_ENABLED. With the + # consumer on and workers at 0 the real enqueuer still writes + # outbound_activities/deliveries, so intent accumulates durably and starts + # draining the moment workers are raised. That is the intended staging + # order: consumer first, watch the queue, then workers. + OUTBOUND_WORKERS: ${OUTBOUND_WORKERS:-0} + + # Kill switches (decision 19). A blocked delivery is PARKED — it stays + # pending and resumes when the switch clears (internal/outbound/worker.go + # :236-242) — never poisoned, never cancelled, so engaging one loses + # nothing. DRY_RUN parks too, after translating and logging. + # + # ALL OF THESE ARE READ ONCE AT BOOT (config.Load, main.go:109); there is + # no reload signal. Changing one means editing /opt/tidepool/.env and + # recreating the service. + # + # HOSTS is lowercased on load; COMMUNITIES and ACTORS are matched EXACTLY + # and case-sensitively (config.go:519-525). Communities are AP ids + # ("https://lemmy.world/c/comicstrips"), matched against the delivery's + # ordering key; actors are DIDs. There is no allowlist form — a canary is + # expressed by disabling every community except the one under test. + OUTBOUND_DRY_RUN: ${OUTBOUND_DRY_RUN:-false} + OUTBOUND_DISABLED: ${OUTBOUND_DISABLED:-false} + OUTBOUND_DISABLED_HOSTS: ${OUTBOUND_DISABLED_HOSTS:-} + OUTBOUND_DISABLED_COMMUNITIES: ${OUTBOUND_DISABLED_COMMUNITIES:-} + OUTBOUND_DISABLED_ACTORS: ${OUTBOUND_DISABLED_ACTORS:-} + + # The acceptance engine's per-author-per-community flood cap (0 = + # unlimited). Tidepool SIGNS the community's acceptance, so it must not + # let one native account flood a community it vouches for. + ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY: ${ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY:-50} + + # The reconciliation sweep (task 17e). It is wired UNCONDITIONALLY + # (main.go:491-506) and sweeps once at startup before its first tick, so + # this interval is not an on/off switch — durationVar refuses 0 or + # negative (config.go:666-668). It writes nothing, to peers or to our own + # tables, which is what makes an always-on schedule safe. The only way to + # quiet it is a large interval; see DEPLOY.md "Not implemented". + DIVERGENCE_INTERVAL: ${DIVERGENCE_INTERVAL:-15m} + # How long a pending delivery may sit before the sweep calls its + # acceptance stale. The report's one crying-wolf knob; 12h is derived from + # the retry schedule plus the causal wait budget (internal/ingest/ + # divergence.go:72-81), not picked. + DIVERGENCE_ACCEPTANCE_STALE_AFTER: ${DIVERGENCE_ACCEPTANCE_STALE_AFTER:-12h} volumes: - ./:/repo:ro networks: