From 18597ccb2ac255d8345fba5342d7cff4431f7967 Mon Sep 17 00:00:00 2001 From: Bretton Date: Fri, 14 Aug 2026 23:07:09 -0700 Subject: [PATCH] docs(deploy): a runbook that describes the system as it is (18b) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Production could not boot. `stringVar` errors in production when a variable is unset, `AP_USER_ORIGIN` is read through it, and prod compose set none of the v2 environment — so today's compose applied to today's HEAD stops at startup with "config: AP_USER_ORIGIN is required in production". Failing closed is correct and stays; the fix is that the compose file now carries the variables the binary actually requires. A trailing slash is also fatal, since CanonicalizeOrigin refuses any path, and `https://coves.social/` is the spelling most people would type. The governing rule for this pass was to document what the code does rather than what the plan intended, and applying it changed several answers. MINT_RATE_PER_MINUTE IS NOT AN ANNOUNCEMENT THROTTLE, though this commit was briefed to describe it as one. The gate is wired into the materializer and governs INBOUND DID minting for bridged Lemmy actors; the consumer's persona minter is handed to startConsumer directly and never passes through it, and internal/outbound contains no limiter, sleep, or throttle at all. So the runbook names it for what it does, and says plainly that the canary IS the throttle: one community, one worker. Recorded as missing rather than described, because an operator reaching for a lever mid-rollout is exactly who a wrong runbook hurts. Two related facts shape that canary and are now written down. The scoped kill switches are DENYLIST-ONLY — there is no allowlist form — so canarying one community means enumerating the others by AP id, and a community added to the follow list later silently escapes the canary. And configuration is read once at startup with no reload path, so every switch costs a container recreate. No AP path currently reaches Tidepool. The coves.social site sends /.well-known/* to a static file server and everything else to the appview, so the Caddy section is written as a cross-repo instruction set against the real Coves Caddyfile — webfinger, nodeinfo, /ap/*, and an apex Accept matcher, since instance.go says in its own comment that content negotiation is deliberately Caddy's job. It carries the inode warning, and a hard note that no `header_up Host` may be set on those proxies because the host router keys on r.Host. Documented as MISSING rather than described, each with its blast radius: KEK and signing-key rotation (no path exists, and the one function whose name suggests it manages the PLC escrow key, itself sealed under the KEK), backup and restore (the mount exists, nothing writes it, and no database backup covers the KEK in .env), a divergence off switch (the sweep is unconditional, sweeps at startup, and durationVar refuses 0 — a large interval is the only lever), and periodic vote re-seed. Also corrected two claims already in the README: that everything in the config table is required in production (only the stringVar values are), and that flipping CONSUMER_ENABLED lands on stubs from tasks 15-17 (they are wired, which makes the old text actively misleading about what the flag does). And JETSTREAM_URL's example pointed at a port nothing serves. Unresolved and flagged rather than papered over: the e2e image pins Lemmy 0.19.19 while decision 19 and the task docs name 0.19.20 as the target and strictness ceiling, and the behaviours those tasks verified were read from 0.19.20 source. The support matrix says 0.19.x pending that call. Co-Authored-By: Claude Opus 5 (1M context) --- .env.prod.example | 79 ++++- DEPLOY.md | 701 ++++++++++++++++++++++++++++++++++++++++ FOLLOWUPS.md | 59 +++- README.md | 99 +++++- SELF_HOSTED_RELAY.md | 11 +- docker-compose.prod.yml | 85 +++++ 6 files changed, 1007 insertions(+), 27 deletions(-) create mode 100644 DEPLOY.md diff --git a/.env.prod.example b/.env.prod.example index f6ae7b0..27d8513 100644 --- a/.env.prod.example +++ b/.env.prod.example @@ -16,9 +16,29 @@ POSTGRES_PASSWORD=CHANGE_ME # proxying CDN edge breaks. BRIDGE_HOSTNAME=tdpl.io +# Origin the NATIVE Coves users' ActivityPub actors are served under — the +# second origin on the bridge's listener, split from BRIDGE_HOSTNAME by a Host +# router. REQUIRED IN PRODUCTION: unset, the process refuses to start with +# config: AP_USER_ORIGIN is required in production +# (this is why a HEAD deploy onto a pre-v2 .env fails at boot — fail closed). +# Must be https here, and must not be BRIDGE_HOSTNAME or a subdomain of it. +# +# Baked into every actor_id minted under it. Serving reads the stored actor_id, +# never this value, so changing it does NOT move existing actors — it only +# changes where new ones appear. Treat it as permanent. +# Caddy must proxy this origin's AP paths to the bridge; see DEPLOY.md. +# +# NO TRAILING SLASH. It must be a bare scheme://host — any path (including a +# lone "/"), query, fragment, or userinfo is refused at boot. +AP_USER_ORIGIN=https://coves.social + # 32-byte key-encryption key sealing per-actor signing keys at rest # (AES-256-GCM). Generate with: openssl rand -hex 32 # Losing it means losing every bridged repo's signing keys — back it up. +# THERE IS NO ROTATION PATH. Nothing in the codebase re-seals existing +# ciphertext under a new KEK; changing this value orphans every sealed key +# (every bridged identity and the PLC escrow rotation key) with no recovery. +# See DEPLOY.md "Not implemented" before you touch it. BRIDGE_KEK=CHANGE_ME # Bearer token for the /admin API. Generate with: openssl rand -hex 32 @@ -38,17 +58,56 @@ RELAY_HOSTS=https://bsky.network RELAY_ADMIN_PASSWORD=CHANGE_ME # Task-14 Jetstream consumer: native Coves users' opt-outs, profiles, posts, -# comments and votes flowing outward. DEFAULT OFF and left off until the -# outbound-delivery/acceptance/echo-moderation seams (tasks 15-17) land: with -# it on, the pipeline runs but delivery is a logging noop, postv2 is skipped -# (no acceptance engine), and opt-out deleteRemote + account deletions are only -# recorded, not acted on. Flip to true (and set JETSTREAM_URL) only once those -# seams are wired. +# comments and votes flowing outward. DEFAULT OFF. The seams it feeds are now +# all wired — outbound delivery (15), the acceptance engine (16), echo +# suppression and moderation (17) — so turning it on runs the REAL pipeline: +# postv2 is admitted and a community-signed acceptance written, opt-out +# deleteRemote and confirmed account deletions actually purge at peers. It is +# still default-off because it writes durable outbound state, and because +# nothing about being able to run it makes it safe to run it un-canaried. +# Read the staged rollout in DEPLOY.md before flipping this. CONSUMER_ENABLED=false # The self-hosted Jetstream the consumer subscribes to (ws:// or wss://). # REQUIRED when CONSUMER_ENABLED=true; validated at boot whenever set, so a bad -# value fails fast instead of becoming a silent reconnect loop. Points at the -# jetstream service in docker-compose.prod.yml (host-local). Safe to stage here -# ahead of flipping CONSUMER_ENABLED. -# JETSTREAM_URL=ws://jetstream:6008/subscribe +# value fails fast instead of becoming a silent reconnect loop. Safe to stage +# ahead of flipping CONSUMER_ENABLED — which is the point of staging it. +# +# Points at the jetstream service in docker-compose.prod.yml, which listens on +# :8080 (JETSTREAM_ADDR). :6060 is its debug/metrics listener and serves no +# /subscribe. Compose already defaults this; uncomment only to override. +# JETSTREAM_URL=ws://tidepool-prod-jetstream:8080/subscribe + +# Delivery workers. 0 = OFF, and AND-gated on CONSUMER_ENABLED — there is no +# OUTBOUND_ENABLED. With the consumer on and this at 0, outbound intent +# accumulates durably and nothing is POSTed; raising it starts draining that +# backlog immediately, so raise it deliberately. +OUTBOUND_WORKERS=0 + +# Kill switches. A blocked delivery is PARKED — still pending, resumed when the +# switch clears — never dropped. Read once at boot; changing one needs a +# service recreate, not a signal. +# OUTBOUND_DRY_RUN translate + log, POST nothing +# OUTBOUND_DISABLED global stop +# OUTBOUND_DISABLED_HOSTS comma-separated inbox hosts (lowercased) +# OUTBOUND_DISABLED_COMMUNITIES comma-separated community AP ids, EXACT +# match ("https://lemmy.world/c/comicstrips") +# OUTBOUND_DISABLED_ACTORS comma-separated actor DIDs, EXACT match +# There is no allowlist form: a one-community canary is spelled by disabling +# every OTHER community in the follow list. See DEPLOY.md. +# OUTBOUND_DRY_RUN=false +# OUTBOUND_DISABLED=false +# OUTBOUND_DISABLED_HOSTS= +# OUTBOUND_DISABLED_COMMUNITIES= +# OUTBOUND_DISABLED_ACTORS= + +# Acceptance-engine flood cap: how many posts one native author may have +# accepted into one bridged community. 0 = unlimited. Tidepool signs the +# community's acceptance, so this bounds what it vouches for. +# ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY=50 + +# Reconciliation sweep cadence. NOT an on/off switch: the sweep is wired +# unconditionally, runs once at startup, and 0 is refused. It is read-only +# (reports divergence, repairs nothing), so leaving it alone is correct. +# DIVERGENCE_INTERVAL=15m +# DIVERGENCE_ACCEPTANCE_STALE_AFTER=12h diff --git a/DEPLOY.md b/DEPLOY.md new file mode 100644 index 0000000..0e9b0e3 --- /dev/null +++ b/DEPLOY.md @@ -0,0 +1,701 @@ +# Deploying Tidepool v2 + +The production runbook for the v2 surface: native Coves users federating +outward (tasks 13–17), on top of the v1 inbound bridge that is already running. + +**This document describes what is in the code, not what was planned.** Where a +lever does not exist, it says so under [Not implemented](#not-implemented) +rather than describing a plausible one. A runbook that names a knob nobody +built is worse than one admitting the knob is missing, because it sends an +operator hunting for it during an incident. Every claim below cites the file it +came from; if you change the code, change the citation. + +Related: [`SELF_HOSTED_RELAY.md`](SELF_HOSTED_RELAY.md) (the relay + Jetstream +ingest path), [`README.md`](README.md#configuration) (the full config table and +admin route table). + +--- + +## 1. Read this before deploying HEAD + +**A HEAD deploy onto a pre-v2 `/opt/tidepool/.env` does not start.** It exits +during `config.Load` with: + +``` +config: AP_USER_ORIGIN is required in production +``` + +`AP_USER_ORIGIN` is read through `stringVar`, which falls back to a dev default +in development and returns an error in production +(`internal/config/config.go:797-806`, read at `:414`). The pre-v2 env file has +no such variable. + +**This is correct behaviour and must not be "fixed" by defaulting it away.** +The origin is baked into every `actor_id` this deployment mints; serving then +derives every URL from the *stored* `actor_id` and never from config +(`internal/personas/serving.go:20-23`). A wrong value is therefore not a +restart away from repaired — it is a set of federated identities pointing at +the wrong place, permanently. Two further boot-time guards exist for the same +reason: it must be `https` in production (`config.go:782-784`), and its host +must not be `BRIDGE_HOSTNAME` or a subdomain of it, which would shadow the +bridged-handle namespace (`config.go:786-790`). + +**It must also be a bare origin — no trailing slash.** +`personas.CanonicalizeOrigin` refuses any path, query, fragment, or userinfo +(`internal/personas/origin.go:42-46`), and `https://coves.social/` has a path +of `/`. That is a boot failure, not a normalization. + +`docker-compose.prod.yml` now supplies `AP_USER_ORIGIN` with a +`${AP_USER_ORIGIN:-https://coves.social}` default, so a `git pull` plus the +normal deploy command is sufficient. Set it explicitly in `.env` anyway if this +deployment is not tdpl.io/coves.social. + +### Minimum `.env` delta + +Nothing else is *required* — every other v2 variable has a real default in +every environment. See `.env.prod.example` for the full annotated set. + +```sh +# required; the rest of the v2 knobs default safely +AP_USER_ORIGIN=https://coves.social +``` + +### Deploy command + +Unchanged, and still targeted — never a bare `up -d`: + +```sh +cd /opt/tidepool && git pull +docker compose -f docker-compose.prod.yml up -d --build tidepool +``` + +The migrate one-shot and the server share one locally-built image and the +server only starts after `service_completed_successfully` on the migration, so +a failed migration holds the old container up rather than starting a new one +against an unmigrated schema. + +Confirm the boot actually got past config: + +```sh +docker logs --tail 50 tidepool-prod | grep -E 'listening|config:' +curl -sf localhost:8091/xrpc/_health +``` + +--- + +## 2. The v2 flag topology + +Everything v2 is off by default and turns on in one order. Nothing here is a +runtime toggle: **config is read exactly once**, at +`cmd/tidepool/main.go:109`, and there is no reload signal — changing any value +means editing `.env` and recreating the service. + +``` +CONSUMER_ENABLED=false ← nothing v2 runs. Today's state. + │ + │ ON: the Jetstream consumer dials JETSTREAM_URL. The REAL enqueuer + │ persists outbound intent; the acceptance engine admits postv2 + │ and writes community-signed acceptances; opt-out deleteRemote + │ and confirmed account deletions purge at peers. + │ /admin/admissions* comes into existence. + ▼ +OUTBOUND_WORKERS=0 ← intent accumulates, nothing is POSTed. + │ + │ >0: delivery workers drain the queue onto the wire. + ▼ +OUTBOUND_DISABLED / _HOSTS / _COMMUNITIES / _ACTORS / OUTBOUND_DRY_RUN + ← per-scope parking, applied per delivery. +``` + +Three properties of that diagram are load-bearing and are the ones operators +get wrong: + +**There is no `OUTBOUND_ENABLED`.** Delivery starts only when +`OUTBOUND_WORKERS > 0` **AND** `CONSUMER_ENABLED` +(`cmd/tidepool/main.go:716` and `:843`, inside `startConsumer`, which is only +called under the consumer flag at `:537`). `OUTBOUND_WORKERS` also defaults to +**0**, via `intVarNonNegative` rather than `intVar`, precisely so that 0 is a +legal value and not a config error (`config.go:501`, `:688-702`). + +**Consumer-on / workers-0 is a designed staging step, not a broken state.** +With the consumer running, the real persisting enqueuer is always wired — it +writes `outbound_activities`/`outbound_deliveries` inside the consumer's gate +transaction, so intent past the gate is never dropped +(`cmd/tidepool/main.go:695-713`). Raising `OUTBOUND_WORKERS` later drains +whatever accumulated, immediately. Budget for that. + +**`OUTBOUND_DISABLED` parks, it does not fail.** A blocked delivery stays +`pending` and resumes when the switch clears; it is never poisoned and never +cancelled (`internal/outbound/worker.go:235-242`, +`internal/outbound/switches.go:5-12`). Engaging a kill switch loses nothing. +`OUTBOUND_DRY_RUN` parks the same way, after translating and logging +(`worker.go:243-248`). + +Scope matching, from `config.go:519-525`: + +| Switch | Matching | +|---|---| +| `OUTBOUND_DISABLED_HOSTS` | inbox host, **lowercased** on load and compared case-insensitively — a kill switch must fail closed on case | +| `OUTBOUND_DISABLED_COMMUNITIES` | community **AP id**, e.g. `https://lemmy.world/c/comicstrips`, **exact and case-sensitive**, compared against the delivery's ordering key (`worker.go:237-240`) | +| `OUTBOUND_DISABLED_ACTORS` | actor **DID**, exact and case-sensitive | + +There is no allowlist form of any of these. See +[Staged rollout](#5-staged-rollout) for what that means for a canary. + +--- + +## 3. The admin surface at incident time + +Full route table with request bodies: [README](README.md#the-admin-api). All +routes are bearer-authenticated (`Authorization: Bearer $ADMIN_TOKEN`) behind +one middleware group (`internal/ingest/follow.go:148-162`, +`internal/accept/admin.go:65-71`). Production publishes `127.0.0.1:8091` on the +box so admin calls skip Caddy entirely. + +Four behaviours worth knowing *before* you need them, because each one reads +like a fault and is not: + +**`POST /admin/communities/reconcile` → 501 means `FOLLOW_LIST_PATH` is +unset.** The follow reconciler is only constructed when the path is configured +(`cmd/tidepool/main.go:468-483`), and the handler nil-checks it +(`internal/ingest/follow.go:575-579`). Production *does* set +`FOLLOW_LIST_PATH=/repo/communities.yaml`, so a 501 there means the env changed, +not that reconciliation broke. + +**`/admin/admissions` and `/admin/admissions/readmit` → 404 means the consumer +is off.** Those routes are registered inside the `if cfg.ConsumerEnabled` block +(`cmd/tidepool/main.go:537-556`) because a force re-admit needs the acceptance +engine, which only exists there. A 404 is "the consumer is not running", never +"the endpoint is broken". Every other `/admin` route answers regardless. + +**`POST /admin/outbound/redrive` refuses an unscoped redrive.** With no +`activity`, no `community`, and no explicit `{"all":true}` it returns 400 +(`internal/ingest/follow.go:210-213`) — an unscoped redrive would re-attempt +every poisoned delivery at once, and a malformed body must not become a silent +fleet-wide replay. `POST /admin/outbound/cancel` likewise demands exactly one +of `actor` or `community` (`follow.go:236-239`). + +**`GET /admin/outbound` answers even with the consumer off.** The deliveries +store is wired unconditionally (`cmd/tidepool/main.go:454`), so an empty +`by_state` map is the truth about the queue, not a symptom of misconfiguration. + +```sh +T="Authorization: Bearer $ADMIN_TOKEN" +curl -s -H "$T" localhost:8091/admin/outbound # queue depth by state +curl -s -H "$T" localhost:8091/admin/divergence # one sweep, synchronously +curl -s -H "$T" localhost:8091/admin/metrics # tidepool* expvars +curl -s -H "$T" 'localhost:8091/admin/admissions?status=rejected' +``` + +--- + +## 4. coves.social Caddy — a CROSS-REPO change + +**There is no Caddyfile in this repository.** TLS for both `tdpl.io` and +`coves.social` terminates in the **Coves** Caddy (`coves-prod-caddy`), which +reaches this stack over the shared external `coves-prod-network`. The work +below is an edit to `~/Code/coves/Caddyfile` — a Coves-repo change, deployed on +the Coves side. + +### Why Caddy has to do this at all + +`AP_USER_ORIGIN=https://coves.social` puts the native users' ActivityPub surface +on the **Coves** hostname, served by the Tidepool process. Tidepool's Host +router sends requests whose `Host` is `coves.social` to the persona surface, and +everything under `tdpl.io` to the bridge (`internal/personas/hostrouter.go:89-115`). +Caddy currently proxies `coves.social` entirely to the AppView, so **none of +those AP paths reach Tidepool today.** + +The apex needs a content-negotiated split, and Tidepool deliberately cannot do +it itself: `internal/personas/instance.go:32-34` states in its own comment that +content negotiation is absent *because Caddy owns it in production*, and that a +peer sending `ld+json`, `activity+json`, or no `Accept` at all must still get +the actor. Peers fetch `GET /` to find the instance actor; browsers fetch +`GET /` to find the web app. Only the edge can tell them apart. + +### ⚠️ The Caddyfile inode trap — read this first + +**Edit the Caddyfile in place. Never replace the file.** The production Caddy +mounts the Caddyfile as a single-file bind mount, and Docker pins a single-file +mount to the **inode present when the container started**. Anything that +replaces the file rather than writing through it — `git checkout`, `mv`, +`sed -i`, most editors' atomic-save, `scp` of a new file — leaves the container +serving the **old** contents forever, while `cat` on the host shows the new +ones. This has bitten this project before. It is the same trap that made +Tidepool's own `communities.yaml` a *directory* mount +(`docker-compose.prod.yml:18-24`). + +Safe edits: `nano`/`vim` with backupcopy=yes, `cat > file`, `tee`. After any +edit, confirm the container sees it: + +```sh +docker exec coves-prod-caddy cat /etc/caddy/Caddyfile | grep -c '/ap/' +docker exec coves-prod-caddy caddy validate --config /etc/caddy/Caddyfile \ + --adapter caddyfile +docker exec coves-prod-caddy caddy reload --config /etc/caddy/Caddyfile \ + --adapter caddyfile +``` + +If the container's copy does not show your change, the inode moved: recreate +the Caddy container (`docker compose -f docker-compose.prod.yml up -d +--force-recreate caddy` in `/opt/coves`). + +### The change + +Inside the **existing `coves.social { … }` site block** — do not add a second +block for the same hostname — add the AP routes. `handle` blocks in one site +are mutually exclusive and sorted by path specificity, so +`/.well-known/webfinger` wins over the existing `/.well-known/*` static block, +and `/ap/*` is disjoint from everything already there. + +```caddyfile +coves.social { + # ── Tidepool's native-user AP surface (AP_USER_ORIGIN) ────────────── + # These paths belong to the bridge, not the AppView. More specific than + # the /.well-known/* static block below, so they win the handle sort. + # + # NOTE: no `header_up Host` on these proxies. Caddy v2 forwards the + # original Host by default, and Tidepool's Host router keys on it to + # choose the persona surface over the bridge surface + # (internal/personas/hostrouter.go). Rewriting Host to the upstream + # address would 421 every one of these requests. + handle /.well-known/webfinger { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + handle /.well-known/nodeinfo { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + handle /nodeinfo/2.0 { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + # /ap/actor/{did}, /ap/actor/{did}/outbox, /ap/object/*, /ap/activity/*, + # and POST /ap/inbox — the shared inbox for this origin. + handle /ap/* { + reverse_proxy tidepool:80 { + header_up X-Real-IP {remote_host} + } + } + + # ── Apex: split by Accept ─────────────────────────────────────────── + # `handle /` matches the apex EXACTLY (a Caddy path matcher is exact + # unless it ends in *), so this replaces only the bare "/" case that + # the catch-all used to serve. Same-name directives run in Caddyfile + # order, so the matched reverse_proxy is tried before the fallback. + # + # Two Accept lines, not one substring: Lemmy and Mastodon send + # application/activity+json, but the AS2 spec form is + # application/ld+json;profile="…activitystreams", which shares no + # useful substring with the first. Values for the SAME header field + # are OR'ed. + handle / { + @ap { + header Accept *application/activity+json* + header Accept *application/ld+json* + } + reverse_proxy @ap tidepool:80 { + header_up X-Real-IP {remote_host} + } + # Fallback: the web app, with the SAME upstream options as the + # catch-all below — copy them verbatim, header_up Host included + # (DPoP htu matching depends on it). + reverse_proxy appview:8080 { + health_uri /xrpc/_health + health_interval 30s + health_timeout 5s + header_up Host {host} + header_up X-Real-IP {remote_host} + header_up X-Forwarded-For {remote_host} + header_up X-Forwarded-Proto {scheme} + header_up X-Forwarded-Host {host} + } + } + + # … existing handle /.well-known/*, /client-metadata.json, /img/*, + # catch-all handle, headers, CSP, encode — all unchanged … +} +``` + +### Verify from outside + +```sh +# instance actor (Application), not the web app +curl -s -H 'Accept: application/activity+json' https://coves.social/ | jq '.type,.id' +# expect: "Application" "https://coves.social/" + +# the AS2 spelling must reach the same document +curl -s -H 'Accept: application/ld+json; profile="https://www.w3.org/ns/activitystreams"' \ + https://coves.social/ | jq '.type' + +# a browser must still get the web app +curl -sI -H 'Accept: text/html' https://coves.social/ | head -1 + +curl -s 'https://coves.social/.well-known/webfinger?resource=acct:alice@coves.social' | jq . +curl -s https://coves.social/.well-known/nodeinfo | jq . +curl -s https://coves.social/nodeinfo/2.0 | jq '.software.name' # "tidepool" +``` + +A 421 from any of these means the `Host` header was rewritten on the way +through. A 404 on webfinger for a user who has never federated is correct — +actors are minted lazily on first federating interaction. + +### Not changing + +`tdpl.io`, the per-instance wildcard blocks, and the on-demand catch-all are +untouched by v2. The `on_demand_tls ask` gate still points at +`http://tidepool:80/.well-known/tidepool-tls-ask`, which is served on the +bridge Host (`cmd/tidepool/main.go:165`) and is unaffected by anything above. + +--- + +## 5. Staged rollout + +Decision 19's canary. Each step is a separate `.env` edit plus +`docker compose -f docker-compose.prod.yml up -d tidepool`, because config is +read once at boot. + +### Step 0 — Caddy first + +Do section 4 **before** enabling the consumer. A native actor that federates +outward triggers Lemmy to fetch its webfinger, its actor document, and the +origin apex, all at `coves.social`. If Caddy is not routing those to Tidepool +yet, Lemmy gets the web app or a 404 and caches the failure. + +### Step 1 — consumer on, delivery off + +``` +CONSUMER_ENABLED=true +OUTBOUND_WORKERS=0 +``` + +Nothing reaches any peer. Watch for a day: + +```sh +curl -s -H "$T" localhost:8091/admin/metrics | jq '{ + cursor_age: .tidepool_consumer_cursor_age_seconds, + last_event_age: .tidepool_consumer_last_event_age_seconds, + connected: .tidepool_consumer_connected, + dead_letters: .tidepool_consumer_dead_letters, + reconnects: .tidepool_consumer_reconnects +}' +curl -s -H "$T" localhost:8091/admin/outbound # intent accumulating +curl -s -H "$T" 'localhost:8091/admin/admissions?status=rejected' | jq +``` + +`tidepool_consumer_cursor_age_seconds` and +`tidepool_consumer_last_event_age_seconds` are what make a *stalled* consumer +visible: the process stays up and the healthcheck stays green while events +quietly stop arriving (`cmd/tidepool/main.go:827-830`). A climbing +`tidepool_consumer_dead_letters` means events are failing and being parked for +the redriver, not lost. + +Two reading rules for those keys: + +- **The whole `tidepool_consumer_*` family only exists while the consumer + runs** — `PublishMetrics` is called inside `startConsumer` + (`cmd/tidepool/main.go:830`). Absent keys mean the consumer is off, not that + it is broken and silent. The `tidepool_divergence_*` gauges are the opposite: + registered at package init, so they are always present. +- **`tidepool_consumer_dead_letters: -1` is not a count.** It is the sentinel + for "storage could not be read" (`internal/consume/metrics.go:26-30`), chosen + because a `0` would claim the backlog is empty at exactly the moment nobody + can tell. + +Check the rejections before letting anything out. A misconfigured +`ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY`, a stale community mapping, or a +consent state you did not expect all surface here as `decisionCode`s while the +blast radius is still zero. + +### Step 2 — one community + +There is **no allowlist**, so a one-community canary is spelled as a denylist +over the rest of `communities.yaml`. With today's four entries, canarying +`!comicstrips@lemmy.world` means: + +``` +OUTBOUND_WORKERS=1 +OUTBOUND_DISABLED_COMMUNITIES=https://lemmy.world/c/selfhosted,https://lemmy.world/c/fediverse,https://lemmy.ml/c/linux +``` + +AP ids, exact case — not the `!name@host` spelling `communities.yaml` uses. +Confirm the ids you are about to paste rather than constructing them; the +`community` field of the list response **is** the AP group id +(`internal/ingest/follow.go:283-300`): + +```sh +curl -s -H "$T" localhost:8091/admin/communities | jq -r '.communities[].community' +``` + +**Adding a community to `communities.yaml` later does not add it to this +denylist.** The reconciler will subscribe it and it will start federating on +the next sweep. Whenever the follow list grows during a canary, extend +`OUTBOUND_DISABLED_COMMUNITIES` in the same change. + +Optionally precede this with `OUTBOUND_DRY_RUN=true` for one cycle: every +delivery is translated and logged, nothing is POSTed, and the parked deliveries +resume when you clear it. That validates the translator against real records +without touching a peer. + +### On announcement throttling — what actually exists + +Decision 19 asks for "deliberate throttling of initial actor announcements". +**Read this carefully, because the obvious knob is the wrong one.** + +`MINT_RATE_PER_MINUTE` / `MINT_BURST` (production overrides them to **10/20**, +`docker-compose.prod.yml`, because the public `plc.directory` 429s mint bursts +during community backfill) gate **inbound** DID minting only. They are wired +into `ingest.NewMintGate`, whose sole consumer is the materializer's minter +(`cmd/tidepool/main.go:271-276`, `:314`) — the path where an unseen *Lemmy* +author gets an atproto DID. + +**There is no rate limiter on the outbound side.** `internal/outbound` contains +no `rate.Limiter`, no sleep, and no throttle of any kind; the persona actors +the consumer mints do not pass through `mintGate` (`main.go:539` hands +`personasService` directly to `startConsumer`). The only levers on outbound +volume are: + +- `OUTBOUND_WORKERS` — worker **concurrency**, not a rate. Workers poll with a + 1s idle interval (`main.go:53`); with a full queue they run flat out. +- the scoped kill switches — which communities/actors/hosts may deliver at all. + +So the canary *is* the throttle: `OUTBOUND_WORKERS=1` plus a one-community +scope. Do not go looking for an announcement rate knob; nobody built one, and +this is filed under [Not implemented](#not-implemented). + +### Step 3 — widen + +Remove entries from `OUTBOUND_DISABLED_COMMUNITIES` one at a time, raising +`OUTBOUND_WORKERS` as the queue justifies. Before each widening: + +```sh +# poison depth — the queue's own verdict +curl -s -H "$T" localhost:8091/admin/outbound | jq .by_state + +# echo suppression: bridge-origin content must never re-materialize +curl -s -H "$T" localhost:8091/admin/metrics | jq '{ + mapped_object: .tidepool_echo_drops_mapped_object, + local_activity: .tidepool_echo_drops_local_activity, + local_actor: .tidepool_echo_drops_local_actor, + ancestor: .tidepool_echo_drops_ancestor_short_circuit +}' + +# the reconciliation report +curl -s -H "$T" localhost:8091/admin/divergence | jq '{counts, truncated}' +``` + +What each one means: + +- **`by_state.poisoned` climbing** — deliveries exhausting their retries. + Inspect, fix, then `redrive` **scoped** to the affected community. +- **echo drop counters rising steadily** — expected and healthy: our own + content arriving back from Lemmy and being correctly refused. A counter at + **zero** while native content is flowing is the alarming case; it means + classification is not firing and re-materialization is possible. +- **`tidepool_divergence_*` gauges** — the read-only reconciliation sweep + (`GET /admin/divergence` runs one synchronously). Watch + `acceptance_undelivered_stale` (a delivery pending past + `DIVERGENCE_ACCEPTANCE_STALE_AFTER`, default 12h), + `acceptance_undelivered_poisoned`, and `vote_recast_undelivered`. Also watch + `tidepool_divergence_sweep_failures` and + `tidepool_divergence_sweep_age_seconds`: a failed sweep publishes **nothing** + and the gauges keep their previous values + (`internal/ingest/divergence.go:584-591`), so a stale age with quiet gauges + is a *silent* failure mode. The sweep never repairs anything — it reports. + +### Rollback + +Rollback is a **kill switch, not an un-deploy**. In escalation order, each +step being one `.env` edit plus `up -d tidepool`: + +1. **`OUTBOUND_DISABLED_COMMUNITIES=`** — park one community. + Everything else keeps flowing; the parked deliveries resume when you clear + it. +2. **`OUTBOUND_DISABLED=true`** — park everything outbound. The consumer keeps + running and keeps recording intent; nothing reaches any peer. This is the + big red button and it is **lossless**. +3. **`OUTBOUND_WORKERS=0`** — stop the workers entirely. Equivalent effect to + (2) for delivery; prefer (2), because a parked delivery carries a recorded + reason and a stopped worker does not. +4. **`CONSUMER_ENABLED=false`** — stop consuming. Intent stops being recorded. + The consumer resumes from its stored cursor when re-enabled, so this is + recoverable, but it is the only step that stops *observing*, and + `/admin/admissions` disappears with it — taking your triage view down at the + moment you most want it. Prefer 1–3. + +Do **not** roll back by cancelling deliveries. `POST /admin/outbound/cancel` +is for consent withdrawal and community removal; a cancelled delivery is not +resumable the way a parked one is. + +Only redeploy the previous image if the failure is a code defect rather than a +federation outcome. The kill switches address the latter faster and without a +schema-version question. + +--- + +## 6. Not implemented + +Everything in this section is a real operational gap. None of it has a +mechanism in the code today. It is written down because an operator who +assumes one of these exists will look for it during the exact incident where +looking costs the most. + +### Key rotation — `BRIDGE_KEK` and per-actor RSA keys + +**No rotation path exists. Not partial, not manual, not scripted.** + +Every mention of `BRIDGE_KEK` in this repository is a warning, never a +procedure. `internal/config/config.go:49-53` documents the value; the sealing +itself is AES-256-GCM in `internal/identity/keys.go`. There is nothing that +re-seals existing ciphertext under a new key: the binary has exactly two +subcommands, `tidepool` and `tidepool migrate` +(`cmd/tidepool/main.go:69-78`). + +Blast radius of losing or changing it: every per-actor RSA signing key for +every bridged identity is sealed under it, plus the PLC **escrow rotation +key** (`internal/identity/keys.go:140-144`). Change the KEK and every one of +those ciphertexts becomes undecryptable — no bridged actor can sign, and the +escrow key that could recover the DIDs is itself sealed under the key you just +replaced. Approximately 950 identities. There is no recovery. + +*Naming trap:* `LoadOrCreateRotationKey` is **not** KEK rotation. It loads or +generates the did:plc escrow/recovery key — an atproto identity concept — +which is itself sealed under the KEK. Do not read that symbol as evidence that +rotation is implemented. + +What a real rotation would require, none of which exists: a key-version +column or KEK-id alongside each sealed blob; a dual-read custodian that tries +the new KEK then the old; an online re-seal pass over `bridged_actors` and +`service_keys`; and a cutover that retires the old KEK only after the pass +completes. The ciphertext does carry a one-byte version prefix +(`internal/identity/keys.go:127`), which is a hook someone could build on, but +nothing reads it as a key selector today. + +Per-actor **RSA** rotation is equally undefined: rotating an actor's key means +republishing `publicKey` in its actor document and having every peer that +cached it re-fetch, with no grace-overlap mechanism in the code to publish two +keys at once. + +**Until this is built, treat `BRIDGE_KEK` as immutable, and back it up +somewhere that survives the loss of the server.** + +### Backup and restore + +**No procedure exists, and no tooling.** `docker-compose.prod.yml:46` mounts +`./backups:/backups` into the Postgres container. Nothing writes to it. There +is no cron, no `pg_dump` wrapper, no restore drill, and no documented RPO/RTO. + +What is at risk, in order of irreplaceability: + +1. **`BRIDGE_KEK`** — lives in `/opt/tidepool/.env`, not in Postgres, and is + not covered by any database backup. Losing it is unrecoverable (above). +2. **`bridged_actors` / `service_keys`** — the sealed signing keys. Losing + these loses the identities even if the KEK survives. +3. **Repo blocks and commits** — the atproto repos themselves. Re-derivable + from upstream only by re-bridging, which mints new DIDs; the old at-uris do + not come back. +4. **`outbound_*`, `admissions`** — in-flight federation state. Losing it + double-sends or drops deliveries. + +Anything actually built here should be a separate task with a **restore +drill**, since an unverified backup is a claim, not a capability. + +### A divergence off switch + +**There is none.** The reconciliation sweep is wired unconditionally — the +constructor and `go divergence.Run(ctx)` sit outside every flag at +`cmd/tidepool/main.go:485-506` — and `Run` performs one sweep **immediately**, +before its first tick (`internal/ingest/divergence.go:561-565`). Setting +`DIVERGENCE_INTERVAL=0` does not disable it: `durationVar` rejects zero and +negative values and the process refuses to start +(`internal/config/config.go:666-668`). + +The only available lever is a **large interval** — `DIVERGENCE_INTERVAL=8760h` +quiets the background pass. The startup sweep still runs, once, on every boot, +and `GET /admin/divergence` still works. + +This is defensible: the sweep writes nothing, to peers or to our own tables, +which is exactly what makes an always-on schedule safe. But if a sweep is ever +implicated in an incident (lock pressure, a long multi-table scan — two of its +legs scan `outbound_activities` in full, per `FOLLOWUPS.md`), **the runbook +answer is "raise the interval and restart", not "disable it", because disabling +is not possible.** + +### Periodic vote re-seed + +**Nothing re-seeds vote aggregates on a schedule.** `SeedPostCounts` has one +caller, on the community-backfill path. + +This makes one documented self-healing claim conditional in a way that matters: +baseline-only voters can drift when a later flip or clear has no per-voter +baseline row to retract, and the standing note is that "a re-seed heals the +aggregate" (`FOLLOWUPS.md`, *Votes*). For a **quiet community that is never +backfilled again, the next re-seed is never.** The drift is permanent, silently, +and nothing reports it — the divergence sweep compares atproto state against +outbound state, not vote aggregates against the origin instance. + +The manual lever is a backfill of the affected community +(`POST /admin/communities/backfill`), which re-seeds as a side effect. That is +a workaround, not a scheduled heal, and it does other work besides. + +--- + +## 7. Support matrix + +| Peer | Status | Notes | +|---|---|---| +| **Lemmy 0.19.x** | **Targeted.** Strictness ceiling and e2e target | See the version discrepancy below | +| **PieFed** | Best-effort | No pinned instance in the harness; conformance is held by captured-wire fixtures. Its votes arrive from anonymous per-user actors, which is fine for tallies but means no per-voter identity | +| **Lemmy 1.0-beta** | Tracked, **not targeted** | Vote `FederationMode`, inbox collapsing, and `NoteWrapper` all change behaviour we depend on. No harness coverage | +| **Mastodon** | Incidental | The `security/v1` context is published so its parser accepts our `publicKey`, and its hosts appear in production handle subdomains. Not a target; not tested | + +### ⚠️ Version discrepancy — unresolved, needs a decision + +The tree contradicts itself about which Lemmy version is pinned: + +- `e2e/lemmy/Dockerfile:35` pins **`ARG LEMMY_VERSION=0.19.19`**. This is what + `make e2e` actually builds and tests against. +- `PLAN.md:430` (decision 19) says "**Lemmy 0.19.20** is the pinned strictness + ceiling and e2e target". `tasks/18-e2e-deploy.md:21` repeats it, and the + 0.19.20 source is cited as the authority for specific verified behaviours + across `tasks/13`, `14`, `15`, `17` — the `Delete`-summary convention, the + `check_bot_account` rule, `Instance`-enum strictness. + +**The matrix above says "0.19.x" deliberately, because writing either number +alone would be a claim the tree does not support.** The behaviours we verified +were read from 0.19.20 source; the behaviours we *test* are 0.19.19's. + +This is not resolvable from the docs — it needs a call: + +1. bump `e2e/lemmy/Dockerfile` to `0.19.20` so the tested version matches the + decided one (preferred; the e2e stack is owned by a separate task and this + file is out of scope for this change), **or** +2. amend decision 19 to name 0.19.19 as the pin and re-verify the source + claims against that tag. + +Until one of those happens, do not cite a specific patch version as "the +supported one" in operator-facing material. + +--- + +## 8. Known operational gaps carried forward + +Not new, but they shape what the runbook above can promise: + +- **Per-IP rate limits degrade to global ones behind Caddy.** Every in-process + limiter keys on `RemoteAddr` and ignores `X-Forwarded-For`, so from behind + the proxy they see only Caddy's container IP. Rate limiting at the edge is + the fix; tracked in `FOLLOWUPS.md`. +- **`ENVIRONMENT=production` has never been exercised end to end.** The e2e + harness runs in development mode (migrations-on-start, HTTP, private fetch, + strict lexicon validation). The production-only refusals — `BRIDGE_SCHEME=http`, + `ALLOW_PRIVATE_FETCH`, `AP_HOST_FALLTHROUGH_DEV`, and the `AP_USER_ORIGIN` + https rule — are unit-covered, not harness-covered. Section 1 exists because + of this. +- **Production lexicon validation records and writes rather than failing.** + A strict-first rollout should wait until + `tidepool_lexicon_validation_failures` stays at zero in production. diff --git a/FOLLOWUPS.md b/FOLLOWUPS.md index 3127a00..0b361d6 100644 --- a/FOLLOWUPS.md +++ b/FOLLOWUPS.md @@ -279,17 +279,59 @@ task documents and git history rather than this list. ## Production rollout -- Bluesky's public relay accepts new PDS hosts, but its documented default - allowance is only 100 accounts, 50 repo-stream events/second, 2,600/hour, - and 21,000/day. Tidepool mints one repo DID per bridged actor/community and - will exceed the account cap quickly. Arrange a relay limit increase before - broad subscriptions, or operate a suitably bootstrapped relay. +RESOLVED — Bluesky's public-relay account cap (100 accounts, 50 ev/s) is no +longer load-bearing: the self-hosted relay + Jetstream in +`docker-compose.prod.yml` are the app's ingest path and carry only our two PDS +hosts with an effectively unlimited account limit. `bsky.network` remains the +wider-visibility path only. Runbook: `SELF_HOSTED_RELAY.md`. + +DOCUMENTED, not resolved — the v2 deploy gaps below now have a written home in +`DEPLOY.md` (§6 "Not implemented") with their blast radius. Writing them down +is not building them; they stay open here: + +- **No `BRIDGE_KEK` / per-actor RSA rotation path.** Nothing re-seals existing + ciphertext under a new KEK, and the binary's only subcommand is `migrate` + (no args = serve). Changing the KEK orphans every bridged identity's signing key + *and* the PLC escrow rotation key sealed under it (~950 identities, no + recovery). Would need a key-version selector on each sealed blob, a + dual-read custodian, an online re-seal pass, and a cutover. +- **No backup or restore procedure.** `docker-compose.prod.yml` mounts + `./backups` into the Postgres container and nothing writes to it. Note + `BRIDGE_KEK` lives in `.env` and is not covered by any database backup at + all. Whatever is built needs a restore *drill* — an unverified backup is a + claim, not a capability. +- **No divergence off switch.** The sweep is wired unconditionally, sweeps once + at startup, and `DIVERGENCE_INTERVAL=0` is refused by `durationVar`. The only + lever is a large interval. Defensible while the sweep stays read-only; revisit + if it is ever implicated in lock pressure. +- **No outbound announcement rate limiter.** Decision 19 asks for "deliberate + throttling of initial actor announcements"; `internal/outbound` has no rate + limiter, and `MINT_RATE_PER_MINUTE`/`MINT_BURST` gate the **inbound** mint + path only (`ingest.NewMintGate`, materializer-only consumer). Today the + throttle is the canary itself: `OUTBOUND_WORKERS=1` plus a one-community + scope. +- **Kill switches and every other knob are boot-time only.** Config is read once + at `cmd/tidepool/main.go:109` with no reload signal, so engaging a kill switch + during an incident requires a container recreate. A SIGHUP reload (or an + admin-write switch table) would cut that latency. +- **Scoped kill switches are denylist-only.** There is no allowlist form, so a + one-community canary must enumerate every *other* subscribed community — and + adding a community to `communities.yaml` silently escapes an existing canary. + +Still open, unchanged: + - `ENVIRONMENT=production` has not been exercised end-to-end. The harness uses development mode for migrations-on-start, HTTP/private fetching, and strict lexicon validation. - Production lexicon validation currently records a metric and writes the record instead of failing it. A strict-first rollout should happen only after `tidepool_lexicon_validation_failures` remains zero in production. +- **Lemmy pin contradiction.** `e2e/lemmy/Dockerfile:35` pins + `LEMMY_VERSION=0.19.19`; PLAN decision 19 (`PLAN.md:430`) and + `tasks/18-e2e-deploy.md:21` name **0.19.20** as the pinned strictness ceiling + and e2e target, and the 0.19.20 source is cited as the authority for verified + behaviours across tasks 13/14/15/17. Either bump the Dockerfile or amend the + decision; until then, operator-facing docs say `0.19.x`. ## Sync surface @@ -319,8 +361,11 @@ task documents and git history rather than this list. deadlock-avoidance property has no true concurrency test. - Vote subject resolution occurs outside the mutation transaction, leaving a narrow race with deletion. -- Baseline-only voters can temporarily drift: a later flip or clear lacks a - per-voter baseline row to retract. A re-seed heals the aggregate. +- Baseline-only voters can drift: a later flip or clear lacks a per-voter + baseline row to retract. A re-seed heals the aggregate — but nothing + re-seeds on a schedule (see "Nothing re-seeds periodically" above), so for a + quiet community that is never backfilled again the drift is permanent and + unreported. "Temporarily" was the wrong word. ## Materializer and storage diff --git a/README.md b/README.md index 2f024fb..554a135 100644 --- a/README.md +++ b/README.md @@ -242,12 +242,24 @@ relays, or public Lemmy instances. ## Configuration -Environment variables with logged dev defaults (see -`internal/config/config.go`); everything below is **required in -production**: +All configuration is environment variables, read **once at process start** +(`config.Load`) — there is no reload signal, so changing any value below means +recreating the container. + +Two classes, and the difference matters at boot: + +- **Required in production** — `DATABASE_URL`, `LISTEN_ADDR`, + `BRIDGE_HOSTNAME`, `PLC_DIRECTORY_URL`, `BRIDGE_KEK`, `ADMIN_TOKEN`, + `AP_USER_ORIGIN`. These have *dev defaults only*; unset with + `ENVIRONMENT=production` the process refuses to start (`config: is + required in production`). Fail-closed on purpose: every one of them is baked + into identities or authority, where a defaulted guess is worse than no boot. +- **Tuning knobs** — everything else. Real defaults in every environment, + logged when applied. | Variable | Dev default | Meaning | |---|---|---| +| `ENVIRONMENT` | `development` | `development` or `production`; any other value is refused at boot. Development enables migrations-on-start, dev defaults, and strict lexicon validation. Production additionally *refuses* `BRIDGE_SCHEME=http`, `ALLOW_PRIVATE_FETCH`, `ALLOW_DEV_REQUEST_CRAWL`, and `AP_HOST_FALLTHROUGH_DEV` | | `DATABASE_URL` | local dev postgres | bridge state | | `LISTEN_ADDR` | `:8091` | HTTP bind address | | `BRIDGE_HOSTNAME` | `localhost` | public domain of the bridge; anchors handles and the PDS endpoint in minted DID docs | @@ -258,6 +270,8 @@ production**: | `USER_AGENT` | derived | outbound HTTP user agent | | `ALLOW_PRIVATE_FETCH` | off | dev-only: disables the SSRF egress guard (AP fetches **and** PLC directory requests) so localhost targets work | | `FIREHOSE_RETENTION` | `72h` | how long `firehose_events` rows are kept for `subscribeRepos` cursor replay (Go duration; a background pruner trims older events hourly) | +| `MAX_BLOB_BYTES` | `5242880` (5 MiB) | outer transport budget for remote media (avatars, banners, post images) the materializer downloads per blob. Individual lexicon slots impose tighter caps (avatars 1 MB); this is the ceiling over all of them. Fails closed — oversized media is dropped, never truncated | +| `PROFILE_REFRESH_TTL` | `24h` | how stale a bridged actor's materialized profile may get before the materializer re-fetches it. `Update{Person\|Group}` refreshes immediately regardless — this covers what Lemmy never federates (bio edits, and so the `#nobridge` marker) | | `RELAY_HOSTS` | *(optional)* | comma-separated relays to send `com.atproto.sync.requestCrawl` to on startup (each retried on a bounded budget — the relay calls back into `describeServer` before subscribing, which can race process start); in development the request is logged, never sent, unless `ALLOW_DEV_REQUEST_CRAWL` opts in | | `ALLOW_DEV_REQUEST_CRAWL` | off | dev-only: actually SEND `requestCrawl` to `RELAY_HOSTS` in development (exists for the e2e stack's local BigSky); refused in production, where sending is already the behavior | | `ADMIN_TOKEN` | `dev-admin-token` | bearer token protecting the `/admin` API | @@ -280,13 +294,51 @@ production**: | `SYNC_MAX_SUBSCRIBERS` | `100` | concurrent `subscribeRepos` connection cap | | `FOLLOW_LIST_PATH` | *(optional)* | declarative follow list (see below); unset = the `/admin` API is the only subscription control | | `FOLLOW_LIST_INTERVAL` | `15m` | follow-list reconciler sweep cadence | -| `CONSUMER_ENABLED` | **off** | turns on the task-14 Jetstream consumer (native users' opt-outs, profiles, posts, comments, votes → durable outbound state). Default off: it writes durable state and hands work to delivery seams that are still stubbed until tasks 15–17 land, so a deployment that has not been wired end to end should not silently start accumulating it. Enabling it now runs the pipeline with a logging-noop enqueuer (nothing is delivered), a nil acceptance engine (postv2 skipped), and nil destructive/terminal tiers (opt-out `deleteRemote` and account deletions recorded, not acted on) | +| `DIVERGENCE_INTERVAL` | `15m` | cadence of the reconciliation sweep (task 17e) that compares atproto state against outbound state and publishes the `tidepool_divergence_*` gauges. **Not an on/off switch:** the sweep is wired unconditionally, runs once at startup before its first tick, and `0` is refused — it is read-only (it reports, never repairs), which is what makes an always-on schedule safe. `GET /admin/divergence` runs one on demand | +| `DIVERGENCE_ACCEPTANCE_STALE_AFTER` | `12h` | how long a pending delivery may sit before the sweep reports its acceptance as **stale**. The report's one crying-wolf knob — shorter and every in-flight post is a finding, longer and a queue that stopped this morning is not in tonight's report. 12h is derived from the retry schedule (~2–3h to poison) plus the causal wait budget (6h), not picked | +| `CONSUMER_ENABLED` | **off** | turns on the Jetstream consumer (task 14): native users' opt-outs, profiles, posts, comments and votes flowing outward. Default off because it writes durable outbound state, and because a deployment that has not been canaried should not start accumulating it — not because the seams behind it are stubbed. They are wired: with it on, the **real** enqueuer persists outbound intent, the acceptance engine admits postv2 and writes community-signed acceptances, and opt-out `deleteRemote` / confirmed account deletions actually purge at peers. It also gates two other things — the `OUTBOUND_WORKERS` AND, and whether `/admin/admissions*` exists at all | | `JETSTREAM_URL` | *(optional)* | the self-hosted Jetstream the consumer subscribes to (`ws://` or `wss://`); **required** when `CONSUMER_ENABLED`, and validated at boot whenever set so a typo fails fast instead of becoming a reconnect loop. May be staged ahead of the flag | - -## Subscribing to communities (admin API) - -Community subscriptions are operator-driven, over bearer-token-protected -endpoints (`Authorization: Bearer $ADMIN_TOKEN`): +| `OUTBOUND_WORKERS` | `0` (**off**) | how many delivery workers run. **There is no `OUTBOUND_ENABLED`:** delivery starts only when this is `>0` *and* `CONSUMER_ENABLED`. With the consumer on and this at `0`, outbound intent still accumulates durably and nothing is POSTed — which is the intended staging step, not a broken state. Raising it drains the accumulated backlog immediately | +| `OUTBOUND_DISABLED` | off | global delivery kill switch. A blocked delivery is **parked** — it stays `pending` and resumes when the switch clears — never poisoned, never cancelled. Engaging it loses nothing; it stops the wire | +| `OUTBOUND_DISABLED_HOSTS` | *(empty)* | comma-separated inbox **hosts** to park. Lowercased on load and compared case-insensitively — a kill switch must fail closed on case | +| `OUTBOUND_DISABLED_COMMUNITIES` | *(empty)* | comma-separated community **AP ids** to park (`https://lemmy.world/c/comicstrips`), matched **exactly and case-sensitively** against the delivery's ordering key. There is no allowlist form: a one-community canary is spelled by disabling every other community | +| `OUTBOUND_DISABLED_ACTORS` | *(empty)* | comma-separated actor **DIDs** to park, exact match | +| `OUTBOUND_DRY_RUN` | off | translate and log every delivery, POST nothing. Parks like the kill switches, so nothing is lost — the difference is that the translation ran and is in the log | +| `ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY` | `50` | acceptance-engine flood cap: how many posts one native author may have accepted into one bridged community. `0` = unlimited. Tidepool signs the community's acceptance, so this bounds what it vouches for | + +## The admin API + +Every route below is bearer-protected (`Authorization: Bearer $ADMIN_TOKEN`) +and mounted on the bridge's own `Host`. Production publishes port +`127.0.0.1:8091` on the box for exactly this — admin calls do not round-trip +through Caddy. + +| Route | Purpose | Answers 501/404 when | +|---|---|---| +| `POST /admin/communities` | subscribe (WebFinger → Group → materialize → signed `Follow`) | — | +| `DELETE /admin/communities` | unsubscribe (`Undo{Follow}`; records kept, content stops) | — | +| `GET /admin/communities` | list subscriptions and their state | — | +| `POST /admin/communities/backfill` | on-demand outbox backfill | backfill unconfigured (**501**) | +| `POST /admin/communities/reconcile` | force one follow-list sweep | **501** — `FOLLOW_LIST_PATH` is unset. The reconciler is only built when the path is set, so this is "no follow list configured", not a broken endpoint | +| `POST /admin/reemit` | re-emit a repo's records as delete+create pairs (relay cold-start gap) | repo manager unconfigured (**501**) | +| `POST /admin/objects/sweep-deleted` | origin-verified cleanup of missed deletes | sweeper unconfigured (**501**) | +| `GET /admin/divergence` | run one reconciliation sweep synchronously and return the report | always wired in a normal deployment | +| `GET /admin/outbound` | delivery queue depth by state (`pending`/`poisoned`/`cancelled`/…) | — the store is always wired, so this answers even with the consumer off and workers at 0. An empty queue then is the truth, not a misconfiguration | +| `POST /admin/outbound/redrive` | reset poisoned deliveries to pending | — | +| `POST /admin/outbound/cancel` | park an actor's or a community's pending deliveries as cancelled | — | +| `GET /admin/admissions` | list acceptance decisions with `status`, `decisionCode`, `evaluatedCid`; filter by `?status=`/`?community=` | **404** — these routes are registered **only** when `CONSUMER_ENABLED`. A 404 here means the consumer is off, not that the endpoint is broken | +| `POST /admin/admissions/readmit` | force re-admit a rejected post | **404**, same reason | +| `GET /admin/metrics` | expvar counters filtered to the `tidepool*` prefix (never Go's `cmdline`/`memstats`) | — | + +`POST /admin/outbound/redrive` **refuses an unscoped redrive**: send +`{"activity":"…"}`, `{"community":"…"}`, or an explicit `{"all":true}`. A +missing filter is a 400, never a silent fleet-wide replay of every poisoned +delivery. `POST /admin/outbound/cancel` requires exactly one of +`{"actor":"did:…"}` or `{"community":"https://…"}`. + +### Subscribing to communities + +Community subscriptions are operator-driven: ```sh # follow: WebFinger → fetch Group → materialize community → signed Follow @@ -622,6 +674,35 @@ Note TLS: a single wildcard certificate only covers one label level, while bridged handles sit two levels below `BRIDGE_HOSTNAME` — terminate TLS with on-demand certificate issuance (e.g. Caddy) or per-instance wildcard certs. +## Production deploy + +The runbook is **[`DEPLOY.md`](DEPLOY.md)**: the boot-time config gate, the v2 +flag topology (`CONSUMER_ENABLED` → `OUTBOUND_WORKERS` → kill switches), the +cross-repo Caddy change that puts the native-user AP surface on +`coves.social`, the staged canary and its rollback order, and — explicitly — +the things that have **no** mechanism today (KEK/RSA rotation, backup/restore, +a divergence off switch, a periodic vote re-seed). +[`SELF_HOSTED_RELAY.md`](SELF_HOSTED_RELAY.md) covers the relay + Jetstream +ingest path. + +One thing to know before deploying HEAD onto an existing box: `AP_USER_ORIGIN` +is required in production and was not in the pre-v2 env file, so the process +fails at startup rather than coming up half-configured. That is the intended +behaviour — see DEPLOY.md §1. + +### Support matrix + +| Peer | Status | +|---|---| +| **Lemmy 0.19.x** | targeted — the e2e target and strictness ceiling | +| **PieFed** | best-effort, behind captured-wire conformance (votes arrive from anonymous per-user actors: fine for tallies, no per-voter identity) | +| **Lemmy 1.0-beta** | tracked, not targeted (vote `FederationMode`, inbox collapsing, `NoteWrapper`) | +| **Mastodon** | incidental — the `security/v1` context is published so its parser accepts our `publicKey`; not a target, not tested | + +The patch version is deliberately written as `0.19.x`: `e2e/lemmy/Dockerfile` +pins `0.19.19` while decision 19 and the task docs name `0.19.20`. That +contradiction is unresolved — see DEPLOY.md §7. + ## License Tidepool is licensed under the [GNU Affero General Public License v3.0](LICENSE) diff --git a/SELF_HOSTED_RELAY.md b/SELF_HOSTED_RELAY.md index 9d273fb..e356854 100644 --- a/SELF_HOSTED_RELAY.md +++ b/SELF_HOSTED_RELAY.md @@ -30,7 +30,16 @@ at bsky.network, so native Coves signup #101 would silently vanish from the app exactly like the bridged commenters did. Design decisions and their reasons live as comments on the compose -services; this file is the runbook. +services; this file is the runbook for the **ingest** path only. The v2 +deploy — config gate, the `CONSUMER_ENABLED`/`OUTBOUND_WORKERS` staging +order, the coves.social Caddy change, rollout and rollback — is +[`DEPLOY.md`](DEPLOY.md). + +Note the direction: this file describes the relay and Jetstream that feed +**Coves**. Tidepool's own consumer reads the same Jetstream from the other +side, over `JETSTREAM_URL` +(`ws://tidepool-prod-jetstream:8080/subscribe`, :8080 = `JETSTREAM_ADDR`; +:6060 is the debug listener and serves no `/subscribe`). ## Not in scope here (Phase 2+, Coves repo) diff --git a/docker-compose.prod.yml b/docker-compose.prod.yml index 1249b33..9481667 100644 --- a/docker-compose.prod.yml +++ b/docker-compose.prod.yml @@ -23,6 +23,13 @@ # plain `git pull` is enough; the reconciler picks the new file up on its # next sweep (FOLLOW_LIST_INTERVAL, default 15m) — no container restart. # +# v2 env: the `tidepool` service below gained the task-13/14/15/16/17 knobs +# (AP_USER_ORIGIN, CONSUMER_ENABLED, JETSTREAM_URL, OUTBOUND_*, the admission +# quota, DIVERGENCE_*). AP_USER_ORIGIN is REQUIRED in production, so deploying +# HEAD onto a pre-v2 /opt/tidepool/.env fails at boot rather than coming up +# half-configured. The staged-rollout order, the kill switches and the +# coves.social Caddy changes are in DEPLOY.md. +# # Ops note (README "Sync surface"): the bridge's per-IP rate limiters key on # RemoteAddr and see only Caddy's container IP from behind the proxy — # per-IP limits degrade to global ones until rate limiting exists at the @@ -104,6 +111,84 @@ services: # instead (SELF_HOSTED_RELAY.md). RELAY_HOSTS: ${RELAY_HOSTS:-https://bsky.network} FOLLOW_LIST_PATH: /repo/communities.yaml + + # --- v2 (tasks 13-17): the native-user surface and the outbound half --- + # + # REQUIRED IN PRODUCTION. config.stringVar refuses an unset value outside + # development (internal/config/config.go:797-806), so a HEAD deploy onto + # an env file without this dies at boot with + # config: AP_USER_ORIGIN is required in production + # — fail-closed, and deliberately so: this origin is baked into every + # actor_id this deployment mints, so a wrong value is not a restart away + # from fixed, it is a set of federated identities pointing at the wrong + # place forever. Must be https in production (config.go:782-784) and must + # not be BRIDGE_HOSTNAME or a subdomain of it (config.go:786-790), which + # would shadow the bridged-handle namespace. + # + # Serving derives every URL from the stored actor_id, never from this + # value — so changing it later does NOT migrate existing actors, it only + # changes where NEW ones are minted. + AP_USER_ORIGIN: ${AP_USER_ORIGIN:-https://coves.social} + + # The Jetstream consumer: native Coves users' records flowing OUTWARD. + # Default off. Staged here ahead of the flag on purpose — JETSTREAM_URL is + # validated whenever it is set (config.go:481-489), not only when the + # consumer is on, so a typo is a boot failure with a clear message today + # rather than a silent reconnect loop on the day someone flips the switch. + # + # Container NAME, not the `jetstream` service alias, for the same reason + # DATABASE_URL uses tidepool-prod-postgres: this service also joins + # coves-prod-network, and a bare alias is only unambiguous until the Coves + # stack grows a service by the same name. :8080 is JETSTREAM_ADDR below; + # :6060 is its debug listener and serves no /subscribe. + CONSUMER_ENABLED: ${CONSUMER_ENABLED:-false} + JETSTREAM_URL: ${JETSTREAM_URL:-ws://tidepool-prod-jetstream:8080/subscribe} + + # Delivery workers. 0 = OFF, and it is AND-gated on CONSUMER_ENABLED + # (cmd/tidepool/main.go:716, 843) — there is no OUTBOUND_ENABLED. With the + # consumer on and workers at 0 the real enqueuer still writes + # outbound_activities/deliveries, so intent accumulates durably and starts + # draining the moment workers are raised. That is the intended staging + # order: consumer first, watch the queue, then workers. + OUTBOUND_WORKERS: ${OUTBOUND_WORKERS:-0} + + # Kill switches (decision 19). A blocked delivery is PARKED — it stays + # pending and resumes when the switch clears (internal/outbound/worker.go + # :236-242) — never poisoned, never cancelled, so engaging one loses + # nothing. DRY_RUN parks too, after translating and logging. + # + # ALL OF THESE ARE READ ONCE AT BOOT (config.Load, main.go:109); there is + # no reload signal. Changing one means editing /opt/tidepool/.env and + # recreating the service. + # + # HOSTS is lowercased on load; COMMUNITIES and ACTORS are matched EXACTLY + # and case-sensitively (config.go:519-525). Communities are AP ids + # ("https://lemmy.world/c/comicstrips"), matched against the delivery's + # ordering key; actors are DIDs. There is no allowlist form — a canary is + # expressed by disabling every community except the one under test. + OUTBOUND_DRY_RUN: ${OUTBOUND_DRY_RUN:-false} + OUTBOUND_DISABLED: ${OUTBOUND_DISABLED:-false} + OUTBOUND_DISABLED_HOSTS: ${OUTBOUND_DISABLED_HOSTS:-} + OUTBOUND_DISABLED_COMMUNITIES: ${OUTBOUND_DISABLED_COMMUNITIES:-} + OUTBOUND_DISABLED_ACTORS: ${OUTBOUND_DISABLED_ACTORS:-} + + # The acceptance engine's per-author-per-community flood cap (0 = + # unlimited). Tidepool SIGNS the community's acceptance, so it must not + # let one native account flood a community it vouches for. + ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY: ${ADMISSION_MAX_PER_AUTHOR_PER_COMMUNITY:-50} + + # The reconciliation sweep (task 17e). It is wired UNCONDITIONALLY + # (main.go:491-506) and sweeps once at startup before its first tick, so + # this interval is not an on/off switch — durationVar refuses 0 or + # negative (config.go:666-668). It writes nothing, to peers or to our own + # tables, which is what makes an always-on schedule safe. The only way to + # quiet it is a large interval; see DEPLOY.md "Not implemented". + DIVERGENCE_INTERVAL: ${DIVERGENCE_INTERVAL:-15m} + # How long a pending delivery may sit before the sweep calls its + # acceptance stale. The report's one crying-wolf knob; 12h is derived from + # the retry schedule plus the causal wait budget (internal/ingest/ + # divergence.go:72-81), not picked. + DIVERGENCE_ACCEPTANCE_STALE_AFTER: ${DIVERGENCE_ACCEPTANCE_STALE_AFTER:-12h} volumes: - ./:/repo:ro networks: -- 2.51.2