From 26bee30a55af6f071b20f5baf2547739824182ce Mon Sep 17 00:00:00 2001 From: Bretton <36870434+BrettM86@users.noreply.github.com> Date: Tue, 21 Jul 2026 20:33:30 -0700 Subject: [PATCH] =?UTF-8?q?chore:=20prep=20for=20public=20mirror=20?= =?UTF-8?q?=E2=80=94=20untrack=20build-loop=20internals,=20point=20links?= =?UTF-8?q?=20at=20GitHub?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Untrack LOOP_STATE.md and tasks/ (internal AI build-loop state and specs; files remain local, now gitignored) - Trim PLAN.md to the architecture and locked decisions; drop the loop iteration tables and protocol - Point Coves links (README, CI lexicon-drift clone) at github.com/BrettM86/coves Co-Authored-By: Claude Fable 5 --- .github/workflows/ci.yml | 2 +- .gitignore | 6 + LOOP_STATE.md | 650 ----------------------------------- PLAN.md | 45 --- README.md | 7 +- tasks/01-scaffold-storage.md | 56 --- tasks/02-ap-protocol.md | 53 --- tasks/03-identity-repos.md | 57 --- tasks/04-sync-firehose.md | 50 --- tasks/05-materializer.md | 72 ---- tasks/06-ingestion.md | 63 ---- tasks/07-vote-aggregates.md | 47 --- tasks/08-e2e-harness.md | 57 --- tasks/09-e2e-relay.md | 72 ---- tasks/10-e2e-scenarios.md | 55 --- tasks/11-hardening.md | 70 ---- tasks/12-perf-scale.md | 48 --- 17 files changed, 10 insertions(+), 1400 deletions(-) delete mode 100644 LOOP_STATE.md delete mode 100644 tasks/01-scaffold-storage.md delete mode 100644 tasks/02-ap-protocol.md delete mode 100644 tasks/03-identity-repos.md delete mode 100644 tasks/04-sync-firehose.md delete mode 100644 tasks/05-materializer.md delete mode 100644 tasks/06-ingestion.md delete mode 100644 tasks/07-vote-aggregates.md delete mode 100644 tasks/08-e2e-harness.md delete mode 100644 tasks/09-e2e-relay.md delete mode 100644 tasks/10-e2e-scenarios.md delete mode 100644 tasks/11-hardening.md delete mode 100644 tasks/12-perf-scale.md diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 13a054a..42c438f 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -51,7 +51,7 @@ jobs: # Tidepool vendors Coves' schemas. Comparing only against the local # manifest cannot detect a Coves-only lexicon change, so CI checks the # current canonical Coves tree as well. - run: git clone --depth 1 https://tangled.org/bretton.dev/coves /tmp/coves + run: git clone --depth 1 https://github.com/BrettM86/coves /tmp/coves - name: Lexicon manifest and upstream drift check run: ./scripts/check-lexicons.sh /tmp/coves - name: Unit & integration tests diff --git a/.gitignore b/.gitignore index ba736d6..e8959fb 100644 --- a/.gitignore +++ b/.gitignore @@ -17,7 +17,13 @@ coverage.out # Postgres dump target of docker-compose.prod.yml (server-side only) /backups/ +# Internal build-process docs (AI build-loop state and task specs) — kept +# locally, not published +/LOOP_STATE.md +/tasks/ + # Local Claude Code session/tooling state — but the deploy command is + # repo-committed (like the Coves one; it uses a placeholder so # the SSH target stays out of git history) .claude/* diff --git a/LOOP_STATE.md b/LOOP_STATE.md deleted file mode 100644 index b5e5b63..0000000 --- a/LOOP_STATE.md +++ /dev/null @@ -1,650 +0,0 @@ -# Tidepool build-loop state - -Protocol: PLAN.md §Loop protocol. One task per iteration: -implement (Fable agent) → /second-opinion review → fix → verify → commit → -update this file → schedule next. Stop the loop when every task is `done`. - -| # | Task | Status | Commit | Notes | -|---|------|--------|--------|-------| -| 1 | 01-scaffold-storage | done | (see git log) | reviewed by 7 reviewers, 18 fixes applied | -| 2 | 02-ap-protocol | done | (see git log) | 5 reviewers incl. security; 14 fixes (critical: actor-id binding; high: SSRF, webfinger host confusion) | -| 3 | 03-identity-repos | done | (see git log) | 8 reviewers (5 Claude + codex/gemini/glm); 16 fixes (genesis race, seq ordering, MST-corruption-as-NotFound, KeyUse deletes, TID micro-fill) + 7 new tests | -| 4 | 04-sync-firehose | done | (see git log) | 7/8 reviewers (glm watchdog-killed); fixes: ping starvation, prune-mid-replay OutdatedCursor, broadcaster closed-channel, SeqBounds dirty-read, pruner fail-closed; consent-on-firehose deferred to 06 | -| 5 | 05-materializer | done | (see git log) | 8 reviewers; fixes: id-authority binding, Note-root panic, create-after-delete, nobridge scrub, embedded-actor trust, byte caps, uri scheme, Group-type check + 9 regression tests | -| 6 | 06-ingestion | done | (see git log) | 8 reviewers (5 Claude + codex/gemini; glm wandered, no JSON); fixes: announced-Delete/Undo scoped to announcer authority (+actor-delete only self), bare Update{Person/Group} no-mint gate, announce content community-authority check, Undo{Delete} restore compensation, handleAccept pending-only, queue lease fencing token + shutdown-cancel handling + processed/poisoned exclusivity, backfillReplies tombstone check, truncation leaves resumable, activityID rand-fail propagates + 14 regression tests | -| 7 | 07-vote-aggregates | done | (see git log) | 6/8 reviewers (Gemini perm-denied, glm watchdog-killed); fixes: announced-vote subject↔community binding (post mapping-DID / comment reply.root), bare Undo{Like} signer binding, RetractVote id-targeted undo, dup-id 0/0 aggregate-row leak, seeder zero-clobber presence check, limiter sweep-throttle + 50k fail-closed cap, at-uri validation + ~20 regression tests | -| 8 | 08-e2e-harness (infra: Dockerfile, compose, Lemmy federation, Makefile, CI, lexicon-sync) | done | (see git log) | 7/7 reviewers (4 Claude + codex/gemini/glm, first full external panel since 03); fixes: PRODUCTION https→http redirect-downgrade guard (codex unique catch), webfinger fallback narrowed to transport failures + both-legs errors + 4 tests, minter PDS-endpoint scheme threading, PLC image commit-pin, --wait-timeout + CI logs if:always() + Makefile up-failure cleanup/teardown-status, loopback-only host binds, check-lexicons fail-open holes, sync-lexicons bridge-nesting guard, 2 false compose-header claims rewritten (invented env var, wrong --wait semantics) | -| 9 | 08-e2e-harness (tests: tests/e2e helpers + 8 scenarios, FOLLOWUPS.md, README) | done | (see git log) | 8/8 reviewers (5 Claude + codex/gemini/glm; gemini zero-issue "excellent", codex sharpest); fixes: drain() dead-listener vacuous-pass (5/8 flagged), centralized vetEvent (unknown-collection Fatalf + lexicon-validate every consumed create/update, suite-wide locked-decision-7 enforcement), scenario-7 backfill-completion poll + gap-post cursor-resume proof + op-agnostic dup keys, readLoop goroutine join, subscribe fail-fast on explicit reject, embed.external e2e coverage, seeder e2e assertion. NEW ASSERTION CAUGHT REAL BUG: Lemmy vote-clear federates Undo with RECONSTRUCTED inner vote (fresh id, type Like even for live dislike; flips are bare opposite votes, no Undo) → id-targeted RetractVote no-oped every production vote-clear; fixed with known-id replay probe + live-vote fallback + 3 unit tests | - -v1 loop (tasks 01–08) COMPLETE — `make e2e` green (8/8 scenarios, ~100s), full unit suite green. - -## v1.1 loop (tasks 09–12) — added 2026-07-10 - -Goal: work the FOLLOWUPS.md backlog and put a real relay in the e2e -pipeline. Locked requirement: every state has a full e2e pipeline of -Lemmy → PDS record → firehose ingestion (where applicable) — task 09 -re-points Jetstream through the relay so every scenario transits it. -Same protocol: implement (Fable agent) → /second-opinion → fix → verify → -commit → update this file → next task. Votes-as-records was deliberately -NOT scheduled — see FOLLOWUPS.md "Design revisits" (decide with the -write-back design; tasks 11–12 are its prerequisites). - -| # | Task | Status | Commit | Notes | -|---|------|--------|--------|-------| -| 10 | 09-e2e-relay | done | (see git log) | 7/7 reviewers (4 Claude emulated + codex/gemini/glm); fixes: dev requestCrawl PUBLIC-relay dial guard (codex unique catch — NewPrivateOnlyHTTPClient, inverse SSRF guard), terminal-error classification made pre-flight-only (whole-chain IsValidation was abandoning a relay on attempt 1 for transient DNS), 10s per-attempt timeout (budget arithmetic was 14min worst-case, not 2min), vacuous validation-no-retry test rewritten + 400-is-retried pin, vetEvent per-DID rev-monotonicity (restores per-repo ordering assertion suite-wide), drain() returns+clears pending (closes task-10 vacuous-pass trap), relay poll robustness + pagination cap, doc corrections (RESOLVE_ADDRESS overstatement, spec BGS_CRAWL_INSECURE_WS annotation, FOLLOWUPS 16th-failure off-by-one). KEPT DELIBERATE over 3 reviewers' objection: all wire errors incl. 4xx retried — bigsky answers the describeServer callback race with HTTP 400 (comment + test pin it). Final clean make e2e: 10/10, 96.7s | -| 11 | 10-e2e-scenarios | done | (see git log) | 6/7 reviewers (glm watchdog-killed); UNANIMOUS 6/6 finding: tombstone confirm-fetch transient failure → definitive 401 permanently lost legitimate account deletions → fixed with three-way taxonomy (tombstone→202, alive/validation/404→401, transport/5xx→503 defer) + test; codex unique: confirmation fetch followed cross-authority redirects (open-redirect → forged 410) → FetchActorSameAuthority pins every hop; security: unauthenticated durable-write path flagged → encoded into task 11 rate-limit spec; also: zz-sweep replay floor + honest bounds (sentinel-only pass was vacuous), Delete(Actor) over-scrub drain, actor!=object + Announce{Delete} 401 pins, GET / route-level test, cursor 0→1 doc fixes, vote-hammer header de-overclaimed. TASK ITSELF: 7 scenarios + 2 PRODUCTION fixes (apex instance actor — Lemmy silently never delivers Delete{Person} without a Site actor row; tombstone-verified self-delete acceptance). Final clean make e2e: 17/17, 239s | -| 12 | 11-hardening | done | (see git log) | 6/7 reviewers (glm watchdog-killed on 5k-line diff); NO high-sev confirmed (gemini's "carry-forward type assertion always fails" was a FALSE POSITIVE — GetRecord returns typed atdata.Blob, test green). Fixes: FollowRetrier atomic UPDATE...RETURNING claim (list-then-update raced Accept + burned attempts on transient failure + silent exhaustion), rate-limit refusal observability (expvar counters + sampled Warn — mistuned limit silently dropped all traffic), community-DID blob orphan now retryable (was swallowed → served forever; required delete-before-soft-delete reorder), ScrubVoter DELETE...RETURNING recompute (phantom-count lost update), carry-forward drops on permanent 404/410 vs carries on transient, DeleteActor terminal-state fixpoint (no double #account), migration-011 CHECK tightened + raw-insert test, /admin/metrics scoped expvar, internal/prune fail-closed test, proxy XFF ops note. TASK: inbox+sync admission control, #account{active:false,status:deleted} frame verified purging repo from bigsky, follow auto-retry, 3 pruners, one-tx record+mapping, blob/vote scrubs, service_keys rename. delete-before-create: README was RIGHT, FOLLOWUPS stale (task 06 already closed it). Final clean make e2e: 17/17, 232s | -| 13 | 12-perf-scale | done | (see git log) | 7/8 reviewers (5 Claude emulated + codex/gemini; glm watchdog-killed again); NO high-sev code findings — gemini zero-issue "excellent". Fixes: mid-stream getRepo failure now panics http.ErrAbortHandler so a truncated CAR is a transport-level failure, not a clean 200 (4/7 flagged — the diff's sharpest catch); client-disconnect logging downgraded to Debug; vanished-repo race → 404; FALSE parenthetical in the gc.go invariant header corrected ("every block a commit references is in newBlocks" → only NEWLY-referenced blocks are; believing the original justified deleting rule (a)); DELETE-time created_at re-check + RR isolation + 256-block batch boundary all made test-load-bearing with fail-then-pass proofs (pr-test-analyzer: "you could delete the race-guard clause and the suite stayed green" — no longer); CAR-slice identity test extended to update/delete ops + ordered comparison + no swallowed reader errors; ExportCARTo pre-first-byte doc overclaim fixed; app↔DB clock-skew assumption documented (codex unique). TASK: per-DID MST cache (PutRecord 140.6ms→3.30ms, 42.6x, on a 2k-record repo), streaming reachable-set getRepo (CAR 10.9x smaller, batch-bounded residency), blocks GC (invariant-first design in gc.go), ClaimNext loose index scan (162.6ms→0.050ms on a 50k backlog). Resumed from a killed predecessor agent's partial tree — its walk was correct but 1-SELECT-per-block (698ms/op, would have been a getRepo latency REGRESSION; batching fixed it). Final clean make e2e: 17/17 | - -v1.1 loop (tasks 09–12) COMPLETE — `make e2e` green (17/17), full unit -suite green, all FOLLOWUPS items scheduled into this loop closed or -explicitly deferred with rationale. - -Statuses: pending → in-progress → review → done (or blocked: ). - -## Cross-task notes for future iterations -(implementation agents & reviewers append surprises, interface changes, -and deferred TODOs here) - -- Reference clones live at ~/Code/bridgy-fed, ~/Code/granary, - ~/Code/arroba (CC0). Coves AppView at ~/Code/coves. -- Coves post consumer requires: repo DID == record.community, community - indexed before post, author user indexed before post. - -### From task 01 (storage layer semantics later tasks MUST know) -- Ports: dev postgres 5442, test postgres 5443, HTTP :8091. Test DB URL: - postgres://tidepool_test:tidepool_test@localhost:5443/tidepool_test. - Containers tidepool-dev-postgres / tidepool-test-postgres. -- `PutMapping` derives at_uri itself — callers supply (DID, collection, - rkey, CID) only. Second ap_id claiming same at_uri → IsAlreadyExists - (deterministic-rkey collision = bug signal). ap_objects has an `origin` - column (fediverse|bridge) for task 06 echo suppression. -- `ResolveStrongRef` has THREE outcomes: found; IsNotFound (missing → - task 05 fetches ancestor chain); IsTombstoned (deleted → task 05 drops - the subtree, consent-relevant). Tombstoned does NOT satisfy IsNotFound. -- `UpsertActor`/`UpsertCommunity`: identity drift (same AP id, different - DID/type/instance) → ConflictError, row untouched. Tombstoned actors - (consent_state=deleted, terminal) are fully frozen — upserts no-op and - return the stored row. Handle/signing key are sticky: omitted values - never clobber stored ones; key mutation is deliberate-only. -- ConsentState zero value ("") is INVALID by design (consent must be - stated explicitly — fails closed). Model field is SigningKeyEncrypted - (task 03 stores AES-GCM ciphertext there, column `signing_key`). -- ap_objects timestamp is `PublishedAt` (AP published), NOT created_at. -- `followed_at` stamps only on transition into accepted; AP-driven - arbitrary transitions otherwise legal. inbox_events queue-consumption - API (ListPending/ordering/attempts) deliberately deferred to task 06. -- ENVIRONMENT=production disables migrations-on-start and dev defaults. -- indigo pinned to pseudo-version v0.0.0-20260202181658-ea3d39eec464 - (same as Coves; @latest needs Go 1.26). Unique constraints have - explicit names, mapped via pq.Error.Constraint in uniqueViolation(). -- pr-review-toolkit plugin agents unavailable in this session — the loop - emulates them with general-purpose agents (works fine; keep doing it). - -### From task 02 (internal/ap protocol layer — tasks 05/06 consume this) -- FetchObject/FetchActor error branches mirror ResolveStrongRef: IsNotFound - (404/401/403) → task 05 fetches ancestor chain; IsTombstoned (410 or - Tombstone body) → drop subtree (consent). SignatureError → IsValidation. -- ap.Object is one universal struct. ap.Time has .OK()/.Valid — a present - but malformed `published` is non-nil but OK()==false; task 05 MUST call - OK() before deriving rkeys/TIDs (zero Time would collide/mis-sort). -- FetchCollection signals ErrCollectionTruncated when it hits the page cap - or a next-loop — task 06 backfill must treat that as "resume needed", - NOT complete. Bare-IRI collection items come back with only ID set - (Type==""); re-fetch them. -- SSRF egress guard is ON by default; config.AllowPrivateAddresses / - ClientOptions.AllowPrivateAddresses (env ALLOW_PRIVATE_FETCH, dev-only) - disables it. ANY test/consumer hitting 127.0.0.1 httptest servers must - set AllowPrivateAddresses=true or fetches are blocked at dial time. -- Task 06 inbox wiring: Verifier.Verify(ctx, req, body) returns the signing - actor id, enforces same-authority binding (actor.ID host == keyId host) - and requires host+date+(request-target)+digest signed. It does ONE - fresh-key refetch on verify failure (key rotation) — task 06 should gate - that retry (only when key came from cache) to bound forgery amplification. - ServiceActor.DocumentJSON() ready to serve at /actor; inbox convention - https://{host}/inbox. service_keys table (migration 005) holds the - bridge's RSA key UNENCRYPTED (documented tradeoff; not user key material). -- Lemmy HTTP-sig facts (activitypub-federation-rust): Digest required on - EVERY request incl. GET; keyId is {actorID}#main-key; hs2019 treated as - rsa-sha256; 1h date-skew window. -- .claude/ is gitignored (session/tooling state, incl. scheduled_tasks.lock). - -### From task 03 (identity + virtual repo layer — tasks 04/05/06 consume this) -- internal/repo.Manager is the ONLY write path into repos. PutRecord/ - DeleteRecord return (*repo.CommitResult, error) — {RecordCID (empty for - deletes), CommitCID, Rev, Seq, NoOp}. NewManager returns (*Manager, - error) (nil db/keys rejected). Records must carry non-empty `$type`. - Identical re-put = idempotent NO-OP: NoOp=true, Seq=0, same cid+rev, - NO new commit/firehose event (deterministic rkeys rely on this). -- repo.DeterministicTID(published time.Time, canonicalAPID string) - (syntax.TID, error) is task 05's rkey function. FAILS CLOSED on zero/ - pre-epoch published (callers still gate on ap.Time.OK()). For second- - precision inputs the microsecond field is filled from sha256(ap_id) — - same-second bulk imports don't birthday-collide the 10 clock-ID bits; - within-second sort order is hash order. GOLDEN-VALUE TESTS pin the - algorithm (tid_test.go) — changing it breaks every persisted at-uri. - Commit revs come from repo_state via NextRev — monotonic per repo - across restarts. ops use typed repo.OpAction consts. -- firehose_events schema for task 04: seq bigserial, did, commit_cid, - prev_data_cid (MST root before commit, NULL on genesis — the sync v1.1 - prevData), since_rev (previous commit's rev, NULL on genesis — the - #commit `since` field), rev, ops jsonb ([{action,path,cid,prev}]), - car bytea (CARv1, ROOT/COMMIT BLOCK FIRST, contains commit + MST-diff + - record blocks), created_at. Appended in the SAME tx as the commit. - Commits are v3, Prev always null. -- Commit serialization: every commitWrite tx takes GLOBAL - pg_advisory_xact_lock(0x7469646570636d) (distinct from testutil's - session lock 0x7469646570 — keep them distinct). This guarantees - seq order == commit-visibility order, so task 04 may tail with naive - `WHERE seq > cursor` — any future writer bypassing repo.Manager breaks - that. Per-DID mutex + repo_state row lock remain as backstops. - [SUPERSEDED by task 12: blocks now has GC (invariant in - internal/repo/gc.go), readers hold REPEATABLE READ snapshots instead - of relying on append-only, and ExportCAR is reachable-set-only.] -- identity.Minter.MintActor mints did:plc via MODERN plc_operation genesis - ops (indigo's plc package only has the deprecated legacy `create` op — - don't use it): rotationKeys=[bridge escrow key], verificationMethods. - atproto=per-actor key, signed by the escrow rotation key (enables later - claiming). Minter does NOT write bridged_actors — callers (05/06) upsert - the returned Identity{DID, Handle, DIDKey, SigningKeyEncrypted}; handle - uniqueness race is caught by the bridged_actors_handle_key index. -- Minting failure semantics (tasks 05/06): a failed mint can leave a - registered DID on the directory (PLC ops are forever); orphans are - slog.Error'd with did+handle. On handle-collision retry, callers should - eventually REUSE the minted DID via a PLC updateHandle op, not re-mint - (deferred). Unrepresentable usernames (all-CJK/emoji) get deterministic - u<10-hex-of-sha256> labels; collision suffixes shorten the base so the - 63-char DNS label limit holds. No mint rate limiting yet — REVISIT in - task 06 when inbound AP activity can trigger minting (abuse vector). -- Key custody: identity.Custodian (AES-256-GCM under 32-byte BRIDGE_KEK, - new env var, dev default is a fixed public key, required in prod). - Ciphertexts are AAD-bound to the DID — copying signing_key between rows - breaks decryption. Escrow rotation key lives ENCRYPTED in service_keys - row "plc-rotation" (unlike the plaintext RSA service key; NOTE the - column is named private_key_pem but holds sealed ciphertext — rename - candidate). identity.ActorKeys implements repo.SigningKeys — - SigningKey(ctx, did, use repo.KeyUse): tombstoned actor + KeyUseWrite → - IsTombstoned (frozen); tombstoned + KeyUseDelete → key RELEASED, so - task 05's Delete(Actor) → scrub-records flow works regardless of - consent-flip ordering. Residual TOCTOU: a consent flip racing an - in-flight commit can let that ONE commit land (consent read is outside - the commit tx — full fix deferred; fine for single-writer v1). -- store.BridgedActors grew GetByHandle (minting collision-suffix + - resolveHandle). Handle scheme: name.instance-with-dashes.BRIDGE_HOSTNAME, - lowercased, non-[a-z0-9-] runs → single dash, collisions get -2/-3/…. - Tombstoned actors' handles do NOT resolve. -- Wired in main.go: GET /xrpc/com.atproto.identity.resolveHandle and - GET /.well-known/atproto-did (resolves from Host header; wildcard DNS - requirement documented in README). Task 04 mounts sync endpoints next to - them. -- PLC egress uses ap.NewGuardedHTTPClient (same SSRF guard as the AP - client; ALLOW_PRIVATE_FETCH relaxes it in dev/tests only). -- Testing: internal/testutil.DB(t) is the shared pg harness — it holds a - postgres advisory lock per test process because store/repo/identity - packages share the test DB and `go test ./...` runs packages in parallel. - New pg-using packages MUST use it. PLC tests hit a LOCAL directory only - (default http://localhost:3002, env TIDEPOOL_TEST_PLC_URL, `make plc-up` - or the running Coves dev PLC); they hard-fail on non-loopback URLs and - skip under -short/unreachable. NEVER point tests at https://plc.directory. -- go.mod grew direct deps: go-cid, go-block-format, go-ipld-format, - go-car (v0, indigo's pinned pseudo-version), go-multihash. -- Test DB URL needs ?sslmode=disable (the bare URL earlier in this file - fails with "SSL is not enabled"); Makefile's TEST_DATABASE_URL has it. - -### From task 04 (sync surface — tasks 05/06/08 consume this) -- internal/sync serves the full relay-facing surface: subscribeRepos WS - (cursor replay from firehose_events then live tail; per-conn outbox reads - the DB log — no in-memory queue; slow consumers evicted by write deadline - and resume by reconnecting with their last cursor), getRepo, - getLatestCommit, getRecord (proof CAR), listRepos, getRepoStatus, - describeServer, /xrpc/_health. Protocol tokens exported: sync. - InfoOutdatedCursor / sync.ErrorFutureCursor. -- repo.Manager owns ALL sync reads (ListEvents, SeqBounds, PruneEvents, - GetRecordProof, ListRepos, GetRepoInfo) — internal/sync has zero SQL. - Broadcaster wakes on pg_notify (emitted IN the commit tx, repo. - FirehoseNotifyChannel) with a poll fallback; wake = "rescan the log", - payloads are never trusted. Pings live in a dedicated per-conn goroutine - (pingPump) — replay stretches must never starve liveness (was a real bug: - healthy consumers evicted every 3×pingInterval during deep backfills). -- Mid-replay pruning is SIGNALED: on a seq gap the outbox re-checks - SeqBounds and emits #info OutdatedCursor if the consumer's position fell - off the retained window (benign nextval gaps stay silent). Deactivated - (consent=deleted) repos refuse getRepo/getRecord/getLatestCommit - (RepoDeactivated) and report active:false in listRepos/getRepoStatus. -- DEFERRED to task 06 (documented in streamEvents): the firehose itself - carries only #commit — consent revocation must eventually emit an - #account{active:false} frame (+ scrub delete commits) so subscribers - purge; tombstoned actors' historical events stay replayable for the - retention window until then. Task 06 must also DECIDE: does nobridge - (vs deleted) deactivate the read surface? (bridgy-fed deletes bridged - content on nobridge discovery; we currently only stop new - materialization.) -- FIREHOSE_RETENTION (Go duration, default 72h) drives RunPruner (hourly; - fails closed on retention<=0). RELAY_HOSTS drives requestCrawl on start - (log-only in dev, SSRF-guarded client in prod). created_at on - firehose_events is clock_timestamp() (commit-visible time), not tx start. -- Jetstream verification is MANUAL (compose profile `jetstream`, port 6018 - + README runbook "Verifying with Jetstream") — the DoD line "Jetstream - consumes without errors" is verified by hand, not CI. Task 08's harness - should automate it. -- Not yet done (pre-internet-facing hardening, flagged by security review): - no connection cap / per-IP rate limit on the public surface; getRepo - buffers full CARs in memory (revisit with block GC / tree cache); - PruneEvents is one unbatched DELETE per sweep. -- Deferred design notes for task 04: repo package should own the sync - read API (GetRecordProof for sync.getRecord, ListEvents(sinceSeq, - limit)) rather than task 04 issuing raw SQL against repo tables — - decide there. MST loads are full-tree, one SELECT per node → PutRecord - is O(repo size); fine now, needs a per-DID tree cache before big - community backfills (task 05). SigningKeys could become a SignCommit - capability (keeps key plaintext inside identity; enables KMS later) — - revisit before the interface calcifies. No OnCommit hook yet: task 04's - broadcaster should LISTEN/NOTIFY or poll seq (CommitResult.Seq exists). - -### From task 05 (materializer — task 06 is THE consumer) -- Entry points task 06 drives: MaterializePost(Page/Article), - MaterializeComment(Note), HandleUpdate(obj) [Update; also the Create - dispatcher — an Update for an unseen object just materializes it], - HandleDelete(apID) [object OR actor], DeleteActor (terminal, sets - consent=deleted), SuppressActor (reversible nobridge scrub), - Ensure/RefreshActor, Ensure/RefreshCommunity. All return *Result - (DID/ATURI/CID/NoOp) or a typed error. IsSkip(err) = log-and-never-retry - (consent/tombstone/cycle/unusable input); any other error = retryable. -- SECURITY CONTRACT task 06 MUST honor: RefreshActor/RefreshCommunity (and - HandleUpdate on Person/Group) TRUST an embedded actor doc — call them - ONLY after verifying the activity signature (signer == actor). Content - paths (Ensure*) always re-fetch by IRI and bind the fetched id to the - fetch authority (ap.SameAuthority), so an inline/forged attributedTo - can't mint under a victim's id. Fetched-object ids are authority-bound at - the materializer boundary (comments.go, actors.go), NOT inside - ap.FetchObject (kept generic; ap.SameAuthority is exported for reuse). -- Consent: nobridge on a previously-bridged actor now SCRUBS existing - records (scrubActorRecords) then sets NoBridge (reversible). deleted is - terminal. commitRecord refuses to resurrect an object whose own mapping - is soft-deleted (unordered create-after-delete). KNOWN GAP for task 06: - a Delete arriving BEFORE the object was ever materialized leaves no - mapping to tombstone → a later Create still materializes. Task 06's inbox - needs dedup/ordering (or a tombstone-of-unseen-ids table) to close it. - Also: Undo(Delete)/restore must explicitly clear the soft-delete. -- Lexicon validation: every record validated against vendored lexicons/ - (gojsonschema-equivalent indigo validator). StrictValidation=true in - dev/tests (fail closed); production logs+writes on failure — task 06 - should wire a metric on that and consider strict-first rollout. Text - fields are capped to BOTH maxGraphemes and maxLength bytes (truncateText). -- Blob store: migration 008 blobs(did,cid,bytes,mime); repo.PutBlob/GetBlob - (content-addressed, CID computed server-side); sync getBlob serves with - nosniff + sandbox CSP; image fetches go through the SSRF-guarded ap - client with per-slot size/type caps. MAX_BLOB_BYTES config exists but is - clamped by ap client's maxResponseBytes (5MiB, no knob) — raising it - above 5MiB is currently a no-op (wire it in task 06/07 if needed). -- DID-MINT AMPLIFICATION (task 06 MUST address when wiring the inbox): a - crafted deep comment thread with a distinct fake author per level mints - up to maxAncestorDepth(50) DIDs per delivered object. No rate limiting in - identity.Minter or materialize yet. Gate inbound minting in task 06. -- Deferred (LOW, noted for later): transient media-fetch failure on refresh - drops existing blobs (no carry-forward); a stale actor behind a 403ing - instance drops content instead of serving stale; commitRecord's - PutRecord→PutMapping isn't one tx (self-heals on retry; a Delete landing - in the crash window logs Warn); DeleteActor scrubs records but not blobs - under community DIDs; communityRef uses a Lemmy /c/ heuristic (Mbin /m/ - later). Test gaps still open: embed.images arm + nsfw label shapes are - never lexicon-validated (only external embed is) — add before trusting - "all records validate" for image posts. - -### From task 06 (ingestion — tasks 07/08 consume this) -- internal/ingest is the inbox→queue→dispatch→materializer pipeline. Entry: - POST /inbox (+/actor/inbox alias) verify-sig → authority-bind signer → - dedupe by activity id → Enqueue. GET /actor, /.well-known/webfinger - (service actor only), /.well-known/nodeinfo + /nodeinfo/2.0 - (software.name "tidepool"). Admin (bearer ADMIN_TOKEN, constant-time): - POST/DELETE/GET /admin/communities, POST /admin/communities/backfill. -- TASK 07 SEAM: implement ingest.VoteAggregator — ApplyVote/RetractVote(ctx, - vote *ap.Object, communityIRI string). Wired as NewNoopVotes in main.go; - swap it. Announce{Like|Dislike}→ApplyVote, Announce{Undo{Like|Dislike}}→ - RetractVote; communityIRI is "" for bare (non-announced) votes. RetractVote - MUST treat nil/bare-IRI vote.Object/vote.Actor as no-ops (don't error/ - retry). Aggregate COUNT SEEDING is NOT expressible via per-vote ApplyVote — - if task 07 seeds historical counts from Lemmy's API, add a separate - SeedAggregates method + Backfill wiring (additive, not breaking). -- QUEUE (store.InboxEvents.Enqueue/ClaimNext/Release/MarkPoisoned/ - MarkProcessed): durable pg work queue, per-community serial ordering key, - FOR UPDATE SKIP LOCKED + NOT-EXISTS older-sibling gate. Outcome contract - (identical on Handler.Process, Processor, queue dispatch): nil/IsSkip → - processed; IsValidation → poison; else retry w/ exp backoff (cap 1h); - attempt-cap → poison. CRITICAL for task 07: the queue is now FENCED — - ClaimNext returns claimed_until as a fencing token; Release/MarkPoisoned/ - MarkProcessed take that token + return (applied bool). A stale worker's - write is a silent no-op (applied=false). This exists SO task 07's vote - counter arithmetic isn't double-applied — do NOT bypass repo.Manager-style - or write vote state outside the queue's outcome path without the same - fencing discipline. HEAD-OF-LINE BLOCKING is intentional: a retrying event - blocks its whole ordering key (community) until success or poison. -- Shutdown: processNext no longer classifies ctx-cancellation as a failure - (lease lapses → redeliver); outcome writes use context.WithoutCancel so - completed work is recorded during shutdown. Residual: a shutdown- - interrupted attempt still consumes its ClaimNext attempt increment (not - decremented; harmful poison/error-record behavior is gone). -- AUTHORIZATION MODEL (task 07/08 must preserve): announced (announcer!="") - Delete/Undo require ap.SameAuthority(target, announcer) AND for actor - targets require target==announcer (a community may delete only itself, not - a co-hosted other actor). Announced CONTENT requires communityIRIFrom(obj) - == announcer (no bridging a non-subscribed same-host community). Bare - (announcer=="") Update{Person|Group} is REFRESH-ONLY: never mints/bridges - an unknown actor (was an SSRF+mint-budget vector); embedded actor doc is - trusted only when signer==actor, else forced re-fetch. Bare Delete still - uses SameAuthority(target, signer) (host-granularity — instance is the - trust unit; documented). Undo{Delete} restores mapping BEFORE re- - materialize (commitRecord won't resurrect a soft-deleted mapping) and - COMPENSATES (re-soft-delete + re-tombstone) if HandleUpdate skips/invalid. -- Consent: nobridge scrubs but repo stays ACTIVE (reversible, firehose- - visible delete commits); only `deleted` deactivates the sync read surface. - Still DEFERRED (task 07/08): firehose #account{active:false} frame on - consent revocation (subscribers currently rely on scrub delete-commits); - ap_tombstones grows unbounded (no pruner — mirror FIREHOSE_RETENTION); - NO per-signer/per-IP rate limit on /inbox (queue-flood DoS via many self- - signed identities — coarse admission limit is the hardening item); - ClaimNext does an O(N) row scan when a community's queue backs up behind a - failing event (per-key serialization cost; revisit at scale). Nudge now - cascades on successful claim. -- Backfill: outbox newest→oldest to BACKFILL_MAX_POSTS (100), replies walk - when advertised, tombstone-checked on all three paths now. RESUMABLE-BY- - REDO: truncation (ErrCollectionTruncated) or failures>0 leave - last_backfill_at UNSET so an un-forced re-trigger re-walks; deterministic - rkeys make redo idempotent. main.go drains backfill.Wait() on shutdown - (bounded); TriggerAsync runs on the root ctx (observes shutdown). -- New env: ADMIN_TOKEN (required in prod, dev default dev-admin-token), - BACKFILL_MAX_POSTS (100), MINT_RATE_PER_MINUTE (60), MINT_BURST (120), - INGEST_WORKERS (4). Migration 009: queue columns on inbox_events - (payload, actor_id, ordering_key, attempts, next_attempt_at, claimed_until, - failed_at) + ap_tombstones table. store grew Tombstones (Record/Exists/ - Remove), APObjects.Restore, and the fenced InboxEvents queue API. -- ap.Client.ResolveKeyDetailed exported + Object.Replies field added. Verify - fresh-key refetch is now cache-gated (fires only when the failing key came - from cache) — bounds forgery amplification. -- Deferred TEST gaps (task 08 harness): concurrent-worker queue stress - (SKIP-LOCKED ordering only tested sequentially); mint-gate "retry via queue - backoff" only asserted at unit level, not end-to-end; webfinger/nodeinfo - vs what real Lemmy actually queries (needs live Lemmy). activityID rand- - fail path is guarded but unit-untestable (Go 1.24+ crypto/rand failure is - a fatal crash, not a returnable error). - -### From task 07 (vote aggregates — task 08 consumes this) -- internal/votes: Aggregator implements ingest.VoteAggregator. - NewAggregator(db, apObjects, communities, records, logger) — `records` is - a narrow votes.RecordReader (satisfied by repo.Manager, wired in main.go). - XRPC: GET /xrpc/social.coves.bridge.getVoteAggregates (lexicon at - lexicons/social/coves/bridge/, README documents it as the AppView - integration point). ≤100 uris counted PRE-dedupe; malformed at-uri → - InvalidRequest 400 (validated via syntax.ParseATURI); unknown uris - omitted; response preserves request order; Cache-Control public/max-age=30; - per-IP token bucket on RemoteAddr only (XFF deliberately ignored), - FAIL-CLOSED at 50k buckets (unknown IPs get 429 during rotation floods), - sweep throttled to 1/min, injectable clock for tests. -- Vote state machine: append-only vote_events, invariant ≤1 non-undone row - per (voter,subject) is APP-enforced only — every mutation MUST take the - vote_aggregates row lock FIRST (lockAggregate upsert / SELECT FOR UPDATE), - then recompute-in-tx. Dedupe by activity_id (unique, first-writer-wins — - forged-id squatting narrowed by authority binding but ids remain - unauthenticated). RetractVote targets the undone activity's own id when - vote.ID != "" (replay-proof), direction-match fallback for id-less undos. - Dup activity id is probed BEFORE creating an aggregate row (no 0/0 rows). -- AUTHORITY MODEL for votes: announced votes require the announcer to - resolve to a followed community AND subject∈that community — posts via - mapping.DID == community DID (NOT SameAuthority: Lemmy hosts post objects - on the AUTHOR's instance), comments via one GetRecord read of reply.root's - DID. Bare Like/Dislike: inbox already binds actor↔signer (403). - Bare Undo{vote}: Handler.authorizeBareVote (consent.go, host-granularity - SameAuthority like bare Delete). All mismatches drop at debug as nil — - queue outcome contract (nil=processed) preserved everywhere; only real DB - errors are retryable. -- Seeding: seeded_upvotes/seeded_downvotes baseline columns (beyond spec — - deliberate: served = seeded + live, re-seed idempotent, never clobbers - live votes). LemmySeeder.SeedPostCounts posts-only via SSRF-guarded - client, presence-checked decode (missing post_view/counts → error, old - baseline survives; a wrong-shape 200 can NOT zero-clobber). Wired as - optional ingest.CountSeeder in backfill (best-effort, Warn on failure, - never affects run outcome). SEED_COUNTS_FROM_API (default on; strict bool - parse rejects typos). Known accepted drift: a baseline voter who flips - sends Undo for a like we never saw (no-op) — stale until re-seed. -### From task 09 (relay pipeline — tasks 10/11 MUST know) -- The e2e pipeline is now Lemmy → tidepool → BigSky relay → Jetstream: - EVERY scenario's events transit relay validation (DID resolution against - the local PLC, per-commit sig verification). jetstream-direct (compose - profile `direct`, host :6038) is a debug tap only. Relay on host :2480 - (admin key e2e-relay-admin-key); helpers: relayPDSList/relayListRepos/ - relayGetLatestCommit in tests/e2e/helpers.go. -- CRITICAL for scenario design: bigsky's parallel indexer (100 workers - keyed by repo DID) preserves per-repo order but NOT cross-repo order — - author profiles and community posts routinely swap. The e2e listener now - BUFFERS consumed-but-unmatched commit events and rescans them on later - awaits; await predicates MUST be PURE (side-effecting accounting belongs - in drain loops — scenario 8 was rewritten that way). Cross-repo ordering - assertions were removed from scenarios 2/3/6; presence + linkage is what - survives a relay. Carry to Coves AppView: it cannot rely on - profile-before-post through relay infra (FOLLOWUPS "Relay pipeline"). -- New config: ALLOW_DEV_REQUEST_CRAWL (dev-only, refused in production) - makes dev actually send requestCrawl; RequestCrawlAll now retries per - relay (5s × 24, vars compressible in tests) because bigsky's requestCrawl - handler calls BACK into the announcing host's describeServer before - subscribing — the first attempt races the bridge's own listener (observed - live: attempt 1 fails, attempt 2 lands). -- bigsky env facts (verified against image --help + source; Coves' compose - stanza is wrong): ATP_PLC_HOST (plc), RELAY_ADMIN_KEY, DATA_DIR, - RESOLVE_ADDRESS (defaults to PUBLIC 1.1.1.1 — set 127.0.0.11), - HANDLE_RESOLVER_HOSTS=tidepool (trial-host resolver GETs - /.well-known/atproto-did with the handle as Host header — handles VERIFY, - 0 failures in a full run), BSKYLOG_LOG_LEVEL; --crawl-insecure-ws has NO - env binding (command arg). Fresh relay refuses ALL non-admin requestCrawl - (new-PDS-per-day limit 0, checked before trusted domains) → one-shot - relay-bootstrap service raises it; tidepool depends_on its completion. - Image has no arm64 manifest (platform: linux/amd64, emulated on Apple - Silicon). Relay gets its own postgres with TWO dbs (bgs + carstore). -- Tombstone visibility through the relay is a DEAD END until task 11: - bigsky has no getRepoStatus, filters tombstoned repos from listRepos, and - only learns account state from #account frames the bridge doesn't emit - yet. When task 11 adds the #account frame, add the relay-side assertion - (repo disappears from relay listRepos after consent revocation). -- Timing budgets: eventTimeout 90s→120s (relay hop + first-sight DID work + - amd64 emulation), burstTimeout 3m (scenario 8 is drain-based now), - crawlTimeout 3m (covers the announce retry window). Suite runs ~114s. -- DEFERRED (task 08+): vote_events grows unbounded (no pruning of - superseded/undone rows — mirror FIREHOSE_RETENTION treatment alongside - ap_tombstones); actor-delete/consent revocation does NOT scrub that - actor's vote_events rows (inconsistent with scrub posture elsewhere; - counts are anonymous on the wire so exposure is low); no concurrency - stress test of the aggregate-row locking claim (sequential tests only); - subject resolution happens outside the mutation tx (narrow TOCTOU with a - racing Delete, documented); no upper sanity cap on seeded counts; - comment count seeding skipped (per-comment API calls would triple - backfill egress). - -### From task 10 (e2e scenario completion — tasks 11/12 MUST know) -- The task shipped TWO BRIDGE-CODE changes, not just tests (both were - prerequisites for the Delete(Actor) scenario to be real; both are - review-worthy): - 1. **Instance (Site) actor at the origin apex** (`GET /`, ingest - inbox.go handleInstanceActor + ap.ServiceActor.InstanceDocumentJSON). - Lemmy's federate worker resolves send-to-all-instances targets - (Delete{Person}!) to the stored remote SITE row's inbox and silently - skips (`no inboxes`) peers without one; Lemmy creates the row by - fetching the peer's origin apex on every Person/Community from_json. - Wire trap: the Instance protocol enum accepts ONLY `Application` — - the OPPOSITE of /actor, where Lemmy's Person enum needs `Service`. - Existing peer Lemmys need an actor re-fetch (24h TTL) + worker - restart before deletions start flowing (worker caches site=None for - its lifetime). - 2. **Tombstone-verified self-Delete acceptance** (ingest inbox.go - tombstonedSelfDelete + ActorFetcher option, 3 regression tests). - Lemmy signs the account-deletion Delete{Person} with a key whose - actor doc its origin ALREADY serves as 410 Gone — unverifiable - unless cached. The inbox accepts a bare self-referential Delete when - verification failed on IsTombstoned AND an independent SSRF-guarded - fetch of the claimed actor's own IRI confirms Gone (origin word = - same host-granularity trust as bare-Delete SameAuthority; forgery ≡ - truth). Any other tombstoned-signer payload → definitive 401 (5xx - would head-of-line block Lemmy's forever-retrying per-instance - queue). Review fixes (post-7-model review): the confirmation is - three-way — origin says Gone → accept; origin answers definitively - otherwise (live actor / not-an-actor / 404) → 401; transport/5xx/ - timeout → 503 so the sender redelivers (a blip must never - permanently drop a deletion) — and the confirmation fetch pins - every redirect hop to the actor IRI's own authority - (ap.FetchActorSameAuthority), so an open redirect on the origin - cannot bounce it to an attacker host serving 410. -- **Task 11 rate limit MUST cover the tombstonedSelfDelete branch:** it is an unauthenticated POST → confirmation fetch → TWO durable writes (inbox_events + ap_tombstones) reachable with a cheap 410-serving endpoint; consider a dedicated per-IP cap (flagged in tasks/11-hardening.md). -- **Jetstream cursor full-replay middle band** (measured with a ws probe): - cursor ≤ newest stored event → precise µs replay; cursor > server-now → - live-tail; cursor BETWEEN newest event and now (any "from now" subscribe - on a quiet stream) → replays the ENTIRE store. Every cursorNow() - listener gets full-history replays; negative assertions must be bounded - by an observed event's time_us (see unsubscribe scenario + cursorNow - doc). Carry to Coves AppView: cursor=now is not a dedupe boundary. -- **Consumer-side lexicon validation of blobs requires - atdata.UnmarshalJSON**, not encoding/json — indigo's SchemaBlob.Validate - demands a typed atdata.Blob ("expected a blob" on raw maps). The e2e - validateLexicon now mirrors the materializer's write-side decode; any - JSON Jetstream consumer that validates records with blobs (Coves!) must - do the same. Found by the image-post scenario — the first blob ever to - cross the suite's validation. -- Lemmy 0.19.19 wire facts (source-verified + observed): save_user_settings - federates NOTHING (no Update{Person} variant exists) — consent marker - changes are discoverable only via TTL re-fetch on the actor's next - activity (e2e compose sets PROFILE_REFRESH_TTL=2s to drive it; an - inactive actor's opt-out is never discovered — task 11 candidate: - periodic consent re-scan); delete_account is POST with required - delete_content and federates ONE Delete{Person} with nonstandard - removeData (no per-object deletes); image-post attachment is - {"type":"Image","url",…,"name":alt} with mediaType DROPPED (type field + - extension fallback are the only discriminators; alt text doesn't survive - the Link fallback); pictrs upload is POST /pictrs/image multipart - "images[]", Bearer jwt. -- GET /admin/communities lists accepted/pending ONLY — an unsubscribed - community disappears from it (don't poll it for follow_state=none). -- e2e suite is now 17 scenarios, ~230s full run, re-runnable against an - accumulated stack. TestZZ_SuiteEndSweep (zz_ file = alphabetically last - = runs last) replays the whole firehose from cursor 1 unfiltered and - re-vets everything incl. per-DID rev monotonicity from history start. - TestDeleteActor pins "relay still lists tombstoned repo" — task 11's - #account frame must FLIP that assertion. -- Task-10 stretch (low MINT_RATE_PER_MINUTE compose variant) skipped as - specced — needs its own compose profile (FOLLOWUPS, Ingestion). - -### From task 11 (hardening — task 12 MUST know) -- **firehose_events now carries TWO event kinds** (migration 011): `commit` - and `account` (nullable commit columns + a shape CHECK per kind). - repo.Event grew Kind/AccountActive/AccountStatus; any code iterating - events must switch on Kind (sync/subscribe.go does; e2e vetEvent skips - non-commit Jetstream kinds already). `repo.AppendAccountEvent` takes the - SAME global commit advisory lock as record commits — seq order still == - visibility order; task 12's perf work must preserve that for both kinds. -- **#account status token is "deleted", and so is the bridge's own - getRepoStatus/listRepos status** (was "deactivated"). Wire facts - (source-verified, pinned bigsky): on #account it re-resolves the DID doc - and REQUIRES the sender to be the DID's authoritative PDS; "deleted" → - tombstoned=true (listRepos filters `NOT tombstoned`) + carstore purge; - "deactivated"/"suspended"/"takendown" filter from listRepos but do NOT - purge. Emitted ONLY by DeleteActor (terminal); nobridge stays frameless - (repo active, reversible). -- **repo.PutRecordTx(… TxSideEffect)** is the new atomic seam: the hook - runs inside the commit tx (also on the NoOp re-put, with res.NoOp=true) - and an error rolls the RECORD back too. The materializer's ap_objects - mapping now rides it (store.APObjects.PutMappingTx). Hooks run under the - global advisory lock — keep them tiny; task 12's MST cache must not - change hook semantics. -- **internal/ratelimit** is the shared keyed token-bucket (extracted from - votes; sweep throttle + 50k fail-closed cap). Consumers: votes XRPC - (429), sync surface (429, _health exempt; subscribeRepos also has a - reserve-then-check connection cap), inbox (per-IP pre-body + per-signer - post-verify + a dedicated tombstone-confirmation cap INSIDE - tombstonedSelfDelete, before the confirmation fetch). INBOX REFUSALS ARE - 503, NEVER 4xx — Lemmy's federation crate retries 5xx but permanently - drops 4xx; the tombstone-cap refusal is a DEFER for the same reason. - Defaults are deliberately generous (suite = canary, ran green on - defaults); all envs in README's config table. -- **internal/prune.Run** is the shared retention loop (fail-closed on - retention<=0). Pruners: firehose (batched now, 1000/statement, oldest-up - so a partial sweep keeps the retained suffix contiguous), ap_tombstones - (30d), undone vote_events (90d — an undone row is also its activity id's - dedupe record, so replay protection now has that horizon; live rows - never pruned). -- **Blob scrub caveat**: blobs are content-addressed per (did,cid) with no - reference tracking — scrubbing a deleted actor's record deletes blob rows - other records in the same repo could share (byte-identical media). - Accepted + commented on repo.DeleteBlob; a blob-refcount would close it - if it ever matters. repo grew DeleteBlob/DeleteBlobsForDID. -- **votes.Aggregator.ScrubVoter** locks ALL affected aggregates in - deterministic subject order (ORDER BY + FOR UPDATE) before deleting — - any future multi-subject vote mutation must lock in the same order or - risk deadlock against it. -- **Follow retrier** (ingest.FollowRetrier, migration 012): every - set-to-pending stamps follow_requested_at + increments follow_attempts - (SetFollowState does it); resend consumes the attempt BEFORE sending so - a hanging peer can't get an unbounded budget; none resets. Default: - pending >2m → resend, 5 total sends, 1m sweep. -- **service_keys.private_key_pem → key_material** (migration 013); - store.ServiceKey.PrivateKeyPEM → KeyMaterial (identity/ap callers - updated). -- **ap.ClientOptions.MaxMediaBytes**: FetchMedia's outer clamp, wired from - MAX_BLOB_BYTES; the JSON-object cap (MaxResponseBytes) stays independent - at 5MiB. -- **Delete-before-create verdict**: README was RIGHT, FOLLOWUPS was stale — - task 06's handleDelete records the ap_tombstones marker before - HandleDelete and materializeContent checks it; - TestCreateAfterDeleteTombstone pins the whole ordering. No code change - needed; the false FOLLOWUPS entry is annotated. -- **Validation-failure metric**: materialize.ValidationFailures (expvar, - "tidepool_lexicon_validation_failures", counted in strict AND - log-and-write modes), served on bearer-protected GET /admin/metrics. - Strict-first production rollout still deferred; this counter is its - precondition. -- e2e: TestDeleteActor's task-09/10 pins FLIPPED — it now asserts the - #account frame on the bridge's own firehose (raw CBOR ws helper - dialBridgeFirehose/readBridgeAccountFrame in tests/e2e/helpers.go) and - the repo DISAPPEARING from the relay's listRepos (polled; bigsky - processes the frame async). - -### From task 12 (perf & scale — anything touching repo storage MUST know) -- **`blocks` is no longer append-only-forever.** The replacement invariant - lives at the top of internal/repo/gc.go (delete only if unreachable from - the current head in one REPEATABLE READ snapshot AND older than - BLOCKS_GC_RETENTION, created_at re-checked inside the DELETE). Every - new `blocks` reader must hold a REPEATABLE READ snapshot (GetRecord/ - GetRecordProof/ExportCARTo do; the commit path's loadTree is covered by - the FOR UPDATE head pin) — audit any new reader against gc.go's header. - `blocks.created_at` now means "last written" (ON CONFLICT refresh), not - "first written"; it is the GC race guard and only ever makes GC more - conservative. Retention must comfortably exceed app↔DB clock skew + a - sweep's compute→delete gap (72h default dwarfs both). -- **A future `since` diff export is now HARDER, not just missing**: it - would read historical blocks, which the GC invariant explicitly does - not guarantee. Documented in FOLLOWUPS; decide GC-interaction first. -- **Per-DID MST tree cache** (internal/repo/treecache.go): take() detaches - the entry and validates against the FOR UPDATE head (ABA impossible — - revs strictly increase and are embedded in the commit CID); re-cached - only after durable commit (or the provably-unchanged NoOp head). All - access under the per-DID mutex inside commitWrite ONLY — never add a - read-path consumer. repo.NewManager is now variadic (...Option, - WithTreeCacheSize; MST_CACHE_SIZE env, default 512; n <= 0 disables). - Cold first commit per DID still pays one full-tree load. -- **getRepo streams** (repo.ExportCARTo): reachable-set-only, batched - ANY() fetches (walkBatchSize 256 — a naive per-block walk measured - 698ms/op vs 27.7ms batched; don't "simplify" it away). Mid-stream - failure panics http.ErrAbortHandler so consumers see a transport error - instead of a clean truncated 200; client disconnects log at Debug. - walkReachable is SHARED between export and GC — reachability is - definitionally "what getRepo serves", they cannot drift. -- **ClaimNext is a recursive-CTE loose index scan** over - idx_inbox_events_queue. The ARRAY(...) head-materialization is - LOAD-BEARING (a plain IN regressed the planner to the O(N) scan — - EXPLAIN-verified, commented in inbox_events.go). Semantics unchanged: - per-key serial ordering, claimed_until fencing, SKIP LOCKED. -- **Perf ceilings that remain (for the votes-as-records design revisit):** - the GLOBAL commit advisory lock serializes ALL repos' commits — per-DID - throughput is now ~300 commits/s (3.3ms each) but it is one writer at a - time bridge-wide; firehose_events volume is untouched. Those are design - decisions, not storage perf — storage prerequisites for - votes-as-records now HOLD (FOLLOWUPS updated). -- Deferred: GC-vs-commit concurrency argued + mechanism-pinned, no - goroutine stress test; cold-start loadTree could reuse GetMany batching - if it ever matters; sync_test.go grew an exportCAR seam (mirrors - onUpgrade) for mid-stream failure injection. diff --git a/PLAN.md b/PLAN.md index c9fa508..133e764 100644 --- a/PLAN.md +++ b/PLAN.md @@ -75,48 +75,3 @@ sanctioned side-channel XRPC. Nothing ever strongRefs a vote. actors for Coves users, key claiming/migration for bridged users (escrow the keys, build claiming later), moderation federation, DMs, PieFed/Mbin quirk testing (target the FEP; verify against Lemmy only). - -## Sections (the loop iterates over these, in order) - -| # | Task file | What | ~LOC | Depends on | -|---|-----------|------|------|-----------| -| 1 | tasks/01-scaffold-storage.md | Repo scaffold, config, errors, migrations, mapping-table store | 900 | — | -| 2 | tasks/02-ap-protocol.md | AP vocab types, WebFinger, signed fetch, collection paging | 1000 | 1 | -| 3 | tasks/03-identity-repos.md | did:plc minting, key custody, virtual repo layer (MST/CAR commits) | 1200 | 1 | -| 4 | tasks/04-sync-firehose.md | `com.atproto.sync.*` XRPC + subscribeRepos WebSocket serving | 1000 | 3 | -| 5 | tasks/05-materializer.md | AP → `social.coves.*` translation, rkeys, strongRef resolution | 1200 | 2,3 | -| 6 | tasks/06-ingestion.md | Inbox + sig verification, Follow lifecycle, Announce handling, backfill, consent | 1200 | 5 | -| 7 | tasks/07-vote-aggregates.md | Vote ingestion, aggregate store, side-channel XRPC | 800 | 6 | -| 8 | tasks/08-e2e-harness.md | docker-compose with real Lemmy + PLC + Jetstream, E2E tests, lexicon conformance | 1000 | 4,6 | - -### v1.1 sections (added 2026-07-10 — FOLLOWUPS backlog + relay pipeline) - -| # | Task file | What | ~LOC | Depends on | -|---|-----------|------|------|-----------| -| 9 | tasks/09-e2e-relay.md | BigSky relay in e2e; full pipeline Lemmy → bridge → relay → Jetstream; real requestCrawl | 600 | 8 | -| 10 | tasks/10-e2e-scenarios.md | Scenario completion: image, consent, Delete(Actor), unsubscribe, community update, vote hammer, suite-end sweep | 800 | 9 | -| 11 | tasks/11-hardening.md | Pre-internet-facing hardening: rate limits, #account frame, ordering gaps, pruners, housekeeping | 1000 | 10 | -| 12 | tasks/12-perf-scale.md | MST cache, getRepo streaming/reachable-set, blocks GC, ClaimNext — prerequisite for any votes-as-records revisit | 800 | 11 | - -## Loop protocol - -Each iteration (one section per iteration): - -1. Read `LOOP_STATE.md`; pick the first task not marked `done`. If all are - `done`, stop the loop. -2. Mark it `in-progress`. Spawn a Fable implementation agent with the task - file, PLAN.md, and pointers to Coves/bridgy-fed reference code. The agent - implements the section, keeps `go build ./... && go vet ./... && go test - ./...` green, and commits nothing itself. -3. Review the working-tree diff with the **second-opinion** skill. -4. Fix confirmed findings (subagents for independent fixes), re-verify build - and tests. -5. Commit with a descriptive message, update `LOOP_STATE.md` (status, commit - hash, notes for the next iteration — surprises, deferred TODOs, interface - changes later sections must know about). -6. Schedule the next iteration. - -Reference material: `~/Code/coves` (the AppView; lexicons at -`internal/atproto/lexicon/social/coves/`, consumers at -`internal/atproto/jetstream/`), `~/Code/bridgy-fed`, `~/Code/granary` -(AP↔atproto translation), `~/Code/arroba` (virtual PDS in Python). diff --git a/README.md b/README.md index bb4761c..96c3f0d 100644 --- a/README.md +++ b/README.md @@ -4,12 +4,11 @@ Tidepool is a read-only ActivityPub → atproto bridge for the threadiverse. It follows Lemmy/PieFed/Mbin communities (FEP-1b12 group federation), materializes their posts, comments, and profiles as `social.coves.*` records in a virtual PDS it operates itself, and serves them over -`com.atproto.sync.*` so the [Coves](https://github.com/coves-social) AppView +`com.atproto.sync.*` so the [Coves](https://github.com/BrettM86/coves) AppView indexes fediverse communities exactly as it indexes native ones. Votes stay bridge-side as aggregates behind one sanctioned XRPC. -See **[PLAN.md](PLAN.md)** for the architecture, locked design decisions, -and the task-by-task build plan (`tasks/`). +See **[PLAN.md](PLAN.md)** for the architecture and locked design decisions. ## Quick start @@ -549,4 +548,4 @@ on-demand certificate issuance (e.g. Caddy) or per-instance wildcard certs. ## License Tidepool is licensed under the [GNU Affero General Public License v3.0](LICENSE) -(AGPL-3.0), the same license as [Coves](https://github.com/coves-social). +(AGPL-3.0), the same license as [Coves](https://github.com/BrettM86/coves). diff --git a/tasks/01-scaffold-storage.md b/tasks/01-scaffold-storage.md deleted file mode 100644 index 0f67b26..0000000 --- a/tasks/01-scaffold-storage.md +++ /dev/null @@ -1,56 +0,0 @@ -# Task 01 — Scaffold, config, storage foundation (~900 LOC) - -## Goal -A buildable Go project with the persistence spine every later section writes -through: the `ap_objects` mapping table, bridged-actor registry, community -registry, and event dedupe — plus config, errors, logging, migrations. - -## Deliverables -- `go.mod` (module `tidepool`, Go 1.25): `bluesky-social/indigo`, - `go-chi/chi/v5`, `pressly/goose/v3`, `lib/pq`, `stretchr/testify`. -- `cmd/tidepool/main.go` — wires config, DB, migrations-on-start (dev only), - chi router with `/healthz`, graceful shutdown. Subsystems register as they - land in later tasks; keep main thin. -- `internal/config/config.go` — env vars with logged dev defaults (Coves - style, no config lib): `DATABASE_URL`, `LISTEN_ADDR`, `BRIDGE_HOSTNAME` - (public domain, e.g. `tidepool.example`), `PLC_DIRECTORY_URL`, - `BRIDGE_SERVICE_DID` (optional pre-provisioned), `USER_AGENT`. -- `internal/errors/errors.go` — sentinel + typed errors mirroring - `coves/internal/core/errors`: `ErrNotFound`, `ErrAlreadyExists`, - `ValidationError`, wrap with `%w`, `IsNotFound()` helpers. -- `internal/db/db.go` — open/ping/pool settings; `internal/db/migrations/` - goose SQL files (`NNN_description.sql`, `-- +goose Up/Down`): - - `ap_objects`: id, ap_id (unique), ap_type, origin_instance, did, - collection, rkey, at_uri (unique), cid, created_at, indexed_at, - deleted_at (soft). Index on (did, collection), origin_instance. - - `bridged_actors`: ap_actor_id (unique), actor_type (person|group), did - (unique), handle, signing_key_multibase (secp256k1 private key — - NOTE: encrypt-at-rest is task 03's concern; column is bytea), - consent_state (ok|nobridge|deleted), profile_synced_at, created_at. - - `communities`: ap_group_id (unique), did, preferred_username, instance, - follow_state (none|pending|accepted), followed_at, last_backfill_at. - - `inbox_events`: activity_id (unique) for dedupe, received_at, type, - processed_at, error text. -- `internal/store/` — one repo per table behind interfaces - (`interfaces.go` per Coves convention), raw parameterized SQL, idempotent - upserts (`ON CONFLICT`), soft deletes. This is the package tasks 03/05/06 - consume; get the interfaces right: `PutMapping`, `GetByAPID`, - `GetByATURI`, `ResolveStrongRef(apID) (uri, cid, error)`. -- `Makefile`: `build`, `test` (spins up postgres-test via compose, runs - goose, `go test ./... -short`), `db-migrate`, `lint`, `fmt`. -- `docker-compose.dev.yml`: postgres + postgres-test (profiles like Coves). -- `.golangci.yml`, `.gitignore`, `README.md` (one screen: what Tidepool is, - pointer to PLAN.md). - -## Definition of done -- `go build ./... && go vet ./... && go test ./...` green. -- Store tests run against real postgres (compose), covering upsert - idempotency, strongRef resolution hit/miss, consent-state transitions. -- Migrations apply cleanly up and down. - -## References -- Conventions: `~/Code/coves/CLAUDE.md`, `~/Code/coves/internal/core/errors/`, - `~/Code/coves/internal/db/migrations/`, `~/Code/coves/Makefile`. -- Schema inspiration: `~/Code/bridgy-fed/models.py` (Object/User keyed by - ap_id with `copies` cross-protocol id list — our `ap_objects` is the - relational version of `copies`). diff --git a/tasks/02-ap-protocol.md b/tasks/02-ap-protocol.md deleted file mode 100644 index edf7dc2..0000000 --- a/tasks/02-ap-protocol.md +++ /dev/null @@ -1,53 +0,0 @@ -# Task 02 — ActivityPub protocol layer, client side (~1000 LOC) - -## Goal -Everything needed to *read* the Lemmy-flavored fediverse: typed AP vocab, -WebFinger resolution, HTTP-signature-signed fetch, and collection paging. -No inbox/server here (task 06); no translation (task 05). - -## Deliverables -- `internal/ap/vocab.go` — structs + JSON (un)marshalling for the objects - Lemmy/FEP-1b12 actually emit. Tolerant parsing (json-ld-lite): fields that - may be string-or-object (`to`, `cc`, `attributedTo`, `object`, `icon`) - handled with custom unmarshallers. Types: `Group`, `Person`, `Page` - (posts), `Note` (comments), `Article`, `Create`, `Update`, `Delete`, - `Announce`, `Follow`, `Accept`, `Undo`, `Like`, `Dislike`, `Tombstone`, - `Image`, `Link/Hashtag`, `Collection`/`OrderedCollection`(+Page), - `source` (markdown), `language`, Lemmy extensions (`sensitive`, - `commentsEnabled`/`postingRestrictedToMods`, `stickied`). -- `internal/ap/webfinger.go` — resolve `!community@instance` / - `user@instance` → actor URL via `/.well-known/webfinger`, and reverse - (fetch actor, read `preferredUsername` + host). -- `internal/ap/httpsig.go` — draft-cavage HTTP signatures exactly as Lemmy - validates them: sign `(request-target) host date digest` with RSA-SHA256 - (Lemmy requires RSA actor keys for AP interop; the bridge's *AP-side* - service key is RSA — distinct from atproto secp256k1 repo keys). Signed - GET (many instances require authorized fetch) and signed POST (used by - task 06 for Follow). Verify() lives here too (task 06 consumes it): - fetch remote actor's publicKeyPem with caching, check date skew, digest. -- `internal/ap/client.go` — fetch AP objects with - `Accept: application/activity+json`, retries with backoff, per-host rate - limiting, 5xx/410 handling (410 → treat as deleted), response size cap, - in-flight dedupe. `FetchActor`, `FetchObject`, `FetchCollection` (page - through `OrderedCollectionPage.next`, cap pages, yield items). -- `internal/ap/service_actor.go` — the bridge's own `Application` actor - document served later at `https://BRIDGE_HOSTNAME/actor` (generation + - RSA keypair persistence in DB; the HTTP route itself can land in task 06 - with the inbox, but the document/key logic lives here). -- Golden-file tests: real captured Lemmy JSON (fetch a handful of live - fixtures from lemmy.world/lemmy.ml at implementation time — a Group, a - Page with image embed, a Note with parent, an Announce{Create{Note}}, - a Like — and commit them under `internal/ap/testdata/`). Parse → assert - fields. Round-trip signing test with a local verifier. - -## Definition of done -- All fixtures parse without error; unknown fields ignored, never fatal. -- Signed GET against a live Lemmy instance works (manual smoke, documented - in the task notes; CI tests use recorded fixtures only). -- `go test ./...` green. - -## References -- `~/Code/granary/granary/as2.py` (AP quirk handling), - `~/Code/bridgy-fed/activitypub.py` (signed fetch, key caching, Lemmy - compat notes), FEP-1b12 spec, Lemmy federation docs - (join-lemmy.org/docs/contributors/05-federation.html). diff --git a/tasks/03-identity-repos.md b/tasks/03-identity-repos.md deleted file mode 100644 index d671173..0000000 --- a/tasks/03-identity-repos.md +++ /dev/null @@ -1,57 +0,0 @@ -# Task 03 — Identity minting + virtual repo layer (~1200 LOC) - -## Goal -The atproto half of the bridge's core: mint did:plc identities for bridged -actors (users and communities), custody their signing keys, and maintain a -real signed Merkle Search Tree repo per DID using indigo primitives. This is -the Go port of arroba's job. - -## Deliverables -- `internal/identity/minter.go` — create did:plc via PLC directory - (`PLC_DIRECTORY_URL`): generate secp256k1 keypair per actor, build + - sign the genesis PLC operation (rotation key = bridge escrow key; - verification key = per-actor key), POST to directory. Handle scheme: - communities `technology.lemmy-world.`, users - `alice.lemmy-world.` (dots for @/!, dashes for host - dots; collision-suffix if needed). PDS endpoint in the DID doc = - `https://BRIDGE_HOSTNAME`. Handle resolution: since bridged handles are - subdomains of the bridge, serve - `/.well-known/atproto-did`-style resolution — implement the - `com.atproto.identity.resolveHandle` XRPC + wildcard-DNS assumption; - document the DNS requirement in README. -- `internal/identity/keys.go` — key custody: per-actor secp256k1 private - keys encrypted at rest (AES-GCM with `BRIDGE_KEK` env key), escrow - rotation key handling. Claiming/migration is out of scope; design the - storage so it's possible later. -- `internal/repo/repo.go` — per-DID repo built on indigo's `mst`, `repo`, - and `carstore`/blockstore packages: `PutRecord(did, collection, rkey, - record) (cid, rev, error)`, `DeleteRecord`, `GetRecord`, each producing a - properly signed commit (v3, `rev` TID monotonic per repo) persisted in - postgres (blocks table: did, cid, bytes; repo_state: did, head_cid, rev). - Serialize writes per DID (per-DID mutex or single writer goroutine). -- `internal/repo/events.go` — every commit also appends to a durable - `firehose_events` table (seq bigserial, did, commit CID, CAR slice of - the commit blocks, ops, rev, time) — task 04 serves this. Emitting here - keeps commit+event atomic in one tx. -- `internal/repo/tid.go` — TID generation: normal clock TIDs for repo - `rev`; deterministic content TIDs for rkeys (timestamp bits from AP - `published`, clock-ID bits from hash of ap_id) — exposed for task 05. -- Migrations: `blocks`, `repo_state`, `firehose_events`, plus - `bridged_actors` alterations if needed. -- Tests: mint against a real PLC container (compose has one in Coves dev — - add `plc` service to our compose); repo round-trip (put N records, - export CAR, re-load with indigo, verify MST root + signatures); rkey - determinism (same AP id + published → same TID; different id → different). - -## Definition of done -- Can mint a DID on a local PLC directory, create its repo, write a - profile record, read it back, and verify the commit signature with the - minted key. -- Deterministic rkeys are stable across process restarts. -- `go test ./...` green (PLC-dependent tests behind `-short` skip). - -## References -- `~/Code/arroba/arroba/` — `repo.py`, `mst.py`, `did.py`, `storage.py` - (the exact semantics being ported). -- indigo: `repo`, `mst`, `atproto/crypto`, `api/atproto` packages. -- Coves compose PLC service: `~/Code/coves/docker-compose.dev.yml`. diff --git a/tasks/04-sync-firehose.md b/tasks/04-sync-firehose.md deleted file mode 100644 index 79894b8..0000000 --- a/tasks/04-sync-firehose.md +++ /dev/null @@ -1,50 +0,0 @@ -# Task 04 — com.atproto.sync.* serving + firehose (~1000 LOC) - -## Goal -Make Tidepool a valid `subscribeRepos` upstream so Jetstream (and any relay) -can consume it. After this task, records written by task 03's repo layer are -visible to the Coves AppView with zero Coves changes. - -## Deliverables -- `internal/sync/server.go` — chi routes for the sync XRPC surface Jetstream - and relays actually use: - - `com.atproto.sync.subscribeRepos` — WebSocket, DAG-CBOR frames - (header `{op:1, t:"#commit"}` + body per the sync spec: seq, rebase, - tooBig, repo, commit, rev, since, blocks (CAR slice), ops, time). - Cursor support: `?cursor=N` replays from `firehose_events` then goes - live; per-connection outbox goroutine with slow-consumer disconnect; - `#info` frame `OutdatedCursor` when cursor < oldest retained seq. - - `com.atproto.sync.getRepo` (full CAR export), `getLatestCommit`, - `getRecord` (proof CAR), `listRepos` (paginated), `getRepoStatus`. - - `com.atproto.server.describeServer` + `_health` — enough for crawlers. - - `com.atproto.identity.resolveHandle` (from task 03, mounted here). -- `internal/sync/broadcast.go` — fan-out: repo layer signals new seq; - broadcaster wakes subscriber outboxes; each reads sequentially from - `firehose_events` (DB-backed, so restarts and slow consumers are safe). -- Event retention: config `FIREHOSE_RETENTION` (default 72h), pruning job. -- Optional-but-cheap: `com.atproto.sync.requestCrawl` client helper to ask - a relay to crawl us (used in prod; log-only in dev). -- Tests: end-to-end within the package — write records via task 03 API, - connect a real WebSocket client, decode frames with indigo's - `events`/`repo` packages, assert ops + CAR blocks verify against the MST; - cursor replay from mid-stream; slow-consumer eviction. -- **Integration proof (the money test, may live in compose profile):** - run Jetstream container pointed at Tidepool - (`--ws-url ws://tidepool:PORT/xrpc/com.atproto.sync.subscribeRepos` per - Coves' jetstream service config), write a `social.coves.community.post`, - assert Jetstream emits the decoded JSON commit with correct collection, - repo DID, and record body. - -## Definition of done -- Jetstream consumes Tidepool without errors and re-emits our records. -- Cursor replay is gapless and ordered (test writes 100 records, connects - at cursor 50, sees exactly 51..100 then live events). -- `go test ./...` green. - -## References -- atproto sync spec (event stream framing): atproto.com/specs/event-stream - and /specs/sync. -- `~/Code/arroba/arroba/xrpc_sync.py` — the reference implementation of - exactly this surface, including subscribeRepos framing. -- indigo `events` package (frame encoding), Jetstream config in - `~/Code/coves/docker-compose.dev.yml`. diff --git a/tasks/05-materializer.md b/tasks/05-materializer.md deleted file mode 100644 index 756e859..0000000 --- a/tasks/05-materializer.md +++ /dev/null @@ -1,72 +0,0 @@ -# Task 05 — Materializer: AP → social.coves.* (~1200 LOC) - -## Goal -The translation heart: given a fetched/delivered AP object, produce the -correct `social.coves.*` record(s) in the correct repo(s), idempotently, -with valid strongRefs. Pure translation + orchestration; transport is -tasks 02/06, repos are task 03. - -## Deliverables -- `internal/materialize/actors.go` — - - Lemmy `Person` → mint DID (task 03) if unseen → - `social.coves.actor.profile` (rkey `self`): displayName from - name/preferredUsername, bio from summary (HTML→text) **plus appended - provenance line** "🌉 bridged from @user@instance by Tidepool", - avatar/banner: fetch image, store as blob (add a minimal blob store to - the repo layer if task 03 didn't: blobs table + `sync.getBlob`), set - createdAt from `published`. Respect consent_state — `nobridge` actors - are never materialized (their content dropped with a logged reason). - - Lemmy `Group` → mint community DID → `social.coves.community.profile` - (rkey `self`): name = preferredUsername, displayName = name, - description = summary, `createdBy` = bridge service DID, `hostedBy` = - bridge service DID, avatar/banner blobs. - - Profile refresh: re-materialize on AP `Update{Person|Group}` or when - profile_synced_at older than config TTL. -- `internal/materialize/posts.go` — Lemmy `Page` → - `social.coves.community.post` **written into the community's repo**, - fields per `~/Code/coves/internal/atproto/lexicon/social/coves/community/post.json`: - community = community DID, author = bridged author DID, title = name, - content from `source.content` (markdown) if present else HTML→markdown, - embed union: image attachment → `embed.images` (blob), external link - `attachment`/`url` → `embed.external` (fetch og-image optional, skip in - v1), `sensitive` → contentLabels/nsfw label, langs from `language`, - createdAt = published. Deterministic rkey (task 03 tid.go). -- `internal/materialize/comments.go` — Lemmy `Note` → - `social.coves.community.comment` **in the author's repo**: content - markdown, `reply.root` + `reply.parent` strongRefs resolved via - `ap_objects` (`inReplyTo` chain). **Missing-parent protocol**: if parent - ap_id unmapped, recursively fetch (task 02) and materialize ancestors - first (depth cap ~50, cycle guard); root = walk to the Page. If an - ancestor's author is nobridge/deleted, materialize a placeholder-free - skip: drop the comment subtree (log). -- `internal/materialize/updates.go` — `Update` → re-put record same rkey; - `Delete`/`Tombstone` → repo DeleteRecord + soft-delete mapping; - `Delete(Actor)` → tombstone all their records + mark consent deleted. -- `internal/materialize/html.go` — HTML→markdown for Lemmy HTML content - (they send both `content` HTML and usually `source` markdown; prefer - source). Strip scripts, cap length per lexicon maxGraphemes, generate - facets for links if cheap (else plain markdown text — check what Coves - post consumer expects; content is markdown per lexicon). -- Validate every produced record against the actual lexicon JSON schemas — - vendor `social/coves/*` lexicon files from Coves into `lexicons/` - (script to re-sync: `scripts/sync-lexicons.sh`) and validate with - `xeipuuv/gojsonschema` like Coves does, at materialization time in dev, - in tests always. -- Golden tests: task 02 fixtures → expected record JSON - (`internal/materialize/testdata/*.golden.json`); idempotency test - (materialize twice → identical at-uri/cid, single mapping row); - missing-parent chain test with a 3-deep fixture thread. - -## Definition of done -- All produced records validate against vendored Coves lexicons. -- Emission ordering guarantee: community profile + author profile commits - always precede the content commit referencing them (assert via - firehose_events seq in a test). -- `go test ./...` green. - -## References -- Coves lexicons + consumers (validation rules to satisfy): - `~/Code/coves/internal/atproto/jetstream/post_consumer.go`, - `comment_consumer.go`. -- `~/Code/granary/granary/as2.py` + `bluesky.py` — field-mapping prior art. -- `~/Code/coves/tests/lexicon-test-data/` — valid/invalid record fixtures. diff --git a/tasks/06-ingestion.md b/tasks/06-ingestion.md deleted file mode 100644 index 0b97c36..0000000 --- a/tasks/06-ingestion.md +++ /dev/null @@ -1,63 +0,0 @@ -# Task 06 — Ingestion pipeline: inbox, follows, backfill, consent (~1200 LOC) - -## Goal -Wire the fediverse to the materializer: subscribe to communities, receive -pushed activities, verify them, unwrap FEP-1b12 Announces, keep ordering -and idempotency, and backfill history. After this task the bridge is -functionally complete end-to-end (minus votes). - -## Deliverables -- `internal/ingest/inbox.go` — POST `/inbox` (shared inbox) + per-actor - inboxes: verify HTTP signature (task 02 `Verify`, actor key fetch with - cache + re-fetch-on-fail once), check `inbox_events` dedupe by activity - id, enqueue. Reject unsigned/expired (>1h skew). Serve the service actor - document at `/actor` + WebFinger for it at `/.well-known/webfinger` - (needed so Lemmy accepts our Follow), and `/.well-known/nodeinfo` - (minimal; Lemmy checks software names for allowlists). -- `internal/ingest/queue.go` — durable work queue on postgres - (`inbox_events` as the queue; worker pool, per-community serial ordering - key so a community's events apply in order, retry w/ backoff, poison - → error column + skip). No external queue dependency. -- `internal/ingest/handler.go` — activity dispatch: - - `Announce{Create|Update{Page|Note}}` (FEP-1b12 group fan-out — the - normal Lemmy path) → materializer. - - Bare `Create/Update/Delete` from user actors (some arrive direct) → - verify the object belongs to a followed community, then materialize. - - `Announce{Like|Dislike}` → hand to vote aggregator (task 07 stub: - define the interface now, no-op impl until 07). - - `Accept{Follow}` → mark community follow_state accepted. - - `Undo`, `Delete(Actor)`, `Update{Group|Person}` → materializer updates. - - **Echo suppression**: drop any activity whose object's ap_id maps to a - record the bridge itself created (future-proofing for write-side; cheap - check against ap_objects origin flag — add `origin` column - (fediverse|bridge) to ap_objects if not present). -- `internal/ingest/follow.go` — community subscription lifecycle: admin/API - trigger `POST /admin/communities {"community":"!tech@lemmy.world"}` → - WebFinger resolve → fetch Group → materialize community → signed - `Follow` from service actor → await Accept; `DELETE` → `Undo{Follow}`. - Simple bearer-token admin auth (`ADMIN_TOKEN` env). -- `internal/ingest/backfill.go` — on follow-accept (and on demand): page - the Group outbox (task 02 FetchCollection), materialize newest→oldest up - to `BACKFILL_MAX_POSTS` (default 100) + each post's replies collection - if advertised; rate-limited, resumable via last_backfill_at. -- `internal/ingest/consent.go` — scan actor summary/tags for - `#nobridge`/`#nobot` on first sight and on profile Update → set - consent_state, tombstone existing content if switching to nobridge; - `Delete(Actor)` handling already in materializer — wire it. Document the - policy in README (mirrors Bridgy Fed norms). -- Tests: signature verify against fixtures (valid, bad digest, expired, - key rotation re-fetch); dedupe; ordering (two comments same community - arrive out of order → both land, parent-fetch path exercised); follow - state machine; echo suppression; consent transitions. Use a fake Lemmy - HTTP server (httptest) serving task 02 fixtures. - -## Definition of done -- With the fake Lemmy server: subscribe → Accept → Announce{Create{Page}} - → post record visible via task 04 firehose; backfill produces mapped - history; nobridge author's comment dropped with logged reason. -- `go test ./...` green. - -## References -- `~/Code/bridgy-fed/activitypub.py` (inbox verification, Announce - unwrapping, quirks), `ids.py` (echo/copies logic), FEP-1b12. -- Lemmy federation docs for Follow/Accept + outbox shapes. diff --git a/tasks/07-vote-aggregates.md b/tasks/07-vote-aggregates.md deleted file mode 100644 index ed6afdc..0000000 --- a/tasks/07-vote-aggregates.md +++ /dev/null @@ -1,47 +0,0 @@ -# Task 07 — Vote aggregation side channel (~800 LOC) - -## Goal -The deliberately-unmaterialized data path: Lemmy Like/Dislike activities -become bridge-side aggregate counts served over one small versioned XRPC. -No PDS records, no firehose traffic. (Materialization principle: nothing -ever strongRefs a vote.) - -## Deliverables -- Migration `vote_aggregates`: subject_ap_id, subject_at_uri, upvotes, - downvotes, updated_at, PK subject_ap_id; plus `vote_events` dedupe table - (activity_id unique, voter_ap_id, subject_ap_id, direction, undone bool) - — needed because Lemmy sends `Undo{Like}` and voters flip votes; count - distinct voters' latest state, don't blindly increment. -- `internal/votes/aggregator.go` — implements the interface stubbed in - task 06: handle `Like`, `Dislike`, `Undo{Like|Dislike}`; recompute or - incrementally maintain counts per subject; only for subjects present in - `ap_objects` (votes on unbridged content: count into a pending bucket or - drop — DROP in v1, log at debug). -- Backfill hook: Lemmy `Page`/`Note` objects don't carry counts via AP - reliably, but the Group outbox Announces historical Likes sparsely. - Accept that backfilled posts start near zero; document the limitation. - (Optional if trivial: seed from Lemmy's public API `counts` field during - task 06 backfill via a `SeedCounts(subject, up, down)` method — gate - behind config `SEED_COUNTS_FROM_API`, default on.) -- `internal/votes/xrpc.go` — the sanctioned side channel: - `GET /xrpc/social.coves.bridge.getVoteAggregates?uris=at://...,at://...` - (≤100 uris) → `{aggregates: [{uri, upvotes, downvotes, updatedAt}]}`. - Write the lexicon JSON for it under `lexicons/social/coves/bridge/` - (new nsid — this is Tidepool's published contract; versioned via the - nsid, breaking changes mean a new name). Public read, cache headers, - rate limit by IP. -- Tests: vote → count; flip vote → net change correct; undo → decrement; - duplicate activity id → no-op; 100-uri batch query; unknown uri → omitted - from response (not an error). - -## Definition of done -- Fake-Lemmy E2E: Announce{Like} ×3 + Dislike ×1 + Undo ×1 → XRPC returns - {up:2, down:1} (or per fixture design). -- Lexicon file validates and is documented in README as the AppView - integration point. -- `go test ./...` green. - -## References -- Coves vote semantics: `~/Code/coves/internal/atproto/lexicon/social/coves/feed/vote.json`, - `vote_consumer.go` (what natives do; we deliberately bypass it). -- PLAN.md decision 7. diff --git a/tasks/08-e2e-harness.md b/tasks/08-e2e-harness.md deleted file mode 100644 index ac4ef84..0000000 --- a/tasks/08-e2e-harness.md +++ /dev/null @@ -1,57 +0,0 @@ -# Task 08 — E2E harness: real Lemmy → Tidepool → Jetstream (~1000 LOC) - -## Goal -Prove the whole read path against real infrastructure, Coves-style ("E2E -tests must test REAL infrastructure - not mocks"): a real Lemmy container -federating with Tidepool, records flowing out our firehose, decoded by a -real Jetstream, validating against real Coves lexicons. - -## Deliverables -- `docker-compose.e2e.yml` (or profile in dev compose): postgres, PLC - directory, Tidepool (built from Dockerfile — write the Dockerfile, - multi-stage, distroless-ish), Jetstream pointed at Tidepool, **Lemmy** - (lemmy + lemmy postgres + pictrs; use dessalines/lemmy image with - federation enabled, allowlist tidepool host; both services on one - compose network with hostnames — Lemmy requires HTTPS in prod mode but - supports a debug/local federation mode; if TLS is unavoidable, add a - caddy with internal CA like Coves' Caddyfile.dev pattern). Getting - Lemmy↔bridge federation working in compose is the hard 60% of this task - — budget accordingly; crib from Lemmy's own `docker/federation` compose - in github.com/LemmyNet/lemmy which runs multi-instance federation - locally over HTTP. -- `tests/e2e/helpers.go` — Lemmy API client for test setup only (create - site/admin, community, post, comment, vote via Lemmy's HTTP API); - Jetstream WebSocket listener with per-collection matchers + timeouts; - Tidepool admin client (subscribe community). -- `tests/e2e/bridge_test.go` — the scenarios: - 1. Subscribe `!testing@lemmy` → community.profile appears on firehose, - validates against Coves lexicon. - 2. Lemmy user posts → actor.profile then community.post (in the - community DID's repo, author = user DID) appear, in that order. - 3. Comment + nested reply → comment records with correct root/parent - strongRefs (resolve them: parent uri/cid match earlier events). - 4. Edit post → update event; delete comment → delete op on firehose. - 5. Votes → getVoteAggregates XRPC reflects them; no vote records on - the firehose. - 6. Backfill: pre-existing posts appear after subscribe. - 7. Idempotency/restart: restart Tidepool container mid-test, replay - causes no duplicate rkeys and Jetstream cursor resume works. -- Lexicon conformance suite: every record type Tidepool emits validated - against `~/Code/coves` lexicons synced by `scripts/sync-lexicons.sh` - (already vendored in task 05 — add a CI check that they're in sync). -- Makefile: `make e2e` (compose up, wait-for-healthy, run - `go test ./tests/e2e/... -tags e2e`, compose down), `make e2e-logs`. -- CI workflow file (GitHub Actions) running unit tests always, e2e on - demand/label (it's heavy). -- README "Running the stack" section + architecture diagram refresh. - -## Definition of done -- `make e2e` passes from a clean checkout on this machine. -- A written FOLLOWUPS.md capturing everything discovered but deferred - (PieFed quirks, relay requestCrawl in prod, key claiming, write-side). - -## References -- `~/Code/coves/docker-compose.dev.yml`, `Caddyfile.dev`, `Makefile`, - `tests/integration/helpers.go` (harness conventions). -- LemmyNet/lemmy `docker/federation/` (local federation compose), - `api_tests/` (their own federation test patterns). diff --git a/tasks/09-e2e-relay.md b/tasks/09-e2e-relay.md deleted file mode 100644 index dfb778c..0000000 --- a/tasks/09-e2e-relay.md +++ /dev/null @@ -1,72 +0,0 @@ -# Task 09 — E2E relay: full pipeline Lemmy → Tidepool → relay → Jetstream (~600 LOC) - -## Goal -Put a real atproto relay (indigo BigSky) between the bridge and Jetstream -in the e2e stack, so every scenario's records transit the strictest -consumer that exists: DID resolution against the local PLC, per-commit -signature verification, and initial `getRepo` crawls. Closes the FOLLOWUPS -item "relay `requestCrawl` has never been exercised against a real relay" -and turns the sync surface's correctness claims (CAR exports, proofs, -cursor semantics) into things a hostile consumer verifies on every run. - -## Locked decisions -- **One pipeline, through the relay.** Jetstream consumes from the relay, - not the bridge, so every existing and future scenario exercises relay - validation for free (locked requirement: every state has a full - Lemmy → PDS record → firehose pipeline where applicable). Keep the - direct bridge→Jetstream wiring available behind a compose profile for - debugging, not as the tested path. -- **The bridge's own `RequestCrawlAll` must send the real request.** The - e2e stack runs `ENVIRONMENT=development`, where `RELAY_HOSTS` is - log-only. Add a narrowly-scoped dev override (mirror the - `ALLOW_PRIVATE_FETCH` pattern: dev-only, refused/ignored in production - where sending is already the behavior) so the harness drives the real - code path. The relay's admin API (`POST /admin/pds/requestCrawl`, - Coves' pattern) is an acceptable suite *fallback/bootstrap*, but the - DoD is bridge-originated crawl. -- Local-only invariant holds: loopback-only host binds, nothing ever - contacts a public relay or plc.directory. - -## Deliverables -- `docker-compose.e2e.yml`: relay service — the pinned bigsky image Coves - uses (`ghcr.io/bluesky-social/indigo:bigsky-0a2d4173e6e89e49b448f6bb0a6e1ab58d12b385`, - bump deliberately if needed), its own postgres (or a second database in - an existing one — decide and comment), `BGS_CRAWL_INSECURE_WS=true` - (ws:// upstream) *[verified outcome: this env var does not exist in the - pinned image — `--crawl-insecure-ws` is arg-only and is passed as a - command argument instead; see FOLLOWUPS.md]*, admin key, healthcheck, - loopback-only host port. - **Identity resolution must point at the compose `plc` service** — find - bigsky's actual PLC-host flag/env in the indigo source (verify, don't - guess) or the relay will try the public directory and fail closed. -- Jetstream re-pointed at the relay's firehose; existing scenarios pass - unchanged through the longer pipeline (expect timing budgets to need - loosening — crawl + validation adds latency). -- Bridge sends `requestCrawl` to the relay on startup via the real - `RequestCrawlAll` path (`RELAY_HOSTS=` + the new dev override). -- `tests/e2e`: relay-specific assertions — the bridge host appears in the - relay's crawled-PDS state; repos are listed/crawled; the restart/replay - scenario still proves cursor resume *through the relay*; tombstoned - (`active:false`) repo status observed through the relay where its API - exposes it. -- **Known rock, investigate and document:** bridged handles - (`alice.lemmy.tidepool`) do not resolve in compose DNS — determine - whether bigsky's handle verification failure is non-fatal (likely: - marks handle invalid, keeps repo) and document the posture; add - network aliases only if actually required. -- README stack diagram + runbook update; FOLLOWUPS updates for anything - discovered and deferred. - -## Definition of done -- `make e2e` green from a clean checkout with all scenarios transiting - Lemmy → tidepool → relay → Jetstream. -- Bridge-originated `requestCrawl` observed in relay state/logs (asserted, - not eyeballed). -- Unit suite still green; nothing contacts public infrastructure. - -## References -- `~/Code/coves/docker-compose.dev.yml` relay stanza (lines ~248–296): - image pin, `BGS_CRAWL_INSECURE_WS`, admin requestCrawl curl. -- indigo source (`cmd/bigsky`) for flags/env: PLC host, admin API routes. -- `internal/sync/crawl.go` `RequestCrawlAll`; `cmd/tidepool/main.go` - dev-mode skip. diff --git a/tasks/10-e2e-scenarios.md b/tasks/10-e2e-scenarios.md deleted file mode 100644 index 02729b4..0000000 --- a/tasks/10-e2e-scenarios.md +++ /dev/null @@ -1,55 +0,0 @@ -# Task 10 — E2E scenario completion: every state, full pipeline (~800 LOC) - -## Goal -Close the FOLLOWUPS "scenario ideas not yet covered" list so every -user-visible state transition the bridge implements is proven end-to-end -(Lemmy → tidepool → relay → Jetstream, task 09's pipeline), with lexicon -validation on the wire. After this task, "the e2e suite passes" means -every materialization arm and lifecycle flow has crossed real -infrastructure. - -## Deliverables (each a scenario in `tests/e2e`, using existing helpers) -1. **Image post**: upload via pictrs → post with image → blob fetched and - stored → `embed.images` (+ nsfw label shape if expressible via Lemmy) - crosses the wire and lexicon-validates. Closes the materializer test - gap ("embed.images never appears on the wire"). -2. **Consent — `#nobridge`**: Lemmy user with `#nobridge` in bio posts → - nothing materializes (and no DID mint); remove the marker → bridging - resumes on profile refresh. Then the reverse: bridged actor adds the - marker → scrub delete-commits observed on the firehose, repo stays - active (reversible posture). -3. **`Delete(Actor)`**: Lemmy account deletion → all their records - scrubbed (delete ops on firehose), repo terminally tombstoned — - `getRepoStatus`/`listRepos` report `active:false`, handle stops - resolving, content endpoints refuse. Assert through the relay where - its API exposes repo state. -4. **Unsubscribe**: `DELETE /admin/communities` → `Undo{Follow}` → new - Lemmy posts in that community produce NO bridge output (bounded - negative assertion), while a still-subscribed community keeps flowing - (positive control in the same window). -5. **Community profile update**: rename/description change in Lemmy → - `community.profile` update event with rkey `self`. -6. **Vote concurrency hammer**: many voters, one post, delivered - concurrently (parallel inbox deliveries / multiple Lemmy users voting - in a burst) → final `getVoteAggregates` exactly correct. First real - exercise of the aggregate-row locking claim beyond unit level. -7. **Suite-end sweep**: after all scenarios, replay the entire firehose - from cursor 0 with an unfiltered listener and assert no event ever - carried a collection outside the four emitted ones and every - create/update lexicon-validates — closes the "events emitted while no - unfiltered listener was subscribed" gap. -- Stretch (skip if it needs its own compose profile): low - `MINT_RATE_PER_MINUTE` variant driving the mint gate end-to-end. - -## Definition of done -- `make e2e` green, all scenarios through the full relay pipeline. -- Every new record shape that crosses the wire is lexicon-validated by - the consuming listener (vetEvent path), not by unit fixtures. -- FOLLOWUPS updated: covered items removed, new discoveries added. - -## References -- `tests/e2e/bridge_test.go`, `helpers.go` (vetEvent, drain, listener - conventions — extend, don't fork). -- FOLLOWUPS.md "E2E harness itself" + materializer/vote test-gap items. -- Lemmy API: pictrs upload, account deletion, community edit endpoints - (0.19.x, `/api/v3`). diff --git a/tasks/11-hardening.md b/tasks/11-hardening.md deleted file mode 100644 index 570f6d6..0000000 --- a/tasks/11-hardening.md +++ /dev/null @@ -1,70 +0,0 @@ -# Task 11 — Hardening: the pre-internet-facing FOLLOWUPS (~1000 LOC) - -## Goal -Work the correctness/security backlog in FOLLOWUPS.md so the bridge can -face the open internet: admission control on every public surface, the -missing firehose account signal, ordering/tombstone gaps, unbounded table -growth, and the small housekeeping items. Verified under the task 09/10 -harness — the strictest pipeline we have. - -## Deliverables (FOLLOWUPS is the checklist; this is the triage) -Security / admission control: -- **Per-signer AND per-IP rate limit on `/inbox`** (top item: - queue-flood DoS via many self-signed identities). Token-bucket, - fail-closed cap, mirroring the votes XRPC limiter's discipline. - MUST specifically cover the `tombstonedSelfDelete` branch (task 10's - deleted-actor acceptance): it is an UNAUTHENTICATED POST that costs - the bridge an outbound confirmation fetch and — when the claimed - origin answers 410, which any cheap attacker-run endpoint can — - TWO durable writes (`inbox_events` + `ap_tombstones`). Consider a - dedicated per-IP cap on that branch, tighter than the general inbox - limit. -- **Connection cap + per-IP rate limit on the public sync surface** - (subscribeRepos, getRepo and friends). -- Seeded-count upper sanity cap (a hostile origin API can't inject - absurd baselines). - -Protocol correctness: -- **`#account{active:false}` firehose frame** on `Delete(Actor)` / - consent revocation, so subscribers purge instead of relying on scrub - delete-commits. Assert it in an e2e scenario (task 10's Delete(Actor) - scenario gains the frame check). -- **Delete-before-Create**: reconcile the README claim ("remembered via - `ap_tombstones`") against the FOLLOWUPS gap ("a Delete arriving before - its object was ever materialized leaves nothing to tombstone") — - verify which is true, close the gap, kill the false doc either way. -- **Automatic Follow re-send** when a subscription stays `pending` past - a threshold (Lemmy first-contact Accept race — currently only the test - harness retries; production operators shouldn't have to). -- Actor-delete / consent revocation scrubs that actor's `vote_events` - rows (consistency with scrub posture elsewhere). - -Unbounded growth / housekeeping: -- Pruners for `ap_tombstones` and superseded/undone `vote_events` rows - (mirror `FIREHOSE_RETENTION` treatment); batch the `PruneEvents` - DELETE while there. -- `MAX_BLOB_BYTES` above 5 MiB: wire the AP client's response cap to the - config value instead of silently clamping. -- Transient media-fetch failure on profile refresh carries forward - existing blobs instead of dropping them; `DeleteActor` scrubs blobs - stored under community DIDs. -- `commitRecord` PutRecord→PutMapping in one tx. -- Delete dead `ingest.NewNoopVotes`; rename - `service_keys.private_key_pem` (it holds sealed ciphertext for the - plc-rotation row) via migration. -- Strict-validation failure metric (production logs-and-writes today — - make it observable; strict-first rollout stays deferred). - -## Definition of done -- Full unit suite + `make e2e` green. -- Every FOLLOWUPS item this task closes is deleted from FOLLOWUPS.md; - anything triaged out is annotated with why. -- New public-surface limits have tests proving both enforcement and - non-interference with legitimate load (the e2e suite itself is the - canary — it must pass under the new limits). - -## References -- FOLLOWUPS.md (Ingestion, Sync surface, Votes, Materializer, - Storage/housekeeping sections). -- `internal/votes/xrpc.go` limiter (pattern to reuse), LOOP_STATE task - 06/07 notes (queue fencing, outcome contract — do not violate). diff --git a/tasks/12-perf-scale.md b/tasks/12-perf-scale.md deleted file mode 100644 index 0b43612..0000000 --- a/tasks/12-perf-scale.md +++ /dev/null @@ -1,48 +0,0 @@ -# Task 12 — Perf & scale: big-community readiness (~800 LOC) - -## Goal -Remove the known scaling cliffs before any big community hits them — and -as the explicit prerequisite for ever revisiting votes-as-records -(FOLLOWUPS "Design revisits": write amplification is only viable after -this task). Every change here is behavior-preserving; the e2e suite and -golden-value tests (deterministic TIDs, at-uris) must not move. - -## Deliverables -- **Per-DID MST tree cache**: `PutRecord` is currently O(repo size) with - one SELECT per node, full-tree, per commit. Cache the tree per DID with - invalidation tied to the commit path (single-writer discipline + the - global advisory lock make this tractable). Benchmark before/after with - a realistic big-repo fixture (thousands of records) and record numbers - in the commit message. -- **`getRepo` memory + reachable-set**: stream the CAR instead of - buffering it whole; export the reachable set from the current commit - rather than every historical block. Coordinate with GetRecord's - read-consistency dependence on append-only `blocks` (LOOP_STATE task - 03/04 notes) — reachable-set-only export must not break proof reads. -- **`blocks` GC**: prune blocks unreachable from any live commit, as an - explicit background sweep with a retention guard. This is the - load-bearing append-only table — design the invariant first (what do - GetRecord/getRepo/replay need?), write it down in the code, then - implement to it. If a safe GC needs the sync `since`/diff-export work, - say so and defer that half explicitly. -- **`ClaimNext` O(N) scan** when one community's queue backs up behind a - failing event: bound the scan (indexed skip / per-key cursor) without - breaking the per-community serial-ordering guarantee or the fencing - contract. -- Cheap wins while in the area: `getRepo` `since` param (diff export) if - the reachable-set work makes it nearly free — otherwise leave the - documented full-CAR fallback. - -## Definition of done -- Full unit suite + `make e2e` green; golden TID/at-uri tests untouched. -- Before/after benchmarks for PutRecord (big repo) and getRepo (memory) - recorded in the commit message. -- FOLLOWUPS updated (items closed/annotated); the votes-as-records - design-revisit note updated to reflect which prerequisites now hold. - -## References -- FOLLOWUPS.md Sync surface + Storage sections; LOOP_STATE task 03/04 - notes (commit serialization, advisory locks, why blocks is - append-only). -- `internal/repo/` (MST load path, ExportCAR), `internal/store` - (ClaimNext). -- 2.51.2