diff --git a/LOOP_STATE.md b/LOOP_STATE.md index 0aa8cd0..944ad06 100644 --- a/LOOP_STATE.md +++ b/LOOP_STATE.md @@ -11,7 +11,7 @@ update this file → schedule next. Stop the loop when every task is `done`. | 3 | 03-identity-repos | done | (see git log) | 8 reviewers (5 Claude + codex/gemini/glm); 16 fixes (genesis race, seq ordering, MST-corruption-as-NotFound, KeyUse deletes, TID micro-fill) + 7 new tests | | 4 | 04-sync-firehose | done | (see git log) | 7/8 reviewers (glm watchdog-killed); fixes: ping starvation, prune-mid-replay OutdatedCursor, broadcaster closed-channel, SeqBounds dirty-read, pruner fail-closed; consent-on-firehose deferred to 06 | | 5 | 05-materializer | done | (see git log) | 8 reviewers; fixes: id-authority binding, Note-root panic, create-after-delete, nobridge scrub, embedded-actor trust, byte caps, uri scheme, Group-type check + 9 regression tests | -| 6 | 06-ingestion | pending | | | +| 6 | 06-ingestion | done | (see git log) | 8 reviewers (5 Claude + codex/gemini; glm wandered, no JSON); fixes: announced-Delete/Undo scoped to announcer authority (+actor-delete only self), bare Update{Person/Group} no-mint gate, announce content community-authority check, Undo{Delete} restore compensation, handleAccept pending-only, queue lease fencing token + shutdown-cancel handling + processed/poisoned exclusivity, backfillReplies tombstone check, truncation leaves resumable, activityID rand-fail propagates + 14 regression tests | | 7 | 07-vote-aggregates | pending | | | | 8 | 08-e2e-harness | pending | | | @@ -265,3 +265,80 @@ and deferred TODOs here) later). Test gaps still open: embed.images arm + nsfw label shapes are never lexicon-validated (only external embed is) — add before trusting "all records validate" for image posts. + +### From task 06 (ingestion — tasks 07/08 consume this) +- internal/ingest is the inbox→queue→dispatch→materializer pipeline. Entry: + POST /inbox (+/actor/inbox alias) verify-sig → authority-bind signer → + dedupe by activity id → Enqueue. GET /actor, /.well-known/webfinger + (service actor only), /.well-known/nodeinfo + /nodeinfo/2.0 + (software.name "tidepool"). Admin (bearer ADMIN_TOKEN, constant-time): + POST/DELETE/GET /admin/communities, POST /admin/communities/backfill. +- TASK 07 SEAM: implement ingest.VoteAggregator — ApplyVote/RetractVote(ctx, + vote *ap.Object, communityIRI string). Wired as NewNoopVotes in main.go; + swap it. Announce{Like|Dislike}→ApplyVote, Announce{Undo{Like|Dislike}}→ + RetractVote; communityIRI is "" for bare (non-announced) votes. RetractVote + MUST treat nil/bare-IRI vote.Object/vote.Actor as no-ops (don't error/ + retry). Aggregate COUNT SEEDING is NOT expressible via per-vote ApplyVote — + if task 07 seeds historical counts from Lemmy's API, add a separate + SeedAggregates method + Backfill wiring (additive, not breaking). +- QUEUE (store.InboxEvents.Enqueue/ClaimNext/Release/MarkPoisoned/ + MarkProcessed): durable pg work queue, per-community serial ordering key, + FOR UPDATE SKIP LOCKED + NOT-EXISTS older-sibling gate. Outcome contract + (identical on Handler.Process, Processor, queue dispatch): nil/IsSkip → + processed; IsValidation → poison; else retry w/ exp backoff (cap 1h); + attempt-cap → poison. CRITICAL for task 07: the queue is now FENCED — + ClaimNext returns claimed_until as a fencing token; Release/MarkPoisoned/ + MarkProcessed take that token + return (applied bool). A stale worker's + write is a silent no-op (applied=false). This exists SO task 07's vote + counter arithmetic isn't double-applied — do NOT bypass repo.Manager-style + or write vote state outside the queue's outcome path without the same + fencing discipline. HEAD-OF-LINE BLOCKING is intentional: a retrying event + blocks its whole ordering key (community) until success or poison. +- Shutdown: processNext no longer classifies ctx-cancellation as a failure + (lease lapses → redeliver); outcome writes use context.WithoutCancel so + completed work is recorded during shutdown. Residual: a shutdown- + interrupted attempt still consumes its ClaimNext attempt increment (not + decremented; harmful poison/error-record behavior is gone). +- AUTHORIZATION MODEL (task 07/08 must preserve): announced (announcer!="") + Delete/Undo require ap.SameAuthority(target, announcer) AND for actor + targets require target==announcer (a community may delete only itself, not + a co-hosted other actor). Announced CONTENT requires communityIRIFrom(obj) + == announcer (no bridging a non-subscribed same-host community). Bare + (announcer=="") Update{Person|Group} is REFRESH-ONLY: never mints/bridges + an unknown actor (was an SSRF+mint-budget vector); embedded actor doc is + trusted only when signer==actor, else forced re-fetch. Bare Delete still + uses SameAuthority(target, signer) (host-granularity — instance is the + trust unit; documented). Undo{Delete} restores mapping BEFORE re- + materialize (commitRecord won't resurrect a soft-deleted mapping) and + COMPENSATES (re-soft-delete + re-tombstone) if HandleUpdate skips/invalid. +- Consent: nobridge scrubs but repo stays ACTIVE (reversible, firehose- + visible delete commits); only `deleted` deactivates the sync read surface. + Still DEFERRED (task 07/08): firehose #account{active:false} frame on + consent revocation (subscribers currently rely on scrub delete-commits); + ap_tombstones grows unbounded (no pruner — mirror FIREHOSE_RETENTION); + NO per-signer/per-IP rate limit on /inbox (queue-flood DoS via many self- + signed identities — coarse admission limit is the hardening item); + ClaimNext does an O(N) row scan when a community's queue backs up behind a + failing event (per-key serialization cost; revisit at scale). Nudge now + cascades on successful claim. +- Backfill: outbox newest→oldest to BACKFILL_MAX_POSTS (100), replies walk + when advertised, tombstone-checked on all three paths now. RESUMABLE-BY- + REDO: truncation (ErrCollectionTruncated) or failures>0 leave + last_backfill_at UNSET so an un-forced re-trigger re-walks; deterministic + rkeys make redo idempotent. main.go drains backfill.Wait() on shutdown + (bounded); TriggerAsync runs on the root ctx (observes shutdown). +- New env: ADMIN_TOKEN (required in prod, dev default dev-admin-token), + BACKFILL_MAX_POSTS (100), MINT_RATE_PER_MINUTE (60), MINT_BURST (120), + INGEST_WORKERS (4). Migration 009: queue columns on inbox_events + (payload, actor_id, ordering_key, attempts, next_attempt_at, claimed_until, + failed_at) + ap_tombstones table. store grew Tombstones (Record/Exists/ + Remove), APObjects.Restore, and the fenced InboxEvents queue API. +- ap.Client.ResolveKeyDetailed exported + Object.Replies field added. Verify + fresh-key refetch is now cache-gated (fires only when the failing key came + from cache) — bounds forgery amplification. +- Deferred TEST gaps (task 08 harness): concurrent-worker queue stress + (SKIP-LOCKED ordering only tested sequentially); mint-gate "retry via queue + backoff" only asserted at unit level, not end-to-end; webfinger/nodeinfo + vs what real Lemmy actually queries (needs live Lemmy). activityID rand- + fail path is guarded but unit-untestable (Go 1.24+ crypto/rand failure is + a fatal crash, not a returnable error). diff --git a/README.md b/README.md index cf4aecf..a597b35 100644 --- a/README.md +++ b/README.md @@ -50,6 +50,70 @@ production**: | `ALLOW_PRIVATE_FETCH` | off | dev-only: disables the SSRF egress guard (AP fetches **and** PLC directory requests) so localhost targets work | | `FIREHOSE_RETENTION` | `72h` | how long `firehose_events` rows are kept for `subscribeRepos` cursor replay (Go duration; a background pruner trims older events hourly) | | `RELAY_HOSTS` | *(optional)* | comma-separated relays to send `com.atproto.sync.requestCrawl` to on startup; in development the request is logged, never sent | +| `ADMIN_TOKEN` | `dev-admin-token` | bearer token protecting the `/admin` API | +| `BACKFILL_MAX_POSTS` | `100` | posts materialized per community backfill run | +| `MINT_RATE_PER_MINUTE` / `MINT_BURST` | `60` / `120` | rate gate on inbound DID minting (PLC registrations are forever; unseen authors in delivered content trigger mints) | +| `INGEST_WORKERS` | `4` | inbox queue worker-pool size | + +## Subscribing to communities (admin API) + +Community subscriptions are operator-driven, over bearer-token-protected +endpoints (`Authorization: Bearer $ADMIN_TOKEN`): + +```sh +# follow: WebFinger → fetch Group → materialize community → signed Follow +curl -X POST localhost:8091/admin/communities \ + -H "Authorization: Bearer dev-admin-token" \ + -d '{"community":"!technology@lemmy.world"}' + +# state is `pending` until the community's Accept arrives at /inbox, which +# flips it to `accepted` and triggers an outbox backfill automatically. + +curl localhost:8091/admin/communities \ + -H "Authorization: Bearer dev-admin-token" # list +curl -X POST localhost:8091/admin/communities/backfill \ + -H "Authorization: Bearer dev-admin-token" \ + -d '{"community":"!technology@lemmy.world"}' # on-demand backfill +curl -X DELETE localhost:8091/admin/communities \ + -H "Authorization: Bearer dev-admin-token" \ + -d '{"community":"!technology@lemmy.world"}' # Undo{Follow} +``` + +The bridge's AP face lives next to the inbox: the service actor document at +`/actor`, WebFinger for it at `/.well-known/webfinger`, and a minimal +nodeinfo 2.0 (`software.name: "tidepool"` — what Lemmy instance allowlists +match against). + +## Consent policy (#nobridge / #nobot) + +Tidepool mirrors the [Bridgy Fed](https://fed.brid.gy/docs#opt-out) opt-out +norms, enforced fail-closed in the materializer and wired to live AP +activity by the ingestion layer: + +- An actor whose profile summary or hashtags carry **`#nobridge`** or + **`#nobot`** is never bridged: no DID is minted, and every post or comment + they author is dropped with the reason logged. +- If a **previously bridged** actor adds the marker (seen on a profile + `Update` or any profile re-fetch), every record they authored is deleted + from the bridged repos and new materialization stops. This state is + **reversible**: removing the marker upstream restores bridging on the next + profile refresh. +- **`Delete(Actor)`** (account deletion upstream) scrubs all their records + and tombstones the bridged repo **terminally**; the sync surface reports + the repo `active: false`. +- Object-level `Delete`s tombstone the mapped record; a `Delete` arriving + before its object was ever seen is remembered (`ap_tombstones`), so an + out-of-order or re-delivered `Create` cannot resurrect deleted content. + `Undo{Delete}` restores by re-fetching the object from its origin. +- Bridged profiles are visibly labeled: bios/descriptions end with a + "bridged from … by Tidepool" provenance line, and community profiles set + `hostedBy` to the bridge's service DID. + +Deliveries are accepted only over valid draft-cavage HTTP signatures +(rsa-sha256, 1h date-skew window, digest required); content is accepted only +from communities the operator subscribed to, and embedded objects are +re-fetched from their origin instance whenever the delivering signer lacks +authority over the object's id. ## Sync surface (what relays and Jetstream consume) diff --git a/cmd/tidepool/main.go b/cmd/tidepool/main.go index 08776c1..20b5720 100644 --- a/cmd/tidepool/main.go +++ b/cmd/tidepool/main.go @@ -21,6 +21,8 @@ import ( "tidepool/internal/config" "tidepool/internal/db" "tidepool/internal/identity" + "tidepool/internal/ingest" + "tidepool/internal/materialize" "tidepool/internal/repo" "tidepool/internal/store" tidepoolsync "tidepool/internal/sync" @@ -134,8 +136,142 @@ func run(logger *slog.Logger) error { } } - // Later tasks register here: AP inbox + WebFinger (02/06), - // vote aggregates XRPC (07). + // The ingestion pipeline (task 06): AP inbox + signature verification, + // the durable work queue, activity dispatch into the materializer, the + // community Follow lifecycle, outbox backfill, and consent enforcement. + serviceKeys := store.NewServiceKeys(database) + serviceActor, err := ap.LoadOrCreateServiceActor(ctx, serviceKeys, cfg.BridgeHostname) + if err != nil { + return err + } + apClient := ap.NewClient(ap.ClientOptions{ + UserAgent: cfg.UserAgent, + Signer: serviceActor.Signer(), + AllowPrivateAddresses: cfg.AllowPrivateAddresses, + }) + + rotationKey, err := identity.LoadOrCreateRotationKey(ctx, serviceKeys, custodian) + if err != nil { + return err + } + minter, err := identity.NewMinter(identity.MinterOptions{ + PLCDirectoryURL: cfg.PLCDirectoryURL, + BridgeHostname: cfg.BridgeHostname, + RotationKey: rotationKey, + Custodian: custodian, + Actors: actors, + HTTPClient: ap.NewGuardedHTTPClient(cfg.AllowPrivateAddresses, 30*time.Second), + UserAgent: cfg.UserAgent, + Logger: logger, + }) + if err != nil { + return err + } + // Inbound AP activity can trigger DID minting (unseen authors), so the + // materializer's minter goes through the rate gate. + mintGate, err := ingest.NewMintGate(minter, cfg.MintRatePerMinute, cfg.MintBurst, logger) + if err != nil { + return err + } + + // The bridge's own DID (community.profile createdBy/hostedBy). An + // operator may pre-provision one; otherwise the bridge identifies as + // did:web on its own hostname. + serviceDID := cfg.BridgeServiceDID + if serviceDID == "" { + serviceDID = "did:web:" + cfg.BridgeHostname + logger.Info("BRIDGE_SERVICE_DID not set, deriving from hostname", "did", serviceDID) + } + + objects := store.NewAPObjects(database) + communities := store.NewCommunities(database) + tombstones := store.NewTombstones(database) + inboxEvents := store.NewInboxEvents(database) + + materializer, err := materialize.New(materialize.Options{ + Fetcher: apClient, + Objects: objects, + Actors: actors, + Communities: communities, + Repos: repoManager, + Minter: mintGate, + ServiceDID: serviceDID, + ProfileRefreshTTL: cfg.ProfileRefreshTTL, + MaxBlobBytes: cfg.MaxBlobBytes, + StrictValidation: cfg.IsDevelopment(), + Logger: logger, + }) + if err != nil { + return err + } + + backfill, err := ingest.NewBackfill(ingest.BackfillOptions{ + Fetcher: apClient, + Materializer: materializer, + Communities: communities, + Tombstones: tombstones, + MaxPosts: cfg.BackfillMaxPosts, + // Async runs derive from the run context so a mid-run backfill stops + // pulling remote pages once shutdown starts; the drain below waits for + // it, and an interrupted run leaves last_backfill_at unset (resumable). + BaseContext: ctx, + Logger: logger, + }) + if err != nil { + return err + } + handler, err := ingest.NewHandler(ingest.HandlerOptions{ + Materializer: materializer, + Fetcher: apClient, + Objects: objects, + Communities: communities, + Tombstones: tombstones, + Votes: ingest.NewNoopVotes(logger), // task 07 replaces this + Backfill: backfill, + ServiceActorID: serviceActor.ID, + Logger: logger, + }) + if err != nil { + return err + } + queue, err := ingest.NewQueue(ingest.QueueOptions{ + Events: inboxEvents, + Processor: handler, + Workers: cfg.IngestWorkers, + Logger: logger, + }) + if err != nil { + return err + } + go queue.Run(ctx) + + inbox, err := ingest.NewInbox(ingest.InboxOptions{ + Verifier: ap.NewVerifier(apClient), + Events: inboxEvents, + Queue: queue, + Service: serviceActor, + Logger: logger, + }) + if err != nil { + return err + } + inbox.Routes(router) + + admin, err := ingest.NewAdmin(ingest.AdminOptions{ + Token: cfg.AdminToken, + Client: apClient, + Materializer: materializer, + Communities: communities, + Service: serviceActor, + Backfill: backfill, + Logger: logger, + }) + if err != nil { + return err + } + admin.Routes(router) + + // Task 07 registers here: vote aggregates XRPC. server := &http.Server{ Addr: cfg.ListenAddr, @@ -168,6 +304,17 @@ func run(logger *slog.Logger) error { if err := server.Shutdown(shutdownCtx); err != nil { return fmt.Errorf("graceful shutdown: %w", err) } + // Drain in-flight backfill. ctx is already cancelled, so async runs + // are unwinding (they stop pulling remote pages and leave + // last_backfill_at unset, which is resumable). Wait bounded so a stuck + // run can't hold shutdown open past the deadline. + drained := make(chan struct{}) + go func() { backfill.Wait(); close(drained) }() + select { + case <-drained: + case <-shutdownCtx.Done(): + logger.Warn("backfill drain timed out; abandoning in-flight run (resumable on restart)") + } // ListenAndServe has returned by now (Shutdown guarantees it); // drain its error so a bind failure racing the signal still exits // non-zero instead of being lost in the buffered channel. diff --git a/internal/ap/client.go b/internal/ap/client.go index 874eb73..5a33664 100644 --- a/internal/ap/client.go +++ b/internal/ap/client.go @@ -589,13 +589,24 @@ func (c *Client) SendActivity(ctx context.Context, inboxURL string, activity any // document is fetched from the keyId minus fragment (Lemmy convention: // "{actor}#main-key"). func (c *Client) ResolveKey(ctx context.Context, keyID string) (*rsa.PublicKey, string, error) { + key, ownerID, _, err := c.ResolveKeyDetailed(ctx, keyID) + return key, ownerID, err +} + +// ResolveKeyDetailed is ResolveKey plus cache provenance: fromCache reports +// whether the key was served from the positive cache. The Verifier uses it +// to gate its rotation retry — only a cached key can be stale, so only a +// cached key that fails verification earns a fresh re-fetch (bounding the +// outbound-fetch amplification a forged delivery can cause). +func (c *Client) ResolveKeyDetailed(ctx context.Context, keyID string) (key *rsa.PublicKey, ownerID string, fromCache bool, err error) { c.mu.Lock() cached, ok := c.keyCache[keyID] c.mu.Unlock() if ok && c.now().Before(cached.expiresAt) { - return cached.key, cached.ownerID, nil + return cached.key, cached.ownerID, true, nil } - return c.resolveKeyUncached(ctx, keyID) + key, ownerID, err = c.resolveKeyUncached(ctx, keyID) + return key, ownerID, false, err } // resolveKeyUncached fetches and validates the actor for keyID, bypassing the diff --git a/internal/ap/httpsig.go b/internal/ap/httpsig.go index 339d789..8ae13f9 100644 --- a/internal/ap/httpsig.go +++ b/internal/ap/httpsig.go @@ -230,6 +230,17 @@ type freshKeyResolver interface { ResolveKeyFresh(ctx context.Context, keyID string) (*rsa.PublicKey, string, error) } +// cacheAwareKeyResolver is an optional extension of KeyResolver that reports +// whether the resolved key was served from a cache. The Verifier uses it to +// GATE the fresh-key retry: a key that was just fetched from the network +// cannot be stale, so a signature failing against it is a forgery (or +// garbage) and a re-fetch gains nothing — it only lets an attacker turn one +// bogus delivery into two outbound fetches (amplification). The Client +// implements it via ResolveKeyDetailed. +type cacheAwareKeyResolver interface { + ResolveKeyDetailed(ctx context.Context, keyID string) (key *rsa.PublicKey, ownerID string, fromCache bool, err error) +} + // KeyResolverFunc adapts a function to the KeyResolver interface. type KeyResolverFunc func(ctx context.Context, keyID string) (*rsa.PublicKey, string, error) @@ -341,7 +352,23 @@ func (v *Verifier) Verify(ctx context.Context, req *http.Request, body []byte) ( } } - key, ownerID, err := v.resolver.ResolveKey(ctx, keyID) + // Resolve the signing key, learning whether it came from a cache when + // the resolver can tell us (the Client can): only a cached key may be + // stale, so only a cached key earns the rotation retry below. + var ( + key *rsa.PublicKey + ownerID string + fromCache bool + ) + if aware, ok := v.resolver.(cacheAwareKeyResolver); ok { + key, ownerID, fromCache, err = aware.ResolveKeyDetailed(ctx, keyID) + } else { + // Resolvers that cannot report cache provenance keep the old + // behavior (retry allowed) — better a rare extra fetch than + // blackholing a rotated key for the cache TTL. + fromCache = true + key, ownerID, err = v.resolver.ResolveKey(ctx, keyID) + } if err != nil { return "", fmt.Errorf("ap: resolve signing key %q: %w", keyID, err) } @@ -359,12 +386,15 @@ func (v *Verifier) Verify(ctx context.Context, req *http.Request, body []byte) ( if rsa.VerifyPKCS1v15(key, crypto.SHA256, hashed[:], signature) == nil { return ownerID, nil } - // The (possibly cached) key did not verify. If the resolver can bypass its - // cache, try exactly once with a freshly-fetched key: a remote key rotation - // leaves a stale key in cache that fails every signature until it expires. - // We only re-verify when the fresh key actually differs, so a forged - // signature does not gain extra verification attempts. - if fresh, ok := v.resolver.(freshKeyResolver); ok { + // The (possibly cached) key did not verify. If the key came from a cache + // and the resolver can bypass it, try exactly once with a freshly-fetched + // key: a remote key rotation leaves a stale key in cache that fails every + // signature until it expires. The retry is gated on fromCache — a key + // fetched fresh from the network moments ago cannot be stale, and + // re-fetching it would just let forged deliveries amplify outbound + // traffic. We also only re-verify when the fresh key actually differs, so + // a forged signature does not gain extra verification attempts. + if fresh, ok := v.resolver.(freshKeyResolver); ok && fromCache { freshKey, freshOwner, ferr := fresh.ResolveKeyFresh(ctx, keyID) if ferr == nil && !freshKey.Equal(key) && rsa.VerifyPKCS1v15(freshKey, crypto.SHA256, hashed[:], signature) == nil { diff --git a/internal/ap/vocab.go b/internal/ap/vocab.go index f0bf207..5a8b46b 100644 --- a/internal/ap/vocab.go +++ b/internal/ap/vocab.go @@ -105,6 +105,9 @@ type Object struct { Image *Object `json:"image,omitempty"` Published *Time `json:"published,omitempty"` Updated *Time `json:"updated,omitempty"` + // Replies, when advertised, is the object's replies collection (a bare + // IRI or an inline collection). Task 06's backfill pages through it. + Replies *Object `json:"replies,omitempty"` // Lemmy emits a single language object on posts/comments but an ARRAY // of language objects on Group actors; Languages accepts both. diff --git a/internal/config/config.go b/internal/config/config.go index de260ec..73846aa 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -77,6 +77,24 @@ type Config struct { // tests that hit httptest servers on 127.0.0.1. Set ALLOW_PRIVATE_FETCH=1 // to enable. AllowPrivateAddresses bool + // AdminToken is the bearer token protecting the /admin API (community + // subscribe/unsubscribe/backfill). ADMIN_TOKEN; required in production, + // dev default is a fixed, publicly known value. + AdminToken string + // BackfillMaxPosts caps how many posts one community backfill run + // materializes (BACKFILL_MAX_POSTS, default 100). + BackfillMaxPosts int + // MintRatePerMinute is the sustained rate cap on inbound DID minting + // (MINT_RATE_PER_MINUTE, default 60). Inbound AP activity can trigger + // minting (post/comment authors), so it must be bounded — PLC + // registrations are forever. + MintRatePerMinute float64 + // MintBurst is the mint token bucket's burst size (MINT_BURST, default + // 120) — sized to absorb a community backfill's author spike. + MintBurst int + // IngestWorkers is the inbox queue worker-pool size (INGEST_WORKERS, + // default 4). + IngestWorkers int } // Load reads configuration from the environment. logger must not be nil; @@ -191,6 +209,32 @@ func Load(logger *slog.Logger) (*Config, error) { return nil, fmt.Errorf("config: ALLOW_PRIVATE_FETCH must not be set in production") } + // Admin API auth: like the KEK, the dev default is fixed and public — + // required in production. + cfg.AdminToken, err = stringVar(logger, isDevelopment, "ADMIN_TOKEN", "dev-admin-token") + if err != nil { + return nil, err + } + + // Ingestion tuning knobs: real defaults in every environment. + cfg.BackfillMaxPosts, err = intVar(logger, "BACKFILL_MAX_POSTS", 100) + if err != nil { + return nil, err + } + mintRate, err := intVar(logger, "MINT_RATE_PER_MINUTE", 60) + if err != nil { + return nil, err + } + cfg.MintRatePerMinute = float64(mintRate) + cfg.MintBurst, err = intVar(logger, "MINT_BURST", 120) + if err != nil { + return nil, err + } + cfg.IngestWorkers, err = intVar(logger, "INGEST_WORKERS", 4) + if err != nil { + return nil, err + } + defaultUserAgent := fmt.Sprintf("tidepool/0.1 (+https://%s)", cfg.BridgeHostname) cfg.UserAgent = os.Getenv("USER_AGENT") if cfg.UserAgent == "" { @@ -227,6 +271,22 @@ func decodeKEK(encoded string) ([]byte, error) { return raw, nil } +// intVar returns a positive-integer environment variable, falling back to a +// logged default in every environment (tuning knob semantics, like +// FIREHOSE_RETENTION). +func intVar(logger *slog.Logger, name string, fallback int) (int, error) { + raw := os.Getenv(name) + if raw == "" { + logger.Info(name+" not set, using default", "value", fallback) + return fallback, nil + } + parsed, err := strconv.Atoi(raw) + if err != nil || parsed <= 0 { + return 0, fmt.Errorf("config: %s must be a positive integer, got %q", name, raw) + } + return parsed, nil +} + // boolVar reports whether an environment variable is set to a truthy value // ("1", "true", "yes", case-insensitive). func boolVar(name string) bool { diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 79df715..787f8dc 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -18,6 +18,8 @@ func clearConfigEnv(t *testing.T) { for _, name := range []string{ "ENVIRONMENT", "DATABASE_URL", "LISTEN_ADDR", "BRIDGE_HOSTNAME", "PLC_DIRECTORY_URL", "BRIDGE_SERVICE_DID", "USER_AGENT", "BRIDGE_KEK", + "ADMIN_TOKEN", "BACKFILL_MAX_POSTS", "MINT_RATE_PER_MINUTE", + "MINT_BURST", "INGEST_WORKERS", } { t.Setenv(name, "") } @@ -78,6 +80,7 @@ func TestLoad_ProductionWithAllValues(t *testing.T) { t.Setenv("BRIDGE_HOSTNAME", "tidepool.example") t.Setenv("PLC_DIRECTORY_URL", "https://plc.directory") t.Setenv("BRIDGE_KEK", "sfDrM4bIeCJp01ZBTArLPJXNQlD7pcYFsod2An6UAF0=") // base64 form + t.Setenv("ADMIN_TOKEN", "prod-admin-token") cfg, err := Load(discardLogger()) require.NoError(t, err) @@ -85,6 +88,25 @@ func TestLoad_ProductionWithAllValues(t *testing.T) { assert.Equal(t, "tidepool/0.1 (+https://tidepool.example)", cfg.UserAgent, "user agent default derives from the bridge hostname") assert.Len(t, cfg.BridgeKEK, 32, "base64 KEK decodes to 32 bytes") + assert.Equal(t, "prod-admin-token", cfg.AdminToken) + assert.Equal(t, 100, cfg.BackfillMaxPosts, "tuning knobs default in every environment") + assert.Equal(t, float64(60), cfg.MintRatePerMinute) + assert.Equal(t, 120, cfg.MintBurst) + assert.Equal(t, 4, cfg.IngestWorkers) +} + +func TestLoad_ProductionRequiresAdminToken(t *testing.T) { + clearConfigEnv(t) + t.Setenv("ENVIRONMENT", EnvironmentProduction) + t.Setenv("DATABASE_URL", "postgres://prod/db") + t.Setenv("LISTEN_ADDR", ":8080") + t.Setenv("BRIDGE_HOSTNAME", "tidepool.example") + t.Setenv("PLC_DIRECTORY_URL", "https://plc.directory") + t.Setenv("BRIDGE_KEK", "sfDrM4bIeCJp01ZBTArLPJXNQlD7pcYFsod2An6UAF0=") + + _, err := Load(discardLogger()) + require.Error(t, err, "production must never run on the public dev-default admin token") + assert.Contains(t, err.Error(), "ADMIN_TOKEN") } func TestLoad_ProductionRequiresKEK(t *testing.T) { diff --git a/internal/db/migrations/009_ingestion_queue.sql b/internal/db/migrations/009_ingestion_queue.sql new file mode 100644 index 0000000..9daf486 --- /dev/null +++ b/internal/db/migrations/009_ingestion_queue.sql @@ -0,0 +1,50 @@ +-- +goose Up +-- Task 06 turns inbox_events into the durable ingestion work queue (PLAN +-- decision: postgres is the queue, no external broker). The dedupe-by- +-- activity-id contract from migration 004 is unchanged; these columns add +-- what the worker pool needs: +-- payload the raw activity JSON as delivered (verified body) +-- actor_id the signature-verified AP actor the activity binds to +-- ordering_key per-community serialization key: events sharing a key +-- are processed strictly in arrival (id) order +-- attempts how many times a worker claimed the event +-- next_attempt_at retry-with-backoff schedule; claimable when <= now +-- claimed_until worker lease; a crashed worker's claim expires +-- failed_at poison marker: permanently failed events are skipped +-- (error column keeps the reason) and stop blocking their +-- ordering key +ALTER TABLE inbox_events + ADD COLUMN payload BYTEA, + ADD COLUMN actor_id TEXT NOT NULL DEFAULT '', + ADD COLUMN ordering_key TEXT NOT NULL DEFAULT '', + ADD COLUMN attempts INTEGER NOT NULL DEFAULT 0, + ADD COLUMN next_attempt_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP, + ADD COLUMN claimed_until TIMESTAMPTZ, + ADD COLUMN failed_at TIMESTAMPTZ; + +-- The claim query scans pending events per ordering key in id order. +CREATE INDEX idx_inbox_events_queue + ON inbox_events (ordering_key, id) + WHERE processed_at IS NULL AND failed_at IS NULL; + +-- ap_tombstones remembers Deletes that arrived for AP ids we had never +-- materialized. AP delivery is unordered: without this, a Create arriving +-- (or being re-delivered) after its Delete would happily materialize content +-- the origin already removed (the create-after-delete gap flagged in task +-- 05). Undo{Delete} removes the row. +CREATE TABLE ap_tombstones ( + ap_id TEXT PRIMARY KEY CHECK (ap_id <> ''), + deleted_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP +); + +-- +goose Down +DROP TABLE IF EXISTS ap_tombstones; +DROP INDEX IF EXISTS idx_inbox_events_queue; +ALTER TABLE inbox_events + DROP COLUMN IF EXISTS failed_at, + DROP COLUMN IF EXISTS claimed_until, + DROP COLUMN IF EXISTS next_attempt_at, + DROP COLUMN IF EXISTS attempts, + DROP COLUMN IF EXISTS ordering_key, + DROP COLUMN IF EXISTS actor_id, + DROP COLUMN IF EXISTS payload; diff --git a/internal/ingest/backfill.go b/internal/ingest/backfill.go new file mode 100644 index 0000000..2a202df --- /dev/null +++ b/internal/ingest/backfill.go @@ -0,0 +1,359 @@ +package ingest + +import ( + "context" + stderrors "errors" + "fmt" + "log/slog" + "sync" + "time" + + "tidepool/internal/ap" + "tidepool/internal/errors" + "tidepool/internal/materialize" + "tidepool/internal/store" +) + +// Backfill defaults. +const ( + // DefaultBackfillMaxPosts caps how many posts one backfill run + // materializes (BACKFILL_MAX_POSTS). + DefaultBackfillMaxPosts = 100 + // defaultBackfillMinInterval is how fresh last_backfill_at must be for + // an un-forced trigger (a re-delivered Accept) to be skipped. + defaultBackfillMinInterval = time.Hour + // maxRepliesPerPost bounds one post's replies-collection walk. + maxRepliesPerPost = 200 +) + +// BackfillFetcher is the slice of *ap.Client backfill needs. +type BackfillFetcher interface { + FetchActor(ctx context.Context, iri string) (*ap.Object, error) + FetchObject(ctx context.Context, iri string) (*ap.Object, error) + FetchCollection(ctx context.Context, iri string, visit func(*ap.Object) error) error +} + +// BackfillOptions configures NewBackfill. Fetcher, Materializer, +// Communities, and Tombstones are required. +type BackfillOptions struct { + Fetcher BackfillFetcher + Materializer Materializer + Communities store.Communities + Tombstones store.Tombstones + // MaxPosts caps posts per run (default 100, config.BackfillMaxPosts). + MaxPosts int + // MinInterval is the freshness window for un-forced triggers + // (default 1h). + MinInterval time.Duration + // BaseContext is the root context TriggerAsync derives async runs from. + // Wiring the server/run context here lets a mid-run backfill observe + // shutdown (stop pulling remote pages) instead of running on an + // unstoppable context.Background(). Defaults to context.Background(). + BaseContext context.Context + Logger *slog.Logger +} + +// Backfill pages a community's outbox after a Follow is accepted (and on +// demand), materializing history newest→oldest. Rate limiting rides on the +// AP client's per-host limiter; runs are serialized per community and +// resumable: a partial run leaves last_backfill_at unset so the next +// trigger re-walks (deterministic rkeys make re-materialization free). +type Backfill struct { + fetcher BackfillFetcher + mat Materializer + communities store.Communities + tombstones store.Tombstones + maxPosts int + minInterval time.Duration + baseCtx context.Context + logger *slog.Logger + + mu sync.Mutex + running map[string]bool + // wg tracks async runs so tests (and shutdown) can drain them. + wg sync.WaitGroup +} + +// NewBackfill validates options and builds a Backfill. +func NewBackfill(opts BackfillOptions) (*Backfill, error) { + if opts.Fetcher == nil { + return nil, errors.NewValidationError("fetcher", "must not be nil") + } + if opts.Materializer == nil { + return nil, errors.NewValidationError("materializer", "must not be nil") + } + if opts.Communities == nil { + return nil, errors.NewValidationError("communities", "must not be nil") + } + if opts.Tombstones == nil { + return nil, errors.NewValidationError("tombstones", "must not be nil") + } + logger := opts.Logger + if logger == nil { + logger = slog.Default() + } + baseCtx := opts.BaseContext + if baseCtx == nil { + baseCtx = context.Background() + } + b := &Backfill{ + fetcher: opts.Fetcher, + mat: opts.Materializer, + communities: opts.Communities, + tombstones: opts.Tombstones, + maxPosts: opts.MaxPosts, + minInterval: opts.MinInterval, + baseCtx: baseCtx, + logger: logger, + running: map[string]bool{}, + } + if b.maxPosts <= 0 { + b.maxPosts = DefaultBackfillMaxPosts + } + if b.minInterval <= 0 { + b.minInterval = defaultBackfillMinInterval + } + return b, nil +} + +// TriggerAsync starts a backfill run in the background (the follow-accept +// hook and the admin endpoint). Duplicate triggers for a community already +// running are dropped. +func (b *Backfill) TriggerAsync(community *store.Community, force bool) { + b.wg.Add(1) + go func() { + defer b.wg.Done() + if err := b.Run(b.baseCtx, community, force); err != nil { + b.logger.Error("backfill failed", "community", community.APGroupID, "error", err) + } + }() +} + +// Wait blocks until every async run finishes (shutdown/tests). +func (b *Backfill) Wait() { b.wg.Wait() } + +// Run performs one backfill synchronously. force bypasses the +// last-backfill freshness check. Runs for the same community are +// serialized; a concurrent duplicate returns immediately. +func (b *Backfill) Run(ctx context.Context, community *store.Community, force bool) error { + if !force && community.LastBackfillAt != nil && + time.Since(*community.LastBackfillAt) < b.minInterval { + b.logger.Info("backfill skipped: recently completed", + "community", community.APGroupID, "last_backfill_at", community.LastBackfillAt) + return nil + } + + b.mu.Lock() + if b.running[community.APGroupID] { + b.mu.Unlock() + b.logger.Info("backfill already running", "community", community.APGroupID) + return nil + } + b.running[community.APGroupID] = true + b.mu.Unlock() + defer func() { + b.mu.Lock() + delete(b.running, community.APGroupID) + b.mu.Unlock() + }() + + group, err := b.fetcher.FetchActor(ctx, community.APGroupID) + if err != nil { + return fmt.Errorf("ingest: backfill fetch group %s: %w", community.APGroupID, err) + } + if group.Outbox == "" { + b.logger.Warn("backfill: group advertises no outbox", "community", community.APGroupID) + return nil + } + + // Collect first, materialize second: outboxes are newest-first, and the + // spec wants newest→oldest materialization up to the cap. Lemmy + // outboxes are a single OrderedCollection page; Mastodon-style paged + // collections walk first/next. + var items []*ap.Object + walkTruncated := false + err = b.fetcher.FetchCollection(ctx, group.Outbox, func(item *ap.Object) error { + if len(items) >= b.maxPosts { + return ap.ErrStop + } + clone := *item + items = append(items, &clone) + return nil + }) + if err != nil { + var truncated *ap.CollectionTruncatedError + if stderrors.As(err, &truncated) { + // The walk hit the page cap or a paging loop before the natural + // end: everything collected is still materialized, but the run is + // NOT a clean completion. We leave last_backfill_at unset (like the + // failures>0 branch) so an un-forced re-trigger actually re-walks + // rather than skipping inside the freshness window. + walkTruncated = true + b.logger.Warn("backfill: outbox walk truncated; older history not reached", + "community", community.APGroupID, "pages", truncated.Pages, "next", truncated.Next) + } else { + return fmt.Errorf("ingest: backfill outbox %s: %w", group.Outbox, err) + } + } + + posts, failures := 0, 0 + for _, item := range items { + if ctx.Err() != nil { + return ctx.Err() + } + ok, err := b.materializeOutboxItem(ctx, item, community.APGroupID) + switch { + case err == nil: + if ok { + posts++ + } + case materialize.IsSkip(err): + b.logger.Info("backfill item skipped", "community", community.APGroupID, "reason", err.Error()) + default: + failures++ + b.logger.Warn("backfill item failed", "community", community.APGroupID, "error", err) + } + } + + if failures > 0 { + // Leave last_backfill_at untouched so the next trigger retries the + // walk (resumable-by-redo; deterministic rkeys make it idempotent). + return fmt.Errorf("ingest: backfill %s: %d of %d items failed", + community.APGroupID, failures, len(items)) + } + if walkTruncated { + // Truncated walks are not clean completions: leave last_backfill_at + // unset so an un-forced re-trigger re-walks from the top instead of + // skipping inside the freshness window. Collected items already + // materialized above; deterministic rkeys make the redo idempotent. + b.logger.Info("backfill partial: truncated walk left resumable", + "community", community.APGroupID, "posts", posts) + return nil + } + if err := b.communities.SetLastBackfill(ctx, community.APGroupID, time.Now()); err != nil { + return fmt.Errorf("ingest: record backfill for %s: %w", community.APGroupID, err) + } + b.logger.Info("backfill complete", "community", community.APGroupID, "posts", posts) + return nil +} + +// materializeOutboxItem unwraps one outbox entry (Announce{Create{Page}}, +// Create{Page}, or a bare Page) and materializes the post plus its +// advertised replies. ok reports whether a post actually landed. +func (b *Backfill) materializeOutboxItem(ctx context.Context, item *ap.Object, communityIRI string) (bool, error) { + obj := item + for obj != nil && (obj.Type == ap.TypeAnnounce || obj.Type == ap.TypeCreate || obj.Type == ap.TypeUpdate) { + obj = obj.Object + } + if obj == nil { + return false, skip(item.ID, "outbox item carries no object") + } + if obj.ID == "" { + return false, skip(item.ID, "outbox item object has no id") + } + + // Same funnel rules as live deliveries: never resurrect deleted + // content, trust embedded bodies only on the outbox host's authority. + tombstoned, err := b.tombstones.Exists(ctx, obj.ID) + if err != nil { + return false, fmt.Errorf("ingest: tombstone check for %s: %w", obj.ID, err) + } + if tombstoned { + return false, skip(obj.ID, "object was deleted upstream") + } + obj, err = b.resolveEmbedded(ctx, obj, communityIRI) + if err != nil { + return false, err + } + + switch obj.Type { + case ap.TypePage, ap.TypeArticle: + if _, err := b.mat.MaterializePost(ctx, obj); err != nil { + return false, err + } + b.backfillReplies(ctx, obj) + return true, nil + case ap.TypeNote: + if _, err := b.mat.MaterializeComment(ctx, obj); err != nil { + return false, err + } + return true, nil + default: + return false, skip(obj.ID, "unsupported outbox object type "+obj.Type) + } +} + +// backfillReplies pages a post's advertised replies collection. Failures +// are logged, never fatal — replies are best-effort garnish on backfill. +func (b *Backfill) backfillReplies(ctx context.Context, post *ap.Object) { + if post.Replies == nil || post.Replies.ID == "" { + // Not advertised (or inline-only, which Lemmy never emits). + return + } + repliesIRI := post.Replies.ID + count := 0 + err := b.fetcher.FetchCollection(ctx, repliesIRI, func(item *ap.Object) error { + if count >= maxRepliesPerPost { + return ap.ErrStop + } + count++ + note := *item + resolved, err := b.resolveEmbedded(ctx, ¬e, repliesIRI) + if err != nil { + b.logger.Info("backfill reply skipped", "post", post.ID, "error", err.Error()) + return nil + } + if resolved.Type != ap.TypeNote { + return nil + } + // Same funnel rule as the live path and materializeOutboxItem: a reply + // with a recorded Delete must never be resurrected, even if it still + // lingers in the origin's replies collection (delivery/collection race). + tombstoned, err := b.tombstones.Exists(ctx, resolved.ID) + if err != nil { + b.logger.Warn("backfill reply tombstone check failed", "post", post.ID, "reply", resolved.ID, "error", err) + return nil + } + if tombstoned { + b.logger.Info("backfill reply skipped", "post", post.ID, "reason", "reply was deleted upstream") + return nil + } + if _, err := b.mat.MaterializeComment(ctx, resolved); err != nil { + if materialize.IsSkip(err) { + b.logger.Info("backfill reply skipped", "post", post.ID, "reason", err.Error()) + return nil + } + b.logger.Warn("backfill reply failed", "post", post.ID, "error", err) + } + return nil + }) + if err != nil && !stderrors.Is(err, ap.ErrCollectionTruncated) { + b.logger.Warn("backfill replies walk failed", "post", post.ID, "error", err) + } +} + +// resolveEmbedded applies the embedded-object trust rule to collection +// items: bodies whose id lives on the collection's own authority are used +// as-is, everything else is re-fetched from its origin (with the +// self-asserted-id binding). +func (b *Backfill) resolveEmbedded(ctx context.Context, obj *ap.Object, sourceIRI string) (*ap.Object, error) { + if obj.Type != "" && ap.SameAuthority(obj.ID, sourceIRI) { + return obj, nil + } + fetched, err := b.fetcher.FetchObject(ctx, obj.ID) + switch { + case err == nil: + case errors.IsTombstoned(err): + return nil, skip(obj.ID, "object is tombstoned upstream") + case errors.IsNotFound(err): + return nil, skip(obj.ID, "object is unavailable upstream") + default: + return nil, fmt.Errorf("ingest: fetch %s: %w", obj.ID, err) + } + if fetched.ID == "" { + fetched.ID = obj.ID + } else if !ap.SameAuthority(fetched.ID, obj.ID) { + return nil, skip(obj.ID, fmt.Sprintf("fetched object served a cross-authority id %s", fetched.ID)) + } + return fetched, nil +} diff --git a/internal/ingest/backfill_test.go b/internal/ingest/backfill_test.go new file mode 100644 index 0000000..082d827 --- /dev/null +++ b/internal/ingest/backfill_test.go @@ -0,0 +1,283 @@ +package ingest + +import ( + "context" + "encoding/json" + "net/http" + "os" + "path/filepath" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "tidepool/internal/errors" + "tidepool/internal/materialize" +) + +const secondPageID = "https://lemmy.world/post/49122698" + +// newBackfill builds a real Backfill over the harness stack. +func newBackfill(t *testing.T, h *harness, maxPosts int) *Backfill { + t.Helper() + b, err := NewBackfill(BackfillOptions{ + Fetcher: h.client, + Materializer: h.mat, + Communities: h.communities, + Tombstones: h.tombstones, + MaxPosts: maxPosts, + }) + require.NoError(t, err) + return b +} + +// serveOutboxFixtures registers the outbox fixture plus everything its +// items need: both pages' authors and a replies collection advertised on +// the first page. +func serveOutboxFixtures(t *testing.T, h *harness) { + t.Helper() + h.serveLemmyWorldContent() + h.serveObject("/u/Aweigh", person("https://lemmy.world/u/Aweigh", "Aweigh", nil)) + + outboxRaw, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "outbox_lemmy_world.json")) + require.NoError(t, err) + var outbox map[string]any + require.NoError(t, json.Unmarshal(outboxRaw, &outbox)) + // Advertise a replies collection on the newest page (Lemmy doesn't + // today, but the spec covers "if advertised"). + items := outbox["orderedItems"].([]any) + page := items[0].(map[string]any)["object"].(map[string]any)["object"].(map[string]any) + page["replies"] = pageID + "/replies" + h.serveObject("/c/technology/outbox", outbox) + + replyID := "https://lemmy.world/comment/3001" + h.serveObject("/u/replier", person("https://lemmy.world/u/replier", "replier", nil)) + h.serveObject("/post/49131386/replies", map[string]any{ + "type": "OrderedCollection", + "id": pageID + "/replies", + "totalItems": 1, + "orderedItems": []any{note(replyID, "https://lemmy.world/u/replier", pageID, "a backfilled reply", "2026-07-07T05:00:00.000000Z")}, + }) +} + +// TestBackfillProducesMappedHistory is the DoD backfill check: paging the +// group outbox materializes the posts (newest first) and each advertised +// replies collection, then stamps last_backfill_at. +func TestBackfillProducesMappedHistory(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + serveOutboxFixtures(t, h) + b := newBackfill(t, h, 10) + ctx := context.Background() + + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.NoError(t, b.Run(ctx, community, true)) + + // Both outbox posts landed. + for _, id := range []string{pageID, secondPageID} { + mapping, err := h.objects.GetByAPID(ctx, id) + require.NoError(t, err, "outbox post %s must be materialized", id) + assert.Equal(t, materialize.CollectionPost, mapping.Collection) + } + // The advertised reply landed too. + replyMapping, err := h.objects.GetByAPID(ctx, "https://lemmy.world/comment/3001") + require.NoError(t, err, "advertised replies must be backfilled") + assert.Equal(t, materialize.CollectionComment, replyMapping.Collection) + + community, err = h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.NotNil(t, community.LastBackfillAt, "a clean run stamps last_backfill_at") + + // A fresh un-forced trigger is skipped (resumable-freshness contract). + before := *community.LastBackfillAt + require.NoError(t, b.Run(ctx, community, false)) + community, err = h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + assert.WithinDuration(t, before, *community.LastBackfillAt, time.Second, + "an un-forced re-trigger inside the freshness window must not re-run") +} + +// TestBackfillHonorsMaxPosts: the post cap stops the walk early. +func TestBackfillHonorsMaxPosts(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + serveOutboxFixtures(t, h) + b := newBackfill(t, h, 1) + ctx := context.Background() + + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.NoError(t, b.Run(ctx, community, true)) + + _, err = h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err, "the newest post lands") + _, err = h.objects.GetByAPID(ctx, secondPageID) + assert.True(t, errors.IsNotFound(err), "posts past BACKFILL_MAX_POSTS are not materialized") +} + +// TestBackfillSkipsTombstonedReplies: a reply still advertised in the origin's +// replies collection but with a recorded Delete (delivery/collection race) must +// NOT be resurrected during backfill — the same funnel rule the live path and +// materializeOutboxItem enforce (Finding J). +func TestBackfillSkipsTombstonedReplies(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + serveOutboxFixtures(t, h) + b := newBackfill(t, h, 10) + ctx := context.Background() + + const replyID = "https://lemmy.world/comment/3001" + require.NoError(t, h.tombstones.Record(ctx, replyID)) + + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.NoError(t, b.Run(ctx, community, true)) + + // The post carrying the replies collection still lands. + _, err = h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err, "the post still materializes") + // The tombstoned reply is not resurrected. + _, err = h.objects.GetByAPID(ctx, replyID) + assert.True(t, errors.IsNotFound(err), "a tombstoned reply must not be backfilled") +} + +// serveTruncatingOutbox re-serves the group outbox as a next-pointer loop so +// FetchCollection returns a CollectionTruncatedError after collecting both +// posts. Requires serveOutboxFixtures to have registered the authors/pages. +func serveTruncatingOutbox(t *testing.T, h *harness) { + t.Helper() + outboxRaw, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "outbox_lemmy_world.json")) + require.NoError(t, err) + var outbox map[string]any + require.NoError(t, json.Unmarshal(outboxRaw, &outbox)) + items := outbox["orderedItems"].([]any) + + const ( + p1 = "https://lemmy.world/c/technology/outbox/page1" + p2 = "https://lemmy.world/c/technology/outbox/page2" + ) + // Header points at page1; page1→page2→page1 loops, so the walk truncates + // after both items are collected (never reaching a natural end). + h.serveObject("/c/technology/outbox", map[string]any{ + "type": "OrderedCollection", + "id": "https://lemmy.world/c/technology/outbox", + "first": p1, + }) + h.serveObject("/c/technology/outbox/page1", map[string]any{ + "type": "OrderedCollectionPage", "id": p1, "next": p2, + "orderedItems": []any{items[0]}, + }) + h.serveObject("/c/technology/outbox/page2", map[string]any{ + "type": "OrderedCollectionPage", "id": p2, "next": p1, + "orderedItems": []any{items[1]}, + }) +} + +// TestBackfillTruncatedWalkLeavesResumable: a truncated outbox walk still +// materializes everything collected, but is NOT a clean completion — +// last_backfill_at stays nil so an un-forced re-trigger actually re-walks +// instead of skipping inside the freshness window (Finding L). +func TestBackfillTruncatedWalkLeavesResumable(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + serveOutboxFixtures(t, h) + serveTruncatingOutbox(t, h) + b := newBackfill(t, h, 100) + ctx := context.Background() + + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.NoError(t, b.Run(ctx, community, true), "truncation is not a fatal error") + + // Both collected posts still materialized. + for _, id := range []string{pageID, secondPageID} { + _, err := h.objects.GetByAPID(ctx, id) + require.NoError(t, err, "collected post %s must still materialize", id) + } + // The truncated run is not a completion: last_backfill_at stays nil. + community, err = h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.Nil(t, community.LastBackfillAt, + "a truncated walk must leave last_backfill_at unset so a re-trigger re-walks") + + // An un-forced re-trigger is NOT skipped (nil last_backfill_at) — it + // re-walks the outbox, proving resumability. + before := h.hitCount("/c/technology/outbox") + require.NoError(t, b.Run(ctx, community, false)) + assert.Greater(t, h.hitCount("/c/technology/outbox"), before, + "an un-forced re-trigger must re-walk when last_backfill_at is unset") +} + +// TestBackfillPartialFailureLeavesResumable: a hard (non-skip) item failure +// leaves last_backfill_at unset so the next trigger retries the walk. +func TestBackfillPartialFailureLeavesResumable(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + serveOutboxFixtures(t, h) + // One good item (the real fixture's fully-embedded page) plus one broken + // item whose cross-authority object 500s: a hard failure (not a skip), so + // failures>0. + outboxRaw, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "outbox_lemmy_world.json")) + require.NoError(t, err) + var outbox map[string]any + require.NoError(t, json.Unmarshal(outboxRaw, &outbox)) + goodItem := outbox["orderedItems"].([]any)[0] + + const brokenID = "https://other.test/post/broken" + h.mux.HandleFunc("GET /post/broken", func(w http.ResponseWriter, _ *http.Request) { + http.Error(w, "boom", http.StatusInternalServerError) + }) + h.serveObject("/c/technology/outbox", map[string]any{ + "type": "OrderedCollection", + "id": "https://lemmy.world/c/technology/outbox", + "totalItems": 2, + "orderedItems": []any{ + goodItem, + // A broken item: cross-authority id with no embedded body forces a + // fetch that 500s. + map[string]any{ + "type": "Create", + "actor": personID, + "object": map[string]any{"id": brokenID}, + }, + }, + }) + b := newBackfill(t, h, 100) + ctx := context.Background() + + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + err = b.Run(ctx, community, true) + require.Error(t, err, "a hard item failure surfaces as a run error") + + // The good post still landed. + _, err = h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err, "the healthy item still materializes") + + community, err = h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.Nil(t, community.LastBackfillAt, + "a partial-failure run must leave last_backfill_at unset (resumable)") +} + +// TestBackfillSkipsTombstonedObjects: deleted-upstream markers hold during +// backfill too. +func TestBackfillSkipsTombstonedObjects(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + serveOutboxFixtures(t, h) + b := newBackfill(t, h, 10) + ctx := context.Background() + + require.NoError(t, h.tombstones.Record(ctx, pageID)) + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + require.NoError(t, b.Run(ctx, community, true)) + + _, err = h.objects.GetByAPID(ctx, pageID) + assert.True(t, errors.IsNotFound(err), "a tombstoned id must not be backfilled") + _, err = h.objects.GetByAPID(ctx, secondPageID) + require.NoError(t, err, "other posts still land") +} diff --git a/internal/ingest/consent.go b/internal/ingest/consent.go new file mode 100644 index 0000000..b5e1103 --- /dev/null +++ b/internal/ingest/consent.go @@ -0,0 +1,213 @@ +// Consent enforcement (PLAN.md locked decision 6, mirroring Bridgy Fed +// norms — the policy is documented in README "Consent policy"): +// +// - #nobridge / #nobot in an actor's summary or tags blocks +// materialization. The materializer scans on FIRST SIGHT (an opted-out +// actor is never minted) and on every profile refresh; a previously +// bridged actor who adds the marker gets their existing records +// scrubbed and consent set to nobridge (reversible: removing the +// marker restores bridging on the next refresh). +// - Delete(Actor) tombstones the bridged repo: every record scrubbed, +// consent set to deleted (terminal). +// +// This file wires the activity-side triggers: profile Updates (the "on +// profile Update" scan — RefreshActor/RefreshCommunity re-run the marker +// scan), Delete dispatch, and Undo{Delete} restores. The signature contract +// from task 05 is enforced here: an embedded actor document is trusted only +// when the verified signer IS that actor; anything else is re-fetched from +// its origin. + +package ingest + +import ( + "context" + "fmt" + + "tidepool/internal/ap" + "tidepool/internal/errors" + "tidepool/internal/materialize" +) + +// applyProfileUpdate handles Update{Person|Group}. The embedded document is +// passed to the materializer (which trusts it) ONLY when the verified +// signer is the updated actor itself; otherwise the update degrades to a +// forced re-fetch by IRI — same effect, no trust in the delivered body. +// Refresh* re-runs the consent-marker scan, so a profile that gained +// #nobridge is scrubbed and suppressed on this path. +// +// A BARE (non-announced) profile Update is refresh-only: it may re-materialize +// an actor/community the bridge ALREADY bridged, but must never mint or bridge +// a previously-unknown one. Otherwise any self-signed actor could drive an +// outbound fetch to an arbitrary target IRI (an SSRF / fetch-oracle) and burn +// the permanent-PLC mint budget, defeating the subscription-trust model. +// Announced profile updates ride the followed community's trust (handleAnnounce +// already authorized the announcer) and may ensure/refresh unknown actors. +func (h *Handler) applyProfileUpdate(ctx context.Context, actorDoc *ap.Object, signer, announcer string) error { + if announcer == "" { + known, err := h.isBridged(ctx, actorDoc.ID) + if err != nil { + return err + } + if !known { + return skip(actorDoc.ID, + "bare profile update for an actor we have not bridged (refresh-only, never mint)") + } + } + ref := actorDoc + if actorDoc.ID != signer { + // Not self-signed: use a bare reference so the materializer + // re-fetches the document from the actor's own origin. + ref = &ap.Object{ID: actorDoc.ID} + } + var err error + if actorDoc.Type == ap.TypeGroup { + _, err = h.mat.RefreshCommunity(ctx, ref) + } else { + _, err = h.mat.RefreshActor(ctx, ref) + } + return err +} + +// handleDelete processes Delete{object-or-actor}. announcer is the +// announcing community's AP id ("" when delivered bare). +// +// Authorization: a Delete announced by a followed community is trusted (the +// community moderates its own content — Lemmy only announces deletes for +// objects in its communities). A bare Delete must come from the deleted +// id's own authority (the actor deleting their content/account, or their +// instance acting for them). Anything else is dropped. +func (h *Handler) handleDelete(ctx context.Context, del *ap.Object, signer, announcer string) error { + targetID := refID(del.Object) + if targetID == "" { + return errors.NewValidationError("delete", "delete carries no object id") + } + if err := h.authorizeDelete(ctx, del.ID, targetID, signer, announcer); err != nil { + return err + } + + // Record the tombstone marker BEFORE deleting: if this is a Delete for + // an object we never materialized, the marker is the only thing + // stopping a later (re-delivered, out-of-order) Create from + // resurrecting it. + if err := h.tombstones.Record(ctx, targetID); err != nil { + return fmt.Errorf("ingest: record tombstone for %s: %w", targetID, err) + } + // HandleDelete branches actor vs object itself: a known actor id runs + // the full Delete(Actor) scrub (records tombstoned, consent → deleted, + // terminal); an object id deletes the record and soft-deletes its + // mapping; an unknown id is a logged no-op. + if err := h.mat.HandleDelete(ctx, targetID); err != nil { + return err + } + return nil +} + +// handleUndo processes Undo{Like|Dislike|Delete|Follow}. +func (h *Handler) handleUndo(ctx context.Context, undo *ap.Object, signer, announcer string) error { + inner := undo.Object + if inner == nil || inner.Type == "" { + // A bare-IRI undo target is unactionable: we cannot know what kind + // of activity is being undone without its body. + return skip(undo.ID, "undo carries no inline activity") + } + switch inner.Type { + case ap.TypeLike, ap.TypeDislike: + return h.votes.RetractVote(ctx, inner, announcer) + case ap.TypeDelete: + return h.handleUndoDelete(ctx, undo, inner, signer, announcer) + case ap.TypeFollow: + // A remote undoing a follow of us — the bridge has no followers in + // v1 (read-only), nothing to do. + return skip(undo.ID, "undo of a follow is not applicable to the bridge") + default: + return skip(undo.ID, "unsupported undone activity type "+inner.Type) + } +} + +// handleUndoDelete restores an object whose Delete was previously applied: +// clear the create-after-delete marker, re-fetch the object from its origin +// (never trust the undo body), clear the mapping's soft delete, and +// re-materialize. Idempotent; a restore for content that is still gone +// upstream is a skip. +func (h *Handler) handleUndoDelete(ctx context.Context, undo, del *ap.Object, signer, announcer string) error { + targetID := refID(del.Object) + if targetID == "" { + return skip(undo.ID, "undone delete carries no object id") + } + // Same authorization rule as the delete itself (a restore must not let a + // cross-authority or co-hosted actor un-delete a victim's content). + if err := h.authorizeDelete(ctx, undo.ID, targetID, signer, announcer); err != nil { + return err + } + + // The origin must actually serve the object again — a restore is only + // as real as the content behind it. + restored, err := h.fetchBound(ctx, targetID) + if err != nil { + return err + } + + if err := h.tombstones.Remove(ctx, targetID); err != nil { + return fmt.Errorf("ingest: clear tombstone for %s: %w", targetID, err) + } + if err := h.objects.Restore(ctx, targetID); err != nil && !errors.IsNotFound(err) { + return fmt.Errorf("ingest: restore mapping for %s: %w", targetID, err) + } + h.logger.Info("object restored upstream; re-materializing", "ap_id", targetID) + if _, err = h.mat.HandleUpdate(ctx, restored); err != nil { + // Compensation: the mapping's soft delete is already cleared and its + // record was deleted from the repo. If re-materialization declines the + // object (a skip — nobridge/deleted author/tombstoned ancestor — or a + // validation error), leaving the mapping live would strand it WITHOUT a + // record (downstream parent-lookup/echo/GetByAPID would treat it as + // materialized). Roll back to the pre-undo state: re-soft-delete the + // mapping and re-record the tombstone. + if materialize.IsSkip(err) || errors.IsValidation(err) { + if rerr := h.objects.SoftDelete(ctx, targetID); rerr != nil && !errors.IsNotFound(rerr) { + return fmt.Errorf("ingest: re-soft-delete after declined restore of %s: %w", targetID, rerr) + } + if rerr := h.tombstones.Record(ctx, targetID); rerr != nil { + return fmt.Errorf("ingest: re-record tombstone after declined restore of %s: %w", targetID, rerr) + } + h.logger.Info("restore re-materialization declined; rolled back to deleted state", + "ap_id", targetID, "reason", err) + } + return err + } + return nil +} + +// authorizeDelete enforces who may Delete (or Undo{Delete}) a target id. +// +// - Bare (unannounced): only the target id's OWN authority may delete it — +// the actor removing their content/account, or their instance acting for +// them. +// - Announced by a followed community: the target must live on the +// announcing community's own authority (a community moderates only its own +// instance's content). For a target that is itself a bridged ACTOR — the +// terminal DeleteActor scrub — the community may delete only ITSELF, never +// a co-hosted OTHER actor whose bridged presence spans other communities. +func (h *Handler) authorizeDelete(ctx context.Context, activityID, targetID, signer, announcer string) error { + if announcer == "" { + if !ap.SameAuthority(targetID, signer) { + return skip(activityID, fmt.Sprintf( + "delete of %s signed by cross-authority actor %s", targetID, signer)) + } + return nil + } + if !ap.SameAuthority(targetID, announcer) { + return skip(activityID, fmt.Sprintf( + "announced delete of %s by cross-authority community %s", targetID, announcer)) + } + if targetID != announcer { + isActor, err := h.targetIsBridgedActor(ctx, targetID) + if err != nil { + return err + } + if isActor { + return skip(activityID, fmt.Sprintf( + "community %s may not delete co-hosted actor %s", announcer, targetID)) + } + } + return nil +} diff --git a/internal/ingest/follow.go b/internal/ingest/follow.go new file mode 100644 index 0000000..acb2c91 --- /dev/null +++ b/internal/ingest/follow.go @@ -0,0 +1,391 @@ +package ingest + +import ( + "context" + "crypto/rand" + "encoding/hex" + "encoding/json" + "fmt" + "log/slog" + "net/http" + "strings" + + "github.com/go-chi/chi/v5" + + "tidepool/internal/ap" + "tidepool/internal/errors" + "tidepool/internal/materialize" + "tidepool/internal/store" +) + +// FollowClient is the slice of *ap.Client the follow lifecycle needs. +type FollowClient interface { + ResolveHandle(ctx context.Context, handle string) (string, error) + FetchActor(ctx context.Context, iri string) (*ap.Object, error) + SendActivity(ctx context.Context, inboxURL string, activity any) error +} + +// AdminOptions configures NewAdmin. Everything except Logger is required. +type AdminOptions struct { + // Token is the bearer token protecting /admin (config.AdminToken). + Token string + // Client resolves handles, fetches Group actors, and delivers the + // signed Follow/Undo activities. + Client FollowClient + // Materializer bridges the community before we follow it (profile + // record first, PLAN decision 3). + Materializer Materializer + Communities store.Communities + // Service is the AP actor the Follow is issued by. + Service *ap.ServiceActor + // Backfill serves the on-demand backfill endpoint (optional). + Backfill Backfiller + Logger *slog.Logger +} + +// Admin is the operator API driving the community subscription lifecycle: +// +// POST /admin/communities {"community":"!tech@lemmy.world"} +// DELETE /admin/communities {"community":"!tech@lemmy.world"} +// GET /admin/communities +// POST /admin/communities/backfill {"community":"!tech@lemmy.world"} +// +// All endpoints require "Authorization: Bearer $ADMIN_TOKEN". +type Admin struct { + token string + client FollowClient + mat Materializer + communities store.Communities + service *ap.ServiceActor + backfill Backfiller + logger *slog.Logger +} + +// NewAdmin validates options and builds the Admin API. +func NewAdmin(opts AdminOptions) (*Admin, error) { + if opts.Token == "" { + return nil, errors.NewValidationError("token", "must not be empty") + } + if opts.Client == nil { + return nil, errors.NewValidationError("client", "must not be nil") + } + if opts.Materializer == nil { + return nil, errors.NewValidationError("materializer", "must not be nil") + } + if opts.Communities == nil { + return nil, errors.NewValidationError("communities", "must not be nil") + } + if opts.Service == nil { + return nil, errors.NewValidationError("service", "must not be nil") + } + logger := opts.Logger + if logger == nil { + logger = slog.Default() + } + return &Admin{ + token: opts.Token, + client: opts.Client, + mat: opts.Materializer, + communities: opts.Communities, + service: opts.Service, + backfill: opts.Backfill, + logger: logger, + }, nil +} + +// Routes mounts the admin API on a chi router. +func (a *Admin) Routes(r chi.Router) { + r.Route("/admin", func(r chi.Router) { + r.Use(requireBearer(a.token, a.logger)) + r.Post("/communities", a.handleSubscribe) + r.Delete("/communities", a.handleUnsubscribe) + r.Get("/communities", a.handleList) + r.Post("/communities/backfill", a.handleBackfill) + }) +} + +// communityRequest is the shared request body: a Lemmy-style handle +// ("!tech@lemmy.world") or the Group's AP id URL. +type communityRequest struct { + Community string `json:"community"` +} + +// communityResponse reports a community's bridge state. +type communityResponse struct { + Community string `json:"community"` + DID string `json:"did"` + FollowState string `json:"follow_state"` + LastBackfillAt string `json:"last_backfill_at,omitempty"` +} + +func communityJSON(c *store.Community) communityResponse { + resp := communityResponse{ + Community: c.APGroupID, + DID: c.DID, + FollowState: string(c.FollowState), + } + if c.LastBackfillAt != nil { + resp.LastBackfillAt = c.LastBackfillAt.UTC().Format("2006-01-02T15:04:05Z") + } + return resp +} + +// handleSubscribe resolves, bridges, and follows a community: +// WebFinger → fetch Group → materialize community profile → signed Follow +// from the service actor → follow_state pending (Accept arrives via the +// inbox and flips it to accepted, which triggers backfill). +func (a *Admin) handleSubscribe(w http.ResponseWriter, r *http.Request) { + ctx := r.Context() + var req communityRequest + if err := decodeJSONBody(r, &req); err != nil || strings.TrimSpace(req.Community) == "" { + http.Error(w, `body must be {"community":"!name@instance"}`, http.StatusBadRequest) + return + } + + groupIRI, err := a.resolveCommunity(ctx, req.Community) + if err != nil { + a.writeResolveError(w, req.Community, err) + return + } + + // Bridge the community first: DID minted, community.profile committed — + // content referencing it can land the moment announces start. + community, err := a.mat.EnsureCommunity(ctx, &ap.Object{ID: groupIRI}) + if err != nil { + if materialize.IsSkip(err) { + a.logger.Warn("subscribe refused", "community", groupIRI, "reason", err.Error()) + http.Error(w, "community cannot be bridged: "+err.Error(), http.StatusUnprocessableEntity) + return + } + a.logger.Error("subscribe: materialize community", "community", groupIRI, "error", err) + http.Error(w, "failed to bridge community", http.StatusBadGateway) + return + } + + if community.FollowState == store.FollowStateAccepted { + // Already subscribed; idempotent success. + writeJSON(w, http.StatusOK, communityJSON(community)) + return + } + + group, err := a.client.FetchActor(ctx, groupIRI) + if err != nil { + a.logger.Error("subscribe: fetch group", "community", groupIRI, "error", err) + http.Error(w, "failed to fetch community actor", http.StatusBadGateway) + return + } + inbox := group.SharedInboxOrInbox() + if inbox == "" { + http.Error(w, "community actor advertises no inbox", http.StatusUnprocessableEntity) + return + } + + follow, err := a.buildFollow(groupIRI) + if err != nil { + a.logger.Error("subscribe: build follow", "community", groupIRI, "error", err) + http.Error(w, "failed to build Follow", http.StatusInternalServerError) + return + } + if err := a.client.SendActivity(ctx, inbox, follow); err != nil { + a.logger.Error("subscribe: deliver follow", "community", groupIRI, "error", err) + http.Error(w, "failed to deliver Follow", http.StatusBadGateway) + return + } + if err := a.communities.SetFollowState(ctx, groupIRI, store.FollowStatePending); err != nil { + a.logger.Error("subscribe: record pending follow", "community", groupIRI, "error", err) + http.Error(w, "failed to record follow state", http.StatusInternalServerError) + return + } + a.logger.Info("follow sent; awaiting accept", "community", groupIRI, "did", community.DID) + + community.FollowState = store.FollowStatePending + writeJSON(w, http.StatusAccepted, communityJSON(community)) +} + +// handleUnsubscribe sends Undo{Follow} and clears the follow state. The +// local state is cleared even when the remote can no longer be reached — +// unsubscribing from a dead instance must succeed. +func (a *Admin) handleUnsubscribe(w http.ResponseWriter, r *http.Request) { + ctx := r.Context() + var req communityRequest + if err := decodeJSONBody(r, &req); err != nil || strings.TrimSpace(req.Community) == "" { + http.Error(w, `body must be {"community":"!name@instance"}`, http.StatusBadRequest) + return + } + + groupIRI, err := a.resolveCommunity(ctx, req.Community) + if err != nil { + a.writeResolveError(w, req.Community, err) + return + } + community, err := a.communities.GetByAPGroupID(ctx, groupIRI) + if errors.IsNotFound(err) { + http.Error(w, "community is not bridged", http.StatusNotFound) + return + } + if err != nil { + a.logger.Error("unsubscribe: look up community", "community", groupIRI, "error", err) + http.Error(w, "internal error", http.StatusInternalServerError) + return + } + + // Best-effort remote notification. + if group, err := a.client.FetchActor(ctx, groupIRI); err == nil { + if inbox := group.SharedInboxOrInbox(); inbox != "" { + undo, err := a.buildUndoFollow(groupIRI) + if err != nil { + // A local crypto/rand failure: skip the notification rather + // than emitting a zero-entropy id, but still clear local state + // (unsubscribing must succeed even if the Undo can't be sent). + a.logger.Warn("unsubscribe: build undo failed (state cleared anyway)", + "community", groupIRI, "error", err) + } else if err := a.client.SendActivity(ctx, inbox, undo); err != nil { + a.logger.Warn("unsubscribe: deliver undo failed (state cleared anyway)", + "community", groupIRI, "error", err) + } + } + } else { + a.logger.Warn("unsubscribe: fetch group failed (state cleared anyway)", + "community", groupIRI, "error", err) + } + + if err := a.communities.SetFollowState(ctx, groupIRI, store.FollowStateNone); err != nil { + a.logger.Error("unsubscribe: clear follow state", "community", groupIRI, "error", err) + http.Error(w, "failed to clear follow state", http.StatusInternalServerError) + return + } + a.logger.Info("community unfollowed", "community", groupIRI) + community.FollowState = store.FollowStateNone + writeJSON(w, http.StatusOK, communityJSON(community)) +} + +// handleList reports every community in accepted or pending state. +func (a *Admin) handleList(w http.ResponseWriter, r *http.Request) { + ctx := r.Context() + out := []communityResponse{} + for _, state := range []store.FollowState{store.FollowStateAccepted, store.FollowStatePending} { + communities, err := a.communities.ListByFollowState(ctx, state) + if err != nil { + a.logger.Error("list communities", "state", state, "error", err) + http.Error(w, "internal error", http.StatusInternalServerError) + return + } + for _, c := range communities { + out = append(out, communityJSON(c)) + } + } + writeJSON(w, http.StatusOK, map[string]any{"communities": out}) +} + +// handleBackfill triggers an on-demand backfill for a subscribed community. +func (a *Admin) handleBackfill(w http.ResponseWriter, r *http.Request) { + if a.backfill == nil { + http.Error(w, "backfill is not configured", http.StatusNotImplemented) + return + } + ctx := r.Context() + var req communityRequest + if err := decodeJSONBody(r, &req); err != nil || strings.TrimSpace(req.Community) == "" { + http.Error(w, `body must be {"community":"!name@instance"}`, http.StatusBadRequest) + return + } + groupIRI, err := a.resolveCommunity(ctx, req.Community) + if err != nil { + a.writeResolveError(w, req.Community, err) + return + } + community, err := a.communities.GetByAPGroupID(ctx, groupIRI) + if errors.IsNotFound(err) { + http.Error(w, "community is not bridged", http.StatusNotFound) + return + } + if err != nil { + a.logger.Error("backfill: look up community", "community", groupIRI, "error", err) + http.Error(w, "internal error", http.StatusInternalServerError) + return + } + a.backfill.TriggerAsync(community, true) + writeJSON(w, http.StatusAccepted, communityJSON(community)) +} + +// resolveCommunity turns the request's community reference into the Group's +// AP id: URLs pass through, handles go through WebFinger. +func (a *Admin) resolveCommunity(ctx context.Context, ref string) (string, error) { + ref = strings.TrimSpace(ref) + if strings.HasPrefix(ref, "https://") || strings.HasPrefix(ref, "http://") { + return ref, nil + } + return a.client.ResolveHandle(ctx, ref) +} + +func (a *Admin) writeResolveError(w http.ResponseWriter, ref string, err error) { + switch { + case errors.IsValidation(err): + http.Error(w, err.Error(), http.StatusBadRequest) + case errors.IsNotFound(err): + http.Error(w, "community not found: "+ref, http.StatusNotFound) + default: + a.logger.Error("resolve community", "community", ref, "error", err) + http.Error(w, "failed to resolve community", http.StatusBadGateway) + } +} + +// buildFollow constructs the signed Follow activity (delivery signing +// happens in the client; this is the payload Lemmy validates). +func (a *Admin) buildFollow(groupIRI string) (*ap.Object, error) { + id, err := a.activityID("follow") + if err != nil { + return nil, err + } + return &ap.Object{ + Context: json.RawMessage(`"https://www.w3.org/ns/activitystreams"`), + ID: id, + Type: ap.TypeFollow, + Actor: &ap.Object{ID: a.service.ID}, + Object: &ap.Object{ID: groupIRI}, + To: ap.Audience{groupIRI}, + }, nil +} + +// buildUndoFollow constructs Undo{Follow}. Lemmy matches the undo by the +// inner Follow's actor and object. +func (a *Admin) buildUndoFollow(groupIRI string) (*ap.Object, error) { + inner, err := a.buildFollow(groupIRI) + if err != nil { + return nil, err + } + inner.Context = nil + id, err := a.activityID("undo") + if err != nil { + return nil, err + } + return &ap.Object{ + Context: json.RawMessage(`"https://www.w3.org/ns/activitystreams"`), + ID: id, + Type: ap.TypeUndo, + Actor: &ap.Object{ID: a.service.ID}, + Object: inner, + To: ap.Audience{groupIRI}, + }, nil +} + +// activityID mints a unique bridge-side activity id. A crypto/rand failure +// is propagated rather than swallowed: proceeding with the zero buffer would +// emit a constant, remote-deduped id (…/kind/000…0), and Lemmy dedupes by +// activity id — so the first such activity would land and every later +// Follow/Undo{Follow} would be silently ignored while the admin API still +// reports success. The buildFollow/buildUndoFollow callers map this to a 5xx. +func (a *Admin) activityID(kind string) (string, error) { + var buf [16]byte + if _, err := rand.Read(buf[:]); err != nil { + return "", fmt.Errorf("ingest: mint activity id: %w", err) + } + return fmt.Sprintf("https://%s/activities/%s/%s", a.service.Hostname, kind, hex.EncodeToString(buf[:])), nil +} + +// writeJSON writes a JSON response body. +func writeJSON(w http.ResponseWriter, status int, body any) { + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(status) + _ = json.NewEncoder(w).Encode(body) +} diff --git a/internal/ingest/follow_test.go b/internal/ingest/follow_test.go new file mode 100644 index 0000000..213ecab --- /dev/null +++ b/internal/ingest/follow_test.go @@ -0,0 +1,175 @@ +package ingest + +import ( + "bytes" + "context" + "encoding/json" + "net/http" + "net/http/httptest" + "os" + "path/filepath" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "tidepool/internal/ap" + "tidepool/internal/store" +) + +// Note (Finding K): activityID's crypto/rand failure path is not unit-tested. +// On Go 1.24+, crypto/rand.Read treats an entropy failure as a fatal, +// unrecoverable error (it never returns a non-nil error to the caller), so a +// failing-reader injection would crash the test binary rather than exercise +// the guard. Per the finding's guidance we keep the code-level guard (which +// propagates any returned error and, critically, stops emitting the constant +// zero-entropy id that Lemmy would dedupe) without contorting the design for +// this essentially-never path. The propagation wiring through +// buildFollow/buildUndoFollow into the 5xx handler paths is covered by the +// normal follow lifecycle tests below. + +// TestFollowLifecycle drives the full state machine: subscribe → pending → +// Accept → accepted (backfill triggered) → unsubscribe (Undo delivered) → +// none. Redelivered Accepts must not re-trigger backfill. +func TestFollowLifecycle(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() // asserts pending → accepted internally + ctx := context.Background() + + assert.Equal(t, 1, h.backfills.count(), "accept must trigger a backfill") + + // A re-delivered Accept (new activity id) is idempotent: state stays + // accepted, no second backfill. + status := h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/accept/follow-redelivered", + "type": "Accept", + "actor": groupID, + "object": map[string]any{"type": "Follow", "actor": h.service.ID, "object": groupID}, + }) + require.Equal(t, http.StatusAccepted, status) + h.drain() + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + assert.Equal(t, store.FollowStateAccepted, community.FollowState) + assert.Equal(t, 1, h.backfills.count(), "a redelivered accept must not re-trigger backfill") + + // Unsubscribe: Undo{Follow} delivered to the community inbox, state + // cleared. + deliveriesBefore := len(h.inboxLog) + rec := h.adminRequest(http.MethodDelete, "/admin/communities", + map[string]any{"community": groupID}) + require.Equal(t, http.StatusOK, rec.Code, rec.Body.String()) + + h.mu.Lock() + require.Greater(t, len(h.inboxLog), deliveriesBefore, "unsubscribe must deliver an Undo") + undoRaw := h.inboxLog[len(h.inboxLog)-1] + h.mu.Unlock() + undo, err := ap.ParseObject(undoRaw) + require.NoError(t, err) + assert.Equal(t, ap.TypeUndo, undo.Type) + require.NotNil(t, undo.Object) + assert.Equal(t, ap.TypeFollow, undo.Object.Type) + assert.Equal(t, groupID, undo.Object.Object.ID) + + community, err = h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + assert.Equal(t, store.FollowStateNone, community.FollowState) +} + +// TestFollowRejected: the community answers Reject{Follow} → state none. +func TestFollowRejected(t *testing.T) { + h := newHarness(t) + group := h.technologyGroup() + webfinger, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "webfinger_group.json")) + require.NoError(t, err) + h.serveJSON("/.well-known/webfinger", webfinger) + + rec := h.adminRequest(http.MethodPost, "/admin/communities", + map[string]any{"community": "!technology@lemmy.world"}) + require.Equal(t, http.StatusAccepted, rec.Code, rec.Body.String()) + + status := h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/reject/follow-1", + "type": "Reject", + "actor": groupID, + "object": map[string]any{"type": "Follow", "actor": h.service.ID, "object": groupID}, + }) + require.Equal(t, http.StatusAccepted, status) + h.drain() + + community, err := h.communities.GetByAPGroupID(context.Background(), groupID) + require.NoError(t, err) + assert.Equal(t, store.FollowStateNone, community.FollowState) + assert.Equal(t, 0, h.backfills.count(), "a rejected follow must not backfill") +} + +// TestAcceptFromWrongSignerDropped: an Accept for our Follow signed by an +// actor that is not the community must not flip the state. +func TestAcceptFromWrongSignerDropped(t *testing.T) { + h := newHarness(t) + h.technologyGroup() + webfinger, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "webfinger_group.json")) + require.NoError(t, err) + h.serveJSON("/.well-known/webfinger", webfinger) + + rec := h.adminRequest(http.MethodPost, "/admin/communities", + map[string]any{"community": "!technology@lemmy.world"}) + require.Equal(t, http.StatusAccepted, rec.Code, rec.Body.String()) + + // A different lemmy.world actor (same authority, passes the inbox + // binding) tries to accept on the community's behalf. + impostor := h.newRemoteActor("https://lemmy.world/u/impostor", + person("https://lemmy.world/u/impostor", "impostor", nil)) + status := h.deliver(impostor, map[string]any{ + "id": "https://lemmy.world/activities/accept/forged", + "type": "Accept", + "actor": impostor.id, + "object": map[string]any{"type": "Follow", "actor": h.service.ID, "object": groupID}, + }) + require.Equal(t, http.StatusAccepted, status) + h.drain() + + community, err := h.communities.GetByAPGroupID(context.Background(), groupID) + require.NoError(t, err) + assert.Equal(t, store.FollowStatePending, community.FollowState, + "only the community itself may accept its follow") +} + +// TestAdminAuth: /admin requires the bearer token. +func TestAdminAuth(t *testing.T) { + h := newHarness(t) + + body := map[string]any{"community": "!technology@lemmy.world"} + raw, err := json.Marshal(body) + require.NoError(t, err) + + // No token. + req := httptest.NewRequest(http.MethodGet, "https://"+bridgeHost+"/admin/communities", nil) + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + assert.Equal(t, http.StatusUnauthorized, rec.Code) + + // Wrong token. + req = httptest.NewRequest(http.MethodPost, "https://"+bridgeHost+"/admin/communities", bytes.NewReader(raw)) + req.Header.Set("Authorization", "Bearer wrong-token") + rec = httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + assert.Equal(t, http.StatusUnauthorized, rec.Code) +} + +// TestAdminList reports subscribed communities. +func TestAdminList(t *testing.T) { + h := newHarness(t) + h.subscribeTechnology() + + rec := h.adminRequest(http.MethodGet, "/admin/communities", nil) + require.Equal(t, http.StatusOK, rec.Code) + var out struct { + Communities []communityResponse `json:"communities"` + } + require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &out)) + require.Len(t, out.Communities, 1) + assert.Equal(t, groupID, out.Communities[0].Community) + assert.Equal(t, "accepted", out.Communities[0].FollowState) + assert.Equal(t, testDIDFor("technology", "lemmy.world"), out.Communities[0].DID) +} diff --git a/internal/ingest/handler.go b/internal/ingest/handler.go new file mode 100644 index 0000000..98356bd --- /dev/null +++ b/internal/ingest/handler.go @@ -0,0 +1,508 @@ +package ingest + +import ( + "context" + "fmt" + "log/slog" + + "tidepool/internal/ap" + "tidepool/internal/errors" + "tidepool/internal/materialize" + "tidepool/internal/store" +) + +// Materializer is the slice of *materialize.Materializer the dispatcher +// drives (task 05's entry points). +type Materializer interface { + MaterializePost(ctx context.Context, page *ap.Object) (*materialize.Result, error) + MaterializeComment(ctx context.Context, note *ap.Object) (*materialize.Result, error) + HandleUpdate(ctx context.Context, obj *ap.Object) (*materialize.Result, error) + HandleDelete(ctx context.Context, apID string) error + RefreshActor(ctx context.Context, actorRef *ap.Object) (*store.BridgedActor, error) + RefreshCommunity(ctx context.Context, groupRef *ap.Object) (*store.Community, error) + EnsureCommunity(ctx context.Context, groupRef *ap.Object) (*store.Community, error) +} + +// Fetcher is the slice of *ap.Client the dispatcher uses to re-fetch +// objects it must not trust from a delivery. +type Fetcher interface { + FetchObject(ctx context.Context, iri string) (*ap.Object, error) +} + +// Backfiller is notified when a community's Follow is accepted (the +// backfill trigger). *Backfill implements it; tests inject recorders. +type Backfiller interface { + TriggerAsync(community *store.Community, force bool) +} + +// HandlerOptions configures NewHandler. Materializer, Fetcher, Objects, +// Communities, Tombstones, Votes, and ServiceActorID are required; +// Backfill and Logger are optional. +type HandlerOptions struct { + Materializer Materializer + Fetcher Fetcher + Objects store.APObjects + Communities store.Communities + Tombstones store.Tombstones + Votes VoteAggregator + Backfill Backfiller + // ServiceActorID is the bridge's own AP actor id; Accepts must wrap a + // Follow issued by it. + ServiceActorID string + Logger *slog.Logger +} + +// Handler dispatches verified, deduplicated inbox activities to the +// materializer, the vote aggregator, and the follow state machine. It is +// the queue's processor: a nil or IsSkip return marks the event processed, +// a validation error poisons it, anything else is retried with backoff. +type Handler struct { + mat Materializer + fetcher Fetcher + objects store.APObjects + communities store.Communities + tombstones store.Tombstones + votes VoteAggregator + backfill Backfiller + serviceID string + logger *slog.Logger +} + +// NewHandler validates options and builds a Handler. +func NewHandler(opts HandlerOptions) (*Handler, error) { + if opts.Materializer == nil { + return nil, errors.NewValidationError("materializer", "must not be nil") + } + if opts.Fetcher == nil { + return nil, errors.NewValidationError("fetcher", "must not be nil") + } + if opts.Objects == nil { + return nil, errors.NewValidationError("objects", "must not be nil") + } + if opts.Communities == nil { + return nil, errors.NewValidationError("communities", "must not be nil") + } + if opts.Tombstones == nil { + return nil, errors.NewValidationError("tombstones", "must not be nil") + } + if opts.Votes == nil { + return nil, errors.NewValidationError("votes", "must not be nil") + } + if opts.ServiceActorID == "" { + return nil, errors.NewValidationError("service_actor_id", "must not be empty") + } + logger := opts.Logger + if logger == nil { + logger = slog.Default() + } + return &Handler{ + mat: opts.Materializer, + fetcher: opts.Fetcher, + objects: opts.Objects, + communities: opts.Communities, + tombstones: opts.Tombstones, + votes: opts.Votes, + backfill: opts.Backfill, + serviceID: opts.ServiceActorID, + logger: logger, + }, nil +} + +// Process handles one claimed queue event. The error contract mirrors the +// materializer's: nil or IsSkip → processed (skips are logged with their +// reason and never retried); IsValidation → poisoned; anything else → +// retryable. +func (h *Handler) Process(ctx context.Context, event *store.InboxEvent) error { + if len(event.Payload) == 0 { + return errors.NewValidationError("payload", "event carries no activity payload") + } + activity, err := ap.ParseObject(event.Payload) + if err != nil { + return errors.NewValidationError("payload", err.Error()) + } + // The inbox verified the HTTP signature and bound the activity's actor + // to the signer's authority; event.ActorID is that bound actor id. + signer := event.ActorID + if signer == "" { + return errors.NewValidationError("actor_id", "event carries no verified actor") + } + + switch activity.Type { + case ap.TypeAnnounce: + return h.handleAnnounce(ctx, activity, signer) + case ap.TypeCreate, ap.TypeUpdate: + return h.handleBareCreateUpdate(ctx, activity, signer) + case ap.TypeDelete: + return h.handleDelete(ctx, activity, signer, "") + case ap.TypeUndo: + return h.handleUndo(ctx, activity, signer, "") + case ap.TypeAccept: + return h.handleAccept(ctx, activity, signer) + case ap.TypeReject: + return h.handleReject(ctx, activity, signer) + case ap.TypeLike, ap.TypeDislike: + // Bare votes (rare; Lemmy normally announces them via the group). + return h.votes.ApplyVote(ctx, activity, "") + default: + return skip(activity.ID, "unsupported activity type "+activity.Type) + } +} + +// handleAnnounce unwraps FEP-1b12 group fan-out: Announce{Create|Update| +// Delete|Undo|Like|Dislike|...} from a community we follow. +func (h *Handler) handleAnnounce(ctx context.Context, announce *ap.Object, signer string) error { + // Only communities the bridge subscribed to may push content. pending is + // accepted too: Lemmy can start announcing before we processed its + // Accept (both arrive on the same ordering key, but a re-delivered + // Announce may overtake). + community, err := h.communities.GetByAPGroupID(ctx, signer) + if errors.IsNotFound(err) { + return skip(announce.ID, "announce from actor we do not follow: "+signer) + } + if err != nil { + return fmt.Errorf("ingest: look up announcing community %s: %w", signer, err) + } + if community.FollowState == store.FollowStateNone { + return skip(announce.ID, "announce from unfollowed community "+signer) + } + + inner := announce.Object + if inner == nil || (inner.ID == "" && inner.Type == "") { + return errors.NewValidationError("announce", "announce carries no object") + } + // A bare-IRI announce (object is just an id): fetch it. FetchObject + // fetches exactly inner.ID, and resolveDelivered below re-checks the + // body's self-asserted id, so cross-host forgery cannot slip in. + if inner.Type == "" { + fetched, err := h.fetchBound(ctx, inner.ID) + if err != nil { + return err + } + inner = fetched + } + + switch inner.Type { + case ap.TypeCreate: + return h.materializeContent(ctx, inner.Object, signer, false, signer) + case ap.TypeUpdate: + return h.materializeContent(ctx, inner.Object, signer, true, signer) + case ap.TypePage, ap.TypeArticle, ap.TypeNote: + // Some implementations announce the object itself, not the Create. + return h.materializeContent(ctx, inner, signer, false, signer) + case ap.TypeLike, ap.TypeDislike: + return h.votes.ApplyVote(ctx, inner, signer) + case ap.TypeDelete: + return h.handleDelete(ctx, inner, signer, signer) + case ap.TypeUndo: + return h.handleUndo(ctx, inner, signer, signer) + default: + // Lock, Add, Remove, Block, ... — moderation activities the bridge + // does not translate in v1. + return skip(announce.ID, "unsupported announced activity type "+inner.Type) + } +} + +// handleBareCreateUpdate processes a Create/Update delivered directly by a +// user (or community) actor rather than through group fan-out. +func (h *Handler) handleBareCreateUpdate(ctx context.Context, activity *ap.Object, signer string) error { + obj := activity.Object + if obj == nil || (obj.ID == "" && obj.Type == "") { + return errors.NewValidationError(activity.Type, "activity carries no object") + } + isUpdate := activity.Type == ap.TypeUpdate + return h.materializeContent(ctx, obj, signer, isUpdate, "") +} + +// materializeContent is the single content funnel: echo suppression, +// create-after-delete tombstones, embedded-object trust, followed-community +// checks, then the materializer. announcer is the announcing community's AP +// id ("" when the activity arrived bare). +func (h *Handler) materializeContent(ctx context.Context, obj *ap.Object, signer string, isUpdate bool, announcer string) error { + if obj == nil || obj.ID == "" { + return errors.NewValidationError("object", "content object carries no id") + } + + // Profile updates ride the same rails (Announce{Update{Group}}, bare + // Update{Person}) but have their own trust rule; nothing below applies. + if obj.Type == ap.TypePerson || obj.Type == ap.TypeGroup { + return h.applyProfileUpdate(ctx, obj, signer, announcer) + } + + // Echo suppression: an activity whose object the bridge itself emitted + // must never round-trip back in (write-side future-proofing; ap_objects + // carries the origin flag since task 01). + if mapping, err := h.objects.GetByAPID(ctx, obj.ID); err == nil { + if mapping.Origin == store.OriginBridge { + return skip(obj.ID, "echo of a bridge-authored object") + } + } else if !errors.IsNotFound(err) { + return fmt.Errorf("ingest: echo check for %s: %w", obj.ID, err) + } + + // Create-after-delete: a Delete for this id may have arrived before any + // materialization (no mapping to tombstone — task 05's known gap). The + // ap_tombstones marker closes it here, in the ingest layer. + tombstoned, err := h.tombstones.Exists(ctx, obj.ID) + if err != nil { + return fmt.Errorf("ingest: tombstone check for %s: %w", obj.ID, err) + } + if tombstoned { + return skip(obj.ID, "object was deleted upstream before it was ever materialized") + } + + obj, err = h.resolveDelivered(ctx, obj, signer) + if err != nil { + return err + } + + // Bare deliveries must belong to a community the bridge follows; the + // announce path already established that for its signer. + if announcer == "" { + communityIRI := communityIRIFrom(obj) + if communityIRI == "" { + return skip(obj.ID, "bare delivery names no community (no audience group IRI)") + } + community, err := h.communities.GetByAPGroupID(ctx, communityIRI) + if errors.IsNotFound(err) { + return skip(obj.ID, "bare delivery for a community we do not follow: "+communityIRI) + } + if err != nil { + return fmt.Errorf("ingest: look up community %s: %w", communityIRI, err) + } + if community.FollowState == store.FollowStateNone { + return skip(obj.ID, "bare delivery for unfollowed community "+communityIRI) + } + } else { + // Announced content must belong to the announcing community itself: a + // followed community may fan out only its own content, never inject + // into a DIFFERENT community's repo (even one co-hosted on the same + // instance). The materializer derives the target community from the + // object's own audience and EnsureCommunity()s it, so without this the + // announcer could name any community it likes. + if objCommunity := communityIRIFrom(obj); objCommunity != "" && objCommunity != announcer { + return skip(obj.ID, fmt.Sprintf( + "announced object names community %s but was announced by %s", objCommunity, announcer)) + } + } + + switch obj.Type { + case ap.TypePage, ap.TypeArticle: + if isUpdate { + _, err = h.mat.HandleUpdate(ctx, obj) + } else { + _, err = h.mat.MaterializePost(ctx, obj) + } + case ap.TypeNote: + if isUpdate { + _, err = h.mat.HandleUpdate(ctx, obj) + } else { + _, err = h.mat.MaterializeComment(ctx, obj) + } + default: + return skip(obj.ID, "unsupported content type "+obj.Type) + } + return err +} + +// resolveDelivered decides whether a delivered (embedded) object may be +// used as-is or must be re-fetched. The rule: an embedded copy is trusted +// only when its id lives on the signer's own authority — a community can +// vouch for content on its own instance, but content whose canonical id is +// on ANOTHER instance (a lemmy.zip post announced by a lemmy.world +// community, the normal federation case) is re-fetched from its origin so a +// malicious instance cannot forge bodies under a victim's id. Fetched +// bodies get the same self-asserted-id authority binding task 05 applies. +func (h *Handler) resolveDelivered(ctx context.Context, obj *ap.Object, signer string) (*ap.Object, error) { + if obj.Type != "" && ap.SameAuthority(obj.ID, signer) { + return obj, nil + } + return h.fetchBound(ctx, obj.ID) +} + +// fetchBound fetches an object by IRI and binds the body's self-asserted id +// to the fetch authority (empty ids inherit the request IRI). Unavailable +// and tombstoned objects are skips: content that cannot be verified at its +// origin is dropped, not retried. +func (h *Handler) fetchBound(ctx context.Context, iri string) (*ap.Object, error) { + fetched, err := h.fetcher.FetchObject(ctx, iri) + switch { + case err == nil: + case errors.IsTombstoned(err): + return nil, skip(iri, "object is tombstoned upstream") + case errors.IsNotFound(err): + return nil, skip(iri, "object is unavailable upstream") + default: + return nil, fmt.Errorf("ingest: fetch %s: %w", iri, err) + } + if fetched.ID == "" { + fetched.ID = iri + } else if !ap.SameAuthority(fetched.ID, iri) { + return nil, skip(iri, fmt.Sprintf("fetched object served a cross-authority id %s", fetched.ID)) + } + return fetched, nil +} + +// handleAccept marks a community's Follow accepted. Lemmy signs the Accept +// with the community actor itself, so the verified signer must BE the +// community the embedded Follow names. +func (h *Handler) handleAccept(ctx context.Context, accept *ap.Object, signer string) error { + communityID, err := h.followCommunity(ctx, accept, signer) + if err != nil { + return err + } + community, err := h.communities.GetByAPGroupID(ctx, communityID) + if errors.IsNotFound(err) { + return skip(accept.ID, "accept for a community we never followed: "+communityID) + } + if err != nil { + return fmt.Errorf("ingest: look up community %s: %w", communityID, err) + } + switch community.FollowState { + case store.FollowStateAccepted: + // Idempotent re-delivery: already accepted, no state change and no + // fresh backfill. + return nil + case store.FollowStateNone: + // We unsubscribed (state cleared, Undo{Follow} sent) — Lemmy can retry + // an Accept for hours. A late Accept must not silently re-subscribe us + // nor trigger a backfill. + return skip(accept.ID, "accept for a community we unfollowed: "+communityID) + } + // Only pending → accepted is a real transition (and the sole backfill + // trigger). + if err := h.communities.SetFollowState(ctx, communityID, store.FollowStateAccepted); err != nil { + return fmt.Errorf("ingest: mark follow accepted for %s: %w", communityID, err) + } + h.logger.Info("community follow accepted", "community", communityID) + if h.backfill != nil { + community.FollowState = store.FollowStateAccepted + h.backfill.TriggerAsync(community, false) + } + return nil +} + +// handleReject clears a community's Follow (the remote refused or revoked +// the subscription). +func (h *Handler) handleReject(ctx context.Context, reject *ap.Object, signer string) error { + communityID, err := h.followCommunity(ctx, reject, signer) + if err != nil { + return err + } + if _, err := h.communities.GetByAPGroupID(ctx, communityID); errors.IsNotFound(err) { + return skip(reject.ID, "reject for a community we never followed: "+communityID) + } else if err != nil { + return fmt.Errorf("ingest: look up community %s: %w", communityID, err) + } + if err := h.communities.SetFollowState(ctx, communityID, store.FollowStateNone); err != nil { + return fmt.Errorf("ingest: mark follow rejected for %s: %w", communityID, err) + } + h.logger.Warn("community follow rejected", "community", communityID) + return nil +} + +// followCommunity extracts and authorizes the community a Follow response +// (Accept/Reject) refers to: the embedded Follow must be ours (actor == the +// service actor) and the responding signer must be the community itself. +func (h *Handler) followCommunity(_ context.Context, response *ap.Object, signer string) (string, error) { + follow := response.Object + communityID := signer + if follow != nil && follow.Type == ap.TypeFollow { + if follow.Actor != nil && follow.Actor.ID != "" && follow.Actor.ID != h.serviceID { + return "", skip(response.ID, + "embedded follow was issued by "+follow.Actor.ID+", not the bridge") + } + if follow.Object != nil && follow.Object.ID != "" { + communityID = follow.Object.ID + } + } + if communityID != signer { + return "", skip(response.ID, fmt.Sprintf( + "follow response signed by %s for community %s (signer must be the community)", + signer, communityID)) + } + return communityID, nil +} + +// isBridged reports whether an AP id already has an ap_objects mapping — the +// "already bridged?" signal used to keep bare profile Updates refresh-only +// (never mint). Every bridged actor/community has a profile mapping row (rkey +// "self"), and every bridged post/comment a content mapping, so a hit means +// the id is known; a miss means it was never materialized. +func (h *Handler) isBridged(ctx context.Context, apID string) (bool, error) { + if _, err := h.objects.GetByAPID(ctx, apID); err == nil { + return true, nil + } else if errors.IsNotFound(err) { + return false, nil + } else { + return false, fmt.Errorf("ingest: check bridged state for %s: %w", apID, err) + } +} + +// targetIsBridgedActor reports whether an AP id is a bridged ACTOR (its +// mapping is a profile record), as opposed to a content object. A bridged +// actor always has a profile mapping (EnsureActor commits rkey "self"), so a +// harmful Delete(Actor) against a bridged victim is always detectable here; +// an unbridged actor is a materializer no-op regardless. +func (h *Handler) targetIsBridgedActor(ctx context.Context, apID string) (bool, error) { + mapping, err := h.objects.GetByAPID(ctx, apID) + if errors.IsNotFound(err) { + return false, nil + } + if err != nil { + return false, fmt.Errorf("ingest: classify delete target %s: %w", apID, err) + } + return mapping.Collection == materialize.CollectionActorProfile || + mapping.Collection == materialize.CollectionCommunityProfile, nil +} + +// refID returns the id of a possibly-nil object reference. +func refID(obj *ap.Object) string { + if obj == nil { + return "" + } + return obj.ID +} + +// communityIRIFrom finds the community Group IRI an object/activity is +// addressed to: Lemmy sets `audience` (FEP-1b12); older objects carry the +// group in to/cc next to the public collection (heuristic: Lemmy community +// IRIs live under /c/). +func communityIRIFrom(obj *ap.Object) string { + for _, iri := range obj.Audience { + if iri != "" && !isPublicIRI(iri) { + return iri + } + } + for _, list := range []ap.Audience{obj.To, obj.Cc} { + for _, iri := range list { + if iri == "" || isPublicIRI(iri) { + continue + } + if containsCommunityPath(iri) { + return iri + } + } + } + return "" +} + +func isPublicIRI(iri string) bool { + return iri == ap.PublicAudience || iri == "as:Public" || iri == "Public" +} + +func containsCommunityPath(iri string) bool { + // Lemmy community IRIs are https://host/c/name; keep the same heuristic + // the materializer uses (Mbin /m/ deferred with it). + for i := 0; i+3 <= len(iri); i++ { + if iri[i] == '/' && iri[i+1] == 'c' && iri[i+2] == '/' { + return true + } + } + return false +} + +// skip builds the shared log-and-never-retry error (the materializer's +// SkipError, so the queue's IsSkip check covers both layers). +func skip(apID, reason string) error { + return &materialize.SkipError{APID: apID, Reason: reason} +} diff --git a/internal/ingest/handler_test.go b/internal/ingest/handler_test.go new file mode 100644 index 0000000..959d381 --- /dev/null +++ b/internal/ingest/handler_test.go @@ -0,0 +1,795 @@ +package ingest + +import ( + "context" + "net/http" + "strings" + "testing" + + "github.com/ipfs/go-cid" + "github.com/multiformats/go-multihash" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "tidepool/internal/ap" + "tidepool/internal/errors" + "tidepool/internal/materialize" + "tidepool/internal/store" +) + +// mintCount returns how many identities the fake minter has minted so far. +func (f *fakeMinter) mintCount() int { + f.mu.Lock() + defer f.mu.Unlock() + return f.mints +} + +// followedCommunity registers a second subscribed community (accepted) on an +// arbitrary host that can sign inbox deliveries — the attacker's community in +// the announced-authorization tests. It skips the WebFinger/Follow lifecycle +// and writes the communities row directly. +func (h *harness) followedCommunity(id, username, instance string) *remoteActor { + h.t.Helper() + actor := h.newRemoteActor(id, map[string]any{ + "type": "Group", + "id": id, + "preferredUsername": username, + "inbox": id + "/inbox", + "published": "2024-01-01T00:00:00.000000Z", + }) + ctx := context.Background() + _, err := h.communities.UpsertCommunity(ctx, store.Community{ + APGroupID: id, + DID: testDIDFor(username, instance), + PreferredUsername: username, + Instance: instance, + }) + require.NoError(h.t, err) + require.NoError(h.t, h.communities.SetFollowState(ctx, id, store.FollowStateAccepted)) + return actor +} + +// TestAnnounceCreatePageEndToEnd is the definition-of-done flow: subscribe +// → Accept → Announce{Create{Page}} delivered by the fake Lemmy → the post +// record is visible on the task-04 firehose. +func TestAnnounceCreatePageEndToEnd(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + + status := h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json")) + require.Equal(t, http.StatusAccepted, status) + h.drain() + + // The event processed cleanly. + event, err := h.events.GetEvent(context.Background(), + "https://lemmy.world/activities/announce/create/6a91b0d9-c1e5-45d6-be8b-ce248306867e") + require.NoError(t, err) + require.NotNil(t, event.ProcessedAt) + assert.Empty(t, event.Error) + + // The post is mapped and lives in the community repo. + mapping, err := h.objects.GetByAPID(context.Background(), pageID) + require.NoError(t, err) + communityDID := testDIDFor("technology", "lemmy.world") + assert.Equal(t, communityDID, mapping.DID) + assert.Equal(t, materialize.CollectionPost, mapping.Collection) + + // ... and is visible via the firehose (task 04 reads the same log). + var postOps int + for _, path := range h.firehoseOps() { + if strings.HasPrefix(path, materialize.CollectionPost+"/") { + postOps++ + } + } + assert.Equal(t, 1, postOps, "the post commit must be on the firehose") +} + +// TestAnnounceFromUnfollowedCommunityDropped: content pushed by an actor +// the bridge never subscribed to is skipped, not materialized. +func TestAnnounceFromUnfollowedCommunityDropped(t *testing.T) { + h := newHarness(t) + // The group actor exists (and can sign), but no subscription happened. + group := h.technologyGroup() + h.serveLemmyWorldContent() + + status := h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json")) + require.Equal(t, http.StatusAccepted, status) + h.drain() + + event, err := h.events.GetEvent(context.Background(), + "https://lemmy.world/activities/announce/create/6a91b0d9-c1e5-45d6-be8b-ce248306867e") + require.NoError(t, err) + assert.NotNil(t, event.ProcessedAt, "skips are processed, not retried") + _, err = h.objects.GetByAPID(context.Background(), pageID) + assert.True(t, errors.IsNotFound(err), "unfollowed content must not be materialized") +} + +// TestOutOfOrderComments: a child comment arrives before its parent — the +// parent-fetch path materializes the whole chain, and the later parent +// delivery is an idempotent no-op. Both comments land. +func TestOutOfOrderComments(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + + parentID := "https://lemmy.world/comment/1001" + childID := "https://lemmy.world/comment/1002" + author := "https://lemmy.world/u/threadster" + h.serveObject("/u/threadster", person(author, "threadster", nil)) + parentDoc := note(parentID, author, pageID, "parent comment", "2026-07-07T04:00:00.000000Z") + childDoc := note(childID, author, parentID, "child comment", "2026-07-07T04:05:00.000000Z") + h.serveObject("/comment/1001", parentDoc) + h.serveObject("/comment/1002", childDoc) + + announce := func(id string, inner map[string]any) map[string]any { + return map[string]any{ + "id": id, + "type": "Announce", + "actor": groupID, + "to": []any{ap.PublicAudience}, + "audience": groupID, + "object": map[string]any{ + "id": id + "/create", + "type": "Create", + "actor": author, + "audience": groupID, + "object": inner, + }, + } + } + + // Child first (out of order): the ancestor walk fetches the parent + // comment and the page. + require.Equal(t, http.StatusAccepted, + h.deliver(group, announce("https://lemmy.world/activities/announce/child", childDoc))) + // Parent second. + require.Equal(t, http.StatusAccepted, + h.deliver(group, announce("https://lemmy.world/activities/announce/parent", parentDoc))) + h.drain() + + ctx := context.Background() + parentMapping, err := h.objects.GetByAPID(ctx, parentID) + require.NoError(t, err, "parent must land via the ancestor-fetch path") + childMapping, err := h.objects.GetByAPID(ctx, childID) + require.NoError(t, err) + pageMapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err, "the thread's page is materialized as the root") + + // The child's reply refs point at the parent and root at the page. + record, _, err := h.manager.GetRecord(ctx, childMapping.DID, childMapping.Collection, childMapping.RKey) + require.NoError(t, err) + reply, ok := record["reply"].(map[string]any) + require.True(t, ok) + parentRef, _ := reply["parent"].(map[string]any) + rootRef, _ := reply["root"].(map[string]any) + assert.Equal(t, parentMapping.ATURI, parentRef["uri"]) + assert.Equal(t, pageMapping.ATURI, rootRef["uri"]) + + // Both processed, exactly one comment record each (idempotent re-put). + var commentOps int + for _, path := range h.firehoseOps() { + if strings.HasPrefix(path, materialize.CollectionComment+"/") { + commentOps++ + } + } + assert.Equal(t, 2, commentOps, "each comment commits exactly once") +} + +// TestAnnounceLikeRoutedToVotes: Announce{Like} goes to the vote-aggregator +// seam (task 07), not the materializer. +func TestAnnounceLikeRoutedToVotes(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + + status := h.deliver(group, loadFixture(t, "announce_like.json")) + require.Equal(t, http.StatusAccepted, status) + h.drain() + + require.Len(t, h.votes.applied, 1) + assert.Equal(t, "Like "+pageID, h.votes.applied[0]) +} + +// TestEchoSuppression: an inbound activity whose object maps to a record +// the bridge itself created (origin=bridge) is dropped. +func TestEchoSuppression(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + // Simulate a future write-side record the bridge emitted for this AP id. + sum, err := multihash.Sum([]byte("bridge-emitted record"), multihash.SHA2_256, -1) + require.NoError(t, err) + communityDID := testDIDFor("technology", "lemmy.world") + _, err = h.objects.PutMapping(ctx, store.APObjectMapping{ + APID: pageID, + APType: "Page", + OriginInstance: "lemmy.world", + Origin: store.OriginBridge, + DID: communityDID, + Collection: materialize.CollectionPost, + RKey: "3lbzzzzzzzz2a", + CID: cid.NewCidV1(cid.DagCBOR, sum).String(), + }) + require.NoError(t, err) + + opsBefore := len(h.firehoseOps()) + status := h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json")) + require.Equal(t, http.StatusAccepted, status) + h.drain() + + event, err := h.events.GetEvent(ctx, + "https://lemmy.world/activities/announce/create/6a91b0d9-c1e5-45d6-be8b-ce248306867e") + require.NoError(t, err) + assert.NotNil(t, event.ProcessedAt, "echoes are processed (skipped), not retried") + assert.Equal(t, opsBefore, len(h.firehoseOps()), "an echo must not produce commits") + + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.Equal(t, store.OriginBridge, mapping.Origin, "the bridge-origin mapping is untouched") +} + +// TestCreateAfterDeleteTombstone closes the task-05 gap: a Delete arriving +// for a never-materialized object must prevent a later Create from +// resurrecting it; Undo{Delete} restores. +func TestCreateAfterDeleteTombstone(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + author := h.newRemoteActor(personID, person(personID, "LeftLeaningFreedomFighters", nil)) + ctx := context.Background() + + // 1. Delete arrives FIRST (out of order), for an object never seen. + deleteActivity := loadFixture(t, "delete_page.json") + require.Equal(t, http.StatusAccepted, h.deliver(author, deleteActivity)) + h.drain() + + tombstoned, err := h.tombstones.Exists(ctx, pageID) + require.NoError(t, err) + assert.True(t, tombstoned, "a delete of an unseen id must leave a tombstone marker") + + // 2. The Create arrives late: it must NOT materialize. + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + _, err = h.objects.GetByAPID(ctx, pageID) + assert.True(t, errors.IsNotFound(err), "create-after-delete must not resurrect the object") + + // 3. Undo{Delete}: the origin restored the post; the bridge re-fetches + // and materializes it. + require.Equal(t, http.StatusAccepted, h.deliver(author, map[string]any{ + "id": "https://lemmy.world/activities/undo/delete-1", + "type": "Undo", + "actor": personID, + "object": deleteActivity, + })) + h.drain() + + tombstoned, err = h.tombstones.Exists(ctx, pageID) + require.NoError(t, err) + assert.False(t, tombstoned, "undo must clear the tombstone marker") + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err, "the restored object must be materialized") + assert.False(t, mapping.IsDeleted()) +} + +// TestDeleteMaterializedObject: the normal delete flow — record removed, +// mapping soft-deleted, re-delivered Create cannot resurrect it. +func TestDeleteMaterializedObject(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + author := h.newRemoteActor(personID, person(personID, "LeftLeaningFreedomFighters", nil)) + ctx := context.Background() + + announce := loadFixture(t, "announce_create_page_lemmy_world.json") + require.Equal(t, http.StatusAccepted, h.deliver(group, announce)) + h.drain() + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + + require.Equal(t, http.StatusAccepted, h.deliver(author, loadFixture(t, "delete_page.json"))) + h.drain() + + mapping, err = h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.True(t, mapping.IsDeleted(), "the mapping must be soft-deleted") + _, _, err = h.manager.GetRecord(ctx, mapping.DID, mapping.Collection, mapping.RKey) + assert.Error(t, err, "the record must be deleted from the repo") + + // A re-delivered Create (new activity id, same object) must not revive it. + announce["id"] = "https://lemmy.world/activities/announce/create/redelivery" + require.Equal(t, http.StatusAccepted, h.deliver(group, announce)) + h.drain() + mapping, err = h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.True(t, mapping.IsDeleted(), "a re-delivered create must not resurrect deleted content") +} + +// TestCrossAuthorityDeleteDropped: a bare Delete signed by an actor on a +// different instance than the object is unauthorized and dropped. +func TestCrossAuthorityDeleteDropped(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + mallory := h.newRemoteActor("https://evil.example/u/mallory", person("https://evil.example/u/mallory", "mallory", nil)) + ctx := context.Background() + + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + + require.Equal(t, http.StatusAccepted, h.deliver(mallory, map[string]any{ + "id": "https://evil.example/activities/delete/1", + "type": "Delete", + "actor": mallory.id, + "object": pageID, + })) + h.drain() + + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.False(t, mapping.IsDeleted(), "a cross-authority delete must be dropped") +} + +// TestNobridgeCommentDropped is the DoD consent check: a comment whose +// author opted out via #nobridge is dropped with the reason logged (the +// skip), and no identity is minted for them. +func TestNobridgeCommentDropped(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + optOutID := "https://lemmy.world/u/optout" + h.serveObject("/u/optout", person(optOutID, "optout", map[string]any{ + "summary": "

please leave me alone #nobridge

", + })) + commentID := "https://lemmy.world/comment/2001" + commentDoc := note(commentID, optOutID, pageID, "my private opinion", "2026-07-07T06:00:00.000000Z") + h.serveObject("/comment/2001", commentDoc) + + require.Equal(t, http.StatusAccepted, h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/announce/nobridge-comment", + "type": "Announce", + "actor": groupID, + "audience": groupID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/create/nobridge-comment", + "type": "Create", + "actor": optOutID, + "audience": groupID, + "object": commentDoc, + }, + })) + h.drain() + + event, err := h.events.GetEvent(ctx, "https://lemmy.world/activities/announce/nobridge-comment") + require.NoError(t, err) + assert.NotNil(t, event.ProcessedAt, "the skip is logged and the event completed") + + _, err = h.objects.GetByAPID(ctx, commentID) + assert.True(t, errors.IsNotFound(err), "the comment must not be materialized") + _, err = h.actors.GetByAPActorID(ctx, optOutID) + assert.True(t, errors.IsNotFound(err), "no identity is minted for an opted-out actor") +} + +// TestConsentTransitionOnProfileUpdate: a bridged author adds #nobridge; +// the self-signed Update{Person} scrubs their content and flips consent +// (reversible nobridge, not deleted). +func TestConsentTransitionOnProfileUpdate(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + author := h.newRemoteActor(personID, person(personID, "LeftLeaningFreedomFighters", nil)) + ctx := context.Background() + + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + actorRow, err := h.actors.GetByAPActorID(ctx, personID) + require.NoError(t, err) + require.Equal(t, store.ConsentStateOK, actorRow.ConsentState) + + // The author edits their profile to carry #nobridge; their instance + // delivers Update{Person} signed by the author. + updated := person(personID, "LeftLeaningFreedomFighters", map[string]any{ + "summary": "

done with bridges #nobridge

", + }) + require.Equal(t, http.StatusAccepted, h.deliver(author, map[string]any{ + "id": "https://lemmy.world/activities/update/person-1", + "type": "Update", + "actor": personID, + "object": updated, + })) + h.drain() + + actorRow, err = h.actors.GetByAPActorID(ctx, personID) + require.NoError(t, err) + assert.Equal(t, store.ConsentStateNoBridge, actorRow.ConsentState, + "profile update with #nobridge must flip consent (reversibly)") + + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.True(t, mapping.IsDeleted(), "existing content must be scrubbed on opt-out") +} + +// TestDeleteActorTombstonesRepo: Delete(Actor) scrubs and terminally +// tombstones the bridged actor. +func TestDeleteActorTombstonesRepo(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + author := h.newRemoteActor(personID, person(personID, "LeftLeaningFreedomFighters", nil)) + ctx := context.Background() + + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + + require.Equal(t, http.StatusAccepted, h.deliver(author, map[string]any{ + "id": "https://lemmy.world/activities/delete/actor-1", + "type": "Delete", + "actor": personID, + "object": personID, + })) + h.drain() + + actorRow, err := h.actors.GetByAPActorID(ctx, personID) + require.NoError(t, err) + assert.Equal(t, store.ConsentStateDeleted, actorRow.ConsentState, "Delete(Actor) is terminal") + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.True(t, mapping.IsDeleted(), "the deleted actor's content is scrubbed") +} + +// TestForgedEmbeddedContentRefetched: content embedded in an announce whose +// id lives on ANOTHER instance is re-fetched from its origin — the +// delivered (potentially forged) body is never trusted. +func TestForgedEmbeddedContentRefetched(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + // The announce embeds a lemmy.zip comment with FORGED content; the real + // note (served by lemmy.zip) says something else. Chain fixtures for the + // ancestor walk: parent note on sh.itjust.works replying to the page. + noteDoc := loadFixture(t, "note_lemmy_zip.json") + h.serveObject("/comment/27485395", noteDoc) + h.serveObject("/u/tixooo", person(noteAuthor, "tixooo", nil)) + h.serveObject("/comment/26248018", + note(parentNote, parentActor, pageID, "the parent comment", "2026-07-07T04:30:00.000000Z")) + h.serveObject("/u/DemandtheOxfordComma", person(parentActor, "DemandtheOxfordComma", nil)) + + announce := loadFixture(t, "announce_create_note.json") + forged := announce["object"].(map[string]any)["object"].(map[string]any) + forged["content"] = "

FORGED BODY

" + forged["source"] = map[string]any{"content": "FORGED BODY", "mediaType": "text/markdown"} + + require.Equal(t, http.StatusAccepted, h.deliver(group, announce)) + h.drain() + + mapping, err := h.objects.GetByAPID(ctx, noteID) + require.NoError(t, err, "the comment must land (from its origin)") + record, _, err := h.manager.GetRecord(ctx, mapping.DID, mapping.Collection, mapping.RKey) + require.NoError(t, err) + content, _ := record["content"].(string) + assert.NotContains(t, content, "FORGED", "the forged embedded body must be discarded") + assert.Contains(t, content, "theme for all of human history", "the origin's body is used") +} + +// TestAnnouncedDeleteOfOwnPost (Finding 1, positive): a followed community +// announcing a Delete of its OWN post (same host, same authority) is +// authorized and soft-deletes the record. +func TestAnnouncedDeleteOfOwnPost(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + require.False(t, mapping.IsDeleted()) + + require.Equal(t, http.StatusAccepted, h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/announce/delete-own", + "type": "Announce", + "actor": groupID, + "audience": groupID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/delete/own-post", + "type": "Delete", + "actor": groupID, + "object": pageID, + }, + })) + h.drain() + + mapping, err = h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.True(t, mapping.IsDeleted(), "a community may delete its own announced post") +} + +// TestCrossAuthorityAnnouncedDeleteDropped (Finding 1, negative): a followed +// community on one instance cannot delete content hosted on ANOTHER instance +// by wrapping the Delete in an Announce. The old code skipped all authority +// checks whenever announcer != "". +func TestCrossAuthorityAnnouncedDeleteDropped(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + + // A different subscribed community, on evil.example, announces a Delete of + // the lemmy.world post — content it does not host. + evil := h.followedCommunity("https://evil.example/c/foo", "foo", "evil.example") + require.Equal(t, http.StatusAccepted, h.deliver(evil, map[string]any{ + "id": "https://evil.example/activities/announce/delete-victim", + "type": "Announce", + "actor": "https://evil.example/c/foo", + "audience": "https://evil.example/c/foo", + "object": map[string]any{ + "id": "https://evil.example/activities/delete/victim", + "type": "Delete", + "actor": "https://evil.example/c/foo", + "object": pageID, + }, + })) + h.drain() + + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.False(t, mapping.IsDeleted(), "a community must not delete another instance's content") + tombstoned, err := h.tombstones.Exists(ctx, pageID) + require.NoError(t, err) + assert.False(t, tombstoned, "an unauthorized announced delete must not record a tombstone") +} + +// TestAnnouncedActorDeleteOfCoHostedActorDropped (Finding 1, actor path): even +// on its own authority, a community may delete only ITSELF, never a co-hosted +// OTHER actor whose bridged presence spans other communities (the terminal +// DeleteActor scrub). +func TestAnnouncedActorDeleteOfCoHostedActorDropped(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + // The post author is bridged as a co-hosted actor on lemmy.world. + require.Equal(t, http.StatusAccepted, + h.deliver(group, loadFixture(t, "announce_create_page_lemmy_world.json"))) + h.drain() + actorRow, err := h.actors.GetByAPActorID(ctx, personID) + require.NoError(t, err) + require.Equal(t, store.ConsentStateOK, actorRow.ConsentState) + + require.Equal(t, http.StatusAccepted, h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/announce/delete-cohost", + "type": "Announce", + "actor": groupID, + "audience": groupID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/delete/cohost-actor", + "type": "Delete", + "actor": groupID, + "object": personID, + }, + })) + h.drain() + + actorRow, err = h.actors.GetByAPActorID(ctx, personID) + require.NoError(t, err) + assert.Equal(t, store.ConsentStateOK, actorRow.ConsentState, + "a community must not tombstone a co-hosted actor via announce") + mapping, err := h.objects.GetByAPID(ctx, pageID) + require.NoError(t, err) + assert.False(t, mapping.IsDeleted(), "the co-hosted actor's content stays live") +} + +// TestBareProfileUpdateForUnknownActorDropped (Finding 2): a self-signed evil +// actor sends a bare Update{Person} naming an unknown target IRI. The old code +// reached RefreshActor, which fetched the target (SSRF/fetch-oracle) and minted +// a PLC DID. The refresh-only gate drops it before any fetch or mint. +func TestBareProfileUpdateForUnknownActorDropped(t *testing.T) { + h := newHarness(t) + ctx := context.Background() + + evil := h.newRemoteActor("https://evil.example/u/mallory", + person("https://evil.example/u/mallory", "mallory", nil)) + + // The attacker-chosen target: an unknown actor on a victim host, served + // with a hit counter so we can prove it is never fetched. + target := "https://victim.example/u/target" + h.serveObject("/u/target", person(target, "target", nil)) + mintsBefore := h.minter.mintCount() + + require.Equal(t, http.StatusAccepted, h.deliver(evil, map[string]any{ + "id": "https://evil.example/activities/update/ssrf", + "type": "Update", + "actor": evil.id, + "object": person(target, "target", nil), + })) + h.drain() + + assert.Equal(t, 0, h.hitCount("/u/target"), "the unknown target IRI must never be fetched") + assert.Equal(t, mintsBefore, h.minter.mintCount(), "no identity minted for an unknown actor") + _, err := h.actors.GetByAPActorID(ctx, target) + assert.True(t, errors.IsNotFound(err), "the unknown target must not be bridged") +} + +// TestAnnouncedObjectForDifferentCommunityDropped (Finding 3): a followed +// community may fan out only its OWN content. An Announce whose embedded object +// names a different community (even same-host) must not inject into that +// community's repo. +func TestAnnouncedObjectForDifferentCommunityDropped(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + ctx := context.Background() + + otherCommunity := "https://lemmy.world/c/otherthing" + injectedID := "https://lemmy.world/post/88888" + author := "https://lemmy.world/u/injector" + h.serveObject("/u/injector", person(author, "injector", nil)) + page := map[string]any{ + "type": "Page", + "id": injectedID, + "attributedTo": author, + "to": []any{ap.PublicAudience}, + "audience": otherCommunity, // NOT the announcing community + "name": "injected post", + "source": map[string]any{"content": "injected", "mediaType": "text/markdown"}, + "published": "2026-07-07T07:00:00.000000Z", + } + h.serveObject("/post/88888", page) + + require.Equal(t, http.StatusAccepted, h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/announce/inject", + "type": "Announce", + "actor": groupID, + "audience": groupID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/create/inject", + "type": "Create", + "actor": author, + "object": page, + }, + })) + h.drain() + + _, err := h.objects.GetByAPID(ctx, injectedID) + assert.True(t, errors.IsNotFound(err), "content for a non-announced community must be dropped") + _, err = h.communities.GetByAPGroupID(ctx, otherCommunity) + assert.True(t, errors.IsNotFound(err), "the other community must not be bridged") +} + +// TestUndoDeleteRollsBackWhenRematerializeSkips (Finding 4): if the restore's +// re-materialization is declined (a skip), the compensation must re-soft-delete +// the mapping and re-record the tombstone — never leave a live mapping without +// a record. +func TestUndoDeleteRollsBackWhenRematerializeSkips(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + h.serveLemmyWorldContent() + author := h.newRemoteActor(personID, person(personID, "LeftLeaningFreedomFighters", nil)) + ctx := context.Background() + + postID := "https://lemmy.world/post/70001" + buildPage := func(withCommunity bool) map[string]any { + p := map[string]any{ + "type": "Page", + "id": postID, + "attributedTo": personID, + "to": []any{ap.PublicAudience}, + "name": "a post", + "source": map[string]any{"content": "body", "mediaType": "text/markdown"}, + "published": "2026-07-07T08:00:00.000000Z", + } + if withCommunity { + p["audience"] = groupID + } + return p + } + h.serveObject("/post/70001", buildPage(true)) + + require.Equal(t, http.StatusAccepted, h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/announce/create-70001", + "type": "Announce", + "actor": groupID, + "audience": groupID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/create/70001", + "type": "Create", + "actor": personID, + "audience": groupID, + "object": buildPage(true), + }, + })) + h.drain() + mapping, err := h.objects.GetByAPID(ctx, postID) + require.NoError(t, err) + require.False(t, mapping.IsDeleted()) + + // Delete it (bare, on the author's own authority). + require.Equal(t, http.StatusAccepted, h.deliver(author, map[string]any{ + "id": "https://lemmy.world/activities/delete/70001", + "type": "Delete", + "actor": personID, + "object": postID, + })) + h.drain() + mapping, err = h.objects.GetByAPID(ctx, postID) + require.NoError(t, err) + require.True(t, mapping.IsDeleted()) + tombstoned, err := h.tombstones.Exists(ctx, postID) + require.NoError(t, err) + require.True(t, tombstoned) + + // The origin re-serves the post but now WITHOUT a community, so + // re-materialization skips ("post names no community"). + h.serveObject("/post/70001", buildPage(false)) + + require.Equal(t, http.StatusAccepted, h.deliver(author, map[string]any{ + "id": "https://lemmy.world/activities/undo/70001", + "type": "Undo", + "actor": personID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/delete/70001", + "type": "Delete", + "actor": personID, + "object": postID, + }, + })) + h.drain() + + mapping, err = h.objects.GetByAPID(ctx, postID) + require.NoError(t, err) + assert.True(t, mapping.IsDeleted(), "a declined restore must re-soft-delete the mapping") + tombstoned, err = h.tombstones.Exists(ctx, postID) + require.NoError(t, err) + assert.True(t, tombstoned, "a declined restore must retain the tombstone") +} + +// TestLateAcceptAfterUnfollowIgnored (Finding 5): after an operator +// unsubscribes (state → none), a late/re-delivered Accept must not flip the +// community back to accepted nor trigger a fresh backfill. +func TestLateAcceptAfterUnfollowIgnored(t *testing.T) { + h := newHarness(t) + group := h.subscribeTechnology() + ctx := context.Background() + + // The real Accept during subscribe triggered exactly one backfill. + require.Equal(t, 1, h.backfills.count()) + + // The operator unsubscribes. + require.NoError(t, h.communities.SetFollowState(ctx, groupID, store.FollowStateNone)) + + // Lemmy re-delivers a late Accept for the follow we already withdrew. + require.Equal(t, http.StatusAccepted, h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/accept/follow-late", + "type": "Accept", + "actor": groupID, + "object": map[string]any{ + "id": "https://lemmy.world/activities/follow/late", + "type": "Follow", + "actor": h.service.ID, + "object": groupID, + }, + })) + h.drain() + + community, err := h.communities.GetByAPGroupID(ctx, groupID) + require.NoError(t, err) + assert.Equal(t, store.FollowStateNone, community.FollowState, + "a late Accept must not re-subscribe an unfollowed community") + assert.Equal(t, 1, h.backfills.count(), "no fresh backfill for an unfollowed community") +} diff --git a/internal/ingest/inbox.go b/internal/ingest/inbox.go new file mode 100644 index 0000000..0378e8b --- /dev/null +++ b/internal/ingest/inbox.go @@ -0,0 +1,277 @@ +// Package ingest wires the fediverse to the materializer: the shared AP +// inbox (HTTP-signature verified, deduplicated, durably queued), the +// FEP-1b12 Announce dispatcher, the community Follow lifecycle, outbox +// backfill, and consent enforcement. After this package the bridge is +// functionally complete end-to-end (minus vote aggregates, task 07). +package ingest + +import ( + "crypto/subtle" + "encoding/json" + "fmt" + "io" + "log/slog" + "net/http" + + "github.com/go-chi/chi/v5" + + "tidepool/internal/ap" + "tidepool/internal/errors" + "tidepool/internal/store" +) + +// maxInboxBodyBytes caps inbound activity payloads. Lemmy activities are a +// few KB; anything approaching a megabyte is abuse. +const maxInboxBodyBytes = 1 << 20 + +// softwareName is what nodeinfo reports; Lemmy admins allowlist by this +// name. +const ( + softwareName = "tidepool" + softwareVersion = "0.1.0" +) + +// InboxOptions configures NewInbox. Verifier, Events, Service, and Queue +// are required. +type InboxOptions struct { + // Verifier checks inbound HTTP signatures (task 02). + Verifier *ap.Verifier + // Events is the dedupe + queue store. + Events store.InboxEvents + // Queue is nudged after every enqueue. + Queue *Queue + // Service is the bridge's AP service actor (document served at /actor). + Service *ap.ServiceActor + Logger *slog.Logger +} + +// Inbox is the HTTP face of ingestion: POST /inbox (+ the actor inbox +// alias), the service actor document, WebFinger for it, and nodeinfo. +type Inbox struct { + verifier *ap.Verifier + events store.InboxEvents + queue *Queue + service *ap.ServiceActor + logger *slog.Logger +} + +// NewInbox validates options and builds the Inbox. +func NewInbox(opts InboxOptions) (*Inbox, error) { + if opts.Verifier == nil { + return nil, errors.NewValidationError("verifier", "must not be nil") + } + if opts.Events == nil { + return nil, errors.NewValidationError("events", "must not be nil") + } + if opts.Queue == nil { + return nil, errors.NewValidationError("queue", "must not be nil") + } + if opts.Service == nil { + return nil, errors.NewValidationError("service", "must not be nil") + } + logger := opts.Logger + if logger == nil { + logger = slog.Default() + } + return &Inbox{ + verifier: opts.Verifier, + events: opts.Events, + queue: opts.Queue, + service: opts.Service, + logger: logger, + }, nil +} + +// Routes mounts the AP-facing surface on a chi router. +func (ib *Inbox) Routes(r chi.Router) { + r.Post("/inbox", ib.handleInbox) + // The service actor's own inbox: Lemmy delivers Accepts wherever the + // Follow actor's document says; both spellings land here. + r.Post("/actor/inbox", ib.handleInbox) + r.Get(ap.ServiceActorPath, ib.handleActor) + r.Get("/.well-known/webfinger", ib.handleWebFinger) + r.Get("/.well-known/nodeinfo", ib.handleNodeInfoDiscovery) + r.Get("/nodeinfo/2.0", ib.handleNodeInfo) +} + +// handleInbox receives one AP delivery: verify the HTTP signature, bind the +// activity's actor to the signer, dedupe by activity id, enqueue for the +// worker pool, 202. Everything heavier happens async — remote instances +// time deliveries and treat slow inboxes as dead. +func (ib *Inbox) handleInbox(w http.ResponseWriter, r *http.Request) { + body, err := io.ReadAll(http.MaxBytesReader(w, r.Body, maxInboxBodyBytes)) + if err != nil { + http.Error(w, "request body too large or unreadable", http.StatusRequestEntityTooLarge) + return + } + + actorID, err := ib.verifier.Verify(r.Context(), r, body) + if err != nil { + if errors.IsValidation(err) { + // Bad/expired/missing signature: a definitive rejection. + ib.logger.Warn("inbox delivery rejected: signature", "error", err) + http.Error(w, "signature verification failed", http.StatusUnauthorized) + return + } + // Key resolution failed (network, remote down): tell the sender to + // retry rather than swallowing the delivery. + ib.logger.Warn("inbox delivery deferred: key resolution", "error", err) + http.Error(w, "signature key unavailable", http.StatusServiceUnavailable) + return + } + + activity, err := ap.ParseObject(body) + if err != nil { + http.Error(w, "payload is not an AP activity", http.StatusBadRequest) + return + } + if activity.ID == "" || activity.Type == "" { + http.Error(w, "activity must carry id and type", http.StatusBadRequest) + return + } + + // Bind the activity's claimed actor to the verified signer. Exact + // equality is the common case (Lemmy signs as the acting actor); same + // authority tolerates instance-actor signing (Mastodon secure-mode + // relays) without letting host A speak for host B. The QUEUED actor id + // is the activity's actor — downstream authorization (followed + // community, delete authority) keys off it. + boundActor := actorID + if claimed := refID(activity.Actor); claimed != "" { + if !ap.SameAuthority(claimed, actorID) { + ib.logger.Warn("inbox delivery rejected: actor/signer authority mismatch", + "activity_actor", claimed, "signer", actorID) + http.Error(w, "activity actor does not match signature", http.StatusForbidden) + return + } + boundActor = claimed + } + + isNew, err := ib.events.Enqueue(r.Context(), store.InboxEvent{ + ActivityID: activity.ID, + Type: activity.Type, + Payload: body, + ActorID: boundActor, + OrderingKey: orderingKeyFor(activity, boundActor), + }) + if err != nil { + ib.logger.Error("enqueue inbox event", "activity_id", activity.ID, "error", err) + http.Error(w, "failed to record delivery", http.StatusInternalServerError) + return + } + if !isNew { + // Duplicate delivery: acknowledged, not re-queued. + w.WriteHeader(http.StatusOK) + return + } + ib.queue.Nudge() + w.WriteHeader(http.StatusAccepted) +} + +// orderingKeyFor derives the per-community serialization key without any +// network round-trip: the community IRI when the activity (or its nested +// objects) is addressed to one, otherwise the bound actor — so one +// community's events apply in order, and unrelated actors never block each +// other. +func orderingKeyFor(activity *ap.Object, boundActor string) string { + for probe := activity; probe != nil; probe = probe.Object { + if iri := communityIRIFrom(probe); iri != "" { + return iri + } + } + return boundActor +} + +// handleActor serves the bridge's Application actor document (Lemmy fetches +// it to validate our Follow signatures). +func (ib *Inbox) handleActor(w http.ResponseWriter, _ *http.Request) { + doc, err := ib.service.DocumentJSON() + if err != nil { + ib.logger.Error("render service actor document", "error", err) + http.Error(w, "internal error", http.StatusInternalServerError) + return + } + w.Header().Set("Content-Type", ap.ContentTypeActivityJSON) + _, _ = w.Write(doc) +} + +// handleWebFinger answers for the service actor only (needed so Lemmy +// accepts our Follow: it resolves the follower's account). The bridged +// handle space is atproto-side and never surfaces here. +func (ib *Inbox) handleWebFinger(w http.ResponseWriter, r *http.Request) { + resource := r.URL.Query().Get("resource") + acct := fmt.Sprintf("acct:%s@%s", ib.service.Hostname, ib.service.Hostname) + if resource != acct && resource != ib.service.ID { + http.Error(w, "resource not found", http.StatusNotFound) + return + } + response := ap.WebFingerResponse{ + Subject: acct, + Aliases: []string{ib.service.ID}, + Links: []ap.WebFingerLink{{ + Rel: "self", + Type: ap.ContentTypeActivityJSON, + Href: ib.service.ID, + }}, + } + w.Header().Set("Content-Type", "application/jrd+json") + _ = json.NewEncoder(w).Encode(response) +} + +// handleNodeInfoDiscovery serves the nodeinfo well-known discovery document. +func (ib *Inbox) handleNodeInfoDiscovery(w http.ResponseWriter, _ *http.Request) { + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(map[string]any{ + "links": []any{map[string]any{ + "rel": "http://nodeinfo.diaspora.software/ns/schema/2.0", + "href": "https://" + ib.service.Hostname + "/nodeinfo/2.0", + }}, + }) +} + +// handleNodeInfo serves a minimal nodeinfo 2.0 document. Lemmy reads +// software.name for its instance allow/block lists, so the bridge must +// identify itself here. +func (ib *Inbox) handleNodeInfo(w http.ResponseWriter, _ *http.Request) { + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(map[string]any{ + "version": "2.0", + "software": map[string]any{ + "name": softwareName, + "version": softwareVersion, + }, + "protocols": []any{"activitypub"}, + "services": map[string]any{"inbound": []any{}, "outbound": []any{}}, + "openRegistrations": false, + "usage": map[string]any{"users": map[string]any{}}, + "metadata": map[string]any{}, + }) +} + +// requireBearer is the admin-auth middleware shared with follow.go: a +// constant-time bearer-token check against ADMIN_TOKEN. +func requireBearer(token string, logger *slog.Logger) func(http.Handler) http.Handler { + return func(next http.Handler) http.Handler { + return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + const prefix = "Bearer " + header := r.Header.Get("Authorization") + if len(header) <= len(prefix) || header[:len(prefix)] != prefix || + subtle.ConstantTimeCompare([]byte(header[len(prefix):]), []byte(token)) != 1 { + logger.Warn("admin request rejected: bad bearer token", "path", r.URL.Path) + http.Error(w, "unauthorized", http.StatusUnauthorized) + return + } + next.ServeHTTP(w, r) + }) + } +} + +// decodeJSONBody decodes a small JSON request body into dst. +func decodeJSONBody(r *http.Request, dst any) error { + defer func() { _, _ = io.Copy(io.Discard, r.Body) }() + decoder := json.NewDecoder(io.LimitReader(r.Body, 1<<16)) + if err := decoder.Decode(dst); err != nil { + return fmt.Errorf("decode request body: %w", err) + } + return nil +} diff --git a/internal/ingest/inbox_test.go b/internal/ingest/inbox_test.go new file mode 100644 index 0000000..a7b4c11 --- /dev/null +++ b/internal/ingest/inbox_test.go @@ -0,0 +1,246 @@ +package ingest + +import ( + "bytes" + "context" + "encoding/json" + "net/http" + "net/http/httptest" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "tidepool/internal/ap" + "tidepool/internal/errors" +) + +// likeActivity is a small, valid activity for signature-path tests (it +// dispatches to the recording vote stub, no fixtures needed). +func likeActivity(id string, actor string) map[string]any { + return map[string]any{ + "id": id, + "type": "Like", + "actor": actor, + "object": pageID, + } +} + +// TestInboxValidSignature: a correctly signed delivery is accepted (202), +// recorded, and processed by the queue. +func TestInboxValidSignature(t *testing.T) { + h := newHarness(t) + alice := h.newRemoteActor("https://lemmy.world/u/alice", person("https://lemmy.world/u/alice", "alice", nil)) + + status := h.deliver(alice, likeActivity("https://lemmy.world/activities/like/1", alice.id)) + require.Equal(t, http.StatusAccepted, status) + + event, err := h.events.GetEvent(context.Background(), "https://lemmy.world/activities/like/1") + require.NoError(t, err) + assert.Equal(t, "Like", event.Type) + assert.Equal(t, alice.id, event.ActorID) + + h.drain() + event, err = h.events.GetEvent(context.Background(), "https://lemmy.world/activities/like/1") + require.NoError(t, err) + assert.NotNil(t, event.ProcessedAt, "valid deliveries must be processed") + assert.Len(t, h.votes.applied, 1, "bare Like goes to the vote aggregator") +} + +// TestInboxActorInboxAlias: the per-actor inbox path accepts the same +// deliveries as the shared inbox. +func TestInboxActorInboxAlias(t *testing.T) { + h := newHarness(t) + alice := h.newRemoteActor("https://lemmy.world/u/alice2", person("https://lemmy.world/u/alice2", "alice2", nil)) + + status := h.deliverTo("/actor/inbox", alice, likeActivity("https://lemmy.world/activities/like/alias", alice.id)) + require.Equal(t, http.StatusAccepted, status) +} + +// TestInboxBadDigest: a body that does not match the signed Digest header +// is rejected before any key resolution. +func TestInboxBadDigest(t *testing.T) { + h := newHarness(t) + alice := h.newRemoteActor("https://lemmy.world/u/badd", person("https://lemmy.world/u/badd", "badd", nil)) + + signed, err := json.Marshal(likeActivity("https://lemmy.world/activities/like/2", alice.id)) + require.NoError(t, err) + tampered, err := json.Marshal(likeActivity("https://lemmy.world/activities/like/TAMPERED", alice.id)) + require.NoError(t, err) + + req := httptest.NewRequest(http.MethodPost, "https://"+bridgeHost+"/inbox", bytes.NewReader(tampered)) + req.Header.Set("Content-Type", ap.ContentTypeActivityJSON) + require.NoError(t, alice.signer().SignRequest(req, signed)) // digest over the ORIGINAL body + + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + assert.Equal(t, http.StatusUnauthorized, rec.Code) + + _, err = h.events.GetEvent(context.Background(), "https://lemmy.world/activities/like/TAMPERED") + assert.True(t, errors.IsNotFound(err), "rejected deliveries must not be enqueued") +} + +// TestInboxExpiredDate: a Date header outside the 1h skew window is +// rejected (Lemmy's EXPIRES_AFTER). +func TestInboxExpiredDate(t *testing.T) { + h := newHarness(t) + alice := h.newRemoteActor("https://lemmy.world/u/late", person("https://lemmy.world/u/late", "late", nil)) + + body, err := json.Marshal(likeActivity("https://lemmy.world/activities/like/3", alice.id)) + require.NoError(t, err) + req := httptest.NewRequest(http.MethodPost, "https://"+bridgeHost+"/inbox", bytes.NewReader(body)) + req.Header.Set("Content-Type", ap.ContentTypeActivityJSON) + require.NoError(t, alice.signer().SignRequest(req, body)) + // Backdate the request past the skew window. (The stale Date also breaks + // the signature, but the skew check runs first and is what must reject.) + req.Header.Set("Date", time.Now().Add(-2*time.Hour).UTC().Format(http.TimeFormat)) + + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + assert.Equal(t, http.StatusUnauthorized, rec.Code) +} + +// TestInboxUnsignedRejected: no Signature header → 401. +func TestInboxUnsignedRejected(t *testing.T) { + h := newHarness(t) + body, err := json.Marshal(likeActivity("https://lemmy.world/activities/like/4", "https://lemmy.world/u/ghost")) + require.NoError(t, err) + req := httptest.NewRequest(http.MethodPost, "https://"+bridgeHost+"/inbox", bytes.NewReader(body)) + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + assert.Equal(t, http.StatusUnauthorized, rec.Code) +} + +// TestInboxKeyRotationRefetch: a delivery signed with a ROTATED key still +// verifies — the stale cached key fails, the verifier re-fetches once +// (because the key came from cache) and recovers. +func TestInboxKeyRotationRefetch(t *testing.T) { + h := newHarness(t) + rotator := h.newRemoteActor("https://lemmy.world/u/rotator", person("https://lemmy.world/u/rotator", "rotator", nil)) + + // Prime the key cache with a valid delivery under the ORIGINAL key. + status := h.deliver(rotator, likeActivity("https://lemmy.world/activities/like/rot-1", rotator.id)) + require.Equal(t, http.StatusAccepted, status) + fetchesAfterPrime := h.hitCount("/u/rotator") + + // Rotate: publish a new key and sign the next delivery with it. + newKey, err := ap.GenerateRSAKey() + require.NoError(t, err) + h.serveActorDoc(rotator.id, person(rotator.id, "rotator", nil), &newKey.PublicKey) + rotated := &remoteActor{id: rotator.id, key: newKey} + + status = h.deliver(rotated, likeActivity("https://lemmy.world/activities/like/rot-2", rotator.id)) + assert.Equal(t, http.StatusAccepted, status, + "a rotated key must recover via the one-shot fresh re-fetch") + assert.Equal(t, fetchesAfterPrime+1, h.hitCount("/u/rotator"), + "rotation recovery costs exactly one extra actor fetch") +} + +// TestInboxFreshKeyFailureNoRefetch is the amplification gate: when the +// failing key was NOT served from cache (it was just fetched), the verifier +// must NOT fetch again — a forged delivery costs at most one outbound +// request. +func TestInboxFreshKeyFailureNoRefetch(t *testing.T) { + h := newHarness(t) + victim := h.newRemoteActor("https://lemmy.world/u/victim", person("https://lemmy.world/u/victim", "victim", nil)) + + // Sign with a key the victim never published; nothing cached yet. + forgerKey, err := ap.GenerateRSAKey() + require.NoError(t, err) + forger := &remoteActor{id: victim.id, key: forgerKey} + + status := h.deliver(forger, likeActivity("https://lemmy.world/activities/like/forged", victim.id)) + assert.Equal(t, http.StatusUnauthorized, status) + assert.Equal(t, 1, h.hitCount("/u/victim"), + "a signature failing against a FRESH key must not trigger a second fetch") +} + +// TestInboxActorSignerMismatch: an activity claiming an actor on a +// different authority than the verified signer is rejected. +func TestInboxActorSignerMismatch(t *testing.T) { + h := newHarness(t) + mallory := h.newRemoteActor("https://evil.example/u/mallory", person("https://evil.example/u/mallory", "mallory", nil)) + + status := h.deliver(mallory, likeActivity("https://evil.example/activities/like/5", "https://lemmy.world/u/alice")) + assert.Equal(t, http.StatusForbidden, status, + "host A must not deliver activities attributed to host B") +} + +// TestInboxDedupe: re-delivering the same activity id acknowledges (200) +// without re-enqueueing or re-processing. +func TestInboxDedupe(t *testing.T) { + h := newHarness(t) + alice := h.newRemoteActor("https://lemmy.world/u/dedupe", person("https://lemmy.world/u/dedupe", "dedupe", nil)) + activity := likeActivity("https://lemmy.world/activities/like/once", alice.id) + + require.Equal(t, http.StatusAccepted, h.deliver(alice, activity)) + h.drain() + require.Len(t, h.votes.applied, 1) + + // Second delivery: acknowledged but not re-processed. + assert.Equal(t, http.StatusOK, h.deliver(alice, activity)) + h.drain() + assert.Len(t, h.votes.applied, 1, "duplicate deliveries must not re-process") + + // Third delivery straight to the DB layer: still a single row. + isNew, err := h.events.RecordEvent(context.Background(), "https://lemmy.world/activities/like/once", "Like") + require.NoError(t, err) + assert.False(t, isNew) +} + +// TestServiceActorEndpoints: /actor, WebFinger, and nodeinfo serve what +// Lemmy needs to accept a Follow from the bridge. +func TestServiceActorEndpoints(t *testing.T) { + h := newHarness(t) + + get := func(path string) *httptest.ResponseRecorder { + req := httptest.NewRequest(http.MethodGet, "https://"+bridgeHost+path, nil) + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + return rec + } + + // Actor document with the publicKey block Lemmy requires. + rec := get("/actor") + require.Equal(t, http.StatusOK, rec.Code) + actor, err := ap.ParseObject(rec.Body.Bytes()) + require.NoError(t, err) + assert.Equal(t, h.service.ID, actor.ID) + assert.Equal(t, ap.TypeApplication, actor.Type) + require.NotNil(t, actor.PublicKey) + assert.Equal(t, h.service.KeyID(), actor.PublicKey.ID) + assert.NotEmpty(t, actor.PublicKey.PublicKeyPem) + + // WebFinger resolves the service actor (and only it). + rec = get("/.well-known/webfinger?resource=acct:" + bridgeHost + "@" + bridgeHost) + require.Equal(t, http.StatusOK, rec.Code) + var jrd ap.WebFingerResponse + require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &jrd)) + require.NotEmpty(t, jrd.Links) + assert.Equal(t, h.service.ID, jrd.Links[0].Href) + assert.Equal(t, http.StatusNotFound, get("/.well-known/webfinger?resource=acct:nobody@example.com").Code) + + // nodeinfo discovery + document (Lemmy matches software.name). + rec = get("/.well-known/nodeinfo") + require.Equal(t, http.StatusOK, rec.Code) + var discovery struct { + Links []struct { + Href string `json:"href"` + } `json:"links"` + } + require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &discovery)) + require.NotEmpty(t, discovery.Links) + + rec = get("/nodeinfo/2.0") + require.Equal(t, http.StatusOK, rec.Code) + var nodeinfo struct { + Software struct { + Name string `json:"name"` + } `json:"software"` + Protocols []string `json:"protocols"` + } + require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &nodeinfo)) + assert.Equal(t, "tidepool", nodeinfo.Software.Name) + assert.Contains(t, nodeinfo.Protocols, "activitypub") +} diff --git a/internal/ingest/ingest_test.go b/internal/ingest/ingest_test.go new file mode 100644 index 0000000..5bb5236 --- /dev/null +++ b/internal/ingest/ingest_test.go @@ -0,0 +1,562 @@ +package ingest + +import ( + "bytes" + "context" + "crypto/rsa" + "crypto/sha256" + "encoding/base32" + "encoding/json" + "net/http" + "net/http/httptest" + "os" + "path/filepath" + "strings" + "sync" + "testing" + "time" + + "github.com/bluesky-social/indigo/atproto/atcrypto" + "github.com/go-chi/chi/v5" + "github.com/stretchr/testify/require" + + "tidepool/internal/ap" + "tidepool/internal/identity" + "tidepool/internal/materialize" + "tidepool/internal/repo" + "tidepool/internal/store" + "tidepool/internal/testutil" +) + +// Fixture ids (internal/ap/testdata, captured off live Lemmy). +const ( + groupID = "https://lemmy.world/c/technology" + personID = "https://lemmy.world/u/LeftLeaningFreedomFighters" + pageID = "https://lemmy.world/post/49131386" + noteID = "https://lemmy.zip/comment/27485395" + noteAuthor = "https://lemmy.zip/u/tixooo" + parentNote = "https://sh.itjust.works/comment/26248018" + parentActor = "https://sh.itjust.works/u/DemandtheOxfordComma" + + bridgeHost = "bridge.test" + testServiceDID = "did:web:bridge.test" + testAdminToken = "test-admin-token" +) + +// testKEK seals test signing keys. +var testKEK = []byte("0123456789abcdef0123456789abcdef") + +// pngBytes is a valid 1x1 transparent PNG (blob fetches). +var pngBytes = []byte{ + 0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a, 0x00, 0x00, 0x00, 0x0d, + 0x49, 0x48, 0x44, 0x52, 0x00, 0x00, 0x00, 0x01, 0x00, 0x00, 0x00, 0x01, + 0x08, 0x06, 0x00, 0x00, 0x00, 0x1f, 0x15, 0xc4, 0x89, 0x00, 0x00, 0x00, + 0x0d, 0x49, 0x44, 0x41, 0x54, 0x78, 0x9c, 0x62, 0x00, 0x01, 0x00, 0x00, + 0x05, 0x00, 0x01, 0x0d, 0x0a, 0x2d, 0xb4, 0x00, 0x00, 0x00, 0x00, 0x49, + 0x45, 0x4e, 0x44, 0xae, 0x42, 0x60, 0x82, +} + +var jpegBytes = []byte{0xff, 0xd8, 0xff, 0xe0, 0x00, 0x10, 0x4a, 0x46, 0x49, 0x46, 0x00, 0xff, 0xd9} + +// rewriteTransport routes every outbound request to the local fixture +// server, preserving the original URL via the Host header (same trick as +// the materialize tests): fixture JSON keeps its real fediverse ids while +// tests never touch the network. +type rewriteTransport struct{ target string } + +func (rt rewriteTransport) RoundTrip(req *http.Request) (*http.Response, error) { + clone := req.Clone(req.Context()) + clone.Host = req.URL.Host + clone.URL.Scheme = "http" + clone.URL.Host = rt.target + return http.DefaultTransport.RoundTrip(clone) +} + +// fakeMinter mints deterministic local identities (no PLC directory). +type fakeMinter struct { + custodian *identity.Custodian + mu sync.Mutex + mints int +} + +func (f *fakeMinter) MintActor(_ context.Context, req identity.MintRequest) (*identity.Identity, error) { + f.mu.Lock() + f.mints++ + f.mu.Unlock() + key, err := atcrypto.GeneratePrivateKeyK256() + if err != nil { + return nil, err + } + pub, err := key.PublicKey() + if err != nil { + return nil, err + } + sum := sha256.Sum256([]byte(req.PreferredUsername + "@" + req.Instance)) + did := "did:plc:" + strings.ToLower(base32.StdEncoding.EncodeToString(sum[:]))[:24] + sealed, err := f.custodian.EncryptActorKey(did, key) + if err != nil { + return nil, err + } + handle := strings.ToLower(req.PreferredUsername) + "." + + strings.ReplaceAll(req.Instance, ".", "-") + "." + bridgeHost + return &identity.Identity{ + DID: did, Handle: handle, DIDKey: pub.DIDKey(), SigningKeyEncrypted: sealed, + }, nil +} + +// testDIDFor mirrors fakeMinter's derivation for assertions. +func testDIDFor(username, instance string) string { + sum := sha256.Sum256([]byte(username + "@" + instance)) + return "did:plc:" + strings.ToLower(base32.StdEncoding.EncodeToString(sum[:]))[:24] +} + +// remoteActor is a fake fediverse actor that can sign deliveries to the +// bridge's inbox. +type remoteActor struct { + id string + key *rsa.PrivateKey +} + +func (a *remoteActor) signer() *ap.Signer { return ap.NewSigner(a.id+"#main-key", a.key) } + +// recordingBackfill captures follow-accepted triggers. +type recordingBackfill struct { + mu sync.Mutex + triggers []string +} + +func (r *recordingBackfill) TriggerAsync(c *store.Community, _ bool) { + r.mu.Lock() + defer r.mu.Unlock() + r.triggers = append(r.triggers, c.APGroupID) +} + +func (r *recordingBackfill) count() int { + r.mu.Lock() + defer r.mu.Unlock() + return len(r.triggers) +} + +// recordingVotes captures vote-aggregator hand-offs. +type recordingVotes struct { + mu sync.Mutex + applied []string + retracted []string +} + +func (v *recordingVotes) ApplyVote(_ context.Context, vote *ap.Object, _ string) error { + v.mu.Lock() + defer v.mu.Unlock() + v.applied = append(v.applied, vote.Type+" "+refID(vote.Object)) + return nil +} + +func (v *recordingVotes) RetractVote(_ context.Context, vote *ap.Object, _ string) error { + v.mu.Lock() + defer v.mu.Unlock() + v.retracted = append(v.retracted, vote.Type+" "+refID(vote.Object)) + return nil +} + +type harness struct { + t *testing.T + + router chi.Router + queue *Queue + handler *Handler + inbox *Inbox + admin *Admin + client *ap.Client + mat *materialize.Materializer + manager *repo.Manager + minter *fakeMinter + service *ap.ServiceActor + objects store.APObjects + actors store.BridgedActors + communities store.Communities + tombstones store.Tombstones + events store.InboxEvents + backfills *recordingBackfill + votes *recordingVotes + + mux *http.ServeMux + fixtures *httptest.Server + mu sync.Mutex + served map[string][]byte + hits map[string]int + inboxLog [][]byte // POSTs captured at the fake Lemmy shared inbox +} + +// serviceRSAKey is generated once: the bridge service actor's key is not +// under test and RSA keygen is the slowest part of harness setup. +var ( + serviceRSAKey *rsa.PrivateKey + serviceRSAKeyOnce sync.Once +) + +func testServiceActor(t *testing.T) *ap.ServiceActor { + t.Helper() + serviceRSAKeyOnce.Do(func() { + key, err := ap.GenerateRSAKey() + require.NoError(t, err) + serviceRSAKey = key + }) + return &ap.ServiceActor{ + ID: "https://" + bridgeHost + "/actor", + Hostname: bridgeHost, + Key: serviceRSAKey, + } +} + +func newHarness(t *testing.T) *harness { + t.Helper() + database := testutil.DB(t) + testutil.Truncate(t, database, + "ap_objects", "bridged_actors", "communities", "inbox_events", + "ap_tombstones", "blocks", "repo_state", "firehose_events", "blobs") + + custodian, err := identity.NewCustodian(testKEK) + require.NoError(t, err) + actors := store.NewBridgedActors(database) + objects := store.NewAPObjects(database) + communities := store.NewCommunities(database) + tombstones := store.NewTombstones(database) + events := store.NewInboxEvents(database) + manager, err := repo.NewManager(database, identity.NewActorKeys(actors, custodian), nil) + require.NoError(t, err) + + h := &harness{ + t: t, + objects: objects, + actors: actors, + communities: communities, + tombstones: tombstones, + events: events, + manager: manager, + served: map[string][]byte{}, + hits: map[string]int{}, + } + h.mux = http.NewServeMux() + h.fixtures = httptest.NewServer(h.mux) + t.Cleanup(h.fixtures.Close) + + // The fake Lemmy shared inbox: captures the bridge's outbound + // deliveries (Follow, Undo{Follow}). + h.mux.HandleFunc("POST /inbox", func(w http.ResponseWriter, r *http.Request) { + body, _ := readAll(r) + h.mu.Lock() + h.inboxLog = append(h.inboxLog, body) + h.mu.Unlock() + w.WriteHeader(http.StatusOK) + }) + h.mux.HandleFunc("/pictrs/image/", func(w http.ResponseWriter, r *http.Request) { + if strings.HasSuffix(r.URL.Path, ".jpeg") || strings.HasSuffix(r.URL.Path, ".jpg") { + w.Header().Set("Content-Type", "image/jpeg") + _, _ = w.Write(jpegBytes) + return + } + w.Header().Set("Content-Type", "image/png") + _, _ = w.Write(pngBytes) + }) + + h.service = testServiceActor(t) + h.client = ap.NewClient(ap.ClientOptions{ + HTTPClient: &http.Client{Transport: rewriteTransport{target: h.fixtures.Listener.Addr().String()}}, + Signer: h.service.Signer(), + AllowPrivateAddresses: true, + PerHostRPS: 100000, + PerHostBurst: 100000, + MaxAttempts: 1, + }) + + h.minter = &fakeMinter{custodian: custodian} + h.mat, err = materialize.New(materialize.Options{ + Fetcher: h.client, + Objects: objects, + Actors: actors, + Communities: communities, + Repos: manager, + Minter: h.minter, + ServiceDID: testServiceDID, + StrictValidation: true, + }) + require.NoError(t, err) + + h.backfills = &recordingBackfill{} + h.votes = &recordingVotes{} + h.handler, err = NewHandler(HandlerOptions{ + Materializer: h.mat, + Fetcher: h.client, + Objects: objects, + Communities: communities, + Tombstones: tombstones, + Votes: h.votes, + Backfill: h.backfills, + ServiceActorID: h.service.ID, + }) + require.NoError(t, err) + + h.queue, err = NewQueue(QueueOptions{ + Events: events, + Processor: h.handler, + Workers: 1, + MaxAttempts: 3, + Lease: time.Minute, + }) + require.NoError(t, err) + + h.inbox, err = NewInbox(InboxOptions{ + Verifier: ap.NewVerifier(h.client), + Events: events, + Queue: h.queue, + Service: h.service, + }) + require.NoError(t, err) + + h.admin, err = NewAdmin(AdminOptions{ + Token: testAdminToken, + Client: h.client, + Materializer: h.mat, + Communities: communities, + Service: h.service, + Backfill: h.backfills, + }) + require.NoError(t, err) + + router := chi.NewRouter() + h.inbox.Routes(router) + h.admin.Routes(router) + h.router = router + return h +} + +func readAll(r *http.Request) ([]byte, error) { + defer func() { _ = r.Body.Close() }() + var buf bytes.Buffer + _, err := buf.ReadFrom(r.Body) + return buf.Bytes(), err +} + +// serveJSON serves (or re-serves) body at path, counting hits. +func (h *harness) serveJSON(path string, body []byte) { + h.mu.Lock() + defer h.mu.Unlock() + if _, registered := h.served[path]; !registered { + h.mux.HandleFunc("GET "+path, func(w http.ResponseWriter, r *http.Request) { + h.mu.Lock() + current := h.served[path] + h.hits[path]++ + h.mu.Unlock() + w.Header().Set("Content-Type", ap.ContentTypeActivityJSON) + _, _ = w.Write(current) + }) + } + h.served[path] = body +} + +func (h *harness) serveObject(path string, obj map[string]any) { + h.t.Helper() + body, err := json.Marshal(obj) + require.NoError(h.t, err) + h.serveJSON(path, body) +} + +func (h *harness) hitCount(path string) int { + h.mu.Lock() + defer h.mu.Unlock() + return h.hits[path] +} + +// loadFixture parses a testdata fixture into a Go map for tweaking. +func loadFixture(t *testing.T, file string) map[string]any { + t.Helper() + body, err := os.ReadFile(filepath.Join("..", "ap", "testdata", file)) + require.NoError(t, err) + var doc map[string]any + require.NoError(t, json.Unmarshal(body, &doc)) + return doc +} + +// urlPath extracts the path of an absolute AP id. +func urlPath(t *testing.T, id string) string { + t.Helper() + idx := strings.Index(id, "//") + require.GreaterOrEqual(t, idx, 0) + rest := id[idx+2:] + slash := strings.IndexByte(rest, '/') + require.Greater(t, slash, 0) + return rest[slash:] +} + +// newRemoteActor generates an RSA key for an actor, serves its document +// (synthetic or fixture-derived) with the matching public key, and returns +// the signing handle for deliveries. +func (h *harness) newRemoteActor(id string, doc map[string]any) *remoteActor { + h.t.Helper() + key, err := ap.GenerateRSAKey() + require.NoError(h.t, err) + h.serveActorDoc(id, doc, &key.PublicKey) + return &remoteActor{id: id, key: key} +} + +// serveActorDoc publishes an actor document whose publicKey is pub. +func (h *harness) serveActorDoc(id string, doc map[string]any, pub *rsa.PublicKey) { + h.t.Helper() + publicPEM, err := ap.EncodePublicKeyPEM(pub) + require.NoError(h.t, err) + doc["publicKey"] = map[string]any{ + "id": id + "#main-key", + "owner": id, + "publicKeyPem": string(publicPEM), + } + h.serveObject(urlPath(h.t, id), doc) +} + +// person builds a minimal synthetic Person document. +func person(id, username string, extra map[string]any) map[string]any { + doc := map[string]any{ + "type": "Person", + "id": id, + "preferredUsername": username, + "inbox": id + "/inbox", + "published": "2024-01-01T00:00:00.000000Z", + } + for k, v := range extra { + doc[k] = v + } + return doc +} + +// note builds a minimal synthetic Lemmy comment. +func note(id, author, inReplyTo, markdown, published string) map[string]any { + return map[string]any{ + "type": "Note", + "id": id, + "attributedTo": author, + "to": []any{ap.PublicAudience}, + "audience": groupID, + "content": "

" + markdown + "

", + "source": map[string]any{"content": markdown, "mediaType": "text/markdown"}, + "published": published, + "inReplyTo": inReplyTo, + } +} + +// technologyGroup returns the lemmy.world group fixture as a remote actor +// (its published RSA key replaced with a locally generated one so the actor +// can sign deliveries). +func (h *harness) technologyGroup() *remoteActor { + h.t.Helper() + doc := loadFixture(h.t, "group_lemmy_world.json") + return h.newRemoteActor(groupID, doc) +} + +// serveLemmyWorldContent registers the standard page/person fixtures. +func (h *harness) serveLemmyWorldContent() { + h.t.Helper() + page, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "page_lemmy_world.json")) + require.NoError(h.t, err) + h.serveJSON("/post/49131386", page) + personDoc, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "person_lemmy_world.json")) + require.NoError(h.t, err) + h.serveJSON("/u/LeftLeaningFreedomFighters", personDoc) +} + +// deliver signs and posts an activity to the bridge inbox, returning the +// HTTP status. +func (h *harness) deliver(actor *remoteActor, activity map[string]any) int { + h.t.Helper() + return h.deliverTo("/inbox", actor, activity) +} + +func (h *harness) deliverTo(path string, actor *remoteActor, activity map[string]any) int { + h.t.Helper() + body, err := json.Marshal(activity) + require.NoError(h.t, err) + req := httptest.NewRequest(http.MethodPost, "https://"+bridgeHost+path, bytes.NewReader(body)) + req.Header.Set("Content-Type", ap.ContentTypeActivityJSON) + require.NoError(h.t, actor.signer().SignRequest(req, body)) + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + return rec.Code +} + +// drain synchronously processes queued events until the queue idles. +func (h *harness) drain() { + h.t.Helper() + for h.queue.processNext(context.Background()) { + } +} + +// adminRequest performs an authenticated admin API call. +func (h *harness) adminRequest(method, path string, body map[string]any) *httptest.ResponseRecorder { + h.t.Helper() + var reader *bytes.Reader + if body != nil { + raw, err := json.Marshal(body) + require.NoError(h.t, err) + reader = bytes.NewReader(raw) + } else { + reader = bytes.NewReader(nil) + } + req := httptest.NewRequest(method, "https://"+bridgeHost+path, reader) + req.Header.Set("Authorization", "Bearer "+testAdminToken) + rec := httptest.NewRecorder() + h.router.ServeHTTP(rec, req) + return rec +} + +// subscribeTechnology drives the whole follow lifecycle with the fixture +// community: admin subscribe → Follow captured → Accept delivered → +// follow_state accepted. Returns the community's signing handle. +func (h *harness) subscribeTechnology() *remoteActor { + h.t.Helper() + group := h.technologyGroup() + webfinger, err := os.ReadFile(filepath.Join("..", "ap", "testdata", "webfinger_group.json")) + require.NoError(h.t, err) + h.serveJSON("/.well-known/webfinger", webfinger) + + rec := h.adminRequest(http.MethodPost, "/admin/communities", + map[string]any{"community": "!technology@lemmy.world"}) + require.Equal(h.t, http.StatusAccepted, rec.Code, rec.Body.String()) + + // The bridge delivered a Follow to the fake Lemmy inbox. + h.mu.Lock() + require.NotEmpty(h.t, h.inboxLog, "subscribe must deliver a Follow") + followRaw := h.inboxLog[len(h.inboxLog)-1] + h.mu.Unlock() + follow, err := ap.ParseObject(followRaw) + require.NoError(h.t, err) + require.Equal(h.t, ap.TypeFollow, follow.Type) + require.Equal(h.t, h.service.ID, follow.Actor.ID) + require.Equal(h.t, groupID, follow.Object.ID) + + // Lemmy answers with Accept{Follow}, signed by the community. + status := h.deliver(group, map[string]any{ + "id": "https://lemmy.world/activities/accept/follow-1", + "type": "Accept", + "actor": groupID, + "object": map[string]any{"id": follow.ID, "type": "Follow", "actor": h.service.ID, "object": groupID}, + }) + require.Equal(h.t, http.StatusAccepted, status) + h.drain() + + community, err := h.communities.GetByAPGroupID(context.Background(), groupID) + require.NoError(h.t, err) + require.Equal(h.t, store.FollowStateAccepted, community.FollowState) + return group +} + +// firehoseOps flattens all firehose event op paths. +func (h *harness) firehoseOps() []string { + h.t.Helper() + events, err := h.manager.ListEvents(context.Background(), 0, 1000) + require.NoError(h.t, err) + var paths []string + for _, evt := range events { + for _, op := range evt.Ops { + paths = append(paths, op.Path) + } + } + return paths +} diff --git a/internal/ingest/mintgate.go b/internal/ingest/mintgate.go new file mode 100644 index 0000000..7b3bf56 --- /dev/null +++ b/internal/ingest/mintgate.go @@ -0,0 +1,66 @@ +package ingest + +import ( + "context" + stderrors "errors" + "fmt" + "log/slog" + + "golang.org/x/time/rate" + + "tidepool/internal/errors" + "tidepool/internal/identity" + "tidepool/internal/materialize" +) + +// ErrMintRateExceeded marks a mint refused by the rate gate. It is a plain +// retryable error (not a skip, not a validation error): the queue backs the +// delivery off and the mint succeeds on a later attempt once the bucket +// refills. +var ErrMintRateExceeded = stderrors.New("ingest: DID mint rate exceeded") + +// MintGate rate-limits inbound DID minting. Without it, a crafted deep +// comment thread with a distinct fake author per level mints up to the +// materializer's ancestor-depth cap (50) of PLC identities per delivered +// object — and PLC registrations are forever (the amplification vector +// flagged in task 05). The token bucket bounds sustained mint throughput +// while the burst absorbs legitimate spikes (community backfills). +type MintGate struct { + inner materialize.ActorMinter + limiter *rate.Limiter + logger *slog.Logger +} + +// NewMintGate wraps a minter with a token bucket of perMinute sustained +// mints and the given burst. +func NewMintGate(inner materialize.ActorMinter, perMinute float64, burst int, logger *slog.Logger) (*MintGate, error) { + if inner == nil { + return nil, errors.NewValidationError("minter", "must not be nil") + } + if perMinute <= 0 { + return nil, errors.NewValidationError("per_minute", "must be positive") + } + if burst <= 0 { + return nil, errors.NewValidationError("burst", "must be positive") + } + if logger == nil { + logger = slog.Default() + } + return &MintGate{ + inner: inner, + limiter: rate.NewLimiter(rate.Limit(perMinute/60.0), burst), + logger: logger, + }, nil +} + +// MintActor implements materialize.ActorMinter. A refused mint fails fast +// (never blocks a queue worker) with ErrMintRateExceeded. +func (g *MintGate) MintActor(ctx context.Context, req identity.MintRequest) (*identity.Identity, error) { + if !g.limiter.Allow() { + g.logger.Warn("DID mint refused by rate gate", + "preferred_username", req.PreferredUsername, "instance", req.Instance) + return nil, fmt.Errorf("mint identity for %s@%s: %w", + req.PreferredUsername, req.Instance, ErrMintRateExceeded) + } + return g.inner.MintActor(ctx, req) +} diff --git a/internal/ingest/mintgate_test.go b/internal/ingest/mintgate_test.go new file mode 100644 index 0000000..f9cbc0c --- /dev/null +++ b/internal/ingest/mintgate_test.go @@ -0,0 +1,53 @@ +package ingest + +import ( + "context" + stderrors "errors" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "tidepool/internal/identity" + "tidepool/internal/materialize" +) + +type countingMinter struct{ mints int } + +func (m *countingMinter) MintActor(_ context.Context, _ identity.MintRequest) (*identity.Identity, error) { + m.mints++ + return &identity.Identity{DID: "did:plc:fake", Handle: "fake.handle"}, nil +} + +// TestMintGateBoundsAmplification: the deep-thread attack scenario — many +// distinct authors arriving at once — is capped at the burst size, refused +// mints are retryable (not skips), and the inner minter is never called for +// a refused mint. +func TestMintGateBoundsAmplification(t *testing.T) { + inner := &countingMinter{} + gate, err := NewMintGate(inner, 60, 3, nil) + require.NoError(t, err) + + ctx := context.Background() + var refused int + for i := 0; i < 50; i++ { + _, err := gate.MintActor(ctx, identity.MintRequest{PreferredUsername: "troll", Instance: "evil.example"}) + if err != nil { + require.True(t, stderrors.Is(err, ErrMintRateExceeded)) + assert.False(t, materialize.IsSkip(err), "a refused mint must be retryable, not a skip") + refused++ + } + } + assert.Equal(t, 3, inner.mints, "mints are capped at the burst size") + assert.Equal(t, 47, refused) +} + +func TestMintGateRejectsBadConfig(t *testing.T) { + inner := &countingMinter{} + _, err := NewMintGate(nil, 60, 10, nil) + assert.Error(t, err) + _, err = NewMintGate(inner, 0, 10, nil) + assert.Error(t, err) + _, err = NewMintGate(inner, 60, 0, nil) + assert.Error(t, err) +} diff --git a/internal/ingest/queue.go b/internal/ingest/queue.go new file mode 100644 index 0000000..dd68b28 --- /dev/null +++ b/internal/ingest/queue.go @@ -0,0 +1,291 @@ +package ingest + +import ( + "context" + "log/slog" + "sync" + "time" + + "tidepool/internal/errors" + "tidepool/internal/materialize" + "tidepool/internal/store" +) + +// Queue defaults, overridable through QueueOptions. +const ( + defaultWorkers = 4 + defaultMaxAttempts = 5 + defaultRetryBaseDelay = 30 * time.Second + defaultMaxRetryDelay = time.Hour + defaultLease = 2 * time.Minute + defaultPollInterval = time.Second + // bookkeepingTimeout bounds the outcome-recording writes, which run on a + // context detached from the worker's (see processNext) so a completed + // event is durably acked even as the pool is shutting down. + bookkeepingTimeout = 10 * time.Second +) + +// Processor consumes one claimed inbox event. *Handler implements it. +// Return contract: nil or IsSkip → processed; IsValidation → poisoned; +// anything else → retried with backoff until the attempt cap, then +// poisoned. +type Processor interface { + Process(ctx context.Context, event *store.InboxEvent) error +} + +// QueueOptions configures NewQueue. Events and Processor are required. +type QueueOptions struct { + Events store.InboxEvents + Processor Processor + // Workers is the pool size (default 4). Ordering stays correct at any + // size: ClaimNext never hands out an event whose ordering key has an + // older pending sibling. + Workers int + // MaxAttempts poisons an event after this many failed claims + // (default 5). + MaxAttempts int + // RetryBaseDelay is the first backoff step, doubling per attempt up to + // an hour (default 30s). + RetryBaseDelay time.Duration + // Lease is how long a claim lasts before a crashed worker's event + // becomes claimable again; it also bounds one event's processing time + // (default 2m). + Lease time.Duration + // PollInterval is the idle re-check period (default 1s). The inbox also + // nudges the pool on every enqueue, so this is only the fallback for + // retries and missed nudges. + PollInterval time.Duration + Logger *slog.Logger +} + +// Queue is the durable ingestion work loop: a worker pool draining +// inbox_events (the postgres-backed queue — no external broker), with +// per-community serial ordering, retry with backoff, and poison handling. +type Queue struct { + events store.InboxEvents + processor Processor + workers int + maxAttempts int + retryBaseDelay time.Duration + lease time.Duration + pollInterval time.Duration + logger *slog.Logger + + nudge chan struct{} + // now is a test seam for backoff scheduling. + now func() time.Time +} + +// NewQueue validates options and builds a Queue. +func NewQueue(opts QueueOptions) (*Queue, error) { + if opts.Events == nil { + return nil, errors.NewValidationError("events", "must not be nil") + } + if opts.Processor == nil { + return nil, errors.NewValidationError("processor", "must not be nil") + } + logger := opts.Logger + if logger == nil { + logger = slog.Default() + } + q := &Queue{ + events: opts.Events, + processor: opts.Processor, + workers: opts.Workers, + maxAttempts: opts.MaxAttempts, + retryBaseDelay: opts.RetryBaseDelay, + lease: opts.Lease, + pollInterval: opts.PollInterval, + logger: logger, + nudge: make(chan struct{}, 1), + now: time.Now, + } + if q.workers <= 0 { + q.workers = defaultWorkers + } + if q.maxAttempts <= 0 { + q.maxAttempts = defaultMaxAttempts + } + if q.retryBaseDelay <= 0 { + q.retryBaseDelay = defaultRetryBaseDelay + } + if q.lease <= 0 { + q.lease = defaultLease + } + if q.pollInterval <= 0 { + q.pollInterval = defaultPollInterval + } + return q, nil +} + +// Nudge wakes the worker pool without waiting for the poll interval. The +// inbox calls it after every enqueue (same process, no broker round-trip). +func (q *Queue) Nudge() { + select { + case q.nudge <- struct{}{}: + default: + } +} + +// Run drives the worker pool until ctx is cancelled. +func (q *Queue) Run(ctx context.Context) { + var wg sync.WaitGroup + for i := 0; i < q.workers; i++ { + wg.Add(1) + go func() { + defer wg.Done() + q.workerLoop(ctx) + }() + } + wg.Wait() +} + +// workerLoop drains the queue, then sleeps until a nudge or the poll tick. +func (q *Queue) workerLoop(ctx context.Context) { + ticker := time.NewTicker(q.pollInterval) + defer ticker.Stop() + for { + // Drain: keep claiming while there is work. + for { + if ctx.Err() != nil { + return + } + claimed := q.processNext(ctx) + if !claimed { + break + } + } + select { + case <-ctx.Done(): + return + case <-q.nudge: + case <-ticker.C: + } + } +} + +// processNext claims and processes one event, reporting whether anything +// was claimed (false = queue idle). Every outcome is recorded on the row; +// bookkeeping failures are logged, never fatal — the lease expiry +// re-delivers the event. +func (q *Queue) processNext(ctx context.Context) bool { + event, err := q.events.ClaimNext(ctx, q.lease) + if err != nil { + if !errors.IsNotFound(err) && ctx.Err() == nil { + q.logger.Error("claim next inbox event", "error", err) + } + return false + } + + // Cascade the wakeup: this worker is now busy, so nudge an idle peer to + // claim the next event instead of leaving a burst of enqueues to drain + // one worker at a time. ClaimNext's NotFound path (empty queue) does not + // nudge, so this self-terminates. + q.Nudge() + + // The lease ClaimNext stamped is this attempt's fencing token; the + // outcome writes below present it so a stale worker cannot clobber a + // re-claim. + token := claimToken(event) + + // Bound processing by the lease: a worker must not outlive its claim, + // or a second worker could process the same event concurrently. + procCtx, cancel := context.WithTimeout(ctx, q.lease) + err = q.processor.Process(procCtx, event) + cancel() + + // Shutdown is not a processing failure. If the parent context was + // cancelled, a non-nil err is (almost certainly) cancellation, not a + // real fault: do not classify it — leave the row untouched so the lease + // lapses and the event is redelivered intact, rather than burning the + // attempt into a retry/poison. A successful Process (err == nil) still + // falls through so its outcome is durably acked below. + if err != nil && ctx.Err() != nil { + q.logger.Info("shutdown interrupted inbox event processing; leaving for redelivery", + "activity_id", event.ActivityID, "type", event.Type) + return true + } + + // Record the outcome on a context detached from the worker's, so an + // event that just completed is durably acked even though shutdown has + // cancelled ctx. Bounded so a stuck DB cannot hang shutdown forever. + bookCtx, bookCancel := context.WithTimeout(context.WithoutCancel(ctx), bookkeepingTimeout) + defer bookCancel() + + switch { + case err == nil: + q.finish(bookCtx, event, token) + case materialize.IsSkip(err): + // Deliberate drops: log the reason, mark processed, never retry. + q.logger.Info("inbox event skipped", + "activity_id", event.ActivityID, "type", event.Type, "reason", err.Error()) + q.finish(bookCtx, event, token) + case errors.IsValidation(err): + // Malformed payloads never get better; poison immediately. + q.logger.Warn("inbox event poisoned (invalid payload)", + "activity_id", event.ActivityID, "type", event.Type, "error", err) + q.poison(bookCtx, event, token, err.Error()) + case event.Attempts >= q.maxAttempts: + q.logger.Error("inbox event poisoned (attempt cap reached)", + "activity_id", event.ActivityID, "type", event.Type, + "attempts", event.Attempts, "error", err) + q.poison(bookCtx, event, token, err.Error()) + default: + delay := q.backoff(event.Attempts) + q.logger.Warn("inbox event failed; scheduling retry", + "activity_id", event.ActivityID, "type", event.Type, + "attempt", event.Attempts, "retry_in", delay, "error", err) + applied, rerr := q.events.Release(bookCtx, event.ActivityID, err.Error(), q.now().Add(delay), token) + if rerr != nil { + q.logger.Error("release inbox event", "activity_id", event.ActivityID, "error", rerr) + } else if !applied { + q.logger.Warn("stale claim; retry release discarded", "activity_id", event.ActivityID) + } + } + return true +} + +func (q *Queue) finish(ctx context.Context, event *store.InboxEvent, token time.Time) { + applied, err := q.events.MarkProcessed(ctx, event.ActivityID, token) + if err != nil { + q.logger.Error("mark inbox event processed", "activity_id", event.ActivityID, "error", err) + return + } + if !applied { + q.logger.Warn("stale claim; processed outcome discarded", "activity_id", event.ActivityID) + } +} + +func (q *Queue) poison(ctx context.Context, event *store.InboxEvent, token time.Time, message string) { + applied, err := q.events.MarkPoisoned(ctx, event.ActivityID, message, token) + if err != nil { + q.logger.Error("poison inbox event", "activity_id", event.ActivityID, "error", err) + return + } + if !applied { + q.logger.Warn("stale claim; poison outcome discarded", "activity_id", event.ActivityID) + } +} + +// claimToken extracts the fencing token (the lease deadline ClaimNext +// stamped) from a freshly claimed event. A claimed row always has +// ClaimedUntil set; the zero fallback is defensive and never matches a live +// row, so a would-be write is a safe no-op. +func claimToken(event *store.InboxEvent) time.Time { + if event.ClaimedUntil == nil { + return time.Time{} + } + return *event.ClaimedUntil +} + +// backoff doubles the base delay per completed attempt, capped at an hour. +func (q *Queue) backoff(attempts int) time.Duration { + delay := q.retryBaseDelay + for i := 1; i < attempts && delay < defaultMaxRetryDelay; i++ { + delay *= 2 + } + if delay > defaultMaxRetryDelay { + delay = defaultMaxRetryDelay + } + return delay +} diff --git a/internal/ingest/queue_test.go b/internal/ingest/queue_test.go new file mode 100644 index 0000000..9a97632 --- /dev/null +++ b/internal/ingest/queue_test.go @@ -0,0 +1,423 @@ +package ingest + +import ( + "context" + stderrors "errors" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "tidepool/internal/errors" + "tidepool/internal/store" + "tidepool/internal/testutil" +) + +// queueHarness is a lighter setup for queue-semantics tests: real +// inbox_events store, scriptable processor. +type queueHarness struct { + events store.InboxEvents +} + +func newQueueHarness(t *testing.T) *queueHarness { + t.Helper() + database := testutil.DB(t) + testutil.Truncate(t, database, "inbox_events") + return &queueHarness{events: store.NewInboxEvents(database)} +} + +func (qh *queueHarness) enqueue(t *testing.T, id, orderingKey string) { + t.Helper() + isNew, err := qh.events.Enqueue(context.Background(), store.InboxEvent{ + ActivityID: id, + Type: "Announce", + Payload: []byte(`{"id":"` + id + `","type":"Announce"}`), + ActorID: "https://lemmy.world/c/x", + OrderingKey: orderingKey, + }) + require.NoError(t, err) + require.True(t, isNew) +} + +// TestQueueOrderingPerKey: events sharing an ordering key process strictly +// serially and in arrival order; different keys interleave freely. +func TestQueueOrderingPerKey(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + qh.enqueue(t, "a-1", "community-a") + qh.enqueue(t, "a-2", "community-a") + qh.enqueue(t, "b-1", "community-b") + + // First claim: the oldest event overall. + first, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + assert.Equal(t, "a-1", first.ActivityID) + assert.Equal(t, 1, first.Attempts) + + // While a-1 is in flight, a-2 (same key) is invisible; b-1 (other key) + // is claimable. + second, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + assert.Equal(t, "b-1", second.ActivityID, + "a community's next event must wait for the in-flight one; other communities proceed") + + // Queue is now drained from a claimer's perspective. + _, err = qh.events.ClaimNext(ctx, time.Minute) + assert.True(t, errors.IsNotFound(err)) + + // Completing a-1 releases a-2. The claim holder presents its fencing + // token (the ClaimedUntil ClaimNext stamped). + require.NotNil(t, first.ClaimedUntil) + applied, err := qh.events.MarkProcessed(ctx, "a-1", *first.ClaimedUntil) + require.NoError(t, err) + require.True(t, applied) + third, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + assert.Equal(t, "a-2", third.ActivityID) +} + +// TestQueueRetryBackoff: a released event is invisible until its +// next_attempt_at, then claimable again with the attempt count preserved. +func TestQueueRetryBackoff(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + qh.enqueue(t, "retry-1", "community-a") + + event, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.Equal(t, 1, event.Attempts) + require.NotNil(t, event.ClaimedUntil) + + // Fail with a short (but comfortably-margined) backoff: releasing + // requires the claim's fencing token. Not claimable until it elapses... + backoff := 150 * time.Millisecond + applied, err := qh.events.Release(ctx, "retry-1", "remote 503", time.Now().Add(backoff), *event.ClaimedUntil) + require.NoError(t, err) + require.True(t, applied) + _, err = qh.events.ClaimNext(ctx, time.Minute) + assert.True(t, errors.IsNotFound(err), "backed-off events must not be claimable") + + // ...and it still blocks younger events on the same key (serial order). + qh.enqueue(t, "retry-2", "community-a") + _, err = qh.events.ClaimNext(ctx, time.Minute) + assert.True(t, errors.IsNotFound(err), + "a backing-off event must keep blocking its ordering key") + + // Once the backoff elapses it is claimable again, attempts accumulate, + // and the last failure is recorded. + time.Sleep(2 * backoff) + event, err = qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + assert.Equal(t, "retry-1", event.ActivityID) + assert.Equal(t, 2, event.Attempts) + assert.Equal(t, "remote 503", event.Error, "the last failure is recorded") +} + +// TestQueueFencingToken: a worker whose lease expired and was re-claimed by +// a second worker cannot overwrite the newer attempt's outcome. This is the +// correctness guarantee task 07's vote arithmetic depends on. +func TestQueueFencingToken(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + qh.enqueue(t, "fence-1", "community-a") + + // Worker A claims with a short lease, then its lease expires. + claimA, err := qh.events.ClaimNext(ctx, 60*time.Millisecond) + require.NoError(t, err) + require.NotNil(t, claimA.ClaimedUntil) + tokenA := *claimA.ClaimedUntil + + time.Sleep(120 * time.Millisecond) + + // Worker B re-claims the same event (a new attempt, a new token). + claimB, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.Equal(t, "fence-1", claimB.ActivityID) + require.NotNil(t, claimB.ClaimedUntil) + tokenB := *claimB.ClaimedUntil + assert.Equal(t, 2, claimB.Attempts) + assert.False(t, tokenB.Equal(tokenA), "re-claim must mint a distinct token") + + // A's stale outcomes are all no-ops: it no longer holds the claim. + applied, err := qh.events.MarkProcessed(ctx, "fence-1", tokenA) + require.NoError(t, err) + assert.False(t, applied, "stale MarkProcessed must be discarded") + + applied, err = qh.events.Release(ctx, "fence-1", "A late failure", time.Now().Add(time.Hour), tokenA) + require.NoError(t, err) + assert.False(t, applied, "stale Release must be discarded") + + applied, err = qh.events.MarkPoisoned(ctx, "fence-1", "A late poison", tokenA) + require.NoError(t, err) + assert.False(t, applied, "stale MarkPoisoned must be discarded") + + // B still holds the event: A's writes changed nothing. + stored, err := qh.events.GetEvent(ctx, "fence-1") + require.NoError(t, err) + assert.Nil(t, stored.ProcessedAt) + assert.Nil(t, stored.FailedAt) + assert.Empty(t, stored.Error) + + // B's outcome wins. + applied, err = qh.events.MarkProcessed(ctx, "fence-1", tokenB) + require.NoError(t, err) + assert.True(t, applied, "the current claim holder's outcome must be applied") + + // And A cannot un-process what B completed. + applied, err = qh.events.Release(ctx, "fence-1", "A even later", time.Now().Add(time.Hour), tokenA) + require.NoError(t, err) + assert.False(t, applied) + stored, err = qh.events.GetEvent(ctx, "fence-1") + require.NoError(t, err) + assert.NotNil(t, stored.ProcessedAt, "B's success must stand") +} + +// TestQueueMarkProcessedAndPoisonMutuallyExclusive: a stale finish cannot +// stamp processed_at onto a poisoned row (the failed_at guard), keeping the +// two terminal states disjoint for ops/metrics. +func TestQueueMarkProcessedAndPoisonMutuallyExclusive(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + qh.enqueue(t, "excl-1", "community-a") + + claim, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.NotNil(t, claim.ClaimedUntil) + token := *claim.ClaimedUntil + + applied, err := qh.events.MarkPoisoned(ctx, "excl-1", "bad payload", token) + require.NoError(t, err) + require.True(t, applied) + + // A late finish (same worker, same token) must not overwrite the poison. + applied, err = qh.events.MarkProcessed(ctx, "excl-1", token) + require.NoError(t, err) + assert.False(t, applied, "a finish must not resurrect a poisoned row") + + stored, err := qh.events.GetEvent(ctx, "excl-1") + require.NoError(t, err) + assert.NotNil(t, stored.FailedAt) + assert.Nil(t, stored.ProcessedAt, "poisoned and processed must stay mutually exclusive") +} + +// TestQueuePoison: a poisoned event is never claimed again, keeps its +// error, and stops blocking its ordering key. +func TestQueuePoison(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + qh.enqueue(t, "poison-1", "community-a") + qh.enqueue(t, "poison-2", "community-a") + + event, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.Equal(t, "poison-1", event.ActivityID) + require.NotNil(t, event.ClaimedUntil) + applied, err := qh.events.MarkPoisoned(ctx, "poison-1", "unparseable payload", *event.ClaimedUntil) + require.NoError(t, err) + require.True(t, applied) + + stored, err := qh.events.GetEvent(ctx, "poison-1") + require.NoError(t, err) + assert.NotNil(t, stored.FailedAt) + assert.Equal(t, "unparseable payload", stored.Error) + assert.Nil(t, stored.ProcessedAt) + + // The poisoned head no longer blocks the key. + next, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + assert.Equal(t, "poison-2", next.ActivityID) + require.NotNil(t, next.ClaimedUntil) + applied, err = qh.events.MarkProcessed(ctx, "poison-2", *next.ClaimedUntil) + require.NoError(t, err) + require.True(t, applied) + + _, err = qh.events.ClaimNext(ctx, time.Minute) + assert.True(t, errors.IsNotFound(err), "poisoned events are never re-claimed") +} + +// TestQueueLeaseExpiry: a crashed worker's claim expires and the event is +// re-delivered to another worker. +func TestQueueLeaseExpiry(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + qh.enqueue(t, "lease-1", "community-a") + + _, err := qh.events.ClaimNext(ctx, 50*time.Millisecond) + require.NoError(t, err) + _, err = qh.events.ClaimNext(ctx, time.Minute) + require.True(t, errors.IsNotFound(err), "a live lease blocks re-claiming") + + time.Sleep(80 * time.Millisecond) + event, err := qh.events.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + assert.Equal(t, "lease-1", event.ActivityID) + assert.Equal(t, 2, event.Attempts, "the re-delivery counts as a new attempt") +} + +// scriptedProcessor fails a configurable number of times, then succeeds. +type scriptedProcessor struct { + failures int + processed []string + errs int +} + +func (p *scriptedProcessor) Process(_ context.Context, event *store.InboxEvent) error { + if p.errs < p.failures { + p.errs++ + return stderrors.New("transient failure") + } + p.processed = append(p.processed, event.ActivityID) + return nil +} + +// TestQueueWorkerRetryThenSuccess drives the worker loop itself: a +// transient failure schedules a retry (visible via next_attempt_at), and a +// later pass processes the event. +func TestQueueWorkerRetryThenSuccess(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + processor := &scriptedProcessor{failures: 1} + q, err := NewQueue(QueueOptions{ + Events: qh.events, + Processor: processor, + MaxAttempts: 3, + RetryBaseDelay: time.Millisecond, + Lease: time.Minute, + }) + require.NoError(t, err) + + qh.enqueue(t, "flaky-1", "community-a") + require.True(t, q.processNext(ctx), "first pass claims and fails") + stored, err := qh.events.GetEvent(ctx, "flaky-1") + require.NoError(t, err) + assert.Nil(t, stored.ProcessedAt) + assert.Equal(t, "transient failure", stored.Error) + + // After the (tiny) backoff, the retry succeeds. + time.Sleep(20 * time.Millisecond) + require.True(t, q.processNext(ctx)) + stored, err = qh.events.GetEvent(ctx, "flaky-1") + require.NoError(t, err) + assert.NotNil(t, stored.ProcessedAt) + assert.Equal(t, []string{"flaky-1"}, processor.processed) +} + +// TestQueueWorkerPoisonsAfterMaxAttempts: persistent failures hit the +// attempt cap and the event is poisoned with the error recorded. +func TestQueueWorkerPoisonsAfterMaxAttempts(t *testing.T) { + qh := newQueueHarness(t) + ctx := context.Background() + processor := &scriptedProcessor{failures: 99} + q, err := NewQueue(QueueOptions{ + Events: qh.events, + Processor: processor, + MaxAttempts: 2, + RetryBaseDelay: time.Millisecond, + Lease: time.Minute, + }) + require.NoError(t, err) + + qh.enqueue(t, "doomed-1", "community-a") + require.True(t, q.processNext(ctx)) // attempt 1 → retry scheduled + time.Sleep(20 * time.Millisecond) + require.True(t, q.processNext(ctx)) // attempt 2 → cap reached → poison + + stored, err := qh.events.GetEvent(ctx, "doomed-1") + require.NoError(t, err) + assert.NotNil(t, stored.FailedAt, "the event must be poisoned at the attempt cap") + assert.Equal(t, "transient failure", stored.Error) + assert.False(t, q.processNext(ctx), "a poisoned event is not claimable") +} + +// blockingProcessor signals when Process is entered, then blocks until the +// given context is cancelled, returning that cancellation error. It models a +// worker whose in-flight Process is interrupted by shutdown. +type blockingProcessor struct { + entered chan struct{} +} + +func (p *blockingProcessor) Process(ctx context.Context, _ *store.InboxEvent) error { + close(p.entered) + <-ctx.Done() + return ctx.Err() +} + +// TestQueueShutdownDoesNotClassifyOutcome (Finding E-a): when the parent +// context is cancelled mid-Process, the interrupted work is NOT classified +// as a failure — the row is left untouched (no error recorded, not poisoned, +// not rescheduled) so the lease lapses and the event is redelivered intact. +func TestQueueShutdownDoesNotClassifyOutcome(t *testing.T) { + qh := newQueueHarness(t) + processor := &blockingProcessor{entered: make(chan struct{})} + q, err := NewQueue(QueueOptions{ + Events: qh.events, + Processor: processor, + MaxAttempts: 3, + Lease: time.Minute, + }) + require.NoError(t, err) + + qh.enqueue(t, "shutdown-1", "community-a") + + ctx, cancel := context.WithCancel(context.Background()) + done := make(chan bool, 1) + go func() { done <- q.processNext(ctx) }() + + <-processor.entered // Process is now in flight + cancel() // graceful shutdown + assert.True(t, <-done, "processNext still reports it claimed an event") + + // The event must be untouched: no error, not poisoned, not processed, + // and no extra attempt burned beyond the one the claim counted. + stored, err := qh.events.GetEvent(context.Background(), "shutdown-1") + require.NoError(t, err) + assert.Nil(t, stored.ProcessedAt, "cancellation must not mark processed") + assert.Nil(t, stored.FailedAt, "cancellation must not poison") + assert.Empty(t, stored.Error, "cancellation must not record a failure") + assert.Equal(t, 1, stored.Attempts, "cancellation must not burn a retry as a failure") +} + +// gatedSuccessProcessor blocks until released, then reports success. It +// models a Process that finishes right as shutdown cancels the parent ctx. +type gatedSuccessProcessor struct { + entered chan struct{} + release chan struct{} +} + +func (p *gatedSuccessProcessor) Process(_ context.Context, _ *store.InboxEvent) error { + close(p.entered) + <-p.release + return nil +} + +// TestQueueSuccessDuringShutdownIsAcked (Finding E-b): an event that +// SUCCEEDS just as shutdown cancels ctx is still durably marked processed, +// because the outcome write runs on a context detached from the worker's. +func TestQueueSuccessDuringShutdownIsAcked(t *testing.T) { + qh := newQueueHarness(t) + processor := &gatedSuccessProcessor{entered: make(chan struct{}), release: make(chan struct{})} + q, err := NewQueue(QueueOptions{ + Events: qh.events, + Processor: processor, + MaxAttempts: 3, + Lease: time.Minute, + }) + require.NoError(t, err) + + qh.enqueue(t, "shutdown-ok-1", "community-a") + + ctx, cancel := context.WithCancel(context.Background()) + done := make(chan bool, 1) + go func() { done <- q.processNext(ctx) }() + + <-processor.entered + cancel() // shutdown hits... + close(processor.release) // ...just as Process completes successfully + assert.True(t, <-done) + + stored, err := qh.events.GetEvent(context.Background(), "shutdown-ok-1") + require.NoError(t, err) + assert.NotNil(t, stored.ProcessedAt, "completed work must be acked despite shutdown") + assert.Empty(t, stored.Error) +} diff --git a/internal/ingest/votes.go b/internal/ingest/votes.go new file mode 100644 index 0000000..10b7509 --- /dev/null +++ b/internal/ingest/votes.go @@ -0,0 +1,52 @@ +package ingest + +import ( + "context" + "log/slog" + + "tidepool/internal/ap" +) + +// VoteAggregator consumes Like/Dislike activities. Votes never become +// records (PLAN.md locked decision 7): task 07 implements this interface +// with the bridge-side aggregate store behind the +// social.coves.bridge.getVoteAggregates XRPC. Task 06 only defines the +// seam and hands activities over. +type VoteAggregator interface { + // ApplyVote records one Like or Dislike. vote is the (possibly + // Announce-unwrapped) activity: Type is Like or Dislike, Actor is the + // voter, Object is the voted-on AP object. communityIRI is the + // announcing community's AP id ("" for a bare, un-announced vote). + ApplyVote(ctx context.Context, vote *ap.Object, communityIRI string) error + + // RetractVote undoes a previously applied vote (Undo{Like|Dislike}); + // vote is the inner activity being undone. + RetractVote(ctx context.Context, vote *ap.Object, communityIRI string) error +} + +// noopVotes is the task-06 placeholder implementation: it logs at debug and +// drops the vote. Task 07 replaces it with the real aggregator. +type noopVotes struct { + logger *slog.Logger +} + +// NewNoopVotes returns a VoteAggregator that discards votes (logged at +// debug). Wired until task 07 lands the aggregate store. +func NewNoopVotes(logger *slog.Logger) VoteAggregator { + if logger == nil { + logger = slog.Default() + } + return &noopVotes{logger: logger} +} + +func (v *noopVotes) ApplyVote(_ context.Context, vote *ap.Object, communityIRI string) error { + v.logger.Debug("vote dropped (aggregator lands in task 07)", + "type", vote.Type, "object", refID(vote.Object), "community", communityIRI) + return nil +} + +func (v *noopVotes) RetractVote(_ context.Context, vote *ap.Object, communityIRI string) error { + v.logger.Debug("vote retraction dropped (aggregator lands in task 07)", + "type", vote.Type, "object", refID(vote.Object), "community", communityIRI) + return nil +} diff --git a/internal/store/ap_objects.go b/internal/store/ap_objects.go index b26f0f2..66f0076 100644 --- a/internal/store/ap_objects.go +++ b/internal/store/ap_objects.go @@ -161,6 +161,28 @@ func (r *postgresAPObjects) SoftDelete(ctx context.Context, apID string) error { return nil } +func (r *postgresAPObjects) Restore(ctx context.Context, apID string) error { + // Clearing an already-clear deleted_at is harmless, so affected == 0 + // can only mean the row does not exist. + query := ` + UPDATE ap_objects + SET deleted_at = NULL + WHERE ap_id = $1` + + result, err := r.db.ExecContext(ctx, query, apID) + if err != nil { + return fmt.Errorf("restore ap_object %q: %w", apID, err) + } + affected, err := result.RowsAffected() + if err != nil { + return fmt.Errorf("restore ap_object %q: rows affected: %w", apID, err) + } + if affected == 0 { + return errors.NewNotFoundError("ap_object", apID) + } + return nil +} + // validateMapping checks the atproto identifiers with indigo's syntax // package, defaults Origin to fediverse, and derives ATURI from // (DID, Collection, RKey). diff --git a/internal/store/inbox_events.go b/internal/store/inbox_events.go index aa160e1..e082c09 100644 --- a/internal/store/inbox_events.go +++ b/internal/store/inbox_events.go @@ -5,6 +5,7 @@ import ( "database/sql" stderrors "errors" "fmt" + "time" "tidepool/internal/errors" ) @@ -19,47 +20,178 @@ func NewInboxEvents(db *sql.DB) InboxEvents { } func (r *postgresInboxEvents) RecordEvent(ctx context.Context, activityID, activityType string) (bool, error) { - if activityID == "" { + return r.Enqueue(ctx, InboxEvent{ActivityID: activityID, Type: activityType}) +} + +func (r *postgresInboxEvents) Enqueue(ctx context.Context, event InboxEvent) (bool, error) { + if event.ActivityID == "" { return false, errors.NewValidationError("activity_id", "must not be empty") } - if activityType == "" { + if event.Type == "" { return false, errors.NewValidationError("type", "must not be empty") } query := ` - INSERT INTO inbox_events (activity_id, type) - VALUES ($1, $2) + INSERT INTO inbox_events (activity_id, type, payload, actor_id, ordering_key) + VALUES ($1, $2, $3, $4, $5) ON CONFLICT (activity_id) DO NOTHING` - result, err := r.db.ExecContext(ctx, query, activityID, activityType) + result, err := r.db.ExecContext(ctx, query, + event.ActivityID, event.Type, event.Payload, event.ActorID, event.OrderingKey) if err != nil { - return false, fmt.Errorf("record inbox event %q: %w", activityID, err) + return false, fmt.Errorf("enqueue inbox event %q: %w", event.ActivityID, err) } affected, err := result.RowsAffected() if err != nil { - return false, fmt.Errorf("record inbox event %q: rows affected: %w", activityID, err) + return false, fmt.Errorf("enqueue inbox event %q: rows affected: %w", event.ActivityID, err) } return affected == 1, nil } -func (r *postgresInboxEvents) MarkProcessed(ctx context.Context, activityID string) error { +// eventColumns is the SELECT list shared by every query that scans a full +// InboxEvent row. +const eventColumns = `id, activity_id, type, COALESCE(payload, ''::bytea), actor_id, + ordering_key, attempts, next_attempt_at, claimed_until, failed_at, + received_at, processed_at, COALESCE(error, '')` + +func (r *postgresInboxEvents) ClaimNext(ctx context.Context, lease time.Duration) (*InboxEvent, error) { + if lease <= 0 { + return nil, errors.NewValidationError("lease", "must be positive") + } + + // The candidate subquery picks the oldest processable event: + // - unprocessed, not poisoned, past its retry schedule; + // - unleased, or leased by a worker whose lease expired (crash); + // - with NO older unprocessed, unpoisoned sibling on the same + // ordering key — this is the per-community serialization: while an + // older event is pending (claimed, backing off, or simply queued), + // every younger event on that key is invisible to workers. A + // poisoned sibling stops blocking (poison → skip). + // FOR UPDATE SKIP LOCKED lets concurrent workers race without + // serializing on row locks; the claiming UPDATE stamps the lease and + // counts the attempt atomically. query := ` UPDATE inbox_events - SET processed_at = CURRENT_TIMESTAMP, error = NULL - WHERE activity_id = $1` + SET claimed_until = CURRENT_TIMESTAMP + make_interval(secs => $1), + attempts = attempts + 1 + WHERE id = ( + SELECT e.id FROM inbox_events e + WHERE e.processed_at IS NULL + AND e.failed_at IS NULL + AND e.next_attempt_at <= CURRENT_TIMESTAMP + AND (e.claimed_until IS NULL OR e.claimed_until <= CURRENT_TIMESTAMP) + AND NOT EXISTS ( + SELECT 1 FROM inbox_events prior + WHERE prior.ordering_key = e.ordering_key + AND prior.id < e.id + AND prior.processed_at IS NULL + AND prior.failed_at IS NULL) + ORDER BY e.id + LIMIT 1 + FOR UPDATE SKIP LOCKED) + RETURNING ` + eventColumns - result, err := r.db.ExecContext(ctx, query, activityID) + event, err := scanInboxEvent(r.db.QueryRowContext(ctx, query, lease.Seconds())) if err != nil { - return fmt.Errorf("mark inbox event %q processed: %w", activityID, err) + if stderrors.Is(err, sql.ErrNoRows) { + return nil, errors.NewNotFoundError("inbox_event", "next claimable") + } + return nil, fmt.Errorf("claim next inbox event: %w", err) } - affected, err := result.RowsAffected() - if err != nil { - return fmt.Errorf("mark inbox event %q processed: rows affected: %w", activityID, err) + return event, nil +} + +func (r *postgresInboxEvents) MarkProcessed(ctx context.Context, activityID string, claimToken time.Time) (bool, error) { + // Fencing: only the worker that still holds the claim (claimed_until == + // claimToken, the value ClaimNext stamped) may record the outcome. A + // stale worker whose lease expired and was re-claimed by a second worker + // no longer matches claimed_until, writes 0 rows, and is reported as a + // no-op (applied=false) so it cannot clobber the newer attempt. The + // failed_at guard keeps processed and poisoned mutually exclusive, so a + // late finish can never stamp processed_at onto a poisoned row. + query := ` + WITH updated AS ( + UPDATE inbox_events + SET processed_at = CURRENT_TIMESTAMP, error = NULL, claimed_until = NULL + WHERE activity_id = $1 + AND processed_at IS NULL + AND failed_at IS NULL + AND claimed_until = $2 + RETURNING 1 + ) + SELECT + EXISTS (SELECT 1 FROM inbox_events WHERE activity_id = $1), + EXISTS (SELECT 1 FROM updated)` + + var exists, applied bool + if err := r.db.QueryRowContext(ctx, query, activityID, claimToken).Scan(&exists, &applied); err != nil { + return false, fmt.Errorf("mark inbox event %q processed: %w", activityID, err) } - if affected == 0 { - return errors.NewNotFoundError("inbox_event", activityID) + if !exists { + return false, errors.NewNotFoundError("inbox_event", activityID) } - return nil + return applied, nil +} + +func (r *postgresInboxEvents) Release(ctx context.Context, activityID, message string, nextAttempt time.Time, claimToken time.Time) (bool, error) { + if message == "" { + return false, errors.NewValidationError("error", "must not be empty") + } + + // Fencing + already-processed guard: only the current claim holder may + // reschedule (claimed_until == claimToken), and a stale worker's late + // release must not reschedule an event a successful retry completed + // (processed_at IS NULL). A mismatch writes 0 rows → applied=false. + query := ` + WITH updated AS ( + UPDATE inbox_events + SET error = $2, claimed_until = NULL, next_attempt_at = $3 + WHERE activity_id = $1 AND processed_at IS NULL AND claimed_until = $4 + RETURNING 1 + ) + SELECT + EXISTS (SELECT 1 FROM inbox_events WHERE activity_id = $1), + EXISTS (SELECT 1 FROM updated)` + + var exists, applied bool + if err := r.db.QueryRowContext(ctx, query, activityID, message, nextAttempt.UTC(), claimToken).Scan(&exists, &applied); err != nil { + return false, fmt.Errorf("release inbox event %q: %w", activityID, err) + } + if !exists { + return false, errors.NewNotFoundError("inbox_event", activityID) + } + return applied, nil +} + +func (r *postgresInboxEvents) MarkPoisoned(ctx context.Context, activityID, message string, claimToken time.Time) (bool, error) { + if message == "" { + return false, errors.NewValidationError("error", "must not be empty") + } + + // Fencing + already-processed guard: only the current claim holder may + // poison (claimed_until == claimToken), and a stale worker must not + // poison an event a successful retry completed (processed_at IS NULL). A + // mismatch writes 0 rows → applied=false. + query := ` + WITH updated AS ( + UPDATE inbox_events + SET error = $2, claimed_until = NULL, + failed_at = COALESCE(failed_at, CURRENT_TIMESTAMP) + WHERE activity_id = $1 AND processed_at IS NULL AND claimed_until = $3 + RETURNING 1 + ) + SELECT + EXISTS (SELECT 1 FROM inbox_events WHERE activity_id = $1), + EXISTS (SELECT 1 FROM updated)` + + var exists, applied bool + if err := r.db.QueryRowContext(ctx, query, activityID, message, claimToken).Scan(&exists, &applied); err != nil { + return false, fmt.Errorf("poison inbox event %q: %w", activityID, err) + } + if !exists { + return false, errors.NewNotFoundError("inbox_event", activityID) + } + return applied, nil } func (r *postgresInboxEvents) MarkFailed(ctx context.Context, activityID string, message string) error { @@ -91,20 +223,30 @@ func (r *postgresInboxEvents) MarkFailed(ctx context.Context, activityID string, } func (r *postgresInboxEvents) GetEvent(ctx context.Context, activityID string) (*InboxEvent, error) { - query := ` - SELECT id, activity_id, type, received_at, processed_at, COALESCE(error, '') - FROM inbox_events WHERE activity_id = $1` + query := `SELECT ` + eventColumns + ` FROM inbox_events WHERE activity_id = $1` - var event InboxEvent - err := r.db.QueryRowContext(ctx, query, activityID).Scan( - &event.ID, &event.ActivityID, &event.Type, - &event.ReceivedAt, &event.ProcessedAt, &event.Error, - ) + event, err := scanInboxEvent(r.db.QueryRowContext(ctx, query, activityID)) if err != nil { if stderrors.Is(err, sql.ErrNoRows) { return nil, errors.NewNotFoundError("inbox_event", activityID) } return nil, fmt.Errorf("get inbox event %q: %w", activityID, err) } + return event, nil +} + +func scanInboxEvent(row rowScanner) (*InboxEvent, error) { + var event InboxEvent + err := row.Scan( + &event.ID, &event.ActivityID, &event.Type, &event.Payload, &event.ActorID, + &event.OrderingKey, &event.Attempts, &event.NextAttemptAt, &event.ClaimedUntil, + &event.FailedAt, &event.ReceivedAt, &event.ProcessedAt, &event.Error, + ) + if err != nil { + return nil, err + } + if len(event.Payload) == 0 { + event.Payload = nil + } return &event, nil } diff --git a/internal/store/inbox_events_test.go b/internal/store/inbox_events_test.go index 4f12a54..54111bf 100644 --- a/internal/store/inbox_events_test.go +++ b/internal/store/inbox_events_test.go @@ -3,6 +3,7 @@ package store import ( "context" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -40,14 +41,23 @@ func TestInboxEvents_MarkProcessed(t *testing.T) { _, err := repo.RecordEvent(ctx, testActivityID, "Announce") require.NoError(t, err) - require.NoError(t, repo.MarkProcessed(ctx, testActivityID)) + // The worker must hold the claim to record an outcome: the fencing token + // is the ClaimedUntil ClaimNext stamped. + claimed, err := repo.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.NotNil(t, claimed.ClaimedUntil) + token := *claimed.ClaimedUntil + + applied, err := repo.MarkProcessed(ctx, testActivityID, token) + require.NoError(t, err) + assert.True(t, applied, "the claim holder's outcome must be applied") event, err := repo.GetEvent(ctx, testActivityID) require.NoError(t, err) assert.NotNil(t, event.ProcessedAt) assert.Empty(t, event.Error) - err = repo.MarkProcessed(ctx, "https://lemmy.world/activities/missing") + _, err = repo.MarkProcessed(ctx, "https://lemmy.world/activities/missing", token) assert.True(t, errors.IsNotFound(err), "expected IsNotFound, got %v", err) } @@ -58,6 +68,13 @@ func TestInboxEvents_MarkFailedThenRecovers(t *testing.T) { _, err := repo.RecordEvent(ctx, testActivityID, "Create") require.NoError(t, err) + // Claim so the event is leased (MarkFailed leaves the lease intact, so a + // later MarkProcessed by the same worker still holds a valid token). + claimed, err := repo.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.NotNil(t, claimed.ClaimedUntil) + token := *claimed.ClaimedUntil + require.NoError(t, repo.MarkFailed(ctx, testActivityID, "parent object fetch timed out")) event, err := repo.GetEvent(ctx, testActivityID) @@ -66,7 +83,9 @@ func TestInboxEvents_MarkFailedThenRecovers(t *testing.T) { assert.Equal(t, "parent object fetch timed out", event.Error) // A later successful retry clears the error. - require.NoError(t, repo.MarkProcessed(ctx, testActivityID)) + applied, err := repo.MarkProcessed(ctx, testActivityID, token) + require.NoError(t, err) + assert.True(t, applied) event, err = repo.GetEvent(ctx, testActivityID) require.NoError(t, err) assert.NotNil(t, event.ProcessedAt) @@ -82,7 +101,12 @@ func TestInboxEvents_MarkFailedOnProcessedEventIsNoOp(t *testing.T) { _, err := repo.RecordEvent(ctx, testActivityID, "Create") require.NoError(t, err) - require.NoError(t, repo.MarkProcessed(ctx, testActivityID)) + claimed, err := repo.ClaimNext(ctx, time.Minute) + require.NoError(t, err) + require.NotNil(t, claimed.ClaimedUntil) + applied, err := repo.MarkProcessed(ctx, testActivityID, *claimed.ClaimedUntil) + require.NoError(t, err) + require.True(t, applied) // A late failure report from a stale worker must not un-process an // event a successful retry already completed: no-op success. diff --git a/internal/store/interfaces.go b/internal/store/interfaces.go index 6f8a597..40d6d34 100644 --- a/internal/store/interfaces.go +++ b/internal/store/interfaces.go @@ -55,6 +55,13 @@ type APObjects interface { // preserves the original tombstone time; a missing mapping is an error // satisfying errors.IsNotFound. SoftDelete(ctx context.Context, apID string) error + + // Restore clears a soft delete (Undo{Delete}/restore): the explicit + // counterpart task 05 required so an un-deleted object can be + // re-materialized (commitRecord refuses to resurrect a mapping whose + // deleted_at is set). Restoring a live mapping is a no-op success; a + // missing mapping is an error satisfying errors.IsNotFound. + Restore(ctx context.Context, apID string) error } // BridgedActors registers fediverse actors bridged into atproto and their @@ -145,18 +152,59 @@ type ServiceKeys interface { Get(ctx context.Context, name string) (*ServiceKey, error) } -// InboxEvents deduplicates inbound AP activities and records processing -// outcomes. The queue-consumption side (ListPending and friends) is -// deliberately deferred to task 06, which owns the processing loop. +// InboxEvents deduplicates inbound AP activities and doubles as the durable +// postgres work queue the ingestion worker pool consumes (task 06). Rows are +// never deleted: a processed row IS the dedupe record for re-deliveries. type InboxEvents interface { // RecordEvent inserts the activity if it has not been seen before. // It returns isNew=false (and no error) when the activity id was // already recorded — the caller should drop the duplicate delivery. RecordEvent(ctx context.Context, activityID, activityType string) (isNew bool, err error) - // MarkProcessed stamps the event as successfully processed and clears - // any recorded error. - MarkProcessed(ctx context.Context, activityID string) error + // Enqueue inserts a full queue item (ActivityID, Type, Payload, ActorID, + // OrderingKey are honored; the rest is defaulted). Like RecordEvent it + // returns isNew=false when the activity id was already recorded, so a + // duplicate delivery never re-enqueues work. + Enqueue(ctx context.Context, event InboxEvent) (isNew bool, err error) + + // ClaimNext atomically claims the oldest processable event and + // increments its attempt counter. An event is processable when it is + // unprocessed, not poisoned, past its next_attempt_at, unleased (or the + // lease expired), and — the per-community ordering guarantee — no older + // unprocessed, unpoisoned event shares its ordering key. An empty queue + // returns an error satisfying errors.IsNotFound. + // + // The returned event's ClaimedUntil is the fencing/claim token for this + // claim: MarkProcessed/Release/MarkPoisoned require it so a worker whose + // lease expired and was re-claimed by another cannot overwrite the newer + // attempt's outcome. + ClaimNext(ctx context.Context, lease time.Duration) (*InboxEvent, error) + + // MarkProcessed stamps the event as successfully processed and clears any + // recorded error. claimToken must equal the ClaimedUntil returned by the + // ClaimNext that produced this attempt. It returns applied=true when the + // outcome was written; applied=false (no error) means the claim was stale + // — the lease expired and another worker re-claimed the event, or the row + // is already in a terminal state — so this outcome was discarded. A + // missing event is an error satisfying errors.IsNotFound. + MarkProcessed(ctx context.Context, activityID string, claimToken time.Time) (applied bool, err error) + + // Release records a processing failure and schedules the retry: error + // message stored, lease cleared, next_attempt_at set. claimToken must + // equal the claim's ClaimedUntil. It returns applied=true when the retry + // was scheduled; applied=false (no error) means the claim was stale or + // the event already completed, so the release was discarded. A missing + // event is an error satisfying errors.IsNotFound. + Release(ctx context.Context, activityID, message string, nextAttempt time.Time, claimToken time.Time) (applied bool, err error) + + // MarkPoisoned permanently fails the event: failed_at stamped, error + // recorded, lease cleared. Poisoned events are skipped by ClaimNext and + // stop blocking their ordering key. claimToken must equal the claim's + // ClaimedUntil. It returns applied=true when the event was poisoned; + // applied=false (no error) means the claim was stale or the event already + // completed, so the poison was discarded. A missing event is an error + // satisfying errors.IsNotFound. + MarkPoisoned(ctx context.Context, activityID, message string, claimToken time.Time) (applied bool, err error) // MarkFailed records a processing error on an unprocessed event, // leaving it unprocessed so it can be retried or inspected. The message @@ -168,3 +216,19 @@ type InboxEvents interface { // GetEvent returns the event for an activity id. GetEvent(ctx context.Context, activityID string) (*InboxEvent, error) } + +// Tombstones remembers AP object ids whose Delete arrived before (or +// without) a materialization — the create-after-delete gap: a Create +// delivered after its Delete must not resurrect content the origin removed. +// Undo{Delete} removes the marker. +type Tombstones interface { + // Record idempotently marks an AP id as deleted upstream. + Record(ctx context.Context, apID string) error + + // Exists reports whether the AP id carries a tombstone marker. + Exists(ctx context.Context, apID string) (bool, error) + + // Remove clears the marker (Undo{Delete}/restore). Removing a missing + // marker is a no-op success. + Remove(ctx context.Context, apID string) error +} diff --git a/internal/store/models.go b/internal/store/models.go index 14b8c05..d4ed204 100644 --- a/internal/store/models.go +++ b/internal/store/models.go @@ -143,12 +143,33 @@ type ServiceKey struct { CreatedAt time.Time } -// InboxEvent is a received AP activity, recorded for dedupe and -// processing bookkeeping. +// InboxEvent is a received AP activity: the dedupe record AND the durable +// work-queue item task 06's worker pool consumes. type InboxEvent struct { - ID int64 - ActivityID string // AP activity id (URL) — the dedupe key - Type string // AP activity type: Announce, Create, Like, ... + ID int64 + ActivityID string // AP activity id (URL) — the dedupe key + Type string // AP activity type: Announce, Create, Like, ... + // Payload is the raw activity JSON exactly as delivered (the + // signature-verified request body). Nil on legacy rows. + Payload []byte + // ActorID is the AP actor the activity is bound to; the inbox verified + // that it shares authority with the HTTP-signature signer. + ActorID string + // OrderingKey serializes processing: events sharing a key (normally the + // community IRI) are handled strictly in arrival order. + OrderingKey string + // Attempts counts how many times a worker claimed this event. + Attempts int + // NextAttemptAt is the retry-backoff schedule; claimable when <= now. + NextAttemptAt time.Time + // ClaimedUntil is the current worker lease; nil/past means unclaimed. It + // also serves as the fencing/claim token: ClaimNext stamps a fresh value + // on every claim, and MarkProcessed/Release/MarkPoisoned require it so a + // worker whose lease expired cannot overwrite a re-claim's outcome. + ClaimedUntil *time.Time + // FailedAt marks a poisoned event: permanently failed, skipped by the + // queue, no longer blocking its ordering key. + FailedAt *time.Time ReceivedAt time.Time ProcessedAt *time.Time Error string // last processing error; empty if none diff --git a/internal/store/tombstones.go b/internal/store/tombstones.go new file mode 100644 index 0000000..ed5b4db --- /dev/null +++ b/internal/store/tombstones.go @@ -0,0 +1,55 @@ +package store + +import ( + "context" + "database/sql" + "fmt" + + "tidepool/internal/errors" +) + +type postgresTombstones struct { + db *sql.DB +} + +// NewTombstones creates the postgres-backed ap_tombstones repository. +func NewTombstones(db *sql.DB) Tombstones { + return &postgresTombstones{db: db} +} + +func (r *postgresTombstones) Record(ctx context.Context, apID string) error { + if apID == "" { + return errors.NewValidationError("ap_id", "must not be empty") + } + query := ` + INSERT INTO ap_tombstones (ap_id) + VALUES ($1) + ON CONFLICT (ap_id) DO NOTHING` + if _, err := r.db.ExecContext(ctx, query, apID); err != nil { + return fmt.Errorf("record tombstone %q: %w", apID, err) + } + return nil +} + +func (r *postgresTombstones) Exists(ctx context.Context, apID string) (bool, error) { + if apID == "" { + return false, errors.NewValidationError("ap_id", "must not be empty") + } + var exists bool + query := `SELECT EXISTS (SELECT 1 FROM ap_tombstones WHERE ap_id = $1)` + if err := r.db.QueryRowContext(ctx, query, apID).Scan(&exists); err != nil { + return false, fmt.Errorf("check tombstone %q: %w", apID, err) + } + return exists, nil +} + +func (r *postgresTombstones) Remove(ctx context.Context, apID string) error { + if apID == "" { + return errors.NewValidationError("ap_id", "must not be empty") + } + if _, err := r.db.ExecContext(ctx, + `DELETE FROM ap_tombstones WHERE ap_id = $1`, apID); err != nil { + return fmt.Errorf("remove tombstone %q: %w", apID, err) + } + return nil +}