From 3b1030ec4278843e9f3c17f3b3ad25d92aafb47c Mon Sep 17 00:00:00 2001 From: Bretton Date: Sat, 15 Aug 2026 18:04:53 -0700 Subject: [PATCH 1/3] =?UTF-8?q?test(e2e):=20a=20reference=20PDS=20joins=20?= =?UTF-8?q?the=20stack=20=E2=80=94=20the=20first=20repo=20host=20we=20didn?= =?UTF-8?q?'t=20write?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every atproto repo in the harness was one we implement, so a spec misreading shared by our virtual PDS and our consumer was invisible: both halves agree, and the suite goes green on the agreement rather than the spec. ghcr.io/bluesky-social/pds (Coves CI's digest pin) now hosts a native test user, crawled by the stack's BigSky through the same validation every bridged repo crosses. The PDS shares the relay's network namespace (network_mode service:relay): bigsky's createExternalUser keys the peering on the DID DOCUMENT's service host and dials it from its own netns (bgs.go:1013-1045, with an explicit localhost: http special-case), so PDS_HOSTNAME=localhost is truthful only from inside that namespace — a hostname-per-service shape fails silently (fedmgr logs, cursor advances, RepoCount stays 0). Ports therefore publish on the relay (3081; Coves dev owns 3001), and the bootstrap one-shot gates on createSession, not createAccount, because one-shots re-run on every e2e-up against a persisting account. Smoke scenario: postv2 written via the reference PDS is observed on Jetstream and in relay repo state (getLatestCommit). Full suite green around it; the global vetEvent-whitelist tripwire for future native- repo scenarios is documented in the test header and FOLLOWUPS. Co-Authored-By: Claude Fable 5 --- FOLLOWUPS.md | 35 +++++ Makefile | 2 +- README.md | 68 +++++++--- docker-compose.e2e.yml | 204 +++++++++++++++++++++++++++- tests/e2e/native_pds_test.go | 251 +++++++++++++++++++++++++++++++++++ 5 files changed, 539 insertions(+), 21 deletions(-) create mode 100644 tests/e2e/native_pds_test.go diff --git a/FOLLOWUPS.md b/FOLLOWUPS.md index d01be9d..1c3410b 100644 --- a/FOLLOWUPS.md +++ b/FOLLOWUPS.md @@ -454,6 +454,41 @@ Still open, unchanged: - The PLC directory image is pinned by commit in `e2e/plc/Dockerfile`; bump it deliberately when upstream fixes are needed. +- **The stack now carries a reference PDS** (`ghcr.io/bluesky-social/pds`, + digest-pinned to the build Coves' CI pins; host port 127.0.0.1:3081, + published on the `relay` service because the PDS shares its network + namespace). One smoke scenario uses it — + `tests/e2e/native_pds_test.go` writes a `community.postv2` into the native + account's own repo and follows it through relay → Jetstream — which buys + the wire, not the record: a repo host we did not write, crossing the same + infrastructure, so a misreading of the spec shared by our commit writer and + our commit reader can no longer stay invisible. The task-18 scenarios + (native posts, acceptance records, votes both ways, Lemmy-side visibility) + are still open; this is deliverable 1 only. +- **TRIPWIRE — the collection whitelist is global.** `vetEvent` + (`tests/e2e/helpers.go`) checks every event on the firehose against a + single `expectedCollections` allowlist, whatever repo it came from, and a + violation fails the WHOLE suite (including the end-of-run sweep, which + re-vets the replay). `community.postv2` is already on that list, which is + the only reason the native repo passes today. **The first scenario that + writes any other collection into the reference PDS's repo takes every other + scenario down with it** — `bridge.federation`, `feed.vote`, + `actor.profile` on a native DID, anything. Rescoping the whitelist by repo class (bridge-hosted + author DIDs / bridge-hosted community DIDs / native-PDS DIDs, each with + independent rev tracking) is task 18's sweep item and was deliberately left + out of the reference-PDS landing. Related: the e2e stack leaves + `CONSUMER_ENABLED` at its default of false, so the acceptance engine never + reacts to native records — a scenario that needs an acceptance record must + turn the consumer on and point `JETSTREAM_URL` at the compose Jetstream + first. +- The relay does not survive being recreated on its own: its disk persister + records log-file refs in `relay-postgres` (which persists) while the files + live in the container's writable layer (which does not), so bigsky exits at + boot with `open /data/persister/evts-0: no such file or directory`. Any + edit to the `relay` service in `docker-compose.e2e.yml` triggers exactly + that on the next `up`; `make e2e-down` and start clean. A named volume for + the relay's `/data` would remove the trap, at the cost of one more thing + `down -v` has to take away in lockstep with the databases. - Jetstream exits when its upstream disconnects; compose's `restart: unless-stopped` supplies recovery. Remove the workaround if Jetstream gains reconnect support. diff --git a/Makefile b/Makefile index 5b0e535..2bb1b18 100644 --- a/Makefile +++ b/Makefile @@ -117,7 +117,7 @@ e2e: ## Full e2e run: build + start the stack, run the suite, tear down (-v) e2e-up: ## Start the e2e stack and leave it running (for iterating on tests) @$(E2E_COMPOSE) up -d --build --wait --wait-timeout 600 - @echo "$(GREEN)✓ e2e stack up: tidepool 127.0.0.1:8092, lemmy 127.0.0.1:8541, relay 127.0.0.1:2480, jetstream 127.0.0.1:6028$(RESET)" + @echo "$(GREEN)✓ e2e stack up: tidepool 127.0.0.1:8092, lemmy 127.0.0.1:8541, relay 127.0.0.1:2480, reference pds 127.0.0.1:3081, jetstream 127.0.0.1:6028$(RESET)" e2e-test: ## Run the e2e suite against an already-running stack (make e2e-up) @go test -tags e2e -count=1 -v -timeout 20m ./tests/e2e/... diff --git a/README.md b/README.md index ab43d70..c1cfdfa 100644 --- a/README.md +++ b/README.md @@ -87,8 +87,8 @@ your directory is on a different port. The end-to-end harness runs the whole read path against **real infrastructure** — a real Lemmy federating with the bridge, a real did:plc directory backing DID minting, a real atproto relay (indigo **BigSky**) -crawling the bridge, and a real Jetstream decoding the **relay's** firehose -— in one compose network: +crawling the bridge, a real reference PDS hosting a native account, and a +real Jetstream decoding the **relay's** firehose — in one compose network: ``` docker-compose.e2e.yml @@ -103,9 +103,13 @@ crawling the bridge, and a real Jetstream decoding the **relay's** firehose │ ┌─────────────▼──┐ │ │request │ │ │ plc (did:plc │◀────┐ │ │Crawl │ │ │ + plc-postgres)│ │ ┌──▼─┴───────┐ │ - │ └────────────────┘ DID │ │ relay │ │ - │ resolution └──│ (BigSky) │ │ - │ │ + postgres │ │ + │ └──┬─────────────┘ DID │ │ relay │ │ + │ ▲ resolution └──│ (BigSky) │ │ + │ │ │ + postgres │ │ + │ ┌──┴─────────────┐ │ + the │ │ + │ │ pds (reference │ CBOR │ reference │ │ + │ │ @atproto/pds) ├───────▶│ pds, in │ │ + │ └────────────────┘ frames │ this netns │ │ │ └──────┬─────┘ │ │ CBOR frames│ │ │ ┌───────▼─────┐ │ @@ -114,7 +118,7 @@ crawling the bridge, and a real Jetstream decoding the **relay's** firehose │ └───────┬─────┘ │ └───────────────────────────────────────────────────────────────┼───────┘ host (127.0.0.1 only): tidepool :8092, lemmy :8541, │ - relay :2480, jetstream :6028 ◀──────┘ + relay :2480, pds :3081, jetstream :6028 ◀──────┘ tests/e2e (go test -tags e2e) ``` @@ -130,14 +134,37 @@ relay's new-PDS-per-day limit is 0 (refuses all non-admin `requestCrawl`), so a one-shot `relay-bootstrap` service raises the limit over the admin API before the bridge starts — the announcement itself is still bridge-originated, and the suite asserts it landed (`relay_test.go`). -Bridged handles verify for real too: `HANDLE_RESOLVER_HOSTS=tidepool` -points bigsky's trial-host resolver at the bridge's Host-header-keyed -`/.well-known/atproto-did` (compose-DNS-invisible names like -`alice.lemmy.tidepool` would otherwise fail handle verification, which -bigsky treats as non-fatal). For debugging, a direct bridge→Jetstream tap +Bridged handles verify for real too: +`HANDLE_RESOLVER_HOSTS=tidepool,localhost:3001` points bigsky's trial-host +resolver at the bridge's Host-header-keyed `/.well-known/atproto-did`, and +at the reference PDS's (below) for its own handles (compose-DNS-invisible +names like `alice.lemmy.tidepool` would otherwise fail handle verification, +which bigsky treats as non-fatal). For debugging, a direct bridge→Jetstream tap exists behind the `direct` compose profile (`jetstream-direct`, host port 6038) — it is not the tested path. +The relay also crawls a **reference PDS** — `ghcr.io/bluesky-social/pds`, +pinned by digest — hosting a native account (`native-alice.pds.test`) whose +DID is minted in the same local PLC. It is there because every other repo on +this firehose is the bridge's: our commits, read back by our own +understanding of the spec at both ends, so a misreading shared by the writer +and the reader is invisible and the suite goes green on the agreement. A +repo host we did not write cannot share it. The `pds` container runs in the +**relay's network namespace** (`network_mode: "service:relay"`), which is +not a convenience: `@atproto/pds` derives its public URL from +`PDS_HOSTNAME`, special-casing `localhost` to `http://localhost:$PDS_PORT` +and turning every other hostname into `https://`, which this plain-HTTP +harness cannot serve — and bigsky peers a PDS by the host in its DID +*documents*, calling `describeServer` at that url from inside its own +namespace. Sharing the namespace is what makes `http://localhost:3001` true +for both. Consequences: the PDS's host port is published on the `relay` +service (a container-network-mode service cannot publish its own), and it is +3081 rather than 3001 because the Coves dev stack already owns 3001. A +one-shot `pds-bootstrap` creates the account over the plain public +`createAccount` (no admin API — the PDS runs with invites off), treats an +already-taken handle as success, gates on `createSession` returning 200, and +then `requestCrawl`s `localhost:3001` at the relay. + One ordering consequence worth knowing: bigsky indexes its inbound firehose with a parallel scheduler keyed by repo DID, so **per-repo event order survives the relay but cross-repo order does not**. Since the author-owned @@ -152,7 +179,7 @@ proves the rule: they sat in the community's repo with no separate attestation, so there was no cross-repo pair to reorder — those records still exist and are not migrated. -The host ports bind **loopback-only** (`127.0.0.1:8092/8541/2480/6028`): +The host ports bind **loopback-only** (`127.0.0.1:8092/8541/2480/3081/6028`): the stack carries admin tokens and runs with `ALLOW_PRIVATE_FETCH=1`, so it must not be reachable from the local network. @@ -176,7 +203,7 @@ plain-HTTP AP ids to match. See the header comments in `docker-compose.e2e.yml` and `e2e/lemmy/Dockerfile` for the full story. The suite (`tests/e2e/`, build tag `e2e`: `bridge_test.go`, -`lifecycle_test.go`, `media_test.go`, `relay_test.go`, +`lifecycle_test.go`, `media_test.go`, `native_pds_test.go`, `relay_test.go`, `votes_hammer_test.go`, `zz_sweep_test.go`) covers: subscribe → `community.profile` on the firehose; a link post → `actor.profile` and `community.post` (presence + author linkage; arrival @@ -227,11 +254,22 @@ sentinel alone. Every negative assertion ("nothing bridged") is bounded by a positive control in the same window — never a bare sleep-and-assert-nothing. +Task 18 adds the **native-PDS smoke scenario** (`native_pds_test.go`): the +suite authenticates as the bootstrap account on the reference PDS, writes +one `community.postv2` into that account's own repo, and follows it through +the relay to Jetstream and into the relay's own `getLatestCommit` state. +Deliberately a smoke test, not a feature test — what it buys is the wire, +not the record. Note before extending it: `vetEvent`'s collection whitelist +is global, so a second scenario writing any other collection into the native +repo fails the whole suite until that whitelist is rescoped by repo class +(see FOLLOWUPS.md). + Every create/update the tests consume from Jetstream has passed the relay's signature verification AND is validated against the vendored Coves lexicons -on the consumer side of the wire, and any collection outside the four the -bridge emits fails the suite immediately (votes must never become records). Two scripts keep the vendored +on the consumer side of the wire, and any collection outside the +`expectedCollections` allowlist fails the suite immediately, in any repo +(votes must never become records). Two scripts keep the vendored lexicons honest: `scripts/sync-lexicons.sh` copies them from a Coves checkout; `scripts/check-lexicons.sh` verifies the committed manifest and byte-compares against the current Coves checkout. CI clones the canonical Coves diff --git a/docker-compose.e2e.yml b/docker-compose.e2e.yml index c6e47b2..34a7ed4 100644 --- a/docker-compose.e2e.yml +++ b/docker-compose.e2e.yml @@ -2,11 +2,13 @@ # did:plc directory backing DID minting, a real atproto relay (indigo # BigSky) crawling the bridge like a hostile consumer — DID resolution # against the local PLC, per-commit signature verification, getRepo crawls — -# and a real Jetstream decoding the RELAY's firehose. `make e2e` builds it, +# a real reference PDS (@atproto/pds) hosting a native account, and a real +# Jetstream decoding the RELAY's firehose. `make e2e` builds it, # waits for health, runs tests/e2e (tagged `e2e`) from the host, and tears # it down. The tested pipeline is: # # Lemmy ──AP──▶ tidepool ──subscribeRepos──▶ relay (BigSky) ──▶ jetstream +# pds ──subscribeRepos──▶ ┘ # # so every scenario's records transit relay validation for free. A direct # bridge→Jetstream tap exists behind the `direct` compose profile for @@ -80,11 +82,17 @@ # 127.0.0.1:8092 Tidepool HTTP (admin API, XRPC, healthz) # 127.0.0.1:8541 Lemmy HTTP API (8541 nods to lemmy_alpha in upstream's federation compose) # 127.0.0.1:2480 Relay (BigSky) XRPC + admin API +# 127.0.0.1:3081 Reference PDS XRPC — PUBLISHED ON THE `relay` SERVICE (the +# pds shares its network namespace and so cannot publish its +# own ports); 3081 rather than the PDS's natural 3001 +# because the Coves dev stack already owns 3001 here # 127.0.0.1:6028 Jetstream WebSocket (/subscribe) # # LOCAL-ONLY: nothing here talks to plc.directory, public relays, or public -# Lemmy instances. All egress stays on the compose network (image pulls and -# build-time package/source downloads aside). +# Lemmy instances — the reference PDS mints its DIDs against the compose PLC +# (PDS_DID_PLC_URL) and announces itself only to the compose relay. All +# egress stays on the compose network (image pulls and build-time +# package/source downloads aside). name: tidepool-e2e @@ -263,17 +271,33 @@ services: # bridge's own /.well-known/atproto-did (Host-header keyed) — bigsky's # trial-host resolver speaks exactly that. Without this, handle # verification fails non-fatally (repo kept, handle marked invalid). - HANDLE_RESOLVER_HOSTS: tidepool + # Comma-separated (a StringSliceFlag): the second host is the reference + # PDS, which serves the same well-known for its `.pds.test` handles and + # is reachable at localhost:3001 from THIS namespace (see the `pds` + # service). Nice-to-have, not a gate — the failure mode is non-fatal. + HANDLE_RESOLVER_HOSTS: tidepool,localhost:3001 RELAY_ADMIN_KEY: e2e-relay-admin-key # Default is 100 repos per PDS — a long-lived stack under repeated # suite runs mints past that. RELAY_DEFAULT_REPO_LIMIT: "10000" # Disk persister gives real event playback (with TTL) for downstream - # cursor replay instead of the DB persister's unbounded table. + # cursor replay instead of the DB persister's unbounded table. Note + # that it keeps its log-file REFS in relay-postgres while the files + # themselves live in this container's writable layer, so recreating + # this container alone (any edit to this service does) makes bigsky + # exit at boot on "open /data/persister/evts-0: no such file or + # directory". Recreate the stack with `make e2e-down` first. RELAY_PERSISTER_DIR: /data/persister BSKYLOG_LOG_LEVEL: info ports: - "127.0.0.1:${RELAY_E2E_PORT:-2480}:2470" + # The reference PDS's port, published HERE because the `pds` service + # runs in this container's network namespace and docker refuses + # `ports` on a container-network-mode service — at container CREATE + # time, which `docker compose config` does not catch. Publishing it + # here is the same listener: in a shared namespace there is one + # loopback and one port space (see the `pds` service header). + - "127.0.0.1:${PDS_E2E_PORT:-3081}:3001" networks: [tidepool-e2e] depends_on: relay-postgres: @@ -308,6 +332,165 @@ services: relay: condition: service_healthy + # ── Reference PDS (@atproto/pds — a repo host we did NOT write) ───────── + # + # Every other repo on this firehose is the bridge's: our commits, read back + # by our own understanding of the spec at both ends, so a shared misreading + # is invisible. This is the third party that cannot share it — a record it + # commits crosses the same relay, the same firehose and the same Jetstream + # as a bridged one (tests/e2e/native_pds_test.go). + # + # WHY IT SHARES THE RELAY'S NETWORK NAMESPACE + # @atproto/pds derives its public URL from PDS_HOSTNAME and special-cases + # `localhost` → http://localhost:${PDS_PORT}; ANY other hostname becomes + # https://…, which this plain-HTTP harness cannot serve. So the DID + # documents this PDS writes name http://localhost:3001 — and that string + # has to be TRUE for the relay, because bigsky's createExternalUser + # (indigo bgs/bgs.go:1013-1045) takes the service endpoint out of the DID + # DOCUMENT, looks the PDS peering up by THAT host, and calls + # describeServer against THAT url from inside its own namespace + # (bgs.go:1035-1037 special-cases a `localhost:` prefix back to http — + # indigo is built for exactly this topology), then refuses commits from a + # PDS that is not authoritative for the repo (bgs.go:823-836). + # `network_mode: "service:relay"` makes localhost:3001 the same listener + # for both containers, so the DID document tells the truth. The obvious + # alternative — a `pds` hostname of its own on the compose network — fails + # SILENTLY and expensively: the relay accepts the crawl, registers the + # host, then logs and advances its cursor forever (fedmgr.go:556-563) with + # RepoCount 0 and not one event delivered. + # + # Consequences encoded elsewhere in this file: + # * The host port is published on the `relay` service, not here. + # * `relay` must NOT depends_on this service — that is a cycle, since + # this service's namespace IS the relay's. + # * In-network callers reach this PDS at relay:3001; the crawl + # announcement says `localhost:3001`, because the relay keys the + # peering on the DID-doc host and that match IS the mechanism. + # * Restarting or recreating the RELAY container tears this container's + # networking down with it (docker cannot re-enter a namespace that + # was replaced) — recreate `pds` alongside it. The suite's restart + # scenario bounces only `tidepool`, so it is safe today. + # + # Unlike the relay image, this digest is multi-arch — arm64 hosts run it + # natively and it needs no `platform:` pin. + # + # Expected noise, not a fault: at account creation this PDS emits a #sync + # frame that the pinned bigsky predates, and the relay logs one + # `ERROR failed handling event … err="invalid fed event"` for it. The + # #identity and #account frames around it are consumed normally, the + # handle verifies, and every #commit that follows is indexed. + pds: + # Pinned by digest — the SAME digest Coves' CI stack pins + # (coves/docker-compose.ci.yml), so both repos exercise one PDS build. + # sha256:1fa8bbce… was resolved 2026-07-28 from ghcr.io/bluesky-social/pds:latest. + image: ghcr.io/bluesky-social/pds@sha256:1fa8bbceabb65d8e1710749b1ea92c1c20a7489ca38da4a0a5f64c0c10a70c29 + network_mode: "service:relay" + environment: + # localhost:3001 is a claim about the relay's namespace, not a + # loopback-only bind — see the header. Changing either value without + # the other breaks peering. + PDS_HOSTNAME: localhost + PDS_PORT: 3001 + PDS_DATA_DIRECTORY: /pds + PDS_BLOBSTORE_DISK_LOCATION: /pds/blocks + # LOCAL-ONLY: DIDs are minted in the compose PLC, not the public + # plc.directory (the default). Resolved through the relay's namespace, + # where the compose DNS name `plc` works exactly as it does elsewhere. + PDS_DID_PLC_URL: http://plc:3000 + PDS_JWT_SECRET: e2e-pds-jwt-secret + PDS_ADMIN_PASSWORD: e2e-pds-admin-password + # Distinct from Coves' key: two hosts sharing a rotation key could + # rewrite each other's identities, which is not the topology modelled. + PDS_PLC_ROTATION_KEY_K256_PRIVATE_KEY_HEX: a65e9c570ef61d87f7a4404d3b6a4763ed537b7683d9848b9a73208689c14bda + # `.test` is RFC-6761 reserved and can never resolve on the public + # internet — the same reasoning as the `.invalid` fixtures in the unit + # suite. native-alice.pds.test (the bootstrap account, and a constant + # in tests/e2e/native_pds_test.go) lives under it. + PDS_SERVICE_HANDLE_DOMAINS: .pds.test + PDS_DEV_MODE: "true" + # No invite codes, so pds-bootstrap creates its account over the plain + # public createAccount endpoint rather than the admin API — the same + # call any client would make. + PDS_INVITE_REQUIRED: "false" + NODE_ENV: development + LOG_ENABLED: "true" + LOG_LEVEL: info + volumes: + # Not durability — EXISTENCE. The image ships no /pds, and @atproto/pds + # opens its sqlite files in PDS_DATA_DIRECTORY without creating the + # directory first ("Cannot open database because the directory does not + # exist", verified on this digest). + - pds-data:/pds + depends_on: + # Not an ordering preference: the namespace this container joins does + # not exist until the relay's container does. + relay: + condition: service_healthy + plc: + condition: service_healthy + healthcheck: + test: ["CMD", "wget", "--spider", "-q", "http://localhost:3001/xrpc/_health"] + interval: 2s + timeout: 5s + retries: 30 + start_period: 20s + + # One-shot: create the native test account on the reference PDS, prove its + # credentials work, and tell the relay to crawl it. + # + # NOT the admin API — PDS_INVITE_REQUIRED is false, so this is the same + # public com.atproto.server.createAccount any client would call, which is + # the point: nothing about the native account is privileged. + # + # IDEMPOTENT BY CONSTRUCTION, because `make e2e-up` starts (and therefore + # re-runs) exited one-shots, and the second run always finds the handle + # taken. createAccount's failure is tolerated and the GATE is createSession: + # "these credentials authenticate" is the claim worth failing the stack on, + # and it is exactly the claim tests/e2e/native_pds_test.go opens with. It is + # also the only claim BusyBox wget can check — it cannot read a 4xx body, so + # "handle already taken" is indistinguishable from "the PDS rejected our + # configuration" at the createAccount call alone. + # + # Then requestCrawl, which is the PUBLIC sync endpoint (no admin bearer): + # relay-bootstrap has already raised the new-PDS-per-day limit that would + # otherwise refuse it, and the hostname announced is `localhost:3001` — + # the host in the DID documents this PDS writes, which is what the relay + # keys its peering on. The relay validates it by calling describeServer + # back at that url; no race to lose here, since `pds` is healthy first. + # + # Runs in its own namespace on the compose network, so the PDS is + # relay:3001 from here — the same listener, seen from outside. + pds-bootstrap: + image: ghcr.io/bluesky-social/pds@sha256:1fa8bbceabb65d8e1710749b1ea92c1c20a7489ca38da4a0a5f64c0c10a70c29 + entrypoint: ["/bin/sh", "-ec"] + command: + - | + echo '▶ creating native-alice.pds.test on the reference PDS' + wget -q -O- --header='Content-Type: application/json' \ + --post-data='{"handle":"native-alice.pds.test","email":"native-alice@pds.test","password":"native-alice-e2e-pass"}' \ + http://relay:3001/xrpc/com.atproto.server.createAccount \ + || echo ' createAccount failed — expected when the account already exists; createSession below is the real gate' + echo + echo '▶ authenticating as native-alice.pds.test (the gate)' + wget -q -O- --header='Content-Type: application/json' \ + --post-data='{"identifier":"native-alice.pds.test","password":"native-alice-e2e-pass"}' \ + http://relay:3001/xrpc/com.atproto.server.createSession > /dev/null + echo ' ✓ credentials valid' + echo '▶ announcing localhost:3001 (the DID-doc host) to the relay' + wget -q -O- --header='Content-Type: application/json' \ + --post-data='{"hostname":"localhost:3001"}' \ + http://relay:2470/xrpc/com.atproto.sync.requestCrawl + echo ' ✓ relay is crawling the reference PDS' + restart: "no" + networks: [tidepool-e2e] + depends_on: + pds: + condition: service_healthy + # A fresh relay refuses every non-admin requestCrawl until this + # one-shot raises new_pds_per_day_limit off its default of 0. + relay-bootstrap: + condition: service_completed_successfully + # ── Jetstream (consumes the RELAY's firehose — the tested pipeline) ───── jetstream: image: ghcr.io/bluesky-social/jetstream:sha-306e463693365e21a5ffd3ec051a5a7920000214 @@ -427,3 +610,14 @@ networks: tidepool-e2e: driver: bridge name: tidepool-e2e-network + +volumes: + # The reference PDS's repos, blobs and account database (why it must exist + # at all: see the `pds` service's volumes). Named rather than anonymous so + # the native account survives a `docker compose up` that recreates the + # container — `pds-bootstrap` is idempotent either way. `make e2e` / + # `make e2e-down` remove it with -v, in the same breath as the PLC's + # database: the account's did:plc is minted there, so keeping one without + # the other leaves a repo whose DID nothing can resolve, and the relay + # rejects every commit it signs. + pds-data: diff --git a/tests/e2e/native_pds_test.go b/tests/e2e/native_pds_test.go new file mode 100644 index 0000000..4e9dc87 --- /dev/null +++ b/tests/e2e/native_pds_test.go @@ -0,0 +1,251 @@ +//go:build e2e + +package e2e + +// Task 18 deliverable 1: the harness gains an atproto repo host we did NOT +// write. Every other scenario in this suite drives OUR PDS implementation — +// the bridge's virtual repos — through OUR reading of the atproto spec, at +// both ends. A shared misreading between the thing that emits commits and +// the thing that consumes them is therefore invisible: both halves agree, +// and the suite goes green on the agreement rather than on the spec. The +// reference PDS (ghcr.io/bluesky-social/pds) is the third party that cannot +// share our misreadings. When a record written by an implementation nobody +// here controls crosses the same relay, the same firehose, and the same +// Jetstream that the bridge's records cross, the pipeline has been proven +// against the protocol instead of against itself. +// +// This scenario is deliberately a smoke test, not a feature test: it writes +// one postv2 into a native repo and watches for it downstream. What it +// buys is the wire, not the record. +// +// TRIPWIRE for anyone extending this file: vetEvent (helpers.go) fails the +// WHOLE suite on any collection outside expectedCollections, globally. A +// second scenario writing some other collection into the native repo takes +// every other scenario down with it — rescoping that whitelist is task 18's +// sweep item and is out of scope here. + +import ( + "bytes" + "encoding/json" + "fmt" + "io" + "net/http" + "testing" + "time" + + "github.com/bluesky-social/indigo/atproto/syntax" +) + +// ── Bootstrap contract ───────────────────────────────────────────────────── + +// The reference PDS's bootstrap account. These three constants ARE the +// contract between this test and the compose stack: the `pds-bootstrap` +// one-shot in docker-compose.e2e.yml must create exactly this handle with +// exactly this password (com.atproto.server.createAccount, invites off), +// and the handle's domain must be one of the PDS's configured +// PDS_SERVICE_HANDLE_DOMAINS. Change either side without the other and the +// scenario fails at createSession with an authentication error rather than +// anything about federation. +const ( + nativeHandle = "native-alice.pds.test" + nativePassword = "native-alice-e2e-pass" +) + +// nativeCommunityDID is a syntactically valid did:plc that nothing in this +// stack resolves — the smoke record needs the postv2 lexicon's required +// `community` field to PARSE as a DID, and nothing more. No acceptance +// record is ever written for this post (CONSUMER_ENABLED is off in the e2e +// stack, so the acceptance engine never sees it), so the community is a +// name on a record, not a participant. +const nativeCommunityDID = "did:plc:e2enativepdssmokecommun2" + +// pdsURL is the reference PDS's host endpoint. It is published on the RELAY +// service's ports (the pds container shares the relay's network namespace, +// so it cannot publish its own), and on 3081 rather than the PDS's natural +// 3001 because the Coves dev stack already owns 3001 on this machine. +func pdsURL() string { return envOr("PDS_E2E_URL", "http://localhost:3081") } + +// ── Reference-PDS client ─────────────────────────────────────────────────── + +// pdsClient speaks the two XRPC endpoints this scenario needs against the +// reference PDS. Test-only, and deliberately not routed through any of +// tidepool's own atproto client code: a client that shared production's +// request-building would inherit production's assumptions about the server, +// which is the exact thing an outside implementation is here to test. +type pdsClient struct { + http *http.Client + accessJWT string + did string + handle string +} + +// do issues one JSON XRPC call, decoding into out when non-nil. Errors name +// the endpoint and carry the server's body: a reference PDS reports lexicon +// and auth failures in prose that is worth reading verbatim. +func (c *pdsClient) do(method, nsid string, body, out any) error { + var reqBody io.Reader + if body != nil { + raw, err := json.Marshal(body) + if err != nil { + return err + } + reqBody = bytes.NewReader(raw) + } + req, err := http.NewRequest(method, pdsURL()+"/xrpc/"+nsid, reqBody) + if err != nil { + return err + } + if body != nil { + req.Header.Set("Content-Type", "application/json") + } + if c.accessJWT != "" { + req.Header.Set("Authorization", "Bearer "+c.accessJWT) + } + resp, err := c.http.Do(req) + if err != nil { + return err + } + defer func() { _ = resp.Body.Close() }() + raw, err := io.ReadAll(resp.Body) + if err != nil { + return err + } + if resp.StatusCode/100 != 2 { + return fmt.Errorf("pds %s: status %d: %s", nsid, resp.StatusCode, truncate(raw, 300)) + } + if out != nil { + if err := json.Unmarshal(raw, out); err != nil { + return fmt.Errorf("pds %s: decode: %w (%s)", nsid, err, truncate(raw, 200)) + } + } + return nil +} + +// createSession authenticates as the bootstrap account and remembers the +// DID the PDS minted for it — the DID this scenario then follows through +// the relay. We never hardcode that DID: it is whatever the local PLC +// issued when the bootstrap ran, and only the PDS can tell us. +func (c *pdsClient) createSession(handle, password string) error { + var out struct { + DID string `json:"did"` + Handle string `json:"handle"` + AccessJWT string `json:"accessJwt"` + } + err := c.do(http.MethodPost, "com.atproto.server.createSession", + map[string]any{"identifier": handle, "password": password}, &out) + if err != nil { + return err + } + if out.DID == "" || out.AccessJWT == "" { + return fmt.Errorf("pds createSession(%s): session missing did/accessJwt (did=%q)", handle, out.DID) + } + c.did, c.handle, c.accessJWT = out.DID, out.Handle, out.AccessJWT + return nil +} + +// createRecord writes into the session's OWN repo. Returns the at-uri the +// PDS assigned, which the caller checks against the rkey it asked for. +func (c *pdsClient) createRecord(collection, rkey string, record map[string]any) (uri, cid string, err error) { + var out struct { + URI string `json:"uri"` + CID string `json:"cid"` + } + err = c.do(http.MethodPost, "com.atproto.repo.createRecord", map[string]any{ + "repo": c.did, + "collection": collection, + "rkey": rkey, + "record": record, + }, &out) + if err != nil { + return "", "", err + } + return out.URI, out.CID, nil +} + +// ── Scenario ─────────────────────────────────────────────────────────────── + +// TestNativePDS_RecordTransitsRelayToJetstream is the acceptance test for a +// reference PDS in the stack: a record written by @atproto/pds — an +// implementation this repo did not write and cannot quietly agree with — +// reaches the suite's Jetstream through the same relay every bridged commit +// crosses. +// +// The chain each assertion pins, in order: the reference PDS is reachable +// and the bootstrap account exists (createSession); the PDS accepts and +// commits our record (createRecord); the relay had already crawled the PDS +// and validated the commit's signature against the DID it resolved from the +// local PLC, and Jetstream decoded the relay's firehose (the awaited event — +// Jetstream's upstream IS the relay, so arrival there is the whole transit +// proven at once); and the relay built repo state from it rather than merely +// forwarding frames (getLatestCommit). +func TestNativePDS_RecordTransitsRelayToJetstream(t *testing.T) { + h := newHarness(t) + + pds := &pdsClient{http: h.http} + if err := pds.createSession(nativeHandle, nativePassword); err != nil { + t.Fatalf("createSession as the native bootstrap account %q at %s: %v\n"+ + "the e2e stack has no reference PDS (or its bootstrap did not run) — see docker-compose.e2e.yml: "+ + "without a repo host we did not implement, every commit this suite validates was produced and "+ + "consumed by our own reading of the spec, and a shared misreading stays invisible", + nativeHandle, pdsURL(), err) + } + t.Logf("native account: handle=%s did=%s (minted by the reference PDS against the local PLC)", pds.handle, pds.did) + + // One rkey per run, time-derived: the PDS's repo outlives a single + // `make e2e-test` (the stack is brought up separately and re-tested), + // and a fixed rkey would 400 as already-existing on the second run — + // which would look like a transit failure and is not one. + rkey := syntax.NewTIDNow(0).String() + + // Cursor BEFORE the write, so a listener that finishes dialing after the + // commit still replays it. + cursor := cursorNow() + l := h.newListener(t, cursor, colPostV2) + + record := map[string]any{ + "$type": colPostV2, + "community": nativeCommunityDID, + "title": "native pds smoke " + rkey, + "createdAt": time.Now().UTC().Format(time.RFC3339), + } + uri, cid, err := pds.createRecord(colPostV2, rkey, record) + if err != nil { + t.Fatalf("createRecord %s/%s in the native repo %s: %v", colPostV2, rkey, pds.did, err) + } + if want := fmt.Sprintf("at://%s/%s/%s", pds.did, colPostV2, rkey); uri != want { + t.Fatalf("the PDS committed the record at %q, not the %q we asked for — the rest of this scenario follows the rkey we chose", uri, want) + } + t.Logf("wrote %s (cid %s)", uri, cid) + + ev := l.await("native postv2 through the relay", func(e *jsEvent) bool { + return e.Did == pds.did && e.Commit.Collection == colPostV2 && + e.Commit.RKey == rkey && e.Commit.Operation == opCreate + }) + if got := recordField(t, ev.Commit.Record, "community"); got != nativeCommunityDID { + t.Errorf("jetstream delivered community %q, want %q — the record changed shape in transit", got, nativeCommunityDID) + } + + // Relay-side state, not just passthrough: the relay indexes + // asynchronously, so poll it out — only the deadline is fatal. + deadline := time.Now().Add(eventTimeout) + var lastErr error + for { + headCID, rev, err := h.relayGetLatestCommit(pds.did) + if err != nil { + lastErr = err + t.Logf("relay getLatestCommit(%s): %v (retrying)", pds.did, err) + } else if headCID != "" && rev != "" { + if rev < ev.Commit.Rev { + t.Errorf("relay head rev %q is OLDER than the rev %q it already emitted for this repo — relay state lagging its own firehose", rev, ev.Commit.Rev) + } + t.Logf("relay serves the native repo: head=%s rev=%s", headCID, rev) + return + } + if time.Now().After(deadline) { + t.Fatalf("the relay never served a commit for the native repo %s within %s (last error: %v) — "+ + "the event reached Jetstream, so the relay forwarded it without building repo state from it", + pds.did, eventTimeout, lastErr) + } + time.Sleep(time.Second) + } +} -- 2.51.2 From 2bd2552821e3397b4f7ce3b2f837b8ea0cb596c2 Mon Sep 17 00:00:00 2001 From: Bretton Date: Sat, 15 Aug 2026 20:01:20 -0700 Subject: [PATCH 2/3] fix(e2e): make the smoke test's failures name the broken hop, and the docs stop overclaiming MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Second-opinion findings. The relay-head poll now runs BEFORE the Jetstream await (peering proof first — a dead crawl fails there, naming pds-bootstrap, instead of surfacing as a generic transit timeout), treats a stale rev as keep-polling rather than an instant failure (the repo carries a pre-existing head across runs, so the old branch could fail on a one-second indexing lag), and never reports "last error: ". The Jetstream assertion pins the commit CID against what createRecord returned — content addressing covers the whole record, not one field. waitForStack probes the reference PDS, making its own docstring true. Docs: the whitelist tripwire is now stated with both halves (fails closed on unknown collections, fails OPEN on allowlisted collections in the wrong repo class — the actor.profile example was exactly backwards); bootstrap echoes admit what they cannot know (createAccount's cause) and claim only what a 200 proves (crawl accepted, not crawling); stale "four-collection" and stack-inventory comments updated; indigo citations anchored to the snapshot they were read at. Co-Authored-By: Claude Fable 5 --- FOLLOWUPS.md | 29 +++++----- README.md | 11 ++-- docker-compose.e2e.yml | 43 +++++++++----- tests/e2e/helpers.go | 25 ++++++-- tests/e2e/native_pds_test.go | 109 +++++++++++++++++++++++++---------- 5 files changed, 150 insertions(+), 67 deletions(-) diff --git a/FOLLOWUPS.md b/FOLLOWUPS.md index 1c3410b..1406a5d 100644 --- a/FOLLOWUPS.md +++ b/FOLLOWUPS.md @@ -455,7 +455,7 @@ Still open, unchanged: - The PLC directory image is pinned by commit in `e2e/plc/Dockerfile`; bump it deliberately when upstream fixes are needed. - **The stack now carries a reference PDS** (`ghcr.io/bluesky-social/pds`, - digest-pinned to the build Coves' CI pins; host port 127.0.0.1:3081, + digest-pinned to the same build Coves' CI pins; host port 127.0.0.1:3081, published on the `relay` service because the PDS shares its network namespace). One smoke scenario uses it — `tests/e2e/native_pds_test.go` writes a `community.postv2` into the native @@ -465,18 +465,21 @@ Still open, unchanged: our commit reader can no longer stay invisible. The task-18 scenarios (native posts, acceptance records, votes both ways, Lemmy-side visibility) are still open; this is deliverable 1 only. -- **TRIPWIRE — the collection whitelist is global.** `vetEvent` - (`tests/e2e/helpers.go`) checks every event on the firehose against a - single `expectedCollections` allowlist, whatever repo it came from, and a - violation fails the WHOLE suite (including the end-of-run sweep, which - re-vets the replay). `community.postv2` is already on that list, which is - the only reason the native repo passes today. **The first scenario that - writes any other collection into the reference PDS's repo takes every other - scenario down with it** — `bridge.federation`, `feed.vote`, - `actor.profile` on a native DID, anything. Rescoping the whitelist by repo class (bridge-hosted - author DIDs / bridge-hosted community DIDs / native-PDS DIDs, each with - independent rev tracking) is task 18's sweep item and was deliberately left - out of the reference-PDS landing. Related: the e2e stack leaves +- **TRIPWIRE — the collection whitelist is repo-class-blind, and it cuts both + ways.** `vetEvent` (`tests/e2e/helpers.go`) checks every event on the + firehose against a single `expectedCollections` allowlist with no notion of + which repo the record came from. It therefore fails CLOSED on a collection + outside the list — a scenario writing, say, a vote record into the native + repo fails its own test and, deterministically, the `zz_sweep` replay that + re-vets every retained event — and fails OPEN on one inside it: a + `social.coves.actor.profile` written into the native repo is + indistinguishable, to the whitelist, from the bridge writing that same + collection into a repo it owns. `community.postv2` being allowlisted is the + only reason the native repo passes today; the fails-open half is the real + argument for rescoping by repo CLASS (bridge-hosted author DIDs / + bridge-hosted community DIDs / native-PDS DIDs, each with its own allowlist + and independent rev tracking) — task 18's sweep item, deliberately left out + of the reference-PDS landing. Related: the e2e stack leaves `CONSUMER_ENABLED` at its default of false, so the acceptance engine never reacts to native records — a scenario that needs an acceptance record must turn the consumer on and point `JETSTREAM_URL` at the compose Jetstream diff --git a/README.md b/README.md index c1cfdfa..70c4878 100644 --- a/README.md +++ b/README.md @@ -259,10 +259,13 @@ suite authenticates as the bootstrap account on the reference PDS, writes one `community.postv2` into that account's own repo, and follows it through the relay to Jetstream and into the relay's own `getLatestCommit` state. Deliberately a smoke test, not a feature test — what it buys is the wire, -not the record. Note before extending it: `vetEvent`'s collection whitelist -is global, so a second scenario writing any other collection into the native -repo fails the whole suite until that whitelist is rescoped by repo class -(see FOLLOWUPS.md). +not the record. Note before extending it: `vetEvent`'s `expectedCollections` +allowlist knows nothing about which repo a record came from, so it fails +**closed** on a collection outside the list (a vote record in the native repo +fails its own test and the `zz_sweep` replay) and **open** on one inside it +(an `actor.profile` in the native repo passes silently, indistinguishable +from the bridge writing its own). That second half is why the allowlist wants +rescoping by repo class — task 18's sweep item, see FOLLOWUPS.md. Every create/update the tests consume from Jetstream has passed the relay's diff --git a/docker-compose.e2e.yml b/docker-compose.e2e.yml index 34a7ed4..6c7777a 100644 --- a/docker-compose.e2e.yml +++ b/docker-compose.e2e.yml @@ -151,7 +151,7 @@ services: # bridgedStats field via a real UPDATE commit on the firehose. The # suite's zz-sweep lexicon-validates every create AND update, so the # synced #bridgedStats def must make those pass. Every emitted update is - # a post/comment in the four-collection whitelist, so no negative + # a post/comment in the suite's expectedCollections allowlist, so no negative # assertion is tripped (see votes_hammer_test.go: it awaits the vote XRPC # counts and never asserts "no further commits"). STATS_REFRESH_INTERVAL: 2s @@ -345,19 +345,23 @@ services: # `localhost` → http://localhost:${PDS_PORT}; ANY other hostname becomes # https://…, which this plain-HTTP harness cannot serve. So the DID # documents this PDS writes name http://localhost:3001 — and that string - # has to be TRUE for the relay, because bigsky's createExternalUser - # (indigo bgs/bgs.go:1013-1045) takes the service endpoint out of the DID - # DOCUMENT, looks the PDS peering up by THAT host, and calls - # describeServer against THAT url from inside its own namespace - # (bgs.go:1035-1037 special-cases a `localhost:` prefix back to http — - # indigo is built for exactly this topology), then refuses commits from a - # PDS that is not authoritative for the repo (bgs.go:823-836). + # has to be TRUE for the relay, because bigsky's createExternalUser takes + # the service endpoint out of the DID DOCUMENT, looks the PDS peering up + # by THAT host, and calls describeServer against THAT url from inside its + # own namespace — special-casing a `localhost:` prefix back to http, + # because indigo is built for exactly this topology — then refuses commits + # from a PDS that is not authoritative for the repo. (Line numbers as of + # indigo 8f2296e — 2025-10-10, ~11 days after the pinned image's commit: + # bgs/bgs.go:1013-1045 createExternalUser, :1035-1037 the localhost case, + # :823-836 the authority check. Neither the image's own commit nor the + # indigo in go.mod is that snapshot, so treat these as landmarks to search + # for, not addresses to trust.) # `network_mode: "service:relay"` makes localhost:3001 the same listener # for both containers, so the DID document tells the truth. The obvious # alternative — a `pds` hostname of its own on the compose network — fails # SILENTLY and expensively: the relay accepts the crawl, registers the - # host, then logs and advances its cursor forever (fedmgr.go:556-563) with - # RepoCount 0 and not one event delivered. + # host, then logs and advances its cursor forever (bgs/fedmgr.go:556-563 at + # the same snapshot) with RepoCount 0 and not one event delivered. # # Consequences encoded elsewhere in this file: # * The host port is published on the `relay` service, not here. @@ -457,6 +461,12 @@ services: # the host in the DID documents this PDS writes, which is what the relay # keys its peering on. The relay validates it by calling describeServer # back at that url; no race to lose here, since `pds` is healthy first. + # A 200 here means the relay ACCEPTED the request, which is not the same + # as crawling: the silent RepoCount-0 mode in the `pds` header above + # begins with exactly this 200. What proves the crawl is the smoke test's + # pre-write relay-head poll (tests/e2e/native_pds_test.go), which demands + # the relay already serve a head for the native repo before the scenario + # writes anything. # # Runs in its own namespace on the compose network, so the PDS is # relay:3001 from here — the same listener, seen from outside. @@ -466,10 +476,15 @@ services: command: - | echo '▶ creating native-alice.pds.test on the reference PDS' - wget -q -O- --header='Content-Type: application/json' \ + if ! wget -q -O- --header='Content-Type: application/json' \ --post-data='{"handle":"native-alice.pds.test","email":"native-alice@pds.test","password":"native-alice-e2e-pass"}' \ - http://relay:3001/xrpc/com.atproto.server.createAccount \ - || echo ' createAccount failed — expected when the account already exists; createSession below is the real gate' + http://relay:3001/xrpc/com.atproto.server.createAccount; then + echo ' createAccount failed with the status above, and BusyBox wget cannot read the' + echo ' response body — this one-shot does not know why. Benign IF the account already' + echo ' exists, which is every re-run of the stack. If the createSession below fails' + echo ' too, THIS was the real failure: check the pds logs for PLC reachability and' + echo ' that the handle sits under PDS_SERVICE_HANDLE_DOMAINS.' + fi echo echo '▶ authenticating as native-alice.pds.test (the gate)' wget -q -O- --header='Content-Type: application/json' \ @@ -480,7 +495,7 @@ services: wget -q -O- --header='Content-Type: application/json' \ --post-data='{"hostname":"localhost:3001"}' \ http://relay:2470/xrpc/com.atproto.sync.requestCrawl - echo ' ✓ relay is crawling the reference PDS' + echo ' ✓ relay accepted the crawl request' restart: "no" networks: [tidepool-e2e] depends_on: diff --git a/tests/e2e/helpers.go b/tests/e2e/helpers.go index a266916..176afcd 100644 --- a/tests/e2e/helpers.go +++ b/tests/e2e/helpers.go @@ -2,11 +2,18 @@ // Package e2e drives the docker-compose.e2e.yml stack end to end: a real // Lemmy (debug build, plain-HTTP federation) federating with Tidepool, a -// real did:plc directory backing DID minting, a real BigSky relay crawling -// the bridge (DID resolution against the local PLC, per-commit signature -// verification), and a real Jetstream decoding the RELAY's firehose — every +// real did:plc directory backing DID minting, a real reference PDS +// (ghcr.io/bluesky-social/pds) hosting a natively-registered test account, a +// real BigSky relay crawling BOTH repo hosts — the bridge and that PDS — +// with DID resolution against the local PLC and per-commit signature +// verification, and a real Jetstream decoding the RELAY's firehose — every // event the suite consumes has therefore survived relay validation. // +// The reference PDS is the one repo host in the stack this repo did not +// write (see native_pds_test.go): without it, every commit the suite +// validates was produced and consumed by a single reading of the atproto +// spec, and a misreading shared by both ends is invisible. +// // Coves-style: E2E tests test REAL infrastructure, not mocks. Run them with // `make e2e` (compose up --build → wait for health → go test -tags e2e → // compose down -v), or against an already-running stack with @@ -270,6 +277,9 @@ func waitForStack() error { {"relay /xrpc/_health", func() error { return probeHTTP(client, relayURL()+"/xrpc/_health") }}, + {"reference pds /xrpc/_health", func() error { + return probeHTTP(client, pdsURL()+"/xrpc/_health") + }}, {"jetstream /subscribe", func() error { u := jetstreamURL() + "/subscribe?cursor=" + fmt.Sprint(time.Now().UnixMicro()) conn, _, err := websocket.DefaultDialer.Dial(u, nil) @@ -1374,8 +1384,13 @@ func (h *harness) relayGetLatestCommit(did string) (cid, rev string, err error) // ── Jetstream WebSocket listener ─────────────────────────────────────────── -// jsEvent is Jetstream's JSON event shape (kind "commit" only — the bridge -// emits no identity/account frames yet). +// jsEvent is Jetstream's JSON event shape. Only the "commit" kind is +// decoded in full: the bridge still emits no identity/account frames, but +// the reference PDS does — registering the native account puts #identity and +// #account on the firehose (plus a #sync frame Jetstream does not surface as +// a kind of its own) — so a listener can no longer assume every frame is a +// commit. Commit is a POINTER for exactly that reason: it is nil on those +// kinds, vetEvent returns early on them, and await skips them. type jsEvent struct { Did string `json:"did"` TimeUs int64 `json:"time_us"` diff --git a/tests/e2e/native_pds_test.go b/tests/e2e/native_pds_test.go index 4e9dc87..2171237 100644 --- a/tests/e2e/native_pds_test.go +++ b/tests/e2e/native_pds_test.go @@ -18,11 +18,16 @@ package e2e // one postv2 into a native repo and watches for it downstream. What it // buys is the wire, not the record. // -// TRIPWIRE for anyone extending this file: vetEvent (helpers.go) fails the -// WHOLE suite on any collection outside expectedCollections, globally. A -// second scenario writing some other collection into the native repo takes -// every other scenario down with it — rescoping that whitelist is task 18's -// sweep item and is out of scope here. +// TRIPWIRE for anyone extending this file, in both directions. vetEvent +// (helpers.go) checks collections against expectedCollections globally, with +// no notion of which repo a record came from, so it fails CLOSED on a +// collection outside the whitelist — a scenario writing, say, a vote record +// into the native repo takes every other scenario down with it — and fails +// OPEN on one inside it: a social.coves.actor.profile written into the +// native repo is indistinguishable, to the whitelist, from the bridge +// writing the same collection into a repo it owns. That second half is why +// the whitelist wants rescoping by repo CLASS rather than by collection +// alone; it is task 18's sweep item and out of scope here. import ( "bytes" @@ -170,14 +175,18 @@ func (c *pdsClient) createRecord(collection, rkey string, record map[string]any) // reaches the suite's Jetstream through the same relay every bridged commit // crosses. // -// The chain each assertion pins, in order: the reference PDS is reachable -// and the bootstrap account exists (createSession); the PDS accepts and -// commits our record (createRecord); the relay had already crawled the PDS -// and validated the commit's signature against the DID it resolved from the -// local PLC, and Jetstream decoded the relay's firehose (the awaited event — -// Jetstream's upstream IS the relay, so arrival there is the whole transit -// proven at once); and the relay built repo state from it rather than merely -// forwarding frames (getLatestCommit). +// The chain each assertion pins, in causal order — which is also the order +// they run in, so a break is reported by the hop that broke rather than by +// the last hop to time out. The reference PDS is reachable and the bootstrap +// account exists (createSession); the relay has crawled the PDS and built +// repo state for the account (a pre-write getLatestCommit — the +// account-creation commit is already a head, so this isolates "peering +// works" from "our record transits"); the PDS accepts and commits our record +// (createRecord); the relay validated the commit's signature against the DID +// it resolved from the local PLC and Jetstream decoded the relay's firehose +// (the awaited event — Jetstream's upstream IS the relay, so arrival there +// is the transit proven at once), carrying the exact CID the PDS assigned; +// and the relay advanced its own repo state to the commit it emitted. func TestNativePDS_RecordTransitsRelayToJetstream(t *testing.T) { h := newHarness(t) @@ -191,6 +200,19 @@ func TestNativePDS_RecordTransitsRelayToJetstream(t *testing.T) { } t.Logf("native account: handle=%s did=%s (minted by the reference PDS against the local PLC)", pds.handle, pds.did) + // Peering first, BEFORE anything is written. The account-creation commit + // already gave this repo a head, so a head served here means the relay + // crawled the PDS, resolved the DID at the local PLC, and verified a + // commit signature — all the machinery our record will need, checked + // while nothing about our record is in question yet. Without this hop a + // broken requestCrawl surfaces only as a generic Jetstream timeout two + // steps later, pointing at the wrong service. + preRev := h.awaitRelayHead(t, pds.did, "", + "the reference PDS was never crawled — check that pds-bootstrap's requestCrawl reached the relay "+ + "(`docker compose -f docker-compose.e2e.yml logs pds-bootstrap`) and that it announced the same "+ + "host string the PDS's DID document names; that match IS the peering mechanism") + t.Logf("relay already tracks the native repo at rev %s (the account-creation commit) — peering is live", preRev) + // One rkey per run, time-derived: the PDS's repo outlives a single // `make e2e-test` (the stack is brought up separately and re-tested), // and a fixed rkey would 400 as already-existing on the second run — @@ -221,30 +243,55 @@ func TestNativePDS_RecordTransitsRelayToJetstream(t *testing.T) { return e.Did == pds.did && e.Commit.Collection == colPostV2 && e.Commit.RKey == rkey && e.Commit.Operation == opCreate }) - if got := recordField(t, ev.Commit.Record, "community"); got != nativeCommunityDID { - t.Errorf("jetstream delivered community %q, want %q — the record changed shape in transit", got, nativeCommunityDID) + + // Content addressing checks the whole record at once, and checks it the + // way atproto itself does: the CID is the hash of the DAG-CBOR the PDS + // committed, so an identical CID at the far end means the bytes Jetstream + // re-encoded to JSON are the bytes the PDS signed. A field-by-field + // comparison would only cover the fields someone thought to list. + if ev.Commit.CID != cid { + t.Errorf("jetstream delivered %s at cid %q, but the PDS committed cid %q — "+ + "the record did not survive the PDS → relay → jetstream re-encoding intact", + uri, ev.Commit.CID, cid) } - // Relay-side state, not just passthrough: the relay indexes - // asynchronously, so poll it out — only the deadline is fatal. + // The relay must also ADVANCE its own repo state to the commit it just + // emitted, not merely forward the frame. Its indexing is asynchronous, so + // a rev still behind the emitted one is lag to wait out, never an + // immediate failure — the deadline is the only verdict. + postRev := h.awaitRelayHead(t, pds.did, ev.Commit.Rev, + "the commit reached Jetstream, so the relay forwarded it without ever building repo state from it") + t.Logf("relay head advanced %s → %s across the write", preRev, postRev) +} + +// awaitRelayHead polls the relay's view of a repo head until it serves a +// non-empty cid/rev at or after minRev ("" for any head), returning that +// rev. Every non-answer — a transport error, a 200 carrying an empty head, +// a rev still behind minRev — is transient by default and recorded as the +// reason to report if the deadline arrives; nothing here fails early, +// because every one of those states is a normal moment in the relay's +// asynchronous indexing. hint says what a real timeout would MEAN at this +// call site, which is the part a reader chasing the failure needs. +func (h *harness) awaitRelayHead(t *testing.T, did, minRev, hint string) string { + t.Helper() deadline := time.Now().Add(eventTimeout) - var lastErr error + lastErr := "the poll never completed an attempt" for { - headCID, rev, err := h.relayGetLatestCommit(pds.did) - if err != nil { - lastErr = err - t.Logf("relay getLatestCommit(%s): %v (retrying)", pds.did, err) - } else if headCID != "" && rev != "" { - if rev < ev.Commit.Rev { - t.Errorf("relay head rev %q is OLDER than the rev %q it already emitted for this repo — relay state lagging its own firehose", rev, ev.Commit.Rev) - } - t.Logf("relay serves the native repo: head=%s rev=%s", headCID, rev) - return + headCID, rev, err := h.relayGetLatestCommit(did) + switch { + case err != nil: + lastErr = err.Error() + t.Logf("relay getLatestCommit(%s): %v (retrying)", did, err) + case headCID == "" || rev == "": + lastErr = fmt.Sprintf("getLatestCommit returned 200 with empty cid/rev (%q/%q)", headCID, rev) + case rev < minRev: + lastErr = fmt.Sprintf("relay head rev %q is still older than rev %q, which it had already emitted for this repo", rev, minRev) + default: + return rev } if time.Now().After(deadline) { - t.Fatalf("the relay never served a commit for the native repo %s within %s (last error: %v) — "+ - "the event reached Jetstream, so the relay forwarded it without building repo state from it", - pds.did, eventTimeout, lastErr) + t.Fatalf("the relay never served a usable head for repo %s within %s: %s — %s", + did, eventTimeout, lastErr, hint) } time.Sleep(time.Second) } -- 2.51.2 From 6e0d3c48372674eddcd915a229e5d36b11e4ef75 Mon Sep 17 00:00:00 2001 From: Bretton Date: Sat, 15 Aug 2026 20:08:01 -0700 Subject: [PATCH 3/3] docs: the README suite list gains moderation_test.go, which existed all along (review) Co-Authored-By: Claude Fable 5 --- README.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 70c4878..6ca5a6b 100644 --- a/README.md +++ b/README.md @@ -203,8 +203,9 @@ plain-HTTP AP ids to match. See the header comments in `docker-compose.e2e.yml` and `e2e/lemmy/Dockerfile` for the full story. The suite (`tests/e2e/`, build tag `e2e`: `bridge_test.go`, -`lifecycle_test.go`, `media_test.go`, `native_pds_test.go`, `relay_test.go`, -`votes_hammer_test.go`, `zz_sweep_test.go`) +`lifecycle_test.go`, `media_test.go`, `moderation_test.go`, +`native_pds_test.go`, `relay_test.go`, `votes_hammer_test.go`, +`zz_sweep_test.go`) covers: subscribe → `community.profile` on the firehose; a link post → `actor.profile` and `community.post` (presence + author linkage; arrival order across the two repos is relay-dependent, see above), with the shared -- 2.51.2