diff --git a/README.md b/README.md index c0699ec..8dc2951 100644 --- a/README.md +++ b/README.md @@ -10,7 +10,7 @@ The project is experimental. The connectors, Jazz persistence layer, and process - **Jazz** stores events, cursors, consumer progress, execution records, and rebuildable projections. - **Consumers** subscribe with a narrow query, process matching events, and append outputs and lifecycle records. - **Dispatchers** independently accumulate completed candidate activity, render destination-specific batches, apply channel policy and velocity limits, perform external actions, and append delivery receipts. -- **Inspector** exposes local read-only views of events, executions, source health, and lineage. +- **Inspector** exposes local views of events, executions, source health, lineage, and blinded Review items. It is read-only unless the separately configured OAuth-only Review capability is active; even then, its only mutation is one append-only decision route. Producers and consumers communicate through Jazz. There is no central matcher, event dispatch queue, or lease scheduler. External egress uses separate dispatcher processes so source ingestion and consumer throughput never wait on notification policy or destination rate limits. @@ -64,11 +64,71 @@ Runtime data is stored under `.thoughtstream/`. Set `THOUGHTSTREAM_ROOT` to use | `pnpm thought event ` | Show one event. | | `pnpm thought runs` | List consumer executions. | | `pnpm thought run ` | Show one execution and its trace. | +| `pnpm thought review-prompt --file --external-id ` | Append one complete versioned Review prompt. Public export preauthorization is accepted only with `--privacy public-source`. | +| `pnpm thought review-item --prompt-event --candidate-runs ,` | Materialize one immutable blinded pair from two exact completed same-trigger candidate runs. | +| `pnpm thought review-queue` | Inspect the projected Review queue and active append-only decisions. | | `pnpm thought judgment --kind ` | Append an explicit quality judgment. External use requires `--external-export-eligible`; sensitive/private material also requires `--authorize-sensitive-external-export`. | | `pnpm thought training-export --output ` | Export only externally eligible, entirely public-source judgments into privacy-minimized JSONL and a content-addressed manifest. Sensitive/private export additionally requires both explicit private-export flags and a non-Git destination. | | `pnpm serve` | Start the local inspector on port 4317. | +| `pnpm configure:inspector-oauth -- --origin --did --handle ` | Generate owner-only OAuth client/store configuration while retaining Basic fallback. | +| `pnpm configure:inspector-review` | Generate one owner-only proxy-to-inspector Review capability. Generation does not restart or activate either service. | | `pnpm test` | Run the test suite. | +The inspector is deliberately loopback-only. `scripts/serve-inspector-proxy.ts` +serves a fixed public landing/documentation allowlist and forwards only +authenticated `GET`/`HEAD` requests below `/inspector` to the private upstream. +Public routes have no Jazz handle, manifest reader, runtime-state reader, +directory listing, or arbitrary file fallback. The proxy strips credentials +before forwarding. + +ATProto OAuth uses the official Node client for PKCE, PAR, DPoP, nonce handling, +identity resolution, and refresh. Public client metadata and JWKS live at +`/oauth/client-metadata.json` and `/oauth/jwks.json`; the callback is +`/oauth/callback`. Only `OAUTH_ALLOWED_DID` may receive an inspector session. +SDK protocol state, browser application state, DPoP keys, access/refresh tokens, +and opaque browser-session records are AES-256-GCM encrypted in an owner-only +directory outside the runtime root. The two state values are intentionally +distinct. Every store has hard entry/byte limits and one process owner; a second +proxy cannot open the fixed `~/.local/share/thoughtstream-inspector-auth` +directory, and custom store paths are rejected because systemd cannot write +them. Browser cookies contain random ids only. Login and callback are +rate-limited at nginx and in-process, callback query strings are excluded from +access logs, and `www` redirects to the canonical origin before application +routing. The authoritative watchdog covers SDK exchange, application-state +consumption, generation promotion, browser-session persistence, and cleanup. +Timeout synchronously makes the attempt non-promotable before advancing the +callback queue. Persistent writes recheck authority before atomic rename. +Promoted sessions and browser cookies carry per-DID generations, so application +cleanup can remove only its own local generation and does not request remote +revocation. The unmodified SDK may still revoke after issuer or session-store +failure; provider-side effects are outside the local authority guarantee. +Never-settling callback quarantines are capped at eight; exhaustion returns an +operator-recycle-required response until attempts settle or the process restarts. + +Basic Auth remains an independent break-glass path until a real HTTPS OAuth +login, private read, logout/local session deletion, and Basic rollback read have +been observed. When enabled it authorizes inspector reads without consulting +OAuth and is stripped before forwarding. When disabled it is ignored and not +advertised. The service refuses to start if neither authentication path exists. + +Review writes remain narrower than inspector authentication. Basic is always read-only. An allowlisted OAuth browser can append a decision only when both inspector processes load the same separately generated Review capability. The proxy verifies the server-side browser session and CSRF token, signs the exact method/path/body with a fresh nonce, and forwards no cookie, Authorization header, CSRF value, OAuth token, or DID. The inspector verifies that one-time envelope and accepts only the fixed Review decision schema. It cannot create prompts, run models, export datasets, activate adapters, publish, or perform a generic Jazz mutation. + +Run `scripts/configure-inspector-credentials.sh` to create or rotate the Basic +fallback without putting it in shell history. After the domain and exact DID are +known, run: + +```sh +pnpm configure:inspector-oauth -- --origin https://thought.stream --did did:plc:REPLACE_ME --handle cameron.stream +``` + +That command creates owner-only client/store keys and leaves +`PROXY_BASIC_FALLBACK_ENABLED=1`. It does not install units, reload nginx, +restart the proxy, or prove OAuth. Deployment templates live under +`deploy/systemd/` and `deploy/nginx/`; activate them only after DNS, TLS, the +HTTP-to-HTTPS redirect, external metadata/JWKS fetches, and rollback copies are +verified. Removing the Basic fallback is a later explicit operation, not part of +OAuth installation. + ## Connectors | Connector | Current support | @@ -157,9 +217,33 @@ Pi declarations select a trusted provider profile rather than supplying endpoint Pi declarations may opt into the read-only `atproto.fetch-markdown` and `web.download-image` tools. The trusted parent executes those bounded reads before inference and passes only their evidence into the sandbox. The image tool accepts only URLs discovered in the current source record or fetched Markdown, rejects non-public network destinations, bounds response size, and stores content-addressed artifacts under `.thoughtstream/artifacts/`. Durable tool outcomes contain status, field names, counts, and hashes rather than arguments, source bodies, image bytes, or arbitrary errors. [`agents/bluesky-enrichment-observer.yaml`](agents/bluesky-enrichment-observer.yaml) is disabled by default because it requires a configured Tinker credential and deliberate source/actor policy. +Output contracts are versioned registry entries rather than one universal observation shape. [`agents/conceptualizer.example.yaml`](agents/conceptualizer.example.yaml) shows an inactive Tinker-backed consumer using `stream.thought.output.conceptualization@1`. A successful run settles one bounded private `stream.thought.derived.concept.graph` event with source/run lineage. It does not publish ATProto records or mutate an external graph. + +[`agents/review-candidate-a.example.yaml`](agents/review-candidate-a.example.yaml) and [`agents/review-candidate-b.example.yaml`](agents/review-candidate-b.example.yaml) show the disabled two-candidate path. Both consume the same `stream.thought.source.review.prompt`, emit strict `stream.thought.output.review-response@1` outputs, and have no tools or external actions. `review-item` freezes the exact completed pair before browser grading. `spec/review.md` owns judgeability, correction, blinding, supersession, browser authority, and export semantics. + +A local public-safe canary starts from the reviewed payload fixture: + +```sh +pnpm thought review-prompt \ + --file fixtures/review/prompt.example.json \ + --external-id response-quality-canary-001 \ + --privacy public-source +``` + +Enable two concrete declarations only after replacing their model or immutable adapter selections, run the ordinary consumer process, then materialize the two completed run ids: + +```sh +pnpm thought review-item --prompt-event --candidate-runs , +pnpm thought review-queue +``` + +Import and model execution remain CLI/runtime operations. The browser can append decisions but cannot create campaigns or trigger inference. + +Learned models use an immutable startup catalog: release manifests under `adapters/releases/` plus one `adapters/deployment.yaml` selecting an exact active release for each declaration. Startup resolves checkpoint environment references into process-local private bindings and persists only public release, binding digest, catalog digest, and generation. Activation or retirement requires stopping all adapter-egress services, atomically installing a new catalog generation, restarting, verifying PID/start identity and loaded digest, and running a canary. There is no hot Jazz lifecycle or file-lock authority. The tracked Julia files are inert examples and contain no real checkpoint. + Agent outputs do not become training data merely because a run completed. `judgment` writes an explicit accept, reject, correction, or pairwise preference event with separate quality and external-export eligibility. Receipt-bound `👍` and `👎` reactions provide quality evidence only; their reaction, delivery, run, output, source root, supersession, and retraction lineage remain private and durable. -Default `training-export` includes only active, externally eligible judgments whose complete source/output chain is `public-source`. The v2 export contains validated chosen/rejected fields, minimal source classification, trace type/order, output-contract identity, model/adapter provenance, prompt hash, and judgment criterion. It omits source payloads, actor/route/external/correlation/idempotency identifiers, event/run/delivery ids, source and trace-content hashes, trace timestamps, and arbitrary context fields. Sensitive/private export requires explicit authority at judgment creation, both `--include-sensitive-private` and `--authorize-sensitive-private-export`, a file destination outside every Git worktree and configured public-content root, and owner-only atomic dataset/manifest files. +Default `training-export` includes only active, externally eligible judgments whose complete source/output chain is `public-source`. Legacy judgments remain `thoughtstream.training-example.v3`. Reviewed preferences and corrections use v4, adding the exact preauthorized public prompt/evidence, bounded criterion metadata, one chosen response, one or two rejected candidates, campaign identity, and exact candidate model/adapter/catalog provenance. The dataset manifest is v4 and records mixed example-format counts and Review campaigns. Adapter privacy/export policy is checked independently on every participating run. Export omits notes, browser submission ids, source actor/route/external/correlation/idempotency identifiers, event/run/delivery ids, source and trace-content hashes, trace timestamps, arbitrary context fields, and private checkpoints. Sensitive/private export requires explicit authority at judgment creation, both `--include-sensitive-private` and `--authorize-sensitive-private-export`, a file destination outside every Git worktree and configured public-content root, and owner-only atomic dataset/manifest files. The browser cannot declassify private Review material. Eligible terminal output-validation failures append one deterministic repair request. The separately declared `output-repair` Pi consumer regenerates the original bounded context, runs through the same Bubblewrap/broker boundary, and may append one contract-valid correction proposal. A proposal is inert until an `accept` or `correct` judgment names its repair run; rejection, supersession, or retraction is preserved append-only. See [`spec/repairs.md`](spec/repairs.md) for eligibility, privacy, authority, and training rules. diff --git a/adapters/deployment.example.yaml b/adapters/deployment.example.yaml new file mode 100644 index 0000000..7200994 --- /dev/null +++ b/adapters/deployment.example.yaml @@ -0,0 +1,16 @@ +# Coordinated-deployment example only. Install one completed bundle, stop all +# adapter-egress services, restart them, verify loaded digest/PID identity, then canary. +schemaVersion: 1 +generation: 1 +expectedProcesses: + - conceptualizer +selections: + - declaration: + id: conceptualizer + version: 1 + release: + id: julia-adapter + version: 1 + # Replace with the digest of the release manifest selected at deployment time. + manifestSha256: "0000000000000000000000000000000000000000000000000000000000000000" + state: candidate diff --git a/adapters/releases/julia-adapter-v1.example.yaml b/adapters/releases/julia-adapter-v1.example.yaml new file mode 100644 index 0000000..3ee05cb --- /dev/null +++ b/adapters/releases/julia-adapter-v1.example.yaml @@ -0,0 +1,23 @@ +# Example only: this is not a live Tinker run and contains no checkpoint value. +schemaVersion: 1 +id: julia-adapter +version: 1 +description: Example Julia adapter release with a private environment-bound checkpoint. +releasedAt: "2026-01-01T00:00:00.000Z" +providerProfile: tinker-default +baseModel: tinker/julia-base +checkpoint: + env: THOUGHTSTREAM_JULIA_ADAPTER_CHECKPOINT +dataset: + id: example/julia-training-set + sha256: "0000000000000000000000000000000000000000000000000000000000000000" +evals: + - id: example/julia-eval-set + sha256: "1111111111111111111111111111111111111111111111111111111111111111" +capabilities: + - conceptualization +privacyClass: private +exportClass: restricted +sourceManifest: + id: example/julia-source-manifest + sha256: "2222222222222222222222222222222222222222222222222222222222222222" diff --git a/agents/conceptualizer.example.yaml b/agents/conceptualizer.example.yaml new file mode 100644 index 0000000..6d2e894 --- /dev/null +++ b/agents/conceptualizer.example.yaml @@ -0,0 +1,55 @@ +id: conceptualizer +version: 1 +name: Conceptualizer +description: Extracts one bounded private concept graph from a canonical event batch. +enabled: false +outputContract: + id: stream.thought.output.conceptualization + version: 1 +subscribe: + types: + - stream.thought.derived.event.batch + sources: + - batch:cameron-atproto + privacy: + - public-source +context: + strategy: atproto-batch + maxEvents: 1 + maxChars: 64000 + atprotoObject: true +runner: + kind: pi + profile: openai-json-default + model: gpt-4.1-mini + maxOutputTokens: 2000 + timeoutMs: 120000 +accounting: + leaseMs: 180000 + reservation: + inputTokens: 40000 + outputTokens: 2000 + costMicrousd: 200000 + limits: + - window: rolling + durationMs: 300000 + maxCalls: 2 + maxInputTokens: 80000 + maxOutputTokens: 4000 + maxCostMicrousd: 400000 + - window: hour + maxCalls: 12 + maxInputTokens: 480000 + maxOutputTokens: 24000 + maxCostMicrousd: 2400000 + - window: day + maxCalls: 100 + maxInputTokens: 4000000 + maxOutputTokens: 200000 + maxCostMicrousd: 20000000 +prompt: prompts/conceptualizer.md +emit: + - stream.thought.derived.concept.graph +policy: + tools: [] + externalActions: false diff --git a/agents/review-candidate-a.example.yaml b/agents/review-candidate-a.example.yaml new file mode 100644 index 0000000..d980772 --- /dev/null +++ b/agents/review-candidate-a.example.yaml @@ -0,0 +1,44 @@ +id: review-candidate-a +version: 1 +name: Review candidate A +description: Generates one contract-bound candidate response for a complete Review prompt. +enabled: false +outputContract: + id: stream.thought.output.review-response + version: 1 +subscribe: + types: + - stream.thought.source.review.prompt + sources: + - review:prompt + privacy: + - public-source +context: + strategy: single-event + maxEvents: 1 + maxChars: 220000 +runner: + kind: pi + profile: tinker-default + model: REPLACE_WITH_FROZEN_PARENT_OR_CANDIDATE_A + maxOutputTokens: 8000 + timeoutMs: 120000 +accounting: + leaseMs: 180000 + reservation: + inputTokens: 120000 + outputTokens: 8000 + costMicrousd: 400000 + limits: + - window: rolling + durationMs: 300000 + maxCalls: 2 + maxInputTokens: 240000 + maxOutputTokens: 16000 + maxCostMicrousd: 800000 +prompt: prompts/review-candidate.md +emit: + - stream.thought.derived.review.response +policy: + tools: [] + externalActions: false diff --git a/agents/review-candidate-b.example.yaml b/agents/review-candidate-b.example.yaml new file mode 100644 index 0000000..82d22fe --- /dev/null +++ b/agents/review-candidate-b.example.yaml @@ -0,0 +1,44 @@ +id: review-candidate-b +version: 1 +name: Review candidate B +description: Generates the comparison response for the same complete Review prompt. +enabled: false +outputContract: + id: stream.thought.output.review-response + version: 1 +subscribe: + types: + - stream.thought.source.review.prompt + sources: + - review:prompt + privacy: + - public-source +context: + strategy: single-event + maxEvents: 1 + maxChars: 220000 +runner: + kind: pi + profile: tinker-default + model: REPLACE_WITH_FROZEN_PARENT_OR_CANDIDATE_B + maxOutputTokens: 8000 + timeoutMs: 120000 +accounting: + leaseMs: 180000 + reservation: + inputTokens: 120000 + outputTokens: 8000 + costMicrousd: 400000 + limits: + - window: rolling + durationMs: 300000 + maxCalls: 2 + maxInputTokens: 240000 + maxOutputTokens: 16000 + maxCostMicrousd: 800000 +prompt: prompts/review-candidate.md +emit: + - stream.thought.derived.review.response +policy: + tools: [] + externalActions: false diff --git a/deploy/nginx/thought.stream.bootstrap.conf b/deploy/nginx/thought.stream.bootstrap.conf new file mode 100644 index 0000000..78bb235 --- /dev/null +++ b/deploy/nginx/thought.stream.bootstrap.conf @@ -0,0 +1,12 @@ +# Bootstrap this virtual host only after thought.stream resolves to this nginx +# host. It exists only so Certbot can prove domain control. It deliberately does +# not expose the password proxy over plaintext HTTP. +server { + listen 80; + listen [::]:80; + server_name thought.stream www.thought.stream; + + location / { + return 404; + } +} diff --git a/deploy/nginx/thought.stream.conf b/deploy/nginx/thought.stream.conf new file mode 100644 index 0000000..f3ccd1b --- /dev/null +++ b/deploy/nginx/thought.stream.conf @@ -0,0 +1,125 @@ +limit_req_zone $binary_remote_addr zone=thoughtstream_oauth_login:10m rate=6r/m; +limit_req_zone $binary_remote_addr zone=thoughtstream_oauth_callback:10m rate=12r/m; +limit_req_zone $binary_remote_addr zone=thoughtstream_review:10m rate=60r/m; +log_format thoughtstream_no_query '$remote_addr [$time_local] "$request_method $uri $server_protocol" $status $body_bytes_sent'; + +server { + listen 80; + listen [::]:80; + server_name thought.stream www.thought.stream; + access_log /var/log/nginx/thought.stream.access.log thoughtstream_no_query; + error_log /dev/null crit; + return 308 https://thought.stream$request_uri; +} + +server { + listen 443 ssl http2; + listen [::]:443 ssl http2; + server_name www.thought.stream; + + ssl_certificate /etc/letsencrypt/live/thought.stream/fullchain.pem; + ssl_certificate_key /etc/letsencrypt/live/thought.stream/privkey.pem; + include /etc/letsencrypt/options-ssl-nginx.conf; + ssl_dhparam /etc/letsencrypt/ssl-dhparams.pem; + + access_log /var/log/nginx/thought.stream.access.log thoughtstream_no_query; + error_log /dev/null crit; + add_header Strict-Transport-Security "max-age=31536000" always; + return 308 https://thought.stream$request_uri; +} + +server { + listen 443 ssl http2; + listen [::]:443 ssl http2; + server_name thought.stream; + + ssl_certificate /etc/letsencrypt/live/thought.stream/fullchain.pem; + ssl_certificate_key /etc/letsencrypt/live/thought.stream/privkey.pem; + include /etc/letsencrypt/options-ssl-nginx.conf; + ssl_dhparam /etc/letsencrypt/ssl-dhparams.pem; + + access_log /var/log/nginx/thought.stream.access.log thoughtstream_no_query; + add_header Strict-Transport-Security "max-age=31536000" always; + + location = /oauth/login { + client_max_body_size 8k; + limit_req zone=thoughtstream_oauth_login burst=2 nodelay; + limit_req_status 429; + limit_except GET HEAD POST { deny all; } + proxy_pass http://127.0.0.1:4319; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + proxy_set_header Connection ""; + proxy_buffering off; + proxy_read_timeout 30s; + proxy_send_timeout 30s; + } + + location = /oauth/callback { + access_log off; + error_log /dev/null crit; + limit_req zone=thoughtstream_oauth_callback burst=4 nodelay; + limit_req_status 429; + limit_except GET HEAD { deny all; } + proxy_pass http://127.0.0.1:4319; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + proxy_set_header Connection ""; + proxy_buffering off; + proxy_read_timeout 30s; + proxy_send_timeout 30s; + } + + location = /oauth/logout { + client_max_body_size 8k; + limit_except GET HEAD POST { deny all; } + proxy_pass http://127.0.0.1:4319; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + proxy_set_header Connection ""; + proxy_buffering off; + proxy_read_timeout 30s; + proxy_send_timeout 30s; + } + + location ~ ^/inspector/api/reviews/[^/]+/decisions$ { + client_max_body_size 100k; + limit_req zone=thoughtstream_review burst=10 nodelay; + limit_req_status 429; + limit_except POST { deny all; } + proxy_pass http://127.0.0.1:4319; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + proxy_set_header Connection ""; + proxy_buffering off; + proxy_read_timeout 30s; + proxy_send_timeout 30s; + } + + location / { + client_max_body_size 8k; + limit_except GET HEAD { deny all; } + proxy_pass http://127.0.0.1:4319; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + proxy_set_header Connection ""; + proxy_buffering off; + proxy_read_timeout 30s; + proxy_send_timeout 30s; + } +} diff --git a/deploy/systemd/thoughtstream-inspector-proxy.service b/deploy/systemd/thoughtstream-inspector-proxy.service new file mode 100644 index 0000000..1ae8725 --- /dev/null +++ b/deploy/systemd/thoughtstream-inspector-proxy.service @@ -0,0 +1,32 @@ +[Unit] +Description=ThoughtStream public site and authenticated inspector proxy +After=thoughtstream-inspector.service +Requires=thoughtstream-inspector.service + +[Service] +# Singleton contract: never template or replicate this unit against one OAUTH_STORE_DIR. +# The process also holds an owner-only store lock and refuses a second live owner. +Type=simple +WorkingDirectory=%h/code/thought-stream +EnvironmentFile=%h/.config/thoughtstream/credentials/inspector-proxy.env +EnvironmentFile=-%h/.config/thoughtstream/credentials/inspector-oauth.env +EnvironmentFile=-%h/.config/thoughtstream/credentials/inspector-review.env +ExecStart=%h/.nvm/versions/node/v22.19.0/bin/node --import tsx %h/code/thought-stream/scripts/serve-inspector-proxy.ts +Restart=on-failure +RestartSec=5 +KillMode=control-group +NoNewPrivileges=true +PrivateTmp=true +ProtectSystem=strict +ProtectHome=read-only +ReadWritePaths=-%h/.local/share/thoughtstream-inspector-auth +ProtectKernelTunables=true +ProtectControlGroups=true +RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX +RestrictRealtime=true +RestrictSUIDSGID=true +LockPersonality=true +UMask=0077 + +[Install] +WantedBy=default.target diff --git a/deploy/systemd/thoughtstream-inspector.service b/deploy/systemd/thoughtstream-inspector.service new file mode 100644 index 0000000..967829c --- /dev/null +++ b/deploy/systemd/thoughtstream-inspector.service @@ -0,0 +1,28 @@ +[Unit] +Description=ThoughtStream private inspector and bounded Review decision sink +After=network-online.target +Wants=network-online.target + +[Service] +Type=simple +WorkingDirectory=%h/code/thought-stream +Environment=THOUGHTSTREAM_ROOT=%h/.local/share/thoughtstream/live +EnvironmentFile=-%h/.config/thoughtstream/credentials/inspector-review.env +ExecStart=%h/.nvm/versions/node/v22.19.0/bin/node --import tsx %h/code/thought-stream/src/cli.ts serve --host 127.0.0.1 --port 4317 +Restart=on-failure +RestartSec=5 +NoNewPrivileges=true +PrivateTmp=true +ProtectSystem=strict +ProtectHome=read-only +ReadWritePaths=%h/.local/share/thoughtstream/live +ProtectKernelTunables=true +ProtectControlGroups=true +RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX +RestrictRealtime=true +RestrictSUIDSGID=true +LockPersonality=true +UMask=0077 + +[Install] +WantedBy=default.target diff --git a/fixtures/review/prompt.example.json b/fixtures/review/prompt.example.json new file mode 100644 index 0000000..2e6fc0c --- /dev/null +++ b/fixtures/review/prompt.example.json @@ -0,0 +1,32 @@ +{ + "campaign": { + "id": "response-quality-canary", + "version": 1, + "label": "Response quality canary" + }, + "prompt": "Given the supplied critique and evidence, explain which mechanism should change and why.", + "evidence": "Replace this synthetic public evidence with the complete facts required to make the comparison judgeable.", + "criterion": { + "id": "mechanism-changing-response", + "version": 1, + "label": "Mechanism-changing response quality", + "instructions": "Prefer the response that identifies the causal mechanism, proposes a testable change, and preserves uncertainty where evidence is incomplete.", + "reasonCodes": [ + "mechanism-identified", + "mechanism-missed", + "evidence-used", + "unsupported-claim" + ], + "responseTags": [ + "specific", + "vague", + "overconfident", + "well-calibrated" + ] + }, + "candidateAgentIds": [ + "review-candidate-a", + "review-candidate-b" + ], + "externalExportEligible": true +} diff --git a/package.json b/package.json index b160f8c..f2a1df0 100644 --- a/package.json +++ b/package.json @@ -14,7 +14,9 @@ "build:harness-image": "pnpm build:harness && docker build -f docker/pi-coding-harness.Dockerfile -t thoughtstream/pi-coding-harness:local .", "canary:letta-agent-sdk": "tsx scripts/letta-agent-sdk-canary.ts", "provision:letta-resident": "tsx scripts/provision-letta-resident.ts", + "configure:inspector-oauth": "tsx scripts/configure-inspector-oauth.ts", "split:service-credentials": "tsx scripts/split-service-credentials.ts", + "configure:inspector-review": "tsx scripts/configure-inspector-review.ts", "test:harness-container": "pnpm build:harness-image && THOUGHTSTREAM_RUN_CONTAINER_TESTS=1 vitest run test/harness-container.test.ts", "check": "tsc --noEmit", "pretest": "pnpm build:sandbox", @@ -30,6 +32,8 @@ "demo": "tsx src/cli.ts demo --root fixtures/vault --source filesystem:fixture" }, "dependencies": { + "@atproto/jwk-jose": "0.2.4", + "@atproto/oauth-client-node": "0.4.9", "@earendil-works/pi-agent-core": "^0.80.6", "@earendil-works/pi-ai": "^0.80.6", "@earendil-works/pi-coding-agent": "0.80.6", @@ -37,7 +41,7 @@ "chokidar": "^4.0.3", "diff": "^8.0.2", "fast-glob": "^3.3.3", - "fast-xml-parser": "^5.10.0", + "fast-xml-parser": "^5.10.1", "jazz-napi": "2.0.0-alpha.53", "jazz-tools": "2.0.0-alpha.53", "typebox": "1.1.38", @@ -57,5 +61,11 @@ "engines": { "node": ">=22.19.0" }, + "pnpm": { + "overrides": { + "brace-expansion@<=5.0.7": "5.0.8", + "protobufjs@>=8.0.0 <8.4.1": "8.4.1" + } + }, "packageManager": "pnpm@10.20.0" } diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index caaf55e..1385de2 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -4,10 +4,20 @@ settings: autoInstallPeers: true excludeLinksFromLockfile: false +overrides: + brace-expansion@<=5.0.7: 5.0.8 + protobufjs@>=8.0.0 <8.4.1: 8.4.1 + importers: .: dependencies: + '@atproto/jwk-jose': + specifier: 0.2.4 + version: 0.2.4 + '@atproto/oauth-client-node': + specifier: 0.4.9 + version: 0.4.9 '@earendil-works/pi-agent-core': specifier: ^0.80.6 version: 0.80.6(ws@8.21.0)(zod@4.4.3) @@ -30,8 +40,8 @@ importers: specifier: ^3.3.3 version: 3.3.3 fast-xml-parser: - specifier: ^5.10.0 - version: 5.10.0 + specifier: ^5.10.1 + version: 5.10.1 jazz-napi: specifier: 2.0.0-alpha.53 version: 2.0.0-alpha.53 @@ -88,6 +98,94 @@ packages: zod: optional: true + '@atproto-labs/did-resolver@0.3.5': + resolution: {integrity: sha512-0dMM+hj40VQiD/EJhlC1UMQgPXRwKeqM7NgJte7fVuYMv5b3P0W6+Lu3iDumHULcSjMmMXxJZzoi3i493Y0gCA==} + engines: {node: '>=22'} + + '@atproto-labs/fetch-node@0.3.5': + resolution: {integrity: sha512-fVoniRexly08D5Htzmdd6bvS2Y1tFMWVAiZCUbFBRojm5QFfhzYVvZRwBUR9zL3+8Uh+bEikZ66MFvzbwhsx0g==} + engines: {node: '>=22'} + + '@atproto-labs/fetch@0.3.4': + resolution: {integrity: sha512-YxXwi8HMk2HHDd5rljPGqxZ8bSeU78sFNz511Y222D0FraEq98p2r1ifkIgI23KQOfy0k0VzafeJAiGDh1CvfQ==} + engines: {node: '>=22'} + + '@atproto-labs/handle-resolver-node@0.2.6': + resolution: {integrity: sha512-hsez5HllGLeWBsToLe4GpDwae1QGcQUVwrgPnurSoNhncIOi44xVDcmTLvf6Jjokc2rHisCbiveju+Xf/0fzow==} + engines: {node: '>=22'} + + '@atproto-labs/handle-resolver@0.4.6': + resolution: {integrity: sha512-a3ZoQ0xpIowFuIhvjXwaloJwT3He6NJ02FfrKaRW0tdT2GvEH6nLaVHr46ps8A4ddCz+Q4oUha9atfr4O+0P2g==} + engines: {node: '>=22'} + + '@atproto-labs/identity-resolver@0.4.5': + resolution: {integrity: sha512-N/v4vZ4z8hFkudBKe5S1MIUP4nI+Qrr2FVVoGJ1ED+j4tT05y6JB8XFC34jgdZUaeQhi5wLaaSMoqUD269S6tw==} + engines: {node: '>=22'} + + '@atproto-labs/pipe@0.2.4': + resolution: {integrity: sha512-n67jCcrC+ouAeO10cWkpPzzLMlDi/lDCU30Us+LGqhOPhT6c4t5ASdBLQi9W3jUQtRzQBt3G9zipF+xKWNvVbw==} + engines: {node: '>=22'} + + '@atproto-labs/simple-store-memory@0.2.4': + resolution: {integrity: sha512-xAAUlOP9etqP9GGmJPq9gY1bRuJlZFdsC6wtAJSwUwMTwLRdenR7s8Qg5nxhbth5XL1XOtGABoD59NW3p+sl4A==} + engines: {node: '>=22'} + + '@atproto-labs/simple-store@0.4.4': + resolution: {integrity: sha512-3YH03xg99ZUS6Fq/6jcKRiZwUbzDb58xmS7ibg0y4+dukwrHa92cV/md76RpQJkAo9KYBvwLETdvF6fNtwIoGw==} + engines: {node: '>=22'} + + '@atproto/common-web@0.5.6': + resolution: {integrity: sha512-5Y4MIK9dpkJPiKiE6u7iEHitxj+g3aAU2GfGL686JlKE2zDKD4y18BJb+uVek6nXQKb5XOdNVvw+7BHarpn8Fw==} + engines: {node: '>=22'} + + '@atproto/did@0.5.4': + resolution: {integrity: sha512-BlnwQ+obL+4ZA71KH/EzZ3TY+cpSxnLiUI85mjBJIQwvDF/oN2sQA18wIj9jpduviIt2b/cMtlJuzHjzBkfXvw==} + engines: {node: '>=22'} + + '@atproto/jwk-jose@0.2.4': + resolution: {integrity: sha512-gzDoA0JTwnc0ZJOBLM7WX9xFxtynRS2K1Bofb8epzoMWDQvyvfbPcfkdPrKFM7NXCFUVpGpBnsCB8KFPTf1rCg==} + engines: {node: '>=22'} + + '@atproto/jwk-webcrypto@0.3.4': + resolution: {integrity: sha512-UsFIUozqnRecXPo6HgKV4PW4FqYHxX1V3iAe0rRV6Q2RSfYD8V2mZ89pv8NpvJynnaqJArBQ7HZlcg0F4tRYhA==} + engines: {node: '>=22'} + + '@atproto/jwk@0.7.4': + resolution: {integrity: sha512-tq7TUDmNfe1yDfpRgdGQMJdl9TUlJmREQNCag9yg5w8Evu+TOiFiLgiOCbo7X4ouRPSgd1DpOzXbUa8UyKKMZA==} + engines: {node: '>=22'} + + '@atproto/lex-data@0.1.5': + resolution: {integrity: sha512-TEM6GHuYpNm4O90LjNgbYq1Gmcr875S+BHrDxkg4PB5w/nlsz4HlbkGQG/WP/xbIV5O8TF5nNHDpCZ5ezTRrFA==} + engines: {node: '>=22'} + + '@atproto/lex-json@0.1.4': + resolution: {integrity: sha512-ENR2cWkVrES+UL6TovbCRdX9BJOyHHJUS8jYx3Lxp3j4vEphjn/u+DW7bWloST2O8ID2OuVlt6+28ftNEEjmQQ==} + engines: {node: '>=22'} + + '@atproto/lexicon@0.7.7': + resolution: {integrity: sha512-92VH2oEsJdrIVNy7WY8rGn99ANNVglyUffoN7GJc0mxKi+fXN5iVlJ2cOyTfYABhr2ldSiptZWXEPEdBaIR9/A==} + engines: {node: '>=22'} + + '@atproto/oauth-client-node@0.4.9': + resolution: {integrity: sha512-pLsghLG4vmm9h0FEyw+1+wSr+KH8vJjn68CEDtLabB4SPGtzwp6phcJDkrbKWrhndWMDdqEr9X+mAFSgb1Q9lg==} + engines: {node: '>=22'} + + '@atproto/oauth-client@0.7.11': + resolution: {integrity: sha512-kCoQxT2CXEKT2LCZA3nZFd5VjbZN6WQ7z8asUfRICuhm79Q/Ovbs8eFftdcxM75O7M/9/EdcZ6Opppi0cq5kjw==} + engines: {node: '>=22'} + + '@atproto/oauth-types@0.7.5': + resolution: {integrity: sha512-x75O0HsKB1IGfBikAQrrTX6EL8Rt4Q0+wMcwhZQRGPk/N/WqbYbsW3Powj4R8ZJKSfWCpVfaw32Piu1pTi891Q==} + engines: {node: '>=22'} + + '@atproto/syntax@0.7.2': + resolution: {integrity: sha512-tZ1Tr0R9pK4bI4Zs69t29cjMlCFQvRNBeNkqJC5pGeNuCt64D0eoF7s/AlqeVmYppClmQ0xJquggGgsMDP6j7w==} + engines: {node: '>=22'} + + '@atproto/xrpc@0.8.6': + resolution: {integrity: sha512-yVfKrlwZBBm44Ft9jDvHcTCwQ1ElqSlzXHWhT/KcqE+u24EAhWVFosfABfGbUByB/PmN8eLJ06FsWo5pUQcMNQ==} + engines: {node: '>=22'} + '@aws-crypto/sha256-browser@5.2.0': resolution: {integrity: sha512-AXfN/lGotSQwu6HNcEsIASo7kWXZ5HYWvfOmSNKDsEqC4OashTp8alTmaz+F7TC2L083SFv5RdB+qU3Vs1kZqw==} @@ -1076,8 +1174,8 @@ packages: resolution: {integrity: sha512-IYqDGiTXab6FniAgnSdZwgWbomxpy9FtYvLKs7wCUs2a8RkITG+DFGO1DM9cr+E3/RgADRpFjrKVaJ1z6sjtEg==} engines: {node: '>= 20.19.0'} - '@nodable/entities@2.2.0': - resolution: {integrity: sha512-9uGyhaQavEUMC8AIddIjau4NsnsXhou+j5sBAGojCM1oxmQpVKTWR/9JxABD6UAv12vpIms55fPZKFQEhG6uBg==} + '@nodable/entities@3.0.0': + resolution: {integrity: sha512-8L9xFeTYKhm49xfIypoe2W5wV1m/3Z58kT+7kR9A8OyFxcPduI4VmxaUMQyKYrRjUoLLSXv6EKKID5Tvj9cUVw==} '@nodelib/fs.scandir@2.1.5': resolution: {integrity: sha512-vq24Bq3ym5HEQm2NKCr3yXDwjc7vTsEThRDnkp2DK9p1uqLR+DHurm/NOTo0KG7HYHU7eppKZj3MyqYuMBf62g==} @@ -1213,9 +1311,6 @@ packages: '@protobufjs/float@1.0.2': resolution: {integrity: sha512-Ddb+kVXlXst9d+R9PfTIxh1EdNkgoRe5tOX6t01f1lYWOvJnSPDBlG241QLzcyPdoNTsblLUdujGSE4RzrTZGQ==} - '@protobufjs/inquire@1.1.2': - resolution: {integrity: sha512-pa0vFRuws4wkvaXKK1uXZMAwAX4/t8ANaJo45iw/oQHNQ9q5xUzwgFmVJGXiga2BeN+zpX7Vf9vmsiIa2J+MUw==} - '@protobufjs/path@1.1.2': resolution: {integrity: sha512-6JOcJ5Tm08dOHAbdR3GrvP+yUUfkjG5ePsHYczMFLq3ZmMkAD98cDgcT2iA1lJ9NVwFd4tH/iSSoe44YWkltEA==} @@ -1616,9 +1711,9 @@ packages: bowser@2.14.1: resolution: {integrity: sha512-tzPjzCxygAKWFOJP011oxFHs57HzIhOEracIgAePE4pqB3LikALKnSzUyU4MGs9/iCEUuHlAJTjTc5M+u7YEGg==} - brace-expansion@5.0.7: - resolution: {integrity: sha512-7oFy703dxfY3/NLxC1fh2SUCQ0H9rmAY+5EpDVfXjUTTs+HEwR2nYaqLv+GWcTsumwxPfiz6CzCNkwXwBUwqCA==} - engines: {node: 18 || 20 || >=22} + brace-expansion@5.0.8: + resolution: {integrity: sha512-JZyDyq3D4AUifKTPOB7DELf6XsB3WdPuNxCtob1vFXPsSXhdAiHBWJ/tJ8HAc9aH84BK+5JFZLNkJKx3G9kzQg==} + engines: {node: 20 || >=22} braces@3.0.3: resolution: {integrity: sha512-yQbXgO/OSZVD2IsiLlro+7Hf6Q18EJrKSEsdoMzKePKXct3gvD8oLcOQdIzGupr5Fj+EDe8gO/lxc1BzfMpxvA==} @@ -1683,6 +1778,9 @@ packages: resolution: {integrity: sha512-rcQ1bsQO9799wq24uE5AM2tAILy4gXGIK/njFWcVQkGNZ96edlpY+A7bjwvzjYvLDyzmG1MmMLZhpcsb+klNMQ==} engines: {node: ^12.20.0 || ^14.13.1 || >=16.0.0} + core-js@3.49.0: + resolution: {integrity: sha512-es1U2+YTtzpwkxVLwAFdSpaIMyQaq0PBgm3YD1W3Qpsn1NAmO3KSgZfu+oGSWVu6NvLHoHCV/aYcsE5wiB7ALg==} + cross-spawn@7.0.6: resolution: {integrity: sha512-uV2QOWP2nWzsy2aMp8aRibhi9dlzF5Hgh5SHaB9OiTGEyDTiJJyx0uy51QXdyWbtAHNua4XJzUKca3OzKUd3vA==} engines: {node: '>= 8'} @@ -1745,8 +1843,8 @@ packages: es-module-lexer@1.7.0: resolution: {integrity: sha512-jEQoCwk8hyb2AZziIOLhDqpm5+2ww5uIE6lkO/6jcOCusfk6LhMHpXXfBLXTZ7Ydyt0j4VoUQv6uGNYbdW+kBA==} - es-toolkit@1.49.0: - resolution: {integrity: sha512-G5iZ6Pc/FNRY/soKZHC+TxGDD83rHUDXxzaWhGCX44vAv/tMs56WMusnm/KMNK+luUPsgA9U28cGr4RDlSzL2g==} + es-toolkit@1.50.0: + resolution: {integrity: sha512-OyZKhUVvEep9ITEiwHn8GKnMRQIVqoSIX7WnRbkWgJkllCujilqP2rD0u979tkl8wqyc8ICwlc1UBVv/Sl1G6w==} esbuild@0.21.5: resolution: {integrity: sha512-mg3OPMV4hXywwpoDxu3Qda5xCKQi+vCTZq8S9J/EpkhB2HzKXq4SNFZE3+NK93JYxc8VMSep+lOUSC/RVKaBqw==} @@ -1789,8 +1887,8 @@ packages: fast-xml-builder@1.3.0: resolution: {integrity: sha512-F74cZEdCvuw9P41GAC3rod4X04jjWGM1JPEv/GWSqFTWLsdyMSBMBMlm9Hk3GLBgLBbdBNY8yee0pQh2RBVESQ==} - fast-xml-parser@5.10.0: - resolution: {integrity: sha512-SLhnTEqE5QpJHq/6zl9bsmImEP2adv+y6Wy+cJa7nVTRzQh1OZfCe9k29M5xN74LWnu0xa1zrUrq3KnOKl92Fg==} + fast-xml-parser@5.10.1: + resolution: {integrity: sha512-IEMIf7298kXuZSRFoGfMYrl7is8LpavODgbNz1cwIudv7KwVFnuU+UsMporfq6PD6aXSlawZlARiA3UywCTfMw==} hasBin: true fastq@1.20.1: @@ -1899,6 +1997,10 @@ packages: react-devtools-core: optional: true + ipaddr.js@2.4.0: + resolution: {integrity: sha512-9VGk3HGanVE6JoZXHiCpnGy5X0jYDnN4EA4lntFPj+1vIWlFhIylq2CrrCOJH9EAhc5CYhq18F2Av2tgoAPsYQ==} + engines: {node: '>= 10'} + is-docker@3.0.0: resolution: {integrity: sha512-eljcgEDlEns/7AXFosB5K/2nCM4P7FQPkGc/DWLy5rmFEWvZayGrik1d9/QIY5nJ4f9YsVvBkA6kJpHn9rISdQ==} engines: {node: ^12.20.0 || ^14.13.1 || >=16.0.0} @@ -1940,6 +2042,9 @@ packages: isexe@2.0.0: resolution: {integrity: sha512-RHxMLp9lnKHGHRng9QFhRCMbYAcVpn69smSGcq3f36xjgVVWThj4qqLbTLlq7Ssj8B+fIQ1EuCEGI2lKsyQeIw==} + iso-datestring-validator@2.2.2: + resolution: {integrity: sha512-yLEMkBbLZTlVQqOnQ4FiMujR6T4DEcCb1xizmvXS+OxuhwcbtynoosRzdMA69zZCShCNAbi+gJ71FxZBBXx1SA==} + jazz-napi@2.0.0-alpha.53: resolution: {integrity: sha512-5+EEUoahaNA2d0IqvYDlmrx9bjhLjpvkbt5Nab+C09EeKGnJcz7JJbofC0xuJDe++9I7/8cnYps7YcAYjdK3Ug==} @@ -1987,6 +2092,9 @@ packages: resolution: {integrity: sha512-AC/7JofJvZGrrneWNaEnJeOLUx+JlGt7tNa0wZiRPT4MY1wmfKjt2+6O2p2uz2+skll8OZZmJMNqeke7kKbNgQ==} hasBin: true + jose@5.10.0: + resolution: {integrity: sha512-s+3Al/p9g32Iq+oqXxkW//7jk2Vig6FF1CFqzVXoTUXt2qz89YWbL+OwS17NFYEvxC35n0FKeGO2LGYSxeM2Gg==} + jose@6.2.3: resolution: {integrity: sha512-YYVDInQKFJfR/xa3ojUTl8c2KoTwiL1R5Wg9YCydwH0x0B9grbzlg5HC7mMjCtUJjbQ/YnGEZIhI5tCgfTb4Hw==} @@ -2016,6 +2124,9 @@ packages: loupe@3.2.1: resolution: {integrity: sha512-CdzqowRJCeLU72bHvWqwRBBlLcMEtIvGrlvef74kMnV2AolS9Y8xUv1I0U/MNAWMhBlKIoyuEgoJ0t/bbwHbLQ==} + lru-cache@10.4.3: + resolution: {integrity: sha512-JNAzZcXrCt42VGLuYz0zfAzDfAvJWW6AfYlDBQyDV5DClI2m5sAmK+OIO7s59XfsRsWHp02jAJrRadPRGTt6SQ==} + lru-cache@11.5.2: resolution: {integrity: sha512-4pfM1Ff0x50o0tQwb5ucw/RzNyD0/YJME6IVcStalZuMWxdt3sR3huStTtxz4PUmvZfRguvDejasvQ2kifR11g==} engines: {node: 20 || >=22} @@ -2072,6 +2183,9 @@ packages: ms@2.1.3: resolution: {integrity: sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA==} + multiformats@13.4.2: + resolution: {integrity: sha512-eh6eHCrRi1+POZ3dA+Dq1C6jhP1GNtr9CRINMb67OKzqW9I5DUuZM/3jLPlzhgpGeiNUlEGEbkCYChXMCc/8DQ==} + nanoid@3.3.16: resolution: {integrity: sha512-bzlKTyNJ7+LdGIIwy8ijFpIqEQIvafahV7eYykJ8Cvh42EdJeODoJ6gUJXpQJvej1BddH8OqTXZNE/KfbWAu8Q==} engines: {node: ^10 || ^12 || ^13.7 || ^14 || >=15.0.1} @@ -2173,8 +2287,8 @@ packages: resolution: {integrity: sha512-/FPD0nUc9jH6rfFjji9IBqOz4pcSE3CsT1m7Ep6Mdb0LxSUMj8hgl6GomOvZzpNpAqqGaXA0P3VSrZLFzIhQrw==} engines: {node: '>=12.0.0'} - protobufjs@8.0.1: - resolution: {integrity: sha512-NWWCCscLjs+cOKF/s/XVNFRW7Yih0fdH+9brffR5NZCy8k42yRdl5KlWKMVXuI1vfCoy4o1z80XR/W/QUb3V3w==} + protobufjs@8.4.1: + resolution: {integrity: sha512-oXf2UgIty8jnwfN4yvL1x79VLhL5uiKjZJbSGXGCIUmHmItTP4eS/UIlWDCeNx3seg+ujfn9vDlPMSrsh7wO+Q==} engines: {node: '>=12.0.0'} queue-microtask@1.2.3: @@ -2378,10 +2492,25 @@ packages: undici-types@6.21.0: resolution: {integrity: sha512-iwDZqg0QAGrg9Rav5H4n0M64c3mkR59cJ6wQp+7C4nI0gsmExaedaYLNO44eT4AtBBwjbTiGPMlt2Md0T9H9JQ==} + undici@6.28.0: + resolution: {integrity: sha512-LIY910g9TI13YS95lrMFrs8Rm/u/irgHeTWoKCoteeJ04CUJ92eEfj0rVn+7VKMPBpUPiUoBKfhNyLI23EE/KA==} + engines: {node: '>=18.17'} + + undici@7.29.0: + resolution: {integrity: sha512-IDxfleLmmbSskfWSUATiN1nfn2rDuvnMOqb5CWR92iIfojA0Ud+ulOAAEQ57LPr9rWmsreUyf5lwyao+7GNNVw==} + engines: {node: '>=20.18.1'} + undici@8.5.0: resolution: {integrity: sha512-xamtWoB1EshgjpmlXd7GGm2VfdDtw1+rD8uhry8pSNW3If6S8E0m2T2+orSKeZXEn/aPJMviCpDBA65WJt8zhg==} engines: {node: '>=22.19.0'} + undici@8.9.0: + resolution: {integrity: sha512-aWZpUj7XoGonMClx4gdDRfgBjqeA+F473aDmROQQbM9n6PRfK/u1q/a0X4wMTgcHfT8H6fpbt98PFuDUwFg2YA==} + engines: {node: '>=22.19.0'} + + unicode-segmenter@0.14.5: + resolution: {integrity: sha512-jHGmj2LUuqDcX3hqY12Ql+uhUTn8huuxNZGq7GvtF6bSybzH3aFgedYu/KTzQStEgt1Ra2F3HxadNXsNjb3m3g==} + unist-util-is@6.0.1: resolution: {integrity: sha512-LsiILbtBETkDz8I9p1dQ0uyRUWuaQzd/cuEeS1hoRSyW5E5XGmTzlwY1OrNzzakGowI9Dr/I8HVaw4hTtnxy8g==} @@ -2502,6 +2631,18 @@ packages: utf-8-validate: optional: true + ws@8.21.1: + resolution: {integrity: sha512-+0NTnW77fFN/DjQi6k/Sq/Yvk4Sgajw7urW8V+asjXnRgDs9gyGkdb7EzgfhA4goXsRIZKE28fzIXBHEzhuiWw==} + engines: {node: '>=10.0.0'} + peerDependencies: + bufferutil: ^4.0.1 + utf-8-validate: '>=5.0.2' + peerDependenciesMeta: + bufferutil: + optional: true + utf-8-validate: + optional: true + wsl-utils@0.1.0: resolution: {integrity: sha512-h3Fbisa2nKGPxCpm89Hk33lBLsnaGBvctQopaBSOW/uIs6FTe1ATyAnKFJrzVs9vpGdsTe73WF3V4lIsk4Gacw==} engines: {node: '>=18'} @@ -2523,6 +2664,9 @@ packages: peerDependencies: zod: ^3.25.28 || ^4 + zod@3.25.76: + resolution: {integrity: sha512-gzUt/qt81nXsFGKIFcC3YnfEAx5NkunCfnDlvuBSSFS02bcXu4Lmea0AFIUwbLWxWPx3d9p8S5QoaujKcNQxcQ==} + zod@4.4.3: resolution: {integrity: sha512-ytENFjIJFl2UwYglde2jchW2Hwm4GJFLDiSXWdTrJQBIN9Fcyp7n4DhxJEiWNAJMV1/BqWfW/kkg71UDcHJyTQ==} @@ -2542,6 +2686,144 @@ snapshots: optionalDependencies: zod: 4.4.3 + '@atproto-labs/did-resolver@0.3.5': + dependencies: + '@atproto-labs/fetch': 0.3.4 + '@atproto-labs/pipe': 0.2.4 + '@atproto-labs/simple-store': 0.4.4 + '@atproto-labs/simple-store-memory': 0.2.4 + '@atproto/did': 0.5.4 + zod: 3.25.76 + + '@atproto-labs/fetch-node@0.3.5': + dependencies: + '@atproto-labs/fetch': 0.3.4 + '@atproto-labs/pipe': 0.2.4 + ipaddr.js: 2.4.0 + undici_v6: undici@6.28.0 + undici_v7: undici@7.29.0 + undici_v8: undici@8.9.0 + + '@atproto-labs/fetch@0.3.4': + dependencies: + '@atproto-labs/pipe': 0.2.4 + + '@atproto-labs/handle-resolver-node@0.2.6': + dependencies: + '@atproto-labs/fetch-node': 0.3.5 + '@atproto-labs/handle-resolver': 0.4.6 + '@atproto/did': 0.5.4 + + '@atproto-labs/handle-resolver@0.4.6': + dependencies: + '@atproto-labs/simple-store': 0.4.4 + '@atproto-labs/simple-store-memory': 0.2.4 + '@atproto/did': 0.5.4 + zod: 3.25.76 + + '@atproto-labs/identity-resolver@0.4.5': + dependencies: + '@atproto-labs/did-resolver': 0.3.5 + '@atproto-labs/handle-resolver': 0.4.6 + + '@atproto-labs/pipe@0.2.4': {} + + '@atproto-labs/simple-store-memory@0.2.4': + dependencies: + '@atproto-labs/simple-store': 0.4.4 + lru-cache: 10.4.3 + + '@atproto-labs/simple-store@0.4.4': {} + + '@atproto/common-web@0.5.6': + dependencies: + '@atproto/lex-data': 0.1.5 + '@atproto/lex-json': 0.1.4 + '@atproto/syntax': 0.7.2 + zod: 3.25.76 + + '@atproto/did@0.5.4': + dependencies: + zod: 3.25.76 + + '@atproto/jwk-jose@0.2.4': + dependencies: + '@atproto/jwk': 0.7.4 + jose: 5.10.0 + + '@atproto/jwk-webcrypto@0.3.4': + dependencies: + '@atproto/jwk': 0.7.4 + '@atproto/jwk-jose': 0.2.4 + zod: 3.25.76 + + '@atproto/jwk@0.7.4': + dependencies: + multiformats: 13.4.2 + zod: 3.25.76 + + '@atproto/lex-data@0.1.5': + dependencies: + multiformats: 13.4.2 + tslib: 2.8.1 + unicode-segmenter: 0.14.5 + + '@atproto/lex-json@0.1.4': + dependencies: + '@atproto/lex-data': 0.1.5 + tslib: 2.8.1 + + '@atproto/lexicon@0.7.7': + dependencies: + '@atproto/common-web': 0.5.6 + '@atproto/syntax': 0.7.2 + multiformats: 13.4.2 + zod: 3.25.76 + + '@atproto/oauth-client-node@0.4.9': + dependencies: + '@atproto-labs/did-resolver': 0.3.5 + '@atproto-labs/handle-resolver-node': 0.2.6 + '@atproto-labs/simple-store': 0.4.4 + '@atproto/did': 0.5.4 + '@atproto/jwk': 0.7.4 + '@atproto/jwk-jose': 0.2.4 + '@atproto/jwk-webcrypto': 0.3.4 + '@atproto/oauth-client': 0.7.11 + '@atproto/oauth-types': 0.7.5 + + '@atproto/oauth-client@0.7.11': + dependencies: + '@atproto-labs/did-resolver': 0.3.5 + '@atproto-labs/fetch': 0.3.4 + '@atproto-labs/handle-resolver': 0.4.6 + '@atproto-labs/identity-resolver': 0.4.5 + '@atproto-labs/simple-store': 0.4.4 + '@atproto-labs/simple-store-memory': 0.2.4 + '@atproto/did': 0.5.4 + '@atproto/jwk': 0.7.4 + '@atproto/oauth-types': 0.7.5 + '@atproto/xrpc': 0.8.6 + core-js: 3.49.0 + multiformats: 13.4.2 + zod: 3.25.76 + + '@atproto/oauth-types@0.7.5': + dependencies: + '@atproto/did': 0.5.4 + '@atproto/jwk': 0.7.4 + zod: 3.25.76 + + '@atproto/syntax@0.7.2': + dependencies: + iso-datestring-validator: 2.2.2 + tslib: 2.8.1 + + '@atproto/xrpc@0.8.6': + dependencies: + '@atproto/lexicon': 0.7.7 + zod: 3.25.76 + '@aws-crypto/sha256-browser@5.2.0': dependencies: '@aws-crypto/sha256-js': 5.2.0 @@ -3376,7 +3658,7 @@ snapshots: '@noble/hashes@2.2.0': {} - '@nodable/entities@2.2.0': {} + '@nodable/entities@3.0.0': {} '@nodelib/fs.scandir@2.1.5': dependencies: @@ -3441,7 +3723,7 @@ snapshots: '@opentelemetry/sdk-logs': 0.216.0(@opentelemetry/api@1.9.1) '@opentelemetry/sdk-metrics': 2.7.1(@opentelemetry/api@1.9.1) '@opentelemetry/sdk-trace-base': 2.7.1(@opentelemetry/api@1.9.1) - protobufjs: 8.0.1 + protobufjs: 8.4.1 '@opentelemetry/resources@2.7.1(@opentelemetry/api@1.9.1)': dependencies: @@ -3520,8 +3802,6 @@ snapshots: '@protobufjs/float@1.0.2': {} - '@protobufjs/inquire@1.1.2': {} - '@protobufjs/path@1.1.2': {} '@protobufjs/pool@1.1.0': {} @@ -3887,7 +4167,7 @@ snapshots: bowser@2.14.1: {} - brace-expansion@5.0.7: + brace-expansion@5.0.8: dependencies: balanced-match: 4.0.4 @@ -3944,6 +4224,8 @@ snapshots: convert-to-spaces@2.0.1: {} + core-js@3.49.0: {} + cross-spawn@7.0.6: dependencies: path-key: 3.1.1 @@ -3987,7 +4269,7 @@ snapshots: es-module-lexer@1.7.0: {} - es-toolkit@1.49.0: {} + es-toolkit@1.50.0: {} esbuild@0.21.5: optionalDependencies: @@ -4125,9 +4407,9 @@ snapshots: path-expression-matcher: 1.6.2 xml-naming: 0.3.0 - fast-xml-parser@5.10.0: + fast-xml-parser@5.10.1: dependencies: - '@nodable/entities': 2.2.0 + '@nodable/entities': 3.0.0 fast-xml-builder: 1.3.0 is-unsafe: 2.0.0 path-expression-matcher: 1.6.2 @@ -4259,7 +4541,7 @@ snapshots: cli-cursor: 4.0.0 cli-truncate: 6.1.1 code-excerpt: 4.0.0 - es-toolkit: 1.49.0 + es-toolkit: 1.50.0 indent-string: 5.0.0 is-in-ci: 2.0.0 patch-console: 2.0.0 @@ -4274,12 +4556,14 @@ snapshots: type-fest: 5.8.0 widest-line: 6.0.0 wrap-ansi: 10.0.0 - ws: 8.21.0 + ws: 8.21.1 yoga-layout: 3.2.1 transitivePeerDependencies: - bufferutil - utf-8-validate + ipaddr.js@2.4.0: {} + is-docker@3.0.0: {} is-extglob@2.1.1: {} @@ -4308,6 +4592,8 @@ snapshots: isexe@2.0.0: {} + iso-datestring-validator@2.2.2: {} + jazz-napi@2.0.0-alpha.53: optionalDependencies: '@garden-co/jazz-napi-darwin-arm64': 2.0.0-alpha.53 @@ -4339,6 +4625,8 @@ snapshots: jiti@2.7.0: {} + jose@5.10.0: {} + jose@6.2.3: {} js-tokens@4.0.0: {} @@ -4371,6 +4659,8 @@ snapshots: loupe@3.2.1: {} + lru-cache@10.4.3: {} + lru-cache@11.5.2: {} lru_map@0.4.1: {} @@ -4421,12 +4711,14 @@ snapshots: minimatch@10.2.5: dependencies: - brace-expansion: 5.0.7 + brace-expansion: 5.0.8 minipass@7.1.3: {} ms@2.1.3: {} + multiformats@13.4.2: {} + nanoid@3.3.16: {} node-addon-api@7.1.1: {} @@ -4523,19 +4815,8 @@ snapshots: '@types/node': 22.20.1 long: 5.3.2 - protobufjs@8.0.1: + protobufjs@8.4.1: dependencies: - '@protobufjs/aspromise': 1.1.2 - '@protobufjs/base64': 1.1.2 - '@protobufjs/codegen': 2.0.5 - '@protobufjs/eventemitter': 1.1.1 - '@protobufjs/fetch': 1.1.1 - '@protobufjs/float': 1.0.2 - '@protobufjs/inquire': 1.1.2 - '@protobufjs/path': 1.1.2 - '@protobufjs/pool': 1.1.0 - '@protobufjs/utf8': 1.1.2 - '@types/node': 22.20.1 long: 5.3.2 queue-microtask@1.2.3: {} @@ -4770,8 +5051,16 @@ snapshots: undici-types@6.21.0: {} + undici@6.28.0: {} + + undici@7.29.0: {} + undici@8.5.0: {} + undici@8.9.0: {} + + unicode-segmenter@0.14.5: {} + unist-util-is@6.0.1: dependencies: '@types/unist': 3.0.3 @@ -4892,6 +5181,8 @@ snapshots: ws@8.21.0: {} + ws@8.21.1: {} + wsl-utils@0.1.0: dependencies: is-wsl: 3.1.1 @@ -4906,6 +5197,8 @@ snapshots: dependencies: zod: 4.4.3 + zod@3.25.76: {} + zod@4.4.3: {} zwitch@2.0.4: {} diff --git a/prompts/conceptualizer.md b/prompts/conceptualizer.md new file mode 100644 index 0000000..c913cdd --- /dev/null +++ b/prompts/conceptualizer.md @@ -0,0 +1,9 @@ +# Conceptualizer + +Extract a small concept graph from the supplied canonical ThoughtStream event packet. + +Concepts should name durable ideas present in the evidence, not people, handles, URLs, platforms, or incidental nouns. Use lowercase one-to-three-word phrases. Prefer a few specific concepts over a large generic cloud. + +Each concept's `relationship` describes how that concept relates to the source packet. Use indexed `links` only for a meaningful directional relationship between two extracted concepts. Do not add a link merely because two concepts co-occur. + +The output is a private proposal. It is not an ATProto record, publication request, database mutation, or claim that any external action occurred. Return only the strict conceptualization JSON contract supplied by the runtime. diff --git a/prompts/review-candidate.md b/prompts/review-candidate.md new file mode 100644 index 0000000..057c373 --- /dev/null +++ b/prompts/review-candidate.md @@ -0,0 +1,7 @@ +# Review candidate + +The current source event contains one complete review prompt, optional evidence, and review metadata. + +Answer the payload's `prompt` using the supplied `evidence` when present. Do not discuss the fact that this is a comparison, predict the human preference, name candidate A or B, or optimize for the review rubric as a separate task. Produce the best direct response you would have given to the prompt itself. + +Return only the strict review-response JSON contract supplied by the runtime. External actions and tools are unavailable. diff --git a/public/docs/architecture.md b/public/docs/architecture.md new file mode 100644 index 0000000..9b7b932 --- /dev/null +++ b/public/docs/architecture.md @@ -0,0 +1,23 @@ +# Architecture + +ThoughtStream is organized as a set of narrow processes joined by durable Jazz records. + +## Ingress + +Connectors observe sources and append canonical events with source-local sequence, idempotency, privacy, and lineage. Observation never grants authority to act on a source. + +## Consumers + +Consumers subscribe through bounded Jazz queries. A consumer receives a bounded context packet, runs a deterministic rule or isolated model cell, validates its declared output contract, and settles output, lifecycle evidence, and progress together. + +## Projections + +Projections rebuild query-oriented views from canonical events. They make the stream inspectable without becoming a second source of truth. + +## Egress + +Dispatchers independently select eligible completed work, enforce destination policy, perform one narrow action, and append started, delivered, or failed receipts. A completed model run is not evidence that anything was sent. + +## Recovery + +Connector cursors advance only after durable events. Consumer progress advances only after terminal settlement. Interrupted attempts remain visible and retries preserve their evidence rather than rewriting history. diff --git a/public/docs/index.md b/public/docs/index.md new file mode 100644 index 0000000..6517be3 --- /dev/null +++ b/public/docs/index.md @@ -0,0 +1,12 @@ +# Documentation + +ThoughtStream is a spec-driven event system. Its central design claim is simple: an event, a model run, and an external action are different facts and require different receipts. + +## Read next + +- Architecture: components, authority, and the path from ingress to durable evidence. +- Security: containment, privacy, and why the public website cannot read the private stream. + +## Current boundary + +This documentation is public and static. Operational data remains private. The private inspector requires an allowlisted ATProto identity or the explicitly configured break-glass Basic credential. diff --git a/public/docs/security.md b/public/docs/security.md new file mode 100644 index 0000000..6b0116b --- /dev/null +++ b/public/docs/security.md @@ -0,0 +1,21 @@ +# Security + +ThoughtStream has broad read access, so containment is a product property rather than deployment polish. + +## Capability separation + +Ingress can observe. Model cells can receive bounded context and produce typed proposals. Dispatchers hold narrow, explicit action capabilities. No layer inherits another layer's authority by proximity. + +## Private data + +Events carry a privacy class. Public-source derivations remain private by default because selection, routing, and model context can reveal private interests. Credentials, raw provider bodies, model reasoning, and malformed output are excluded from durable evidence. + +## Public website + +The public server renders only a fixed list of reviewed Markdown files. It has no Jazz handle, event query, manifest reader, trace reader, directory listing, arbitrary file route, or private upstream fallback. Unknown routes stop locally. + +## Inspector authentication + +The private inspector accepts an allowlisted ATProto identity through the official OAuth client. Protocol state and tokens stay encrypted on the server; the browser receives only a short-lived opaque session cookie. A separately configured Basic credential remains available as a break-glass path during activation. + +OAuth proves identity only. It grants no write, publish, connector, model, dispatcher, or PDS-record authority. diff --git a/public/index.md b/public/index.md new file mode 100644 index 0000000..d03b2b2 --- /dev/null +++ b/public/index.md @@ -0,0 +1,13 @@ +# thought stream + +A durable event fabric for agents, sources, and the evidence between them. + +ThoughtStream separates observation, model execution, projection, and action. Sources become append-only events. Consumers run against bounded context. Outputs settle with lineage and terminal receipts. External actions remain separate capabilities. + +## What this site contains + +- Public architecture and security documentation compiled from reviewed repository files. +- ATProto OAuth endpoints used to authenticate the private inspector. +- No live events, source names, manifests, traces, credentials, service metadata, or runtime state. + +The inspector is private. Public documentation describes the system; it is not a window into the system's data. diff --git a/scripts/configure-inspector-credentials.sh b/scripts/configure-inspector-credentials.sh new file mode 100755 index 0000000..dcec9fc --- /dev/null +++ b/scripts/configure-inspector-credentials.sh @@ -0,0 +1,45 @@ +#!/usr/bin/env bash +set -euo pipefail + +credentials_dir="${THOUGHTSTREAM_CREDENTIALS_DIR:-$HOME/.config/thoughtstream/credentials}" +destination="$credentials_dir/inspector-proxy.env" + +if [[ -L "$credentials_dir" ]]; then + printf 'Refusing symlinked ThoughtStream credentials directory: %s\n' "$credentials_dir" >&2 + exit 1 +fi +mkdir -p -m 0700 "$credentials_dir" +chmod 0700 "$credentials_dir" + +read -r -p 'Username [cameron]: ' username +username="${username:-cameron}" +if [[ ! "$username" =~ ^[A-Za-z0-9._-]+$ ]]; then + printf 'Username may use only letters, numbers, dot, underscore, or hyphen.\n' >&2 + exit 1 +fi + +read -r -s -p 'Password (20+ bytes): ' password +printf '\n' +read -r -s -p 'Confirm password: ' confirmation +printf '\n' +if [[ "$password" != "$confirmation" ]]; then + printf 'Passwords do not match.\n' >&2 + exit 1 +fi +if (( $(printf '%s' "$password" | wc -c) < 20 )); then + printf 'Password must contain at least 20 UTF-8 bytes.\n' >&2 + exit 1 +fi + +umask 077 +temporary="$(mktemp "$credentials_dir/.inspector-proxy.env.tmp.XXXXXX")" +trap 'rm -f "$temporary"; unset password confirmation encoded' EXIT +encoded="$(printf '%s' "$password" | base64 -w 0)" +printf 'PROXY_USER=%s\nPROXY_PASSWORD_B64=%s\nPROXY_HOST=127.0.0.1\nPROXY_PORT=4319\nPROXY_UPSTREAM_HOST=127.0.0.1\nPROXY_UPSTREAM_PORT=4317\nPROXY_BASIC_FALLBACK_ENABLED=1\n' \ + "$username" "$encoded" > "$temporary" +chmod 0600 "$temporary" +mv -f "$temporary" "$destination" +unset password confirmation encoded +trap - EXIT +printf 'Wrote owner-only inspector credentials to %s.\n' "$destination" +printf 'Restart with: systemctl --user restart thoughtstream-inspector-proxy.service\n' diff --git a/scripts/configure-inspector-oauth.ts b/scripts/configure-inspector-oauth.ts new file mode 100644 index 0000000..5ad1082 --- /dev/null +++ b/scripts/configure-inspector-oauth.ts @@ -0,0 +1,85 @@ +import { randomBytes } from "node:crypto"; +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { JoseKey } from "@atproto/oauth-client-node"; + +const args = parseArgs(process.argv.slice(2)); +const origin = new URL(required(args.origin, "--origin")); +if (origin.protocol !== "https:" || origin.pathname !== "/" || origin.search || origin.hash || origin.username || origin.password) { + throw new Error("--origin must be one bare HTTPS origin, for example https://thought.stream"); +} +const did = required(args.did, "--did"); +if (!/^did:(plc|web):/.test(did)) throw new Error("--did must be an ATProto did:plc or did:web identity"); +const handle = required(args.handle, "--handle"); +if (!/^[a-z0-9][a-z0-9.-]+$/i.test(handle)) throw new Error("--handle is invalid"); +const credentialsDirectory = process.env.THOUGHTSTREAM_CREDENTIALS_DIR + ?? path.join(os.homedir(), ".config", "thoughtstream", "credentials"); +if (process.env.THOUGHTSTREAM_OAUTH_STORE_DIR) { + throw new Error("THOUGHTSTREAM_OAUTH_STORE_DIR is unsupported; the systemd sandbox permits only ~/.local/share/thoughtstream-inspector-auth"); +} +const storeDirectory = path.join(os.homedir(), ".local", "share", "thoughtstream-inspector-auth"); +const destination = path.join(credentialsDirectory, "inspector-oauth.env"); + +await refuseSymlink(credentialsDirectory); +await fs.mkdir(credentialsDirectory, { recursive: true, mode: 0o700 }); +await fs.chmod(credentialsDirectory, 0o700); +await refuseSymlink(storeDirectory); +await fs.mkdir(storeDirectory, { recursive: true, mode: 0o700 }); +await fs.chmod(storeDirectory, 0o700); +if (!args.force) { + const exists = await fs.stat(destination).then(() => true, (error: NodeJS.ErrnoException) => error.code === "ENOENT" ? false : Promise.reject(error)); + if (exists) throw new Error(`OAuth configuration already exists at ${destination}; use --force only for deliberate key rotation`); +} + +const key = await JoseKey.generate(["ES256"], `thoughtstream-${new Date().toISOString().slice(0, 10)}`); +if (!key.privateJwk) throw new Error("Generated OAuth client key has no private material"); +const lines = [ + "OAUTH_ENABLED=1", + `OAUTH_PUBLIC_ORIGIN=${origin.origin}`, + `OAUTH_ALLOWED_DID=${did}`, + `OAUTH_EXPECTED_HANDLE=${handle}`, + `OAUTH_STORE_DIR=${storeDirectory}`, + `OAUTH_STORE_KEY_B64=${randomBytes(32).toString("base64")}`, + `OAUTH_PRIVATE_JWK_B64=${Buffer.from(JSON.stringify(key.privateJwk), "utf8").toString("base64")}`, + "OAUTH_SESSION_TTL_MS=43200000", + "PROXY_BASIC_FALLBACK_ENABLED=1", + "", +]; +const temporary = `${destination}.tmp-${process.pid}-${randomBytes(6).toString("hex")}`; +await fs.writeFile(temporary, lines.join("\n"), { mode: 0o600, flag: "wx" }); +await fs.rename(temporary, destination); +await fs.chmod(destination, 0o600); +process.stdout.write(`Wrote owner-only OAuth configuration to ${destination}\n`); +process.stdout.write("Basic fallback remains enabled. Do not disable it until a real HTTPS OAuth login, inspector read, and logout are verified.\n"); + +function parseArgs(values: string[]): Record { + const result: Record = {}; + for (let index = 0; index < values.length; index += 1) { + const current = values[index]!; + if (current === "--") continue; + if (current === "--force") { + result.force = true; + continue; + } + if (!["--origin", "--did", "--handle"].includes(current)) throw new Error(`Unknown argument: ${current}`); + const value = values[index + 1]; + if (!value) throw new Error(`Missing value for ${current}`); + result[current.slice(2)] = value; + index += 1; + } + return result; +} + +function required(value: string | boolean | undefined, label: string): string { + if (typeof value !== "string" || !value.trim()) throw new Error(`${label} is required`); + return value.trim(); +} + +async function refuseSymlink(target: string): Promise { + const stat = await fs.lstat(target).catch((error: NodeJS.ErrnoException) => { + if (error.code === "ENOENT") return undefined; + throw error; + }); + if (stat?.isSymbolicLink()) throw new Error(`Refusing symlinked directory: ${target}`); +} diff --git a/scripts/configure-inspector-review.ts b/scripts/configure-inspector-review.ts new file mode 100644 index 0000000..274e6f0 --- /dev/null +++ b/scripts/configure-inspector-review.ts @@ -0,0 +1,38 @@ +import { randomBytes } from "node:crypto"; +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; + +const force = process.argv.slice(2).includes("--force"); +if (process.argv.slice(2).some((value) => value !== "--force")) { + throw new Error("Usage: configure-inspector-review [--force]"); +} +const credentialsDirectory = process.env.THOUGHTSTREAM_CREDENTIALS_DIR + ?? path.join(os.homedir(), ".config", "thoughtstream", "credentials"); +const destination = path.join(credentialsDirectory, "inspector-review.env"); +await refuseSymlink(credentialsDirectory); +await fs.mkdir(credentialsDirectory, { recursive: true, mode: 0o700 }); +await fs.chmod(credentialsDirectory, 0o700); +if (!force) { + const exists = await fs.stat(destination).then(() => true, (error: NodeJS.ErrnoException) => + error.code === "ENOENT" ? false : Promise.reject(error)); + if (exists) throw new Error(`Review capability already exists at ${destination}; use --force only for deliberate rotation`); +} +const temporary = `${destination}.tmp-${process.pid}-${randomBytes(6).toString("hex")}`; +await fs.writeFile( + temporary, + `THOUGHTSTREAM_REVIEW_CAPABILITY_B64=${randomBytes(32).toString("base64")}\n`, + { mode: 0o600, flag: "wx" }, +); +await fs.rename(temporary, destination); +await fs.chmod(destination, 0o600); +process.stdout.write(`Wrote one owner-only Review capability to ${destination}.\n`); +process.stdout.write("Both inspector services must load the same file. Generation does not restart or activate either service.\n"); + +async function refuseSymlink(target: string): Promise { + const stat = await fs.lstat(target).catch((error: NodeJS.ErrnoException) => { + if (error.code === "ENOENT") return undefined; + throw error; + }); + if (stat?.isSymbolicLink()) throw new Error(`Refusing symlinked directory: ${target}`); +} diff --git a/scripts/letta-agent-sdk-canary.ts b/scripts/letta-agent-sdk-canary.ts index daf36f2..dac32ef 100644 --- a/scripts/letta-agent-sdk-canary.ts +++ b/scripts/letta-agent-sdk-canary.ts @@ -182,7 +182,7 @@ function executionReceipt(traces: Array<{ kind: string; data: unknown }>) { return { backend: "cloud", sdkVersion: LETTA_AGENT_SDK_PACKAGE_VERSION, - adapterRevision: LETTA_AGENT_SDK_ADAPTER_REVISION, + executionAdapterRevision: LETTA_AGENT_SDK_ADAPTER_REVISION, conversationId: stringField(resultRecord, "conversationId") || stringField(sentRecord, "conversationId"), sdkRunIds: Array.isArray(resultRecord?.runIds) ? resultRecord.runIds.filter((value): value is string => typeof value === "string") diff --git a/scripts/serve-inspector-proxy.ts b/scripts/serve-inspector-proxy.ts new file mode 100755 index 0000000..ead34cc --- /dev/null +++ b/scripts/serve-inspector-proxy.ts @@ -0,0 +1,41 @@ +import { + authenticatedProxyOptionsFromEnv, + startAuthenticatedInspectorProxy, +} from "../src/web/authenticated-proxy.js"; +import { + createInspectorOAuthAuth, + oauthConfigurationFromEnv, +} from "../src/web/oauth-auth.js"; + +const options = authenticatedProxyOptionsFromEnv(); +const oauthConfiguration = oauthConfigurationFromEnv(); +const oauth = oauthConfiguration.enabled + ? await createInspectorOAuthAuth({ + publicOrigin: oauthConfiguration.publicOrigin!, + allowedDid: oauthConfiguration.allowedDid!, + expectedHandle: oauthConfiguration.expectedHandle!, + storeDirectory: oauthConfiguration.storeDirectory!, + storeKey: oauthConfiguration.storeKey!, + privateJwk: oauthConfiguration.privateJwk!, + ...(oauthConfiguration.sessionTtlMs ? { sessionTtlMs: oauthConfiguration.sessionTtlMs } : {}), + }) + : undefined; + +let server: Awaited>; +try { + server = await startAuthenticatedInspectorProxy({ ...options, ...(oauth ? { oauth } : {}) }); +} catch (error) { + await oauth?.close(); + throw error; +} +process.stdout.write(`thought stream web proxy: http://${options.host}:${options.port} oauth=${oauth ? "enabled" : "disabled"} basicFallback=${options.basicFallbackEnabled !== false ? "enabled" : "disabled"}\n`); + +for (const signal of ["SIGINT", "SIGTERM"] as const) { + process.once(signal, () => { + server.closeAllConnections(); + server.close(() => { + void oauth?.close().finally(() => process.exit(0)); + }); + setTimeout(() => process.exit(1), 2_000).unref(); + }); +} diff --git a/spec/README.md b/spec/README.md index fb19a69..b488bca 100644 --- a/spec/README.md +++ b/spec/README.md @@ -15,11 +15,13 @@ The core local milestone is implemented and exercised in `test/agent-runtime.tes - [`agents.md`](agents.md): consumer declarations, subscriptions, execution, outputs, and traces. - [`harnesses.md`](harnesses.md): generic container-harness contract, isolation profiles, persistent workspace/session leases, and the Pi coding reference adapter. - [`repairs.md`](repairs.md): deterministic repair eligibility, sandboxed correction proposals, judgment authority, effective-output rebuilding, and training boundaries. +- [`review.md`](review.md): complete review prompts, blinded candidate pairs, judgeability, append-only human decisions, OAuth-only browser writes, and training-data custody. - [`incidents.md`](incidents.md): content-dark operational incident projection, private ledger, and independent Telegram alert policy. - [`tinker.md`](tinker.md): Tinker model and adapter boundary. - [`security.md`](security.md): privacy, credentials, authority, and prompt-injection boundaries. - [`recovery.md`](recovery.md): producer cursors, consumer progress, retries, replay, and terminal evidence. - [`ui.md`](ui.md): thin root-activity interface. +- [`web-auth.md`](web-auth.md): public/private route matrix, OAuth boundary, threat model, and activation proof. - [`testing.md`](testing.md): required tests and acceptance scenario. ## First implementation milestone diff --git a/spec/agents.md b/spec/agents.md index 99d1837..571c8cf 100644 --- a/spec/agents.md +++ b/spec/agents.md @@ -30,6 +30,8 @@ policy: externalActions: false ``` +A Pi runner selects one exact learned Tinker release through `adapter: { id, version }`. Startup resolves it from the validated immutable release and deployment catalogs before the declaration becomes a Jazz row. A learned adapter is mutually exclusive with `runner.model` and `runner.tier`, must match the trusted provider profile, and receives a runtime binding only when its deployment entry is active. The deeply frozen declaration contains the public content-addressed release/binding and catalog identity. Only the resolved checkpoint value remains in module-private weak bindings keyed by the original compiled adapter/declaration objects. Cloning, rehydration, manual construction, or mutation cannot recreate dispatch authority. Declarations without a learned adapter remain backward compatible. + Every `pi` and `letta-agent-sdk` declaration also requires an `accounting` policy. It declares a conservative per-call reservation, a lease longer than the runner timeout, and one or more rolling/hour/day limits. Calls, input tokens, and output tokens are always reserved; micro-US-dollar cost reservation and limits are optional. A cost limit without a cost reservation is invalid because the runtime cannot enforce a dimension it did not reserve. A declaration that cannot admit one complete reservation in every tracked dimension is invalid. Deterministic consumers cannot declare inference accounting. The resident keeps measured token accounting and a conservative full-conversation reservation of 60,000 input / 2,000 output tokens per call, rather than estimating from the current packet size. Its tracked limits are emergency-only circuit breakers: 100 calls / 100,000,000 input / 10,000,000 output per rolling five minutes; 1,000 calls / 1,000,000,000 input / 100,000,000 output per hour; and 10,000 calls / 1,000,000,000 input / 1,000,000,000 output per day. Dollar cost is absent. This accounting is operational telemetry plus a catastrophic-runaway guard, not a budget or thrift mechanism; ordinary or extreme post, like, and Semble activity should not approach the ceilings. @@ -62,17 +64,21 @@ The packet contains exact ids and bounded rendered content: - Explicitly requested neighboring events. - Agent prompt revision. - Output schemas. -- Runtime/model/adapter revision. +- Runtime model, execution-harness adapter revision, and optional learned model-adapter identity as separate fields. - Privacy and tool capability statement. The packet records omitted/truncated content. Silent truncation is forbidden. +The Pi conceptualizer is the narrow non-resident user of `atproto-batch`. It binds the conceptualization output contract to `stream.thought.derived.concept.graph`, subscribes to one ATProto batch event type, uses no tools or external actions, and receives member-expanded context from the trusted parent. Other Pi declarations cannot opt into ATProto object expansion by configuration alone. + ## Model cells and agent harnesses The trusted consumer runtime owns context selection, provider authorization, persistence, output validation, accounting, concurrency, and retry policy. A model runtime or agent harness is never the event store and receives no external-action capability. Ordinary observation and repair consumers use the `observer-v1` model cell: Pi agent core runs with `tools: []` in the existing disposable Bubblewrap sandbox. It has no workspace and a single-use broker capability permits one bounded provider request. This remains an inference-only profile; it is not evidence that a coding harness or arbitrary Pi extension is safe. +Provider JSON response formatting is a capability claim, not a request decoration. The built-in Tinker profile does not advertise it because observed Tinker routes accept `response_format: {"type":"json_object"}` without enforcing a JSON object. Strict JSON output on that profile therefore means explicit prompt instructions plus complete parent-side contract validation and rejection; it must not be described as constrained decoding. The separate `openai-json-default` profile fixes the OpenAI endpoint, credential name, and allowlisted model set and binds the conceptualization declaration to one trusted strict JSON Schema response format. The schema constrains object fields and enums at decoding time; parent validation still enforces cross-field graph invariants and remains authoritative. + Tool-using or persistent agent runtimes use the separate generic contract in `harnesses.md`. The first reference adapter is `pi-coding@1` in the `workspace-v1` container profile. Declarations select only an allowlisted adapter/profile and trusted model tier. They cannot supply an image, host mount, executable extension, provider URL, credential reference, broker budget, or container flag. Any container, broker, or lease setup failure leaves source progress unchanged for a later retry. Neither profile has an in-process or trusted-host fallback. The `letta-agent-sdk` runner is a distinct stateful harness adapter. Its first supported profile is `letta-cloud-v1`: the trusted consumer uses the Letta Agent SDK Cloud backend, while Letta supplies the managed sandbox and persistent agent/conversation state. Jazz remains authoritative for source events, consumer progress, lifecycle evidence, accepted outputs, and channel delivery. The Letta agent owns its conversational continuity and agent memory. ThoughtStream must not rebuild a synthetic transcript and send it again on every turn. @@ -116,7 +122,9 @@ Before the broker or sandbox can dispatch a provider request, the trusted parent Observer-cell tools are explicit declaration capabilities executed by the trusted parent before sandbox launch. The initial read-only set can dereference the current ATProto record through `atproto.md` and download only image URLs discovered in that record or its fetched Markdown. The runtime validates public destinations, follows a bounded redirect chain, caps response bytes and image count, and persists content-addressed image artifacts. A declaration without a tool name receives no corresponding evidence. Workspace-harness tools execute only inside the leased container boundary and are fixed by its adapter/profile pair. -The runner maps sandbox and Pi events to metadata-only trace chunks. Durable traces keep event type, role/model/stop metadata, content counts, and content hashes. They do not keep prompt text, provider thinking, final model text, tool arguments, provider bodies, image bytes, source bodies, or arbitrary provider errors. The provider response remains process-local long enough to validate the typed final part, then leaves no raw-content copy in the trace or lifecycle event stream. +For a learned adapter, startup compilation validates one exact active deployment selection and keeps its checkpoint in a process-local private binding. Read-only evidence prefetch completes before provider egress. The Pi runner resolves the checkpoint only from the original compiler-bound adapter object, then gives it to the trusted broker as the provider model. One serialized admission section rechecks expiry and atomically reserves request count plus cumulative request/response bytes before fetch; response reservation converts to actual usage in the same critical section. Durable broker traces render the public base model, learned-adapter identity, and catalog digest/generation, never the private checkpoint value. + +The runner maps sandbox and Pi events to metadata-only trace chunks. Durable traces keep event type, role/model/stop metadata, content counts, and content hashes. Incremental `pi.message_update` callbacks are redundant with the final message receipt and may number in the hundreds, so the trusted parent replaces them with one `pi.message_updates_coalesced` count before Jazz persistence. They do not keep prompt text, provider thinking, final model text, tool arguments, provider bodies, image bytes, source bodies, or arbitrary provider errors. The provider response remains process-local long enough to validate the typed final part, then leaves no raw-content copy in the trace or lifecycle event stream. ## Outputs @@ -142,7 +150,7 @@ Completed output is not implicit approval. `stream.thought.judgment.training-exa Default export includes only active judgments with both fields true and an entirely `public-source` source/output chain. Telegram reaction projection writes quality-eligible, externally ineligible judgments. Legacy v1 `exportEligible` records remain readable but can authorize default export only for entirely public chains. Sensitive/private external eligibility requires explicit authorization at judgment creation and a second explicit private-export gate at dataset creation. -`thoughtstream.training-example.v2` contains validated chosen/rejected structured output, a minimal source classification (`type`, schema version, source kind, privacy), sanitized context policy, trace type/order, output-contract identity, and model/adapter provenance. It excludes source payloads, actor/route/external/correlation/idempotency identifiers, event/run/delivery ids, source or trace-content hashes, trace timestamps, arbitrary context fields, prompts, provider content, thinking, tool arguments, quarantine, and legacy raw output. `thoughtstream.training-dataset-manifest.v2` contains the dataset content hash, count, kind distribution, and model set without judgment or event ids. Repair runs remain narrower: only active quality-eligible and externally eligible `accept` or `correct` judgments may export. +Legacy `thoughtstream.training-example.v3` contains validated chosen/rejected structured output, a minimal source classification (`type`, schema version, source kind, joined privacy), sanitized context policy, trace type/order, output-contract identity, exact primary execution/model-adapter provenance, and for preferences a separate exact compared provenance and compared trace. Pairwise projection independently applies privacy and learned-adapter export policy to both runs. It excludes source payloads, actor/route/external/correlation/idempotency identifiers, event/run/delivery ids, source or trace-content hashes, trace timestamps, arbitrary context fields, prompts, provider content, thinking, tool arguments, quarantine, and legacy raw output. Review-derived examples use `thoughtstream.training-example.v4` and may include only the exact preauthorized public review prompt/evidence plus campaign and bounded decision metadata described in [`review.md`](review.md). `thoughtstream.training-dataset-manifest.v4` contains the dataset content hash, count, kind distribution, participating model sets, mixed example-format counts, and Review campaign identities without judgment or event ids. Legacy v1/v2 judgment events remain readable but project into v3; Review decisions project into v4. Repair runs remain narrower: only active quality-eligible and externally eligible `accept` or `correct` judgments may export. ## Lifecycle @@ -154,6 +162,8 @@ Model-backed consumer attempts use: Lifecycle events are authoritative evidence. An optional execution row materializes the current state for inspection; it is not a queue claim. A started attempt that survives a process restart without terminal evidence becomes `status-unknown` or `abandoned` according to consumer policy before a new attempt begins. Repair consumers are stricter: an interrupted repair attempt is terminally abandoned and advances request progress, because one repair request authorizes only one proposal generation. +Every started, completed, failed, blocked, or abandoned model run binds `executionAdapterRevision` and, when selected, the exact `modelAdapter` identity. The output event and context manifest carry the same learned-adapter identity; replay compares that identity rather than re-resolving a later active release. + `completed` means output events, terminal evidence, execution state, and consumed source progress settled together at the configured durability tier. A process exiting after model response but before that transaction settles is not completed and may repeat the external call on recovery. The earlier nonterminal attempt remains visible. ## Escalation diff --git a/spec/architecture.md b/spec/architecture.md index 60ba8ed..fc235c2 100644 --- a/spec/architecture.md +++ b/spec/architecture.md @@ -40,6 +40,8 @@ The resident consumes direct Telegram source events or explicit derived ATProto For model-backed consumers, the runner builds a bounded context packet, invokes either an inference cell or an allowlisted container harness through a capability-scoped provider boundary, captures metadata-only events and usage, validates final structured output, then inserts derived events. The existing Pi/Bubblewrap path is an inference cell. Full agent runtimes use the generic contract in `harnesses.md`; Pi coding-agent is the first adapter. The runner never edits its triggering event. +A conceptualizer consumer is a model-backed consumer that extracts concepts and directional links from one bounded source event. For an ATProto batch trigger, the trusted parent resolves every exact member reference and compiles the same bounded ATProto member views used by the resident path before inference; a list of opaque member ids is not sufficient conceptual evidence. It uses the `stream.thought.output.conceptualization@1` output contract, which validates one complete bounded graph of lowercase concept phrases and typed links. The runner settles that graph atomically as one `stream.thought.derived.concept.graph` event with exact source/run lineage. The conceptualizer has no outbound PDS authority; it produces only a private derived observation. Invalid graphs fail as a whole and emit no derived graph, so a partially parsed model answer cannot become durable state. Tinker is the inference provider; the conceptualizer does not train, sync an external graph, or publish to a PDS. + ### 6. Projections Projectors are consumers with no special delivery path. They follow Jazz subscriptions, recover from per-source progress, and maintain rebuildable Jazz views: root activity, topic index, unresolved recommendations, source health, consumer health, document identity, and trace summaries. Jazz performs filtering, ordering, and bounded pagination. @@ -48,9 +50,9 @@ Projectors are consumers with no special delivery path. They follow Jazz subscri External actions are owned by separate destination-specific dispatcher processes. A dispatcher reads completed candidate activity from Jazz, filters it against channel policy, accumulates and renders batches, applies destination velocity limits, performs the action, and appends started/delivered/failed evidence. Producers and consumers never wait on dispatcher policy or destination throughput. The first implementation supports Telegram delivery only. -### 8. Local interface +### 8. Local interface and Review authority -The first interface is a local server and dense activity page. It reads projections and can open the full event/execution lineage. It is not runtime authority. +The interface is a loopback server and dense activity/Review page. Activity, execution, source-health, lineage, and adapter inventory are read projections. Review adds one separately configured mutation: an allowlisted OAuth browser may append a fixed decision after CSRF validation and a body-bound proxy-to-inspector capability check. Basic remains read-only. The route cannot create prompts, run models, export datasets, activate adapters, publish, or perform arbitrary Jazz mutation. Review evidence affects dataset projection only; it is not deployment authority. ## Data flow @@ -78,7 +80,8 @@ The first interface is a local server and dense activity page. It reads projecti | Agent traces | What the runtime and model emitted | | Derived events | Versioned claims or proposals by a named agent/runtime | | Dispatcher receipts | Durable evidence that a configured external action was claimed, delivered, or failed | -| Projections/UI | Rebuildable convenience views only | +| Review decisions | Append-only human evidence over one immutable prompt and exact candidate pair | +| Projections/UI | Rebuildable convenience views; the UI's only mutation is the fixed Review decision append | ## Process shape @@ -88,4 +91,4 @@ Development may run producers, consumers, projectors, dispatchers, and the HTTP The supervisor is process glue, not a central coordinator. Polling loops never overlap themselves. Consumer processes receive work through Jazz queries/subscriptions rather than an in-memory producer-to-agent handoff queue. Shutdown closes watchers and subscriptions, drains in-flight work, settles required Jazz writes, records terminal runtime evidence, and closes the database context. -Manifest presence does not grant authority. A source declaration can read only through its connector contract; agent declarations remain tool-free and action-free; Telegram ingress cannot send; the inspector remains loopback-only. Network sources are disabled in the repository manifest by default. +Manifest presence does not grant authority. A source declaration can read only through its connector contract; agent declarations remain tool-free and action-free; Telegram ingress cannot send; the inspector remains loopback-only. Review decisions require the separate browser and loopback capabilities described above. Network sources are disabled in the repository manifest by default. diff --git a/spec/events.md b/spec/events.md index 1655a11..0ec6f92 100644 --- a/spec/events.md +++ b/spec/events.md @@ -66,6 +66,7 @@ Examples: - `stream.thought.source.email.observed` - `stream.thought.source.telegram.message` - `stream.thought.source.telegram.reaction` +- `stream.thought.source.review.prompt` ### Connector lifecycle @@ -97,6 +98,10 @@ An operational incident is a deterministic, content-dark projection over connect - `stream.thought.consumer.execution.status-unknown` - `stream.thought.consumer.execution.abandoned` +### Model adapter deployment evidence + +Adapter state is not mutated through domain events. One immutable startup deployment catalog selects an exact release for each declaration. Run, output, failure, judgment, and training evidence carries the public release identity, checkpoint-reference digest, catalog digest, and catalog generation. It never contains the resolved checkpoint value. Activation and retirement are coordinated stop/install/restart operations verified through process and loaded-digest receipts, not Jazz lifecycle events. + ### Repair coordination - `stream.thought.agent.repair.requested` @@ -121,3 +126,11 @@ Derived events are proposals or observations. A post candidate is never a publis - `stream.thought.judgment.training-example.retracted` Judgments are append-only. A changed decision names the prior judgment in `supersedesJudgmentEventId`; removal appends a retraction naming `retractedJudgmentEventId`. Rebuildable training and effective-output projections exclude superseded and retracted judgments without deleting their evidence. A `correct` replacement is appended only after canonical output-contract validation. + +### Human Review + +- `stream.thought.derived.review.response` +- `stream.thought.review.item.created` +- `stream.thought.review.decision` + +The source prompt holds complete bounded evidence. Candidate response events are ordinary contract-valid run outputs with a receipt proving one complete, unprojected, tool-free prompt context. A review item freezes two exact same-trigger completed runs, output ids, training-relevant run-receipt digests, and one deterministic blinded order without copying output. Queue, decision, and export projection revalidate those frozen receipts. A review decision records judgeability, preference, tie, correction, or skip and may supersede an earlier decision. Only the active externally eligible `prefer` or `correct` decision can project into a v4 training example. See `review.md`. diff --git a/spec/jazz.md b/spec/jazz.md index 027c366..b6ecd87 100644 --- a/spec/jazz.md +++ b/spec/jazz.md @@ -69,13 +69,14 @@ Historical rows are caller-id inserts. The deployed Jazz permission policy must - `sources`: producer declarations, source-local sequence state, and nonsecret configuration. - `sourceCursors`: the last durable external position for each producer. - `consumers`: compiled declarative consumer specifications. +- Adapter inventory is compiled from the immutable startup catalog and persisted on consumer/run evidence. Jazz has no mutable model-adapter registry or lifecycle authority. - `consumerProgress`: a managed consumer's durable position within each source namespace it has consumed. - `executions`: optional materialized attempt/status evidence for managed model runs; this is not a global work queue. - `inferenceBudgetAccounts`: per-scope rolling/fixed-window counters and active reservation leases. - `inferenceAccounting`: one privacy-dark reservation/settlement record per model attempt, containing only execution identity, model identity, timestamps, status, and numeric estimated/actual charges for policy-tracked dimensions. Calls and tokens are mandatory; untracked cost is omitted from estimate and charged JSON rather than serialized as zero. - `projections`: rebuildable named views and their durable progress. -Operational rows are mutable. Important changes also emit historical lifecycle events so the inspector can distinguish current state from evidence. Accounting JSON readers accept both existing cost-bearing rows and new rows with omitted cost. Settlement follows the estimate captured on each reservation, so an active legacy reservation remains cost-tracked even after a declaration adopts a token-only policy. +Operational rows are mutable. Adapter release identity and deployment catalog identity are content-addressed outside Jazz, then copied into consumer/run evidence for historical inspection. Activation and retirement require a new catalog generation plus coordinated restart; they are not database transitions. Accounting JSON readers accept both existing cost-bearing rows and new rows with omitted cost. Settlement follows the estimate captured on each reservation, so an active legacy reservation remains cost-tracked even after a declaration adopts a token-only policy. ## Event identity and source sequence @@ -91,6 +92,8 @@ Historical rows use caller-id `insert`, never `upsert`. A duplicate-id rejection Each producer also assigns a monotonically increasing `sourceSequence` within its source namespace. This is not a global database sequence and does not pretend independent sources have one total causal order. It gives consumers a durable, queryable boundary per source. The producer writes new event rows, its next source sequence, and its external cursor as one transaction before acknowledging source progress. +Producer append transactions are serialized per canonical source within one process. This prevents overlapping lifecycle or connector callbacks from reading the same source head and attempting conflicting source-sequence or caller-id writes. Independent sources remain concurrent. Cross-process writers must still have disjoint source ownership under the initial topology. + ## Consumer query and replay A consumer declaration compiles directly to a Jazz event query. Predicates may include event type, source, privacy class, address, and typed indexed payload fields. Jazz performs filtering, source-sequence ordering, bounds, and pagination. Consumers do not load the entire events table and recreate a query engine in TypeScript. @@ -113,9 +116,11 @@ The intended database atomic units are: 1. producer events plus that producer's external cursor and source sequence; 2. consumer output/lifecycle events plus terminal execution evidence and progress for the consumed source sequences; -3. one projection update plus its per-source progress. +3. one projection update plus its per-source progress; 4. one inference reservation plus its budget-account charge, and later one reservation settlement plus the corresponding charge adjustment. Budget-window arithmetic treats absent optional cost as zero internally, while persisted reservation and charge records preserve omission. +Adapter startup does not add another Jazz atomic unit. Each process loads one immutable release/deployment bundle before registration, refuses startup on identity or process-set mismatch, and records the resulting public catalog/release/binding identity on declarations and runs. Catalog change uses stop/install/restart, followed by PID/start-time, loaded-digest, and canary verification. Jazz is evidence for what ran, not consensus for what may run. + Model calls, sockets, and filesystem reads occur outside a transaction. The process records a started attempt, performs external work, then commits accepted outputs, terminal evidence, and consumer progress together. If the process dies during external work, the started attempt remains nonterminal; restart records it as status-unknown or abandoned according to policy and may begin a new attempt. No lease is required because one process owns the consumer identity. Use a Jazz transaction when the pinned runtime proves that all writes in one unit settle together at the required authority. Use a direct batch only when grouped visibility, rather than authority validation, is the actual requirement. These semantics require integration tests, but not compare-and-swap or queue machinery under the single-owner topology. diff --git a/spec/recovery.md b/spec/recovery.md index 754d218..5b67dc1 100644 --- a/spec/recovery.md +++ b/spec/recovery.md @@ -30,9 +30,18 @@ Live cutover uses `replay: now`: first stop resident consumption of raw ATProto, Budget denial still follows the existing terminal semantics: the blocked run and consumer progress settle together. It is therefore not a retry queue, and temporary denial can skip resident processing. This slice intentionally does not silently change that contract without a durable next-attempt identity and backoff design. The resident's very high emergency circuit breakers now make accidental denial implausible during ordinary or extreme activity; retryable denial remains a required follow-up for a catastrophic-runaway trip. Batching is retained for semantic coherence and reduced redundant turns, not as a cost-control measure. +## Model-adapter startup recovery + +- A persisted declaration or run proves what previously executed; it cannot recreate the process-local checkpoint binding. +- Restart reloads the complete immutable release catalog and one deployment catalog, verifies optional expected catalog digest/process identity, resolves checkpoints, and recompiles declarations before provider egress. +- Candidate and retired entries create no runtime binding; any enabled declaration still naming one fails compilation. Missing, ambiguous, digest-mismatched, unresolved active, unallowlisted, or unauthorized selections fail startup. Recovery never floats to a newer release by capability similarity. +- Adapter activation and retirement use stop/install/restart. Operators verify that no old PID/process-start identity remains, every restarted process reports the expected loaded catalog digest, and a bounded conceptualizer canary passes. +- Mixed catalog generations are unsupported. If a process did not restart or reports the wrong digest, the deployment is incomplete and adapter-backed work stays disabled. +- There is no authority file, recovery marker, registry lock, owner election, dispatch lease, or operator mutation API to reconstruct after a crash. + ## Consumer recovery -- Read the consumer declaration and per-source progress. +- Read the consumer declaration and per-source progress. A persisted run contains exact public-safe learned-adapter and catalog provenance, but a persisted or rehydrated declaration does not recreate the process-local checkpoint binding. Restart must recompile through the trusted startup loader; replay never silently floats to another release or binding. - Query matching event rows after each source's last terminally handled sequence. - Treat live subscription deltas as wakeups only. Re-query each installed source at the manifest's bounded reconciliation interval so writes committed by another local process cannot remain invisible merely because the persistent driver's cross-process callback did not fire. - Inspect deterministic execution/lifecycle evidence before invoking external work. @@ -82,6 +91,7 @@ The UI flags contradictions rather than choosing whichever row looks friendlier. - Transient transport/rate limit: retry with bounded exponential backoff and jitter. - Timeout/process loss: preserve the nonterminal attempt, classify it, and retry according to consumer policy. - Invalid output: fail without blind retry; the deterministic repair coordinator appends exactly one versioned request only when `repairs.md` eligibility and evidence checks pass. +- Unknown, missing, candidate, retired, ambiguous-capability, provider-mismatched, non-allowlisted, cloned, rehydrated, mutated-after-compilation, or unbound model adapter: reject declaration registration, run creation, and provider dispatch. Complete canonical compiler-bound identity is checked at every boundary. After read-only prefetch and local broker validation, invoke the opaque run-bound authority capability at the actual provider boundary. It revalidates the exact canonical active binding and holds the token-owned cross-process lock plus lease through the complete bounded provider request. Retirement cannot pass that bound; stale active processes fail before egress. A previously started run keeps its recorded adapter identity through recovery. - Authorization/configuration: block until configuration changes. - Letta Agent SDK `success: false`: settle the attempt's conservative accounting charge, classify the remote failure, and leave source progress unchanged. A later retry first reconciles the deterministic turn marker before sending. - Inference budget exhaustion: terminally block that source event before provider dispatch and continue from the resulting durable consumer progress. diff --git a/spec/repairs.md b/spec/repairs.md index e353806..93b1550 100644 --- a/spec/repairs.md +++ b/spec/repairs.md @@ -12,15 +12,16 @@ A run is eligible only when all of the following are true: - it is a terminal failed run from a declaration whose role is `standard`; - it has no output event and is not itself triggered by a repair request; +- it uses the currently supported `stream.thought.output.observation@1` repair contract; - the original trigger event, failed lifecycle event, exact declaration version, prompt hash, declaration fingerprint, canonical output-contract identity, and sanitized validation evidence are all present and mutually consistent; - its diagnostic is either invalid JSON, a canonical output-contract violation, or a semantic-validation failure whose stable rule id is explicitly allowlisted by the policy; - its validation diagnostic contains only bounded counts, hashes, stable reason/rule ids, and sanitized contract issues. -Timeouts, cancellation, provider outages, rate limits, sandbox/broker/process failures, credential/configuration/authentication failures, incomplete or corrupt evidence, repair-origin runs, and unallowlisted semantic failures are ineligible. The coordinator never reads malformed candidate text or quarantine data while deciding eligibility. +Timeouts, cancellation, provider outages, rate limits, sandbox/broker/process failures, credential/configuration/authentication failures, incomplete or corrupt evidence, repair-origin runs, non-observation output contracts, and unallowlisted semantic failures are ineligible. The coordinator never reads malformed candidate text or quarantine data while deciding eligibility. ## Canonical output contract -All current model-backed observation outputs use `stream.thought.output.observation@1`. The output-contract registry binds its id and version to one canonical definition, validator, prompt description, and SHA-256 hash. Original execution, repair execution, human `correct` replacement validation, effective-output rebuilding, and training export resolve that same identity. A hash mismatch is an evidence failure, not a request to use a nearby schema. +The current repair workflow supports only `stream.thought.output.observation@1`. The output-contract registry binds its id and version to one canonical definition, validator, prompt description, and SHA-256 hash. Original execution, repair execution, human `correct` replacement validation, effective-output rebuilding, and training export resolve that same identity. A hash mismatch is an evidence failure, not a request to use a nearby schema. Other registered contracts, including conceptualization graphs, fail terminally without creating an unusable repair request until a contract-matched repair declaration exists. Every run context manifest and accepted output event records the contract id, version, and hash. Invalid-output diagnostics record the same identity plus sanitized issue codes and paths. They never record rejected values or raw model text. @@ -32,7 +33,7 @@ The trusted parent constructs repair context from: - immutable references to the original run, failed lifecycle event, trigger event, and source root; - the exact original declaration metadata and repository prompt whose hashes match the failed run; -- the original provider and model identity, plus checkpoint and adapter revisions when observed; +- the original provider/model identity, execution-adapter revision, learned-adapter identity, and startup catalog digest/generation; - the original bounded source-event context regenerated from the immutable trigger event; - the canonical output-contract definition and identity; - bounded sanitized validation issues from the request. @@ -43,7 +44,7 @@ A repair request permits one model proposal generation. An ambiguous interrupted ## Proposal and authority -A valid repair appends `stream.thought.derived.output.correction.proposed@1`. The event links the original failed run, repair request, repair run, original trigger and source root, contract identity, and one validated structured output. It is inert by default. +A valid repair appends `stream.thought.derived.output.correction.proposed@1`. The event links the original failed run, repair request, repair run, original trigger and source root, contract identity, separate original/repair model-adapter/catalog provenance, and one validated structured output. Its privacy is the join of request state and the repair declaration's learned-adapter privacy. It is inert by default. Existing append-only judgments are the authority surface: @@ -64,8 +65,8 @@ The rebuildable projection is keyed by original run id and applies this order: 2. original valid output; 3. unresolved failure. -Projection rows contain only contract-validated structured output and provenance ids. They can be dropped and reconstructed from runs and append-only events. Rebuilding after judgment supersession or retraction must produce the same result as uninterrupted processing. +Projection rows contain only contract-validated structured output, joined privacy, provenance ids, and separate original/repair public model-adapter/catalog evidence. They can be dropped and reconstructed from runs and append-only events. Rebuilding after judgment supersession or retraction must produce the same result as uninterrupted processing. ## Training boundary -Repair trajectories enter external training export only when the repair proposal has an active `accept` or `correct` judgment with both quality and external-export eligibility. Default export also requires the original source, repair proposal, judgment, and chosen/rejected chain to be entirely `public-source`; sensitive/private repair material requires the separate declassification and destination gates. Rejected, unresolved, retracted, failed, and merely proposed repairs are excluded. Export contains validated chosen/rejected structured output, minimal source classification, allowlisted trajectory type/order, contract identity, and model/adapter provenance. It excludes malformed candidate text, source bodies, prompts, provider bodies, reasoning, tool arguments, image bytes, arbitrary diagnostics, actor/route/external/correlation/idempotency identifiers, internal provenance ids, source hashes, trace-content hashes, and quarantine material. +Repair trajectories enter external training export only when the repair proposal has an active `accept` or `correct` judgment with both quality and external-export eligibility. Default export also requires the joined original source/run, repair request/run/proposal, judgment, and chosen/rejected chain to be `public-source`; learned-adapter privacy/export policy is applied independently to every participating run. Sensitive/private repair material requires the separate declassification and destination gates. Rejected, unresolved, retracted, failed, and merely proposed repairs are excluded. Export contains validated chosen/rejected structured output, minimal source classification, allowlisted trajectory type/order, contract identity, and public model/adapter/catalog provenance. It excludes malformed candidate text, source bodies, prompts, provider bodies, reasoning, tool arguments, image bytes, arbitrary diagnostics, actor/route/external/correlation/idempotency identifiers, internal provenance ids, source hashes, trace-content hashes, private checkpoints, and quarantine material. diff --git a/spec/review.md b/spec/review.md new file mode 100644 index 0000000..92a13c0 --- /dev/null +++ b/spec/review.md @@ -0,0 +1,120 @@ +# Human review and training-data custody + +## Boundary + +Review is a private, append-only human judgment surface over immutable source prompts and exact completed candidate runs. It is the evidence and dataset-custody layer around training. It does not train models, mutate candidate output, infer missing prompt context, or grant a checkpoint deployment authority. + +Tinker or another explicitly configured trainer owns optimization. The immutable adapter catalog owns released artifacts and coordinated activation. Review owns: + +- complete prompt/evidence custody; +- blinded presentation of exact candidate runs; +- judgeability, pairwise preference, tie, correction, and skip decisions; +- supersession without destructive edits; +- privacy- and provenance-gated training export; +- coverage and exclusion evidence for a review campaign. + +## Review prompt + +`stream.thought.source.review.prompt@1` is a source event. Its payload contains: + +- a campaign id, version, and human label; +- the complete prompt shown to candidate models; +- optional bounded evidence required to judge the response; +- one criterion id/version, label, instructions, allowed reason codes, and allowed response tags; +- the candidate agent ids expected to run over the event; +- whether this already-reviewed synthetic/public prompt is eligible for external training export. + +The event privacy is authoritative. A prompt may set `externalExportEligible: true` only when its event is `public-source`. Private and sensitive prompts can be reviewed but are not declassified by the browser. + +The prompt event is the shared trigger for every candidate run. Pairwise candidates generated from different trigger events are not comparable training examples even when their visible text happens to match. + +## Candidate output + +Review candidates use canonical output contract `stream.thought.output.review-response@1`, containing one bounded `response` string. Candidate runs must be completed, must name the same review-prompt trigger, must come from distinct candidate agents declared by that prompt, and must each have exactly one contract-valid output event. The run and context receipts must prove the prompt was the only input, was included whole, was not projected through `payloadFields`, gained no tool evidence, and had no external action authority. + +The model, provider, execution adapter, learned adapter, catalog generation, prompt hash, and trace remain ordinary run provenance. Candidate output is never copied into the review-item event. + +## Review item + +`stream.thought.review.item.created@1` freezes one exact pair: + +- review prompt event id; +- the two candidate run ids, output event ids, and digests of every training-relevant run/context/provenance field; +- a deterministic blinded display order; +- the criterion identity inherited from the prompt. + +Creation validates all references and joins prompt, run, and output privacy. The item id is deterministic over the prompt, criterion, and unordered pair, so repeated materialization addresses the same event. A later retry or replacement run requires a new review item. + +Every later queue, decision, and export path resolves the frozen output event ids and recomputes the candidate receipt digests. A changed completed-run row, output pointer, context manifest, model, prompt hash, adapter identity, catalog identity, or privacy fails closed instead of silently rebinding a human label. Before the first decision, the Review projection withholds run ids and provenance. This is presentational blinding, not a claim that an operator with direct Jazz access cannot inspect underlying rows. + +## Review decision + +`stream.thought.review.decision@1` records exactly one disposition: + +- `prefer`: one candidate wins, with `slight` or `strong` strength; +- `tie`: both are adequately equivalent for the criterion; +- `correct`: neither is acceptable and a contract-valid replacement response is supplied; +- `underdetermined`: the evidence cannot identify the preferred response; +- `malformed`: the review item or rubric is defective; +- `skip`: no judgment is supplied. + +Every decision may carry bounded confidence, at most two allowlisted reason codes, at most two allowlisted response tags, and an optional bounded note. `prefer` must name one candidate run. `correct` must supply a replacement response. Other dispositions must not smuggle a preference or replacement. + +Decisions append. A later decision may explicitly name the prior active decision it supersedes. The active projection chooses the latest valid decision for an item and treats older concurrent leaves as inactive, so a process race cannot create multiple exported labels. Browser retries carry a random submission id and are idempotent. + +Judgeability is part of the data, not failed labor. `tie`, `underdetermined`, `malformed`, and `skip` never become DPO examples. They remain campaign-quality evidence and can prevent an invalid benchmark from quietly acquiring forced labels. + +## Training eligibility + +Review separates three facts: + +1. a human decision exists; +2. the decision is suitable quality evidence; +3. its full prompt/output/provenance chain is authorized for external training export. + +The browser's **use for training** control can make a `prefer` or `correct` decision quality-eligible. It can make the decision externally eligible only when the prompt was preauthorized for external export and the complete prompt, item, candidate-run, output, and decision privacy join is `public-source`. The browser has no private declassification control. + +The review exporter emits `thoughtstream.training-example.v4` only for active externally eligible `prefer` and `correct` decisions: + +- exact public review prompt and optional evidence; +- criterion identity and bounded human judgment metadata; +- contract-valid chosen response; +- one or two exact rejected candidate responses; +- empty trajectories: Review labels bind frozen prompt/response pairs, not trace projections that can gain later operational chunks; +- complete model, execution-adapter, learned-adapter, and catalog provenance for each candidate; +- campaign id/version. + +It omits source/run/event ids, actor/route/correlation/idempotency ids, notes, browser submission ids, cookies, CSRF tokens, capability signatures, private checkpoint references, and arbitrary source payload fields. Legacy v3 examples remain valid and can coexist in a v4 dataset manifest. + +## Browser authority + +The inspector remains read-only by default. Review writes are a separately configured capability: + +- only the allowlisted OAuth browser session can obtain write state and submit a review decision; +- Basic fallback remains read-only and never receives a CSRF token; +- the public proxy accepts POST only on the exact review-decision route, validates a bounded JSON body and the session CSRF token, then signs the exact method/path/body with a separately injected review capability; +- the loopback inspector verifies the signed request, freshness, and one-time nonce before reading the body or appending an event; +- no cookie, Authorization header, OAuth token, DID, CSRF token, or generic Jazz mutation reaches the inspector; +- absent or mismatched capability configuration leaves every data route GET/HEAD-only. + +The write endpoint accepts only the fixed decision schema. It cannot create prompts, create review items, run models, export datasets, activate adapters, publish, send, or mutate arbitrary events. + +## Interface + +The private **Review** tab shows unresolved items first. Each card presents complete prompt/evidence and two blinded candidates. The first controls are judgeability dispositions. When judgeable, the reviewer can record a directional strength, tie, or correction, then optional bounded tags and confidence. Model and adapter provenance is revealed only after submission or from the reviewed-history view. + +Changing a decision creates a superseding event. The UI shows current decision, previous-decision count, campaign coverage, and whether the active decision is evaluation-only or export-eligible. It does not show aggregate preferences before submission. + +## Initial training workflow + +1. Import fully specified public-safe review prompts split by scenario family. +2. Generate multiple real candidate runs from the frozen parent and candidate configurations. +3. Materialize exact blinded pairs after runs settle. +4. Run a small UI canary with training disabled. +5. Review judgeability and pairwise decisions. +6. Export only active, preauthorized public-source examples. +7. Train outside ThoughtStream. +8. Return candidate checkpoint and held-out evaluation receipts. +9. Promote only through the immutable adapter-release and coordinated-activation contract. + +Hard authorization, routing, privacy, and formatting contracts remain deterministic gates outside the preference objective. diff --git a/spec/security.md b/spec/security.md index 48bb42b..c9e1168 100644 --- a/spec/security.md +++ b/spec/security.md @@ -16,8 +16,8 @@ Action filtering happens at the egress boundary. Producers and consumers continu ## Credentials -- Credentials enter through environment variables, keyring commands, or injected runtime providers. Telegram bot and webhook secrets are referenced by environment-variable name in the manifest and never stored there. -- Jazz stores only credential reference names and configuration fingerprints. +- Credentials and private learned-checkpoint paths enter through environment variables, keyring commands, or injected runtime providers. Telegram bot and webhook secrets are referenced by environment-variable name in the manifest and never stored there. +- Jazz stores only credential reference names, public learned-adapter identity, catalog generation/digest, and SHA-256 of the resolved checkpoint binding. The checkpoint value remains in a process-local `WeakMap` owned by startup compilation and never enters Jazz, declarations, traces, events, inspector output, training data, or errors. - Logs, traces, lifecycle events, operational incidents, the private error ledger, repair requests, correction proposals, Telegram notifications, and training exports never contain credential values, raw prompts, provider bodies, provider thinking, malformed or raw model text, tool arguments, image bytes, source bodies, or quarantine content. Operational incidents, the private ledger, and incident alerts additionally exclude arbitrary error messages and stacks. Durable diagnostics use classifications, counts, hashes over normalized classifications, canonical contract identities, stable rule ids, and bounded issue codes/paths. Historical source-specific failure rows may contain error strings; the incident boundary never copies them. - Test processes explicitly disable ambient `.env` loading unless a live integration test is requested. @@ -27,7 +27,7 @@ Action filtering happens at the egress boundary. Producers and consumers continu - `sensitive`: email bodies, private chats, Obsidian content, attachments, health/financial/relationship material, or explicitly marked sources. - `public-source`: content already public at its source. Derivations may still reveal private interest or context and therefore remain private by default. -Agent declarations specify accepted privacy classes. A public-output candidate can be generated from public sources but is still only a private candidate. +Agent declarations specify accepted privacy classes. A public-output candidate can be generated from public sources but is still only a private candidate. One shared privacy join orders `public-source < private < sensitive`. Learned-adapter privacy participates in run rows, output/failure evidence, repair request/proposal, judgment/retraction, delivery, effective-output projection, and training. Each stage may raise privacy and may never lower it. ## Prompt injection @@ -37,9 +37,9 @@ The persistent Letta resident is an explicit operator-trusted capable agent. Tho ## Model cells and workspace harnesses -The `observer-v1` Pi inference cell currently requires an x86_64 Linux host with Bubblewrap, `prlimit`, and the glibc library layout bound by the launcher. It runs in a disposable Bubblewrap process under a different uid with a new network namespace, an empty environment, a minimal read-only runtime, a writable temporary directory, and explicit CPU, memory, file, descriptor, and wall-clock limits. The worker has no host tools and cannot read provider credentials. Read-only enrichment runs in the trusted parent before the worker starts. +The `observer-v1` Pi inference cell currently requires an x86_64 Linux host with Bubblewrap, `prlimit`, and the glibc library layout bound by the launcher. It runs in a disposable Bubblewrap process under a different uid with a new network namespace, an empty environment, a minimal read-only runtime, a writable temporary directory, and explicit CPU, memory, file, descriptor, and wall-clock limits. The worker has no host tools and cannot read provider credentials. Read-only enrichment runs in the trusted parent before the broker begins any adapter-backed provider request. -The observer worker can reach only a per-run Unix socket. Its single-use capability authorizes one request for one run, model, route, token ceiling, size budget, and deadline. The trusted broker validates the request, injects the provider credential, rejects redirects, bounds the response, and returns only allowlisted headers. Missing Bubblewrap, missing worker artifacts, broker failure, protocol failure, timeout, or resource exhaustion fails closed. There is no trusted-host inference fallback. Repair agents use this exact path; the coordinator cannot invoke a provider and repair declarations cannot weaken sandbox or broker policy. +The observer worker can reach only a per-run Unix socket. Its capability authorizes a bounded request set for one run, model, route, token ceiling, cumulative size budgets, and deadline. The trusted broker serializes admission and atomically reserves request count, cumulative request bytes, and response capacity after rechecking expiry; concurrent sockets cannot pass checks against stale counters. It converts each response reservation into actual usage in the same admission critical section. The broker injects the provider credential, rejects redirects, bounds the response, and returns only allowlisted headers. Missing Bubblewrap, missing worker artifacts, broker failure, protocol failure, timeout, or resource exhaustion fails closed. There is no trusted-host inference fallback. Repair agents use this exact path; the coordinator cannot invoke a provider and repair declarations cannot weaken sandbox or broker policy. The `workspace-v1` profile is separately defined in `harnesses.md`. It runs a disposable rootless-in-container process with a read-only root filesystem, no IP network, no inherited environment, all capabilities dropped, `no-new-privileges`, the runtime's default seccomp policy, cgroup-backed CPU/memory/process limits, bounded tmpfs, and exactly one workspace lease, state lease, and provider socket mount. The provider lease authorizes multiple turns only within explicit request-count and cumulative byte budgets. Container image identity and isolation profile are launch evidence. A passing observer-cell canary does not satisfy the workspace-harness gate. @@ -53,6 +53,30 @@ An unrestricted Cloud permission mode authorizes the Letta harness to use its av The resident's mixed Telegram/ATProto conversation makes the ThoughtStream-to-agent border load-bearing. Public source text and third-party Markdown are bounded, snapshotted, and marked as untrusted data; strong references remain distinguishable from mutable protocol or social renderings. The trusted parent calls only the source-appropriate fixed public services: Bluesky post/like context may use atproto.md plus bsky.md, while Semble collection-link context uses atproto.md for the link, card, and collection and never sends those records to bsky.md. ThoughtStream does not pass source credentials, Jazz credentials, deploy keys, Git credentials, host paths, or public-write authority into the packet. Fetched bodies and context snapshots live under private runtime storage and are forbidden from Git, build artifacts, traces, accounting, Telegram delivery, operational errors, and public projections. Prompt guidance reminds the resident not to expose private continuity, but the Cloud sandbox remains an operator-selected capable-agent environment after that border. +## Web authentication containment + +The public website, OAuth control routes, and private inspector forwarding share a process only for deployment convenience. They do not share data authority. The public router is a closed allowlist and cannot obtain a Jazz store, runtime manifest, source/event/trace reader, arbitrary filesystem path, environment dump, or upstream fallback. Only `/inspector` may reach the loopback inspector, and only after OAuth-session or explicitly enabled Basic fallback authentication. + +ATProto OAuth is implemented by the official Node client rather than a partial local protocol implementation. The SDK performs mandatory PKCE, PAR, DPoP, nonce handling, metadata discovery, identity resolution, token refresh, and request serialization. ThoughtStream additionally enforces an exact DID allowlist after callback and before creating a browser session. OAuth grants inspector read access only; the OAuth token is never used as a general PDS capability by this service. + +OAuth protocol state, application state, DPoP private keys, access tokens, refresh tokens, and browser-session records are encrypted at rest with AES-256-GCM under a separately injected 32-byte key. The encrypted store is outside Git and outside the ThoughtStream live runtime root, with owner-only directory/file modes and atomic replacement. Every store has explicit entry-count and serialized-byte limits. The ES256 confidential-client private JWK is separately injected. Neither key may appear in environment diagnostics, process output, tests, errors, events, traces, or HTTP responses. Browser cookies contain only random identifiers and use `HttpOnly`, `Secure`, `SameSite=Lax`, `Path=/`, bounded lifetime, and a `__Host-` name. + +Login passes random browser-bound application state into the SDK. The SDK generates a distinct OAuth protocol state and owns its one-time validation. Callback requires exactly one bounded protocol-state query value, then compares the SDK-returned application state with the unique cookie and consumes its application record. Failed/timed-out callback settlement, browser expiry, DID mismatch, restore failure, and explicit logout delete only matching local staged or promoted generations and make no application-initiated remote revocation request, because provider-wide semantics cannot be proven safe against a newer concurrent grant. The unmodified SDK may independently revoke after issuer, exchange, or session-store failure; that provider-side residual is not represented as a local authorization guarantee. Logout requires a server-stored CSRF token and `POST`. Generic failures reveal no account, token, state, upstream, or private-object detail. + +OAuth initiation and callback have independent bounded rate limits in nginx and the process. Nginx access logging records path without query and disables both callback access/error logging; the application never logs request URLs or SDK callback exceptions. Authorization discovery receives a request-disconnect abort signal because the SDK supports it. The installed SDK callback API does not support cancellation, so every callback uses attempt-scoped session staging plus one watchdog over the complete settlement path. Timeout synchronously expires both authority layers before advancing the serializer. Application-state consumption, generation promotion, and browser persistence pass authority guards that are rechecked after temporary write and immediately before rename. Cleanup is detached and generation-conditional, so hung promotion or cleanup cannot block a newer callback. Expired staging accepts late SDK session-store writes into quarantine, suppressing one known store-failure revocation path; late state is then deleted locally. It cannot suppress every SDK revocation path without transport interception or killable isolation. The underlying SDK promise may remain unresolved, so inert callback quarantines are capped at eight. Capacity refuses new callbacks with an operator-recycle-required `503` until attempts settle or the singleton process restarts. + +SDK restore is also generation-scoped: get, refresh set, and failure delete are bound through `AsyncLocalStorage` to the browser generation that initiated restore; stale completion is a no-op and a post-restore persisted-generation check runs before authority returns. + +The encrypted JSON stores are single-process stores. Startup acquires an owner-only lock in the fixed `~/.local/share/thoughtstream-inspector-auth` directory and refuses a live owner; systemd runs one non-templated proxy unit and grants write access only to that directory. Environment parsing rejects custom OAuth store paths rather than allowing a configuration the sandbox cannot write. Multiple replicas or manual parallel proxy processes may not share the directory. + +Basic Auth remains a separately configured break-glass path while OAuth is being activated. When enabled it authorizes only read-only inspector forwarding, bypasses OAuth restore, and is stripped before upstream access. When disabled it is ignored and not advertised. Startup refuses a configuration with neither OAuth nor Basic. Basic stays enabled until a real external HTTPS metadata fetch, redirect, callback, allowlisted-DID session, private inspector read, logout/local deletion, and a Basic rollback read are observed. Tests and localhost callbacks are not sufficient evidence to remove it. + +Review is the sole browser mutation exception. Basic remains read-only. An allowlisted OAuth session and CSRF token authorize the public proxy to sign one exact bounded decision body with a separately injected capability. The inspector verifies method, normalized path, body digest, freshness, and one-time nonce before applying the fixed append-only schema. Cookies, OAuth tokens, DIDs, CSRF values, and Basic credentials never reach Jazz or the loopback inspector. Missing capability configuration leaves the inspector entirely read-only. See [`review.md`](review.md) and [`web-auth.md`](web-auth.md). + +## Dependency audit residual + +The July 26, 2026 production audit has one unresolved high-severity finding: `sharp@0.34.5` is inherited through `@letta-ai/letta-agent-sdk -> @letta-ai/letta-code`, while the advisory requires `sharp>=0.35.0`. Review prompt generation, browser grading, capability verification, and JSONL export do not invoke image decoding, so this is not exposed by the Review path. It is still a project-level dependency finding and remains visible until the upstream SDK adopts a compatible patched Sharp version or a separately tested major-minor override is approved. Direct `fast-xml-parser` and compatible transitive `protobufjs`/`brace-expansion` findings are patched; a clean audit must not be claimed while Sharp remains. + ## Filesystem containment - Resolve and verify real paths beneath configured roots. @@ -63,7 +87,7 @@ The resident's mixed Telegram/ATProto conversation makes the ThoughtStream-to-ag ## Audit -Configuration changes, agent activation, connector activation, model tier changes, and any future action capability changes append audit events with actor and revision. +Configuration changes, agent activation, connector activation, model tier changes, adapter catalog installation, and any future action capability changes require an operator-visible deployment receipt. An adapter manifest cannot grant a provider endpoint, credential value, host path, or action capability. It may name one non-secret environment reference and selects only an already trusted provider profile. The public base model and private runtime checkpoint are checked independently against the trusted provider allowlist. Startup compilation is the only path that creates the process-local checkpoint binding; compiled identities are deeply frozen, and clone/rehydration/forgery cannot recover that binding. Adapter activation and retirement require stop/install/restart plus PID/start-time, loaded-digest, and canary receipts. There is no runtime lifecycle mutation API. Every Telegram attempt appends a durable `started` claim before calling the Bot API and then appends `delivered` or `failed` evidence. Claims use deterministic identities so concurrent dispatcher processes cannot intentionally claim the same batch twice. A claimed attempt is never inferred as delivered merely because the process exited cleanly. Failure notifications contain a short receipt and allowlisted diagnostic rendering; old traces containing raw content remain unread by the dispatcher. Operational incident alerts use a separate category policy rather than normal source allowlists, and Telegram delivery failures are never recursively alerted through Telegram. Telegram delivery or reaction can supply judgment evidence for a correction proposal, but delivery itself never changes effective output. diff --git a/spec/testing.md b/spec/testing.md index c25cd9a..013f972 100644 --- a/spec/testing.md +++ b/spec/testing.md @@ -19,6 +19,13 @@ - Private incident-ledger path containment, mode `0600`, bounded canonical JSONL, fsync-before-acknowledgement, and restart deduplication by deterministic incident id. - Incident Telegram alerts ignore normal source allowlists, deduplicate by fingerprint and cooldown, obey a separate destination window, render classifications only, and never recurse on Telegram delivery failure. - Cursor advancement only after durability. +- Public web routing is a closed allowlist: every named public page renders from reviewed public files, traversal and unknown routes return local `404`, and no public request reaches the inspector upstream. +- Inspector forwarding is available only below `/inspector`, strips that prefix, permits only `GET`/`HEAD`, removes authorization/cookie/forwarding headers, and requires either a valid server-side OAuth browser session or explicitly enabled Basic fallback. +- OAuth metadata and JWKS are public and exact; metadata declares HTTPS `client_id`/callback, `authorization_code` plus `refresh_token`, `atproto` scope, DPoP-bound tokens, `private_key_jwt`, and ES256. No private JWK is present in either response. +- OAuth tests model SDK protocol state and browser-bound application state as distinct values and exercise the installed official SDK's real `authorize` path to prove its generated protocol-state store key/PAR value differs from stored `appState`. They prove the query protocol state reaches SDK callback while the SDK-returned application state matches the cookie, plus protocol replay, missing/duplicate/mismatched/expired application state, exact DID allowlisting, opaque and duplicate cookies, browser expiry, restore failure, flow-store failure after SDK success, generation-conditional local deletion on CSRF-checked logout, and generic failures. Watchdog tests cover never-resolving callback, hung promotion, hung non-timeout cleanup, and late callback settlement followed by a successful newer generation. They prove the serializer advances, authority guards reject late persistent writes, and exact-generation cleanup preserves newer browser authority. Deferred restore-success and restore-failure races promote generation 2 while generation 1 is in flight, then prove scoped refresh set/delete cannot mutate generation 2 and post-restore recheck denies the old browser. Quarantine tests accumulate to the hard ceiling, verify fail-closed operator-recycle refusal, then prove recovery after settlement and fresh-process construction. Tests make no claim about provider-side revocation from the unmodified SDK; they never contact a PDS or load live credentials. +- Encrypted OAuth stores use direct malformed, unsupported-version, wrong-key, and oversized-envelope fixtures; retain no plaintext sentinel; enforce owner-only modes; expire state records; reject entry-count and serialized-byte exhaustion atomically; and replace files atomically. Direct commit-point tests prove no fallible chmod or other filesystem operation runs after rename and force authority loss after temporary write to prove the pre-rename guard removes the temporary file without committing, preventing disk/memory rollback or late-authority divergence. A live single-process owner lock prevents a second factory from opening the same store directory and can be reacquired after graceful release. +- Process tests prove OAuth login/callback rate-limit rejection occurs before SDK invocation, browser disconnect aborts the SDK-supported authorization path, and custom OAuth store paths fail before service startup because they fall outside the systemd writable path. Deployment tests prove nginx route limits, no-query access logging, callback access-log suppression, canonical `www` redirect before proxying, singleton systemd shape, and explicit Basic fallback generation. +- A production OAuth acceptance gate separately proves a real external metadata/JWKS fetch, PKCE/PAR/DPoP authorization redirect, allowlisted callback, private inspector read, logout/local generation deletion, and one Basic rollback read. Basic fallback is not removed before this gate passes. ## Jazz integration tests @@ -62,6 +69,16 @@ These are capability gates, not aspirational checks. An API named `transaction`, - Privacy tests seed a unique malformed-output sentinel in process-local provider output and prove it is absent from events, runs, traces, projections, Telegram candidates/receipts, and training export. - Training export includes accepted/corrected repairs only and emits only allowlisted trajectory/provenance fields. - Tinker provider configuration is tested without a real credential by inspecting the built model descriptor and request shape. +- Adapter manifests reject malformed ids/versions/digests, unknown fields, duplicate release identities, duplicate capabilities/evals, unknown provider profiles, missing or non-allowlisted public base/checkpoint resolution, ambiguous declaration selection, and adapter/model/tier combinations. Candidate/retired deployment entries remain unbound, and enabled declarations that name them fail compilation. +- A trusted environment checkpoint resolves only during startup compilation and provider dispatch; declaration JSON, fingerprints, stored consumers, runs, contexts, outputs, events, traces, inspector responses, training examples, test failures, and generated manifests contain no private checkpoint value. The private binding lives in a module-local `WeakMap`. Returned release/binding objects are deeply frozen; mutation fails, and clones, rehydration, and manual forgery cannot recover the checkpoint. +- Startup matrix tests cover missing/empty release catalogs, missing deployment catalog, malformed and duplicate releases, candidate/retired selection, missing or digest-mismatched release, duplicate declaration selection, unresolved/non-allowlisted checkpoint, expected catalog mismatch, and unauthorized process identity. +- Release digest remains stable across deployment generations. Catalog digest changes with deployment selection, verifies its optional expected digest without self-reference, and appears with generation on declaration, run, output/failure evidence, inspector inventory, and training provenance. Declarations without learned adapters remain compatible. +- No dynamic lifecycle machinery remains: no Jazz adapter registry, lifecycle events, authority file, recovery marker, owner election, registry lock/tombstone/reaper, dispatch lease, or operator recovery API. Activation/retirement documentation requires coordinated stop/install/restart plus PID/start-time, loaded-digest, and canary receipts. +- Sensitive-adapter tests cover run evidence, success/failure output, repair request/proposal, judgment/retraction, delivery, effective-output projection, and training. Each path uses the shared privacy join. +- Pairwise training tests gate privacy and export class on both primary and compared runs, reject forbidden pairs, require explicit restricted/private authority, and assert exact compared execution/model-adapter provenance in v3 examples and dataset manifests. +- Review tests freeze exact same-trigger candidate runs; `underdetermined`, `tie`, `malformed`, and `skip` decisions never become preference examples; correction requires a contract-valid replacement; supersession leaves one active export label; and public prompt authority gates v4 input text. Private browser decisions cannot declassify data. +- Review web tests prove Basic remains read-only; OAuth POST requires CSRF and a fresh one-time proxy signature; replay, stale signatures, unknown routes, oversized bodies, malformed decisions, and direct loopback POSTs fail without appending events. +- Replay uses the exact public run binding but recompiles the private checkpoint through the trusted startup loader; it does not float to another release or trust serialized checkpoint authority. - Live Tinker sampling is an opt-in credentialed test and never runs in ordinary CI. - Letta Agent SDK declarations resolve the agent id from the named environment variable, reject non-Cloud backends and credential/base-URL selection, require `single-event` context and one or more bounded concrete sources, reject wildcard source patterns, and reject shared enabled agent ids. - Two ready source namespaces under one enabled Letta declaration execute through one shared agent scheduler key: the test runner observes maximum concurrency one while both per-source progress rows settle independently. @@ -89,7 +106,7 @@ The `workspace-v1` profile is not activatable from a consumer declaration until - A workspace mutation persists only in the workspace lease. A state mutation persists only in the state lease. Fresh-container scratch files do not survive a second run. - CPU, memory, process-count, descriptor, output, wall-clock, total lease-byte, and inode limits terminate hostile fixtures with classified failures and no trusted-host fallback. Per-file limits alone do not satisfy the disk-exhaustion gate. - The launcher rejects missing runtimes, mutable/unresolved image references, symlink or out-of-root leases, wrong file types, duplicate mounts, oversized packets/results, trailing frames, run-id mismatch, nonzero exits, and timeout. -- Broker tests cover capability forgery, run-id mismatch, model and route substitution, unauthorized headers, redirect rejection, response overflow, request-count exhaustion, cumulative request/response-byte exhaustion, expiry, replay, and concurrent use. Counters are consumed before upstream dispatch so concurrent calls cannot overspend the lease. +- Broker tests cover capability forgery, run-id mismatch, model and route substitution, unauthorized headers, redirect rejection, response overflow, request-count exhaustion, cumulative request/response-byte exhaustion, expiry, replay, concurrent use, and opaque learned-adapter authority. Concurrent socket tests with `maxRequests: 1` and one-request byte budgets prove serialized immediate-pre-egress reservation admits exactly one upstream request. Adapter tests prove no authority is acquired during prefetch, clock/state changes are observed at immediate pre-egress validation, upstream execution occurs inside the bound, the bound covers response consumption, live operations survive old lock age/process suspension, and retirement cannot cross the held provider request. - Pi coding-agent runs with extension/skill/template discovery disabled, only the profile's built-in tool allowlist, in-memory credentials/settings, and the broker placeholder key. A malicious `.pi/extensions` fixture is not executed. - A deterministic multi-turn provider fixture makes Pi call a workspace tool, observes the tool receipt, returns a final artifact, and proves that every provider turn crossed the broker. A second run resumes the selected session from `/state`; adapter/profile/workspace/model mismatch fails closed. - Sentinels placed in provider credentials, host environment, host-only files, tool arguments, shell output, provider bodies, and model reasoning are absent from durable run/trace/accounting/notification surfaces. diff --git a/spec/tinker.md b/spec/tinker.md index 6c4b7f1..48e59e9 100644 --- a/spec/tinker.md +++ b/spec/tinker.md @@ -1,43 +1,122 @@ -# Tinker integration +# Tinker model adapters ## Role -Tinker is the model adaptation and checkpoint layer. Pi is the agent loop. thought stream is the durable coordination and evidence layer. +Tinker is the model adaptation and checkpoint layer. Pi and the Letta Agent SDK are execution harnesses. ThoughtStream owns durable coordination and evidence. -The runner must not couple event semantics to one Tinker checkpoint, base model, or training recipe. +Human preference campaigns, blinded review, and training-data custody are owned by [`review.md`](review.md). Tinker consumes an explicitly exported dataset and returns training/checkpoint receipts. It does not own browser judgment state, private review authority, or declassification. A Tinker checkpoint cannot become active merely because its training run completed; release and activation remain governed by the immutable adapter contract in this document. -## Runtime configuration +A learned adapter changes model behavior. It does not change a consumer's subscription, output contract, tools, external authority, or event meaning. -The trusted `tinker-default` provider profile owns the OpenAI-compatible endpoint, `TINKER_API_KEY` reference, route, allowed models, request and response budgets, and provider timeout. Agent declarations may select that profile and a model or capability tier. They cannot replace its endpoint, route, or credential reference. The trusted host may extend the model allowlist with `THOUGHTSTREAM_TINKER_ALLOWED_MODELS`; tier mappings are also allowlisted automatically. +## Identity model -The exact model id sent upstream and a validated checkpoint revision returned by Tinker are written to the run receipt. Secrets are not. Provider JSON/schema controls are advisory unless a concrete endpoint/checkpoint is verified to enforce them; the runtime always applies its own strict final-text parser and schema validator before persistence or event emission. +ThoughtStream keeps three identities separate: -## Model ladder +- `executionAdapterRevision` identifies the code path that executed the model. +- `modelAdapter` identifies one immutable learned release and its process-local deployment binding. +- `adapterCatalogDigest` plus `adapterCatalogGeneration` identifies the exact startup deployment selection. -Agent declarations choose capability tiers rather than hard-coding one global model: +Changing deployment state never changes a release digest. Historical runs keep all three public identities. Runs without a learned adapter omit the learned-adapter and catalog fields. -- `triage-small`: cheap classification, extraction, and routing. -- `reasoning-small`: bounded synthesis and document structure work. -- `escalation`: larger or better-adapted model for sparse difficult cases. +## Immutable release catalog -The tier resolver maps `triage-small` to `Qwen/Qwen3.5-4B` by default. Other Tinker tiers require an explicit trusted-host mapping through `THOUGHTSTREAM_TINKER__MODEL`. A declaration may also name a concrete model, but the provider profile must allowlist it before execution. Historical runs retain the resolved model and observed checkpoint revision. +Release manifests live under `adapters/releases/*.yaml`. Each strict version-1 manifest contains: -## Adapter evolution +- stable lowercase `id` and positive integer `version`; +- human-readable `description` and ISO `releasedAt`; +- one trusted `providerProfile` and public base-model id; +- a checkpoint selector containing one environment-variable name; +- dataset and eval receipt ids plus SHA-256 digests; +- bounded capabilities; +- `privacyClass`: `public`, `private`, or `sensitive`; +- `exportClass`: `public`, `restricted`, or `forbidden`. -Training examples come from explicit judgments over run outputs, not from silently treating every accepted-looking output as correct. A judgment event records: +Unknown fields, malformed ids, duplicate release identities, and noncanonical digests fail startup. The canonical manifest digest covers the complete manifest. It does not contain deployment state. -- Run and output ids. -- Criterion/version. -- Preference, correction, or rejection. -- Optional replacement output. -- Independent quality and external-export eligibility. +The checkpoint environment value is resolved only while loading the startup catalog. The public base model and private checkpoint must both pass the trusted provider profile's allowlist. The resolved value is stored only in a module-private `WeakMap` keyed by the original deeply frozen binding object. JSON cloning, structured cloning, manual construction, or rehydration cannot recreate that binding. -An adapter release records its dataset manifest, base model, training configuration, eval set, and checkpoint path. Activating a new adapter updates the tier mapping; it never rewrites prior run metadata. +The public identity contains the checkpoint selector name and SHA-256 binding digest. It never contains the checkpoint value, provider credential, raw dataset, training job, or arbitrary filesystem path. -The implemented external projection writes JSONL only from active v2 judgments with both `qualityEligible` and `externalExportEligible` set. Legacy v1 `exportEligible` records remain compatible only for entirely `public-source` chains. Default export independently requires public source, output, judgment, compared-output, and repair-source privacy. Corrections become chosen/rejected pairs; preferences join two comparable completed runs; accepted outputs become chosen-only examples; rejected outputs become rejected-only examples. Repair runs are narrower: only active eligible accepts and corrections export, while rejected, unresolved, retracted, failed, and merely proposed repairs are excluded. +## Immutable deployment catalog -External examples use a v2 privacy-minimized envelope: no source payload, actor, route, external/correlation/idempotency identifiers, internal event/run/delivery ids, source hashes, trace-content hashes, trace timestamps, or arbitrary context fields. Sensitive/private examples require an explicitly authorized v2 judgment plus the separate CLI private-export acknowledgement, a non-Git/non-public destination, and private `0600` atomic files. This projection is input material for a later Tinker training job, not proof that training has run. +One `adapters/deployment.yaml` selects releases for exact declaration id/version pairs. It contains: + +- schema version and monotonically increasing generation; +- exact release id/version/manifest digest per declaration; +- deployment state: `candidate`, `active`, or `retired`; +- optional expected process identities; +- optional expected catalog digest. + +Only `active` selections create runtime bindings. Candidate and retired entries remain inspectable catalog state but produce no binding; an enabled declaration that names either fails compilation. Missing releases, manifest mismatches, duplicate declaration selections, unresolved active checkpoints, unallowlisted models, unauthorized process identities, and catalog-digest mismatches fail before any provider egress. + +The catalog digest covers the complete immutable release inventory and deployment selection except its self-referential expected-digest field. Adding the expected digest therefore verifies a previously computed bundle rather than changing it. + +A Pi declaration selects one exact adapter: + +```yaml +runner: + kind: pi + profile: tinker-default + adapter: + id: julia-instruction + version: 1 +``` + +An adapter selection cannot be combined with a direct model or capability tier. The selected release supplies the public base model and trusted provider profile; only the private checkpoint reaches the provider request. + +## Startup and runtime contract + +At process startup: + +1. Load and validate every release manifest. +2. Load exactly one deployment catalog. +3. Verify its optional expected digest and process set. +4. Resolve active checkpoints from environment references. +5. Compile each selected declaration against one exact release. +6. Deep-freeze declaration and adapter identity. +7. Persist only public release, binding digest, deployment generation, and catalog digest. + +Adapter-backed declarations fail if they were cloned or constructed outside that loader because provider dispatch cannot recover the private binding. + +Read-only evidence prefetch completes before the broker begins provider egress. Broker admission is serialized and atomically reserves request count, cumulative request bytes, and response capacity. Response reservation becomes actual usage in the same admission critical section, so delayed trace handling cannot double-count capacity. Concurrent sockets cannot oversubscribe any configured bound. + +Durable traces use the public base model and public adapter identity. They never contain the private checkpoint, prompt text, provider body, reasoning, final raw model text, tool arguments, source bodies, or credentials. + +## Privacy and export + +One privacy join orders `public-source < private < sensitive`. Adapter privacy joins with trigger, output, feedback, delivery, repair, and comparison privacy. No downstream path may lower it. + +Run rows, output/failure evidence, correction/effective-output projections, judgments, inspector views, and training examples preserve the public adapter and catalog provenance. Private checkpoint values remain absent. + +Legacy judgment-derived examples remain `thoughtstream.training-example.v3`; Review-derived preference and correction examples use v4 so the exact preauthorized public prompt/evidence can accompany candidate output. A forbidden adapter on either side of a pair, or on the original run behind a repair, excludes it. A restricted adapter on any participating run requires explicit restricted-adapter projection authority. Private or sensitive state on any participating run requires the ordinary private projection and private-write authorities. Pairwise examples contain separate exact primary and compared provenance; repair examples contain separate repair and original provenance; dataset manifests label every participating model. + +## Activation and retirement + +ThoughtStream does not hot-mutate adapter lifecycle state. It has no Jazz lifecycle authority, authority files, recovery markers, owner election, registry locks, tombstones, reapers, dispatch leases, or operator recovery API. + +Activation and retirement are coordinated deployments: + +1. Stage and validate a new release/deployment catalog bundle. +2. Stop every process capable of adapter-backed provider egress. +3. Atomically install the bundle. +4. Restart affected services. +5. Verify each PID/process-start identity and loaded catalog digest. +6. Run a bounded conceptualizer canary. + +Retirement changes deployment state in a new catalog generation and uses the same stop/install/restart procedure. Historical run identity remains immutable. + +This contract deliberately excludes hot activation, hot retirement during provider work, rolling mixed-generation deployment, independent workers loading different catalogs, and multi-host consensus. If those become real requirements, they need a dedicated transactional authority with fencing tokens. A filesystem plus eventually visible Jazz rows is not that authority. + +## Example release + +`adapters/releases/julia-adapter-v1.example.yaml` is metadata only. It refers to synthetic dataset/eval receipts and an environment key. It contains no checkpoint value and does not claim a live Tinker training or sampling run. `adapters/deployment.example.yaml` demonstrates a candidate selection and is not loaded automatically. + +## Comind boundary + +The canonical conceptualizer lives inside ThoughtStream and may select an adapter through the same startup catalog. It retains its strict versioned graph contract, atomic private graph event, exact source/root/run lineage, storage privacy floor, correction/effective-output/training compatibility, repair exclusion, and lack of PDS authority. + +A future standalone Comind process may validate the same release/conformance artifacts but must authorize and load its own deployment catalog. ThoughtStream catalog selection is not Comind deployment authority, and a Comind PDS receipt is not ThoughtStream lifecycle evidence. Shared artifacts contain schema, canonicalization, identity types, and fixtures only. They do not import Jazz, Pi runtime, ATProto writers, repository paths, or application policy. ## Initial limitation -Tinker OpenAI-compatible sampling is treated as a beta/testing runtime rather than a highly available production service. Tinker consumers therefore use bounded requests and retryable operational failures. Deterministic consumers cover degraded workflows when configured separately; a failed Pi run never changes execution mode or falls back to trusted-host inference. +Tinker OpenAI-compatible sampling remains a beta/testing runtime. Adapter consumers use bounded requests and explicit failure evidence. There is no trusted-host inference fallback and no live activation implied by checked-in examples or passing synthetic tests. diff --git a/spec/ui.md b/spec/ui.md index bca39c0..bc76547 100644 --- a/spec/ui.md +++ b/spec/ui.md @@ -1,4 +1,4 @@ -# Root activity interface +# Activity and Review interface ## Purpose @@ -17,23 +17,25 @@ Each activity row shows: - Derived output count. - Failure or blocked evidence. -The root view has one row per originating observation. Consumer lifecycle and derived events are grouped beneath that root rather than rendered as peer activity rows. The row must summarize the consumer's semantic result in ordinary language; lifecycle NSIDs and record identifiers are execution details. +The root view has one row per originating `stream.thought.source.*` observation. Consumer lifecycle and derived events are grouped beneath that root rather than rendered as peer activity rows. Self-rooted connector, runtime, dispatcher, action, and other operational receipts remain available through source health and detail views but never occupy root-activity rows. The row must summarize the consumer's semantic result in ordinary language; lifecycle NSIDs and record identifiers are execution details. Deterministic transforms are labeled **rules**, not agents. The interface states explicitly whether LLM inference occurred. It must not imply model reasoning when a path-based or other hard-coded transform ran. Filters: source, event family, agent, run status, privacy, time range, and text. +The inspector exposes a read-only adapter inventory with public-safe release metadata, canonical lifecycle status/generation, a separately rendered active deployment binding when one exists, selecting consumers, bound runs, and output event ids. It must distinguish the execution adapter from the learned model adapter, and release status from deployment binding, and must never render checkpoint paths or resolved environment values. + The source-health view shows each connector's durable cursor, last success/failure, current error, poll receipt counts, recovery count, and whether a poll has started without terminal evidence. Health is derived from durable cursor and connector events rather than process liveness. ## Detail view -- Consumer lifecycle and derived-output details lead with a plain-language **what processed this** block: what ran, whether it was a rule or model, whether LLM inference occurred, the actual result, recommendation or proposal, derived-record count, and whether external actions were enabled. A reader must not have to traverse raw lifecycle payloads to discover the semantic result. +- Consumer lifecycle and derived-output details lead with a plain-language **what processed this** block: the declaration display name, actual trigger source, whether it was a rule or model, whether LLM inference occurred, the exact produced summary, any recommendation or proposal that actually exists, derived-record count, and whether external actions were enabled. Technical agent ids and context strategies belong in execution details. Runtime metadata must not be appended to model-authored prose. A reader must not have to traverse raw lifecycle payloads to discover either the semantic result or its provenance. - Canonical envelope and payload. - Source strong reference. - Document version/diff when applicable. - Causal tree from root to derived outputs. - Exact context manifest with truncation markers. -- Model, provider, checkpoint, prompt revision, usage, duration, and attempts. +- Model, provider, observed checkpoint revision, execution-adapter revision, learned model-adapter identity, prompt revision, usage, duration, and attempts. - Trace event summary and raw redacted trace download. - Projection contributions. @@ -41,5 +43,65 @@ The source-health view shows each connector's durable cursor, last success/failu - Localhost by default. - No client-side secrets. -- No send/publish/edit controls in the first release. +- No send, publish, arbitrary edit, adapter activation, or generic Jazz controls. The separately specified Review decision form is the only data mutation. - Streaming UI may use server-sent events; persistence never depends on the browser being open. + +## Authenticated external access + +The inspector itself remains loopback-only and has no authentication or public +network authority. External access is an operator deployment made from three +separate layers: + +1. the loopback inspector reads the private Jazz store; +2. a separate loopback proxy authenticates every request before forwarding only + `GET` and `HEAD` to the inspector; +3. an operator-owned HTTPS reverse proxy terminates TLS and forwards to the + authenticated loopback proxy. + +The authenticated proxy uses a high-entropy HTTP Basic credential supplied by +an owner-only service-specific environment file outside the runtime data root. +It compares a fixed-length digest in constant +time, strips authorization and cookie headers before forwarding, does not log +requests or credentials, returns generic upstream failures, and applies +`no-store`, anti-framing, no-referrer, and content-type security headers to +every response. It may bind only to loopback. Basic authentication is not safe +over plaintext HTTP, so the public route must redirect to HTTPS before it is +considered deployed. + +Authentication gates every `/inspector` route and every inspector `/api/` route. There is no public +health, event-count, source-name, runtime metadata, manifest, trace, or error endpoint. A failed login +must not touch the inspector or disclose whether a requested private object +exists. + +## Public site and route separation + +The same loopback web process may serve an allowlisted public surface, but public and private routing are separate capabilities rather than a shared fallback: + +- `/`, `/docs`, and named `/docs/*` pages render only from an explicit repository-owned public-content allowlist. Route input can never select a filesystem path. +- `/oauth/client-metadata.json`, `/oauth/jwks.json`, `/oauth/login`, `/oauth/callback`, and `/oauth/logout` are the only public authentication routes. +- `/inspector` and `/inspector/*` are the only routes that may forward to the loopback inspector. The prefix is removed before forwarding. +- Unknown routes return a local content-dark `404`; they never fall through to the inspector. + +Public rendering has no Jazz handle, runtime root, manifest reader, event query, trace reader, environment dump, directory listing, or generic file-serving primitive. Public pages are built from reviewed Markdown files under `public/` with a fixed route-to-file map and conservative escaping. They contain architecture and operating concepts only, never source names, counts, event payloads, host paths, credentials, live service metadata, or private project state. + +## ATProto OAuth inspector authentication + +ATProto OAuth is a browser-to-web-service authentication flow distinct from authorization to read the inspector. The official Node OAuth client owns protocol requirements including PKCE S256, PAR, DPoP, nonce handling, token refresh, identity resolution, and authorization-server discovery. The web boundary adds these constraints: + +- Public client metadata is served at its exact HTTPS `client_id` URL with no redirect. JWKS is public; the matching private ES256 key remains outside Git in owner-only service configuration. +- OAuth state, DPoP keys, access tokens, refresh tokens, and browser sessions remain server-side in an encrypted owner-only store outside the ThoughtStream runtime root. Browser cookies contain only random opaque session ids. +- Random application state is short-lived, one-time, stored separately, and bound to an `HttpOnly`, `Secure`, `SameSite=Lax` callback cookie. It is passed into `authorize`; the SDK generates and validates a distinct protocol-state query value. After SDK callback succeeds, the returned application state must match and consume the browser-bound record. +- Successful callbacks are accepted only for the configured allowlisted DID. A different DID is deleted from local staging without promotion or an application-initiated remote revocation and receives no inspector session; independent SDK revocation remains a provider-side residual. +- Inspector sessions are short-lived, carry a per-DID promotion generation, restore only against that exact stored generation, and carry a separate CSRF token for logout. Restore get/set/delete are scoped to the initiating generation and rechecked afterward. Expiry, restore failure, and CSRF-checked logout delete only the matching local generation and expire the cookie; application code sends no provider revocation that could invalidate newer authority. +- SDK state, SDK session, application-flow, and browser-session stores have hard entry and serialized-byte bounds and one process owner. The store path is fixed to the one owner-only directory permitted by systemd; custom paths fail closed. Login and callback have independent nginx and in-process rate limits. +- SDK callback has no abort API, so callback credentials settle first in an attempt-scoped staging store. One watchdog covers callback, application-state consumption, promotion, browser persistence, and cleanup. Timeout marks the attempt non-promotable before advancing the serializer; every authority-bearing persistent write rechecks that status immediately before rename. Late expired staging is deleted locally without application-initiated remote revocation, protecting local newer generations. The unmodified SDK may still revoke on other failure paths. Eight retained attempts trigger a fail-closed operator-recycle `503` until settlement or restart. +- OAuth failures are content-dark. Tokens, DIDs other than the configured allowlist, handles, state values, cookies, provider bodies, and exception detail are not logged or returned. +- HTTP Basic remains an independently configured emergency fallback until an operator records one real successful OAuth login and deliberately changes fallback configuration. Enabled Basic authorizes inspector reads only and bypasses OAuth restore; disabled Basic is ignored and not advertised. Code presence or fixture tests do not count as that receipt. + +OAuth supplies identity and, only when the separate Review capability is configured, access to the one fixed append-only Review-decision route. It grants no generic ThoughtStream write, publish, model, connector, dispatcher, Jazz, filesystem, adapter-activation, or public-post authority. + +## Review view + +The private Review tab is a dense judgment workbench, not a trainer dashboard. It shows complete evidence before controls, uses blinded A/B presentation, asks judgeability before preference, supports tie and correction, and reveals provenance only after a decision. Mobile stacks evidence and candidates; desktop may use two columns. Campaign progress is descriptive and carries no gamified streak or speed target. + +See [`review.md`](review.md) for event, authority, supersession, privacy, and export semantics. diff --git a/spec/vision.md b/spec/vision.md index 80c72ee..5f3c3b1 100644 --- a/spec/vision.md +++ b/spec/vision.md @@ -23,6 +23,7 @@ The system should answer: - Cheap triage agents can label, cluster, extract, or route. More expensive agents can be invoked only when a prior result justifies escalation. - Cameron can watch root activity and inspect complete lineage without being buried in raw trace noise. - Model behavior can improve through versioned prompts, adapters, and Tinker training while old outputs remain attributable to their exact runtime. +- A conceptualizer consumer can extract concepts and directional links from source events, producing private derived observations that are inspectable and rebuildable without any outbound publication authority. ## Explicit non-goals for the first release diff --git a/spec/web-auth.md b/spec/web-auth.md new file mode 100644 index 0000000..0b7fa46 --- /dev/null +++ b/spec/web-auth.md @@ -0,0 +1,103 @@ +# Public web and inspector authentication + +## Assets + +Protected assets are private Jazz events, source identities, manifests, traces, runtime paths and metadata, OAuth state and DPoP keys, access and refresh tokens, the confidential-client private key, browser sessions, and the Basic break-glass credential. + +Public assets are the four reviewed Markdown pages, OAuth client metadata, and the public half of the client JWKS. + +## Trust boundaries + +1. Nginx terminates HTTPS, redirects `www.thought.stream` to the canonical origin before application routing, and forwards only GET, HEAD, the two required OAuth POSTs, and one exact bounded Review-decision POST to a loopback web proxy. +2. The web proxy serves an exact public route allowlist. It has no public generic file handler and no public upstream fallback. +3. Only the authenticated `/inspector/` route family can reach the loopback inspector. Authorization, cookie, forwarding, and hop-by-hop headers are removed first. +4. The official ATProto OAuth client crosses the network to discovered authorization/resource servers with its hardened resolver, PKCE, PAR, DPoP, and nonce handling. +5. OAuth secrets and browser sessions are encrypted in an owner-only store outside Git and outside the ThoughtStream runtime root. + +## Route matrix + +| Route | Methods | Authority | Upstream access | +| --- | --- | --- | --- | +| `/`, `/docs`, `/docs/architecture`, `/docs/security` | GET, HEAD | public reviewed files | none | +| `/oauth/client-metadata.json`, `/oauth/jwks.json` | GET, HEAD | public OAuth discovery | none | +| `/oauth/login` | GET, HEAD, POST | public flow initiation | authorization server only through SDK | +| `/oauth/callback` | GET | one-time browser-bound state | token endpoint only through SDK | +| `/oauth/logout` | GET, HEAD, POST | valid OAuth browser session; CSRF on POST | exact local generation deletion only | +| `/inspector/`, `/inspector/*` except the decision route | GET, HEAD | allowlisted OAuth DID or enabled Basic fallback | loopback inspector | +| `/inspector/api/reviews/:item/decisions` | POST | allowlisted OAuth browser session, session CSRF, and configured proxy-to-inspector Review capability | one fixed append-only Review decision | +| every other route | none | none | none | + +## Threats and controls + +### Public-to-private route confusion + +Encoded traversal, unknown paths, former root `/api` paths, and unsupported methods terminate in the public proxy. They never become arbitrary filesystem paths and never fall through to the inspector. `/inspector` redirects to `/inspector/` only after authentication so the inspector's relative `api/...` requests remain inside the private prefix. + +Public pages and inspector data responses use an inert `script-src 'none'` policy. Authenticated inspector HTML receives a separate route-scoped policy that permits its audited inline loader and same-origin snapshot requests while forbidding forms and framing. The proxy selects this policy from the trusted loopback response content type; public routes never inherit it. + +### Credential forwarding and response smuggling + +Only a small request-header allowlist reaches the inspector. Authorization, cookie, `X-Forwarded-*`, and proxy headers are discarded. Upstream cookies and authentication challenges are discarded. Hop-by-hop headers are never copied. + +### Login CSRF, callback injection, and replay + +Every login creates a random **application state** in a separate expiring flow store and binds it to an `HttpOnly`, `Secure`, `SameSite=Lax` cookie. That application state is passed to the SDK's `authorize` call. The SDK independently generates the OAuth **protocol state**, stores it with PKCE/DPoP material, and sends that distinct value through the authorization request. On callback, ThoughtStream requires exactly one bounded protocol-state query field and one unique application-state cookie, then delegates protocol-state validation and one-time consumption to the SDK. Only the application state returned by the SDK is compared with the cookie and consumed from the flow store. Protocol and application state must not be conflated. + +The login page and its redirect response permit HTTPS form navigation because the atproto profile requires the client to redirect the browser from its local POST to the dynamically discovered Authorization Server after PAR. This exception is route-scoped to `/oauth/login`; other public pages retain self-only form destinations. The browser policy does not choose the destination: the official SDK's hardened identity, resource-server, and authorization-server discovery returns the redirect URL, and HTTP authorization endpoints remain forbidden. + +Missing or duplicate cookies fail before SDK callback. Mismatched or expired application state is detected after the SDK has consumed protocol state; attempt-scoped credentials are then deleted locally without promotion. Application cleanup sends no remote revocation because a provider may define revocation broadly enough to invalidate a newer grant for the same DID; the SDK residual below still applies. Protocol-state replay is rejected by the SDK store. Callback handling is serialized inside the single process so the SDK's separate state-store `get` and `del` calls cannot race each other. These failures create no browser session. + +### Wrong-account authorization + +The callback compares the returned DID to one configured DID using constant-time byte comparison. Credentials for any other authenticated DID are deleted from local staging without promotion or an application-initiated remote revocation. The expected handle is a login hint, not the authorization decision. + +### Token theft and browser-session theft + +OAuth state, DPoP keys, access/refresh tokens, and browser sessions are AES-256-GCM encrypted at rest under a separately injected key. Files and directories are owner-only and replaced atomically. The confidential ES256 client key is injected separately. Browser cookies hold random opaque ids, never a DID or OAuth token, and are short-lived, Secure, HttpOnly, SameSite=Lax, and host-only. + +A stolen browser cookie is still a bearer credential until expiry. The service permits one current browser session, invalidates older browser cookies when a new callback settles, restores the server-side OAuth session only when its persisted per-DID generation matches the browser session, and uses a bounded lifetime. Expiry, DID mismatch, restore failure, and rejected callback settlement delete only their matching local generation. Explicit CSRF-checked logout deletes the verified generation locally and clears the browser cookie. Application logout deliberately sends no remote revocation because provider-wide semantics could invalidate a newer concurrently promoted grant. This does not replace host/browser security. + +### Logout CSRF + +Logout is a POST with a random token stored only in the server-side browser session and rendered only to an authenticated browser. Missing, duplicate, or mismatched tokens fail without revocation or disclosure. + +The same session token protects the exact Review-decision JSON route. The token is returned only from an OAuth-authenticated private session endpoint and is sent in a dedicated request header. Basic authorization never receives it and remains read-only. A successful CSRF check does not reach Jazz directly: the proxy signs the exact method, normalized route, body digest, timestamp, and one-time nonce under a separately injected capability. The inspector verifies that envelope before accepting the fixed Review decision schema. Browser authentication, CSRF, loopback capability, and event validation are distinct gates. + +### SSRF and hostile OAuth metadata + +Authorization-server, resource-server, DID, and handle discovery are delegated to the official Node OAuth client and its hardened fetch/resolver stack. ThoughtStream does not implement permissive metadata fetching or accept operator-supplied authorization endpoints. HTTP is disabled for production metadata. + +### Resource exhaustion and disconnects + +Every encrypted application and SDK store has independent entry-count and serialized-plaintext byte limits. Oversized writes fail atomically and retain the prior document. Encrypted envelope size is checked before read/decrypt. Login and callback are limited independently at both nginx and process layers, with bounded process limiter maps. Browser disconnect aborts SDK authorization discovery/PAR through the SDK-supported `authorize(..., { signal })` path. + +The installed SDK callback API has no `AbortSignal` option. ThoughtStream therefore runs the entire callback settlement path against one authoritative watchdog: SDK exchange, application-state consumption, generation promotion, browser-session persistence, and cleanup. Timeout synchronously marks both the application attempt and staging attempt non-promotable before releasing the global serializer. Every persistent write rechecks authority immediately before atomic rename, and late completion is removed locally through detached cleanup. + +Each promoted DID session receives a monotonically increasing local generation. Browser authority names that generation. SDK restore runs inside an exact-generation capability: get sees only that generation, refresh set and failure delete can mutate only that generation, stale completion becomes a no-op, and a post-restore persisted-generation check runs before browser authority returns. Cleanup deletes only an exact matching generation. + +Expired staging remains quarantined until its SDK promise settles so a late SDK session-store `set` can be absorbed locally instead of taking that particular store-failure path. The staged record is then dropped without an application-initiated remote revocation. This is not a provider-side guarantee: the unmodified SDK may still revoke remotely after issuer, exchange, or other session-store failure, and provider revocation may be grant-wide. Preventing that would require transport interception, a fork, or killable process isolation. The enforceable guarantee is local: stale attempts cannot promote, restore, overwrite, or delete a newer generation. + +A callback that never resolves retains an SDK promise and an inert staging tombstone. Retained attempts are capped at eight. At capacity, new callbacks fail closed with `503`, `Retry-After`, and an explicit operator-recycle-required response. The flow cookie and application/protocol state are preserved so the same callback can be retried after recovery. Capacity recovers when attempts settle and are dropped or when the singleton proxy process restarts. + +### Query-string custody + +The canonical nginx access-log format records `$uri`, never `$request` or `$request_uri`. Callback access logging is disabled and its location error log is discarded so nginx cannot serialize the request line. The Node proxy emits no per-request URL logging and catches OAuth errors without printing SDK exceptions. OAuth callback query strings, codes, issuer parameters, and protocol state must never enter access logs, service output, incidents, or durable ThoughtStream events. + +### Single-process encrypted-store contract + +The encrypted stores are whole-document, in-process serialized stores rather than multi-writer databases. This is a single process storage contract: exactly one OAuth-capable proxy process may own the fixed `~/.local/share/thoughtstream-inspector-auth` directory. Custom OAuth store paths are rejected because the systemd sandbox grants write access only to that path. The application acquires an owner-only PID/token lock in that directory before opening any SDK or application store, refuses a live owner, reclaims only a parseable dead-PID lock, and releases its lock on graceful shutdown. The systemd unit is a singleton service. Horizontal replicas, templated instances, manual parallel starts, and shared network filesystems are unsupported; scale-out requires replacing this store implementation. + +### Availability, Basic break-glass, and lockout + +OAuth configuration is optional at startup. Basic fallback is separately configured and remains enabled through activation. When enabled, a valid Basic credential authorizes only `/inspector/` reads and bypasses OAuth restoration; it is never forwarded upstream, never receives Review session state, and grants no OAuth, PDS, write, or public-route authority. When disabled, Basic credentials are ignored and `WWW-Authenticate: Basic` is not advertised. Startup fails if both OAuth and Basic fallback are disabled. + +Basic fallback is removed only after external metadata/JWKS fetch, real authorization redirect, allowlisted callback, private read, logout/local generation deletion, and a separately exercised Basic rollback read are all observed. A test result or active service status is not that receipt. + +## Residual risks + +- The proxy and inspector run on the same host; host compromise defeats this boundary. +- The official OAuth SDK and its dependency graph remain trusted code. +- OAuth browser sessions reveal highly private data if stolen and, when the separate Review capability is configured, can append one bounded decision under CSRF. They still grant no generic mutation authority. +- Public documentation requires editorial review; a route-safe renderer cannot prevent a human from committing sensitive prose to an allowlisted public file. +- The Basic fallback remains a high-value bearer credential while enabled and must remain HTTPS-only and independently rate-limited at the edge if exposed beyond the single-user activation window. +- Callback token exchange cannot currently be canceled through the official SDK API after it begins; the SDK version exposes abort only for authorization discovery/PAR. The watchdog removes authority and queue blockage, not the underlying unresolved SDK promise. diff --git a/src/adapters/model-adapters.ts b/src/adapters/model-adapters.ts new file mode 100644 index 0000000..2e5bcef --- /dev/null +++ b/src/adapters/model-adapters.ts @@ -0,0 +1,197 @@ +import fs from "node:fs/promises"; +import path from "node:path"; +import YAML from "yaml"; +import { z } from "zod"; +import { canonicalJson, sha256, type JsonObject } from "../core/json.js"; +import { BUILTIN_PROVIDER_PROFILE_IDS, createBuiltinProviderProfileResolver } from "../agents/provider-profiles.js"; + +export type ModelAdapterPrivacyClass = "public" | "private" | "sensitive"; +export type ModelAdapterExportClass = "public" | "restricted" | "forbidden"; + +export interface ModelAdapterDatasetReceipt { id: string; sha256: string; } +export interface ModelAdapterReleaseIdentity { + id: string; + version: number; + description: string; + releasedAt: string; + manifestSha256: string; + checkpointSelector: { kind: "env"; reference: string }; + providerProfile: string; + baseModel: string; + dataset: ModelAdapterDatasetReceipt; + evals: ModelAdapterDatasetReceipt[]; + capabilities: string[]; + privacyClass: ModelAdapterPrivacyClass; + exportClass: ModelAdapterExportClass; + sourceManifest?: ModelAdapterDatasetReceipt | undefined; +} +export interface ModelAdapterIdentity extends ModelAdapterReleaseIdentity { + binding: { checkpointReferenceSha256: string }; +} + +const digest = z.string().regex(/^[a-f0-9]{64}$/); +const identifier = z.string().min(1).max(200).regex(/^[a-z0-9][a-z0-9._/-]*$/); +const publicModelId = z.string().min(3).max(300).regex(/^[A-Za-z0-9][A-Za-z0-9._/-]*$/); +const env = z.string().regex(/^[A-Z_][A-Z0-9_]*$/); +const receipt = z.object({ id: identifier, sha256: digest }).strict(); +const releaseManifestSchema = z.object({ + schemaVersion: z.literal(1), + id: z.string().min(1).max(100).regex(/^[a-z0-9][a-z0-9-]*$/), + version: z.number().int().positive(), + description: z.string().min(1).max(1000), + releasedAt: z.iso.datetime(), + providerProfile: z.enum(BUILTIN_PROVIDER_PROFILE_IDS), + baseModel: publicModelId, + checkpoint: z.object({ env }).strict(), + dataset: receipt, + evals: z.array(receipt).min(1).max(32), + capabilities: z.array(identifier).min(1).max(32), + privacyClass: z.enum(["public", "private", "sensitive"]), + exportClass: z.enum(["public", "restricted", "forbidden"]), + sourceManifest: receipt.optional(), +}).strict().superRefine((value, context) => { + const duplicateCapability = findDuplicate(value.capabilities); + if (duplicateCapability) context.addIssue({ code: "custom", path: ["capabilities"], message: `Duplicate capability ${duplicateCapability}` }); + const duplicateEval = findDuplicate(value.evals.map((entry) => entry.id)); + if (duplicateEval) context.addIssue({ code: "custom", path: ["evals"], message: `Duplicate eval receipt ${duplicateEval}` }); +}); +const deploymentCatalogSchema = z.object({ + schemaVersion: z.literal(1), + generation: z.number().int().positive(), + expectedCatalogSha256: digest.optional(), + expectedProcesses: z.array(z.string().min(1).max(200)).max(128).optional(), + selections: z.array(z.object({ + declaration: z.object({ id: z.string().min(1), version: z.number().int().positive() }).strict(), + release: z.object({ id: z.string().min(1), version: z.number().int().positive(), manifestSha256: digest }).strict(), + state: z.enum(["candidate", "active", "retired"]), + }).strict()).max(256), +}).strict().superRefine((value, context) => { + const duplicateProcess = findDuplicate(value.expectedProcesses ?? []); + if (duplicateProcess) context.addIssue({ code: "custom", path: ["expectedProcesses"], message: `Duplicate expected process ${duplicateProcess}` }); +}); + +type ReleaseManifest = z.infer; +type DeploymentCatalog = z.infer; + +export function modelAdapterReleaseJson(identity: ModelAdapterReleaseIdentity): JsonObject { + const { sourceManifest: _sourceManifest, ...publicIdentity } = identity; + return { ...publicIdentity, checkpointSelector: { ...identity.checkpointSelector }, dataset: { ...identity.dataset }, evals: identity.evals.map((x) => ({ ...x })), capabilities: [...identity.capabilities], ...(identity.sourceManifest ? { sourceManifest: { ...identity.sourceManifest } } : {}) }; +} +export function modelAdapterIdentityJson(identity: ModelAdapterIdentity): JsonObject { return { ...modelAdapterReleaseJson(identity), binding: { ...identity.binding } }; } +export function modelAdapterIdentitySha256(identity: ModelAdapterIdentity): string { return sha256(canonicalJson(modelAdapterIdentityJson(identity))); } +export function adapterReleaseKey(identity: Pick): string { return `${identity.id}@${identity.version}`; } +export const modelAdapterReleaseIdentitySchema = z.object({ + id: z.string(), version: z.number().int().positive(), description: z.string(), releasedAt: z.iso.datetime(), manifestSha256: digest, + checkpointSelector: z.object({ kind: z.literal("env"), reference: env }).strict(), providerProfile: z.enum(BUILTIN_PROVIDER_PROFILE_IDS), baseModel: publicModelId, dataset: receipt, evals: z.array(receipt), capabilities: z.array(z.string()), privacyClass: z.enum(["public", "private", "sensitive"]), exportClass: z.enum(["public", "restricted", "forbidden"]), sourceManifest: receipt.optional(), +}).strict(); +export const modelAdapterIdentitySchema = modelAdapterReleaseIdentitySchema.extend({ binding: z.object({ checkpointReferenceSha256: digest }).strict() }).strict(); + +export interface LoadedAdapterCatalog { + digest: string; + generation: number; + releases: readonly ModelAdapterReleaseIdentity[]; + selectedByDeclaration: ReadonlyMap; +} + +const privateCheckpointBindings = new WeakMap(); + +export async function loadAdapterCatalog(projectRoot: string, environment: NodeJS.ProcessEnv = process.env): Promise { + const root = path.join(projectRoot, "adapters"); + const releaseDirectory = path.join(root, "releases"); + const deploymentPath = path.join(root, "deployment.yaml"); + const releaseDirectoryExists = await fs.stat(releaseDirectory).then((stat) => stat.isDirectory()).catch(() => false); + if (!releaseDirectoryExists) throw new Error("Adapter startup requires adapters/releases"); + await fs.stat(deploymentPath).catch(() => { throw new Error("Adapter startup requires adapters/deployment.yaml"); }); + const entries = await fs.readdir(releaseDirectory, { withFileTypes: true }); + if (!entries.length) throw new Error("Adapter startup requires a complete release catalog"); + const releases = new Map(); + for (const entry of entries.sort((a, b) => a.name.localeCompare(b.name))) { + if (!entry.isFile() || !/\.ya?ml$/i.test(entry.name)) continue; + const manifest = releaseManifestSchema.parse(YAML.parse(await fs.readFile(path.join(releaseDirectory, entry.name), "utf8"))) as ReleaseManifest; + const identity: ModelAdapterReleaseIdentity = deepFreeze({ + id: manifest.id, version: manifest.version, description: manifest.description, releasedAt: manifest.releasedAt, + manifestSha256: sha256(canonicalJson(manifest as unknown as JsonObject)), checkpointSelector: { kind: "env", reference: manifest.checkpoint.env }, providerProfile: manifest.providerProfile, baseModel: manifest.baseModel, + dataset: { ...manifest.dataset }, evals: manifest.evals.map((x) => ({ ...x })), capabilities: [...manifest.capabilities].sort(), privacyClass: manifest.privacyClass, exportClass: manifest.exportClass, ...(manifest.sourceManifest ? { sourceManifest: { ...manifest.sourceManifest } } : {}), + }); + if (releases.has(adapterReleaseKey(identity))) throw new Error(`Duplicate release ${adapterReleaseKey(identity)}`); + releases.set(adapterReleaseKey(identity), identity); + } + if (!releases.size) throw new Error("Adapter startup requires at least one release manifest"); + const deployment = deploymentCatalogSchema.parse(YAML.parse(await fs.readFile(deploymentPath, "utf8"))) as DeploymentCatalog; + const deploymentIdentity: JsonObject = { + schemaVersion: deployment.schemaVersion, + generation: deployment.generation, + ...(deployment.expectedProcesses ? { expectedProcesses: [...deployment.expectedProcesses].sort() } : {}), + selections: [...deployment.selections].sort((left, right) => ( + left.declaration.id.localeCompare(right.declaration.id) + || left.declaration.version - right.declaration.version + )) as unknown as JsonObject["selections"], + }; + const bundle: JsonObject = { + releases: [...releases.values()].map(modelAdapterReleaseJson).sort((a, b) => ( + String(a.id).localeCompare(String(b.id)) || Number(a.version) - Number(b.version) + )), + deployment: deploymentIdentity, + }; + const catalogDigest = sha256(canonicalJson(bundle)); + if (deployment.expectedCatalogSha256 && deployment.expectedCatalogSha256 !== catalogDigest) throw new Error("Deployment catalog digest mismatch"); + if (deployment.expectedProcesses?.length) { + const processIdentity = environment.THOUGHTSTREAM_ADAPTER_PROCESS_ID?.trim(); + if (!processIdentity || !deployment.expectedProcesses.includes(processIdentity)) { + throw new Error("Adapter deployment does not authorize this process identity"); + } + } + const selectedByDeclaration = new Map(); + const seenDeclarations = new Set(); + for (const selected of deployment.selections) { + const declarationKey = `${selected.declaration.id}@${selected.declaration.version}`; + if (seenDeclarations.has(declarationKey)) throw new Error(`Ambiguous selected deployment ${declarationKey}`); + seenDeclarations.add(declarationKey); + const release = releases.get(`${selected.release.id}@${selected.release.version}`); + if (!release || release.manifestSha256 !== selected.release.manifestSha256) throw new Error(`Selected release is missing or has a digest mismatch: ${selected.release.id}@${selected.release.version}`); + if (selected.state !== "active") continue; + const checkpoint = environment[release.checkpointSelector.reference]?.trim(); + if (!checkpoint) throw new Error(`Selected release checkpoint is unresolved: ${release.id}@${release.version}`); + const providerProfiles = createBuiltinProviderProfileResolver(environment); + providerProfiles.resolve(release.providerProfile, release.baseModel); + providerProfiles.resolve(release.providerProfile, checkpoint); + const binding = deepFreeze({ ...release, binding: { checkpointReferenceSha256: sha256(checkpoint) } }); + privateCheckpointBindings.set(binding, checkpoint); + selectedByDeclaration.set(declarationKey, binding); + } + return deepFreeze({ + digest: catalogDigest, + generation: deployment.generation, + releases: [...releases.values()], + selectedByDeclaration: readonlyMap(selectedByDeclaration), + }); +} + +export function privateCheckpointFor(identity: ModelAdapterIdentity): string { + const checkpoint = privateCheckpointBindings.get(identity); + if (!checkpoint || sha256(checkpoint) !== identity.binding.checkpointReferenceSha256) { + throw new Error(`Adapter binding checkpoint is unavailable or mismatched: ${adapterReleaseKey(identity)}`); + } + return checkpoint; +} + +function deepFreeze(value: T): T { if (value && typeof value === "object") { Object.freeze(value); for (const child of Object.values(value as Record)) deepFreeze(child); } return value; } + +function readonlyMap(source: Map): ReadonlyMap { + return new Proxy(source, { + get(target, property) { + if (property === "set" || property === "delete" || property === "clear") return undefined; + const value = Reflect.get(target, property, target); + return typeof value === "function" ? value.bind(target) : value; + }, + }); +} + +function findDuplicate(values: string[]): string | undefined { + const seen = new Set(); + for (const value of values) { + if (seen.has(value)) return value; + seen.add(value); + } + return undefined; +} diff --git a/src/agents/declarations.ts b/src/agents/declarations.ts index 11c9e4c..b767a5d 100644 --- a/src/agents/declarations.ts +++ b/src/agents/declarations.ts @@ -4,11 +4,19 @@ import YAML from "yaml"; import { z } from "zod"; import { canonicalJson, sha256, type JsonObject } from "../core/json.js"; import { createDefaultRegistry } from "../events/registry.js"; -import { createOutputContractRegistry, OBSERVATION_OUTPUT_CONTRACT_ID, OBSERVATION_OUTPUT_CONTRACT_VERSION } from "./output-contracts.js"; +import { + CONCEPTUALIZATION_OUTPUT_CONTRACT_ID, + createOutputContractRegistry, + OBSERVATION_OUTPUT_CONTRACT_ID, + OBSERVATION_OUTPUT_CONTRACT_VERSION, +} from "./output-contracts.js"; import { BUILTIN_PROVIDER_PROFILE_IDS, providerKindForProfile } from "./provider-profiles.js"; import { AGENT_TOOL_NAMES } from "./tools.js"; +import { loadAdapterCatalog, privateCheckpointFor, type LoadedAdapterCatalog } from "../adapters/model-adapters.js"; import type { ThoughtAgentDeclaration } from "./types.js"; +const privateDeclarationCheckpoints = new WeakMap(); + const budgetCounterSchema = z.number().int().positive().max(1_000_000_000); const budgetCostSchema = z.number().int().positive().max(1_000_000_000_000); const budgetLimitFields = { @@ -77,6 +85,7 @@ const piRunnerSchema = z.object({ profile: z.enum(BUILTIN_PROVIDER_PROFILE_IDS).optional(), tier: z.enum(["triage-small", "reasoning-small", "escalation"]).optional(), model: z.string().min(1).max(500).optional(), + adapter: z.object({ id: z.string().min(1), version: z.number().int().positive() }).strict().optional(), outputMode: z.enum(["strict-json", "conversation-text"]).default("strict-json"), ...runnerLimits, }).strict(); @@ -148,10 +157,42 @@ const declarationFileSchema = z.object({ }).strict(), enabled: z.boolean().default(true), }).strict().superRefine((value, context) => { + const conceptualizationOutput = value.outputContract.id === CONCEPTUALIZATION_OUTPUT_CONTRACT_ID; + const conceptGraphEvent = value.emit[0] === "stream.thought.derived.concept.graph"; + const conversationText = (value.runner.kind === "pi" && value.runner.outputMode === "conversation-text") + || (value.runner.kind === "letta-agent-sdk" && value.runner.responseMode === "conversation-text"); + if (value.role === "standard" && conceptualizationOutput !== conceptGraphEvent) { + context.addIssue({ + code: "custom", + path: conceptualizationOutput ? ["emit"] : ["outputContract"], + message: "Standard conceptualization output and stream.thought.derived.concept.graph must be selected together", + }); + } + if (conceptualizationOutput && conversationText) { + context.addIssue({ + code: "custom", + path: ["runner"], + message: "Conceptualization output requires strict JSON mode", + }); + } + if (value.runner.kind === "pi" && value.runner.profile === "openai-json-default" && !conceptualizationOutput) { + context.addIssue({ + code: "custom", + path: ["runner", "profile"], + message: "The strict OpenAI JSON profile is currently bound only to the conceptualization output contract", + }); + } + if (conversationText && value.outputContract.id !== OBSERVATION_OUTPUT_CONTRACT_ID) { + context.addIssue({ + code: "custom", + path: ["outputContract"], + message: "Conversation-text mode requires the observation output contract", + }); + } if (value.runner.kind === "pi" && !value.runner.profile) { context.addIssue({ code: "custom", path: ["runner", "profile"], message: "Pi agents require a trusted provider profile" }); } - if (value.runner.kind === "pi" && !value.runner.model && !value.runner.tier) { + if (value.runner.kind === "pi" && !value.runner.model && !value.runner.tier && !value.runner.adapter) { context.addIssue({ code: "custom", path: ["runner", "model"], message: "Pi agents require a model or capability tier" }); } if (value.runner.kind !== "pi" && value.policy.tools.length > 0) { @@ -202,20 +243,31 @@ const declarationFileSchema = z.object({ message: "Telegram conversation context requires one sensitive Telegram message subscription", }); } - if (value.context.atprotoObject && ( - value.runner.kind !== "letta-agent-sdk" - || !["single-event", "atproto-batch"].includes(value.context.strategy) - || value.context.maxChars < 2_048 - || !value.subscribe.types.some((type) => ( + const lettaAtprotoObjectContext = value.runner.kind === "letta-agent-sdk" + && ["single-event", "atproto-batch"].includes(value.context.strategy) + && value.context.maxEvents === 1 + && value.subscribe.types.some((type) => ( type === "stream.thought.source.atproto.commit" || type === "stream.thought.derived.event.batch" || type === "*" - )) + )); + const conceptualizerAtprotoBatchContext = value.runner.kind === "pi" + && conceptualizationOutput + && conceptGraphEvent + && value.context.strategy === "atproto-batch" + && value.context.maxEvents === 1 + && value.subscribe.types.length === 1 + && value.subscribe.types[0] === "stream.thought.derived.event.batch" + && value.policy.tools.length === 0 + && value.policy.externalActions === false; + if (value.context.atprotoObject && ( + value.context.maxChars < 2_048 + || (!lettaAtprotoObjectContext && !conceptualizerAtprotoBatchContext) )) { context.addIssue({ code: "custom", path: ["context", "atprotoObject"], - message: "ATProto object context requires a Letta Agent SDK source/batch declaration with at least 2048 characters and an ATProto subscription", + message: "ATProto object context requires either a bounded Letta Agent SDK source/batch declaration or a tool-free Pi conceptualizer bound to one batch event", }); } if (value.runner.kind === "letta-agent-sdk") { @@ -251,6 +303,7 @@ const declarationFileSchema = z.object({ export async function loadAgentDeclarations( directory: string, environment: NodeJS.ProcessEnv = process.env, + catalog?: LoadedAdapterCatalog, ): Promise { const registry = createDefaultRegistry(); const outputContracts = createOutputContractRegistry(); @@ -258,6 +311,8 @@ export async function loadAgentDeclarations( if (error.code === "ENOENT") return []; throw error; }); + const projectRoot = path.resolve(directory, ".."); + const adapterCatalog = catalog ?? await loadAdapterCatalogIfPresent(projectRoot, environment); const declarations: ThoughtAgentDeclaration[] = []; for (const entry of entries.sort((left, right) => left.name.localeCompare(right.name))) { if (!entry.isFile() || !/\.(ya?ml|json)$/i.test(entry.name)) continue; @@ -268,7 +323,6 @@ export async function loadAgentDeclarations( for (const type of file.emit) { if (!registry.has(type, 1)) throw new Error(`Agent ${file.id} emits an unregistered event type: ${type}@1`); } - const projectRoot = path.resolve(directory, ".."); const promptPath = path.resolve(projectRoot, file.prompt); if (!isWithin(projectRoot, promptPath)) throw new Error(`Agent prompt escapes project root: ${file.prompt}`); const systemPrompt = await fs.readFile(promptPath, "utf8"); @@ -278,7 +332,28 @@ export async function loadAgentDeclarations( if (compiledEventTypes.length === 0) { throw new Error(`Agent ${file.id} subscription compiles to no registered event types`); } - const concreteModel = resolveRunnerModel(file.runner, environment, file.enabled); + const selectedAdapter = file.runner.kind === "pi" && file.runner.adapter + ? adapterCatalog?.selectedByDeclaration.get(`${file.id}@${file.version}`) + : undefined; + const selectedRelease = selectedAdapter + ? adapterCatalog?.releases.find((release) => ( + release.id === selectedAdapter.id + && release.version === selectedAdapter.version + && release.manifestSha256 === selectedAdapter.manifestSha256 + )) + : undefined; + if (file.runner.kind === "pi" && file.runner.adapter && (!adapterCatalog || !selectedAdapter + || selectedAdapter.id !== file.runner.adapter.id + || selectedAdapter.version !== file.runner.adapter.version)) { + throw new Error(`Agent ${file.id}@${file.version} adapter selection does not match the loaded deployment catalog`); + } + if (selectedAdapter && !selectedRelease) { + throw new Error(`Agent ${file.id}@${file.version} selected release is absent from the loaded release catalog`); + } + if (file.runner.kind === "pi" && file.runner.adapter && (file.runner.model || file.runner.tier)) { + throw new Error(`Agent ${file.id}@${file.version} cannot combine adapter selection with model or tier`); + } + const concreteModel = selectedAdapter?.baseModel ?? resolveRunnerModel(file.runner, environment, file.enabled); const lettaAgentId = file.runner.kind === "letta-agent-sdk" ? resolveLettaAgentId(file.runner.agentIdEnv, environment, file.enabled) : undefined; @@ -291,9 +366,15 @@ export async function loadAgentDeclarations( mode: file.runner.kind, role: file.role, outputContract: { ...outputContract }, - ...(file.runner.kind === "pi" && file.runner.profile ? { - provider: providerKindForProfile(file.runner.profile), - providerProfile: file.runner.profile, + ...(file.runner.kind === "pi" && (selectedAdapter?.providerProfile ?? file.runner.profile) ? { + provider: providerKindForProfile((selectedAdapter?.providerProfile ?? file.runner.profile)!), + providerProfile: (selectedAdapter?.providerProfile ?? file.runner.profile)!, + } : {}), + ...(selectedAdapter ? { + modelAdapterRelease: selectedRelease, + modelAdapter: selectedAdapter, + adapterCatalogDigest: adapterCatalog!.digest, + adapterCatalogGeneration: adapterCatalog!.generation, } : {}), ...(file.runner.kind === "letta-agent-sdk" ? { provider: "letta-cloud" as const, @@ -334,7 +415,9 @@ export async function loadAgentDeclarations( externalActions: file.policy.externalActions, }; declaration.declarationFingerprint = declarationFingerprint(declaration); - declarations.push(declaration); + const compiledDeclaration = deepFreeze(declaration); + if (selectedAdapter) privateDeclarationCheckpoints.set(compiledDeclaration, privateCheckpointFor(selectedAdapter)); + declarations.push(compiledDeclaration); } const duplicate = findDuplicate(declarations.map((declaration) => declaration.id)); if (duplicate) throw new Error(`Duplicate agent declaration id: ${duplicate}`); @@ -343,9 +426,10 @@ export async function loadAgentDeclarations( function resolveRunnerModel(runner: { kind: "deterministic" | "pi" | "letta-agent-sdk"; - profile?: "tinker-default" | "openai-compatible-default" | undefined; + profile?: "tinker-default" | "openai-json-default" | "openai-compatible-default" | undefined; tier?: "triage-small" | "reasoning-small" | "escalation" | undefined; model?: string | undefined; + adapter?: { id: string; version: number } | undefined; }, environment: NodeJS.ProcessEnv, enabled: boolean): string | undefined { if (runner.model) return runner.model; if (!runner.tier) return undefined; @@ -380,6 +464,21 @@ export function declarationFingerprint(declaration: ThoughtAgentDeclaration): st return sha256(canonicalJson(JSON.parse(JSON.stringify(body)) as JsonObject)); } +export function assertTrustedAdapterDeclaration(declaration: ThoughtAgentDeclaration): void { + if (declaration.modelAdapter && !privateDeclarationCheckpoints.has(declaration)) { + throw new Error(`Agent ${declaration.id}@${declaration.version} adapter declaration was not compiled by the trusted declaration loader`); + } +} + +export function privateCheckpointForDeclaration(declaration: ThoughtAgentDeclaration): string { + assertTrustedAdapterDeclaration(declaration); + const checkpoint = privateDeclarationCheckpoints.get(declaration); + if (!checkpoint || !declaration.modelAdapter || sha256(checkpoint) !== declaration.modelAdapter.binding.checkpointReferenceSha256) { + throw new Error(`Agent ${declaration.id}@${declaration.version} has no valid private adapter binding`); + } + return checkpoint; +} + export function agentRole(declaration: ThoughtAgentDeclaration): "standard" | "repair" { return declaration.role ?? "standard"; } @@ -410,3 +509,20 @@ function findDuplicate(values: string[]): string | undefined { } return undefined; } + +function deepFreeze(value: T): T { + if (value && typeof value === "object") { + Object.freeze(value); + for (const child of Object.values(value as Record)) deepFreeze(child); + } + return value; +} + +async function loadAdapterCatalogIfPresent(projectRoot: string, environment: NodeJS.ProcessEnv): Promise { + try { + await fs.access(path.join(projectRoot, "adapters", "deployment.yaml")); + } catch { + return undefined; + } + return loadAdapterCatalog(projectRoot, environment); +} diff --git a/src/agents/execution-adapters.ts b/src/agents/execution-adapters.ts new file mode 100644 index 0000000..4b1f4af --- /dev/null +++ b/src/agents/execution-adapters.ts @@ -0,0 +1,8 @@ +import { LETTA_AGENT_SDK_ADAPTER_REVISION } from "./letta-agent-sdk.js"; +import type { ThoughtAgentDeclaration } from "./types.js"; + +export function executionAdapterRevisionFor(declaration: ThoughtAgentDeclaration): string { + if (declaration.mode === "pi") return "pi-openai-completions-v1"; + if (declaration.mode === "letta-agent-sdk") return LETTA_AGENT_SDK_ADAPTER_REVISION; + return "deterministic-triage-v1"; +} diff --git a/src/agents/letta-agent-sdk.ts b/src/agents/letta-agent-sdk.ts index 910c43e..693117c 100644 --- a/src/agents/letta-agent-sdk.ts +++ b/src/agents/letta-agent-sdk.ts @@ -128,7 +128,7 @@ export class LettaAgentSdkRunner implements AgentRunner { kind: "letta.session.opened", data: { backend: "cloud", - adapterRevision: LETTA_AGENT_SDK_ADAPTER_REVISION, + executionAdapterRevision: LETTA_AGENT_SDK_ADAPTER_REVISION, sdkVersion: LETTA_AGENT_SDK_PACKAGE_VERSION, agentId: config.agentId!, conversation: "main", @@ -370,7 +370,7 @@ export class LettaAgentSdkRunner implements AgentRunner { private parseOutput(declaration: ThoughtAgentDeclaration, text: string): AgentOutput { const identity = outputContractForDeclaration(declaration); const config = declaration.lettaAgent!; - let parsed: ObservationOutput; + let parsed: AgentOutput; if (config.responseMode === "conversation-text") { const summary = text.trim(); if (!summary) throw invalidSdkOutput("empty-final-text", text, declaration); diff --git a/src/agents/output-contracts.ts b/src/agents/output-contracts.ts index 92d9e9b..0ddaa10 100644 --- a/src/agents/output-contracts.ts +++ b/src/agents/output-contracts.ts @@ -4,6 +4,12 @@ import { canonicalJson, sha256, type JsonObject } from "../core/json.js"; export const OBSERVATION_OUTPUT_CONTRACT_ID = "stream.thought.output.observation"; export const OBSERVATION_OUTPUT_CONTRACT_VERSION = 1; +export const CONCEPTUALIZATION_OUTPUT_CONTRACT_ID = "stream.thought.output.conceptualization"; +export const CONCEPTUALIZATION_OUTPUT_CONTRACT_VERSION = 1; + +export const REVIEW_RESPONSE_OUTPUT_CONTRACT_ID = "stream.thought.output.review-response"; +export const REVIEW_RESPONSE_OUTPUT_CONTRACT_VERSION = 1; + const recommendationSchema = z.object({ target: z.string().min(1).max(200), reason: z.string().min(1).max(2_000), @@ -20,6 +26,71 @@ export const observationOutputSchema = z.object({ export type ObservationOutput = z.infer; +const conceptRelationshipEnum = z.enum([ + "RELATES_TO", + "DESCRIBES", + "MENTIONS", + "EXEMPLIFIES", + "CONTRADICTS", + "QUESTIONS", + "SUPPORTS", + "CRITIQUES", +]); + +const conceptItemSchema = z.object({ + text: z.string().min(1).max(60).regex(/^[a-z0-9]+( [a-z0-9]+){0,2}$/, "Concept text must be lowercase 1-3 words with spaces"), + relationship: conceptRelationshipEnum, +}).strict(); + +const conceptLinkSchema = z.object({ + fromIndex: z.number().int().nonnegative(), + toIndex: z.number().int().nonnegative(), + relationship: conceptRelationshipEnum, +}).strict(); + +export const conceptualizationOutputSchema = z.object({ + summary: z.string().min(1).max(2_000), + concepts: z.array(conceptItemSchema).min(0).max(20), + links: z.array(conceptLinkSchema).max(20).optional(), + confidence: z.number().min(0).max(1), +}).strict().superRefine((value, context) => { + for (let index = 0; index < (value.links?.length ?? 0); index += 1) { + const link = value.links![index]!; + if (link.fromIndex >= value.concepts.length) { + context.addIssue({ code: "custom", path: ["links", index, "fromIndex"], message: "Link source is outside the concept array" }); + } + if (link.toIndex >= value.concepts.length) { + context.addIssue({ code: "custom", path: ["links", index, "toIndex"], message: "Link target is outside the concept array" }); + } + } +}); + +export type ConceptualizationOutput = z.infer; + +export function reviewResponseSummary(response: string): string { + return response.length <= 2_000 ? response : `${response.slice(0, 1_999)}…`; +} + +export const reviewResponseOutputSchema = z.object({ + response: z.string().min(1).max(60_000), + summary: z.string().min(1).max(2_000).optional(), + confidence: z.number().min(0).max(1).optional(), +}).strict().superRefine((value, context) => { + if (value.summary !== undefined && value.summary !== reviewResponseSummary(value.response)) { + context.addIssue({ code: "custom", path: ["summary"], message: "Summary must be the canonical bounded response preview" }); + } + if (value.confidence !== undefined && value.confidence !== 1) { + context.addIssue({ code: "custom", path: ["confidence"], message: "Review response confidence is deterministic" }); + } +}).transform((value) => ({ + response: value.response, + summary: reviewResponseSummary(value.response), + confidence: 1, +})); + +export type ReviewResponseOutput = z.infer; +export type SemanticOutput = ObservationOutput | ConceptualizationOutput | ReviewResponseOutput; + export interface OutputContractIdentity { id: string; version: number; @@ -31,11 +102,11 @@ export interface SanitizedOutputIssue { path: Array; } -export interface OutputContractDefinition { +export interface OutputContractDefinition { identity: OutputContractIdentity; definition: JsonObject; prompt: string; - schema: z.ZodType; + schema: z.ZodType; } const observationDefinition: JsonObject = { @@ -67,7 +138,7 @@ const observationIdentity: OutputContractIdentity = { sha256: sha256(canonicalJson(observationDefinition)), }; -export const OBSERVATION_OUTPUT_CONTRACT: OutputContractDefinition = { +export const OBSERVATION_OUTPUT_CONTRACT: OutputContractDefinition = { identity: observationIdentity, definition: observationDefinition, prompt: "Return exactly one raw JSON object and nothing else. Use only the required top-level keys summary, tags, importance, confidence, plus optional recommendation. Minimal valid example: {\"summary\":\"One concise observation\",\"tags\":[],\"importance\":\"low\",\"confidence\":0.5}. The importance value must be exactly one of \"low\", \"normal\", or \"high\"; \"medium\" is invalid. The optional recommendation object may contain only target, reason, and proposedAction. Do not wrap the object, add unknown fields, use Markdown fences, or write text before or after it.", @@ -89,12 +160,12 @@ export class OutputContractValidationError extends Error { export class OutputContractRegistry { private readonly contracts = new Map(); - register(contract: OutputContractDefinition): this { + register(contract: OutputContractDefinition): this { const key = outputContractKey(contract.identity.id, contract.identity.version); if (this.contracts.has(key)) throw new Error(`Output contract already registered: ${key}`); const expectedHash = sha256(canonicalJson(contract.definition)); if (contract.identity.sha256 !== expectedHash) throw new Error(`Output contract definition hash mismatch: ${key}`); - this.contracts.set(key, contract); + this.contracts.set(key, contract as unknown as OutputContractDefinition); return this; } @@ -112,16 +183,145 @@ export class OutputContractRegistry { return contract; } - validate(identity: OutputContractIdentity, value: unknown): ObservationOutput { + validate(identity: OutputContractIdentity, value: unknown): SemanticOutput { + const contract = this.resolve(identity); + const result = contract.schema.safeParse(value); + if (!result.success) throw new OutputContractValidationError(contract.identity, sanitizeOutputIssues(result.error.issues)); + return result.data as SemanticOutput; + } + + canonicalize(identity: OutputContractIdentity, value: unknown): JsonObject { const contract = this.resolve(identity); const result = contract.schema.safeParse(value); if (!result.success) throw new OutputContractValidationError(contract.identity, sanitizeOutputIssues(result.error.issues)); - return result.data; + return JSON.parse(canonicalJson(result.data as JsonObject)) as JsonObject; } } +const conceptualizationDefinition: JsonObject = { + id: CONCEPTUALIZATION_OUTPUT_CONTRACT_ID, + version: CONCEPTUALIZATION_OUTPUT_CONTRACT_VERSION, + type: "object", + unknownFields: "reject", + canonicalizer: { id: "stream.thought.canonical-json-after-schema", version: 1 }, + invariants: [ + "links[].fromIndex and links[].toIndex are integers", + "0 <= links[].fromIndex < concepts.length", + "0 <= links[].toIndex < concepts.length", + ], + fields: { + summary: { type: "string", minChars: 1, maxChars: 2_000, required: true }, + concepts: { + type: "array", + maxItems: 20, + required: true, + items: { + type: "object", + unknownFields: "reject", + fields: { + text: { type: "string", minChars: 1, maxChars: 60, pattern: "^[a-z0-9]+( [a-z0-9]+){0,2}$", required: true }, + relationship: { type: "enum", values: ["RELATES_TO", "DESCRIBES", "MENTIONS", "EXEMPLIFIES", "CONTRADICTS", "QUESTIONS", "SUPPORTS", "CRITIQUES"], required: true }, + }, + }, + }, + links: { + type: "array", + maxItems: 20, + required: false, + items: { + type: "object", + unknownFields: "reject", + fields: { + fromIndex: { type: "integer", minimum: 0, maximumExclusivePath: "concepts.length", required: true }, + toIndex: { type: "integer", minimum: 0, maximumExclusivePath: "concepts.length", required: true }, + relationship: { type: "enum", values: ["RELATES_TO", "DESCRIBES", "MENTIONS", "EXEMPLIFIES", "CONTRADICTS", "QUESTIONS", "SUPPORTS", "CRITIQUES"], required: true }, + }, + }, + }, + confidence: { type: "number", minimum: 0, maximum: 1, required: true }, + }, +}; + +const conceptualizationIdentity: OutputContractIdentity = { + id: CONCEPTUALIZATION_OUTPUT_CONTRACT_ID, + version: CONCEPTUALIZATION_OUTPUT_CONTRACT_VERSION, + sha256: sha256(canonicalJson(conceptualizationDefinition)), +}; + +export const CONCEPTUALIZATION_OUTPUT_JSON_SCHEMA: JsonObject = { + type: "object", + additionalProperties: false, + required: ["summary", "concepts", "links", "confidence"], + properties: { + summary: { type: "string", minLength: 1, maxLength: 2_000 }, + concepts: { + type: "array", + maxItems: 20, + items: { + type: "object", + additionalProperties: false, + required: ["text", "relationship"], + properties: { + text: { type: "string", minLength: 1, maxLength: 60, pattern: "^[a-z0-9]+( [a-z0-9]+){0,2}$" }, + relationship: { type: "string", enum: ["RELATES_TO", "DESCRIBES", "MENTIONS", "EXEMPLIFIES", "CONTRADICTS", "QUESTIONS", "SUPPORTS", "CRITIQUES"] }, + }, + }, + }, + links: { + type: "array", + maxItems: 20, + items: { + type: "object", + additionalProperties: false, + required: ["fromIndex", "toIndex", "relationship"], + properties: { + fromIndex: { type: "integer", minimum: 0, maximum: 19 }, + toIndex: { type: "integer", minimum: 0, maximum: 19 }, + relationship: { type: "string", enum: ["RELATES_TO", "DESCRIBES", "MENTIONS", "EXEMPLIFIES", "CONTRADICTS", "QUESTIONS", "SUPPORTS", "CRITIQUES"] }, + }, + }, + }, + confidence: { type: "number", minimum: 0, maximum: 1 }, + }, +}; + +export const CONCEPTUALIZATION_OUTPUT_CONTRACT: OutputContractDefinition = { + identity: conceptualizationIdentity, + definition: conceptualizationDefinition, + prompt: "Return exactly one raw JSON object and nothing else. Use only the required top-level keys summary, concepts, links, and confidence. The concepts array must contain objects with exactly the keys text and relationship; strings are invalid. Each concept text must be lowercase 1-3 words (letters, numbers, spaces only). Each relationship must be exactly one of: RELATES_TO, DESCRIBES, MENTIONS, EXEMPLIFIES, CONTRADICTS, QUESTIONS, SUPPORTS, CRITIQUES. Links use zero-based fromIndex/toIndex into the concepts array and the same relationship enum. Minimal valid example: {\"summary\":\"Execution receipts support durable agent memory\",\"concepts\":[{\"text\":\"agent memory\",\"relationship\":\"DESCRIBES\"},{\"text\":\"execution receipts\",\"relationship\":\"SUPPORTS\"}],\"links\":[{\"fromIndex\":1,\"toIndex\":0,\"relationship\":\"SUPPORTS\"}],\"confidence\":0.5}. Do not wrap the object, add unknown fields, use Markdown fences, or write text before or after it.", + schema: conceptualizationOutputSchema, +}; + +const reviewResponseDefinition: JsonObject = { + id: REVIEW_RESPONSE_OUTPUT_CONTRACT_ID, + version: REVIEW_RESPONSE_OUTPUT_CONTRACT_VERSION, + type: "object", + unknownFields: "reject", + fields: { + response: { type: "string", minChars: 1, maxChars: 60_000, required: true }, + summary: { type: "string", minChars: 1, maxChars: 2_000, required: false, deterministic: "bounded response preview" }, + confidence: { type: "number", minimum: 1, maximum: 1, required: false, deterministic: true }, + }, +}; + +const reviewResponseIdentity: OutputContractIdentity = { + id: REVIEW_RESPONSE_OUTPUT_CONTRACT_ID, + version: REVIEW_RESPONSE_OUTPUT_CONTRACT_VERSION, + sha256: sha256(canonicalJson(reviewResponseDefinition)), +}; + +export const REVIEW_RESPONSE_OUTPUT_CONTRACT: OutputContractDefinition = { + identity: reviewResponseIdentity, + definition: reviewResponseDefinition, + prompt: "Return exactly one raw JSON object with the single key response and nothing else. Put the complete response to the review prompt in that string. Do not add unknown fields, Markdown fences around the JSON object, or text before or after it.", + schema: reviewResponseOutputSchema, +}; + export function createOutputContractRegistry(): OutputContractRegistry { - return new OutputContractRegistry().register(OBSERVATION_OUTPUT_CONTRACT); + return new OutputContractRegistry() + .register(OBSERVATION_OUTPUT_CONTRACT) + .register(CONCEPTUALIZATION_OUTPUT_CONTRACT) + .register(REVIEW_RESPONSE_OUTPUT_CONTRACT); } export function defaultOutputContractIdentity(): OutputContractIdentity { @@ -152,14 +352,12 @@ export function sanitizeOutputIssues(issues: z.core.$ZodIssue[], limit = 20): Sa })); } -export function structuredOutputJson(output: ObservationOutput): JsonObject { - return { - summary: output.summary, - tags: output.tags, - importance: output.importance, - confidence: output.confidence, - ...(output.recommendation ? { recommendation: { ...output.recommendation } } : {}), - }; +export function canonicalStructuredOutput( + registry: OutputContractRegistry, + identity: OutputContractIdentity, + value: unknown, +): JsonObject { + return registry.canonicalize(identity, value); } export function outputContractKey(id: string, version: number): string { diff --git a/src/agents/pi.ts b/src/agents/pi.ts index 70800bc..141fd8a 100644 --- a/src/agents/pi.ts +++ b/src/agents/pi.ts @@ -2,8 +2,10 @@ import type { AssistantMessage, ImageContent } from "@earendil-works/pi-ai"; import { createHash, randomUUID } from "node:crypto"; import path from "node:path"; import type { JsonObject } from "../core/json.js"; +import { privateCheckpointForDeclaration } from "./declarations.js"; import type { InferenceUsage } from "../store/types.js"; import { + CONCEPTUALIZATION_OUTPUT_JSON_SCHEMA, createOutputContractRegistry, outputContractForDeclaration, outputContractIdentityJson, @@ -11,6 +13,7 @@ import { type ObservationOutput, type OutputContractIdentity, type OutputContractRegistry, + type SemanticOutput, } from "./output-contracts.js"; import { createBuiltinProviderProfileResolver, @@ -47,11 +50,13 @@ export class PiAgentRunner implements AgentRunner { const outputContract = this.outputContracts.resolve(outputContractIdentity); if (!declaration.model) throw new Error(`Pi agent ${declaration.id} has no model`); if (!declaration.providerProfile) throw new Error(`Pi agent ${declaration.id} has no trusted provider profile`); - const profile = this.providerProfiles.resolve(declaration.providerProfile, declaration.model); + const providerModel = declaration.modelAdapter ? privateCheckpointForDeclaration(declaration) : declaration.model; + const profile = this.providerProfiles.resolve(declaration.providerProfile, providerModel); const providerIdentity: JsonObject = { provider: profile.provider, providerProfile: profile.id, model: declaration.model, + ...(declaration.modelAdapter ? { modelAdapter: declaration.modelAdapter as unknown as JsonObject } : {}), }; const acceptsImages = profile.imageInputModels.has(declaration.model); let observedRevision: string | undefined; @@ -82,7 +87,8 @@ export class PiAgentRunner implements AgentRunner { const broker = await startProviderBroker({ profile, - model: declaration.model, + model: providerModel, + traceModel: declaration.model, runId: input.runId, timeoutMs: declaration.timeoutMs, maxOutputTokens: declaration.maxOutputTokens, @@ -105,10 +111,13 @@ export class PiAgentRunner implements AgentRunner { images: prefetched.images, model: { provider: profile.provider, - id: declaration.model, + id: providerModel, reasoning: profile.provider === "tinker", acceptsImages, jsonObjectResponseFormat: declaration.outputMode !== "conversation-text" && profile.jsonObjectResponseFormat, + ...(declaration.outputMode !== "conversation-text" && profile.jsonSchemaResponseFormat + ? { jsonSchemaResponseFormat: CONCEPTUALIZATION_OUTPUT_JSON_SCHEMA } + : {}), contextWindow: 131_072, maxTokens: declaration.maxOutputTokens, }, @@ -124,24 +133,24 @@ export class PiAgentRunner implements AgentRunner { timeoutMs: declaration.timeoutMs + 2_000, }); observedRevision = result.observedRevision; - for (const trace of result.traces) { + const messageUpdateCount = result.traces.filter((trace) => trace.kind === "pi.message_update").length; + for (const trace of result.traces.filter((candidate) => candidate.kind !== "pi.message_update")) { await onTrace({ kind: trace.kind, data: traceMetadata(trace.data) }); } + if (messageUpdateCount > 0) { + await onTrace({ kind: "pi.message_updates_coalesced", data: { count: messageUpdateCount } }); + } const finalMessage = validateFinalAssistant(result.finalMessage, outputContractIdentity); observedUsage = inferenceUsageFromAssistant(finalMessage); const parsed = declaration.outputMode === "conversation-text" ? parseConversationText(finalMessage, outputContractIdentity, this.outputContracts) : parseFinalOutput(finalMessage, outputContractIdentity, this.outputContracts); return { - summary: parsed.summary, - tags: parsed.tags, - importance: parsed.importance, - confidence: parsed.confidence, - ...(parsed.recommendation ? { recommendation: parsed.recommendation } : {}), + ...parsed, model: { provider: profile.provider, id: declaration.model, - ...(observedRevision ? { revision: observedRevision } : {}), + ...(declaration.modelAdapter ? { revision: `sha256:${declaration.modelAdapter.binding.checkpointReferenceSha256}` } : observedRevision ? { revision: observedRevision } : {}), }, ...(toolSet.outcomes.length > 0 ? { enrichments: toolSet.outcomes } : {}), ...(observedUsage ? { usage: observedUsage } : {}), @@ -156,7 +165,9 @@ export class PiAgentRunner implements AgentRunner { diagnostic: { ...(error.diagnostic ?? {}), ...providerIdentity, - ...(observedRevision ? { checkpointRevision: observedRevision } : {}), + ...(declaration.modelAdapter + ? { checkpointRevision: `sha256:${declaration.modelAdapter.binding.checkpointReferenceSha256}` } + : observedRevision ? { checkpointRevision: observedRevision } : {}), }, }); } @@ -322,7 +333,7 @@ function parseFinalOutput( message: AssistantMessage, identity: OutputContractIdentity, registry: OutputContractRegistry, -): ObservationOutput { +): SemanticOutput { const diagnostic = finalOutputDiagnostic(message); const textParts = message.content.filter((part) => part.type === "text"); const authoritativeParts = message.content.filter((part) => part.type !== "thinking"); @@ -365,12 +376,14 @@ function parseConversationText( if (summary.length > MAX_CONVERSATION_TEXT_CHARS) { throw invalidConversationText("final-text-too-large", diagnostic, identity); } - return registry.validate(identity, { + const output = registry.validate(identity, { summary, tags: ["conversation"], importance: "normal", confidence: 0.5, }); + if (!("tags" in output)) throw invalidConversationText("output-contract-invalid", diagnostic, identity); + return output; } function invalidConversationText( diff --git a/src/agents/provider-profiles.ts b/src/agents/provider-profiles.ts index f190769..c1f7e1b 100644 --- a/src/agents/provider-profiles.ts +++ b/src/agents/provider-profiles.ts @@ -11,6 +11,7 @@ export interface ProviderProfile { allowedModels: ReadonlySet; imageInputModels: ReadonlySet; jsonObjectResponseFormat: boolean; + jsonSchemaResponseFormat: boolean; requestTimeoutMs: number; maxRequestBytes: number; maxResponseBytes: number; @@ -23,11 +24,12 @@ export interface ProviderProfileResolver { const profileIdSchema = z.string().min(1).max(100).regex(/^[a-z0-9][a-z0-9.-]*$/); const modelSchema = z.string().min(1).max(500); -export const BUILTIN_PROVIDER_PROFILE_IDS = ["tinker-default", "openai-compatible-default"] as const; +export const BUILTIN_PROVIDER_PROFILE_IDS = ["tinker-default", "openai-json-default", "openai-compatible-default"] as const; export type BuiltinProviderProfileId = typeof BUILTIN_PROVIDER_PROFILE_IDS[number]; export function providerKindForProfile(profileId: string): ProviderKind { if (profileId === "tinker-default") return "tinker"; + if (profileId === "openai-json-default") return "openai-compatible"; if (profileId === "openai-compatible-default") return "openai-compatible"; throw new Error(`Unknown trusted provider profile: ${profileId}`); } @@ -39,7 +41,7 @@ export function createBuiltinProviderProfileResolver(environment: NodeJS.Process modelSchema.parse(model); const profile = buildBuiltinProfile(profileId, environment); if (!profile.allowedModels.has(model)) { - throw new Error(`Model ${model} is not allowlisted by provider profile ${profileId}`); + throw new Error(`Requested model is not allowlisted by provider profile ${profileId}`); } return profile; }, @@ -53,7 +55,7 @@ export function staticProviderProfileResolver(profiles: ProviderProfile[]): Prov const profile = byId.get(profileId); if (!profile) throw new Error(`Unknown trusted provider profile: ${profileId}`); if (!profile.allowedModels.has(model)) { - throw new Error(`Model ${model} is not allowlisted by provider profile ${profileId}`); + throw new Error(`Requested model is not allowlisted by provider profile ${profileId}`); } return profile; }, @@ -72,9 +74,19 @@ function buildBuiltinProfile(profileId: string, environment: NodeJS.ProcessEnv): baseUrl: "https://tinker.thinkingmachines.dev/services/tinker-prod/oai/api/v1", route: "/chat/completions", apiKeyEnv: "TINKER_API_KEY", - allowedModels: new Set(["Qwen/Qwen3.5-4B", "Qwen/Qwen3.6-27B", "thinkingmachines/Inkling", ...configured]), + allowedModels: new Set([ + "Qwen/Qwen3.5-4B", + "Qwen/Qwen3.5-35B-A3B-Base", + "Qwen/Qwen3.6-27B", + "thinkingmachines/Inkling", + ...configured, + ]), imageInputModels: new Set(configuredAllowedModels(environment.THOUGHTSTREAM_TINKER_IMAGE_MODELS)), - jsonObjectResponseFormat: true, + // Tinker accepts response_format=json_object but does not enforce JSON + // across its current model routes. Strictness remains a parent-side + // complete-value validation contract, not a constrained-decoding claim. + jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: false, requestTimeoutMs: 120_000, maxRequestBytes: 2 * 1024 * 1024, maxResponseBytes: 1_500_000, @@ -96,6 +108,23 @@ function buildBuiltinProfile(profileId: string, environment: NodeJS.ProcessEnv): allowedModels, imageInputModels: new Set(configuredAllowedModels(environment.THOUGHTSTREAM_MODEL_IMAGE_MODELS)), jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: false, + requestTimeoutMs: 120_000, + maxRequestBytes: 2 * 1024 * 1024, + maxResponseBytes: 1_500_000, + }); + } + if (profileId === "openai-json-default") { + return validateProfile({ + id: profileId, + provider: "openai-compatible", + baseUrl: "https://api.openai.com/v1", + route: "/chat/completions", + apiKeyEnv: "OPENAI_API_KEY", + allowedModels: new Set(["gpt-4.1-mini"]), + imageInputModels: new Set(), + jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: true, requestTimeoutMs: 120_000, maxRequestBytes: 2 * 1024 * 1024, maxResponseBytes: 1_500_000, @@ -131,7 +160,7 @@ function validateProfile(profile: ProviderProfile): ProviderProfile { for (const model of profile.imageInputModels) { modelSchema.parse(model); if (!profile.allowedModels.has(model)) { - throw new Error(`Provider profile ${profile.id} marks non-allowlisted model ${model} as image-capable`); + throw new Error(`Provider profile ${profile.id} marks a non-allowlisted model as image-capable`); } } return { ...profile, baseUrl: profile.baseUrl.replace(/\/$/, "") }; diff --git a/src/agents/repairs.ts b/src/agents/repairs.ts index 3524868..af291e2 100644 --- a/src/agents/repairs.ts +++ b/src/agents/repairs.ts @@ -2,9 +2,11 @@ import { canonicalJson, sha256, type JsonObject } from "../core/json.js"; import type { ThoughtEvent } from "../events/types.js"; import type { JazzThoughtStore } from "../jazz/store.js"; import type { AgentRun } from "../store/types.js"; +import { joinPrivacy, runPrivacy } from "../security/privacy.js"; import { agentRole, declarationFingerprint } from "./declarations.js"; import { createOutputContractRegistry, + OBSERVATION_OUTPUT_CONTRACT_ID, outputContractForDeclaration, outputContractIdentityJson, parseOutputContractIdentity, @@ -146,7 +148,7 @@ export class RepairRequestCoordinator { rootEventId: evidence.source.rootEventId, parentEventId: evidence.failed.id, correlationId: run.id, - privacy: evidence.source.privacy, + privacy: joinPrivacy(evidence.source.privacy, evidence.failed.privacy, runPrivacy(run)), payload, traceId: run.id, }); @@ -186,6 +188,7 @@ async function repairEvidence( if (stringField(run.contextManifest.agentRole) !== "standard") return undefined; const outputContract = outputContractForDeclaration(declaration); + if (outputContract.id !== OBSERVATION_OUTPUT_CONTRACT_ID) return undefined; try { outputContracts.resolve(outputContract); const manifestContract = parseOutputContractIdentity(run.contextManifest.outputContract); @@ -206,6 +209,10 @@ async function repairEvidence( id: run.model, ...(checkpointRevision ? { checkpointRevision } : {}), ...(run.adapterRevision ? { adapterRevision: run.adapterRevision } : {}), + ...(run.executionAdapterRevision ? { executionAdapterRevision: run.executionAdapterRevision } : {}), + ...(run.modelAdapter ? { modelAdapter: run.modelAdapter as unknown as JsonObject } : {}), + ...(run.adapterCatalogDigest ? { adapterCatalogDigest: run.adapterCatalogDigest } : {}), + ...(run.adapterCatalogGeneration ? { adapterCatalogGeneration: run.adapterCatalogGeneration } : {}), }; const failedEvents = (await store.listEvents({ types: ["stream.thought.agent.run.failed"] })) .filter((event) => event.payload.runId === run.id); @@ -219,7 +226,7 @@ async function repairEvidence( || failed.traceId !== run.id || failed.rootEventId !== source.rootEventId || failed.parentEventId !== source.id - || failed.privacy !== source.privacy + || failed.privacy !== joinPrivacy(source.privacy, runPrivacy(run)) || failed.payload.runId !== run.id || failed.payload.executionKey !== run.executionKey || failed.payload.agentId !== run.agentId diff --git a/src/agents/runtime.ts b/src/agents/runtime.ts index 17084aa..c1a2da1 100644 --- a/src/agents/runtime.ts +++ b/src/agents/runtime.ts @@ -14,17 +14,24 @@ import { buildTelegramConversationContextPacket, type AtprotoObjectContextOptions, } from "./context.js"; -import { declarationFingerprint } from "./declarations.js"; +import { assertTrustedAdapterDeclaration, declarationFingerprint } from "./declarations.js"; +import { executionAdapterRevisionFor } from "./execution-adapters.js"; +import { declarationPrivacy } from "../security/privacy.js"; import { DeterministicTriageRunner } from "./deterministic.js"; import { LETTA_AGENT_SDK_ADAPTER_REVISION, LettaAgentSdkRunner } from "./letta-agent-sdk.js"; import { + canonicalStructuredOutput, + CONCEPTUALIZATION_OUTPUT_CONTRACT_ID, createOutputContractRegistry, + OBSERVATION_OUTPUT_CONTRACT_ID, + REVIEW_RESPONSE_OUTPUT_CONTRACT_ID, + reviewResponseSummary, outputContractForDeclaration, outputContractIdentityJson, OutputContractValidationError, - structuredOutputJson, type OutputContractRegistry, } from "./output-contracts.js"; +import { REVIEW_RESPONSE_EVENT_TYPE } from "../review/types.js"; import { PiAgentRunner } from "./pi.js"; import { RepairRequestCoordinator } from "./repairs.js"; import { ConsumerScheduler } from "./scheduler.js"; @@ -86,6 +93,8 @@ export class ThoughtAgentRuntime { async registerDeclarations(declarations: ThoughtAgentDeclaration[]): Promise { assertUniqueLettaAgentOwners(declarations); for (const declaration of declarations) { + assertTrustedAdapterDeclaration(declaration); + assertOutputContractEventBinding(declaration); assertRuntimeAccountingPolicy(declaration); this.declarationsByVersion.set(`${declaration.id}@${declaration.version}`, declaration); const declarationJson = asJsonObject(declaration); @@ -293,9 +302,18 @@ export class ThoughtAgentRuntime { attempt, provider: declaration.provider ?? declaration.mode, model: declaration.model ?? declaration.mode, - adapterRevision: adapterRevisionFor(declaration), + privacy: outputPrivacy(declaration, event), + adapterRevision: executionAdapterRevisionFor(declaration), + executionAdapterRevision: executionAdapterRevisionFor(declaration), + ...(declaration.modelAdapter ? { modelAdapter: declaration.modelAdapter } : {}), + ...(declaration.adapterCatalogDigest ? { adapterCatalogDigest: declaration.adapterCatalogDigest } : {}), + ...(declaration.adapterCatalogGeneration ? { adapterCatalogGeneration: declaration.adapterCatalogGeneration } : {}), promptHash: sha256(declaration.systemPrompt), - contextManifest: context.manifest, + contextManifest: { + ...context.manifest, + privacy: executionPrivacy(declaration, event), + ...adapterEvidence(declaration), + }, createdAt: startedAt, startedAt, updatedAt: startedAt, @@ -360,12 +378,21 @@ export class ThoughtAgentRuntime { }); }); try { - const { model, enrichments, usage: providerUsage, ...semanticOutput } = candidateOutput; + const { model: reportedModel, enrichments, usage: providerUsage, ...semanticOutput } = candidateOutput; usage = providerUsage; - const structured = this.outputContracts.validate(outputContractForDeclaration(declaration), semanticOutput); + const structured = this.outputContracts.validate( + outputContractForDeclaration(declaration), + semanticOutput, + ); output = { ...structured, - ...(model ? { model } : {}), + ...(declaration.modelAdapter ? { + model: { + provider: declaration.provider ?? declaration.modelAdapter.providerProfile, + id: declaration.modelAdapter.baseModel, + revision: publicAdapterRevision(declaration), + }, + } : reportedModel ? { model: reportedModel } : {}), ...(enrichments ? { enrichments } : {}), ...(providerUsage ? { usage: providerUsage } : {}), }; @@ -386,9 +413,15 @@ export class ThoughtAgentRuntime { if (error instanceof AgentRunFailure && error.usage) usage = error.usage; await this.settleAccountingReservation(run, usage, completedAt); const message = error instanceof AgentRunFailure ? error.message : "Agent runner failed"; - const failureDiagnostic = error instanceof AgentRunFailure + const rawFailureDiagnostic = error instanceof AgentRunFailure ? error.diagnostic : { code: "unclassified-run-failure", stage: "runner" }; + const failureDiagnostic = declaration.modelAdapter ? { + ...(rawFailureDiagnostic ?? {}), + provider: declaration.provider ?? declaration.modelAdapter.providerProfile, + model: declaration.modelAdapter.baseModel, + checkpointRevision: publicAdapterRevision(declaration), + } : rawFailureDiagnostic; const diagnosticProvider = jsonStringField(failureDiagnostic, "provider"); const diagnosticModel = jsonStringField(failureDiagnostic, "model"); const diagnosticRevision = jsonStringField(failureDiagnostic, "checkpointRevision"); @@ -400,6 +433,7 @@ export class ThoughtAgentRuntime { ...(diagnosticModel ? { model: diagnosticModel } : {}), ...(diagnosticRevision ? { checkpointRevision: diagnosticRevision } : {}), ...(failureDiagnostic ? { result: { failureDiagnostic } } : {}), + executionAdapterRevision: executionAdapterRevisionFor(declaration), completedAt, updatedAt: completedAt, }; @@ -444,6 +478,7 @@ export class ThoughtAgentRuntime { ...(output.model.revision ? { checkpointRevision: output.model.revision } : {}), } : {}), result: asJsonObject(output), + executionAdapterRevision: executionAdapterRevisionFor(declaration), completedAt, updatedAt: completedAt, }; @@ -691,6 +726,8 @@ export class ThoughtAgentRuntime { if ((declaration.role ?? "standard") === "repair") { return this.repairProposalCandidate(declaration, event, run, result, at); } + const observation = isObservationOutput(result) ? result : undefined; + const reviewResponse = isReviewResponseOutput(result) ? result : undefined; return { type: declaration.outputEventType, schemaVersion: 1, @@ -703,20 +740,24 @@ export class ThoughtAgentRuntime { rootEventId: event.rootEventId, parentEventId: event.id, correlationId: event.correlationId, - privacy: executionPrivacy(declaration, event), + privacy: outputPrivacy(declaration, event), payload: { runId: run.id, executionKey: run.executionKey, inputEventId: event.id, inputSourceSequence: event.sourceSequence, - summary: result.summary, - ...(result.tags ? { tags: result.tags } : {}), - ...(result.importance ? { importance: result.importance } : {}), - ...(result.recommendation ? { recommendation: result.recommendation as JsonObject } : {}), - confidence: result.confidence, + summary: reviewResponse ? reviewResponseSummary(reviewResponse.response) : result.summary, + ...(observation ? { tags: observation.tags, importance: observation.importance } : {}), + ...(observation?.recommendation ? { recommendation: observation.recommendation as JsonObject } : {}), + ...(reviewResponse ? {} : { confidence: result.confidence }), outputContract: outputContractIdentityJson(outputContractForDeclaration(declaration)), - structuredOutput: structuredOutputJson(result), + structuredOutput: canonicalStructuredOutput( + this.outputContracts, + outputContractForDeclaration(declaration), + semanticOutputForPersistence(result), + ), ...(result.model ? { model: asJsonObject(result.model) } : {}), + ...adapterEvidence(declaration), }, traceId: run.id, }; @@ -753,7 +794,7 @@ export class ThoughtAgentRuntime { rootEventId: sourceRootEventId, parentEventId: event.id, correlationId: originalRunId, - privacy: event.privacy, + privacy: outputPrivacy(declaration, event), payload: { originalRunId, repairRequestEventId: event.id, @@ -761,7 +802,14 @@ export class ThoughtAgentRuntime { originalTriggerEventId, sourceRootEventId, outputContract: declarationContract, - structuredOutput: structuredOutputJson(result), + originalModel: asJsonObject(event.payload.model), + repairModel: runModelEvidence(run, declaration), + structuredOutput: canonicalStructuredOutput( + this.outputContracts, + outputContractForDeclaration(declaration), + semanticOutputForPersistence(result), + ), + ...adapterEvidence(declaration), }, traceId: run.id, }; @@ -809,6 +857,7 @@ export class ThoughtAgentRuntime { outputContract: outputContractIdentityJson(outputContractForDeclaration(declaration)), inputEventIds: [event.id], attempt: run.attempt, + ...adapterEvidence(declaration), ...payload, }, traceId: run.id, @@ -826,6 +875,28 @@ function assertRuntimeAccountingPolicy(declaration: ThoughtAgentDeclaration): vo } } +function assertOutputContractEventBinding(declaration: ThoughtAgentDeclaration): void { + const contractId = outputContractForDeclaration(declaration).id; + const conceptualizationOutput = contractId === CONCEPTUALIZATION_OUTPUT_CONTRACT_ID; + const conceptGraphEvent = declaration.outputEventType === "stream.thought.derived.concept.graph"; + if ((declaration.role ?? "standard") === "standard" && conceptualizationOutput !== conceptGraphEvent) { + throw new Error(`Agent ${declaration.id}@${declaration.version} must bind conceptualization output and concept graph events together`); + } + const reviewResponseOutput = contractId === REVIEW_RESPONSE_OUTPUT_CONTRACT_ID; + const reviewResponseEvent = declaration.outputEventType === REVIEW_RESPONSE_EVENT_TYPE; + if ((declaration.role ?? "standard") === "standard" && reviewResponseOutput !== reviewResponseEvent) { + throw new Error(`Agent ${declaration.id}@${declaration.version} must bind review response output and review response events together`); + } + const conversationText = (declaration.mode === "pi" && declaration.outputMode === "conversation-text") + || (declaration.mode === "letta-agent-sdk" && declaration.lettaAgent?.responseMode === "conversation-text"); + if (conceptualizationOutput && conversationText) { + throw new Error(`Agent ${declaration.id}@${declaration.version} requires strict JSON for conceptualization output`); + } + if (conversationText && contractId !== OBSERVATION_OUTPUT_CONTRACT_ID) { + throw new Error(`Agent ${declaration.id}@${declaration.version} requires the observation contract for conversation-text output`); + } +} + function assertUniqueLettaAgentOwners(declarations: ThoughtAgentDeclaration[]): void { const owners = new Map(); for (const declaration of declarations) { @@ -865,22 +936,62 @@ function executionPrivacy( declaration: ThoughtAgentDeclaration, event: ThoughtEvent, ): ThoughtEvent["privacy"] { - if (declaration.mode !== "letta-agent-sdk") return event.privacy; - const levels: Record = { - "public-source": 0, - private: 1, - sensitive: 2, + return declaration.mode === "letta-agent-sdk" + ? declarationPrivacy(declaration, ...declaration.acceptedPrivacy, event.privacy) + : declarationPrivacy(declaration, event.privacy); +} + +function outputPrivacy( + declaration: ThoughtAgentDeclaration, + event: ThoughtEvent, +): ThoughtEvent["privacy"] { + const privacy = executionPrivacy(declaration, event); + if (outputContractForDeclaration(declaration).id !== CONCEPTUALIZATION_OUTPUT_CONTRACT_ID) return privacy; + return privacy === "public-source" ? "private" : privacy; +} + +function semanticOutputForPersistence(output: AgentOutput): JsonObject { + const { model: _model, enrichments: _enrichments, usage: _usage, ...semanticOutput } = output; + return semanticOutput as JsonObject; +} + +function isObservationOutput( + output: AgentOutput, +): output is Extract { + return "tags" in output && "importance" in output; +} + +function isReviewResponseOutput( + output: AgentOutput, +): output is Extract { + return "response" in output; +} + +function adapterEvidence(declaration: ThoughtAgentDeclaration): JsonObject { + const catalog = declaration.adapterCatalogDigest ? { + adapterCatalogDigest: declaration.adapterCatalogDigest, + ...(declaration.adapterCatalogGeneration ? { adapterCatalogGeneration: declaration.adapterCatalogGeneration } : {}), + } : {}; + if (!declaration.modelAdapter) return { executionAdapterRevision: executionAdapterRevisionFor(declaration), ...catalog }; + return { + executionAdapterRevision: executionAdapterRevisionFor(declaration), + modelAdapter: declaration.modelAdapter as unknown as JsonObject, + ...catalog, + }; +} + +function runModelEvidence(run: AgentRun, declaration: ThoughtAgentDeclaration): JsonObject { + return { + provider: run.provider, + id: run.model, + ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), + ...adapterEvidence(declaration), }; - return [...declaration.acceptedPrivacy, event.privacy] - .reduce((mostPrivate, candidate) => ( - levels[candidate] > levels[mostPrivate] ? candidate : mostPrivate - ), "public-source" as ThoughtEvent["privacy"]); } -function adapterRevisionFor(declaration: ThoughtAgentDeclaration): string { - if (declaration.mode === "pi") return "pi-openai-completions-v1"; - if (declaration.mode === "letta-agent-sdk") return LETTA_AGENT_SDK_ADAPTER_REVISION; - return "deterministic-triage-v1"; +function publicAdapterRevision(declaration: ThoughtAgentDeclaration): string { + if (!declaration.modelAdapter) throw new Error(`Agent ${declaration.id}@${declaration.version} has no model adapter`); + return `sha256:${declaration.modelAdapter.binding.checkpointReferenceSha256}`; } function executionKey(event: ThoughtEvent, declaration: ThoughtAgentDeclaration): string { diff --git a/src/agents/sandbox/protocol.ts b/src/agents/sandbox/protocol.ts index 12d0840..3f04f6f 100644 --- a/src/agents/sandbox/protocol.ts +++ b/src/agents/sandbox/protocol.ts @@ -24,6 +24,7 @@ export const sandboxRunPacketSchema = z.object({ reasoning: z.boolean(), acceptsImages: z.boolean(), jsonObjectResponseFormat: z.boolean(), + jsonSchemaResponseFormat: z.record(z.string(), z.unknown()).optional(), contextWindow: z.number().int().positive().max(2_000_000), maxTokens: z.number().int().positive().max(32_000), }).strict(), diff --git a/src/agents/sandbox/provider-broker.ts b/src/agents/sandbox/provider-broker.ts index 620c688..f513f39 100644 --- a/src/agents/sandbox/provider-broker.ts +++ b/src/agents/sandbox/provider-broker.ts @@ -38,6 +38,7 @@ export interface ProviderBrokerHandle { export async function startProviderBroker(options: { profile: ProviderProfile; model: string; + traceModel?: string; runId: string; timeoutMs: number; maxOutputTokens: number; @@ -50,7 +51,7 @@ export async function startProviderBroker(options: { const credential = process.env[options.profile.apiKeyEnv]; if (!credential) throw new Error(`Provider profile ${options.profile.id} credential is unavailable`); if (!options.profile.allowedModels.has(options.model)) { - throw new Error(`Model ${options.model} is not allowlisted by provider profile ${options.profile.id}`); + throw new Error(`Requested model is not allowlisted by provider profile ${options.profile.id}`); } const directory = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-provider-")); await fs.chmod(directory, 0o711); @@ -74,6 +75,7 @@ export async function startProviderBroker(options: { let requestBytes = 0; let responseBytes = 0; let reservedResponseBytes = 0; + let admissionQueue = Promise.resolve(); let closed = false; const server = net.createServer((socket) => { void handleConnection(socket).catch(() => { @@ -100,13 +102,20 @@ export async function startProviderBroker(options: { } } - async function handleRequest(request: BrokerRequest): Promise { - if (requests >= maxRequests) { - return rejectResponse( - "turn-exhausted", - maxRequests === 1 ? "Provider capability permits exactly one request" : "Provider capability exhausted its request count", - ); + async function withAdmissionLock(operation: () => T | Promise): Promise { + const previous = admissionQueue; + let release!: () => void; + const current = new Promise((resolve) => { release = resolve; }); + admissionQueue = previous.then(() => current); + await previous; + try { + return await operation(); + } finally { + release(); } + } + + async function handleRequest(request: BrokerRequest): Promise { if (Date.now() > expiresAt) return rejectResponse("capability-expired", "Provider capability expired"); if (!request || request.version !== 1) return rejectResponse("protocol-version", "Unsupported broker protocol version"); if (request.capability !== capability || request.runId !== options.runId) { @@ -120,9 +129,6 @@ export async function startProviderBroker(options: { if (typeof request.body !== "string" || bodyBytes > options.profile.maxRequestBytes) { return rejectResponse("request-oversize", "Provider request exceeds its byte limit"); } - if (requestBytes + bodyBytes > maxTotalRequestBytes) { - return rejectResponse("request-budget-exhausted", "Provider capability exhausted its cumulative request-byte budget"); - } if (!request.headers || Object.entries(request.headers).some(([name]) => !ALLOWED_WORKER_HEADERS.has(name.toLowerCase()))) { return rejectResponse("headers-rejected", "Worker request included an unauthorized header"); } @@ -142,76 +148,113 @@ export async function startProviderBroker(options: { return rejectResponse("token-limit-rejected", "Provider request exceeded the declared output-token limit"); } - const remainingResponseBudget = maxTotalResponseBytes - responseBytes - reservedResponseBytes; - if (remainingResponseBudget <= 0) { - return rejectResponse("response-budget-exhausted", "Provider capability cannot reserve another bounded response"); - } - const responseReservation = Math.min(options.profile.maxResponseBytes, MAX_BROKER_RESPONSE_BODY_BYTES, remainingResponseBudget); - - // Consume the lease before the first await. Concurrent connections cannot both - // observe the same remaining request or byte capacity. - requests += 1; - requestBytes += bodyBytes; - reservedResponseBytes += responseReservation; - await options.onTrace?.("sandbox.provider.request", { - profile: options.profile.id, - provider: options.profile.provider, - model: options.model, - requestIndex: requests, - bodyBytes, - bodySha256: createHash("sha256").update(request.body).digest("hex"), - }); - const controller = new AbortController(); - const timeout = setTimeout(() => controller.abort(), Math.min(options.timeoutMs, options.profile.requestTimeoutMs)); - try { - const endpoint = `${options.profile.baseUrl}${options.profile.route}`; - const upstream = await (options.fetchImpl ?? fetch)(endpoint, { - method: "POST", - redirect: "manual", - signal: controller.signal, - headers: { - authorization: `Bearer ${credential}`, - accept: request.headers.accept ?? "text/event-stream", - "content-type": "application/json", - }, - body: request.body, + const providerRequest = async (): Promise => { + // Reserve broker admission at the actual provider boundary after local request + // validation. Concurrent sockets serialize here and all + // cumulative bounds are rechecked against the current counters. + const admission = await withAdmissionLock(() => { + if (Date.now() > expiresAt) { + return { rejection: rejectResponse("capability-expired", "Provider capability expired") }; + } + if (requests >= maxRequests) { + return { + rejection: rejectResponse( + "turn-exhausted", + maxRequests === 1 ? "Provider capability permits exactly one request" : "Provider capability exhausted its request count", + ), + }; + } + if (requestBytes + bodyBytes > maxTotalRequestBytes) { + return { rejection: rejectResponse("request-budget-exhausted", "Provider capability exhausted its cumulative request-byte budget") }; + } + const remainingResponseBudget = maxTotalResponseBytes - responseBytes - reservedResponseBytes; + if (remainingResponseBudget <= 0) { + return { rejection: rejectResponse("response-budget-exhausted", "Provider capability cannot reserve another bounded response") }; + } + const responseReservation = Math.min( + options.profile.maxResponseBytes, + MAX_BROKER_RESPONSE_BODY_BYTES, + remainingResponseBudget, + ); + requests += 1; + requestBytes += bodyBytes; + reservedResponseBytes += responseReservation; + return { requestIndex: requests, responseReservation }; }); - if (upstream.status >= 300 && upstream.status < 400) { - return rejectResponse("redirect-rejected", "Provider redirect was rejected"); + if ("rejection" in admission) return admission.rejection; + const { requestIndex, responseReservation } = admission; + let responseReservationOutstanding = true; + const controller = new AbortController(); + const timeout = setTimeout(() => controller.abort(), Math.min(options.timeoutMs, options.profile.requestTimeoutMs)); + try { + await options.onTrace?.("sandbox.provider.request", { + profile: options.profile.id, + provider: options.profile.provider, + model: options.traceModel ?? options.model, + requestIndex, + bodyBytes, + bodySha256: createHash("sha256").update(request.body).digest("hex"), + }); + const endpoint = `${options.profile.baseUrl}${options.profile.route}`; + const upstream = await (options.fetchImpl ?? fetch)(endpoint, { + method: "POST", + redirect: "manual", + signal: controller.signal, + headers: { + authorization: `Bearer ${credential}`, + accept: request.headers.accept ?? "text/event-stream", + "content-type": "application/json", + }, + body: request.body, + }); + if (upstream.status >= 300 && upstream.status < 400) { + return rejectResponse("redirect-rejected", "Provider redirect was rejected"); + } + const responseBody = await readBoundedBody(upstream, responseReservation); + const actualResponseBytes = Buffer.byteLength(responseBody); + await withAdmissionLock(() => { + reservedResponseBytes -= responseReservation; + responseBytes += actualResponseBytes; + responseReservationOutstanding = false; + }); + const headers = Object.fromEntries([...upstream.headers.entries()].filter(([name]) => SAFE_RESPONSE_HEADERS.has(name.toLowerCase()))); + await options.onTrace?.("sandbox.provider.response", { + profile: options.profile.id, + provider: options.profile.provider, + model: options.traceModel ?? options.model, + status: upstream.status, + bodyBytes: actualResponseBytes, + bodySha256: createHash("sha256").update(responseBody).digest("hex"), + }); + return { + version: 1, + status: "completed", + httpStatus: upstream.status, + headers, + body: responseBody, + }; + } catch (error) { + await options.onTrace?.("sandbox.provider.failed", { + profile: options.profile.id, + provider: options.profile.provider, + model: options.traceModel ?? options.model, + classification: error instanceof Error && error.name === "AbortError" ? "timeout" : "transport", + }); + return rejectResponse( + error instanceof Error && error.name === "AbortError" ? "provider-timeout" : "provider-transport", + "Provider request failed", + ); + } finally { + if (responseReservationOutstanding) { + await withAdmissionLock(() => { + reservedResponseBytes -= responseReservation; + responseReservationOutstanding = false; + }); + } + clearTimeout(timeout); } - const responseBody = await readBoundedBody(upstream, responseReservation); - responseBytes += Buffer.byteLength(responseBody); - const headers = Object.fromEntries([...upstream.headers.entries()].filter(([name]) => SAFE_RESPONSE_HEADERS.has(name.toLowerCase()))); - await options.onTrace?.("sandbox.provider.response", { - profile: options.profile.id, - provider: options.profile.provider, - model: options.model, - status: upstream.status, - bodyBytes: Buffer.byteLength(responseBody), - bodySha256: createHash("sha256").update(responseBody).digest("hex"), - }); - return { - version: 1, - status: "completed", - httpStatus: upstream.status, - headers, - body: responseBody, - }; - } catch (error) { - await options.onTrace?.("sandbox.provider.failed", { - profile: options.profile.id, - provider: options.profile.provider, - model: options.model, - classification: error instanceof Error && error.name === "AbortError" ? "timeout" : "transport", - }); - return rejectResponse( - error instanceof Error && error.name === "AbortError" ? "provider-timeout" : "provider-transport", - "Provider request failed", - ); - } finally { - reservedResponseBytes -= responseReservation; - clearTimeout(timeout); - } + }; + return providerRequest(); } try { diff --git a/src/agents/sandbox/worker.ts b/src/agents/sandbox/worker.ts index 0f89000..aa26d82 100644 --- a/src/agents/sandbox/worker.ts +++ b/src/agents/sandbox/worker.ts @@ -35,9 +35,21 @@ async function main(): Promise { } const body = await request.text(); const parsed = JSON.parse(body) as Record; - const constrainedBody = packet.model.jsonObjectResponseFormat - ? JSON.stringify({ ...parsed, response_format: { type: "json_object" } }) - : body; + const constrainedBody = packet.model.jsonSchemaResponseFormat + ? JSON.stringify({ + ...parsed, + response_format: { + type: "json_schema", + json_schema: { + name: "thoughtstream_conceptualization", + strict: true, + schema: packet.model.jsonSchemaResponseFormat, + }, + }, + }) + : packet.model.jsonObjectResponseFormat + ? JSON.stringify({ ...parsed, response_format: { type: "json_object" } }) + : body; const headers: Record = {}; for (const name of ["accept", "content-type"]) { const value = request.headers.get(name); diff --git a/src/agents/types.ts b/src/agents/types.ts index 294fae9..6f5c3d9 100644 --- a/src/agents/types.ts +++ b/src/agents/types.ts @@ -1,8 +1,9 @@ import type { JsonObject } from "../core/json.js"; +import type { ModelAdapterIdentity, ModelAdapterReleaseIdentity } from "../adapters/model-adapters.js"; import type { ThoughtEvent } from "../events/types.js"; import type { InferenceBudgetPolicy, InferenceUsage } from "../store/types.js"; import type { AgentContextPacket } from "./context.js"; -import type { OutputContractIdentity } from "./output-contracts.js"; +import type { OutputContractIdentity, SemanticOutput } from "./output-contracts.js"; export type AgentMode = "deterministic" | "pi" | "letta-agent-sdk"; @@ -39,6 +40,10 @@ export interface ThoughtAgentDeclaration { providerProfile?: string | undefined; modelTier?: "triage-small" | "reasoning-small" | "escalation" | undefined; model?: string | undefined; + modelAdapterRelease?: ModelAdapterReleaseIdentity | undefined; + modelAdapter?: ModelAdapterIdentity | undefined; + adapterCatalogDigest?: string | undefined; + adapterCatalogGeneration?: number | undefined; lettaAgent?: LettaAgentSdkConfiguration | undefined; outputMode?: "strict-json" | "conversation-text" | undefined; eventTypes: string[]; @@ -71,16 +76,7 @@ export interface AgentRunInput { context: AgentContextPacket; } -export interface AgentOutput { - summary: string; - tags: string[]; - importance: "low" | "normal" | "high"; - confidence: number; - recommendation?: { - target: string; - reason: string; - proposedAction: string; - } | undefined; +export type AgentOutput = SemanticOutput & { model?: { provider: string; id: string; @@ -89,7 +85,7 @@ export interface AgentOutput { enrichments?: EnrichmentOutcome[] | undefined; /** Trusted provider telemetry. Never part of the semantic output contract. */ usage?: InferenceUsage | undefined; -} +}; export interface EnrichmentOutcome { tool: string; diff --git a/src/cli.ts b/src/cli.ts index 8db3773..8c3158e 100644 --- a/src/cli.ts +++ b/src/cli.ts @@ -33,6 +33,13 @@ import { type JudgmentKind, } from "./training/judgments.js"; import { startInspectorServer } from "./web/inspector.js"; +import { decodeReviewCapability } from "./review/web-capability.js"; +import { + appendReviewPrompt, + createReviewItem, + projectReviewQueue, +} from "./review/review.js"; +import type { PrivacyClass } from "./events/types.js"; const projectRoot = process.env.THOUGHTSTREAM_ROOT ?? process.cwd(); if (process.argv[2] === "stream") { @@ -513,6 +520,33 @@ try { const run = await store.getRun(id); if (!run) throw new Error(`Run not found: ${id}`); print({ run, trace: await store.listTrace(id) }); + } else if (command === "review-prompt") { + const file = valueAfter("--file"); + const externalId = valueAfter("--external-id"); + const privacy = valueAfter("--privacy") ?? "private"; + if (!file || !externalId || !["public-source", "private", "sensitive"].includes(privacy)) { + throw new Error("Usage: thought stream review-prompt --file --external-id [--privacy ] [--source review:prompt]"); + } + const prompt = await appendReviewPrompt(store, { + source: valueAfter("--source") ?? "review:prompt", + externalId, + privacy: privacy as PrivacyClass, + payload: jsonObject(JSON.parse(await fs.readFile(path.resolve(file), "utf8")), "--file"), + }); + print({ reviewPrompt: prompt }); + } else if (command === "review-item") { + const promptEventId = valueAfter("--prompt-event"); + const candidateRunIds = commaSeparated(valueAfter("--candidate-runs")); + if (!promptEventId || candidateRunIds.length !== 2) { + throw new Error("Usage: thought stream review-item --prompt-event --candidate-runs ,"); + } + const item = await createReviewItem(store, { + promptEventId, + candidateRunIds: [candidateRunIds[0]!, candidateRunIds[1]!], + }); + print({ reviewItem: item }); + } else if (command === "review-queue") { + print(await projectReviewQueue(store)); } else if (command === "judgment") { const runId = process.argv[3]; const kind = valueAfter("--kind") as JudgmentKind | undefined; @@ -565,18 +599,21 @@ try { const events = await store.listEvents(); const runs = await store.listRuns(); const projection = await rebuildRootActivity(store); + const review = await projectReviewQueue(store); print({ database: path.join(projectRoot, ".thoughtstream", "state", "jazz.sqlite"), events: events.length, runs: runs.length, completedRuns: runs.filter((run) => run.status === "completed").length, failedRuns: runs.filter((run) => run.status === "failed").length, + review: review.counts, projection, }); } else if (command === "serve") { const host = valueAfter("--host") ?? "127.0.0.1"; const port = Number(valueAfter("--port") ?? "4317"); - const server = await startInspectorServer(store, { host, port }); + const reviewCapability = decodeReviewCapability(process.env.THOUGHTSTREAM_REVIEW_CAPABILITY_B64); + const server = await startInspectorServer(store, { host, port, ...(reviewCapability ? { reviewCapability } : {}) }); process.stdout.write(`thought stream inspector: http://${host}:${port}\n`); await waitForShutdown(server); } else { diff --git a/src/events/registry.ts b/src/events/registry.ts index c8d0997..d94f5b7 100644 --- a/src/events/registry.ts +++ b/src/events/registry.ts @@ -1,13 +1,33 @@ import { z } from "zod"; -import { observationOutputSchema } from "../agents/output-contracts.js"; +import { + CONCEPTUALIZATION_OUTPUT_CONTRACT, + conceptualizationOutputSchema, + createOutputContractRegistry, +} from "../agents/output-contracts.js"; import type { JsonObject, JsonValue } from "../core/json.js"; +import { modelAdapterIdentitySchema } from "../adapters/model-adapters.js"; import { OPERATIONAL_INCIDENT_EVENT_TYPE, operationalIncidentPayloadSchema } from "../incidents/types.js"; +import { + REVIEW_DECISION_EVENT_TYPE, + REVIEW_DECISION_SCHEMA_VERSION, + REVIEW_ITEM_EVENT_TYPE, + REVIEW_ITEM_SCHEMA_VERSION, + REVIEW_PROMPT_EVENT_TYPE, + REVIEW_PROMPT_SCHEMA_VERSION, + REVIEW_RESPONSE_EVENT_TYPE, + reviewDecisionPayloadSchema, + reviewItemPayloadSchema, + reviewPromptPayloadSchema, + reviewResponseEventPayloadSchema, +} from "../review/types.js"; +import type { ThoughtEvent } from "./types.js"; export interface RegisteredEventType { type: string; schemaVersion: number; description: string; payload: z.ZodType; + minimumPrivacy?: ThoughtEvent["privacy"] | undefined; } export class EventRegistry { @@ -26,6 +46,20 @@ export class EventRegistry { return definition.payload.parse(payload); } + validateEvent( + type: string, + schemaVersion: number, + privacy: ThoughtEvent["privacy"], + payload: JsonObject, + ): JsonObject { + const definition = this.types.get(registryKey(type, schemaVersion)); + if (!definition) throw new Error(`Unknown event type/version: ${type}@${schemaVersion}`); + if (definition.minimumPrivacy && privacyRank(privacy) < privacyRank(definition.minimumPrivacy)) { + throw new Error(`${type}@${schemaVersion} requires ${definition.minimumPrivacy} or stricter privacy`); + } + return definition.payload.parse(payload); + } + has(type: string, schemaVersion: number): boolean { return this.types.has(registryKey(type, schemaVersion)); } @@ -124,6 +158,10 @@ const repairModelIdentitySchema = z.object({ id: z.string().min(1).max(500), checkpointRevision: z.string().min(1).max(500).optional(), adapterRevision: z.string().min(1).max(500).optional(), + executionAdapterRevision: z.string().min(1).max(500).optional(), + modelAdapter: modelAdapterIdentitySchema.optional(), + adapterCatalogDigest: sha256Schema.optional(), + adapterCatalogGeneration: z.number().int().positive().optional(), }).strict(); const repairFailureSchema = z.object({ code: z.enum(["invalid-final-output", "semantic-output-invalid"]), @@ -172,8 +210,69 @@ const correctionProposalPayload = z.object({ originalTriggerEventId: z.string().min(1), sourceRootEventId: z.string().min(1), outputContract: outputContractIdentitySchema, - structuredOutput: observationOutputSchema, -}).strict() as unknown as z.ZodType; + structuredOutput: objectPayload, + originalModel: repairModelIdentitySchema, + repairModel: repairModelIdentitySchema, + executionAdapterRevision: z.string().min(1).max(500).optional(), + modelAdapter: modelAdapterIdentitySchema.optional(), + adapterCatalogDigest: sha256Schema.optional(), + adapterCatalogGeneration: z.number().int().positive().optional(), +}).strict().superRefine((value, context) => { + try { + createOutputContractRegistry().canonicalize(value.outputContract, value.structuredOutput); + } catch { + context.addIssue({ + code: "custom", + path: ["structuredOutput"], + message: "Correction proposal must satisfy its registered output contract", + }); + } +}) as unknown as z.ZodType; + +const conceptualizationGraphPayload = z.object({ + runId: z.string().min(1), + executionKey: z.string().min(1), + inputEventId: z.string().min(1), + inputSourceSequence: z.number().int().positive(), + summary: z.string().min(1).max(2_000), + confidence: z.number().min(0).max(1), + outputContract: outputContractIdentitySchema, + structuredOutput: conceptualizationOutputSchema, + model: objectPayload.optional(), + executionAdapterRevision: z.string().min(1).max(500).optional(), + modelAdapter: modelAdapterIdentitySchema.optional(), + adapterCatalogDigest: sha256Schema.optional(), + adapterCatalogGeneration: z.number().int().positive().optional(), +}).strict().superRefine((value, context) => { + const expectedContract = CONCEPTUALIZATION_OUTPUT_CONTRACT.identity; + if ( + value.outputContract.id !== expectedContract.id + || value.outputContract.version !== expectedContract.version + || value.outputContract.sha256 !== expectedContract.sha256 + ) { + context.addIssue({ + code: "custom", + path: ["outputContract"], + message: "Concept graph must name the canonical conceptualization contract identity", + }); + } else { + try { + createOutputContractRegistry().canonicalize(value.outputContract, value.structuredOutput); + } catch { + context.addIssue({ + code: "custom", + path: ["structuredOutput"], + message: "Concept graph must satisfy its canonical conceptualization contract", + }); + } + } + if (value.summary !== value.structuredOutput.summary) { + context.addIssue({ code: "custom", path: ["summary"], message: "Summary must equal canonical structured output" }); + } + if (value.confidence !== value.structuredOutput.confidence) { + context.addIssue({ code: "custom", path: ["confidence"], message: "Confidence must equal canonical structured output" }); + } +}) as unknown as z.ZodType; const legacyJudgmentPayload = z.object({ runId: z.string().min(1), @@ -293,6 +392,31 @@ export function createDefaultRegistry(): EventRegistry { payload: judgmentRetractionPayload, }); + registry.register({ + type: REVIEW_PROMPT_EVENT_TYPE, + schemaVersion: REVIEW_PROMPT_SCHEMA_VERSION, + description: "Complete evidence-bearing prompt for one private review campaign", + payload: reviewPromptPayloadSchema, + }); + registry.register({ + type: REVIEW_RESPONSE_EVENT_TYPE, + schemaVersion: 1, + description: "Canonical candidate response generated for one review prompt", + payload: reviewResponseEventPayloadSchema, + }); + registry.register({ + type: REVIEW_ITEM_EVENT_TYPE, + schemaVersion: REVIEW_ITEM_SCHEMA_VERSION, + description: "Immutable blinded pair of exact completed review candidate runs", + payload: reviewItemPayloadSchema, + }); + registry.register({ + type: REVIEW_DECISION_EVENT_TYPE, + schemaVersion: REVIEW_DECISION_SCHEMA_VERSION, + description: "Append-only human judgeability, preference, correction, tie, or skip decision", + payload: reviewDecisionPayloadSchema, + }); + registry.register({ type: "stream.thought.derived.event.batch", schemaVersion: 1, @@ -300,6 +424,14 @@ export function createDefaultRegistry(): EventRegistry { payload: batchPayloadSchema, }); + registry.register({ + type: "stream.thought.derived.concept.graph", + schemaVersion: 1, + description: "Validated private concept graph derived from one bounded source event", + payload: conceptualizationGraphPayload, + minimumPrivacy: "private", + }); + registry.register({ type: OPERATIONAL_INCIDENT_EVENT_TYPE, schemaVersion: 1, @@ -345,3 +477,7 @@ export function createDefaultRegistry(): EventRegistry { function registryKey(type: string, schemaVersion: number): string { return `${type}@${schemaVersion}`; } + +function privacyRank(privacy: ThoughtEvent["privacy"]): number { + return privacy === "public-source" ? 0 : privacy === "private" ? 1 : 2; +} diff --git a/src/jazz/store.ts b/src/jazz/store.ts index 007ee66..6290dc7 100644 --- a/src/jazz/store.ts +++ b/src/jazz/store.ts @@ -4,6 +4,7 @@ import path from "node:path"; import { RowChangeKind, type Db, type TransactionScope } from "jazz-tools"; import { createJazzContext, type JazzContext } from "jazz-tools/backend"; import { canonicalJson, hashJson, parseJsonObject, sha256, type JsonObject, type JsonValue } from "../core/json.js"; +import { modelAdapterIdentitySchema, type ModelAdapterIdentity } from "../adapters/model-adapters.js"; import { createDefaultRegistry, type EventRegistry } from "../events/registry.js"; import type { AppendEventResult, @@ -54,6 +55,7 @@ type JazzTable = { where(input: Record): { limit(count: number) // JazzThoughtStore, but all of them must share the same per-account reservation queue. // Independent processes still enforce independently and may briefly overshoot an aggregate cap. const inferenceAccountQueues = new Map>(); +const producerSourceQueues = new Map>(); export class JazzThoughtStore { private readonly context: JazzContext; @@ -100,6 +102,14 @@ export class JazzThoughtStore { async appendProducerBatch(candidates: EventCandidate[], cursor?: SourceCursor): Promise { if (candidates.length === 0 && !cursor) return { events: [], inserted: [], unchanged: [] }; const source = candidates[0]?.source ?? cursor!.source; + return this.withProducerSourceLock(source, () => this.appendProducerBatchUnlocked(source, candidates, cursor)); + } + + private async appendProducerBatchUnlocked( + source: string, + candidates: EventCandidate[], + cursor?: SourceCursor, + ): Promise { if (candidates.some((candidate) => candidate.source !== source)) { throw new Error("A producer transaction may settle events for only one source"); } @@ -175,6 +185,21 @@ export class JazzThoughtStore { return result.value; } + private async withProducerSourceLock(source: string, operation: () => Promise): Promise { + const previous = producerSourceQueues.get(source) ?? Promise.resolve(); + let release!: () => void; + const lock = new Promise((resolve) => { release = resolve; }); + const queued = previous.then(() => lock); + producerSourceQueues.set(source, queued); + await previous; + try { + return await operation(); + } finally { + release(); + if (producerSourceQueues.get(source) === queued) producerSourceQueues.delete(source); + } + } + async settleDerivedBatch(settlement: { candidate: EventCandidate; progress: ConsumerProgress; @@ -255,7 +280,7 @@ export class JazzThoughtStore { } else { const sequence = Number(sourceSnapshot?.lastSequence ?? 0) + 1; event = buildEvent(candidate, id, sequence, observedAt, this.runtimeRevision, this.registry); - tx.insert(thoughtstreamApp.events, eventToJazz(event), { id: jazzRowId("batch-event-v1", id) }); + tx.insert(thoughtstreamApp.events, eventToJazz(event), { id: jazzRowId("batch-event-v2", id) }); const sourceData = { key: outputSource, kind: candidate.sourceKind, @@ -266,7 +291,7 @@ export class JazzThoughtStore { updatedAt: observedAt, }; if (sourceSnapshot) tx.update(thoughtstreamApp.sources, String(sourceSnapshot.id), sourceData); - else tx.insert(thoughtstreamApp.sources, sourceData, { id: jazzRowId("batch-source-v1", outputSource) }); + else tx.insert(thoughtstreamApp.sources, sourceData, { id: jazzRowId("batch-source-v2", outputSource) }); inserted = true; } await upsertProgressInTransaction(tx, progress, progressSnapshot); @@ -752,10 +777,10 @@ export class JazzThoughtStore { } sequence += 1; const event = buildEvent(candidate, id, sequence, settledAt, this.runtimeRevision, this.registry); - // Jazz may retain an allocated object id after an earlier transaction - // fails before linking the row. Keep the domain event id stable while - // using a revisioned storage-object namespace for consumer settlement. - tx.insert(thoughtstreamApp.events, eventToJazz(event), { id: jazzRowId("consumer-event-v2", id) }); + // v2 and v3 may remain allocated but unlinked after a failed Jazz + // transaction. A failed insert also poisons the active batch, so do not + // probe old namespaces inside this transaction before selecting v4. + tx.insert(thoughtstreamApp.events, eventToJazz(event), { id: jazzRowId("consumer-event-v4", id) }); events.push(event); } const sourceData = { @@ -768,7 +793,7 @@ export class JazzThoughtStore { updatedAt: settledAt, }; if (sourceRow) tx.update(thoughtstreamApp.sources, String(sourceRow.id), sourceData); - else tx.insert(thoughtstreamApp.sources, sourceData, { id: jazzRowId("consumer-source-v2", outputSource) }); + else tx.insert(thoughtstreamApp.sources, sourceData, { id: jazzRowId("consumer-source-v3", outputSource) }); tx.update(thoughtstreamApp.runs, String(runSnapshot.id), runToJazz(run)); if (progress) await upsertProgressInTransaction(tx, progress, progressSnapshot); @@ -1306,8 +1331,16 @@ function runFromJazz(row: Record): AgentRun { agentId: String(row.agentKey), agentVersion: Number(row.agentVersion), status: String(row.status) as AgentRun["status"], inputEventIds, outputEventIds, attempt: Number(row.attempt), provider: String(row.provider), model: String(row.model), + ...(persistedResult.privacy ? { privacy: persistedResult.privacy } : {}), ...(persistedResult.checkpointRevision ? { checkpointRevision: persistedResult.checkpointRevision } : {}), - ...(String(row.adapterRevision ?? "") ? { adapterRevision: String(row.adapterRevision) } : {}), + ...(String(row.adapterRevision ?? "") ? { + adapterRevision: String(row.adapterRevision), + executionAdapterRevision: String(row.adapterRevision), + } : {}), + ...(persistedResult.executionAdapterRevision ? { executionAdapterRevision: persistedResult.executionAdapterRevision } : {}), + ...(persistedResult.modelAdapter ? { modelAdapter: persistedResult.modelAdapter } : {}), + ...(persistedResult.adapterCatalogDigest ? { adapterCatalogDigest: persistedResult.adapterCatalogDigest } : {}), + ...(persistedResult.adapterCatalogGeneration ? { adapterCatalogGeneration: persistedResult.adapterCatalogGeneration } : {}), promptHash: String(row.promptHash), contextManifest: parseJsonObject(String(row.contextManifestJson)), ...(String(row.accountingReservationKey ?? "") ? { accountingReservationId: String(row.accountingReservationKey) } : {}), ...(persistedResult.result ? { result: persistedResult.result } : {}), @@ -1322,25 +1355,68 @@ function runFromJazz(row: Record): AgentRun { const RUN_RESULT_ENVELOPE_KEY = "__thoughtstreamRunPersistenceEnvelopeV1"; function encodeRunResult(run: AgentRun): string { - if (!run.checkpointRevision) return run.result ? canonicalJson(run.result) : ""; + const hasPrivateEnvelope = Boolean( + run.checkpointRevision + || run.privacy + || run.executionAdapterRevision + || run.modelAdapter + || run.adapterCatalogDigest + || run.adapterCatalogGeneration, + ); + if (!hasPrivateEnvelope) return run.result ? canonicalJson(run.result) : ""; return canonicalJson({ [RUN_RESULT_ENVELOPE_KEY]: { - checkpointRevision: run.checkpointRevision, + ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), + ...(run.privacy ? { privacy: run.privacy } : {}), + ...(run.executionAdapterRevision ? { executionAdapterRevision: run.executionAdapterRevision } : {}), + ...(run.modelAdapter ? { modelAdapter: run.modelAdapter as unknown as JsonObject } : {}), + ...(run.adapterCatalogDigest ? { adapterCatalogDigest: run.adapterCatalogDigest } : {}), + ...(run.adapterCatalogGeneration ? { adapterCatalogGeneration: run.adapterCatalogGeneration } : {}), ...(run.result ? { result: run.result } : {}), }, }); } -function decodeRunResult(value: string): { result?: JsonObject; checkpointRevision?: string } { +function decodeRunResult(value: string): { + result?: JsonObject; + privacy?: ThoughtEvent["privacy"]; + checkpointRevision?: string; + executionAdapterRevision?: string; + modelAdapter?: ModelAdapterIdentity; + adapterCatalogDigest?: string; + adapterCatalogGeneration?: number; +} { if (!value) return {}; const parsed = parseJsonObject(value); const envelope = parsed[RUN_RESULT_ENVELOPE_KEY]; if (!envelope || typeof envelope !== "object" || Array.isArray(envelope)) return { result: parsed }; const checkpointRevision = typeof envelope.checkpointRevision === "string" ? envelope.checkpointRevision : undefined; + const privacy = ["public-source", "private", "sensitive"].includes(String(envelope.privacy)) + ? envelope.privacy as ThoughtEvent["privacy"] + : undefined; + const executionAdapterRevision = typeof envelope.executionAdapterRevision === "string" + ? envelope.executionAdapterRevision + : undefined; + const modelAdapter = envelope.modelAdapter === undefined + ? undefined + : modelAdapterIdentitySchema.parse(envelope.modelAdapter); + const adapterCatalogDigest = typeof envelope.adapterCatalogDigest === "string" + ? envelope.adapterCatalogDigest + : undefined; + const adapterCatalogGeneration = typeof envelope.adapterCatalogGeneration === "number" + && Number.isSafeInteger(envelope.adapterCatalogGeneration) + && envelope.adapterCatalogGeneration > 0 + ? envelope.adapterCatalogGeneration + : undefined; const result = envelope.result; return { ...(result && typeof result === "object" && !Array.isArray(result) ? { result: result as JsonObject } : {}), + ...(privacy ? { privacy } : {}), ...(checkpointRevision ? { checkpointRevision } : {}), + ...(executionAdapterRevision ? { executionAdapterRevision } : {}), + ...(modelAdapter ? { modelAdapter } : {}), + ...(adapterCatalogDigest ? { adapterCatalogDigest } : {}), + ...(adapterCatalogGeneration ? { adapterCatalogGeneration } : {}), }; } @@ -1375,7 +1451,12 @@ function buildEvent( runtimeRevision: string, registry: EventRegistry, ): ThoughtEvent { - const payload = registry.validate(candidate.type, candidate.schemaVersion, candidate.payload); + const payload = registry.validateEvent( + candidate.type, + candidate.schemaVersion, + candidate.privacy, + candidate.payload, + ); return { id, sourceSequence, @@ -1492,6 +1573,6 @@ async function upsertProgressInTransaction( } tx.update(thoughtstreamApp.consumerProgress, String(row.id), data); } else { - tx.insert(thoughtstreamApp.consumerProgress, data, { id: jazzRowId("consumer-progress-v2", progress.id) }); + tx.insert(thoughtstreamApp.consumerProgress, data, { id: jazzRowId("consumer-progress-v3", progress.id) }); } } diff --git a/src/projections/activity.ts b/src/projections/activity.ts index 6a872ae..7923169 100644 --- a/src/projections/activity.ts +++ b/src/projections/activity.ts @@ -6,6 +6,7 @@ import type { AgentRun } from "../store/types.js"; export interface ActivityConsumerRun { id: string; agentId: string; + agentVersion: number; status: AgentRun["status"]; kind: "rule" | "model"; description: string; @@ -48,7 +49,10 @@ export async function buildRootActivity(store: JazzThoughtStore, limit = 100): P const rootIds = new Set(run.inputEventIds.map((id) => eventsById.get(id)?.rootEventId).filter((id): id is string => Boolean(id))); for (const rootId of rootIds) runsByRoot.set(rootId, [...(runsByRoot.get(rootId) ?? []), run]); } - const roots = events.filter((event) => event.id === event.rootEventId); + const roots = events.filter((event) => ( + event.id === event.rootEventId + && event.type.startsWith("stream.thought.source.") + )); const items = roots.slice(-limit).reverse().map((event) => ({ id: event.id, type: event.type, @@ -63,6 +67,7 @@ export async function buildRootActivity(store: JazzThoughtStore, limit = 100): P consumerRuns: (runsByRoot.get(event.id) ?? []).map((run) => ({ id: run.id, agentId: run.agentId, + agentVersion: run.agentVersion, status: run.status, kind: runKind(run), description: describeRunResult(run, run.inputEventIds.map((id) => eventsById.get(id)).filter((input): input is ThoughtEvent => Boolean(input))), @@ -102,14 +107,11 @@ export function describeRunResult(run: AgentRun, inputs: ThoughtEvent[] = []): s const resultSummary = typeof run.result.summary === "string" ? run.result.summary : "Completed without a summary"; const firstSentence = classification ? `Classified ${path ?? "the event"} as ${classification}${importance ? ` (${importance} importance)` : ""}.` - : `Produced: ${resultSummary}.`; + : `Produced: ${resultSummary}`; const recommendation = isObject(run.result.recommendation) && typeof run.result.recommendation.proposedAction === "string" ? ` Proposed: ${run.result.recommendation.proposedAction} The proposal was not executed.` - : " No recommendation was produced."; - const externalActions = run.contextManifest.externalActions === false - ? " External actions were disabled." : ""; - return `${firstSentence}${recommendation}${externalActions}`; + return `${firstSentence}${recommendation}`; } function summarizeRoot(event: ThoughtEvent): string { diff --git a/src/projections/effective-output.ts b/src/projections/effective-output.ts index 462fcf1..f0c5956 100644 --- a/src/projections/effective-output.ts +++ b/src/projections/effective-output.ts @@ -1,8 +1,8 @@ import { + canonicalStructuredOutput, createOutputContractRegistry, outputContractIdentityJson, parseOutputContractIdentity, - structuredOutputJson, type OutputContractIdentity, } from "../agents/output-contracts.js"; import { CORRECTION_PROPOSAL_EVENT_TYPE, REPAIR_REQUEST_EVENT_TYPE } from "../agents/repairs.js"; @@ -11,6 +11,7 @@ import { stableKey } from "../core/ids.js"; import type { ThoughtEvent } from "../events/types.js"; import type { JazzThoughtStore } from "../jazz/store.js"; import type { AgentRun, Projection } from "../store/types.js"; +import { joinPrivacy, runPrivacy } from "../security/privacy.js"; export const EFFECTIVE_OUTPUT_PROJECTION_VERSION = 1; @@ -104,7 +105,7 @@ export async function rebuildEffectiveOutput( judgment, proposal, repairRun, - output: structuredOutputJson(registry.validate(contract, candidate)), + output: canonicalStructuredOutput(registry, contract, candidate), }); } catch { // A corrupt historical judgment cannot become effective merely because it exists. @@ -121,6 +122,15 @@ export async function rebuildEffectiveOutput( originalRunId, status: "repair", sourceRootEventId: originalSource.rootEventId, + originalModel: executionProvenance(originalRun), + repairModel: executionProvenance(selectedRepair.repairRun), + privacy: joinPrivacy( + originalSource.privacy, + runPrivacy(originalRun), + selectedRepair.proposal.privacy, + runPrivacy(selectedRepair.repairRun), + selectedRepair.judgment.privacy, + ), outputContract: outputContractIdentityJson(contract), structuredOutput: selectedRepair.output, outputEventId: selectedRepair.proposal.id, @@ -137,6 +147,8 @@ export async function rebuildEffectiveOutput( originalRunId, status: "original", sourceRootEventId: originalSource.rootEventId, + originalModel: executionProvenance(originalRun), + privacy: joinPrivacy(originalSource.privacy, runPrivacy(originalRun), originalOutput.event.privacy), outputContract: outputContractIdentityJson(contract), structuredOutput: originalOutput.output, outputEventId: originalOutput.event.id, @@ -150,6 +162,8 @@ export async function rebuildEffectiveOutput( originalRunId, status: "unresolved", sourceRootEventId: originalSource.rootEventId, + originalModel: executionProvenance(originalRun), + privacy: joinPrivacy(originalSource.privacy, runPrivacy(originalRun), failed?.privacy), outputContract: outputContractIdentityJson(contract), }; lastEventId = failed?.id ?? originalSource.id; @@ -167,6 +181,18 @@ export async function rebuildEffectiveOutput( return projection; } +function executionProvenance(run: AgentRun): JsonObject { + return { + provider: run.provider, + id: run.model, + ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), + ...(run.executionAdapterRevision ? { executionAdapterRevision: run.executionAdapterRevision } : {}), + ...(run.modelAdapter ? { modelAdapter: run.modelAdapter as unknown as JsonObject } : {}), + ...(run.adapterCatalogDigest ? { adapterCatalogDigest: run.adapterCatalogDigest } : {}), + ...(run.adapterCatalogGeneration ? { adapterCatalogGeneration: run.adapterCatalogGeneration } : {}), + }; +} + export async function rebuildAllEffectiveOutputs(store: JazzThoughtStore): Promise { const runs = await store.listRuns(); const originals = runs.filter((run) => run.contextManifest.agentRole !== "repair"); @@ -191,7 +217,7 @@ async function validatedOriginalOutput( if (!event || !sameContract(event.payload.outputContract, contract)) return undefined; const structuredOutput = objectField(event.payload.structuredOutput); if (!structuredOutput) return undefined; - const output = structuredOutputJson(createOutputContractRegistry().validate(contract, structuredOutput)); + const output = canonicalStructuredOutput(createOutputContractRegistry(), contract, structuredOutput); return { event, output }; } diff --git a/src/review/review.ts b/src/review/review.ts new file mode 100644 index 0000000..83522cc --- /dev/null +++ b/src/review/review.ts @@ -0,0 +1,707 @@ +import { timingSafeEqual } from "node:crypto"; +import { + REVIEW_RESPONSE_OUTPUT_CONTRACT, + canonicalStructuredOutput, + createOutputContractRegistry, + outputContractIdentityJson, + parseOutputContractIdentity, +} from "../agents/output-contracts.js"; +import { canonicalJson, sha256, type JsonObject } from "../core/json.js"; +import { stableKey } from "../core/ids.js"; +import type { PrivacyClass, ThoughtEvent } from "../events/types.js"; +import type { JazzThoughtStore } from "../jazz/store.js"; +import { joinPrivacy, runPrivacy } from "../security/privacy.js"; +import type { AgentRun } from "../store/types.js"; +import { + REVIEW_DECISION_EVENT_TYPE, + REVIEW_DECISION_SCHEMA_VERSION, + REVIEW_ITEM_EVENT_TYPE, + REVIEW_ITEM_SCHEMA_VERSION, + REVIEW_PROMPT_EVENT_TYPE, + REVIEW_PROMPT_SCHEMA_VERSION, + REVIEW_RESPONSE_EVENT_TYPE, + REVIEW_RUNTIME, + reviewDecisionPayloadSchema, + reviewDecisionSubmissionSchema, + reviewItemPayloadSchema, + reviewPromptPayloadSchema, + type ReviewDisposition, +} from "./types.js"; + +export interface AppendReviewPromptInput { + source?: string | undefined; + externalId: string; + occurredAt?: string | undefined; + actor?: string | undefined; + correlationId?: string | undefined; + privacy: PrivacyClass; + payload: JsonObject; +} + +export interface CreateReviewItemInput { + promptEventId: string; + candidateRunIds: [string, string]; + actor?: string | undefined; + source?: string | undefined; +} + +export interface RecordReviewDecisionInput { + reviewItemEventId: string; + disposition: ReviewDisposition; + preferredRunId?: string | undefined; + preferenceStrength?: "slight" | "strong" | undefined; + replacementResponse?: string | undefined; + confidence?: "low" | "medium" | "high" | undefined; + reasonCodes?: string[] | undefined; + responseTags?: string[] | undefined; + notes?: string | undefined; + trainingEligible: boolean; + submissionId: string; + supersedesDecisionEventId?: string | undefined; + actor?: string | undefined; + source?: string | undefined; +} + +export interface ActiveReviewDecisionSet { + active: ThoughtEvent[]; + byItem: Map; + inactiveIds: Set; + historyCountByItem: Map; +} + +export interface ActiveReviewTrainingRecord { + decision: ThoughtEvent; + item: ThoughtEvent; + prompt: ThoughtEvent; + runs: [AgentRun, AgentRun]; + outputs: [ThoughtEvent, ThoughtEvent]; +} + +export interface ReviewQueueCandidate { + label: "A" | "B"; + response: string; + provenance?: { + agentId: string; + agentVersion: number; + provider: string; + model: string; + checkpointRevision?: string | undefined; + executionAdapterRevision?: string | undefined; + modelAdapter?: JsonObject | undefined; + adapterCatalogDigest?: string | undefined; + adapterCatalogGeneration?: number | undefined; + } | undefined; +} + +export interface ReviewQueueItem { + id: string; + createdAt: string; + privacy: PrivacyClass; + campaign: JsonObject; + prompt: string; + evidence?: string | undefined; + criterion: JsonObject; + candidates: [ReviewQueueCandidate, ReviewQueueCandidate]; + decision?: JsonObject | undefined; + decisionHistoryCount: number; + trainingExportPreauthorized: boolean; +} + +export class ReviewDecisionConflictError extends Error { + constructor(message = "Review decision changed; reload the item before submitting") { + super(message); + this.name = "ReviewDecisionConflictError"; + } +} + +export async function recordBrowserReviewDecision( + store: JazzThoughtStore, + reviewItemEventId: string, + body: unknown, +): Promise { + const submission = reviewDecisionSubmissionSchema.parse(body); + const item = await requireReviewItem(store, reviewItemEventId); + const displayOrder = tupleStringField(item.payload.displayOrder, "displayOrder"); + const preferredRunId = submission.preferredCandidate + ? displayOrder[submission.preferredCandidate === "A" ? 0 : 1] + : undefined; + return recordReviewDecision(store, { + reviewItemEventId, + disposition: submission.disposition, + ...(preferredRunId ? { preferredRunId } : {}), + ...(submission.preferenceStrength ? { preferenceStrength: submission.preferenceStrength } : {}), + ...(submission.replacementResponse ? { replacementResponse: submission.replacementResponse } : {}), + ...(submission.confidence ? { confidence: submission.confidence } : {}), + reasonCodes: submission.reasonCodes, + responseTags: submission.responseTags, + ...(submission.notes ? { notes: submission.notes } : {}), + trainingEligible: submission.trainingEligible, + submissionId: submission.submissionId, + ...(submission.supersedesDecisionEventId ? { supersedesDecisionEventId: submission.supersedesDecisionEventId } : {}), + }); +} + +export async function appendReviewPrompt( + store: JazzThoughtStore, + input: AppendReviewPromptInput, +): Promise { + const payload = reviewPromptPayloadSchema.parse(input.payload) as JsonObject; + if (payload.externalExportEligible === true && input.privacy !== "public-source") { + throw new Error("Only public-source review prompts may be preauthorized for external training export"); + } + const result = await store.appendEvent({ + type: REVIEW_PROMPT_EVENT_TYPE, + schemaVersion: REVIEW_PROMPT_SCHEMA_VERSION, + source: input.source ?? "review:prompt", + sourceKind: "system", + externalId: input.externalId, + idempotencyKey: stableKey("review-prompt", input.source ?? "review:prompt", input.externalId), + occurredAt: input.occurredAt ?? new Date().toISOString(), + actor: input.actor ?? "operator:local", + correlationId: input.correlationId ?? stringField(objectField(payload.campaign)?.id, "campaign.id"), + privacy: input.privacy, + payload, + createdByRuntime: REVIEW_RUNTIME, + }); + return result.event; +} + +export async function createReviewItem( + store: JazzThoughtStore, + input: CreateReviewItemInput, +): Promise { + if (input.candidateRunIds[0] === input.candidateRunIds[1]) { + throw new Error("Review candidates must be distinct runs"); + } + const prompt = await requireReviewPrompt(store, input.promptEventId); + const promptPayload = reviewPromptPayloadSchema.parse(prompt.payload) as JsonObject; + const allowedAgentIds = new Set(stringArrayField(promptPayload.candidateAgentIds, "candidateAgentIds")); + const runs = await Promise.all(input.candidateRunIds.map((id) => requireReviewRun(store, id, prompt, allowedAgentIds))) as [AgentRun, AgentRun]; + if (runs[0].agentId === runs[1].agentId) throw new Error("Review candidates must come from distinct declared agents"); + const outputs = await Promise.all(runs.map((run) => requireReviewOutput(store, run, prompt))) as [ThoughtEvent, ThoughtEvent]; + const frozen = runs.map((run, index) => ({ run, output: outputs[index]! })) + .sort((left, right) => left.run.id.localeCompare(right.run.id)) as [ + { run: AgentRun; output: ThoughtEvent }, + { run: AgentRun; output: ThoughtEvent }, + ]; + const sortedRunIds: [string, string] = [frozen[0].run.id, frozen[1].run.id]; + const orderSeed = sha256(canonicalJson({ promptEventId: prompt.id, candidateRunIds: sortedRunIds })); + const displayOrder: [string, string] = Number.parseInt(orderSeed.slice(0, 2), 16) % 2 === 0 + ? [sortedRunIds[0], sortedRunIds[1]] + : [sortedRunIds[1], sortedRunIds[0]]; + const criterion = objectField(promptPayload.criterion)!; + const idempotencyKey = stableKey( + "review-item", + prompt.id, + stringField(criterion.id, "criterion.id"), + String(numberField(criterion.version, "criterion.version")), + ...sortedRunIds, + ); + const result = await store.appendEvent({ + type: REVIEW_ITEM_EVENT_TYPE, + schemaVersion: REVIEW_ITEM_SCHEMA_VERSION, + source: input.source ?? "review:item", + sourceKind: "system", + externalId: idempotencyKey, + idempotencyKey, + occurredAt: new Date().toISOString(), + actor: input.actor ?? "operator:local", + rootEventId: prompt.rootEventId, + parentEventId: prompt.id, + correlationId: stringField(objectField(promptPayload.campaign)?.id, "campaign.id"), + privacy: joinPrivacy(prompt.privacy, ...runs.map((run) => runPrivacy(run)), ...outputs.map((event) => event.privacy)), + payload: { + promptEventId: prompt.id, + candidateRunIds: sortedRunIds, + candidateOutputEventIds: [frozen[0].output.id, frozen[1].output.id], + candidateRunReceiptDigests: [ + reviewRunReceiptDigest(frozen[0].run, frozen[0].output.id), + reviewRunReceiptDigest(frozen[1].run, frozen[1].output.id), + ], + displayOrder, + criterion: { + id: stringField(criterion.id, "criterion.id"), + version: numberField(criterion.version, "criterion.version"), + }, + outputContract: outputContractIdentityJson(REVIEW_RESPONSE_OUTPUT_CONTRACT.identity), + }, + createdByRuntime: REVIEW_RUNTIME, + }); + return result.event; +} + +export async function recordReviewDecision( + store: JazzThoughtStore, + input: RecordReviewDecisionInput, +): Promise { + const item = await requireReviewItem(store, input.reviewItemEventId); + const prompt = await requireReviewPrompt(store, stringField(item.payload.promptEventId, "promptEventId")); + const promptPayload = reviewPromptPayloadSchema.parse(prompt.payload) as JsonObject; + const candidateRunIds = tupleStringField(item.payload.candidateRunIds, "candidateRunIds"); + const frozenCandidates = await requireFrozenReviewCandidates(store, item, prompt); + const runs = frozenCandidates.map((candidate) => candidate.run) as [AgentRun, AgentRun]; + const outputs = frozenCandidates.map((candidate) => candidate.output) as [ThoughtEvent, ThoughtEvent]; + const current = (await activeReviewDecisions(store)).byItem.get(item.id); + if (current?.payload.submissionId === input.submissionId) return current; + if (current?.id !== input.supersedesDecisionEventId) throw new ReviewDecisionConflictError(); + if (!current && input.supersedesDecisionEventId) throw new ReviewDecisionConflictError(); + + const criterion = objectField(promptPayload.criterion)!; + const reasonCodes = input.reasonCodes ?? []; + const responseTags = input.responseTags ?? []; + assertAllowed(reasonCodes, stringArrayField(criterion.reasonCodes, "criterion.reasonCodes"), "reason code"); + assertAllowed(responseTags, stringArrayField(criterion.responseTags, "criterion.responseTags"), "response tag"); + if (input.preferredRunId && !candidateRunIds.includes(input.preferredRunId)) { + throw new Error("Preferred run is not a candidate in this review item"); + } + if (input.replacementResponse) { + canonicalStructuredOutput( + createOutputContractRegistry(), + REVIEW_RESPONSE_OUTPUT_CONTRACT.identity, + { response: input.replacementResponse }, + ); + } + + const combinedPrivacy = joinPrivacy( + prompt.privacy, + item.privacy, + ...runs.map((run) => runPrivacy(run)), + ...outputs.map((event) => event.privacy), + current?.privacy, + ); + const externalExportEligible = input.trainingEligible + && promptPayload.externalExportEligible === true + && combinedPrivacy === "public-source"; + const payload = reviewDecisionPayloadSchema.parse({ + reviewItemEventId: item.id, + promptEventId: prompt.id, + disposition: input.disposition, + ...(input.preferredRunId ? { preferredRunId: input.preferredRunId } : {}), + ...(input.preferenceStrength ? { preferenceStrength: input.preferenceStrength } : {}), + ...(input.replacementResponse ? { replacementResponse: input.replacementResponse } : {}), + ...(input.confidence ? { confidence: input.confidence } : {}), + reasonCodes, + responseTags, + ...(input.notes?.trim() ? { notes: input.notes.trim() } : {}), + trainingEligible: input.trainingEligible, + externalExportEligible, + submissionId: input.submissionId, + ...(current ? { supersedesDecisionEventId: current.id } : {}), + }) as JsonObject; + const idempotencyKey = stableKey("review-decision", item.id, input.submissionId); + const result = await store.appendEvent({ + type: REVIEW_DECISION_EVENT_TYPE, + schemaVersion: REVIEW_DECISION_SCHEMA_VERSION, + source: input.source ?? "review:web", + sourceKind: "system", + externalId: idempotencyKey, + idempotencyKey, + occurredAt: new Date().toISOString(), + actor: input.actor ?? "operator:oauth", + rootEventId: prompt.rootEventId, + parentEventId: item.id, + correlationId: item.id, + privacy: combinedPrivacy, + payload, + createdByRuntime: REVIEW_RUNTIME, + }); + return result.event; +} + +export async function activeReviewDecisions(store: JazzThoughtStore): Promise { + const decisions = await store.listEvents({ types: [REVIEW_DECISION_EVENT_TYPE] }); + const byItemHistory = new Map(); + const explicitlyInactive = new Set(); + for (const decision of decisions) { + const itemId = optionalString(decision.payload.reviewItemEventId); + if (!itemId) continue; + const history = byItemHistory.get(itemId) ?? []; + history.push(decision); + byItemHistory.set(itemId, history); + const superseded = optionalString(decision.payload.supersedesDecisionEventId); + if (superseded) explicitlyInactive.add(superseded); + } + const byItem = new Map(); + const inactiveIds = new Set(explicitlyInactive); + const historyCountByItem = new Map(); + for (const [itemId, history] of byItemHistory) { + history.sort(compareEvents); + const activeLeaves = history.filter((event) => !explicitlyInactive.has(event.id)); + const selected = (activeLeaves.length > 0 ? activeLeaves : history).at(-1)!; + byItem.set(itemId, selected); + historyCountByItem.set(itemId, history.length); + for (const event of history) if (event.id !== selected.id) inactiveIds.add(event.id); + } + return { + active: [...byItem.values()].sort(compareEvents), + byItem, + inactiveIds, + historyCountByItem, + }; +} + +export async function projectReviewQueue(store: JazzThoughtStore): Promise<{ + items: ReviewQueueItem[]; + counts: { + total: number; + unresolved: number; + decided: number; + trainingEligible: number; + exportEligible: number; + byDisposition: Record; + }; +}> { + const [items, decisions] = await Promise.all([ + store.listEvents({ types: [REVIEW_ITEM_EVENT_TYPE] }), + activeReviewDecisions(store), + ]); + const projected: ReviewQueueItem[] = []; + for (const item of items) { + const prompt = await requireReviewPrompt(store, stringField(item.payload.promptEventId, "promptEventId")); + const promptPayload = reviewPromptPayloadSchema.parse(prompt.payload) as JsonObject; + const displayOrder = tupleStringField(item.payload.displayOrder, "displayOrder"); + const decision = decisions.byItem.get(item.id); + const frozen = await requireFrozenReviewCandidates(store, item, prompt); + const byRunId = new Map(frozen.map((candidate) => [candidate.run.id, candidate])); + const candidates = displayOrder.map((runId, index): ReviewQueueCandidate => { + const candidate = byRunId.get(runId); + if (!candidate) throw new Error("Review display order references an unfrozen candidate"); + const { run, output } = candidate; + const structured = objectField(output.payload.structuredOutput)!; + return { + label: index === 0 ? "A" : "B", + response: stringField(structured.response, "structuredOutput.response"), + ...(decision ? { provenance: reviewRunProvenance(run) } : {}), + }; + }) as [ReviewQueueCandidate, ReviewQueueCandidate]; + projected.push({ + id: item.id, + createdAt: item.occurredAt, + privacy: item.privacy, + campaign: objectField(promptPayload.campaign)!, + prompt: stringField(promptPayload.prompt, "prompt"), + ...(typeof promptPayload.evidence === "string" ? { evidence: promptPayload.evidence } : {}), + criterion: objectField(promptPayload.criterion)!, + candidates, + ...(decision ? { decision: reviewDecisionForBrowser(decision, displayOrder) } : {}), + decisionHistoryCount: decisions.historyCountByItem.get(item.id) ?? 0, + trainingExportPreauthorized: promptPayload.externalExportEligible === true && item.privacy === "public-source", + }); + } + projected.sort((left, right) => { + const leftDecided = left.decision ? 1 : 0; + const rightDecided = right.decision ? 1 : 0; + return leftDecided - rightDecided || right.createdAt.localeCompare(left.createdAt) || left.id.localeCompare(right.id); + }); + return { + items: projected, + counts: { + total: projected.length, + unresolved: projected.filter((item) => !item.decision).length, + decided: projected.filter((item) => item.decision).length, + trainingEligible: projected.filter((item) => item.decision?.trainingEligible === true).length, + exportEligible: projected.filter((item) => item.decision?.externalExportEligible === true).length, + byDisposition: projected.reduce>((counts, item) => { + if (item.decision) { + const disposition = stringField(item.decision.disposition, "disposition"); + counts[disposition] = (counts[disposition] ?? 0) + 1; + } + return counts; + }, {}), + }, + }; +} + +export async function activeReviewTrainingRecords( + store: JazzThoughtStore, +): Promise { + const decisions = await activeReviewDecisions(store); + const records: ActiveReviewTrainingRecord[] = []; + for (const decision of decisions.active) { + if (decision.payload.trainingEligible !== true || decision.payload.externalExportEligible !== true) continue; + if (decision.payload.disposition !== "prefer" && decision.payload.disposition !== "correct") continue; + const item = await requireReviewItem(store, stringField(decision.payload.reviewItemEventId, "reviewItemEventId")); + const prompt = await requireReviewPrompt(store, stringField(item.payload.promptEventId, "promptEventId")); + const promptPayload = reviewPromptPayloadSchema.parse(prompt.payload) as JsonObject; + if (promptPayload.externalExportEligible !== true) continue; + const displayOrder = tupleStringField(item.payload.displayOrder, "displayOrder"); + const frozen = await requireFrozenReviewCandidates(store, item, prompt); + const byRunId = new Map(frozen.map((candidate) => [candidate.run.id, candidate])); + const displayed = displayOrder.map((runId) => { + const candidate = byRunId.get(runId); + if (!candidate) throw new Error("Review display order references an unfrozen candidate"); + return candidate; + }) as [{ run: AgentRun; output: ThoughtEvent }, { run: AgentRun; output: ThoughtEvent }]; + const runs: [AgentRun, AgentRun] = [displayed[0].run, displayed[1].run]; + const outputs: [ThoughtEvent, ThoughtEvent] = [displayed[0].output, displayed[1].output]; + const privacy = joinPrivacy( + decision.privacy, + item.privacy, + prompt.privacy, + ...runs.map((run) => runPrivacy(run)), + ...outputs.map((output) => output.privacy), + ); + if (privacy !== "public-source") continue; + records.push({ decision, item, prompt, runs, outputs }); + } + return records; +} + +function reviewDecisionForBrowser( + decision: ThoughtEvent, + displayOrder: [string, string], +): JsonObject { + const preferredRunId = optionalString(decision.payload.preferredRunId); + const preferredCandidate = preferredRunId === displayOrder[0] + ? "A" + : preferredRunId === displayOrder[1] ? "B" : undefined; + return { + eventId: decision.id, + disposition: decision.payload.disposition!, + ...(preferredCandidate ? { preferredCandidate } : {}), + ...(decision.payload.preferenceStrength ? { preferenceStrength: decision.payload.preferenceStrength } : {}), + ...(decision.payload.replacementResponse ? { replacementResponse: decision.payload.replacementResponse } : {}), + ...(decision.payload.confidence ? { confidence: decision.payload.confidence } : {}), + reasonCodes: decision.payload.reasonCodes ?? [], + responseTags: decision.payload.responseTags ?? [], + trainingEligible: decision.payload.trainingEligible === true, + externalExportEligible: decision.payload.externalExportEligible === true, + } as JsonObject; +} + +function reviewRunProvenance(run: AgentRun): ReviewQueueCandidate["provenance"] { + return { + agentId: run.agentId, + agentVersion: run.agentVersion, + provider: run.provider, + model: run.model, + ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), + ...(run.executionAdapterRevision ? { executionAdapterRevision: run.executionAdapterRevision } : {}), + ...(run.modelAdapter ? { modelAdapter: JSON.parse(JSON.stringify(run.modelAdapter)) as JsonObject } : {}), + ...(run.adapterCatalogDigest ? { adapterCatalogDigest: run.adapterCatalogDigest } : {}), + ...(run.adapterCatalogGeneration ? { adapterCatalogGeneration: run.adapterCatalogGeneration } : {}), + }; +} + +async function requireReviewPrompt(store: JazzThoughtStore, id: string): Promise { + const event = await requireEvent(store, id); + if (event.type !== REVIEW_PROMPT_EVENT_TYPE || event.schemaVersion !== REVIEW_PROMPT_SCHEMA_VERSION) { + throw new Error(`Event is not a review prompt: ${id}`); + } + reviewPromptPayloadSchema.parse(event.payload); + return event; +} + +async function requireReviewItem(store: JazzThoughtStore, id: string): Promise { + const event = await requireEvent(store, id); + if (event.type !== REVIEW_ITEM_EVENT_TYPE || event.schemaVersion !== REVIEW_ITEM_SCHEMA_VERSION) { + throw new Error(`Event is not a review item: ${id}`); + } + reviewItemPayloadSchema.parse(event.payload); + return event; +} + +async function requireReviewRun( + store: JazzThoughtStore, + id: string, + prompt: ThoughtEvent, + allowedAgentIds: Set, +): Promise { + const run = await requireRun(store, id); + if (run.status !== "completed") throw new Error(`Review candidate run is not completed: ${id}`); + if (run.triggerEventId !== prompt.id) throw new Error("Review candidates must share the exact review prompt trigger"); + if (!allowedAgentIds.has(run.agentId)) throw new Error(`Review candidate agent is not declared by the prompt: ${run.agentId}`); + assertCompleteReviewContext(run, prompt); + const contract = parseOutputContractIdentity(run.contextManifest.outputContract); + if (!sameContract(contract, REVIEW_RESPONSE_OUTPUT_CONTRACT.identity)) { + throw new Error(`Review candidate has the wrong output contract: ${id}`); + } + return run; +} + +async function requireReviewOutput(store: JazzThoughtStore, run: AgentRun, prompt: ThoughtEvent): Promise { + if (run.status !== "completed" || run.outputEventIds.length !== 1) { + throw new Error(`Review candidate must have one completed output: ${run.id}`); + } + return requireReviewOutputById(store, run, run.outputEventIds[0]!, prompt); +} + +async function requireReviewOutputById( + store: JazzThoughtStore, + run: AgentRun, + outputEventId: string, + prompt: ThoughtEvent, +): Promise { + if (run.status !== "completed" || run.outputEventIds.length !== 1 || run.outputEventIds[0] !== outputEventId) { + throw new Error(`Review candidate output pointer changed after materialization: ${run.id}`); + } + const output = await requireEvent(store, outputEventId); + if ( + output.type !== REVIEW_RESPONSE_EVENT_TYPE + || output.payload.runId !== run.id + || output.payload.executionKey !== run.executionKey + || output.payload.inputEventId !== prompt.id + || output.payload.inputSourceSequence !== prompt.sourceSequence + || output.parentEventId !== prompt.id + || output.rootEventId !== prompt.rootEventId + ) { + throw new Error(`Review candidate output lineage is invalid: ${output.id}`); + } + const outputContract = parseOutputContractIdentity(output.payload.outputContract); + if (!sameContract(outputContract, REVIEW_RESPONSE_OUTPUT_CONTRACT.identity)) { + throw new Error(`Review candidate output contract is invalid: ${output.id}`); + } + canonicalStructuredOutput( + createOutputContractRegistry(), + REVIEW_RESPONSE_OUTPUT_CONTRACT.identity, + output.payload.structuredOutput, + ); + return output; +} + +async function requireFrozenReviewCandidates( + store: JazzThoughtStore, + item: ThoughtEvent, + prompt: ThoughtEvent, +): Promise<[ + { run: AgentRun; output: ThoughtEvent }, + { run: AgentRun; output: ThoughtEvent }, +]> { + const runIds = tupleStringField(item.payload.candidateRunIds, "candidateRunIds"); + const outputIds = tupleStringField(item.payload.candidateOutputEventIds, "candidateOutputEventIds"); + const receiptDigests = tupleStringField(item.payload.candidateRunReceiptDigests, "candidateRunReceiptDigests"); + const promptPayload = reviewPromptPayloadSchema.parse(prompt.payload) as JsonObject; + const allowedAgentIds = new Set(stringArrayField(promptPayload.candidateAgentIds, "candidateAgentIds")); + const candidates = await Promise.all(runIds.map(async (runId, index) => { + const run = await requireReviewRun(store, runId, prompt, allowedAgentIds); + const output = await requireReviewOutputById(store, run, outputIds[index]!, prompt); + const actualDigest = reviewRunReceiptDigest(run, output.id); + if (!safeEqual(actualDigest, receiptDigests[index]!)) { + throw new Error(`Review candidate receipt changed after materialization: ${run.id}`); + } + return { run, output }; + })); + if (candidates[0]!.run.agentId === candidates[1]!.run.agentId) { + throw new Error("Review candidates no longer resolve to distinct declared agents"); + } + return candidates as [ + { run: AgentRun; output: ThoughtEvent }, + { run: AgentRun; output: ThoughtEvent }, + ]; +} + +function assertCompleteReviewContext(run: AgentRun, prompt: ThoughtEvent): void { + const manifest = run.contextManifest; + if (canonicalJson(run.inputEventIds) !== canonicalJson([prompt.id])) { + throw new Error(`Review candidate must have exactly one prompt input: ${run.id}`); + } + if ( + canonicalJson(stringArrayField(manifest.inputEventIds, "contextManifest.inputEventIds")) !== canonicalJson([prompt.id]) + || canonicalJson(stringArrayField(manifest.includedEventIds, "contextManifest.includedEventIds")) !== canonicalJson([prompt.id]) + || stringArrayField(manifest.omittedEventIds, "contextManifest.omittedEventIds").length !== 0 + ) { + throw new Error(`Review candidate context does not contain exactly the frozen prompt: ${run.id}`); + } + const originalChars = numberField(manifest.sourceOriginalChars, "contextManifest.sourceOriginalChars"); + const includedChars = numberField(manifest.sourceIncludedChars, "contextManifest.sourceIncludedChars"); + if ( + manifest.truncated !== false + || originalChars !== includedChars + || manifest.payloadFields !== undefined + || manifest.maxEvents !== 1 + || canonicalJson(stringArrayField(manifest.tools, "contextManifest.tools")) !== "[]" + || manifest.externalActions !== false + ) { + throw new Error(`Review candidate context is truncated, projected, enriched, or action-capable: ${run.id}`); + } +} + +function reviewRunReceiptDigest(run: AgentRun, outputEventId: string): string { + const receipt: JsonObject = { + version: 1, + runId: run.id, + executionKey: run.executionKey, + triggerEventId: run.triggerEventId, + agentId: run.agentId, + agentVersion: run.agentVersion, + status: run.status, + inputEventIds: run.inputEventIds, + outputEventId, + attempt: run.attempt, + provider: run.provider, + model: run.model, + privacy: runPrivacy(run), + ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), + ...(run.adapterRevision ? { adapterRevision: run.adapterRevision } : {}), + ...(run.executionAdapterRevision ? { executionAdapterRevision: run.executionAdapterRevision } : {}), + ...(run.modelAdapter ? { modelAdapter: run.modelAdapter as unknown as JsonObject } : {}), + ...(run.adapterCatalogDigest ? { adapterCatalogDigest: run.adapterCatalogDigest } : {}), + ...(run.adapterCatalogGeneration ? { adapterCatalogGeneration: run.adapterCatalogGeneration } : {}), + promptHash: run.promptHash, + contextManifest: run.contextManifest, + }; + return `sha256:${sha256(canonicalJson(receipt))}`; +} + +async function requireRun(store: JazzThoughtStore, id: string): Promise { + const run = await store.getRun(id); + if (!run) throw new Error(`Run not found: ${id}`); + return run; +} + +async function requireEvent(store: JazzThoughtStore, id: string): Promise { + const event = await store.getEvent(id); + if (!event) throw new Error(`Event not found: ${id}`); + return event; +} + +function sameContract(left: { id: string; version: number; sha256: string }, right: { id: string; version: number; sha256: string }): boolean { + return safeEqual(canonicalJson(left as unknown as JsonObject), canonicalJson(right as unknown as JsonObject)); +} + +function safeEqual(left: string, right: string): boolean { + const leftDigest = Buffer.from(sha256(left), "hex"); + const rightDigest = Buffer.from(sha256(right), "hex"); + return timingSafeEqual(leftDigest, rightDigest); +} + +function assertAllowed(values: string[], allowed: string[], label: string): void { + const allowlist = new Set(allowed); + for (const value of values) if (!allowlist.has(value)) throw new Error(`Unknown review ${label}: ${value}`); +} + +function tupleStringField(value: unknown, name: string): [string, string] { + if (!Array.isArray(value) || value.length !== 2 || value.some((part) => typeof part !== "string" || !part)) { + throw new Error(`Review record has invalid ${name}`); + } + return [value[0] as string, value[1] as string]; +} + +function stringArrayField(value: unknown, name: string): string[] { + if (!Array.isArray(value) || value.some((part) => typeof part !== "string" || !part)) { + throw new Error(`Review record has invalid ${name}`); + } + return value as string[]; +} + +function objectField(value: unknown): JsonObject | undefined { + return value && typeof value === "object" && !Array.isArray(value) ? value as JsonObject : undefined; +} + +function stringField(value: unknown, name: string): string { + if (typeof value !== "string" || !value) throw new Error(`Review record has invalid ${name}`); + return value; +} + +function numberField(value: unknown, name: string): number { + if (typeof value !== "number" || !Number.isSafeInteger(value)) throw new Error(`Review record has invalid ${name}`); + return value; +} + +function optionalString(value: unknown): string | undefined { + return typeof value === "string" && value.length > 0 ? value : undefined; +} + +function compareEvents(left: ThoughtEvent, right: ThoughtEvent): number { + return left.observedAt.localeCompare(right.observedAt) || left.id.localeCompare(right.id); +} diff --git a/src/review/types.ts b/src/review/types.ts new file mode 100644 index 0000000..ce50a2e --- /dev/null +++ b/src/review/types.ts @@ -0,0 +1,203 @@ +import { z } from "zod"; +import { + REVIEW_RESPONSE_OUTPUT_CONTRACT, + reviewResponseOutputSchema, + reviewResponseSummary, +} from "../agents/output-contracts.js"; +import { modelAdapterIdentitySchema } from "../adapters/model-adapters.js"; +import type { JsonObject } from "../core/json.js"; + +export const REVIEW_PROMPT_EVENT_TYPE = "stream.thought.source.review.prompt"; +export const REVIEW_ITEM_EVENT_TYPE = "stream.thought.review.item.created"; +export const REVIEW_DECISION_EVENT_TYPE = "stream.thought.review.decision"; +export const REVIEW_RESPONSE_EVENT_TYPE = "stream.thought.derived.review.response"; + +export const REVIEW_PROMPT_SCHEMA_VERSION = 1; +export const REVIEW_ITEM_SCHEMA_VERSION = 1; +export const REVIEW_DECISION_SCHEMA_VERSION = 1; + +export const REVIEW_RUNTIME = "thoughtstream-review-v1"; + +export const reviewCriterionSchema = z.object({ + id: z.string().min(1).max(200).regex(/^[a-z0-9][a-z0-9._-]*$/), + version: z.number().int().positive(), + label: z.string().min(1).max(200), + instructions: z.string().min(1).max(10_000), + reasonCodes: z.array(z.string().min(1).max(100).regex(/^[a-z0-9][a-z0-9._-]*$/)).max(30), + responseTags: z.array(z.string().min(1).max(100).regex(/^[a-z0-9][a-z0-9._-]*$/)).max(30), +}).strict().superRefine((value, context) => { + addDuplicateIssues(value.reasonCodes, "reasonCodes", context); + addDuplicateIssues(value.responseTags, "responseTags", context); +}); + +export const reviewPromptPayloadSchema = z.object({ + campaign: z.object({ + id: z.string().min(1).max(200).regex(/^[a-z0-9][a-z0-9._-]*$/), + version: z.number().int().positive(), + label: z.string().min(1).max(200), + }).strict(), + prompt: z.string().min(1).max(64_000), + evidence: z.string().min(1).max(128_000).optional(), + criterion: reviewCriterionSchema, + candidateAgentIds: z.array(z.string().min(1).max(200)).min(2).max(8), + externalExportEligible: z.boolean(), +}).strict().superRefine((value, context) => { + addDuplicateIssues(value.candidateAgentIds, "candidateAgentIds", context); +}) as unknown as z.ZodType; + +const contractIdentitySchema = z.object({ + id: z.string().min(1), + version: z.number().int().positive(), + sha256: z.string().regex(/^[a-f0-9]{64}$/), +}).strict(); + +export const reviewResponseEventPayloadSchema = z.object({ + runId: z.string().min(1), + executionKey: z.string().min(1), + inputEventId: z.string().min(1), + inputSourceSequence: z.number().int().positive(), + summary: z.string().min(1).max(2_000), + outputContract: contractIdentitySchema, + structuredOutput: reviewResponseOutputSchema, + model: z.record(z.string(), z.unknown()).optional(), + executionAdapterRevision: z.string().min(1).max(500).optional(), + modelAdapter: modelAdapterIdentitySchema.optional(), + adapterCatalogDigest: z.string().regex(/^[a-f0-9]{64}$/).optional(), + adapterCatalogGeneration: z.number().int().positive().optional(), +}).strict().superRefine((value, context) => { + const expected = REVIEW_RESPONSE_OUTPUT_CONTRACT.identity; + if ( + value.outputContract.id !== expected.id + || value.outputContract.version !== expected.version + || value.outputContract.sha256 !== expected.sha256 + ) { + context.addIssue({ code: "custom", path: ["outputContract"], message: "Review response must name the canonical output contract" }); + } + const expectedSummary = reviewResponseSummary(value.structuredOutput.response); + if (value.summary !== expectedSummary) { + context.addIssue({ code: "custom", path: ["summary"], message: "Review response summary must be the canonical bounded preview" }); + } +}) as unknown as z.ZodType; + +export const reviewItemPayloadSchema = z.object({ + promptEventId: z.string().min(1), + candidateRunIds: z.tuple([z.string().min(1), z.string().min(1)]), + candidateOutputEventIds: z.tuple([z.string().min(1), z.string().min(1)]), + candidateRunReceiptDigests: z.tuple([ + z.string().regex(/^sha256:[a-f0-9]{64}$/), + z.string().regex(/^sha256:[a-f0-9]{64}$/), + ]), + displayOrder: z.tuple([z.string().min(1), z.string().min(1)]), + criterion: z.object({ + id: z.string().min(1).max(200), + version: z.number().int().positive(), + }).strict(), + outputContract: contractIdentitySchema, +}).strict().superRefine((value, context) => { + if (value.candidateRunIds[0] === value.candidateRunIds[1]) { + context.addIssue({ code: "custom", path: ["candidateRunIds"], message: "Review candidates must be distinct" }); + } + if (value.candidateOutputEventIds[0] === value.candidateOutputEventIds[1]) { + context.addIssue({ code: "custom", path: ["candidateOutputEventIds"], message: "Review outputs must be distinct" }); + } + if ( + new Set(value.displayOrder).size !== 2 + || !value.candidateRunIds.includes(value.displayOrder[0]) + || !value.candidateRunIds.includes(value.displayOrder[1]) + ) { + context.addIssue({ code: "custom", path: ["displayOrder"], message: "Display order must be a permutation of candidate runs" }); + } +}) as unknown as z.ZodType; + +export const reviewDispositionSchema = z.enum([ + "prefer", + "tie", + "correct", + "underdetermined", + "malformed", + "skip", +]); + +export type ReviewDisposition = z.infer; + +export const reviewDecisionPayloadSchema = z.object({ + reviewItemEventId: z.string().min(1), + promptEventId: z.string().min(1), + disposition: reviewDispositionSchema, + preferredRunId: z.string().min(1).optional(), + preferenceStrength: z.enum(["slight", "strong"]).optional(), + replacementResponse: z.string().min(1).max(60_000).optional(), + confidence: z.enum(["low", "medium", "high"]).optional(), + reasonCodes: z.array(z.string().min(1).max(100)).max(2), + responseTags: z.array(z.string().min(1).max(100)).max(2), + notes: z.string().min(1).max(4_000).optional(), + trainingEligible: z.boolean(), + externalExportEligible: z.boolean(), + submissionId: z.string().min(16).max(200).regex(/^[A-Za-z0-9._~-]+$/), + supersedesDecisionEventId: z.string().min(1).optional(), +}).strict().superRefine((value, context) => { + addDuplicateIssues(value.reasonCodes, "reasonCodes", context); + addDuplicateIssues(value.responseTags, "responseTags", context); + if (value.disposition === "prefer") { + if (!value.preferredRunId) context.addIssue({ code: "custom", path: ["preferredRunId"], message: "Preference requires one candidate" }); + if (!value.preferenceStrength) context.addIssue({ code: "custom", path: ["preferenceStrength"], message: "Preference requires strength" }); + } else if (value.preferredRunId || value.preferenceStrength) { + context.addIssue({ code: "custom", path: ["preferredRunId"], message: "Only preference may identify a winning run" }); + } + if (value.disposition === "correct") { + if (!value.replacementResponse) context.addIssue({ code: "custom", path: ["replacementResponse"], message: "Correction requires a replacement response" }); + } else if (value.replacementResponse) { + context.addIssue({ code: "custom", path: ["replacementResponse"], message: "Only correction may carry a replacement response" }); + } + if (value.trainingEligible && value.disposition !== "prefer" && value.disposition !== "correct") { + context.addIssue({ code: "custom", path: ["trainingEligible"], message: "Only preference or correction may be used for training" }); + } + if (value.externalExportEligible && !value.trainingEligible) { + context.addIssue({ code: "custom", path: ["externalExportEligible"], message: "External export requires training eligibility" }); + } +}) as unknown as z.ZodType; + +export const reviewDecisionSubmissionSchema = z.object({ + disposition: reviewDispositionSchema, + preferredCandidate: z.enum(["A", "B"]).optional(), + preferenceStrength: z.enum(["slight", "strong"]).optional(), + replacementResponse: z.string().min(1).max(60_000).optional(), + confidence: z.enum(["low", "medium", "high"]).optional(), + reasonCodes: z.array(z.string().min(1).max(100)).max(2).default([]), + responseTags: z.array(z.string().min(1).max(100)).max(2).default([]), + notes: z.string().min(1).max(4_000).optional(), + trainingEligible: z.boolean(), + submissionId: z.string().min(16).max(200).regex(/^[A-Za-z0-9._~-]+$/), + supersedesDecisionEventId: z.string().min(1).optional(), +}).strict().superRefine((value, context) => { + addDuplicateIssues(value.reasonCodes, "reasonCodes", context); + addDuplicateIssues(value.responseTags, "responseTags", context); + if (value.disposition === "prefer") { + if (!value.preferredCandidate) context.addIssue({ code: "custom", path: ["preferredCandidate"], message: "Preference requires one candidate" }); + if (!value.preferenceStrength) context.addIssue({ code: "custom", path: ["preferenceStrength"], message: "Preference requires strength" }); + } else if (value.preferredCandidate || value.preferenceStrength) { + context.addIssue({ code: "custom", path: ["preferredCandidate"], message: "Only preference may identify a winning candidate" }); + } + if (value.disposition === "correct") { + if (!value.replacementResponse) context.addIssue({ code: "custom", path: ["replacementResponse"], message: "Correction requires a replacement response" }); + } else if (value.replacementResponse) { + context.addIssue({ code: "custom", path: ["replacementResponse"], message: "Only correction may carry a replacement response" }); + } + if (value.trainingEligible && value.disposition !== "prefer" && value.disposition !== "correct") { + context.addIssue({ code: "custom", path: ["trainingEligible"], message: "Only preference or correction may be used for training" }); + } +}); + +function addDuplicateIssues( + values: string[], + path: string, + context: z.core.$RefinementCtx, +): void { + const seen = new Set(); + for (let index = 0; index < values.length; index += 1) { + if (seen.has(values[index]!)) { + context.addIssue({ code: "custom", path: [path, index], message: "Values must be unique" }); + } + seen.add(values[index]!); + } +} diff --git a/src/review/web-capability.ts b/src/review/web-capability.ts new file mode 100644 index 0000000..2b40f0d --- /dev/null +++ b/src/review/web-capability.ts @@ -0,0 +1,125 @@ +import { createHash, createHmac, randomBytes, timingSafeEqual } from "node:crypto"; + +export const REVIEW_TIMESTAMP_HEADER = "x-thoughtstream-review-timestamp"; +export const REVIEW_NONCE_HEADER = "x-thoughtstream-review-nonce"; +export const REVIEW_SIGNATURE_HEADER = "x-thoughtstream-review-signature"; +export const REVIEW_CSRF_HEADER = "x-thoughtstream-csrf"; + +const DEFAULT_MAX_SKEW_MS = 30_000; +const DEFAULT_MAX_NONCES = 1_024; + +export interface ReviewRequestSignatureInput { + method: string; + path: string; + body: Buffer; + timestamp?: number | undefined; + nonce?: string | undefined; +} + +export interface ReviewRequestSignature { + timestamp: string; + nonce: string; + signature: string; +} + +export function signReviewRequest(key: Buffer, input: ReviewRequestSignatureInput): ReviewRequestSignature { + assertReviewCapabilityKey(key); + const timestamp = String(input.timestamp ?? Date.now()); + const nonce = input.nonce ?? randomBytes(24).toString("base64url"); + if (!/^\d{13}$/.test(timestamp) || !/^[A-Za-z0-9_-]{32}$/.test(nonce)) { + throw new Error("Review request signature inputs are invalid"); + } + return { + timestamp, + nonce, + signature: signatureFor(key, input.method, input.path, input.body, timestamp, nonce), + }; +} + +export class ReviewCapabilityVerifier { + private readonly seen = new Map(); + + constructor( + private readonly key: Buffer, + private readonly options: { + now?: (() => number) | undefined; + maxSkewMs?: number | undefined; + maxNonces?: number | undefined; + } = {}, + ) { + assertReviewCapabilityKey(key); + const skew = options.maxSkewMs ?? DEFAULT_MAX_SKEW_MS; + const maxNonces = options.maxNonces ?? DEFAULT_MAX_NONCES; + if (!Number.isSafeInteger(skew) || skew < 1_000 || skew > 5 * 60_000) throw new Error("Review capability skew bound is invalid"); + if (!Number.isSafeInteger(maxNonces) || maxNonces < 16 || maxNonces > 100_000) throw new Error("Review capability nonce bound is invalid"); + } + + verify( + headers: Record, + method: string, + path: string, + body: Buffer, + ): boolean { + const timestamp = oneHeader(headers[REVIEW_TIMESTAMP_HEADER]); + const nonce = oneHeader(headers[REVIEW_NONCE_HEADER]); + const signature = oneHeader(headers[REVIEW_SIGNATURE_HEADER]); + if (!timestamp || !nonce || !signature) return false; + if (!/^\d{13}$/.test(timestamp) || !/^[A-Za-z0-9_-]{32}$/.test(nonce) || !/^[A-Za-z0-9_-]{43}$/.test(signature)) { + return false; + } + const now = (this.options.now ?? Date.now)(); + const at = Number(timestamp); + const maxSkewMs = this.options.maxSkewMs ?? DEFAULT_MAX_SKEW_MS; + this.prune(now, maxSkewMs); + if (!Number.isSafeInteger(at) || Math.abs(now - at) > maxSkewMs || this.seen.has(nonce)) return false; + const expected = signatureFor(this.key, method, path, body, timestamp, nonce); + if (!safeStringEqual(signature, expected)) return false; + const maxNonces = this.options.maxNonces ?? DEFAULT_MAX_NONCES; + if (this.seen.size >= maxNonces) return false; + this.seen.set(nonce, at); + return true; + } + + private prune(now: number, maxSkewMs: number): void { + for (const [nonce, at] of this.seen) { + if (now - at > maxSkewMs) this.seen.delete(nonce); + } + } +} + +export function decodeReviewCapability(value: string | undefined): Buffer | undefined { + if (!value?.trim()) return undefined; + const normalized = value.replaceAll(/\s+/g, ""); + const key = Buffer.from(normalized, "base64"); + if (key.toString("base64") !== normalized) throw new Error("Review capability must be canonical base64"); + assertReviewCapabilityKey(key); + return key; +} + +function assertReviewCapabilityKey(key: Buffer): void { + if (key.length < 32 || key.length > 128) throw new Error("Review capability must contain 32 to 128 bytes"); +} + +function signatureFor( + key: Buffer, + method: string, + path: string, + body: Buffer, + timestamp: string, + nonce: string, +): string { + const bodyDigest = createHash("sha256").update(body).digest("hex"); + const canonical = `${timestamp}\n${nonce}\n${method.toUpperCase()}\n${path}\n${bodyDigest}`; + return createHmac("sha256", key).update(canonical, "utf8").digest("base64url"); +} + +function safeStringEqual(left: string, right: string): boolean { + const leftDigest = Buffer.from(left, "utf8"); + const rightDigest = Buffer.from(right, "utf8"); + return leftDigest.length === rightDigest.length && timingSafeEqual(leftDigest, rightDigest); +} + +function oneHeader(value: string | string[] | undefined): string | undefined { + if (Array.isArray(value)) return value.length === 1 ? value[0] : undefined; + return value; +} diff --git a/src/security/privacy.ts b/src/security/privacy.ts new file mode 100644 index 0000000..5b52c3d --- /dev/null +++ b/src/security/privacy.ts @@ -0,0 +1,43 @@ +import type { ModelAdapterIdentity, ModelAdapterReleaseIdentity } from "../adapters/model-adapters.js"; +import type { ThoughtAgentDeclaration } from "../agents/types.js"; +import type { PrivacyClass } from "../events/types.js"; +import type { AgentRun } from "../store/types.js"; + +export type PrivacyInput = PrivacyClass | "public" | undefined; + +const privacyRank: Record, number> = { + public: 0, + "public-source": 0, + private: 1, + sensitive: 2, +}; + +export function joinPrivacy(...values: PrivacyInput[]): PrivacyClass { + let result: PrivacyClass = "public-source"; + for (const value of values) { + if (!value) continue; + const normalized: PrivacyClass = value === "public" ? "public-source" : value; + if (privacyRank[normalized] > privacyRank[result]) result = normalized; + } + return result; +} + +export function modelAdapterPrivacy( + adapter: ModelAdapterIdentity | ModelAdapterReleaseIdentity | undefined, +): PrivacyClass { + return adapter ? joinPrivacy(adapter.privacyClass) : "public-source"; +} + +export function declarationPrivacy( + declaration: ThoughtAgentDeclaration, + ...values: PrivacyInput[] +): PrivacyClass { + return joinPrivacy( + ...values, + modelAdapterPrivacy(declaration.modelAdapter ?? declaration.modelAdapterRelease), + ); +} + +export function runPrivacy(run: AgentRun, ...values: PrivacyInput[]): PrivacyClass { + return joinPrivacy(...values, run.privacy, modelAdapterPrivacy(run.modelAdapter)); +} diff --git a/src/store/types.ts b/src/store/types.ts index 2c88c57..c0d645d 100644 --- a/src/store/types.ts +++ b/src/store/types.ts @@ -1,4 +1,5 @@ import type { JsonObject } from "../core/json.js"; +import type { ModelAdapterIdentity } from "../adapters/model-adapters.js"; import type { EventCandidate, ThoughtEvent, @@ -107,8 +108,13 @@ export interface AgentRun { attempt: number; provider: string; model: string; + privacy?: ThoughtEvent["privacy"]; checkpointRevision?: string; adapterRevision?: string; + executionAdapterRevision?: string; + modelAdapter?: ModelAdapterIdentity; + adapterCatalogDigest?: string; + adapterCatalogGeneration?: number; promptHash: string; contextManifest: JsonObject; accountingReservationId?: string; diff --git a/src/training/judgments.ts b/src/training/judgments.ts index 181fd4d..8c05f8b 100644 --- a/src/training/judgments.ts +++ b/src/training/judgments.ts @@ -2,9 +2,10 @@ import path from "node:path"; import { createHash } from "node:crypto"; import { createOutputContractRegistry, + REVIEW_RESPONSE_OUTPUT_CONTRACT, outputContractIdentityJson, parseOutputContractIdentity, - structuredOutputJson, + canonicalStructuredOutput, type OutputContractIdentity, } from "../agents/output-contracts.js"; import { CORRECTION_PROPOSAL_EVENT_TYPE } from "../agents/repairs.js"; @@ -15,6 +16,8 @@ import type { JazzThoughtStore } from "../jazz/store.js"; import { activeJudgments, rebuildEffectiveOutputForRun } from "../projections/effective-output.js"; import type { AgentRun, TraceChunk } from "../store/types.js"; import { assertPrivateDestination, openPrivateDirectory } from "../security/private-files.js"; +import { joinPrivacy, runPrivacy } from "../security/privacy.js"; +import { activeReviewTrainingRecords } from "../review/review.js"; export type JudgmentKind = "accept" | "reject" | "correct" | "prefer"; @@ -52,12 +55,30 @@ export interface TrainingTraceReference { type: string; } +export interface TrainingRunProvenance { + agentVersion: number; + provider: string; + model: string; + checkpointRevision?: string | undefined; + executionAdapterRevision?: string | undefined; + modelAdapter?: JsonObject | undefined; + adapterCatalogDigest?: string | undefined; + adapterCatalogGeneration?: number | undefined; + promptHash: string; + outputContract: JsonObject; + repair?: true | undefined; +} + export interface TrainingExample { - format: "thoughtstream.training-example.v2"; + format: "thoughtstream.training-example.v3" | "thoughtstream.training-example.v4"; kind: JudgmentKind; judgment: { criterion: string; criterionVersion: number; + preferenceStrength?: "slight" | "strong" | undefined; + confidence?: "low" | "medium" | "high" | undefined; + reasonCodes?: string[] | undefined; + responseTags?: string[] | undefined; }; input: { event: { @@ -67,26 +88,28 @@ export interface TrainingExample { privacy: ThoughtEvent["privacy"]; }; contextManifest: JsonObject; + prompt?: string | undefined; + evidence?: string | undefined; }; trajectory: TrainingTraceReference[]; comparedTrajectory?: TrainingTraceReference[] | undefined; chosen?: JsonObject | undefined; rejected?: JsonObject | undefined; - provenance: { - agentVersion: number; - provider: string; - model: string; - checkpointRevision?: string | undefined; - adapterRevision?: string | undefined; - promptHash: string; - outputContract: JsonObject; - repair?: true | undefined; - compared?: true | undefined; - }; + additionalRejected?: JsonObject[] | undefined; + provenance: TrainingRunProvenance; + originalProvenance?: TrainingRunProvenance | undefined; + comparedProvenance?: TrainingRunProvenance | undefined; + review?: { + campaignId: string; + campaignVersion: number; + disposition: "prefer" | "correct"; + candidateCount: 2; + } | undefined; } export interface TrainingProjectionOptions { includeSensitivePrivate?: boolean | undefined; + includeRestrictedModelAdapters?: boolean | undefined; } export interface TrainingWriteOptions { @@ -96,7 +119,7 @@ export interface TrainingWriteOptions { } export interface TrainingDatasetManifest { - format: "thoughtstream.training-dataset-manifest.v2"; + format: "thoughtstream.training-dataset-manifest.v4"; datasetId: string; generatedAt: string; dataFile: string; @@ -104,6 +127,8 @@ export interface TrainingDatasetManifest { examples: number; kinds: Partial>; models: string[]; + exampleFormats: Record; + reviewCampaigns: string[]; } export async function recordJudgment(store: JazzThoughtStore, input: RecordJudgmentInput): Promise { @@ -128,8 +153,13 @@ export async function recordJudgment(store: JazzThoughtStore, input: RecordJudgm } if ( input.externalExportEligible - && [sourceEvent.privacy, outputEvent.privacy, comparedOutputEvent?.privacy] - .some((privacy) => privacy !== undefined && privacy !== "public-source") + && joinPrivacy( + runPrivacy(run), + sourceEvent.privacy, + outputEvent.privacy, + comparedRun ? runPrivacy(comparedRun) : undefined, + comparedOutputEvent?.privacy, + ) !== "public-source" && input.sensitiveExternalExportAuthorized !== true ) { throw new Error("Sensitive/private external export eligibility requires explicit authorization"); @@ -140,12 +170,12 @@ export async function recordJudgment(store: JazzThoughtStore, input: RecordJudgm if (feedbackSourceEvent && feedbackSourceEvent.rootEventId !== sourceEvent.rootEventId) { throw new Error("Judgment feedback source must share the run source root"); } - if (input.deliveryReceiptEventId) { - await requireDeliveryReceipt(store, input.deliveryReceiptEventId, run.id, sourceEvent.rootEventId); - } - if (input.supersedesJudgmentEventId) { - await requireSupersededJudgment(store, input.supersedesJudgmentEventId, run.id, input.criterion, input.criterionVersion); - } + const deliveryReceipt = input.deliveryReceiptEventId + ? await requireDeliveryReceipt(store, input.deliveryReceiptEventId, run.id, sourceEvent.rootEventId) + : undefined; + const supersededJudgment = input.supersedesJudgmentEventId + ? await requireSupersededJudgment(store, input.supersedesJudgmentEventId, run.id, input.criterion, input.criterionVersion) + : undefined; const payload: JsonObject = { runId: run.id, outputEventId, @@ -184,7 +214,16 @@ export async function recordJudgment(store: JazzThoughtStore, input: RecordJudgm rootEventId: sourceEvent.rootEventId, parentEventId: feedbackSourceEvent?.id ?? outputEventId, correlationId: run.id, - privacy: sourceEvent.privacy, + privacy: joinPrivacy( + runPrivacy(run), + sourceEvent.privacy, + outputEvent.privacy, + comparedRun ? runPrivacy(comparedRun) : undefined, + comparedOutputEvent?.privacy, + feedbackSourceEvent?.privacy, + deliveryReceipt?.privacy, + supersededJudgment?.privacy, + ), payload, createdByRuntime: "thoughtstream-judgment-v2", }); @@ -201,13 +240,20 @@ export async function retractJudgment(store: JazzThoughtStore, input: RetractJud } const run = await requireCompletedRun(store, input.runId); const outputEventId = requireOutputEventId(run); + const outputEvent = await requireEvent(store, outputEventId); const sourceEvent = await requireEvent(store, run.triggerEventId); const feedbackSourceEvent = await requireEvent(store, input.feedbackSourceEventId); if (feedbackSourceEvent.rootEventId !== sourceEvent.rootEventId) { throw new Error("Judgment feedback source must share the run source root"); } - await requireDeliveryReceipt(store, input.deliveryReceiptEventId, run.id, sourceEvent.rootEventId); - await requireSupersededJudgment(store, input.retractedJudgmentEventId, run.id, input.criterion, input.criterionVersion); + const deliveryReceipt = await requireDeliveryReceipt(store, input.deliveryReceiptEventId, run.id, sourceEvent.rootEventId); + const retractedJudgment = await requireSupersededJudgment( + store, + input.retractedJudgmentEventId, + run.id, + input.criterion, + input.criterionVersion, + ); const idempotencyKey = stableKey("judgment-retraction", feedbackSourceEvent.id, input.retractedJudgmentEventId); const result = await store.appendEvent({ type: "stream.thought.judgment.training-example.retracted", @@ -221,7 +267,14 @@ export async function retractJudgment(store: JazzThoughtStore, input: RetractJud rootEventId: sourceEvent.rootEventId, parentEventId: feedbackSourceEvent.id, correlationId: run.id, - privacy: sourceEvent.privacy, + privacy: joinPrivacy( + runPrivacy(run), + sourceEvent.privacy, + outputEvent.privacy, + feedbackSourceEvent.privacy, + deliveryReceipt.privacy, + retractedJudgment.privacy, + ), payload: { runId: run.id, outputEventId, @@ -246,6 +299,7 @@ export async function projectTrainingExamples( for (const judgment of judgmentSet.active) { if (!hasExternalExportAuthority(judgment)) continue; const run = await requireCompletedRun(store, stringField(judgment.payload.runId, "runId")); + if (!modelAdapterExportAllowed(run, options)) continue; const outputEvent = await requireEvent(store, requireOutputEventId(run)); let contract: OutputContractIdentity; try { @@ -262,16 +316,18 @@ export async function projectTrainingExamples( let rejected: JsonObject | undefined; let comparedRun: AgentRun | undefined; let comparedEvent: ThoughtEvent | undefined; + let originalRun: AgentRun | undefined; if (kind === "accept") chosen = runOutput; if (kind === "reject") rejected = runOutput; if (kind === "correct") { chosen = replacement - ? structuredOutputJson(createOutputContractRegistry().validate(contract, replacement)) + ? canonicalStructuredOutput(createOutputContractRegistry(), contract, replacement) : undefined; rejected = runOutput; } if (kind === "prefer") { comparedRun = await requireCompletedRun(store, stringField(judgment.payload.comparedRunId, "comparedRunId")); + if (!modelAdapterExportAllowed(comparedRun, options)) continue; comparedEvent = await requireEvent(store, requireOutputEventId(comparedRun)); chosen = runOutput; rejected = validatedEventOutput(comparedEvent, outputContractForRun(comparedRun, comparedEvent)); @@ -280,39 +336,139 @@ export async function projectTrainingExamples( let inputEvent = await requireEvent(store, run.triggerEventId); if (isRepair) { + const originalRunId = stringField(outputEvent.payload.originalRunId, "originalRunId"); + originalRun = await requireRun(store, originalRunId); + if (!modelAdapterExportAllowed(originalRun, options)) continue; const originalTriggerEventId = stringField(outputEvent.payload.originalTriggerEventId, "originalTriggerEventId"); + if (originalRun.triggerEventId !== originalTriggerEventId) continue; inputEvent = await requireEvent(store, originalTriggerEventId); } - const privacyChain = [judgment.privacy, inputEvent.privacy, outputEvent.privacy, comparedEvent?.privacy] - .filter((privacy): privacy is ThoughtEvent["privacy"] => privacy !== undefined); - if (!options.includeSensitivePrivate && privacyChain.some((privacy) => privacy !== "public-source")) continue; + const combinedPrivacy = joinPrivacy( + judgment.privacy, + inputEvent.privacy, + outputEvent.privacy, + runPrivacy(run), + originalRun ? runPrivacy(originalRun) : undefined, + comparedEvent?.privacy, + comparedRun ? runPrivacy(comparedRun) : undefined, + ); + if (!options.includeSensitivePrivate && combinedPrivacy !== "public-source") continue; const trace = await store.listTrace(run.id); examples.push({ - format: "thoughtstream.training-example.v2", + format: "thoughtstream.training-example.v3", kind, judgment: { criterion: stringField(judgment.payload.criterion, "criterion"), criterionVersion: numberField(judgment.payload.criterionVersion, "criterionVersion"), }, input: { - event: trainingEventReference(inputEvent), + event: trainingEventReference(inputEvent, combinedPrivacy), contextManifest: sanitizeContextManifest(run.contextManifest), }, trajectory: (isRepair ? trace.filter(isAllowlistedRepairTrace) : trace).map(redactTrainingTrace), ...(comparedRun ? { comparedTrajectory: (await store.listTrace(comparedRun.id)).map(redactTrainingTrace) } : {}), ...(chosen ? { chosen } : {}), ...(rejected ? { rejected } : {}), - provenance: { - agentVersion: run.agentVersion, - provider: run.provider, - model: run.model, - ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), - ...(run.adapterRevision ? { adapterRevision: run.adapterRevision } : {}), - promptHash: run.promptHash, - outputContract: outputContractIdentityJson(contract), - ...(isRepair ? { repair: true as const } : {}), - ...(comparedRun ? { compared: true as const } : {}), + provenance: trainingRunProvenance(run, contract, isRepair), + ...(originalRun ? { + originalProvenance: trainingRunProvenance( + originalRun, + parseOutputContractIdentity(originalRun.contextManifest.outputContract), + false, + ), + } : {}), + ...(comparedRun && comparedEvent ? { + comparedProvenance: trainingRunProvenance( + comparedRun, + outputContractForRun(comparedRun, comparedEvent), + false, + ), + } : {}), + }); + } + for (const record of await activeReviewTrainingRecords(store)) { + const promptPayload = record.prompt.payload; + const campaign = objectField(promptPayload.campaign); + const criterion = objectField(promptPayload.criterion); + if (!campaign || !criterion) continue; + const disposition = record.decision.payload.disposition; + if (disposition !== "prefer" && disposition !== "correct") continue; + if (record.runs.some((run) => !modelAdapterExportAllowed(run, options))) continue; + const candidateOutputs = record.outputs.map((event) => + canonicalStructuredOutput( + createOutputContractRegistry(), + REVIEW_RESPONSE_OUTPUT_CONTRACT.identity, + event.payload.structuredOutput, + )) as [JsonObject, JsonObject]; + let primaryIndex = 0; + let chosen: JsonObject; + let rejected: JsonObject; + let additionalRejected: JsonObject[] | undefined; + if (disposition === "prefer") { + const preferredRunId = stringField(record.decision.payload.preferredRunId, "preferredRunId"); + primaryIndex = record.runs.findIndex((run) => run.id === preferredRunId); + if (primaryIndex < 0) continue; + chosen = candidateOutputs[primaryIndex]!; + rejected = candidateOutputs[primaryIndex === 0 ? 1 : 0]!; + } else { + chosen = canonicalStructuredOutput( + createOutputContractRegistry(), + REVIEW_RESPONSE_OUTPUT_CONTRACT.identity, + { response: stringField(record.decision.payload.replacementResponse, "replacementResponse") }, + ); + rejected = candidateOutputs[0]; + additionalRejected = [candidateOutputs[1]]; + } + const comparedIndex = primaryIndex === 0 ? 1 : 0; + const primaryRun = record.runs[primaryIndex]!; + const comparedRun = record.runs[comparedIndex]!; + const primaryEvent = record.outputs[primaryIndex]!; + const comparedEvent = record.outputs[comparedIndex]!; + const combinedPrivacy = joinPrivacy( + record.prompt.privacy, + record.item.privacy, + record.decision.privacy, + runPrivacy(primaryRun), + runPrivacy(comparedRun), + primaryEvent.privacy, + comparedEvent.privacy, + ); + if (!options.includeSensitivePrivate && combinedPrivacy !== "public-source") continue; + examples.push({ + format: "thoughtstream.training-example.v4", + kind: disposition, + judgment: { + criterion: stringField(criterion.id, "criterion.id"), + criterionVersion: numberField(criterion.version, "criterion.version"), + ...(record.decision.payload.preferenceStrength ? { + preferenceStrength: record.decision.payload.preferenceStrength as "slight" | "strong", + } : {}), + ...(record.decision.payload.confidence ? { + confidence: record.decision.payload.confidence as "low" | "medium" | "high", + } : {}), + reasonCodes: stringArray(record.decision.payload.reasonCodes), + responseTags: stringArray(record.decision.payload.responseTags), + }, + input: { + event: trainingEventReference(record.prompt, combinedPrivacy), + contextManifest: sanitizeContextManifest(primaryRun.contextManifest), + prompt: stringField(promptPayload.prompt, "prompt"), + ...(typeof promptPayload.evidence === "string" ? { evidence: promptPayload.evidence } : {}), + }, + // Review labels bind prompt/response pairs, not mutable post-run trace projections. + trajectory: [], + comparedTrajectory: [], + chosen, + rejected, + ...(additionalRejected ? { additionalRejected } : {}), + provenance: trainingRunProvenance(primaryRun, outputContractForRun(primaryRun, primaryEvent), false), + comparedProvenance: trainingRunProvenance(comparedRun, outputContractForRun(comparedRun, comparedEvent), false), + review: { + campaignId: stringField(campaign.id, "campaign.id"), + campaignVersion: numberField(campaign.version, "campaign.version"), + disposition, + candidateCount: 2, }, }); } @@ -339,7 +495,7 @@ export async function writeTrainingJsonl( const serialized = content ? `${content}\n` : ""; const sha256 = createHash("sha256").update(serialized).digest("hex"); const manifest: TrainingDatasetManifest = { - format: "thoughtstream.training-dataset-manifest.v2", + format: "thoughtstream.training-dataset-manifest.v4", datasetId: `sha256:${sha256}`, generatedAt: new Date().toISOString(), dataFile: path.basename(absolute), @@ -349,7 +505,18 @@ export async function writeTrainingJsonl( counts[example.kind] = (counts[example.kind] ?? 0) + 1; return counts; }, {}), - models: [...new Set(examples.map((example) => `${example.provenance.provider}:${example.provenance.model}${example.provenance.checkpointRevision ? `#${example.provenance.checkpointRevision}` : ""}${example.provenance.adapterRevision ? `@${example.provenance.adapterRevision}` : ""}`))].sort(), + models: [...new Set(examples.flatMap((example) => [ + trainingModelLabel(example.provenance), + ...(example.originalProvenance ? [trainingModelLabel(example.originalProvenance)] : []), + ...(example.comparedProvenance ? [trainingModelLabel(example.comparedProvenance)] : []), + ]))].sort(), + exampleFormats: examples.reduce>((counts, example) => { + counts[example.format] = (counts[example.format] ?? 0) + 1; + return counts; + }, {}), + reviewCampaigns: [...new Set(examples.flatMap((example) => example.review + ? [`${example.review.campaignId}@${example.review.campaignVersion}`] + : []))].sort(), }; const privateDirectory = await openPrivateDirectory(path.dirname(absolute), { publicContentRoots: options.publicContentRoots, @@ -364,12 +531,61 @@ export async function writeTrainingJsonl( return manifest; } -function trainingEventReference(event: ThoughtEvent): TrainingExample["input"]["event"] { +function modelAdapterExportAllowed(run: AgentRun, options: TrainingProjectionOptions): boolean { + if (run.modelAdapter?.exportClass === "forbidden") return false; + if (run.modelAdapter?.exportClass === "restricted" && options.includeRestrictedModelAdapters !== true) return false; + return true; +} + +function trainingRunProvenance( + run: AgentRun, + contract: OutputContractIdentity, + repair: boolean, +): TrainingRunProvenance { + return { + agentVersion: run.agentVersion, + provider: run.provider, + model: run.model, + ...(run.checkpointRevision ? { checkpointRevision: run.checkpointRevision } : {}), + ...(run.executionAdapterRevision ? { executionAdapterRevision: run.executionAdapterRevision } : {}), + ...(run.modelAdapter ? { modelAdapter: JSON.parse(JSON.stringify(run.modelAdapter)) as JsonObject } : {}), + ...(run.adapterCatalogDigest ? { adapterCatalogDigest: run.adapterCatalogDigest } : {}), + ...(run.adapterCatalogGeneration ? { adapterCatalogGeneration: run.adapterCatalogGeneration } : {}), + promptHash: run.promptHash, + outputContract: outputContractIdentityJson(contract), + ...(repair ? { repair: true as const } : {}), + }; +} + +function trainingModelLabel(provenance: TrainingRunProvenance): string { + const catalog = provenance.adapterCatalogDigest + ? `@catalog:${provenance.adapterCatalogDigest}${provenance.adapterCatalogGeneration ? `:g${provenance.adapterCatalogGeneration}` : ""}` + : ""; + return `${provenance.provider}:${provenance.model}${provenance.checkpointRevision ? `#${provenance.checkpointRevision}` : ""}${provenance.executionAdapterRevision ? `@execution:${provenance.executionAdapterRevision}` : ""}${adapterTrainingLabel(provenance.modelAdapter)}${catalog}`; +} + +function adapterTrainingLabel(value: JsonObject | undefined): string { + if (!value) return ""; + const id = typeof value.id === "string" ? value.id : "unknown"; + const version = typeof value.version === "number" ? value.version : 0; + const manifestSha256 = typeof value.manifestSha256 === "string" ? value.manifestSha256 : "unknown"; + const binding = value.binding; + const checkpointReferenceSha256 = binding && typeof binding === "object" && !Array.isArray(binding) + && typeof binding.checkpointReferenceSha256 === "string" + ? binding.checkpointReferenceSha256 + : "unknown"; + return `@model-adapter:${id}@${version}#${manifestSha256}#${checkpointReferenceSha256}`; +} + +function trainingEventReference( + event: ThoughtEvent, + privacy: ThoughtEvent["privacy"] = event.privacy, +): TrainingExample["input"]["event"] { return { type: event.type, schemaVersion: event.schemaVersion, sourceKind: event.sourceKind, - privacy: event.privacy, + privacy, }; } @@ -388,6 +604,10 @@ function sanitizeContextManifest(manifest: JsonObject): JsonObject { "agentVersion", "agentRole", "outputContract", + "executionAdapterRevision", + "modelAdapter", + "adapterCatalogDigest", + "adapterCatalogGeneration", "tools", "externalActions", ]) { @@ -423,7 +643,7 @@ function isAllowlistedRepairTrace(trace: TraceChunk): boolean { function validatedEventOutput(event: ThoughtEvent, contract: OutputContractIdentity): JsonObject { const candidate = objectField(event.payload.structuredOutput); if (!candidate) throw new Error(`Output event has no canonical structured output: ${event.id}`); - return structuredOutputJson(createOutputContractRegistry().validate(contract, candidate)); + return canonicalStructuredOutput(createOutputContractRegistry(), contract, candidate); } async function rebuildAfterAuthorityChange(store: JazzThoughtStore, run: AgentRun): Promise { @@ -440,7 +660,7 @@ async function rebuildAfterAuthorityChange(store: JazzThoughtStore, run: AgentRu function validateReplacementOutput(run: AgentRun, outputEvent: ThoughtEvent, replacement: JsonObject): JsonObject { const contract = outputContractForRun(run, outputEvent); - return structuredOutputJson(createOutputContractRegistry().validate(contract, replacement)); + return canonicalStructuredOutput(createOutputContractRegistry(), contract, replacement); } function outputContractForRun(run: AgentRun, outputEvent: ThoughtEvent): OutputContractIdentity { @@ -475,6 +695,12 @@ async function requireCompletedRun(store: JazzThoughtStore, id: string): Promise return run; } +async function requireRun(store: JazzThoughtStore, id: string): Promise { + const run = await store.getRun(id); + if (!run) throw new Error(`Run not found: ${id}`); + return run; +} + async function requireEvent(store: JazzThoughtStore, id: string): Promise { const event = await store.getEvent(id); if (!event) throw new Error(`Event not found: ${id}`); @@ -532,3 +758,7 @@ function numberField(value: unknown, name: string): number { function objectField(value: unknown): JsonObject | undefined { return value && typeof value === "object" && !Array.isArray(value) ? value as JsonObject : undefined; } + +function stringArray(value: unknown): string[] { + return Array.isArray(value) && value.every((part) => typeof part === "string") ? value as string[] : []; +} diff --git a/src/web/authenticated-proxy.ts b/src/web/authenticated-proxy.ts new file mode 100644 index 0000000..b580d92 --- /dev/null +++ b/src/web/authenticated-proxy.ts @@ -0,0 +1,617 @@ +import { createHash, timingSafeEqual } from "node:crypto"; +import http, { + type IncomingHttpHeaders, + type IncomingMessage, + type ServerResponse, +} from "node:http"; +import { isIP } from "node:net"; +import { InspectorOAuthAuth, OAuthCallbackQuarantineCapacityError } from "./oauth-auth.js"; +import { createOAuthRouteRateLimiter, type OAuthRouteRateLimiter } from "./rate-limit.js"; +import { loadPublicPages, publicPageForRoute, type PublicPage } from "./public-site.js"; +import { + REVIEW_CSRF_HEADER, + REVIEW_NONCE_HEADER, + REVIEW_SIGNATURE_HEADER, + REVIEW_TIMESTAMP_HEADER, + decodeReviewCapability, + signReviewRequest, +} from "../review/web-capability.js"; + +export interface AuthenticatedInspectorProxyOptions { + host?: string; + port?: number; + upstreamHost?: string; + upstreamPort?: number; + username: string; + password: string; + basicFallbackEnabled?: boolean; + oauth?: InspectorOAuthAuth; + projectRoot?: string; + publicPages?: Map; + oauthRateLimiter?: OAuthRouteRateLimiter; + reviewCapability?: Buffer | undefined; +} + +const SECURITY_HEADERS = { + "cache-control": "no-store", + "content-security-policy": "default-src 'self'; script-src 'none'; style-src 'unsafe-inline'; connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'self'", + "permissions-policy": "camera=(), microphone=(), geolocation=(), payment=(), usb=()", + "referrer-policy": "no-referrer", + "x-content-type-options": "nosniff", + "x-frame-options": "DENY", +} as const; + +const OAUTH_LOGIN_CONTENT_SECURITY_POLICY = + "default-src 'self'; script-src 'none'; style-src 'unsafe-inline'; connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'self' https:"; + +const INSPECTOR_HTML_CONTENT_SECURITY_POLICY = + "default-src 'self'; script-src 'unsafe-inline'; style-src 'unsafe-inline'; connect-src 'self'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'"; + +const FORWARDED_REQUEST_HEADERS = new Set([ + "accept", + "accept-language", + "if-modified-since", + "if-none-match", + "range", + "user-agent", +]); + +export async function startAuthenticatedInspectorProxy( + options: AuthenticatedInspectorProxyOptions, +): Promise { + const host = options.host ?? "127.0.0.1"; + if (!isLoopback(host)) { + throw new Error("thought stream authenticated proxy may only bind to a loopback address"); + } + const port = boundedPort(options.port ?? 4319, "proxy port"); + const upstreamHost = options.upstreamHost ?? "127.0.0.1"; + if (!isLoopback(upstreamHost)) { + throw new Error("thought stream authenticated proxy upstream must be loopback"); + } + const upstreamPort = boundedPort(options.upstreamPort ?? 4317, "upstream port"); + if (!/^[A-Za-z0-9._-]+$/.test(options.username)) { + throw new Error("Inspector proxy username must use only letters, numbers, dot, underscore, or hyphen"); + } + if (Buffer.byteLength(options.password, "utf8") < 20) { + throw new Error("Inspector proxy password must contain at least 20 UTF-8 bytes"); + } + const basicFallbackEnabled = options.basicFallbackEnabled ?? true; + if (!basicFallbackEnabled && !options.oauth) { + throw new Error("Inspector proxy requires OAuth or enabled Basic fallback"); + } + const publicPages = options.publicPages ?? await loadPublicPages(options.projectRoot ?? process.cwd()); + const oauthRateLimiter = options.oauthRateLimiter ?? createOAuthRouteRateLimiter(); + const expectedAuthorizationDigest = digest( + `Basic ${Buffer.from(`${options.username}:${options.password}`, "utf8").toString("base64")}`, + ); + const server = http.createServer((request, response) => { + void handleRequest({ + request, + response, + upstreamHost, + upstreamPort, + publicPages, + ...(options.oauth ? { oauth: options.oauth } : {}), + basicFallbackEnabled, + expectedAuthorizationDigest, + oauthRateLimiter, + reviewCapability: options.reviewCapability, + }).catch(() => { + request.resume(); + if (!response.headersSent) { + send(response, 500, "Request failed.\n", { "content-type": "text/plain; charset=utf-8" }); + } else { + response.destroy(); + } + }); + }); + server.headersTimeout = 10_000; + server.requestTimeout = 30_000; + server.keepAliveTimeout = 5_000; + server.maxRequestsPerSocket = 100; + + await new Promise((resolve, reject) => { + server.once("error", reject); + server.listen(port, host, () => { + server.off("error", reject); + resolve(); + }); + }); + return server; +} + +export function authenticatedProxyOptionsFromEnv( + env: NodeJS.ProcessEnv = process.env, +): AuthenticatedInspectorProxyOptions { + const encodedPassword = env.PROXY_PASSWORD_B64?.trim(); + if (!encodedPassword) throw new Error("PROXY_PASSWORD_B64 is required"); + const passwordBytes = Buffer.from(encodedPassword, "base64"); + if (passwordBytes.length === 0 || passwordBytes.toString("base64") !== normalizeBase64(encodedPassword)) { + throw new Error("PROXY_PASSWORD_B64 must contain canonical base64"); + } + return { + host: env.PROXY_HOST ?? "127.0.0.1", + port: optionalPort(env.PROXY_PORT, 4319, "PROXY_PORT"), + upstreamHost: env.PROXY_UPSTREAM_HOST ?? "127.0.0.1", + upstreamPort: optionalPort(env.PROXY_UPSTREAM_PORT, 4317, "PROXY_UPSTREAM_PORT"), + username: env.PROXY_USER ?? "", + password: passwordBytes.toString("utf8"), + basicFallbackEnabled: env.PROXY_BASIC_FALLBACK_ENABLED !== "0" && env.PROXY_BASIC_FALLBACK_ENABLED !== "false", + projectRoot: env.PROXY_PROJECT_ROOT ?? process.cwd(), + reviewCapability: decodeReviewCapability(env.THOUGHTSTREAM_REVIEW_CAPABILITY_B64), + }; +} + +async function handleRequest(options: { + request: IncomingMessage; + response: ServerResponse; + upstreamHost: string; + upstreamPort: number; + publicPages: Map; + oauth?: InspectorOAuthAuth; + basicFallbackEnabled: boolean; + expectedAuthorizationDigest: Buffer; + oauthRateLimiter: OAuthRouteRateLimiter; + reviewCapability?: Buffer | undefined; +}): Promise { + const url = requestUrl(options.request); + if (!url) { + options.request.resume(); + send(options.response, 400, "Invalid request.\n", { "content-type": "text/plain; charset=utf-8" }); + return; + } + const publicPage = publicPageForRoute(options.publicPages, url.pathname); + if (publicPage) { + if (!isReadMethod(options.request.method)) return methodNotAllowed(options.request, options.response, "GET, HEAD"); + options.request.resume(); + send(options.response, 200, options.request.method === "HEAD" ? "" : publicPage.html, { + "content-type": "text/html; charset=utf-8", + }); + return; + } + if (url.pathname === "/oauth/client-metadata.json" || url.pathname === "/oauth/jwks.json") { + if (!isReadMethod(options.request.method)) return methodNotAllowed(options.request, options.response, "GET, HEAD"); + options.request.resume(); + if (!options.oauth) return notFound(options.response); + const payload = url.pathname.endsWith("client-metadata.json") ? options.oauth.clientMetadata : options.oauth.jwks; + send(options.response, 200, options.request.method === "HEAD" ? "" : `${JSON.stringify(payload)}\n`, { + "content-type": "application/json; charset=utf-8", + }); + return; + } + if (url.pathname === "/oauth/login") { + if (!options.oauth) return notFound(options.response); + if (options.request.method === "GET" || options.request.method === "HEAD") { + options.request.resume(); + const html = loginPage(); + send(options.response, 200, options.request.method === "HEAD" ? "" : html, { + "content-type": "text/html; charset=utf-8", + "content-security-policy": OAUTH_LOGIN_CONTENT_SECURITY_POLICY, + }); + return; + } + if (options.request.method !== "POST") return methodNotAllowed(options.request, options.response, "GET, HEAD, POST"); + const decision = options.oauthRateLimiter.login(rateLimitClientKey(options.request)); + if (!decision.allowed) { + options.request.resume(); + return rateLimited(options.response, decision.retryAfterSeconds); + } + options.request.resume(); + const pending = abortOnDisconnect(options.request, options.response); + try { + const result = await options.oauth.begin(pending.signal); + if (!pending.signal.aborted) { + redirect(options.response, result.redirect.href, [result.setCookie], 302, { + "content-security-policy": OAUTH_LOGIN_CONTENT_SECURITY_POLICY, + }); + } + } catch { + if (!pending.signal.aborted) { + send(options.response, 400, "OAuth login could not start.\n", { "content-type": "text/plain; charset=utf-8" }); + } + } finally { + pending.cleanup(); + } + return; + } + if (url.pathname === "/oauth/callback") { + if (!options.oauth) return notFound(options.response); + if (options.request.method !== "GET") return methodNotAllowed(options.request, options.response, "GET"); + const decision = options.oauthRateLimiter.callback(rateLimitClientKey(options.request)); + if (!decision.allowed) { + options.request.resume(); + return rateLimited(options.response, decision.retryAfterSeconds); + } + options.request.resume(); + try { + const result = await options.oauth.finish(url.searchParams, options.request.headers.cookie); + redirect(options.response, result.redirect, result.setCookies); + } catch (error) { + if (error instanceof OAuthCallbackQuarantineCapacityError) { + send(options.response, 503, "OAuth callback capacity reached. Operator recycle required.\n", { + "content-type": "text/plain; charset=utf-8", + "retry-after": "60", + }); + } else { + send(options.response, 400, "OAuth callback could not be verified.\n", { + "content-type": "text/plain; charset=utf-8", + "set-cookie": [clearFlowCookie()], + }); + } + } + return; + } + if (url.pathname === "/oauth/logout") { + if (!options.oauth) return notFound(options.response); + const session = await options.oauth.authenticate(options.request.headers.cookie); + if (!session) return authenticationRequired(options.request, options.response, options.oauth !== undefined, options.basicFallbackEnabled); + if (options.request.method === "GET" || options.request.method === "HEAD") { + options.request.resume(); + const html = logoutPage(session.csrfToken); + send(options.response, 200, options.request.method === "HEAD" ? "" : html, { "content-type": "text/html; charset=utf-8" }); + return; + } + if (options.request.method !== "POST") return methodNotAllowed(options.request, options.response, "GET, HEAD, POST"); + try { + const fields = new URLSearchParams(await readFormBody(options.request)); + const clear = await options.oauth.logout(options.request.headers.cookie, oneField(fields, "csrfToken")); + redirect(options.response, "/", [clear], 303); + } catch { + send(options.response, 403, "Logout request could not be verified.\n", { "content-type": "text/plain; charset=utf-8" }); + } + return; + } + if (url.pathname !== "/inspector" && !url.pathname.startsWith("/inspector/")) { + options.request.resume(); + return notFound(options.response); + } + + const reviewWriteRoute = url.search === "" && /^\/inspector\/api\/reviews\/[^/]+\/decisions$/.test(url.pathname); + const basicAuthorized = options.basicFallbackEnabled + && isAuthorized(options.request.headers.authorization, options.expectedAuthorizationDigest); + const oauthAuthorized = (reviewWriteRoute || !basicAuthorized) && options.oauth + ? await options.oauth.authenticate(options.request.headers.cookie) + : undefined; + if (!basicAuthorized && !oauthAuthorized) { + options.request.resume(); + return authenticationRequired(options.request, options.response, options.oauth !== undefined, options.basicFallbackEnabled); + } + if (url.pathname === "/inspector/api/session") { + if (!isReadMethod(options.request.method)) return methodNotAllowed(options.request, options.response, "GET, HEAD"); + options.request.resume(); + const body = oauthAuthorized && options.reviewCapability + ? { reviewWriteEnabled: true, csrfToken: oauthAuthorized.csrfToken } + : { reviewWriteEnabled: false }; + send(options.response, 200, options.request.method === "HEAD" ? "" : JSON.stringify(body), { + "content-type": "application/json; charset=utf-8", + }); + return; + } + if (reviewWriteRoute) { + if (options.request.method !== "POST") return methodNotAllowed(options.request, options.response, "POST"); + if (!oauthAuthorized || !options.reviewCapability) { + options.request.resume(); + send(options.response, 403, "Review write is unavailable.\n", { "content-type": "text/plain; charset=utf-8" }); + return; + } + const csrfToken = uniqueHeader(options.request.headers[REVIEW_CSRF_HEADER]); + if (!csrfToken || !safeEqual(csrfToken, oauthAuthorized.csrfToken)) { + options.request.resume(); + send(options.response, 403, "Review request could not be verified.\n", { "content-type": "text/plain; charset=utf-8" }); + return; + } + let body: Buffer; + try { + body = await readJsonBody(options.request, 98_304); + JSON.parse(body.toString("utf8")); + } catch { + options.request.resume(); + send(options.response, 400, "Review request is invalid.\n", { "content-type": "text/plain; charset=utf-8" }); + return; + } + const upstreamPath = `${url.pathname.slice("/inspector".length)}${url.search}`; + const signature = signReviewRequest(options.reviewCapability, { + method: "POST", + path: upstreamPath, + body, + }); + proxyRequest({ + request: options.request, + response: options.response, + upstreamHost: options.upstreamHost, + upstreamPort: options.upstreamPort, + upstreamPath, + body, + additionalHeaders: { + "content-type": "application/json", + "content-length": String(body.length), + [REVIEW_TIMESTAMP_HEADER]: signature.timestamp, + [REVIEW_NONCE_HEADER]: signature.nonce, + [REVIEW_SIGNATURE_HEADER]: signature.signature, + }, + }); + return; + } + if (!isReadMethod(options.request.method)) return methodNotAllowed(options.request, options.response, "GET, HEAD"); + if (url.pathname === "/inspector") { + options.request.resume(); + return redirect(options.response, `/inspector/${url.search}`, [], 308); + } + const upstreamPath = `${url.pathname.slice("/inspector".length) || "/"}${url.search}`; + proxyRequest({ + request: options.request, + response: options.response, + upstreamHost: options.upstreamHost, + upstreamPort: options.upstreamPort, + upstreamPath, + }); +} + +function proxyRequest(options: { + request: IncomingMessage; + response: ServerResponse; + upstreamHost: string; + upstreamPort: number; + upstreamPath: string; + body?: Buffer | undefined; + additionalHeaders?: Record | undefined; +}): void { + const headers = forwardedHeaders(options.request.headers); + headers.host = `${options.upstreamHost}:${options.upstreamPort}`; + Object.assign(headers, options.additionalHeaders); + const upstream = http.request({ + hostname: options.upstreamHost, + port: options.upstreamPort, + path: options.upstreamPath, + method: options.request.method, + headers, + }, (upstreamResponse) => { + const responseHeaders: Record = {}; + for (const [name, value] of Object.entries(upstreamResponse.headers)) { + if (value === undefined || isHopByHop(name) || name === "set-cookie" || name === "www-authenticate") continue; + responseHeaders[name] = value; + } + Object.assign(responseHeaders, SECURITY_HEADERS); + if (isHtmlContentType(upstreamResponse.headers["content-type"])) { + responseHeaders["content-security-policy"] = INSPECTOR_HTML_CONTENT_SECURITY_POLICY; + } + options.response.writeHead(upstreamResponse.statusCode ?? 502, responseHeaders); + upstreamResponse.pipe(options.response); + }); + upstream.on("error", () => { + if (!options.response.headersSent) { + send(options.response, 502, "Inspector unavailable.\n", { + "content-type": "text/plain; charset=utf-8", + }); + } else { + options.response.destroy(); + } + }); + upstream.end(options.body); +} + +function isHtmlContentType(value: string | string[] | undefined): boolean { + const contentType = Array.isArray(value) ? value[0] : value; + return contentType?.split(";", 1)[0]?.trim().toLowerCase() === "text/html"; +} + +function forwardedHeaders(source: IncomingHttpHeaders): Record { + const headers: Record = {}; + for (const [name, value] of Object.entries(source)) { + if (!FORWARDED_REQUEST_HEADERS.has(name) || value === undefined) continue; + headers[name] = value; + } + return headers; +} + +function isAuthorized(value: string | undefined, expectedDigest: Buffer): boolean { + return timingSafeEqual(digest(typeof value === "string" ? value : ""), expectedDigest); +} + +function digest(value: string): Buffer { + return createHash("sha256").update(value, "utf8").digest(); +} + +function send( + response: ServerResponse, + status: number, + body: string, + headers: Record, +): void { + response.writeHead(status, { ...SECURITY_HEADERS, ...headers }); + response.end(body); +} + +function redirect( + response: ServerResponse, + location: string, + cookies: string[] = [], + status = 302, + headers: Record = {}, +): void { + send(response, status, "", { + location, + ...(cookies.length > 0 ? { "set-cookie": cookies } : {}), + ...headers, + }); +} + +function authenticationRequired( + request: IncomingMessage, + response: ServerResponse, + oauthEnabled: boolean, + basicFallbackEnabled: boolean, +): void { + const acceptsHtml = request.headers.accept?.includes("text/html") ?? false; + if (oauthEnabled && acceptsHtml) return redirect(response, "/oauth/login", [], 303); + send(response, 401, "Authentication required.\n", { + "content-type": "text/plain; charset=utf-8", + ...(basicFallbackEnabled ? { "www-authenticate": 'Basic realm="thought stream", charset="UTF-8"' } : {}), + }); +} + +function rateLimited(response: ServerResponse, retryAfterSeconds: number): void { + send(response, 429, "Too many requests.\n", { + "content-type": "text/plain; charset=utf-8", + "retry-after": String(retryAfterSeconds), + }); +} + +function abortOnDisconnect(request: IncomingMessage, response: ServerResponse): { + signal: AbortSignal; + cleanup: () => void; +} { + const controller = new AbortController(); + const abort = () => controller.abort(); + const close = () => { + if (!response.writableEnded) controller.abort(); + }; + request.once("aborted", abort); + response.once("close", close); + return { + signal: controller.signal, + cleanup: () => { + request.off("aborted", abort); + response.off("close", close); + }, + }; +} + +function rateLimitClientKey(request: IncomingMessage): string { + const remote = normalizeAddress(request.socket.remoteAddress ?? "unknown"); + const forwarded = request.headers["x-real-ip"]; + if (isLoopback(remote) && typeof forwarded === "string" && isIP(forwarded.trim()) !== 0) { + return forwarded.trim(); + } + return remote; +} + +function normalizeAddress(value: string): string { + return value.startsWith("::ffff:") ? value.slice("::ffff:".length) : value; +} + +function methodNotAllowed(request: IncomingMessage, response: ServerResponse, allow: string): void { + request.resume(); + send(response, 405, "Method not allowed.\n", { + allow, + "content-type": "text/plain; charset=utf-8", + }); +} + +function notFound(response: ServerResponse): void { + send(response, 404, "Not found.\n", { "content-type": "text/plain; charset=utf-8" }); +} + +function requestUrl(request: IncomingMessage): URL | undefined { + const raw = request.url ?? "/"; + if (Buffer.byteLength(raw, "utf8") > 8_192) return undefined; + try { + return new URL(raw, "http://thoughtstream.invalid"); + } catch { + return undefined; + } +} + +async function readFormBody(request: IncomingMessage): Promise { + if (!(request.headers["content-type"] ?? "").toLowerCase().startsWith("application/x-www-form-urlencoded")) { + throw new Error("Expected form body"); + } + const parts: Buffer[] = []; + let bytes = 0; + for await (const part of request) { + const buffer = Buffer.isBuffer(part) ? part : Buffer.from(part); + bytes += buffer.length; + if (bytes > 4_096) throw new Error("Form body too large"); + parts.push(buffer); + } + return Buffer.concat(parts).toString("utf8"); +} + +async function readJsonBody(request: IncomingMessage, maxBytes: number): Promise { + const contentType = (request.headers["content-type"] ?? "").toLowerCase().split(";", 1)[0]?.trim(); + if (contentType !== "application/json") throw new Error("Expected JSON body"); + const declared = request.headers["content-length"]; + if (typeof declared === "string" && (!/^\d+$/.test(declared) || Number(declared) > maxBytes)) { + throw new Error("JSON body too large"); + } + const parts: Buffer[] = []; + let bytes = 0; + for await (const part of request) { + const buffer = Buffer.isBuffer(part) ? part : Buffer.from(part); + bytes += buffer.length; + if (bytes > maxBytes) throw new Error("JSON body too large"); + parts.push(buffer); + } + if (bytes === 0) throw new Error("JSON body is empty"); + return Buffer.concat(parts); +} + +function oneField(fields: URLSearchParams, name: string): string | undefined { + const values = fields.getAll(name); + return values.length === 1 ? values[0] : undefined; +} + +function uniqueHeader(value: string | string[] | undefined): string | undefined { + if (Array.isArray(value)) return value.length === 1 ? value[0] : undefined; + if (!value || value.includes(",") || Buffer.byteLength(value, "utf8") > 512) return undefined; + return value; +} + +function safeEqual(left: string, right: string): boolean { + return timingSafeEqual(digest(left), digest(right)); +} + +function loginPage(): string { + return page("Private inspector", "

Authenticate with the configured ATProto identity.

"); +} + +function logoutPage(csrfToken: string): string { + return page("Sign out", `

Revoke this inspector session.

`); +} + +function page(title: string, body: string): string { + return `${escapeHtml(title)}

${escapeHtml(title)}

${body}

Public documentation

`; +} + +function escapeHtml(value: string): string { + return value.replaceAll("&", "&").replaceAll("<", "<").replaceAll(">", ">").replaceAll('"', """).replaceAll("'", "'"); +} + +function clearFlowCookie(): string { + return "__Host-thoughtstream_oauth=; Path=/; Max-Age=0; HttpOnly; Secure; SameSite=Lax"; +} + +function isReadMethod(method: string | undefined): boolean { + return method === "GET" || method === "HEAD"; +} + +function optionalPort(value: string | undefined, fallback: number, label: string): number { + if (value === undefined || value.trim() === "") return fallback; + return boundedPort(Number(value), label); +} + +function boundedPort(value: number, label: string): number { + if (!Number.isSafeInteger(value) || value < 0 || value > 65_535) { + throw new Error(`${label} must be an integer between 0 and 65535`); + } + return value; +} + +function normalizeBase64(value: string): string { + return value.replaceAll(/\s+/g, ""); +} + +function isLoopback(host: string): boolean { + return host === "127.0.0.1" || host === "::1" || host === "localhost"; +} + +function isHopByHop(name: string): boolean { + return name === "connection" + || name === "keep-alive" + || name === "proxy-authenticate" + || name === "proxy-authorization" + || name === "te" + || name === "trailer" + || name === "transfer-encoding" + || name === "upgrade"; +} diff --git a/src/web/inspector.ts b/src/web/inspector.ts index 611bd6d..08eef8e 100644 --- a/src/web/inspector.ts +++ b/src/web/inspector.ts @@ -5,10 +5,18 @@ import { buildSourceHealth } from "../projections/source-health.js"; import { auditRunEvidence } from "../agents/evidence.js"; import type { AgentRun } from "../store/types.js"; import type { ThoughtEvent } from "../events/types.js"; +import { + ReviewDecisionConflictError, + projectReviewQueue, + recordBrowserReviewDecision, +} from "../review/review.js"; +import { ReviewCapabilityVerifier } from "../review/web-capability.js"; export interface InspectorServerOptions { host?: string; port?: number; + reviewCapability?: Buffer | undefined; + reviewVerifier?: ReviewCapabilityVerifier | undefined; } export async function startInspectorServer( @@ -18,8 +26,10 @@ export async function startInspectorServer( const host = options.host ?? "127.0.0.1"; if (!isLoopback(host)) throw new Error("thought stream inspector may only bind to a loopback address"); const port = options.port ?? 4317; + const reviewVerifier = options.reviewVerifier + ?? (options.reviewCapability ? new ReviewCapabilityVerifier(options.reviewCapability) : undefined); const server = http.createServer((request, response) => { - void handleRequest(store, request, response).catch((error) => { + void handleRequest(store, request, response, reviewVerifier).catch((error) => { sendJson(response, 500, { error: error instanceof Error ? error.message : String(error) }); }); }); @@ -33,12 +43,49 @@ export async function startInspectorServer( return server; } -async function handleRequest(store: JazzThoughtStore, request: IncomingMessage, response: ServerResponse): Promise { +async function handleRequest( + store: JazzThoughtStore, + request: IncomingMessage, + response: ServerResponse, + reviewVerifier?: ReviewCapabilityVerifier, +): Promise { + const url = new URL(request.url ?? "/", "http://127.0.0.1"); + const reviewDecisionMatch = url.pathname.match(/^\/api\/reviews\/([^/]+)\/decisions$/); + if (request.method === "POST" && reviewDecisionMatch) { + if (!reviewVerifier) { + request.resume(); + sendJson(response, 405, { error: "Inspector is read-only" }); + return; + } + let body: Buffer; + try { + body = await readBody(request, 98_304); + } catch { + request.resume(); + sendJson(response, 400, { error: "Review request is invalid" }); + return; + } + if (!reviewVerifier.verify(request.headers, "POST", url.pathname, body)) { + sendJson(response, 403, { error: "Review request could not be verified" }); + return; + } + try { + const itemId = decodeURIComponent(reviewDecisionMatch[1]!); + const decision = await recordBrowserReviewDecision(store, itemId, JSON.parse(body.toString("utf8"))); + sendJson(response, 201, { decisionEventId: decision.id, active: true }); + } catch (error) { + if (error instanceof ReviewDecisionConflictError) { + sendJson(response, 409, { error: "Review decision changed; reload before submitting" }); + } else { + sendJson(response, 400, { error: "Review decision is invalid" }); + } + } + return; + } if (request.method !== "GET" && request.method !== "HEAD") { sendJson(response, 405, { error: "Inspector is read-only" }); return; } - const url = new URL(request.url ?? "/", "http://127.0.0.1"); if (url.pathname === "/") { send(response, 200, "text/html; charset=utf-8", renderInspectorHtml()); return; @@ -52,16 +99,37 @@ async function handleRequest(store: JazzThoughtStore, request: IncomingMessage, buildSourceHealth(store), ]); const runEvidence = auditRunEvidence(runs, events); + const adapterSelections = agents + .map((agent) => ({ + id: agent.id, + version: agent.version, + enabled: agent.enabled, + adapter: adapterSelectionFromSpec(agent.spec), + })) + .filter((agent) => agent.adapter !== undefined); + const adapterCatalogs = [...new Map(adapterSelections.map((selection) => [ + `${selection.adapter!.digest}:${selection.adapter!.generation}`, + { digest: selection.adapter!.digest, generation: selection.adapter!.generation }, + ])).values()]; sendJson(response, 200, { activity, runs: runs.slice().reverse(), agents, sources, + adapterInventory: { catalogs: adapterCatalogs, selections: adapterSelections }, runEvidence, evidenceContradictions: runEvidence.filter((report) => !report.consistent).length, }); return; } + if (url.pathname === "/api/reviews") { + sendJson(response, 200, await projectReviewQueue(store)); + return; + } + if (url.pathname === "/api/session") { + sendJson(response, 200, { reviewWriteEnabled: false }); + return; + } const eventMatch = url.pathname.match(/^\/api\/events\/([^/]+)$/); if (eventMatch) { const id = decodeURIComponent(eventMatch[1]!); @@ -90,6 +158,15 @@ async function handleRequest(store: JazzThoughtStore, request: IncomingMessage, sendJson(response, 404, { error: "Not found" }); } +function adapterSelectionFromSpec(spec: Record): { digest: string; generation: number; release: Record } | undefined { + const digest = spec.adapterCatalogDigest; + const generation = spec.adapterCatalogGeneration; + const release = spec.modelAdapter; + if (typeof digest !== "string" || !/^[a-f0-9]{64}$/.test(digest) || !Number.isSafeInteger(generation) || Number(generation) < 1) return undefined; + if (!release || typeof release !== "object" || Array.isArray(release)) return undefined; + return { digest, generation: Number(generation), release: release as Record }; +} + export function renderInspectorHtml(): string { return ` @@ -105,51 +182,62 @@ export function renderInspectorHtml(): string { main { display:grid; grid-template-columns:minmax(320px,.9fr) minmax(360px,1.1fr); min-height:calc(100vh - 56px) } section { min-width:0; border-right:1px solid var(--line) } section:last-child { border-right:0 } .section-head { display:flex; justify-content:space-between; padding:10px 14px; border-bottom:1px solid var(--line); color:var(--muted); text-transform:uppercase; letter-spacing:.08em; font-size:11px } - .item { width:100%; display:block; border:0; border-bottom:1px solid #1b242b; padding:10px 14px; background:transparent; color:inherit; text-align:left; font:inherit; cursor:pointer } - .item:hover,.item.active { background:#152028 } .type { color:var(--cyan); overflow-wrap:anywhere } .summary { margin:3px 0; font-size:14px } .meta { color:var(--muted); font-size:11px; overflow-wrap:anywhere } - .consumer { margin:6px 0 2px; padding-left:9px; border-left:2px solid #35515c; color:#c7d2d9; font-size:12px } .consumer strong { color:var(--cyan) } + .item { width:100%; min-width:0; display:block; border:0; border-bottom:1px solid #1b242b; padding:10px 14px; background:transparent; color:inherit; text-align:left; font:inherit; cursor:pointer; overflow-wrap:anywhere } + .item:hover,.item.active { background:#152028 } .type { color:var(--cyan); overflow-wrap:anywhere } .summary { min-width:0; margin:3px 0; font-size:14px; overflow-wrap:anywhere } .meta { color:var(--muted); font-size:11px; overflow-wrap:anywhere } + .consumer { min-width:0; margin:6px 0 2px; padding-left:9px; border-left:2px solid #35515c; color:#c7d2d9; font-size:12px; overflow-wrap:anywhere } .consumer strong { color:var(--cyan) } .badge { display:inline-block; border:1px solid var(--line); padding:1px 5px; margin-left:5px; color:var(--amber) } .failed { color:var(--red) } #detail { padding:14px; position:sticky; top:57px; max-height:calc(100vh - 57px); overflow:auto } .empty { color:var(--muted); padding:30px 8px } h2 { font-size:14px; margin:0 0 12px } h3 { color:var(--muted); text-transform:uppercase; letter-spacing:.08em; font-size:11px; margin:20px 0 7px } pre { margin:0; padding:10px; border:1px solid var(--line); background:#0d1216; white-space:pre-wrap; overflow-wrap:anywhere; font:11px/1.5 inherit } .lineage { display:grid; gap:6px } .lineage button { color:var(--cyan); background:transparent; border:1px solid var(--line); padding:7px; text-align:left; cursor:pointer; overflow-wrap:anywhere } .work { border:1px solid #35515c; background:#0e181d; padding:12px; margin:8px 0 14px } .work-head { display:flex; align-items:center; gap:8px; margin-bottom:8px } .work-head strong { color:var(--cyan) } - .work-summary { font-size:14px; margin:5px 0 10px } .work-label { color:var(--muted); font-size:10px; text-transform:uppercase; letter-spacing:.08em; margin-top:10px } .proposal { border-left:2px solid var(--amber); padding-left:9px; margin-top:5px } .output-list { display:grid; gap:5px; margin-top:5px } .output-list button { color:var(--cyan); background:transparent; border:1px solid var(--line); padding:7px; text-align:left; cursor:pointer; overflow-wrap:anywhere } + .work-summary { min-width:0; font-size:14px; margin:5px 0 10px; overflow-wrap:anywhere } .work-label { color:var(--muted); font-size:10px; text-transform:uppercase; letter-spacing:.08em; margin-top:10px } .proposal { border-left:2px solid var(--amber); padding-left:9px; margin-top:5px } .output-list { display:grid; gap:5px; margin-top:5px } .output-list button { color:var(--cyan); background:transparent; border:1px solid var(--line); padding:7px; text-align:left; cursor:pointer; overflow-wrap:anywhere } .facts { display:flex; flex-wrap:wrap; gap:5px 12px; color:var(--muted); font-size:11px; margin-top:10px } .facts strong { color:var(--text); font-weight:normal } + .review-prompt { border:1px solid var(--line); background:#0d1216; padding:12px; white-space:pre-wrap; overflow-wrap:anywhere; font:12px/1.55 inherit } + .candidate-grid { display:grid; grid-template-columns:1fr 1fr; gap:10px; margin-top:8px } .candidate { min-width:0; border:1px solid #35515c; background:#0e181d; padding:12px } .candidate h3 { margin:0 0 8px; color:var(--cyan) } .candidate-text { white-space:pre-wrap; overflow-wrap:anywhere; font:12px/1.55 inherit } + .decision { border:1px solid #5b4b2f; background:#18150f; padding:12px; margin-top:12px } .decision strong { color:var(--amber) } + .review-form { border-top:1px solid var(--line); margin-top:16px; padding-top:12px; display:grid; gap:10px } .form-row { display:grid; gap:5px } .form-row label { color:var(--muted); font-size:11px; text-transform:uppercase; letter-spacing:.06em } + select,textarea,input[type=text] { width:100%; background:#0d1216; color:var(--text); border:1px solid var(--line); padding:8px; font:12px/1.45 inherit } textarea { min-height:86px; resize:vertical } .choice-row { display:flex; flex-wrap:wrap; gap:8px 14px } .choice-row label { color:var(--text); text-transform:none; letter-spacing:0 } + .submit-row { display:flex; align-items:center; gap:10px } .submit-row button { border:1px solid var(--cyan); background:#102226; color:var(--text); padding:8px 12px; font:inherit; cursor:pointer } .submit-row button:disabled { opacity:.5; cursor:not-allowed } + .provenance { margin-top:10px; color:var(--muted); font-size:10px; overflow-wrap:anywhere } details { border-top:1px solid var(--line); margin-top:12px; padding-top:9px } summary { color:var(--muted); cursor:pointer; text-transform:uppercase; letter-spacing:.08em; font-size:11px } details pre { margin-top:7px } nav { display:flex; gap:8px } nav button { background:transparent; color:var(--muted); border:0; border-bottom:1px solid transparent; font:inherit; padding:0 0 3px; cursor:pointer } nav button.active { color:var(--text); border-color:var(--cyan) } - @media(max-width:800px){ main{grid-template-columns:1fr} section{border-right:0} #detail{position:static;max-height:none} } + @media(max-width:800px){ main{grid-template-columns:1fr} section{border-right:0} #detail{position:static;max-height:none} .candidate-grid{grid-template-columns:1fr} } -

thought stream

loading
+

thought stream

loading
root activity
Select an observation to see what processed it and what, if anything, happened.
`; } @@ -198,6 +314,23 @@ function send(response: ServerResponse, status: number, contentType: string, bod response.end(body); } +async function readBody(request: IncomingMessage, maxBytes: number): Promise { + const declared = request.headers["content-length"]; + if (typeof declared === "string" && (!/^\d+$/.test(declared) || Number(declared) > maxBytes)) { + throw new Error("Request body too large"); + } + const parts: Buffer[] = []; + let bytes = 0; + for await (const part of request) { + const buffer = Buffer.isBuffer(part) ? part : Buffer.from(part); + bytes += buffer.length; + if (bytes > maxBytes) throw new Error("Request body too large"); + parts.push(buffer); + } + if (bytes === 0) throw new Error("Request body is empty"); + return Buffer.concat(parts); +} + function isLoopback(host: string): boolean { return host === "127.0.0.1" || host === "::1" || host === "localhost"; } diff --git a/src/web/oauth-auth.ts b/src/web/oauth-auth.ts new file mode 100644 index 0000000..676c526 --- /dev/null +++ b/src/web/oauth-auth.ts @@ -0,0 +1,793 @@ +import { AsyncLocalStorage } from "node:async_hooks"; +import { randomBytes, timingSafeEqual } from "node:crypto"; +import os from "node:os"; +import path from "node:path"; +import { + JoseKey, + NodeOAuthClient, + type NodeSavedSession, + type NodeSavedState, +} from "@atproto/oauth-client-node"; +import { SecureJsonStore, decodeKey32, opaqueKey } from "./secure-store.js"; +import { SingleProcessDirectoryLock } from "./single-process-lock.js"; + +const SESSION_COOKIE = "__Host-thoughtstream_session"; +const FLOW_COOKIE = "__Host-thoughtstream_oauth"; +const OAUTH_STATE_TTL_MS = 15 * 60_000; +const DEFAULT_BROWSER_SESSION_TTL_MS = 12 * 60 * 60_000; +const SDK_STATE_MAX_ENTRIES = 128; +const SDK_STATE_MAX_BYTES = 512 * 1024; +const SDK_SESSION_MAX_ENTRIES = 8; +const SDK_SESSION_MAX_BYTES = 512 * 1024; +const APP_FLOW_MAX_ENTRIES = 64; +const APP_FLOW_MAX_BYTES = 32 * 1024; +const BROWSER_SESSION_MAX_ENTRIES = 8; +const BROWSER_SESSION_MAX_BYTES = 32 * 1024; +const DEFAULT_CALLBACK_TIMEOUT_MS = 20_000; + +export interface BrowserSession { + did: string; + generation: number; + csrfToken: string; + expiresAt: number; +} + +export interface OAuthProtocolClient { + readonly clientMetadata: unknown; + readonly jwks: unknown; + authorize(handle: string, options: { state: string; scope: string; signal?: AbortSignal }): Promise; + callback(params: URLSearchParams, attemptId: string): Promise<{ session: { did: string }; state: string | null }>; + expireCallbackAttempt(attemptId: string): void; + dropCallbackAttempt(attemptId: string): Promise; + promoteCallbackSession(attemptId: string, did: string, guard: () => boolean): Promise; + isCurrentGeneration(did: string, generation: number): boolean; + discardPromotedSession(did: string, generation: number): Promise; + restore(did: string, generation: number): Promise<{ did: string }>; +} + +export interface InspectorOAuthAuthOptions { + protocol: OAuthProtocolClient; + allowedDid: string; + expectedHandle: string; + browserSessions: SecureJsonStore; + flowStates: SecureJsonStore<{ createdAt: number }>; + processLock?: SingleProcessDirectoryLock; + sessionTtlMs?: number; + callbackTimeoutMs?: number; + now?: () => number; +} + +export interface OAuthStartResult { + redirect: URL; + setCookie: string; +} + +export interface OAuthCallbackResult { + redirect: string; + setCookies: string[]; +} + +export interface OAuthEnvironmentConfiguration { + enabled: boolean; + publicOrigin?: string; + allowedDid?: string; + expectedHandle?: string; + storeDirectory?: string; + storeKey?: Buffer; + privateJwk?: Record; + sessionTtlMs?: number; +} + +class OAuthCallbackWatchdogError extends Error { + constructor() { + super("OAuth callback timed out"); + } +} + +export interface PersistedOAuthSession { + generation: number; + session?: NodeSavedSession; +} + +interface CallbackAttempt { + status: "active" | "expired" | "promoting"; + sessions: Map; +} + +type SessionOperationContext = + | { kind: "callback"; attemptId: string } + | { kind: "restore"; did: string; generation: number }; + +export class OAuthCallbackQuarantineCapacityError extends Error { + constructor(readonly capacity: number) { + super(`OAuth callback quarantine capacity ${capacity} reached; recycle the proxy process`); + } +} + +export class GenerationSessionStore { + private readonly context = new AsyncLocalStorage(); + private readonly attempts = new Map(); + private readonly currentGenerations = new Map(); + private readonly lastGenerations = new Map(); + + constructor( + private readonly persistent: SecureJsonStore, + private readonly maxRetainedAttempts = 8, + ) { + if (!Number.isSafeInteger(maxRetainedAttempts) || maxRetainedAttempts < 1 || maxRetainedAttempts > 128) { + throw new Error("OAuth callback quarantine capacity is invalid"); + } + } + + async get(did: string): Promise { + const context = this.context.getStore(); + if (context?.kind === "callback") { + return structuredClone(this.attempts.get(context.attemptId)?.sessions.get(did)); + } + const record = await this.persistent.get(did); + if (!record) return undefined; + this.lastGenerations.set(did, Math.max(record.generation, this.lastGenerations.get(did) ?? 0)); + if (context?.kind === "restore" && (context.did !== did || context.generation !== record.generation)) { + return undefined; + } + if (!record.session) return undefined; + this.currentGenerations.set(did, record.generation); + return structuredClone(record.session); + } + + async set(did: string, value: NodeSavedSession): Promise { + const context = this.context.getStore(); + if (context?.kind === "callback") { + const attempt = this.attempts.get(context.attemptId); + if (!attempt) throw new Error("Unknown OAuth callback attempt"); + // Expired attempts remain as local quarantine. This avoids one known + // session-store failure path, but the SDK may still revoke on issuer errors. + attempt.sessions.set(did, structuredClone(value)); + return; + } + if (context?.kind === "restore") { + if (context.did !== did) return; + const replaced = await this.persistent.replaceIf( + did, + (record) => record.generation === context.generation && record.session !== undefined, + { generation: context.generation, session: value }, + ); + if (replaced) { + this.currentGenerations.set(did, context.generation); + this.lastGenerations.set(did, Math.max(context.generation, this.lastGenerations.get(did) ?? 0)); + } + return; + } + const current = await this.persistent.get(did); + if (!current?.session) throw new Error("Cannot update an unknown persistent OAuth session"); + this.currentGenerations.set(did, current.generation); + await this.persistent.set( + did, + { generation: current.generation, session: value }, + () => this.currentGenerations.get(did) === current.generation, + ); + this.currentGenerations.set(did, current.generation); + this.lastGenerations.set(did, Math.max(current.generation, this.lastGenerations.get(did) ?? 0)); + } + + async del(did: string): Promise { + const context = this.context.getStore(); + if (context?.kind === "callback") { + this.attempts.get(context.attemptId)?.sessions.delete(did); + return; + } + if (context?.kind === "restore") { + if (context.did !== did) return; + const replaced = await this.persistent.replaceIf( + did, + (record) => record.generation === context.generation, + { generation: context.generation }, + ); + if (replaced && this.currentGenerations.get(did) === context.generation) { + this.currentGenerations.delete(did); + } + return; + } + const current = await this.persistent.get(did); + if (current) { + await this.persistent.replaceIf( + did, + (record) => record.generation === current.generation, + { generation: current.generation }, + ); + if (this.currentGenerations.get(did) === current.generation) this.currentGenerations.delete(did); + } + } + + run(attemptId: string, operation: () => Promise): Promise { + if (this.attempts.has(attemptId)) throw new Error("Duplicate OAuth callback attempt"); + if (this.attempts.size >= this.maxRetainedAttempts) { + throw new OAuthCallbackQuarantineCapacityError(this.maxRetainedAttempts); + } + this.attempts.set(attemptId, { status: "active", sessions: new Map() }); + return this.context.run({ kind: "callback", attemptId }, operation); + } + + runRestore(did: string, generation: number, operation: () => Promise): Promise { + return this.context.run({ kind: "restore", did, generation }, operation); + } + + async restoreExactGeneration(did: string, generation: number, operation: () => Promise): Promise { + if (!await this.matchesPersistedGeneration(did, generation)) throw new Error("OAuth session generation is stale"); + const result = await this.runRestore(did, generation, operation); + if (!await this.matchesPersistedGeneration(did, generation)) { + throw new Error("OAuth session generation changed during restore"); + } + return result; + } + + quarantineStatus(): { retained: number; capacity: number; recycleRequired: boolean } { + return { + retained: this.attempts.size, + capacity: this.maxRetainedAttempts, + recycleRequired: this.attempts.size >= this.maxRetainedAttempts, + }; + } + + expire(attemptId: string): void { + const attempt = this.attempts.get(attemptId); + if (attempt) attempt.status = "expired"; + } + + async drop(attemptId: string): Promise { + const attempt = this.attempts.get(attemptId); + if (!attempt) return; + attempt.status = "expired"; + attempt.sessions.clear(); + this.attempts.delete(attemptId); + } + + async promote(attemptId: string, did: string, guard: () => boolean): Promise { + const attempt = this.attempts.get(attemptId); + if (!attempt || attempt.status !== "active" || !guard()) { + throw new Error("OAuth callback attempt cannot be promoted"); + } + const session = attempt.sessions.get(did); + if (!session) throw new Error("OAuth callback produced no staged session"); + const current = await this.persistent.get(did); + const generation = Math.max( + current?.generation ?? 0, + this.currentGenerations.get(did) ?? 0, + this.lastGenerations.get(did) ?? 0, + ) + 1; + if (attempt.status !== "active" || !guard()) throw new Error("OAuth callback attempt cannot be promoted"); + attempt.status = "promoting"; + try { + await this.persistent.set( + did, + { generation, session }, + () => attempt.status === "promoting" && guard() && generation > (this.currentGenerations.get(did) ?? 0), + ); + if (attempt.status !== "promoting" || !guard()) { + await this.persistent.replaceIf( + did, + (record) => record.generation === generation, + { generation }, + ); + throw new Error("OAuth callback attempt expired during promotion"); + } + this.currentGenerations.set(did, generation); + this.lastGenerations.set(did, generation); + this.attempts.delete(attemptId); + return generation; + } catch (error) { + attempt.status = "expired"; + throw error; + } + } + + isCurrentGeneration(did: string, generation: number): boolean { + return this.currentGenerations.get(did) === generation; + } + + async matchesPersistedGeneration(did: string, generation: number): Promise { + const record = await this.persistent.get(did); + if (!record?.session || record.generation !== generation) return false; + this.currentGenerations.set(did, generation); + this.lastGenerations.set(did, Math.max(generation, this.lastGenerations.get(did) ?? 0)); + return true; + } + + async discardPromoted(did: string, generation: number): Promise { + await this.persistent.replaceIf( + did, + (record) => record.generation === generation, + { generation }, + ); + if (this.currentGenerations.get(did) === generation) this.currentGenerations.delete(did); + this.lastGenerations.set(did, Math.max(generation, this.lastGenerations.get(did) ?? 0)); + } +} + +export class InspectorOAuthAuth { + private readonly now: () => number; + private readonly sessionTtlMs: number; + private readonly callbackTimeoutMs: number; + private callbackOperation = Promise.resolve(); + + constructor(private readonly options: InspectorOAuthAuthOptions) { + if (!/^did:(plc|web):/.test(options.allowedDid)) throw new Error("OAuth allowlisted DID is invalid"); + if (!/^[a-z0-9][a-z0-9.-]+$/i.test(options.expectedHandle)) throw new Error("OAuth expected handle is invalid"); + this.now = options.now ?? Date.now; + this.sessionTtlMs = options.sessionTtlMs ?? DEFAULT_BROWSER_SESSION_TTL_MS; + if (!Number.isSafeInteger(this.sessionTtlMs) || this.sessionTtlMs < 60_000 || this.sessionTtlMs > 7 * 86_400_000) { + throw new Error("OAuth browser session TTL must be between one minute and seven days"); + } + this.callbackTimeoutMs = options.callbackTimeoutMs ?? DEFAULT_CALLBACK_TIMEOUT_MS; + if (!Number.isSafeInteger(this.callbackTimeoutMs) || this.callbackTimeoutMs < 10 || this.callbackTimeoutMs > 120_000) { + throw new Error("OAuth callback timeout must be between 10ms and two minutes"); + } + } + + get clientMetadata(): unknown { + return this.options.protocol.clientMetadata; + } + + get jwks(): unknown { + return this.options.protocol.jwks; + } + + async begin(signal?: AbortSignal): Promise { + const state = randomBytes(32).toString("base64url"); + await this.options.flowStates.set(state, { createdAt: this.now() }); + let redirect: URL; + try { + redirect = await this.options.protocol.authorize(this.options.expectedHandle, { + state, + scope: "atproto", + ...(signal ? { signal } : {}), + }); + } catch (error) { + await this.options.flowStates.del(state); + throw error; + } + return { + redirect, + setCookie: serializeCookie(FLOW_COOKIE, state, { + maxAgeSeconds: Math.floor(OAUTH_STATE_TTL_MS / 1_000), + }), + }; + } + + async finish(params: URLSearchParams, cookieHeader: string | undefined): Promise { + const previous = this.callbackOperation; + let release!: () => void; + const barrier = new Promise((resolve) => { release = resolve; }); + this.callbackOperation = previous.catch(() => undefined).then(() => barrier); + await previous.catch(() => undefined); + try { + return await this.finishSerialized(params, cookieHeader); + } finally { + release(); + } + } + + private async finishSerialized(params: URLSearchParams, cookieHeader: string | undefined): Promise { + const protocolStates = params.getAll("state"); + const applicationState = readUniqueCookie(cookieHeader, FLOW_COOKIE); + if (protocolStates.length !== 1 || !validProtocolState(protocolStates[0]!) || !validCallbackMultiplicity(params) || !applicationState || !validOpaqueState(applicationState)) { + throw new Error("OAuth callback could not be verified"); + } + const attemptId = randomBytes(24).toString("base64url"); + let authoritative = true; + const guard = () => authoritative; + const settlement = this.settleCallbackAttempt(params, applicationState, attemptId, guard); + const watchdog = callbackWatchdog(this.callbackTimeoutMs); + const settled = await Promise.race([ + settlement.then( + (value) => ({ kind: "result" as const, value }), + (error: unknown) => ({ kind: "failure" as const, error }), + ), + watchdog.promise, + ]); + watchdog.cancel(); + if (settled.kind === "timeout") { + authoritative = false; + this.options.protocol.expireCallbackAttempt(attemptId); + this.observeLateSettlement(settlement); + this.deleteFlowStateDetached(applicationState); + throw new OAuthCallbackWatchdogError(); + } + authoritative = false; + if (settled.kind === "failure") throw settled.error; + return settled.value; + } + + private async settleCallbackAttempt( + params: URLSearchParams, + applicationState: string, + attemptId: string, + guard: () => boolean, + ): Promise { + let did: string | undefined; + let generation: number | undefined; + let browserSessionKey: string | undefined; + try { + const result = await this.options.protocol.callback(params, attemptId); + this.assertCallbackAuthority(guard); + did = result.session.did; + const returnedApplicationState = result.state; + const pending = returnedApplicationState && constantTimeTextEqual(returnedApplicationState, applicationState) + ? await this.options.flowStates.take(applicationState, guard) + : undefined; + this.assertCallbackAuthority(guard); + if (!pending) throw new Error("OAuth callback could not be verified"); + if (!constantTimeTextEqual(did, this.options.allowedDid)) throw new Error("OAuth identity is not authorized"); + this.assertCallbackAuthority(guard); + generation = await this.options.protocol.promoteCallbackSession(attemptId, did, guard); + this.assertCallbackAuthority(guard); + if (!this.options.protocol.isCurrentGeneration(did, generation)) { + throw new Error("OAuth callback generation is no longer current"); + } + const sessionId = randomBytes(32).toString("base64url"); + browserSessionKey = opaqueKey(sessionId); + const browserSession: BrowserSession = { + did, + generation, + csrfToken: randomBytes(32).toString("base64url"), + expiresAt: this.now() + this.sessionTtlMs, + }; + await this.options.browserSessions.replaceAll( + browserSessionKey, + browserSession, + () => guard() && this.options.protocol.isCurrentGeneration(did!, generation!), + ); + this.assertCallbackAuthority(guard); + if (!this.options.protocol.isCurrentGeneration(did, generation)) { + throw new Error("OAuth callback generation is no longer current"); + } + return { + redirect: "/inspector/", + setCookies: [ + serializeCookie(SESSION_COOKIE, sessionId, { maxAgeSeconds: Math.floor(this.sessionTtlMs / 1_000) }), + clearCookie(FLOW_COOKIE), + ], + }; + } catch (error) { + if (error instanceof OAuthCallbackQuarantineCapacityError) throw error; + this.options.protocol.expireCallbackAttempt(attemptId); + if (browserSessionKey && generation !== undefined) { + this.detachCleanup(this.options.browserSessions.deleteIf( + browserSessionKey, + (session) => session.generation === generation, + )); + } + if (did && generation !== undefined) { + this.detachCleanup(this.options.protocol.discardPromotedSession(did, generation)); + } + this.detachCleanup(this.options.protocol.dropCallbackAttempt(attemptId)); + this.deleteFlowStateDetached(applicationState); + throw error; + } + } + + private assertCallbackAuthority(guard: () => boolean): void { + if (!guard()) throw new OAuthCallbackWatchdogError(); + } + + async authenticate(cookieHeader: string | undefined): Promise { + const sessionId = readUniqueCookie(cookieHeader, SESSION_COOKIE); + if (!sessionId || !validOpaqueState(sessionId)) return undefined; + const key = opaqueKey(sessionId); + const stored = await this.options.browserSessions.getWithExpiration(key); + if (!stored) return undefined; + const browserSession = stored.value; + if (stored.expired || browserSession.expiresAt <= this.now() || !constantTimeTextEqual(browserSession.did, this.options.allowedDid)) { + await this.options.browserSessions.del(key); + await this.options.protocol.discardPromotedSession(browserSession.did, browserSession.generation); + return undefined; + } + try { + const restored = await this.options.protocol.restore(browserSession.did, browserSession.generation); + if (!constantTimeTextEqual(restored.did, this.options.allowedDid)) throw new Error("restored DID mismatch"); + return browserSession; + } catch { + await this.options.browserSessions.del(key); + await this.options.protocol.discardPromotedSession(browserSession.did, browserSession.generation); + return undefined; + } + } + + async logout(cookieHeader: string | undefined, csrfToken: string | undefined): Promise { + const sessionId = readUniqueCookie(cookieHeader, SESSION_COOKIE); + if (!sessionId || !csrfToken) throw new Error("Logout request could not be verified"); + const key = opaqueKey(sessionId); + const browserSession = await this.options.browserSessions.get(key); + if (!browserSession || !constantTimeTextEqual(browserSession.csrfToken, csrfToken)) { + throw new Error("Logout request could not be verified"); + } + await this.options.browserSessions.del(key); + await this.options.protocol.discardPromotedSession(browserSession.did, browserSession.generation); + return clearCookie(SESSION_COOKIE); + } + + private observeLateSettlement(settlement: Promise): void { + void settlement.catch(() => undefined); + } + + private detachCleanup(cleanup: Promise): void { + void cleanup.catch(() => undefined); + } + + private deleteFlowStateDetached(applicationState: string): void { + this.detachCleanup(this.options.flowStates.del(applicationState)); + } + + async close(): Promise { + await this.options.processLock?.release(); + } + +} + +export async function createInspectorOAuthAuth( + configuration: Required> & { sessionTtlMs?: number }, +): Promise { + const origin = new URL(configuration.publicOrigin); + if (origin.protocol !== "https:" || origin.username || origin.password || origin.pathname !== "/" || origin.search || origin.hash) { + throw new Error("OAUTH_PUBLIC_ORIGIN must be one bare HTTPS origin"); + } + const key = await JoseKey.fromJWK(configuration.privateJwk); + const processLock = await SingleProcessDirectoryLock.acquire(configuration.storeDirectory); + try { + const stateStore = new SecureJsonStore({ + directory: configuration.storeDirectory, + name: "oauth-state", + key: configuration.storeKey, + maxEntries: SDK_STATE_MAX_ENTRIES, + maxSerializedBytes: SDK_STATE_MAX_BYTES, + ttlMs: OAUTH_STATE_TTL_MS, + validate: objectValue as (value: unknown) => NodeSavedState, + }); + const persistedSessions = new SecureJsonStore({ + directory: configuration.storeDirectory, + name: "oauth-session", + key: configuration.storeKey, + maxEntries: SDK_SESSION_MAX_ENTRIES, + maxSerializedBytes: SDK_SESSION_MAX_BYTES, + validate: (value) => { + const object = objectValue(value); + if (!Number.isSafeInteger(object.generation) || Number(object.generation) < 1) { + throw new Error("Invalid OAuth session generation"); + } + return { + generation: Number(object.generation), + ...(object.session === undefined ? {} : { session: objectValue(object.session) as NodeSavedSession }), + }; + }, + }); + const sessionStore = new GenerationSessionStore(persistedSessions); + const flowStates = new SecureJsonStore<{ createdAt: number }>({ + directory: configuration.storeDirectory, + name: "browser-oauth-flow", + key: configuration.storeKey, + maxEntries: APP_FLOW_MAX_ENTRIES, + maxSerializedBytes: APP_FLOW_MAX_BYTES, + ttlMs: OAUTH_STATE_TTL_MS, + validate: (value) => { + const object = objectValue(value); + if (!Number.isSafeInteger(object.createdAt)) throw new Error("Invalid OAuth flow state"); + return { createdAt: Number(object.createdAt) }; + }, + }); + const browserSessions = new SecureJsonStore({ + directory: configuration.storeDirectory, + name: "browser-session", + key: configuration.storeKey, + maxEntries: BROWSER_SESSION_MAX_ENTRIES, + maxSerializedBytes: BROWSER_SESSION_MAX_BYTES, + validate: browserSessionValue, + }); + await Promise.all([stateStore.initialize(), persistedSessions.initialize(), flowStates.initialize(), browserSessions.initialize()]); + const clientId = new URL("/oauth/client-metadata.json", origin).href; + const callback = new URL("/oauth/callback", origin).href; + const jwksUri = new URL("/oauth/jwks.json", origin).href; + const client = new NodeOAuthClient({ + clientMetadata: { + client_id: clientId, + client_name: "ThoughtStream inspector", + client_uri: origin.href, + redirect_uris: [callback], + grant_types: ["authorization_code", "refresh_token"], + scope: "atproto", + response_types: ["code"], + application_type: "web", + token_endpoint_auth_method: "private_key_jwt", + token_endpoint_auth_signing_alg: "ES256", + dpop_bound_access_tokens: true, + jwks_uri: jwksUri, + }, + keyset: [key], + stateStore, + sessionStore, + requestLock: keyedRuntimeLock(), + }); + const protocol: OAuthProtocolClient = { + clientMetadata: client.clientMetadata, + jwks: client.jwks, + authorize: (handle, options) => client.authorize(handle, options), + callback: (params, attemptId) => sessionStore.run(attemptId, () => client.callback(params)), + expireCallbackAttempt: (attemptId) => sessionStore.expire(attemptId), + dropCallbackAttempt: (attemptId) => sessionStore.drop(attemptId), + promoteCallbackSession: (attemptId, did, guard) => sessionStore.promote(attemptId, did, guard), + isCurrentGeneration: (did, generation) => sessionStore.isCurrentGeneration(did, generation), + discardPromotedSession: (did, generation) => sessionStore.discardPromoted(did, generation), + restore: (did, generation) => sessionStore.restoreExactGeneration( + did, + generation, + () => client.restore(did), + ), + }; + return new InspectorOAuthAuth({ + protocol, + allowedDid: configuration.allowedDid, + expectedHandle: configuration.expectedHandle, + browserSessions, + flowStates, + processLock, + ...(configuration.sessionTtlMs ? { sessionTtlMs: configuration.sessionTtlMs } : {}), + }); + } catch (error) { + await processLock.release(); + throw error; + } +} + +export function oauthConfigurationFromEnv(env: NodeJS.ProcessEnv = process.env): OAuthEnvironmentConfiguration { + const enabled = env.OAUTH_ENABLED === "1" || env.OAUTH_ENABLED === "true"; + if (!enabled) return { enabled: false }; + const privateJwk = parsePrivateJwk(env.OAUTH_PRIVATE_JWK_B64); + const sessionTtlMs = env.OAUTH_SESSION_TTL_MS ? Number(env.OAUTH_SESSION_TTL_MS) : undefined; + const home = env.HOME?.trim() || os.homedir(); + const expectedStoreDirectory = path.resolve(home, ".local", "share", "thoughtstream-inspector-auth"); + const storeDirectory = path.resolve(required(env.OAUTH_STORE_DIR, "OAUTH_STORE_DIR")); + if (storeDirectory !== expectedStoreDirectory) { + throw new Error(`OAUTH_STORE_DIR must be ${expectedStoreDirectory}; the systemd sandbox permits only that owner-only path`); + } + return { + enabled: true, + publicOrigin: required(env.OAUTH_PUBLIC_ORIGIN, "OAUTH_PUBLIC_ORIGIN"), + allowedDid: required(env.OAUTH_ALLOWED_DID, "OAUTH_ALLOWED_DID"), + expectedHandle: required(env.OAUTH_EXPECTED_HANDLE, "OAUTH_EXPECTED_HANDLE"), + storeDirectory, + storeKey: decodeKey32(env.OAUTH_STORE_KEY_B64, "OAUTH_STORE_KEY_B64"), + privateJwk, + ...(sessionTtlMs !== undefined ? { sessionTtlMs } : {}), + }; +} + +export function sessionCookieName(): string { + return SESSION_COOKIE; +} + +export function flowCookieName(): string { + return FLOW_COOKIE; +} + +function callbackWatchdog(timeoutMs: number): { + promise: Promise<{ kind: "timeout" }>; + cancel: () => void; +} { + let timer: NodeJS.Timeout | undefined; + const promise = new Promise<{ kind: "timeout" }>((resolve) => { + timer = setTimeout(() => resolve({ kind: "timeout" }), timeoutMs); + timer.unref(); + }); + return { + promise, + cancel: () => { + if (timer) clearTimeout(timer); + timer = undefined; + }, + }; +} + +function keyedRuntimeLock(): (key: string, fn: () => T | PromiseLike) => Promise { + const locks = new Map>(); + return async (key: string, fn: () => T | PromiseLike): Promise => { + const previous = locks.get(key) ?? Promise.resolve(); + let release!: () => void; + const barrier = new Promise((resolve) => { release = resolve; }); + const chain = previous.catch(() => undefined).then(() => barrier); + locks.set(key, chain); + await previous.catch(() => undefined); + try { + return await fn(); + } finally { + release(); + if (locks.get(key) === chain) locks.delete(key); + } + }; +} + +function serializeCookie(name: string, value: string, options: { maxAgeSeconds: number }): string { + return `${name}=${value}; Path=/; Max-Age=${options.maxAgeSeconds}; HttpOnly; Secure; SameSite=Lax`; +} + +function clearCookie(name: string): string { + return `${name}=; Path=/; Max-Age=0; HttpOnly; Secure; SameSite=Lax`; +} + +function readUniqueCookie(header: string | undefined, name: string): string | undefined { + if (!header || header.length > 8_192) return undefined; + const values = header.split(";").map((part) => part.trim()).flatMap((part) => { + const index = part.indexOf("="); + return index > 0 && part.slice(0, index) === name ? [part.slice(index + 1)] : []; + }); + return values.length === 1 ? values[0] : undefined; +} + +function validCallbackMultiplicity(params: URLSearchParams): boolean { + const codes = params.getAll("code"); + const errors = params.getAll("error"); + return codes.length <= 1 + && errors.length <= 1 + && params.getAll("iss").length <= 1 + && params.getAll("response").length <= 1 + && !(codes.length === 1 && errors.length === 1); +} + +function validProtocolState(value: string): boolean { + return value.length >= 16 && value.length <= 512 && /^[A-Za-z0-9._~-]+$/.test(value); +} + +function validOpaqueState(value: string): boolean { + return /^[A-Za-z0-9_-]{43}$/.test(value); +} + +function constantTimeTextEqual(left: string, right: string): boolean { + const leftBytes = Buffer.from(left, "utf8"); + const rightBytes = Buffer.from(right, "utf8"); + if (leftBytes.length !== rightBytes.length) return false; + return timingSafeEqual(leftBytes, rightBytes); +} + +function parsePrivateJwk(encoded: string | undefined): Record { + if (!encoded) throw new Error("OAUTH_PRIVATE_JWK_B64 is required"); + const compact = encoded.trim(); + const bytes = Buffer.from(compact, "base64"); + if (bytes.length === 0 || bytes.toString("base64") !== compact) { + throw new Error("OAUTH_PRIVATE_JWK_B64 must be canonical base64"); + } + try { + const value = JSON.parse(bytes.toString("utf8")); + if (!value || typeof value !== "object" || Array.isArray(value) || typeof value.d !== "string") throw new Error("shape"); + return value as Record; + } catch { + throw new Error("OAUTH_PRIVATE_JWK_B64 must encode one private JWK object"); + } +} + +function browserSessionValue(value: unknown): BrowserSession { + const object = objectValue(value); + if ( + typeof object.did !== "string" + || !Number.isSafeInteger(object.generation) + || Number(object.generation) < 1 + || typeof object.csrfToken !== "string" + || !Number.isSafeInteger(object.expiresAt) + ) { + throw new Error("Invalid browser session value"); + } + return { + did: object.did, + generation: Number(object.generation), + csrfToken: object.csrfToken, + expiresAt: Number(object.expiresAt), + }; +} + +function objectValue(value: unknown): Record { + if (!value || typeof value !== "object" || Array.isArray(value)) throw new Error("Expected object value"); + return value as Record; +} + +function required(value: string | undefined, label: string): string { + const trimmed = value?.trim(); + if (!trimmed) throw new Error(`${label} is required`); + return trimmed; +} diff --git a/src/web/public-site.ts b/src/web/public-site.ts new file mode 100644 index 0000000..a9c5532 --- /dev/null +++ b/src/web/public-site.ts @@ -0,0 +1,120 @@ +import fs from "node:fs/promises"; +import path from "node:path"; + +export interface PublicPage { + route: string; + title: string; + html: string; +} + +const PUBLIC_PAGE_FILES = new Map([ + ["/", "index.md"], + ["/docs", "docs/index.md"], + ["/docs/architecture", "docs/architecture.md"], + ["/docs/security", "docs/security.md"], +]); + +const MAX_PUBLIC_FILE_BYTES = 64 * 1024; + +export async function loadPublicPages(projectRoot: string): Promise> { + const publicRoot = path.resolve(projectRoot, "public"); + const pages = new Map(); + for (const [route, relativePath] of PUBLIC_PAGE_FILES) { + const filePath = path.resolve(publicRoot, relativePath); + if (filePath !== publicRoot && !filePath.startsWith(`${publicRoot}${path.sep}`)) { + throw new Error("Public content path escapes the allowlisted root"); + } + const stat = await fs.lstat(filePath); + if (!stat.isFile() || stat.isSymbolicLink() || stat.size > MAX_PUBLIC_FILE_BYTES) { + throw new Error(`Invalid public content file: ${relativePath}`); + } + const markdown = await fs.readFile(filePath, "utf8"); + const title = firstHeading(markdown); + pages.set(route, { route, title, html: renderPage(title, markdown) }); + } + return pages; +} + +export function publicPageForRoute(pages: Map, pathname: string): PublicPage | undefined { + const normalized = pathname.length > 1 && pathname.endsWith("/") ? pathname.slice(0, -1) : pathname; + return pages.get(normalized); +} + +function renderPage(title: string, markdown: string): string { + const body = renderMarkdown(markdown); + return ` + + + + +${escapeHtml(title)} · ThoughtStream + + + + +
thought stream
+
${body}
+
Static public documentation. No live stream data is available from these routes.
+ +`; +} + +function renderMarkdown(markdown: string): string { + const blocks: string[] = []; + const paragraph: string[] = []; + const list: string[] = []; + const flushParagraph = () => { + if (paragraph.length === 0) return; + blocks.push(`

${inlineMarkup(paragraph.join(" "))}

`); + paragraph.length = 0; + }; + const flushList = () => { + if (list.length === 0) return; + blocks.push(`
    ${list.map((item) => `
  • ${inlineMarkup(item)}
  • `).join("")}
`); + list.length = 0; + }; + for (const rawLine of markdown.replaceAll("\r\n", "\n").split("\n")) { + const line = rawLine.trim(); + const heading = line.match(/^(#{1,2})\s+(.+)$/); + const item = line.match(/^[-*]\s+(.+)$/); + if (heading) { + flushParagraph(); + flushList(); + const level = heading[1]!.length; + blocks.push(`${inlineMarkup(heading[2]!)}`); + } else if (item) { + flushParagraph(); + list.push(item[1]!); + } else if (line === "") { + flushParagraph(); + flushList(); + } else { + flushList(); + paragraph.push(line); + } + } + flushParagraph(); + flushList(); + return blocks.join("\n"); +} + +function inlineMarkup(value: string): string { + return escapeHtml(value).replaceAll(/`([^`]+)`/g, "$1"); +} + +function firstHeading(markdown: string): string { + const heading = markdown.match(/^#\s+(.+)$/m)?.[1]?.trim(); + if (!heading) throw new Error("Public content requires one level-one heading"); + return heading; +} + +function escapeHtml(value: string): string { + return value + .replaceAll("&", "&") + .replaceAll("<", "<") + .replaceAll(">", ">") + .replaceAll('"', """) + .replaceAll("'", "'"); +} diff --git a/src/web/rate-limit.ts b/src/web/rate-limit.ts new file mode 100644 index 0000000..073f22b --- /dev/null +++ b/src/web/rate-limit.ts @@ -0,0 +1,82 @@ +export interface RateLimitDecision { + allowed: boolean; + retryAfterSeconds: number; +} + +interface Bucket { + startedAt: number; + count: number; +} + +export interface FixedWindowRateLimiterOptions { + maxRequests: number; + windowMs: number; + maxKeys: number; + now?: () => number; +} + +export class FixedWindowRateLimiter { + private readonly buckets = new Map(); + private readonly now: () => number; + + constructor(private readonly options: FixedWindowRateLimiterOptions) { + if (!Number.isSafeInteger(options.maxRequests) || options.maxRequests < 1) throw new Error("Rate limit maxRequests must be positive"); + if (!Number.isSafeInteger(options.windowMs) || options.windowMs < 1_000) throw new Error("Rate limit windowMs must be at least one second"); + if (!Number.isSafeInteger(options.maxKeys) || options.maxKeys < 1 || options.maxKeys > 100_000) throw new Error("Rate limit maxKeys is invalid"); + this.now = options.now ?? Date.now; + } + + check(key: string): RateLimitDecision { + const now = this.now(); + this.prune(now); + let bucket = this.buckets.get(key); + if (!bucket) { + if (this.buckets.size >= this.options.maxKeys) { + return { allowed: false, retryAfterSeconds: Math.ceil(this.options.windowMs / 1_000) }; + } + bucket = { startedAt: now, count: 0 }; + this.buckets.set(key, bucket); + } + const elapsed = now - bucket.startedAt; + if (elapsed >= this.options.windowMs) { + bucket.startedAt = now; + bucket.count = 0; + } + bucket.count += 1; + const retryAfterSeconds = Math.max(1, Math.ceil((bucket.startedAt + this.options.windowMs - now) / 1_000)); + return { allowed: bucket.count <= this.options.maxRequests, retryAfterSeconds }; + } + + private prune(now: number): void { + for (const [key, bucket] of this.buckets) { + if (now - bucket.startedAt >= this.options.windowMs) this.buckets.delete(key); + } + } +} + +export interface OAuthRouteRateLimiter { + login(clientKey: string): RateLimitDecision; + callback(clientKey: string): RateLimitDecision; +} + +export function createOAuthRouteRateLimiter(now: () => number = Date.now): OAuthRouteRateLimiter { + const loginGlobal = new FixedWindowRateLimiter({ maxRequests: 30, windowMs: 10 * 60_000, maxKeys: 1, now }); + const loginClient = new FixedWindowRateLimiter({ maxRequests: 6, windowMs: 10 * 60_000, maxKeys: 1_024, now }); + const callbackGlobal = new FixedWindowRateLimiter({ maxRequests: 60, windowMs: 10 * 60_000, maxKeys: 1, now }); + const callbackClient = new FixedWindowRateLimiter({ maxRequests: 12, windowMs: 10 * 60_000, maxKeys: 1_024, now }); + return { + login(clientKey) { + return combine(loginGlobal.check("global"), loginClient.check(clientKey)); + }, + callback(clientKey) { + return combine(callbackGlobal.check("global"), callbackClient.check(clientKey)); + }, + }; +} + +function combine(global: RateLimitDecision, client: RateLimitDecision): RateLimitDecision { + return { + allowed: global.allowed && client.allowed, + retryAfterSeconds: Math.max(global.retryAfterSeconds, client.retryAfterSeconds), + }; +} diff --git a/src/web/secure-store.ts b/src/web/secure-store.ts new file mode 100644 index 0000000..e138c2d --- /dev/null +++ b/src/web/secure-store.ts @@ -0,0 +1,289 @@ +import { createCipheriv, createDecipheriv, createHash, randomBytes } from "node:crypto"; +import fs from "node:fs/promises"; +import path from "node:path"; + +interface StoredEntry { + value: T; + expiresAt?: number; +} + +interface StoreDocument { + version: 1; + entries: Record>; +} + +interface EncryptedEnvelope { + version: 1; + algorithm: "aes-256-gcm"; + iv: string; + ciphertext: string; + tag: string; +} + +export interface SecureJsonStoreOptions { + directory: string; + name: string; + key: Buffer; + maxEntries: number; + maxSerializedBytes: number; + ttlMs?: number; + now?: () => number; + validate?: (value: unknown) => T; +} + +export interface StoreReadResult { + value: T; + expired: boolean; +} + +export class SecureJsonStore { + private readonly filePath: string; + private readonly now: () => number; + private readonly validate: (value: unknown) => T; + private document: StoreDocument = { version: 1, entries: {} }; + private operation = Promise.resolve(); + private initialization?: Promise; + private initialized = false; + + constructor(private readonly options: SecureJsonStoreOptions) { + if (!path.isAbsolute(options.directory)) throw new Error("Secure store directory must be absolute"); + if (!/^[a-z][a-z0-9-]*$/.test(options.name)) throw new Error("Secure store name is invalid"); + if (options.key.length !== 32) throw new Error("Secure store key must contain exactly 32 bytes"); + if (!Number.isSafeInteger(options.maxEntries) || options.maxEntries < 1 || options.maxEntries > 100_000) { + throw new Error("Secure store maxEntries must be between 1 and 100000"); + } + if (!Number.isSafeInteger(options.maxSerializedBytes) || options.maxSerializedBytes < 1_024 || options.maxSerializedBytes > 64 * 1024 * 1024) { + throw new Error("Secure store maxSerializedBytes must be between 1024 and 67108864"); + } + if (options.ttlMs !== undefined && (!Number.isSafeInteger(options.ttlMs) || options.ttlMs < 1_000)) { + throw new Error("Secure store TTL must be at least one second"); + } + this.filePath = path.join(options.directory, `${options.name}.enc.json`); + this.now = options.now ?? Date.now; + this.validate = options.validate ?? ((value) => value as T); + } + + async initialize(): Promise { + if (this.initialized) return; + this.initialization ??= this.initializeOnce(); + return this.initialization; + } + + async get(key: string): Promise { + const result = await this.getWithExpiration(key); + if (!result) return undefined; + if (result.expired) { + await this.del(key); + return undefined; + } + return result.value; + } + + async getWithExpiration(key: string): Promise | undefined> { + await this.initialize(); + await this.operation; + const entry = this.document.entries[key]; + if (!entry) return undefined; + return { + value: structuredClone(entry.value), + expired: entry.expiresAt !== undefined && entry.expiresAt <= this.now(), + }; + } + + async set(key: string, value: T, guard?: () => boolean): Promise { + await this.mutate(() => { + this.document.entries[key] = { + value: structuredClone(value), + ...(this.options.ttlMs ? { expiresAt: this.now() + this.options.ttlMs } : {}), + }; + }, guard); + } + + async replaceAll(key: string, value: T, guard?: () => boolean): Promise { + await this.mutate(() => { + this.document.entries = { + [key]: { + value: structuredClone(value), + ...(this.options.ttlMs ? { expiresAt: this.now() + this.options.ttlMs } : {}), + }, + }; + }, guard); + } + + async del(key: string): Promise { + await this.mutate(() => { + delete this.document.entries[key]; + }); + } + + async replaceIf(key: string, predicate: (value: T) => boolean, value: T): Promise { + let replaced = false; + await this.mutate(() => { + const entry = this.document.entries[key]; + if (!entry || !predicate(entry.value)) return; + this.document.entries[key] = { + value: structuredClone(value), + ...(this.options.ttlMs ? { expiresAt: this.now() + this.options.ttlMs } : {}), + }; + replaced = true; + }); + return replaced; + } + + async deleteIf(key: string, predicate: (value: T) => boolean): Promise { + let deleted = false; + await this.mutate(() => { + const entry = this.document.entries[key]; + if (!entry || !predicate(entry.value)) return; + delete this.document.entries[key]; + deleted = true; + }); + return deleted; + } + + async take(key: string, guard?: () => boolean): Promise { + let value: T | undefined; + await this.mutate(() => { + const entry = this.document.entries[key]; + if (!entry || (entry.expiresAt !== undefined && entry.expiresAt <= this.now())) { + delete this.document.entries[key]; + return; + } + value = structuredClone(entry.value); + delete this.document.entries[key]; + }, guard); + return value; + } + + async size(): Promise { + await this.initialize(); + await this.operation; + return Object.keys(this.document.entries).length; + } + + private async initializeOnce(): Promise { + await fs.mkdir(this.options.directory, { recursive: true, mode: 0o700 }); + await fs.chmod(this.options.directory, 0o700); + const stat = await fs.stat(this.filePath).catch((error: NodeJS.ErrnoException) => { + if (error.code === "ENOENT") return undefined; + throw error; + }); + if (stat) { + const maxEnvelopeBytes = Math.ceil(this.options.maxSerializedBytes * 1.5) + 4_096; + if (!stat.isFile() || stat.size > maxEnvelopeBytes) throw new Error("Secure store envelope exceeds its configured bound"); + this.document = this.decryptDocument(await fs.readFile(this.filePath, "utf8")); + } + this.removeExpired(); + this.assertBounds(); + this.initialized = true; + } + + private async mutate(change: () => void, guard?: () => boolean): Promise { + await this.initialize(); + const next = this.operation.then(async () => { + const previous = structuredClone(this.document); + try { + if (guard && !guard()) throw new Error("Secure store write lost authority"); + change(); + this.removeExpired(); + this.assertBounds(); + if (guard && !guard()) throw new Error("Secure store write lost authority"); + await this.persist(guard); + } catch (error) { + this.document = previous; + throw error; + } + }); + this.operation = next.catch(() => undefined); + return next; + } + + private removeExpired(): void { + const now = this.now(); + for (const [key, entry] of Object.entries(this.document.entries)) { + if (entry.expiresAt !== undefined && entry.expiresAt <= now) delete this.document.entries[key]; + } + } + + private assertBounds(): void { + const count = Object.keys(this.document.entries).length; + if (count > this.options.maxEntries) throw new Error("Secure store entry limit exceeded"); + const bytes = Buffer.byteLength(JSON.stringify(this.document), "utf8"); + if (bytes > this.options.maxSerializedBytes) throw new Error("Secure store serialized byte limit exceeded"); + } + + private async persist(guard?: () => boolean): Promise { + const iv = randomBytes(12); + const cipher = createCipheriv("aes-256-gcm", this.options.key, iv); + const plaintext = Buffer.from(JSON.stringify(this.document), "utf8"); + if (plaintext.length > this.options.maxSerializedBytes) throw new Error("Secure store serialized byte limit exceeded"); + const ciphertext = Buffer.concat([cipher.update(plaintext), cipher.final()]); + const envelope: EncryptedEnvelope = { + version: 1, + algorithm: "aes-256-gcm", + iv: iv.toString("base64"), + ciphertext: ciphertext.toString("base64"), + tag: cipher.getAuthTag().toString("base64"), + }; + const tempPath = `${this.filePath}.tmp-${process.pid}-${randomBytes(6).toString("hex")}`; + await fs.writeFile(tempPath, `${JSON.stringify(envelope)}\n`, { mode: 0o600, flag: "wx" }); + if (guard && !guard()) { + await fs.unlink(tempPath); + throw new Error("Secure store write lost authority"); + } + await fs.rename(tempPath, this.filePath); + } + + private decryptDocument(raw: string): StoreDocument { + let envelope: EncryptedEnvelope; + try { + envelope = JSON.parse(raw) as EncryptedEnvelope; + } catch { + throw new Error("Secure store envelope is invalid"); + } + if (envelope.version !== 1 || envelope.algorithm !== "aes-256-gcm") { + throw new Error("Secure store envelope version is unsupported"); + } + try { + const decipher = createDecipheriv("aes-256-gcm", this.options.key, Buffer.from(envelope.iv, "base64")); + decipher.setAuthTag(Buffer.from(envelope.tag, "base64")); + const plaintext = Buffer.concat([ + decipher.update(Buffer.from(envelope.ciphertext, "base64")), + decipher.final(), + ]); + if (plaintext.length > this.options.maxSerializedBytes) throw new Error("document bytes"); + const parsed = JSON.parse(plaintext.toString("utf8")) as StoreDocument; + if (parsed.version !== 1 || !parsed.entries || typeof parsed.entries !== "object" || Array.isArray(parsed.entries)) { + throw new Error("document shape"); + } + if (Object.keys(parsed.entries).length > this.options.maxEntries) throw new Error("entry count"); + const entries: Record> = {}; + for (const [key, entry] of Object.entries(parsed.entries)) { + if (!entry || typeof entry !== "object" || Array.isArray(entry) || !("value" in entry)) throw new Error("entry shape"); + const expiresAt = "expiresAt" in entry ? Number(entry.expiresAt) : undefined; + if (expiresAt !== undefined && !Number.isSafeInteger(expiresAt)) throw new Error("entry expiry"); + entries[key] = { + value: this.validate(entry.value), + ...(expiresAt !== undefined ? { expiresAt } : {}), + }; + } + return { version: 1, entries }; + } catch { + throw new Error("Secure store could not be authenticated or decoded"); + } + } +} + +export function decodeKey32(encoded: string | undefined, label: string): Buffer { + if (!encoded) throw new Error(`${label} is required`); + const compact = encoded.trim(); + const bytes = Buffer.from(compact, "base64"); + if (bytes.length !== 32 || bytes.toString("base64") !== compact) { + throw new Error(`${label} must be canonical base64 for exactly 32 bytes`); + } + return bytes; +} + +export function opaqueKey(value: string): string { + return createHash("sha256").update(value, "utf8").digest("hex"); +} diff --git a/src/web/single-process-lock.ts b/src/web/single-process-lock.ts new file mode 100644 index 0000000..763ecbd --- /dev/null +++ b/src/web/single-process-lock.ts @@ -0,0 +1,99 @@ +import { randomBytes } from "node:crypto"; +import fs from "node:fs/promises"; +import path from "node:path"; + +interface LockRecord { + pid: number; + token: string; + createdAt: string; +} + +export class SingleProcessDirectoryLock { + private released = false; + + private constructor( + private readonly filePath: string, + private readonly token: string, + private readonly handle: fs.FileHandle, + ) {} + + static async acquire(directory: string, name = "oauth-store"): Promise { + if (!path.isAbsolute(directory)) throw new Error("Single-process lock directory must be absolute"); + if (!/^[a-z][a-z0-9-]*$/.test(name)) throw new Error("Single-process lock name is invalid"); + await fs.mkdir(directory, { recursive: true, mode: 0o700 }); + const directoryStat = await fs.lstat(directory); + if (!directoryStat.isDirectory() || directoryStat.isSymbolicLink()) { + throw new Error("Single-process lock directory must be a real directory"); + } + await fs.chmod(directory, 0o700); + const filePath = path.join(directory, `${name}.lock`); + for (let attempt = 0; attempt < 2; attempt += 1) { + const token = randomBytes(24).toString("base64url"); + let handle: fs.FileHandle; + try { + handle = await fs.open(filePath, "wx", 0o600); + } catch (error) { + if ((error as NodeJS.ErrnoException).code !== "EEXIST") throw error; + const existing = await readLockRecord(filePath); + if (processIsAlive(existing.pid)) { + throw new Error(`OAuth store is already owned by process ${existing.pid}`); + } + if (attempt > 0) throw new Error("OAuth store lock could not be reclaimed safely"); + await fs.unlink(filePath); + continue; + } + const record: LockRecord = { pid: process.pid, token, createdAt: new Date().toISOString() }; + try { + await handle.writeFile(`${JSON.stringify(record)}\n`, "utf8"); + await handle.sync(); + await fs.chmod(filePath, 0o600); + return new SingleProcessDirectoryLock(filePath, token, handle); + } catch (error) { + await handle.close().catch(() => undefined); + await fs.unlink(filePath).catch(() => undefined); + throw error; + } + } + throw new Error("OAuth store lock could not be acquired"); + } + + async release(): Promise { + if (this.released) return; + this.released = true; + await this.handle.close(); + const existing = await readLockRecord(this.filePath).catch(() => undefined); + if (existing?.token === this.token && existing.pid === process.pid) { + await fs.unlink(this.filePath).catch((error: NodeJS.ErrnoException) => { + if (error.code !== "ENOENT") throw error; + }); + } + } +} + +async function readLockRecord(filePath: string): Promise { + const stat = await fs.lstat(filePath); + if (!stat.isFile() || stat.isSymbolicLink() || stat.size > 4_096) { + throw new Error("OAuth store lock file is invalid; operator review is required"); + } + try { + const parsed = JSON.parse(await fs.readFile(filePath, "utf8")) as Partial; + if (!Number.isSafeInteger(parsed.pid) || Number(parsed.pid) < 1 || typeof parsed.token !== "string" || parsed.token.length < 16) { + throw new Error("shape"); + } + return { pid: Number(parsed.pid), token: parsed.token, createdAt: String(parsed.createdAt ?? "") }; + } catch { + throw new Error("OAuth store lock file is invalid; operator review is required"); + } +} + +function processIsAlive(pid: number): boolean { + try { + process.kill(pid, 0); + return true; + } catch (error) { + const code = (error as NodeJS.ErrnoException).code; + if (code === "ESRCH") return false; + if (code === "EPERM") return true; + throw error; + } +} diff --git a/test/acceptance.test.ts b/test/acceptance.test.ts index a48d831..dcc4653 100644 --- a/test/acceptance.test.ts +++ b/test/acceptance.test.ts @@ -139,7 +139,9 @@ describe("Jazz-native producer and consumer topology", () => { runtime = new ThoughtAgentRuntime(store); const [recovered] = await runtime.consumeBacklog(declarations); - expect(recovered?.output?.recommendation?.target).toBe("charter"); + expect(recovered?.output && "recommendation" in recovered.output + ? recovered.output.recommendation?.target + : undefined).toBe("charter"); expect((await store.getRun(interruptedRunId))?.status).toBe("abandoned"); const latest = await store.latestRunForExecution(executionKey); expect(latest).toMatchObject({ status: "completed", attempt: 2, triggerEventId: changedEvent.id }); diff --git a/test/agent-runtime.test.ts b/test/agent-runtime.test.ts index 0318895..92e6bc0 100644 --- a/test/agent-runtime.test.ts +++ b/test/agent-runtime.test.ts @@ -90,7 +90,7 @@ describe("ThoughtAgentRuntime", () => { }], }); expect(activity.items[0]?.consumerRuns[0]?.description) - .toBe("Classified spec/brief.md as a specification change (high importance). Proposed: Re-evaluate Charter work affected by this specification change; do not edit files automatically. The proposal was not executed. External actions were disabled."); + .toBe("Classified spec/brief.md as a specification change (high importance). Proposed: Re-evaluate Charter work affected by this specification change; do not edit files automatically. The proposal was not executed."); const replay = await runtime.consumeBacklog(declarations); expect(replay).toHaveLength(0); diff --git a/test/authenticated-proxy.test.ts b/test/authenticated-proxy.test.ts new file mode 100644 index 0000000..d836e6c --- /dev/null +++ b/test/authenticated-proxy.test.ts @@ -0,0 +1,746 @@ +import http from "node:http"; +import { afterEach, describe, expect, test } from "vitest"; +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { + authenticatedProxyOptionsFromEnv, + startAuthenticatedInspectorProxy, +} from "../src/web/authenticated-proxy.js"; +import { + InspectorOAuthAuth, + OAuthCallbackQuarantineCapacityError, + type BrowserSession, + type OAuthProtocolClient, +} from "../src/web/oauth-auth.js"; +import { SecureJsonStore } from "../src/web/secure-store.js"; +import { + REVIEW_CSRF_HEADER, + ReviewCapabilityVerifier, +} from "../src/review/web-capability.js"; + +const servers: http.Server[] = []; +const roots: string[] = []; + +afterEach(async () => { + await Promise.all(servers.splice(0).map((server) => new Promise((resolve) => { + server.closeAllConnections(); + server.close(() => resolve()); + }))); + await Promise.all(roots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); +}); + +describe("authenticated inspector proxy", () => { + test("gates every route and forwards only authenticated read requests without credentials", async () => { + const upstreamRequests: Array<{ + method: string | undefined; + authorization: string | undefined; + cookie: string | undefined; + path: string | undefined; + forwardedFor: string | undefined; + forwardedProto: string | undefined; + }> = []; + const upstream = http.createServer((request, response) => { + upstreamRequests.push({ + method: request.method, + authorization: request.headers.authorization, + cookie: request.headers.cookie, + path: request.url, + forwardedFor: request.headers["x-forwarded-for"] as string | undefined, + forwardedProto: request.headers["x-forwarded-proto"] as string | undefined, + }); + response.writeHead(200, { + "content-type": "application/json", + "set-cookie": "upstream-secret=forbidden", + }); + response.end(JSON.stringify({ private: true })); + }); + servers.push(upstream); + const upstreamPort = await listen(upstream); + + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + }); + servers.push(proxy); + const base = baseUrl(proxy); + + const missing = await fetch(`${base}/inspector/api/private-object`); + expect(missing.status).toBe(401); + expect(missing.headers.get("www-authenticate")).toContain("thought stream"); + expect(missing.headers.get("cache-control")).toBe("no-store"); + expect(missing.headers.get("referrer-policy")).toBe("no-referrer"); + expect(upstreamRequests).toHaveLength(0); + + const wrong = await fetch(`${base}/inspector/api/private-object`, { + headers: { authorization: basic("cameron", "wrong-password-with-enough-bytes") }, + }); + expect(wrong.status).toBe(401); + expect(upstreamRequests).toHaveLength(0); + + const authorized = await fetch(`${base}/inspector/api/private-object?detail=1`, { + headers: { + authorization: basic("cameron", "correct-horse-battery-staple-private"), + cookie: "browser-secret=forbidden", + "x-forwarded-for": "198.51.100.4", + "x-forwarded-proto": "https", + }, + }); + expect(authorized.status).toBe(200); + expect(await authorized.json()).toEqual({ private: true }); + expect(authorized.headers.get("set-cookie")).toBeNull(); + expect(authorized.headers.get("x-frame-options")).toBe("DENY"); + expect(upstreamRequests).toEqual([{ + method: "GET", + authorization: undefined, + cookie: undefined, + path: "/api/private-object?detail=1", + forwardedFor: undefined, + forwardedProto: undefined, + }]); + + const head = await fetch(`${base}/inspector/api/private-object`, { + method: "HEAD", + headers: { authorization: basic("cameron", "correct-horse-battery-staple-private") }, + }); + expect(head.status).toBe(200); + expect(await head.text()).toBe(""); + expect(upstreamRequests.at(-1)?.method).toBe("HEAD"); + + const mutation = await fetch(`${base}/inspector/api/private-object`, { + method: "POST", + headers: { authorization: basic("cameron", "correct-horse-battery-staple-private") }, + }); + expect(mutation.status).toBe(405); + expect(mutation.headers.get("allow")).toBe("GET, HEAD"); + expect(mutation.headers.get("content-security-policy")).toContain("frame-ancestors 'none'"); + expect(mutation.headers.get("permissions-policy")).toContain("camera=()"); + expect(upstreamRequests).toHaveLength(2); + }); + + test("permits the authenticated inspector loader only on upstream HTML", async () => { + const upstream = http.createServer((request, response) => { + if (request.url === "/") { + response.writeHead(200, { "content-type": "text/html; charset=utf-8" }); + response.end(""); + return; + } + response.writeHead(200, { "content-type": "application/json; charset=utf-8" }); + response.end("{}\n"); + }); + servers.push(upstream); + const upstreamPort = await listen(upstream); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + }); + servers.push(proxy); + const headers = { authorization: basic("cameron", "correct-horse-battery-staple-private") }; + + const page = await fetch(`${baseUrl(proxy)}/inspector/`, { headers }); + expect(await page.text()).toContain("fetch('api/snapshot')"); + expect(page.headers.get("content-security-policy")).toContain("script-src 'unsafe-inline'"); + expect(page.headers.get("content-security-policy")).toContain("form-action 'none'"); + + const api = await fetch(`${baseUrl(proxy)}/inspector/api/snapshot`, { headers }); + expect(await api.json()).toEqual({}); + expect(api.headers.get("content-security-policy")).toContain("script-src 'none'"); + }); + + test("serves only allowlisted public pages and never falls through to the private upstream", async () => { + let upstreamRequests = 0; + const upstream = http.createServer((_request, response) => { + upstreamRequests += 1; + response.end("private"); + }); + servers.push(upstream); + const upstreamPort = await listen(upstream); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + }); + servers.push(proxy); + const base = baseUrl(proxy); + + const landing = await fetch(base); + expect(landing.status).toBe(200); + expect(await landing.text()).toContain("No live stream data"); + expect(landing.headers.get("content-security-policy")).toContain("form-action 'self'"); + expect(landing.headers.get("content-security-policy")).not.toContain("form-action 'self' https:"); + const architecture = await fetch(`${base}/docs/architecture`); + expect(architecture.status).toBe(200); + expect(await architecture.text()).toContain("Connector cursors advance only after durable events"); + const traversal = await fetch(`${base}/docs/..%2f..%2fetc%2fpasswd`); + expect(traversal.status).toBe(404); + const oldApi = await fetch(`${base}/api/private-object`); + expect(oldApi.status).toBe(404); + const unknown = await fetch(`${base}/anything`); + expect(unknown.status).toBe(404); + expect(upstreamRequests).toBe(0); + }); + + test("completes an OAuth browser session while Basic remains available", async () => { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-proxy-oauth-")); + roots.push(root); + const key = Buffer.alloc(32, 6); + const browserSessions = new SecureJsonStore({ + directory: root, + name: "browser", + key, + maxEntries: 8, + maxSerializedBytes: 32 * 1024, + }); + const flowStates = new SecureJsonStore<{ createdAt: number }>({ + directory: root, + name: "flow", + key, + maxEntries: 64, + maxSerializedBytes: 32 * 1024, + ttlMs: 15 * 60_000, + }); + const protocolState = "sdk-protocol-state-1234567890abcdef"; + let applicationState: string | undefined; + const stagedAttempts = new Map(); + let currentGeneration: number | undefined; + let nextGeneration = 0; + const protocol: OAuthProtocolClient = { + clientMetadata: { client_id: "https://thought.stream/oauth/client-metadata.json", dpop_bound_access_tokens: true }, + jwks: { keys: [{ kty: "EC", kid: "fixture" }] }, + authorize: async (_handle, options) => { + applicationState = options.state; + return new URL("https://pds.example/authorize?request_uri=urn:ietf:params:oauth:request_uri:fixture"); + }, + callback: async (_params, attemptId) => { + stagedAttempts.set(attemptId, "did:plc:allowed"); + return { session: { did: "did:plc:allowed" }, state: applicationState ?? null }; + }, + expireCallbackAttempt: () => undefined, + dropCallbackAttempt: async (attemptId: string) => { stagedAttempts.delete(attemptId); }, + promoteCallbackSession: async (attemptId, did, guard) => { + if (!guard() || stagedAttempts.get(attemptId) !== did) throw new Error("missing staged session"); + stagedAttempts.delete(attemptId); + currentGeneration = ++nextGeneration; + return currentGeneration; + }, + isCurrentGeneration: (_did, generation) => currentGeneration === generation, + discardPromotedSession: async (_did, generation) => { + if (currentGeneration === generation) currentGeneration = undefined; + }, + restore: async (did, generation) => { + if (currentGeneration !== generation) throw new Error("stale generation"); + return { did }; + }, + }; + const oauth = new InspectorOAuthAuth({ + protocol, + allowedDid: "did:plc:allowed", + expectedHandle: "cameron.stream", + browserSessions, + flowStates, + }); + const upstream = http.createServer((_request, response) => response.end("private inspector")); + servers.push(upstream); + const upstreamPort = await listen(upstream); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth, + basicFallbackEnabled: true, + }); + servers.push(proxy); + const base = baseUrl(proxy); + + const metadata = await fetch(`${base}/oauth/client-metadata.json`); + expect(await metadata.json()).toMatchObject({ client_id: "https://thought.stream/oauth/client-metadata.json" }); + const loginPageResponse = await fetch(`${base}/oauth/login`); + expect(loginPageResponse.headers.get("content-security-policy")).toContain("form-action 'self' https:"); + const protectedRoute = await fetch(`${base}/inspector`, { headers: { accept: "text/html" }, redirect: "manual" }); + expect(protectedRoute.status).toBe(303); + expect(protectedRoute.headers.get("location")).toBe("/oauth/login"); + + const login = await fetch(`${base}/oauth/login`, { method: "POST", redirect: "manual" }); + expect(login.status).toBe(302); + expect(login.headers.get("content-security-policy")).toContain("form-action 'self' https:"); + const flowCookie = login.headers.get("set-cookie")!; + const state = cookieFromHeader(flowCookie, "__Host-thoughtstream_oauth"); + const callback = await fetch(`${base}/oauth/callback?state=${encodeURIComponent(protocolState)}&code=fixture`, { + headers: { cookie: `__Host-thoughtstream_oauth=${state}` }, + redirect: "manual", + }); + expect(callback.status).toBe(302); + expect(callback.headers.get("location")).toBe("/inspector/"); + const sessionId = cookieFromHeader(callback.headers.get("set-cookie")!, "__Host-thoughtstream_session"); + + const inspector = await fetch(`${base}/inspector/api/private`, { + headers: { cookie: `__Host-thoughtstream_session=${sessionId}` }, + }); + expect(await inspector.text()).toBe("private inspector"); + const logoutPage = await fetch(`${base}/oauth/logout`, { + headers: { cookie: `__Host-thoughtstream_session=${sessionId}` }, + }); + const csrf = (await logoutPage.text()).match(/name="csrfToken" value="([^"]+)"/)?.[1]; + expect(csrf).toBeDefined(); + const logout = await fetch(`${base}/oauth/logout`, { + method: "POST", + headers: { + cookie: `__Host-thoughtstream_session=${sessionId}`, + "content-type": "application/x-www-form-urlencoded", + }, + body: new URLSearchParams({ csrfToken: csrf! }), + redirect: "manual", + }); + expect(logout.status).toBe(303); + expect(currentGeneration).toBeUndefined(); + }); + + test("keeps Basic read-only and forwards one CSRF-checked OAuth Review mutation under a body-bound capability", async () => { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-review-proxy-")); + roots.push(root); + const fixture = proxyOAuthFixture(root); + const reviewCapability = Buffer.alloc(32, 23); + const upstreamRequests: Array<{ + method: string | undefined; + path: string | undefined; + headers: http.IncomingHttpHeaders; + body: Buffer; + }> = []; + const upstream = http.createServer(async (request, response) => { + const parts: Buffer[] = []; + for await (const part of request) parts.push(Buffer.isBuffer(part) ? part : Buffer.from(part)); + upstreamRequests.push({ + method: request.method, + path: request.url, + headers: request.headers, + body: Buffer.concat(parts), + }); + response.writeHead(201, { "content-type": "application/json" }); + response.end('{"decisionEventId":"fixture"}'); + }); + servers.push(upstream); + const upstreamPort = await listen(upstream); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth: fixture.oauth, + basicFallbackEnabled: true, + reviewCapability, + }); + servers.push(proxy); + const base = baseUrl(proxy); + const basicAuthorization = basic("cameron", "correct-horse-battery-staple-private"); + const basicSession = await fetch(`${base}/inspector/api/session`, { + headers: { authorization: basicAuthorization }, + }); + expect(await basicSession.json()).toEqual({ reviewWriteEnabled: false }); + const basicWrite = await fetch(`${base}/inspector/api/reviews/review%3Aitem/decisions`, { + method: "POST", + headers: { authorization: basicAuthorization, "content-type": "application/json" }, + body: '{"disposition":"skip"}', + }); + expect(basicWrite.status).toBe(403); + expect(upstreamRequests).toHaveLength(0); + + const login = await fetch(`${base}/oauth/login`, { method: "POST", redirect: "manual" }); + const applicationState = cookieFromHeader(login.headers.get("set-cookie")!, "__Host-thoughtstream_oauth"); + const callback = await fetch(`${base}/oauth/callback?state=sdk-protocol-state-1234567890abcdef&code=fixture`, { + headers: { cookie: `__Host-thoughtstream_oauth=${applicationState}` }, + redirect: "manual", + }); + const sessionId = cookieFromHeader(callback.headers.get("set-cookie")!, "__Host-thoughtstream_session"); + const cookie = `__Host-thoughtstream_session=${sessionId}`; + const session = await fetch(`${base}/inspector/api/session`, { headers: { cookie } }); + const sessionBody = await session.json() as { reviewWriteEnabled: boolean; csrfToken: string }; + expect(sessionBody.reviewWriteEnabled).toBe(true); + expect(sessionBody.csrfToken).toMatch(/^[A-Za-z0-9_-]+$/); + + const body = JSON.stringify({ + disposition: "skip", + reasonCodes: [], + responseTags: [], + trainingEligible: false, + submissionId: "submission-proxy-canary-0001", + }); + const missingCsrf = await fetch(`${base}/inspector/api/reviews/review%3Aitem/decisions`, { + method: "POST", + headers: { cookie, "content-type": "application/json" }, + body, + }); + expect(missingCsrf.status).toBe(403); + expect(upstreamRequests).toHaveLength(0); + const written = await fetch(`${base}/inspector/api/reviews/review%3Aitem/decisions`, { + method: "POST", + headers: { + cookie, + "content-type": "application/json", + [REVIEW_CSRF_HEADER]: sessionBody.csrfToken, + }, + body, + }); + expect(written.status).toBe(201); + expect(upstreamRequests).toHaveLength(1); + const forwarded = upstreamRequests[0]!; + expect(forwarded.method).toBe("POST"); + expect(forwarded.path).toBe("/api/reviews/review%3Aitem/decisions"); + expect(forwarded.body.toString("utf8")).toBe(body); + expect(forwarded.headers.authorization).toBeUndefined(); + expect(forwarded.headers.cookie).toBeUndefined(); + expect(forwarded.headers[REVIEW_CSRF_HEADER]).toBeUndefined(); + const verifier = new ReviewCapabilityVerifier(reviewCapability); + expect(verifier.verify(forwarded.headers, "POST", forwarded.path!, forwarded.body)).toBe(true); + }); + + test("keeps Basic break-glass independent and stops advertising it when disabled", async () => { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-break-glass-")); + roots.push(root); + const fixture = proxyOAuthFixture(root); + let upstreamRequests = 0; + const upstream = http.createServer((_request, response) => { + upstreamRequests += 1; + response.end("private"); + }); + servers.push(upstream); + const upstreamPort = await listen(upstream); + const disabled = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth: fixture.oauth, + basicFallbackEnabled: false, + }); + servers.push(disabled); + + const rejectedBasic = await fetch(`${baseUrl(disabled)}/inspector/api/private`, { + headers: { authorization: basic("cameron", "correct-horse-battery-staple-private") }, + }); + expect(rejectedBasic.status).toBe(401); + expect(rejectedBasic.headers.get("www-authenticate")).toBeNull(); + expect(upstreamRequests).toBe(0); + + const enabled = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth: fixture.oauth, + basicFallbackEnabled: true, + }); + servers.push(enabled); + const acceptedBasic = await fetch(`${baseUrl(enabled)}/inspector/api/private`, { + headers: { authorization: basic("cameron", "correct-horse-battery-staple-private") }, + }); + expect(acceptedBasic.status).toBe(200); + expect(await acceptedBasic.text()).toBe("private"); + expect(fixture.protocol.restoreCalls).toEqual([]); + }); + + test("surfaces callback quarantine exhaustion as operator recycle required", async () => { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-quarantine-capacity-")); + roots.push(root); + const fixture = proxyOAuthFixture(root); + fixture.protocol.callbackError = new OAuthCallbackQuarantineCapacityError(8); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth: fixture.oauth, + }); + servers.push(proxy); + const base = baseUrl(proxy); + const login = await fetch(`${base}/oauth/login`, { method: "POST", redirect: "manual" }); + const state = cookieFromHeader(login.headers.get("set-cookie")!, "__Host-thoughtstream_oauth"); + const callback = await fetch(`${base}/oauth/callback?state=sdk-protocol-state-1234567890abcdef&code=fixture`, { + headers: { cookie: `__Host-thoughtstream_oauth=${state}` }, + redirect: "manual", + }); + expect(callback.status).toBe(503); + expect(callback.headers.get("retry-after")).toBe("60"); + expect(callback.headers.get("set-cookie")).toBeNull(); + expect(await callback.text()).toBe("OAuth callback capacity reached. Operator recycle required.\n"); + + fixture.protocol.callbackError = undefined; + const retry = await fetch(`${base}/oauth/callback?state=sdk-protocol-state-1234567890abcdef&code=fixture`, { + headers: { cookie: `__Host-thoughtstream_oauth=${state}` }, + redirect: "manual", + }); + expect(retry.status).toBe(302); + expect(retry.headers.get("location")).toBe("/inspector/"); + }); + + test("rate-limits OAuth initiation and callback before invoking the SDK", async () => { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-rate-limit-")); + roots.push(root); + const fixture = proxyOAuthFixture(root); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth: fixture.oauth, + oauthRateLimiter: { + login: () => ({ allowed: false, retryAfterSeconds: 17 }), + callback: () => ({ allowed: false, retryAfterSeconds: 23 }), + }, + }); + servers.push(proxy); + const base = baseUrl(proxy); + + const login = await fetch(`${base}/oauth/login`, { method: "POST", redirect: "manual" }); + expect(login.status).toBe(429); + expect(login.headers.get("retry-after")).toBe("17"); + const callback = await fetch(`${base}/oauth/callback?state=sdk-protocol-state-1234567890abcdef&code=private-code`); + expect(callback.status).toBe(429); + expect(callback.headers.get("retry-after")).toBe("23"); + expect(fixture.protocol.authorizeCalls).toBe(0); + expect(fixture.protocol.callbackCalls).toBe(0); + }); + + test("aborts OAuth discovery when the initiating browser disconnects", async () => { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-disconnect-")); + roots.push(root); + let authorizeStarted!: () => void; + const started = new Promise((resolve) => { authorizeStarted = resolve; }); + let aborted = false; + const fixture = proxyOAuthFixture(root, { + authorize: async (_handle, options) => { + authorizeStarted(); + await new Promise((_resolve, reject) => { + options.signal?.addEventListener("abort", () => { + aborted = true; + reject(new Error("aborted")); + }, { once: true }); + }); + throw new Error("unreachable"); + }, + }); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + username: "cameron", + password: "correct-horse-battery-staple-private", + oauth: fixture.oauth, + }); + servers.push(proxy); + const address = proxy.address(); + if (!address || typeof address === "string") throw new Error("Missing proxy address"); + const request = http.request({ + hostname: "127.0.0.1", + port: address.port, + path: "/oauth/login", + method: "POST", + }); + request.on("error", () => undefined); + request.end(); + await started; + request.destroy(); + await waitFor(() => aborted); + expect(aborted).toBe(true); + }); + + test("returns a generic content-dark response when the inspector is unavailable", async () => { + const unavailablePort = await unusedPort(); + const proxy = await startAuthenticatedInspectorProxy({ + port: 0, + upstreamPort: unavailablePort, + username: "cameron", + password: "another-correctly-long-private-password", + }); + servers.push(proxy); + const response = await fetch(`${baseUrl(proxy)}/inspector/`, { + headers: { authorization: basic("cameron", "another-correctly-long-private-password") }, + }); + expect(response.status).toBe(502); + expect(await response.text()).toBe("Inspector unavailable.\n"); + }); + + test("refuses public binds, non-loopback upstreams, short passwords, and malformed environment", async () => { + await expect(startAuthenticatedInspectorProxy({ + host: "0.0.0.0", + username: "cameron", + password: "correctly-long-private-password", + })).rejects.toThrow("loopback"); + await expect(startAuthenticatedInspectorProxy({ + upstreamHost: "example.com", + username: "cameron", + password: "correctly-long-private-password", + })).rejects.toThrow("upstream must be loopback"); + await expect(startAuthenticatedInspectorProxy({ + username: "cameron", + password: "too-short", + })).rejects.toThrow("at least 20"); + await expect(startAuthenticatedInspectorProxy({ + username: "cameron:admin", + password: "correctly-long-private-password", + })).rejects.toThrow("username"); + await expect(startAuthenticatedInspectorProxy({ + username: "cameron", + password: "correctly-long-private-password", + basicFallbackEnabled: false, + })).rejects.toThrow("requires OAuth or enabled Basic fallback"); + expect(() => authenticatedProxyOptionsFromEnv({ + PROXY_USER: "cameron", + PROXY_PASSWORD_B64: "not canonical base64 !!!", + })).toThrow("canonical base64"); + expect(() => authenticatedProxyOptionsFromEnv({ PROXY_USER: "cameron" })).toThrow("PROXY_PASSWORD_B64"); + }); + + test("loads a canonical base64 password without retaining the encoded value", () => { + const options = authenticatedProxyOptionsFromEnv({ + PROXY_USER: "cameron", + PROXY_PASSWORD_B64: Buffer.from("correct-horse-battery-staple-private").toString("base64"), + PROXY_PORT: "4319", + PROXY_UPSTREAM_PORT: "4317", + }); + expect(options).toEqual({ + host: "127.0.0.1", + port: 4319, + upstreamHost: "127.0.0.1", + upstreamPort: 4317, + username: "cameron", + password: "correct-horse-battery-staple-private", + basicFallbackEnabled: true, + projectRoot: process.cwd(), + }); + }); +}); + +function proxyOAuthFixture( + root: string, + options: { authorize?: OAuthProtocolClient["authorize"] } = {}, +): { oauth: InspectorOAuthAuth; protocol: ProxyOAuthProtocol } { + const key = Buffer.alloc(32, 12); + const browserSessions = new SecureJsonStore({ + directory: root, + name: "proxy-browser", + key, + maxEntries: 8, + maxSerializedBytes: 32 * 1024, + }); + const flowStates = new SecureJsonStore<{ createdAt: number }>({ + directory: root, + name: "proxy-flow", + key, + maxEntries: 64, + maxSerializedBytes: 32 * 1024, + ttlMs: 15 * 60_000, + }); + const protocol = new ProxyOAuthProtocol(options.authorize); + return { + oauth: new InspectorOAuthAuth({ + protocol, + allowedDid: "did:plc:allowed", + expectedHandle: "cameron.stream", + browserSessions, + flowStates, + }), + protocol, + }; +} + +class ProxyOAuthProtocol implements OAuthProtocolClient { + readonly clientMetadata = { client_id: "https://thought.stream/oauth/client-metadata.json" }; + readonly jwks = { keys: [] }; + authorizeCalls = 0; + callbackCalls = 0; + readonly restoreCalls: string[] = []; + callbackError: Error | undefined; + private applicationState?: string; + private readonly stagedAttempts = new Map(); + private currentGeneration: number | undefined; + private nextGeneration = 0; + + constructor(private readonly customAuthorize?: OAuthProtocolClient["authorize"]) {} + + async authorize(handle: string, options: { state: string; scope: string; signal?: AbortSignal }): Promise { + this.authorizeCalls += 1; + this.applicationState = options.state; + if (this.customAuthorize) return this.customAuthorize(handle, options); + return new URL("https://pds.example/authorize?request_uri=urn:ietf:params:oauth:request_uri:fixture"); + } + + async callback(_params: URLSearchParams, attemptId: string): Promise<{ session: { did: string }; state: string | null }> { + this.callbackCalls += 1; + if (this.callbackError) throw this.callbackError; + this.stagedAttempts.set(attemptId, "did:plc:allowed"); + return { session: { did: "did:plc:allowed" }, state: this.applicationState ?? null }; + } + + expireCallbackAttempt(): void {} + + async dropCallbackAttempt(attemptId: string): Promise { + this.stagedAttempts.delete(attemptId); + } + + async promoteCallbackSession(attemptId: string, did: string, guard: () => boolean): Promise { + if (!guard() || this.stagedAttempts.get(attemptId) !== did) throw new Error("missing staged session"); + this.stagedAttempts.delete(attemptId); + this.currentGeneration = ++this.nextGeneration; + return this.currentGeneration; + } + + isCurrentGeneration(_did: string, generation: number): boolean { + return this.currentGeneration === generation; + } + + async discardPromotedSession(_did: string, generation: number): Promise { + if (this.currentGeneration === generation) this.currentGeneration = undefined; + } + + async restore(did: string, generation: number): Promise<{ did: string }> { + this.restoreCalls.push(did); + if (this.currentGeneration !== generation) throw new Error("stale generation"); + return { did }; + } + + async revoke(): Promise {} + async deleteSession(): Promise {} +} + +async function waitFor(predicate: () => boolean): Promise { + for (let attempt = 0; attempt < 100; attempt += 1) { + if (predicate()) return; + await new Promise((resolve) => setTimeout(resolve, 5)); + } + throw new Error("Timed out waiting for condition"); +} + +function cookieFromHeader(header: string, name: string): string { + const match = header.match(new RegExp(`(?:^|,\\s*)${name}=([^;]+)`)); + if (!match?.[1]) throw new Error(`Missing ${name} cookie`); + return match[1]; +} + +function basic(username: string, password: string): string { + return `Basic ${Buffer.from(`${username}:${password}`).toString("base64")}`; +} + +async function listen(server: http.Server): Promise { + await new Promise((resolve, reject) => { + server.once("error", reject); + server.listen(0, "127.0.0.1", () => resolve()); + }); + const address = server.address(); + if (!address || typeof address === "string") throw new Error("Missing server address"); + return address.port; +} + +function baseUrl(server: http.Server): string { + const address = server.address(); + if (!address || typeof address === "string") throw new Error("Missing server address"); + return `http://127.0.0.1:${address.port}`; +} + +async function unusedPort(): Promise { + const server = http.createServer(); + const port = await listen(server); + await new Promise((resolve) => server.close(() => resolve())); + return port; +} diff --git a/test/conceptualizer.test.ts b/test/conceptualizer.test.ts new file mode 100644 index 0000000..cc08684 --- /dev/null +++ b/test/conceptualizer.test.ts @@ -0,0 +1,319 @@ +import fs from "node:fs/promises"; +import { afterEach, describe, expect, test } from "vitest"; +import { + CONCEPTUALIZATION_OUTPUT_CONTRACT, + createOutputContractRegistry, + OutputContractValidationError, +} from "../src/agents/output-contracts.js"; +import { ThoughtAgentRuntime } from "../src/agents/runtime.js"; +import type { AgentRunner, ThoughtAgentDeclaration } from "../src/agents/types.js"; +import { createDefaultRegistry } from "../src/events/registry.js"; +import type { JazzThoughtStore } from "../src/jazz/store.js"; +import { temporaryProject, testStore } from "./helpers.js"; + +const stores: JazzThoughtStore[] = []; +const roots: string[] = []; + +afterEach(async () => { + await Promise.all(stores.splice(0).map((store) => store.close())); + await Promise.all(roots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); +}); + +describe("conceptualizer output", () => { + test("rejects manually constructed contract and event mismatches at runtime registration", async () => { + const project = await temporaryProject(); + roots.push(project); + const store = testStore(project); + stores.push(store); + const runtime = new ThoughtAgentRuntime(store, []); + const mismatched = conceptualizerDeclaration("conceptualizer-mismatch"); + mismatched.outputEventType = "stream.thought.derived.topics"; + mismatched.emit = ["stream.thought.derived.topics"]; + + await expect(runtime.registerDeclarations([mismatched])) + .rejects.toThrow("must bind conceptualization output and concept graph events together"); + }); + + test("validates one strict bounded graph and rejects invalid indices", () => { + const registry = createOutputContractRegistry(); + const valid = { + summary: "Two related ideas", + concepts: [ + { text: "agent memory", relationship: "DESCRIBES" }, + { text: "durable receipts", relationship: "SUPPORTS" }, + ], + links: [{ fromIndex: 1, toIndex: 0, relationship: "SUPPORTS" }], + confidence: 0.8, + }; + + expect(registry.validate(CONCEPTUALIZATION_OUTPUT_CONTRACT.identity, valid)).toEqual(valid); + expect(() => registry.validate(CONCEPTUALIZATION_OUTPUT_CONTRACT.identity, { + ...valid, + links: [{ fromIndex: 2, toIndex: 0, relationship: "SUPPORTS" }], + })).toThrow(OutputContractValidationError); + expect(() => registry.validate(CONCEPTUALIZATION_OUTPUT_CONTRACT.identity, { + ...valid, + concepts: [{ text: "Agent Memory", relationship: "DESCRIBES" }], + })).toThrow(OutputContractValidationError); + expect(() => registry.validate(CONCEPTUALIZATION_OUTPUT_CONTRACT.identity, { + ...valid, + publicationUri: "at://synthetic/not-a-receipt", + })).toThrow(OutputContractValidationError); + }); + + test("validates correction proposals against the contract named in their payload", () => { + const registry = createDefaultRegistry(); + const payload = { + originalRunId: "run-original", + repairRequestEventId: "evt-request", + repairRunId: "run-repair", + originalTriggerEventId: "evt-source", + sourceRootEventId: "evt-source", + outputContract: { ...CONCEPTUALIZATION_OUTPUT_CONTRACT.identity }, + originalModel: { provider: "tinker", id: "public-base-model" }, + repairModel: { provider: "tinker", id: "repair-model" }, + structuredOutput: { + summary: "One concept", + concepts: [{ text: "agent memory", relationship: "DESCRIBES" }], + confidence: 0.8, + }, + }; + + expect(registry.validate( + "stream.thought.derived.output.correction.proposed", + 1, + payload, + )).toMatchObject(payload); + expect(() => registry.validate( + "stream.thought.derived.output.correction.proposed", + 1, + { + ...payload, + structuredOutput: { + ...payload.structuredOutput, + links: [{ fromIndex: 0, toIndex: 4, relationship: "RELATES_TO" }], + }, + }, + )).toThrow("Correction proposal must satisfy its registered output contract"); + }); + + test("rejects graph envelopes whose common summary disagrees with structured output", () => { + const registry = createDefaultRegistry(); + expect(() => registry.validate("stream.thought.derived.concept.graph", 1, { + runId: "run-graph", + executionKey: "execution-graph", + inputEventId: "evt-source", + inputSourceSequence: 1, + summary: "Contradictory envelope", + confidence: 0.8, + outputContract: { ...CONCEPTUALIZATION_OUTPUT_CONTRACT.identity }, + structuredOutput: { + summary: "Canonical graph", + concepts: [{ text: "agent memory", relationship: "DESCRIBES" }], + confidence: 0.8, + }, + })).toThrow("Summary must equal canonical structured output"); + }); + + test("rejects graph envelopes that claim a different contract identity", () => { + const registry = createDefaultRegistry(); + expect(() => registry.validate("stream.thought.derived.concept.graph", 1, { + runId: "run-graph", + executionKey: "execution-graph", + inputEventId: "evt-source", + inputSourceSequence: 1, + summary: "Canonical graph", + confidence: 0.8, + outputContract: { + ...CONCEPTUALIZATION_OUTPUT_CONTRACT.identity, + sha256: "0".repeat(64), + }, + structuredOutput: { + summary: "Canonical graph", + concepts: [{ text: "agent memory", relationship: "DESCRIBES" }], + confidence: 0.8, + }, + })).toThrow("Concept graph must name the canonical conceptualization contract identity"); + }); + + test("rejects direct public-source graph appends at the durable store boundary", async () => { + const project = await temporaryProject(); + roots.push(project); + const store = testStore(project); + stores.push(store); + const payload = { + runId: "run-direct-graph", + executionKey: "execution-direct-graph", + inputEventId: "evt-source", + inputSourceSequence: 1, + summary: "Canonical graph", + confidence: 0.8, + outputContract: { ...CONCEPTUALIZATION_OUTPUT_CONTRACT.identity }, + structuredOutput: { + summary: "Canonical graph", + concepts: [{ text: "agent memory", relationship: "DESCRIBES" }], + confidence: 0.8, + }, + }; + + await expect(store.appendEvent({ + type: "stream.thought.derived.concept.graph", + schemaVersion: 1, + source: "agent:direct-fixture", + sourceKind: "agent", + externalId: "direct-public-graph", + idempotencyKey: "direct-public-graph", + occurredAt: "2026-07-25T20:00:00.000Z", + actor: "direct-fixture", + correlationId: "direct-public-graph", + privacy: "public-source", + payload, + })).rejects.toThrow("requires private or stricter privacy"); + + const inserted = await store.appendEvent({ + type: "stream.thought.derived.concept.graph", + schemaVersion: 1, + source: "agent:direct-fixture", + sourceKind: "agent", + externalId: "direct-private-graph", + idempotencyKey: "direct-private-graph", + occurredAt: "2026-07-25T20:00:00.000Z", + actor: "direct-fixture", + correlationId: "direct-private-graph", + privacy: "private", + payload, + }); + expect(inserted.event.privacy).toBe("private"); + }); + + test("settles one private graph event atomically with exact source and run lineage", async () => { + const project = await temporaryProject(); + roots.push(project); + const store = testStore(project); + stores.push(store); + const source = (await store.appendEvent({ + type: "stream.thought.source.rss.item", + schemaVersion: 1, + source: "rss:concept-fixture", + sourceKind: "rss", + externalId: "concept-fixture", + idempotencyKey: "concept-fixture", + occurredAt: "2026-07-25T20:00:00.000Z", + actor: "rss:concept-fixture", + correlationId: "concept-fixture", + privacy: "public-source", + payload: { title: "Durable memory requires receipts" }, + })).event; + const runner: AgentRunner = { + mode: "deterministic", + run: async () => ({ + summary: "Receipts make memory changes inspectable", + concepts: [ + { text: "durable memory", relationship: "DESCRIBES" }, + { text: "execution receipts", relationship: "SUPPORTS" }, + ], + links: [{ fromIndex: 1, toIndex: 0, relationship: "SUPPORTS" }], + confidence: 0.9, + }), + }; + const runtime = new ThoughtAgentRuntime(store, [runner]); + + const [result] = await runtime.consumeBacklog([conceptualizerDeclaration("conceptualizer-valid")]); + const graph = result?.derivedEvent; + + expect(result?.error).toBeUndefined(); + expect(graph).toMatchObject({ + type: "stream.thought.derived.concept.graph", + parentEventId: source.id, + rootEventId: source.id, + privacy: "private", + payload: { + inputEventId: source.id, + inputSourceSequence: source.sourceSequence, + structuredOutput: { + concepts: [{ text: "durable memory" }, { text: "execution receipts" }], + links: [{ fromIndex: 1, toIndex: 0, relationship: "SUPPORTS" }], + }, + }, + }); + expect(graph?.payload).not.toHaveProperty("uri"); + expect(graph?.payload).not.toHaveProperty("cid"); + expect((await store.listRuns())[0]?.outputEventIds).toEqual([graph?.id]); + }); + + test("emits no graph when any link falls outside the concept array", async () => { + const project = await temporaryProject(); + roots.push(project); + const store = testStore(project); + stores.push(store); + await store.appendEvent({ + type: "stream.thought.source.rss.item", + schemaVersion: 1, + source: "rss:invalid-concept-fixture", + sourceKind: "rss", + externalId: "invalid-concept-fixture", + idempotencyKey: "invalid-concept-fixture", + occurredAt: "2026-07-25T20:00:00.000Z", + actor: "rss:invalid-concept-fixture", + correlationId: "invalid-concept-fixture", + privacy: "private", + payload: { title: "Invalid graph fixture" }, + }); + const runner: AgentRunner = { + mode: "deterministic", + run: async () => ({ + summary: "Invalid graph", + concepts: [{ text: "one concept", relationship: "DESCRIBES" }], + links: [{ fromIndex: 0, toIndex: 4, relationship: "RELATES_TO" }], + confidence: 0.5, + }), + }; + const runtime = new ThoughtAgentRuntime(store, [runner]); + + const [result] = await runtime.consumeBacklog([ + conceptualizerDeclaration("conceptualizer-invalid", "rss:invalid-concept-fixture"), + ]); + + expect(result).toMatchObject({ error: "Agent final output rejected by canonical contract" }); + expect((await store.listEvents({ types: ["stream.thought.derived.concept.graph"] }))).toHaveLength(0); + expect((await store.listEvents({ types: ["stream.thought.agent.repair.requested"] }))).toHaveLength(0); + expect((await store.listRuns())[0]?.result).toMatchObject({ + failureDiagnostic: { reason: "output-contract-invalid" }, + }); + }); +}); + +function conceptualizerDeclaration( + id: string, + source = "rss:concept-fixture", +): ThoughtAgentDeclaration { + return { + id, + version: 1, + name: id, + description: "Conceptualizer fixture", + mode: "deterministic", + role: "standard", + outputContract: { ...CONCEPTUALIZATION_OUTPUT_CONTRACT.identity }, + declarationFingerprint: `${id}-fingerprint`, + provider: "openai-compatible", + model: "fixture-model", + eventTypes: ["stream.thought.source.rss.item"], + compiledEventTypes: ["stream.thought.source.rss.item"], + sourcePatterns: [source], + acceptedPrivacy: ["public-source", "private"], + initialReplay: "beginning", + outputEventType: "stream.thought.derived.concept.graph", + emit: ["stream.thought.derived.concept.graph"], + promptRef: "prompts/conceptualizer.md", + systemPrompt: "Extract concepts.", + enabled: true, + maxEvents: 1, + maxInputChars: 64_000, + contextStrategy: "single-event", + maxOutputTokens: 2_000, + timeoutMs: 60_000, + tools: [], + externalActions: false, + }; +} diff --git a/test/declarations.test.ts b/test/declarations.test.ts index 431eafb..1471ed1 100644 --- a/test/declarations.test.ts +++ b/test/declarations.test.ts @@ -11,6 +11,55 @@ afterEach(async () => { }); describe("agent declarations", () => { + test("binds conceptualization contracts to graph events and strict JSON mode", async () => { + const source = await fs.readFile( + path.join(process.cwd(), "agents", "conceptualizer.example.yaml"), + "utf8", + ); + const prompt = await fs.readFile( + path.join(process.cwd(), "prompts", "conceptualizer.md"), + "utf8", + ); + const cases = [ + { + name: "wrong-event", + declaration: source.replace( + " - stream.thought.derived.concept.graph", + " - stream.thought.derived.topics", + ), + message: "must be selected together", + }, + { + name: "wrong-contract", + declaration: source.replace( + "stream.thought.output.conceptualization", + "stream.thought.output.observation", + ), + message: "must be selected together", + }, + { + name: "conversation-text", + declaration: source.replace( + " model: gpt-4.1-mini", + " model: gpt-4.1-mini\n outputMode: conversation-text", + ), + message: "strict JSON", + }, + ]; + + for (const current of cases) { + const project = await temporaryProject(`thoughtstream-declaration-${current.name}-`); + roots.push(project); + const agents = path.join(project, "agents"); + const prompts = path.join(project, "prompts"); + await fs.mkdir(agents, { recursive: true }); + await fs.mkdir(prompts, { recursive: true }); + await fs.writeFile(path.join(agents, "conceptualizer.yaml"), current.declaration); + await fs.writeFile(path.join(prompts, "conceptualizer.md"), prompt); + await expect(loadAgentDeclarations(agents, {})).rejects.toThrow(current.message); + } + }); + test("resolves a Tinker capability tier to a concrete persisted model", async () => { const declarations = await loadAgentDeclarations(path.join(process.cwd(), "agents"), testDeclarationEnvironment); expect(declarations.find((declaration) => declaration.id === "bluesky-enrichment-observer")).toMatchObject({ @@ -51,6 +100,23 @@ describe("agent declarations", () => { model: "fixture/escalation-model", tools: [], }); + expect(declarations.find((declaration) => declaration.id === "conceptualizer")).toMatchObject({ + version: 1, + enabled: false, + mode: "pi", + provider: "openai-compatible", + providerProfile: "openai-json-default", + model: "gpt-4.1-mini", + outputContract: { + id: "stream.thought.output.conceptualization", + version: 1, + sha256: expect.stringMatching(/^[a-f0-9]{64}$/), + }, + outputEventType: "stream.thought.derived.concept.graph", + contextStrategy: "atproto-batch", + atprotoObjectContext: true, + tools: [], + }); expect(declarations.find((declaration) => declaration.id === "resident-letta-conversation")).toMatchObject({ version: 3, enabled: false, diff --git a/test/harness-container.test.ts b/test/harness-container.test.ts index 9459d57..d69d818 100644 --- a/test/harness-container.test.ts +++ b/test/harness-container.test.ts @@ -223,6 +223,7 @@ function fixtureProfile(): ProviderProfile { allowedModels: new Set(["fixture-model"]), imageInputModels: new Set(), jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: false, requestTimeoutMs: 5_000, maxRequestBytes: 64 * 1024, maxResponseBytes: 64 * 1024, diff --git a/test/helpers.ts b/test/helpers.ts index 79cc539..0a4b422 100644 --- a/test/helpers.ts +++ b/test/helpers.ts @@ -28,7 +28,9 @@ export function testInferenceAccountingPolicy(maxCalls = 100): InferenceBudgetPo }; } -export function testStore(projectRoot: string): JazzThoughtStore { +export function testStore( + projectRoot: string, +): JazzThoughtStore { return new JazzThoughtStore({ projectRoot, appId: `thoughtstream-test-${randomUUID()}`, diff --git a/test/inspector.test.ts b/test/inspector.test.ts index 7752c01..737f11d 100644 --- a/test/inspector.test.ts +++ b/test/inspector.test.ts @@ -97,7 +97,12 @@ describe("thought stream inspector", () => { const page = await fetch(base); expect(page.status).toBe(200); - expect(await page.text()).toContain("thought stream inspector"); + const pageHtml = await page.text(); + expect(pageHtml).toContain("thought stream inspector"); + expect(pageHtml).toContain("fetch('api/snapshot')"); + expect(pageHtml).not.toContain("fetch('/api/snapshot')"); + expect(pageHtml).toContain("cursor:pointer; overflow-wrap:anywhere"); + expect(pageHtml).toContain("font-size:14px; overflow-wrap:anywhere"); expect(page.headers.get("content-security-policy")).toContain("frame-ancestors 'none'"); const snapshot = await (await fetch(`${base}/api/snapshot`)).json() as { @@ -236,6 +241,19 @@ describe("thought stream inspector", () => { outputEventId: output.id, }, })).event; + const cursor = (await store.appendEvent({ + type: "stream.thought.connector.cursor.advanced", + schemaVersion: 1, + source: "filesystem:test", + sourceKind: "filesystem", + externalId: "cursor:filesystem:test:1", + idempotencyKey: "cursor:filesystem:test:1", + occurredAt: "2026-07-14T00:00:03.000Z", + actor: "filesystem:test", + correlationId: "scan-1", + privacy: "private", + payload: { cursor: { version: 1 } }, + })).event; const server = await startInspectorServer(store, { port: 0 }); servers.push(server); @@ -268,6 +286,7 @@ describe("thought stream inspector", () => { }], }); expect(snapshot.activity.items.some((item) => item.id === completed.id)).toBe(false); + expect(snapshot.activity.items.some((item) => item.id === cursor.id)).toBe(false); const detail = await (await fetch(`${base}/api/events/${encodeURIComponent(completed.id)}`)).json() as { agentActivities: Array<{ @@ -285,6 +304,67 @@ describe("thought stream inspector", () => { ]); }); + test("shows loaded adapter catalog selection without exposing the private checkpoint", async () => { + const project = await temporaryProject(); + roots.push(project); + const store = testStore(project); + stores.push(store); + const privateCheckpoint = "tinker://must-never-reach-inspector"; + const digest = "a".repeat(64); + const modelAdapter = { + id: "julia-adapter", + version: 1, + description: "Julia transformation adapter", + releasedAt: "2026-07-25T20:00:00.000Z", + manifestSha256: "b".repeat(64), + checkpointSelector: { kind: "env", reference: "THOUGHTSTREAM_JULIA_ADAPTER_CHECKPOINT" }, + providerProfile: "tinker-default", + baseModel: "Qwen/Qwen3.5-35B-A3B-Base", + dataset: { id: "julia-dataset", sha256: "c".repeat(64) }, + evals: [{ id: "julia-eval", sha256: "d".repeat(64) }], + capabilities: ["julia-dict-transform"], + privacyClass: "private", + exportClass: "restricted", + binding: { checkpointReferenceSha256: "e".repeat(64) }, + }; + await store.upsertAgent({ + id: "conceptualizer", + version: 1, + enabled: true, + spec: { + id: "conceptualizer", + version: 1, + adapterCatalogDigest: digest, + adapterCatalogGeneration: 7, + modelAdapter, + }, + specHash: "fixture-spec-hash", + updatedAt: "2026-07-25T20:00:00.000Z", + }); + + const server = await startInspectorServer(store, { port: 0 }); + servers.push(server); + const address = server.address(); + if (!address || typeof address === "string") throw new Error("Missing inspector address"); + const snapshot = await (await fetch(`http://127.0.0.1:${address.port}/api/snapshot`)).json() as { + adapterInventory: { + catalogs: Array<{ digest: string; generation: number }>; + selections: Array<{ id: string; version: number; enabled: boolean; adapter: { release: Record } }>; + }; + }; + + expect(snapshot.adapterInventory.catalogs).toEqual([{ digest, generation: 7 }]); + expect(snapshot.adapterInventory.selections).toEqual([expect.objectContaining({ + id: "conceptualizer", + version: 1, + enabled: true, + adapter: expect.objectContaining({ release: modelAdapter }), + })]); + const serialized = JSON.stringify(snapshot.adapterInventory); + expect(serialized).not.toContain(privateCheckpoint); + expect(serialized).not.toContain("/home/"); + }); + test("refuses a non-loopback bind", async () => { const project = await temporaryProject(); roots.push(project); diff --git a/test/jazz-store.test.ts b/test/jazz-store.test.ts index 6e4ad3c..a0ed1f9 100644 --- a/test/jazz-store.test.ts +++ b/test/jazz-store.test.ts @@ -1,5 +1,6 @@ import fs from "node:fs/promises"; import { afterEach, describe, expect, test } from "vitest"; +import type { EventCandidate } from "../src/events/types.js"; import type { JazzThoughtStore } from "../src/jazz/store.js"; import { temporaryProject, testStore } from "./helpers.js"; @@ -12,6 +13,30 @@ afterEach(async () => { }); describe("JazzThoughtStore", () => { + test("serializes concurrent producer appends to one source", async () => { + const root = await temporaryProject(); + roots.push(root); + const store = testStore(root); + stores.push(store); + const candidate = (index: number): EventCandidate => ({ + type: "stream.thought.source.rss.item", + schemaVersion: 1, + source: "rss:concurrent-producer", + sourceKind: "rss", + externalId: `item-${index}`, + idempotencyKey: `item-${index}`, + occurredAt: `2026-07-27T00:00:0${index}.000Z`, + actor: "rss:concurrent-producer", + correlationId: "concurrent-producer", + privacy: "public-source", + payload: { index }, + }); + + const results = await Promise.all([store.appendEvent(candidate(1)), store.appendEvent(candidate(2))]); + expect(results.map((result) => result.event.sourceSequence).sort((left, right) => left - right)).toEqual([1, 2]); + expect((await store.listEvents({ source: "rss:concurrent-producer" }))).toHaveLength(2); + }); + test("appends one durable event for repeated source identity", async () => { const root = await temporaryProject(); roots.push(root); diff --git a/test/judgments.test.ts b/test/judgments.test.ts index 7b7bc39..3acde72 100644 --- a/test/judgments.test.ts +++ b/test/judgments.test.ts @@ -18,7 +18,7 @@ afterEach(async () => { }); describe("training judgments", () => { - test("exports public synthetic judgments through the minimized v2 format in owner-only files", async () => { + test("exports public synthetic judgments through the provenance-complete v3 format in owner-only files", async () => { const fixture = await completedRunFixture("public-source"); const { root, store, runId } = fixture; const replacement: JsonObject = { @@ -46,7 +46,7 @@ describe("training judgments", () => { const examples = await projectTrainingExamples(store); expect(examples).toHaveLength(1); expect(examples[0]).toMatchObject({ - format: "thoughtstream.training-example.v2", + format: "thoughtstream.training-example.v3", kind: "correct", judgment: { criterion: "topic-fidelity", criterionVersion: 1 }, input: { @@ -65,7 +65,7 @@ describe("training judgments", () => { agentVersion: 3, provider: "tinker", model: "Qwen/Qwen3.5-4B", - adapterRevision: "checkpoint-fixture", + executionAdapterRevision: "checkpoint-fixture", }, }); const exported = JSON.stringify(examples); @@ -86,18 +86,18 @@ describe("training judgments", () => { const destination = path.join(root, "exports", "training.jsonl"); const manifest = await writeTrainingJsonl(destination, examples); expect(manifest).toMatchObject({ - format: "thoughtstream.training-dataset-manifest.v2", + format: "thoughtstream.training-dataset-manifest.v4", datasetId: expect.stringMatching(/^sha256:/), examples: 1, kinds: { correct: 1 }, - models: ["tinker:Qwen/Qwen3.5-4B@checkpoint-fixture"], + models: ["tinker:Qwen/Qwen3.5-4B@execution:checkpoint-fixture"], }); expect(JSON.parse((await fs.readFile(destination, "utf8")).trim())).toMatchObject({ - format: "thoughtstream.training-example.v2", + format: "thoughtstream.training-example.v3", chosen: { summary: "Corrected public output" }, }); const storedManifest = JSON.parse(await fs.readFile(`${destination}.manifest.json`, "utf8")) as Record; - expect(storedManifest).toMatchObject({ format: "thoughtstream.training-dataset-manifest.v2", examples: 1 }); + expect(storedManifest).toMatchObject({ format: "thoughtstream.training-dataset-manifest.v4", examples: 1 }); expect(storedManifest).not.toHaveProperty("judgmentEventIds"); expect((await fs.stat(destination)).mode & 0o777).toBe(0o600); expect((await fs.stat(`${destination}.manifest.json`)).mode & 0o777).toBe(0o600); @@ -233,6 +233,119 @@ describe("training judgments", () => { expect(await projectTrainingExamples(fixture.store, { includeSensitivePrivate: true })).toHaveLength(1); }); + test("requires explicit authority for private state on the compared side of a pair", async () => { + const fixture = await completedRunFixture("public-source"); + const compared = await appendComparedRun(fixture, "public", "sensitive-pair", "sensitive"); + await expect(recordJudgment(fixture.store, { + runId: fixture.runId, + comparedRunId: compared.runId, + kind: "prefer", + criterion: "pairwise-sensitive", + criterionVersion: 1, + qualityEligible: true, + externalExportEligible: true, + })).rejects.toThrow("requires explicit authorization"); + const judgment = await recordJudgment(fixture.store, { + runId: fixture.runId, + comparedRunId: compared.runId, + kind: "prefer", + criterion: "pairwise-sensitive", + criterionVersion: 1, + qualityEligible: true, + externalExportEligible: true, + sensitiveExternalExportAuthorized: true, + }); + expect(judgment.privacy).toBe("sensitive"); + expect(await projectTrainingExamples(fixture.store)).toEqual([]); + const examples = await projectTrainingExamples(fixture.store, { includeSensitivePrivate: true }); + expect(examples).toHaveLength(1); + expect(examples[0]).toMatchObject({ + input: { event: { privacy: "sensitive" } }, + comparedProvenance: { modelAdapter: compared.modelAdapter }, + }); + }); + + test("rejects pairwise export when the primary run uses a forbidden adapter", async () => { + const fixture = await completedRunFixture("public-source"); + const compared = await appendComparedRun(fixture, "public", "primary-forbidden-control"); + const primary = await fixture.store.getRun(fixture.runId); + if (!primary) throw new Error("Missing primary run"); + primary.modelAdapter = { + ...compared.modelAdapter, + id: "primary-forbidden-adapter", + manifestSha256: "4".repeat(64), + exportClass: "forbidden", + binding: { checkpointReferenceSha256: "5".repeat(64) }, + }; + await fixture.store.upsertRun(primary); + await recordJudgment(fixture.store, { + runId: fixture.runId, + comparedRunId: compared.runId, + kind: "prefer", + criterion: "pairwise-primary-forbidden", + criterionVersion: 1, + qualityEligible: true, + externalExportEligible: true, + }); + expect(await projectTrainingExamples(fixture.store, { + includeRestrictedModelAdapters: true, + })).toEqual([]); + }); + + test("gates both sides of pairwise export and emits exact compared execution and learned-adapter provenance", async () => { + const fixture = await completedRunFixture("public-source"); + const forbidden = await appendComparedRun(fixture, "forbidden", "forbidden"); + await recordJudgment(fixture.store, { + runId: fixture.runId, + comparedRunId: forbidden.runId, + kind: "prefer", + criterion: "pairwise-forbidden", + criterionVersion: 1, + qualityEligible: true, + externalExportEligible: true, + }); + expect(await projectTrainingExamples(fixture.store, { + includeRestrictedModelAdapters: true, + })).toEqual([]); + + const restricted = await appendComparedRun(fixture, "restricted", "restricted"); + await recordJudgment(fixture.store, { + runId: fixture.runId, + comparedRunId: restricted.runId, + kind: "prefer", + criterion: "pairwise-restricted", + criterionVersion: 1, + qualityEligible: true, + externalExportEligible: true, + }); + expect(await projectTrainingExamples(fixture.store)).toEqual([]); + const [example] = await projectTrainingExamples(fixture.store, { + includeRestrictedModelAdapters: true, + }); + expect(example).toMatchObject({ + format: "thoughtstream.training-example.v3", + kind: "prefer", + provenance: { + executionAdapterRevision: "checkpoint-fixture", + model: "Qwen/Qwen3.5-4B", + }, + comparedProvenance: { + executionAdapterRevision: "compared-execution-restricted", + model: "Qwen/Qwen3.5-35B-A3B-Base", + modelAdapter: restricted.modelAdapter, + adapterCatalogDigest: "7".repeat(64), + adapterCatalogGeneration: 1, + }, + comparedTrajectory: [{ sequence: 1, type: "pi.compared-restricted" }], + }); + const manifest = await writeTrainingJsonl(path.join(fixture.root, "pairwise", "training.jsonl"), [example!]); + expect(manifest.models).toHaveLength(2); + expect(manifest.models).toEqual(expect.arrayContaining([ + expect.stringContaining("@execution:checkpoint-fixture"), + expect.stringContaining(`@execution:compared-execution-restricted@model-adapter:${restricted.modelAdapter.id}@1#${restricted.modelAdapter.manifestSha256}#${restricted.modelAdapter.binding.checkpointReferenceSha256}@catalog:${"7".repeat(64)}:g1`), + ])); + }); + test("does not reinterpret a legacy private exportEligible judgment as declassification", async () => { const fixture = await completedRunFixture("private"); const legacy = await appendLegacyJudgment(fixture, "private"); @@ -247,7 +360,7 @@ describe("training judgments", () => { const examples = await projectTrainingExamples(fixture.store); expect(examples).toHaveLength(1); expect(examples[0]).toMatchObject({ - format: "thoughtstream.training-example.v2", + format: "thoughtstream.training-example.v3", kind: "accept", input: { event: { privacy: "public-source" } }, chosen: fixture.original, @@ -285,6 +398,98 @@ async function appendLegacyJudgment( return result.event; } +async function appendComparedRun( + fixture: Awaited>, + exportClass: "public" | "restricted" | "forbidden", + label: string, + privacy: "public-source" | "private" | "sensitive" = "public-source", +) { + const primary = await fixture.store.getRun(fixture.runId); + if (!primary) throw new Error("Missing primary fixture run"); + const outputContract = outputContractForDeclaration({} as ThoughtAgentDeclaration); + const outputBody: JsonObject = { + summary: `Compared ${label} output`, + tags: [label], + importance: "normal", + confidence: 0.4, + }; + const output = await fixture.store.appendEvent({ + type: "stream.thought.derived.topics", + schemaVersion: 1, + source: `agent:compared-${label}`, + sourceKind: "agent", + externalId: `compared-output-${label}`, + idempotencyKey: `compared-output-${label}`, + occurredAt: "2026-07-15T00:00:01.000Z", + actor: `compared-${label}`, + rootEventId: fixture.sourceEventId, + parentEventId: fixture.sourceEventId, + correlationId: `compared-${label}`, + privacy, + payload: { + outputContract: outputContractIdentityJson(outputContract), + structuredOutput: outputBody, + }, + }); + const release = { + id: `compared-${label}-adapter`, + version: 1, + description: `Compared ${label} adapter`, + releasedAt: "2026-07-25T20:00:00.000Z", + manifestSha256: label === "forbidden" ? "d".repeat(64) : "e".repeat(64), + checkpointSelector: { kind: "env" as const, reference: "THOUGHTSTREAM_TEST_ADAPTER_MODEL" }, + providerProfile: "tinker-default", + baseModel: "Qwen/Qwen3.5-35B-A3B-Base", + dataset: { id: `compared-${label}-dataset`, sha256: "f".repeat(64) }, + evals: [{ id: `compared-${label}-eval`, sha256: "1".repeat(64) }], + capabilities: [`compared-${label}`], + privacyClass: "public" as const, + exportClass, + }; + const modelAdapter = { + ...release, + binding: { checkpointReferenceSha256: label === "forbidden" ? "2".repeat(64) : "3".repeat(64) }, + }; + const runId = `run-compared-${label}-${randomUUID()}`; + await fixture.store.upsertRun({ + id: runId, + executionKey: `execution-compared-${label}`, + triggerEventId: primary.triggerEventId, + agentId: `compared-${label}`, + agentVersion: 1, + status: "completed", + inputEventIds: [primary.triggerEventId], + outputEventIds: [output.event.id], + attempt: 1, + provider: "tinker", + model: release.baseModel, + privacy, + executionAdapterRevision: `compared-execution-${label}`, + modelAdapter, + adapterCatalogDigest: label === "forbidden" ? "6".repeat(64) : "7".repeat(64), + adapterCatalogGeneration: 1, + promptHash: `prompt-${label}`, + contextManifest: { + agentRole: "standard", + outputContract: outputContractIdentityJson(outputContract), + contextStrategy: "single-event", + }, + result: outputBody, + createdAt: "2026-07-15T00:00:00.000Z", + completedAt: "2026-07-15T00:00:02.000Z", + updatedAt: "2026-07-15T00:00:02.000Z", + }); + await fixture.store.appendTrace({ + id: `trace-compared-${label}`, + runId, + sequence: 1, + type: `pi.compared-${label}`, + payload: {}, + createdAt: "2026-07-15T00:00:01.500Z", + }); + return { runId, modelAdapter }; +} + async function completedRunFixture( privacy: "private" | "sensitive" | "public-source", outputPrivacy: "private" | "sensitive" | "public-source" = privacy, @@ -352,7 +557,7 @@ async function completedRunFixture( attempt: 1, provider: "tinker", model: "Qwen/Qwen3.5-4B", - adapterRevision: "checkpoint-fixture", + executionAdapterRevision: "checkpoint-fixture", promptHash: "prompt-hash", contextManifest: { eventIds: [source.event.id], diff --git a/test/letta-agent-sdk-runtime.test.ts b/test/letta-agent-sdk-runtime.test.ts index e8fd049..d13df6e 100644 --- a/test/letta-agent-sdk-runtime.test.ts +++ b/test/letta-agent-sdk-runtime.test.ts @@ -90,7 +90,7 @@ describe("Letta Agent SDK runtime integration", () => { provider: "letta-cloud", model: "agent-default", checkpointRevision: expect.stringContaining(LETTA_AGENT_SDK_ADAPTER_REVISION), - adapterRevision: LETTA_AGENT_SDK_ADAPTER_REVISION, + executionAdapterRevision: LETTA_AGENT_SDK_ADAPTER_REVISION, attempt: 1, }); const traces = await store.listTrace(run!.id); diff --git a/test/model-adapters.test.ts b/test/model-adapters.test.ts new file mode 100644 index 0000000..8e2685f --- /dev/null +++ b/test/model-adapters.test.ts @@ -0,0 +1,500 @@ +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import YAML from "yaml"; +import { afterEach, describe, expect, test } from "vitest"; +import { + loadAdapterCatalog, + modelAdapterIdentityJson, + privateCheckpointFor, +} from "../src/adapters/model-adapters.js"; +import { loadAgentDeclarations } from "../src/agents/declarations.js"; +import { ThoughtAgentRuntime } from "../src/agents/runtime.js"; +import { AgentRunFailure, type AgentRunner } from "../src/agents/types.js"; +import { canonicalJson, sha256, type JsonObject } from "../src/core/json.js"; +import { stableKey } from "../src/core/ids.js"; +import { createDefaultRegistry } from "../src/events/registry.js"; +import type { JazzThoughtStore } from "../src/jazz/store.js"; +import { projectTrainingExamples, recordJudgment } from "../src/training/judgments.js"; +import { testStore } from "./helpers.js"; + +const roots: string[] = []; +const stores: JazzThoughtStore[] = []; +const checkpoint = "tinker://private-fixture-checkpoint"; +const environment = { + THOUGHTSTREAM_TEST_ADAPTER_MODEL: checkpoint, + THOUGHTSTREAM_TINKER_ALLOWED_MODELS: `Qwen/Qwen3.5-4B,${checkpoint}`, +}; + +afterEach(async () => { + await Promise.all(stores.splice(0).map((store) => store.close())); + await Promise.all(roots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); +}); + +describe("immutable model-adapter startup catalog", () => { + test("binds one active release privately and keeps release identity stable across deployment generations", async () => { + const first = await fixtureCatalog({ generation: 1 }); + const second = await fixtureCatalog({ generation: 2 }); + const firstCatalog = await loadAdapterCatalog(first.root, environment); + const secondCatalog = await loadAdapterCatalog(second.root, environment); + const firstBinding = firstCatalog.selectedByDeclaration.get("conceptualizer@1")!; + const secondBinding = secondCatalog.selectedByDeclaration.get("conceptualizer@1")!; + + expect(firstBinding.manifestSha256).toBe(secondBinding.manifestSha256); + expect(firstCatalog.digest).not.toBe(secondCatalog.digest); + expect(firstCatalog.generation).toBe(1); + expect(secondCatalog.generation).toBe(2); + expect(privateCheckpointFor(firstBinding)).toBe(checkpoint); + expect(JSON.stringify(modelAdapterIdentityJson(firstBinding))).not.toContain(checkpoint); + expect(JSON.stringify(firstCatalog)).not.toContain(checkpoint); + expect(Object.isFrozen(firstBinding)).toBe(true); + expect(Object.isFrozen(firstBinding.binding)).toBe(true); + expect(() => { (firstBinding as { description: string }).description = "mutated"; }).toThrow(); + expect(() => (firstCatalog.selectedByDeclaration as Map).set("forged@1", {})).toThrow(); + + const clone = JSON.parse(JSON.stringify(firstBinding)); + expect(() => privateCheckpointFor(clone)).toThrow("unavailable or mismatched"); + }); + + test("validates an expected catalog digest without hashing the digest field into itself", async () => { + const fixture = await fixtureCatalog({ generation: 3 }); + const initial = await loadAdapterCatalog(fixture.root, environment); + await writeDeployment(fixture.root, { + generation: 3, + manifestSha256: fixture.manifestSha256, + expectedCatalogSha256: initial.digest, + }); + await expect(loadAdapterCatalog(fixture.root, environment)).resolves.toMatchObject({ + digest: initial.digest, + generation: 3, + }); + }); + + test.each([ + ["missing release", { releaseId: "missing" }, "missing or has a digest mismatch"], + ["release digest mismatch", { manifestSha256: "0".repeat(64) }, "missing or has a digest mismatch"], + ["unresolved checkpoint", { environment: { THOUGHTSTREAM_TINKER_ALLOWED_MODELS: "Qwen/Qwen3.5-4B" } }, "checkpoint is unresolved"], + ["unallowlisted checkpoint", { environment: { ...environment, THOUGHTSTREAM_TINKER_ALLOWED_MODELS: "Qwen/Qwen3.5-4B" } }, "not allowlisted"], + ["unallowlisted public base", { baseModel: "totally/fabricated-base" }, "not allowlisted"], + ["catalog digest mismatch", { expectedCatalogSha256: "0".repeat(64) }, "catalog digest mismatch"], + ["missing process identity", { expectedProcesses: ["consumer-a"] }, "does not authorize this process identity"], + ["wrong process identity", { expectedProcesses: ["consumer-a"], processIdentity: "consumer-b" }, "does not authorize this process identity"], + ] as Array<[string, FixtureOptions, string]>) ("fails closed for %s", async (_name, options, message) => { + const fixture = await fixtureCatalog(options); + const selectedEnvironment = { + ...(options.environment ?? environment), + ...(options.processIdentity ? { THOUGHTSTREAM_ADAPTER_PROCESS_ID: options.processIdentity } : {}), + }; + await expect(loadAdapterCatalog(fixture.root, selectedEnvironment)).rejects.toThrow(message); + }); + + test.each(["candidate", "retired"] as const)("keeps a %s deployment unbound and rejects an enabled declaration", async (state) => { + const fixture = await fixtureCatalog({ state }); + const catalog = await loadAdapterCatalog(fixture.root, {}); + expect(catalog.selectedByDeclaration.size).toBe(0); + await writeAdapterAgent(fixture.root); + await expect(loadAgentDeclarations(path.join(fixture.root, "agents"), environment, catalog)).rejects.toThrow( + "does not match the loaded deployment catalog", + ); + }); + + test("accepts only an explicitly authorized process identity", async () => { + const fixture = await fixtureCatalog({ expectedProcesses: ["consumer-a", "consumer-b"] }); + await expect(loadAdapterCatalog(fixture.root, { + ...environment, + THOUGHTSTREAM_ADAPTER_PROCESS_ID: "consumer-b", + })).resolves.toMatchObject({ generation: 1 }); + }); + + test("keeps the private checkpoint out of allowlist failures", async () => { + const fixture = await fixtureCatalog(); + const failed = loadAdapterCatalog(fixture.root, { + THOUGHTSTREAM_TEST_ADAPTER_MODEL: checkpoint, + THOUGHTSTREAM_TINKER_ALLOWED_MODELS: "Qwen/Qwen3.5-4B", + }); + await expect(failed).rejects.toThrow("not allowlisted"); + await expect(failed).rejects.not.toThrow(checkpoint); + }); + + test("compiles the production conceptualizer against one exact active adapter selection", async () => { + const fixture = await fixtureCatalog(); + await fs.mkdir(path.join(fixture.root, "agents"), { recursive: true }); + await fs.mkdir(path.join(fixture.root, "prompts"), { recursive: true }); + const declaration = (await fs.readFile(path.join(process.cwd(), "agents", "conceptualizer.example.yaml"), "utf8")) + .replace("enabled: false", "enabled: true") + .replace(" profile: openai-json-default", " profile: tinker-default") + .replace(" model: gpt-4.1-mini", " adapter: { id: fixture-adapter, version: 1 }"); + await fs.writeFile(path.join(fixture.root, "agents", "conceptualizer.yaml"), declaration); + await fs.copyFile( + path.join(process.cwd(), "prompts", "conceptualizer.md"), + path.join(fixture.root, "prompts", "conceptualizer.md"), + ); + + const [compiled] = await loadAgentDeclarations(path.join(fixture.root, "agents"), environment); + expect(compiled).toMatchObject({ + id: "conceptualizer", + enabled: true, + mode: "pi", + model: "Qwen/Qwen3.5-4B", + outputEventType: "stream.thought.derived.concept.graph", + adapterCatalogGeneration: 1, + adapterCatalogDigest: expect.stringMatching(/^[a-f0-9]{64}$/), + modelAdapter: { + id: "fixture-adapter", + version: 1, + manifestSha256: fixture.manifestSha256, + binding: { checkpointReferenceSha256: sha256(checkpoint) }, + }, + }); + expect(Object.isFrozen(compiled)).toBe(true); + expect(Object.isFrozen(compiled!.eventTypes)).toBe(true); + expect(() => { compiled!.description = "mutated"; }).toThrow(); + expect(privateCheckpointFor(compiled!.modelAdapter!)).toBe(checkpoint); + + const store = testStore(fixture.root); + stores.push(store); + const runtime = new ThoughtAgentRuntime(store, []); + await expect(runtime.registerDeclarations([structuredClone(compiled!)])).rejects.toThrow( + "not compiled by the trusted declaration loader", + ); + await expect(runtime.registerDeclarations([compiled!])).resolves.toBeUndefined(); + }); + + test("persists exact adapter and catalog provenance while raising privacy and enforcing export policy", async () => { + const fixture = await fixtureCatalog(); + await writeAdapterAgent(fixture.root); + const declarations = await loadAgentDeclarations(path.join(fixture.root, "agents"), environment); + const declaration = declarations[0]!; + const store = testStore(fixture.root); + stores.push(store); + const source = (await store.appendEvent({ + type: "stream.thought.runtime.notice", + schemaVersion: 1, + source: "fixture:adapter", + sourceKind: "system", + externalId: "adapter-provenance", + idempotencyKey: "adapter-provenance", + occurredAt: "2026-07-25T20:00:00.000Z", + actor: "fixture:adapter", + correlationId: "adapter-provenance", + privacy: "public-source", + payload: { text: "Synthetic adapter provenance fixture" }, + })).event; + const runner: AgentRunner = { + mode: "pi", + run: async () => ({ + summary: "Adapter-backed result", + tags: ["adapter"], + importance: "normal", + confidence: 0.9, + model: { + provider: "tinker", + id: declaration.model!, + revision: checkpoint, + }, + }), + }; + const runtime = new ThoughtAgentRuntime(store, [runner]); + + const [result] = await runtime.consumeBacklog(declarations); + expect(result?.error).toBeUndefined(); + expect(result?.derivedEvent).toMatchObject({ + privacy: "private", + payload: { + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }, + }); + const run = await store.getRun(result!.runId); + expect(run).toMatchObject({ + status: "completed", + privacy: "private", + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + checkpointRevision: `sha256:${sha256(checkpoint)}`, + contextManifest: { + privacy: "private", + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }, + }); + const completed = (await store.listEvents({ types: ["stream.thought.agent.run.completed"] })) + .find((event) => event.payload.runId === run!.id); + expect(completed).toMatchObject({ + privacy: "private", + payload: { + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }, + }); + expect((await store.getProjection(stableKey("effective-output", run!.id)))?.payload.privacy).toBe("private"); + + await recordJudgment(store, { + runId: run!.id, + kind: "accept", + criterion: "adapter-provenance", + criterionVersion: 1, + qualityEligible: true, + externalExportEligible: true, + sensitiveExternalExportAuthorized: true, + }); + expect(await projectTrainingExamples(store, { includeSensitivePrivate: true })).toEqual([]); + const examples = await projectTrainingExamples(store, { + includeSensitivePrivate: true, + includeRestrictedModelAdapters: true, + }); + expect(examples).toHaveLength(1); + expect(examples[0]?.provenance).toMatchObject({ + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }); + expect(examples[0]?.input.contextManifest).toMatchObject({ + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }); + const durable = JSON.stringify({ + source, + agents: await store.listAgents(), + runs: await store.listRuns(), + events: await store.listEvents(), + traces: await store.listTrace(run!.id), + examples, + }); + expect(durable).not.toContain(checkpoint); + }); + + test("keeps adapter-backed output failures repair-eligible with private original provenance", async () => { + const fixture = await fixtureCatalog(); + await writeAdapterAgent(fixture.root); + const declarations = await loadAgentDeclarations(path.join(fixture.root, "agents"), environment); + const declaration = declarations[0]!; + const store = testStore(fixture.root); + stores.push(store); + await store.appendEvent({ + type: "stream.thought.runtime.notice", + schemaVersion: 1, + source: "fixture:adapter", + sourceKind: "system", + externalId: "adapter-repair", + idempotencyKey: "adapter-repair", + occurredAt: "2026-07-25T20:00:00.000Z", + actor: "fixture:adapter", + correlationId: "adapter-repair", + privacy: "public-source", + payload: { text: "Synthetic adapter repair fixture" }, + }); + const runner: AgentRunner = { + mode: "pi", + run: async () => { + throw new AgentRunFailure("Adapter output rejected", { + diagnostic: { + code: "invalid-final-output", + stage: "final-output-validation", + reason: "invalid-json", + checkpointRevision: checkpoint, + outputContract: declaration.outputContract as unknown as JsonObject, + assistantMessages: 1, + textParts: 1, + textChars: 12, + textSha256: sha256("invalid-json"), + thinkingParts: 0, + thinkingChars: 0, + thinkingRedacted: true, + toolCallParts: 0, + otherParts: 0, + stopReason: "stop", + }, + }); + }, + }; + const runtime = new ThoughtAgentRuntime(store, [runner]); + + const [result] = await runtime.consumeBacklog(declarations); + expect(result).toMatchObject({ error: "Adapter output rejected" }); + const run = await store.getRun(result!.runId); + const failed = (await store.listEvents({ types: ["stream.thought.agent.run.failed"] }))[0]; + const request = (await store.listEvents({ types: ["stream.thought.agent.repair.requested"] }))[0]; + expect(run).toMatchObject({ privacy: "private", modelAdapter: declaration.modelAdapter }); + expect(failed).toMatchObject({ privacy: "private" }); + expect(request).toMatchObject({ + privacy: "private", + payload: { + originalRunId: run!.id, + model: { + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }, + }, + }); + expect((await store.getProjection(stableKey("effective-output", run!.id)))?.payload).toMatchObject({ + status: "unresolved", + privacy: "private", + originalModel: { + modelAdapter: declaration.modelAdapter, + adapterCatalogDigest: declaration.adapterCatalogDigest, + adapterCatalogGeneration: 1, + }, + }); + expect(JSON.stringify({ run, failed, request })).not.toContain(checkpoint); + }); + + test("rejects missing, empty, duplicate, and ambiguous catalog structure", async () => { + const missingRoot = await temporaryRoot(); + await expect(loadAdapterCatalog(missingRoot, environment)).rejects.toThrow("adapters/releases"); + + const noDeployment = await temporaryRoot(); + await fs.mkdir(path.join(noDeployment, "adapters", "releases"), { recursive: true }); + await fs.writeFile(path.join(noDeployment, "adapters", "releases", "fixture.yaml"), releaseManifest()); + await expect(loadAdapterCatalog(noDeployment, environment)).rejects.toThrow("adapters/deployment.yaml"); + + const empty = await temporaryRoot(); + await fs.mkdir(path.join(empty, "adapters", "releases"), { recursive: true }); + await fs.writeFile(path.join(empty, "adapters", "deployment.yaml"), deploymentYaml({ manifestSha256: "0".repeat(64) })); + await expect(loadAdapterCatalog(empty, environment)).rejects.toThrow("complete release catalog"); + + const duplicate = await fixtureCatalog(); + await fs.copyFile( + path.join(duplicate.root, "adapters", "releases", "fixture.yaml"), + path.join(duplicate.root, "adapters", "releases", "duplicate.yaml"), + ); + await expect(loadAdapterCatalog(duplicate.root, environment)).rejects.toThrow("Duplicate release"); + + const ambiguous = await fixtureCatalog(); + await writeDeployment(ambiguous.root, { + manifestSha256: ambiguous.manifestSha256, + duplicateSelection: true, + }); + await expect(loadAdapterCatalog(ambiguous.root, environment)).rejects.toThrow("Ambiguous selected deployment"); + + const duplicateCapability = await fixtureCatalog(); + await fs.writeFile( + path.join(duplicateCapability.root, "adapters", "releases", "fixture.yaml"), + releaseManifest().replace("capabilities: [conceptualization]", "capabilities: [conceptualization, conceptualization]"), + ); + await expect(loadAdapterCatalog(duplicateCapability.root, environment)).rejects.toThrow("Duplicate capability"); + + const duplicateEval = await fixtureCatalog(); + await fs.writeFile( + path.join(duplicateEval.root, "adapters", "releases", "fixture.yaml"), + releaseManifest().replace( + `evals: [{ id: fixture-eval, sha256: ${"d".repeat(64)} }]`, + `evals: [{ id: fixture-eval, sha256: ${"d".repeat(64)} }, { id: fixture-eval, sha256: ${"e".repeat(64)} }]`, + ), + ); + await expect(loadAdapterCatalog(duplicateEval.root, environment)).rejects.toThrow("Duplicate eval receipt"); + }); + + test("does not register the removed model-adapter lifecycle event surface", () => { + const registry = createDefaultRegistry(); + for (const type of [ + "stream.thought.model.adapter.registered", + "stream.thought.model.adapter.activated", + "stream.thought.model.adapter.retired", + ]) expect(registry.has(type, 1)).toBe(false); + }); +}); + +interface FixtureOptions { + generation?: number; + state?: "candidate" | "active" | "retired"; + releaseId?: string; + manifestSha256?: string; + expectedCatalogSha256?: string; + expectedProcesses?: string[]; + processIdentity?: string; + environment?: NodeJS.ProcessEnv; + baseModel?: string; +} + +async function fixtureCatalog(options: FixtureOptions = {}) { + const root = await temporaryRoot(); + await fs.mkdir(path.join(root, "adapters", "releases"), { recursive: true }); + const manifest = releaseManifest(options); + const manifestSha256 = sha256(canonicalJson(YAML.parse(manifest))); + await fs.writeFile(path.join(root, "adapters", "releases", "fixture.yaml"), manifest); + await writeDeployment(root, { ...options, manifestSha256: options.manifestSha256 ?? manifestSha256 }); + return { root, manifestSha256 }; +} + +async function writeDeployment(root: string, options: FixtureOptions & { manifestSha256: string; duplicateSelection?: boolean }) { + await fs.mkdir(path.join(root, "adapters"), { recursive: true }); + await fs.writeFile(path.join(root, "adapters", "deployment.yaml"), deploymentYaml(options)); +} + +function releaseManifest(options: FixtureOptions = {}) { + return [ + "schemaVersion: 1", + "id: fixture-adapter", + "version: 1", + "description: Fixture learned adapter", + "releasedAt: 2026-07-25T20:00:00.000Z", + "providerProfile: tinker-default", + `baseModel: ${options.baseModel ?? "Qwen/Qwen3.5-4B"}`, + "checkpoint: { env: THOUGHTSTREAM_TEST_ADAPTER_MODEL }", + `dataset: { id: fixture-dataset, sha256: ${"c".repeat(64)} }`, + `evals: [{ id: fixture-eval, sha256: ${"d".repeat(64)} }]`, + "capabilities: [conceptualization]", + "privacyClass: private", + "exportClass: restricted", + "", + ].join("\n"); +} + +function deploymentYaml(options: FixtureOptions & { manifestSha256: string; duplicateSelection?: boolean }) { + const selection = [ + " - declaration: { id: conceptualizer, version: 1 }", + ` release: { id: ${options.releaseId ?? "fixture-adapter"}, version: 1, manifestSha256: "${options.manifestSha256}" }`, + ` state: ${options.state ?? "active"}`, + ]; + return [ + "schemaVersion: 1", + `generation: ${options.generation ?? 1}`, + ...(options.expectedCatalogSha256 ? [`expectedCatalogSha256: "${options.expectedCatalogSha256}"`] : []), + ...(options.expectedProcesses ? ["expectedProcesses:", ...options.expectedProcesses.map((value) => ` - ${value}`)] : []), + "selections:", + ...selection, + ...(options.duplicateSelection ? selection : []), + "", + ].join("\n"); +} + +async function temporaryRoot() { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-model-adapters-")); + roots.push(root); + return root; +} + +async function writeAdapterAgent(root: string) { + await fs.mkdir(path.join(root, "agents"), { recursive: true }); + await fs.mkdir(path.join(root, "prompts"), { recursive: true }); + await fs.writeFile(path.join(root, "prompts", "adapter.md"), "Observe the synthetic fixture.\n"); + await fs.writeFile(path.join(root, "agents", "adapter.yaml"), [ + "id: conceptualizer", + "version: 1", + "description: Adapter provenance fixture", + "enabled: true", + "subscribe: { types: [stream.thought.runtime.notice], sources: [fixture:adapter], privacy: [public-source] }", + "context: { maxEvents: 1, maxChars: 2000, strategy: single-event }", + "runner:", + " kind: pi", + " profile: tinker-default", + " adapter: { id: fixture-adapter, version: 1 }", + " maxOutputTokens: 200", + " timeoutMs: 60000", + "accounting:", + " leaseMs: 70000", + " reservation: { inputTokens: 1000, outputTokens: 200 }", + " limits: [{ window: hour, maxCalls: 10, maxInputTokens: 10000, maxOutputTokens: 2000 }]", + "prompt: prompts/adapter.md", + "emit: [stream.thought.derived.topics]", + "policy: { tools: [], externalActions: false }", + "", + ].join("\n")); +} diff --git a/test/oauth-auth.test.ts b/test/oauth-auth.test.ts new file mode 100644 index 0000000..81833e9 --- /dev/null +++ b/test/oauth-auth.test.ts @@ -0,0 +1,830 @@ +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { afterEach, describe, expect, test } from "vitest"; +import { JoseKey, NodeOAuthClient, type NodeSavedSession, type NodeSavedState } from "@atproto/oauth-client-node"; +import { + createInspectorOAuthAuth, + flowCookieName, + GenerationSessionStore, + InspectorOAuthAuth, + oauthConfigurationFromEnv, + type BrowserSession, + type OAuthProtocolClient, + type PersistedOAuthSession, + sessionCookieName, +} from "../src/web/oauth-auth.js"; +import { SecureJsonStore } from "../src/web/secure-store.js"; + +const roots: string[] = []; + +afterEach(async () => { + await Promise.all(roots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); +}); + +describe("ATProto OAuth inspector authentication", () => { + test("keeps SDK protocol state distinct from browser-bound application state", async () => { + const fixture = await authFixture(); + const started = await fixture.auth.begin(); + const applicationState = cookieValue(started.setCookie, flowCookieName()); + const protocolState = fixture.protocol.protocolState; + + expect(applicationState).not.toBe(protocolState); + expect(started.redirect.searchParams.get("request_uri")).toBe("urn:ietf:params:oauth:request_uri:fixture"); + expect(fixture.protocol.authorizeCalls).toEqual([{ + handle: "cameron.stream", + applicationState, + scope: "atproto", + signal: undefined, + }]); + + const finished = await fixture.auth.finish( + new URLSearchParams({ state: protocolState, code: "fixture-code" }), + `${flowCookieName()}=${applicationState}`, + ); + expect(finished.redirect).toBe("/inspector/"); + expect(fixture.protocol.callbackStates).toEqual([protocolState]); + const sessionCookie = finished.setCookies.find((value) => value.startsWith(`${sessionCookieName()}=`)); + expect(sessionCookie).toBeDefined(); + expect(sessionCookie).not.toContain("did:plc:allowed"); + expect(sessionCookie).not.toContain("fixture-code"); + + const sessionId = cookieValue(sessionCookie!, sessionCookieName()); + const browser = await fixture.auth.authenticate(`${sessionCookieName()}=${sessionId}`); + expect(browser).toMatchObject({ did: "did:plc:allowed", csrfToken: expect.any(String) }); + expect(fixture.protocol.restoreCalls).toEqual(["did:plc:allowed"]); + + await expect(fixture.auth.finish( + new URLSearchParams({ state: protocolState, code: "replayed-code" }), + `${flowCookieName()}=${applicationState}`, + )).rejects.toThrow("protocol state replayed"); + }); + + test("serializes concurrent callbacks so SDK protocol-state get/delete cannot race", async () => { + const fixture = await authFixture(); + const started = await fixture.auth.begin(); + const applicationState = cookieValue(started.setCookie, flowCookieName()); + let release!: () => void; + fixture.protocol.callbackGate = new Promise((resolve) => { release = resolve; }); + const first = fixture.auth.finish(callbackParams(fixture.protocol), `${flowCookieName()}=${applicationState}`); + const second = fixture.auth.finish(callbackParams(fixture.protocol), `${flowCookieName()}=${applicationState}`); + await new Promise((resolve) => setTimeout(resolve, 5)); + expect(fixture.protocol.callbackCalls).toBe(1); + release(); + const results = await Promise.allSettled([first, second]); + expect(results.filter((result) => result.status === "fulfilled")).toHaveLength(1); + expect(results.filter((result) => result.status === "rejected")).toHaveLength(1); + expect(fixture.protocol.callbackCalls).toBe(2); + }); + + test("watchdog advances past a never-resolving callback and permits a later callback", async () => { + const protocol = new WatchdogProtocol(["never", "success"], true); + const auth = await authWithProtocol(protocol, 20); + const firstStart = await auth.begin(); + const firstAppState = cookieValue(firstStart.setCookie, flowCookieName()); + await expect(auth.finish( + new URLSearchParams({ state: protocol.protocolState(0), code: "first" }), + `${flowCookieName()}=${firstAppState}`, + )).rejects.toThrow("timed out"); + + const secondStart = await auth.begin(); + const secondAppState = cookieValue(secondStart.setCookie, flowCookieName()); + const second = await auth.finish( + new URLSearchParams({ state: protocol.protocolState(1), code: "second" }), + `${flowCookieName()}=${secondAppState}`, + ); + const sessionCookie = second.setCookies.find((value) => value.startsWith(`${sessionCookieName()}=`)); + expect(sessionCookie).toBeDefined(); + expect(protocol.callbackCalls).toBe(2); + expect(protocol.promotedAttempts).toHaveLength(1); + }); + + test("watchdog advances past hung promotion and permits a later generation", async () => { + const protocol = new WatchdogProtocol(["hung-promotion", "success"]); + const auth = await authWithProtocol(protocol, 20); + const firstStart = await auth.begin(); + const firstAppState = cookieValue(firstStart.setCookie, flowCookieName()); + await expect(auth.finish( + new URLSearchParams({ state: protocol.protocolState(0), code: "first" }), + `${flowCookieName()}=${firstAppState}`, + )).rejects.toThrow("timed out"); + + const secondStart = await auth.begin(); + const secondAppState = cookieValue(secondStart.setCookie, flowCookieName()); + const second = await auth.finish( + new URLSearchParams({ state: protocol.protocolState(1), code: "second" }), + `${flowCookieName()}=${secondAppState}`, + ); + expect(second.setCookies.some((value) => value.startsWith(`${sessionCookieName()}=`))).toBe(true); + expect(protocol.currentGeneration).toBe(1); + }); + + test("hung non-timeout cleanup cannot block a subsequent successful callback", async () => { + const protocol = new WatchdogProtocol(["success", "success"], true); + const auth = await authWithProtocol(protocol, 50); + await auth.begin(); + await expect(auth.finish( + new URLSearchParams({ state: protocol.protocolState(0), code: "first" }), + `${flowCookieName()}=${"u".repeat(43)}`, + )).rejects.toThrow("could not be verified"); + + const secondStart = await auth.begin(); + const secondAppState = cookieValue(secondStart.setCookie, flowCookieName()); + const second = await auth.finish( + new URLSearchParams({ state: protocol.protocolState(1), code: "second" }), + `${flowCookieName()}=${secondAppState}`, + ); + expect(second.setCookies.some((value) => value.startsWith(`${sessionCookieName()}=`))).toBe(true); + expect(protocol.currentGeneration).toBe(1); + }); + + test("late callback completion is discarded without revoking newer browser authority", async () => { + const protocol = new WatchdogProtocol(["late", "success"]); + const auth = await authWithProtocol(protocol, 20); + const firstStart = await auth.begin(); + const firstAppState = cookieValue(firstStart.setCookie, flowCookieName()); + await expect(auth.finish( + new URLSearchParams({ state: protocol.protocolState(0), code: "first" }), + `${flowCookieName()}=${firstAppState}`, + )).rejects.toThrow("timed out"); + + const secondStart = await auth.begin(); + const secondAppState = cookieValue(secondStart.setCookie, flowCookieName()); + const second = await auth.finish( + new URLSearchParams({ state: protocol.protocolState(1), code: "second" }), + `${flowCookieName()}=${secondAppState}`, + ); + const sessionCookie = second.setCookies.find((value) => value.startsWith(`${sessionCookieName()}=`))!; + const sessionId = cookieValue(sessionCookie, sessionCookieName()); + expect(await auth.authenticate(`${sessionCookieName()}=${sessionId}`)).toMatchObject({ did: "did:plc:allowed" }); + + protocol.resolveLate(); + await waitFor(() => protocol.lateDiscarded); + expect(await auth.authenticate(`${sessionCookieName()}=${sessionId}`)).toMatchObject({ did: "did:plc:allowed" }); + expect(protocol.persistentDid).toBe("did:plc:allowed"); + expect(protocol.currentGeneration).toBe(1); + expect(protocol.remoteRevocations).toBe(0); + }); + + test("rejects missing, duplicate, expired, or mismatched browser application state", async () => { + const missing = await authFixture(); + const missingStart = await missing.auth.begin(); + await expect(missing.auth.finish( + callbackParams(missing.protocol), + undefined, + )).rejects.toThrow("could not be verified"); + expect(missing.protocol.callbackCalls).toBe(0); + + const duplicate = await authFixture(); + const duplicateStart = await duplicate.auth.begin(); + const duplicateAppState = cookieValue(duplicateStart.setCookie, flowCookieName()); + await expect(duplicate.auth.finish( + callbackParams(duplicate.protocol), + `${flowCookieName()}=${duplicateAppState}; ${flowCookieName()}=${duplicateAppState}`, + )).rejects.toThrow("could not be verified"); + expect(duplicate.protocol.callbackCalls).toBe(0); + + const duplicateProtocol = await authFixture(); + const duplicateProtocolStart = await duplicateProtocol.auth.begin(); + const duplicateProtocolAppState = cookieValue(duplicateProtocolStart.setCookie, flowCookieName()); + const duplicatedParams = callbackParams(duplicateProtocol.protocol); + duplicatedParams.append("state", duplicateProtocol.protocol.protocolState); + await expect(duplicateProtocol.auth.finish( + duplicatedParams, + `${flowCookieName()}=${duplicateProtocolAppState}`, + )).rejects.toThrow("could not be verified"); + expect(duplicateProtocol.protocol.callbackCalls).toBe(0); + + const mismatch = await authFixture(); + await mismatch.auth.begin(); + const unrelatedState = "u".repeat(43); + await expect(mismatch.auth.finish( + callbackParams(mismatch.protocol), + `${flowCookieName()}=${unrelatedState}`, + )).rejects.toThrow("could not be verified"); + expect(mismatch.protocol.callbackCalls).toBe(1); + expect(mismatch.protocol.revoked).toEqual([]); + expect(mismatch.protocol.deletedSessions).toEqual([]); + expect(mismatch.protocol.discardedAttempts).toHaveLength(1); + + let now = 1_000_000; + const expired = await authFixture({ now: () => now }); + const expiredStart = await expired.auth.begin(); + const expiredAppState = cookieValue(expiredStart.setCookie, flowCookieName()); + now += 15 * 60_000 + 1; + await expect(expired.auth.finish( + callbackParams(expired.protocol), + `${flowCookieName()}=${expiredAppState}`, + )).rejects.toThrow("could not be verified"); + expect(expired.protocol.revoked).toEqual([]); + expect(expired.protocol.deletedSessions).toEqual([]); + expect(expired.protocol.discardedAttempts).toHaveLength(1); + }); + + test("deletes staged SDK authority locally when flow-store settlement fails", async () => { + const fixture = await authFixture(); + const started = await fixture.auth.begin(); + const applicationState = cookieValue(started.setCookie, flowCookieName()); + fixture.flowStates.take = async () => { throw new Error("synthetic flow-store disk failure"); }; + + await expect(fixture.auth.finish( + callbackParams(fixture.protocol), + `${flowCookieName()}=${applicationState}`, + )).rejects.toThrow("synthetic flow-store disk failure"); + expect(fixture.protocol.revoked).toEqual([]); + expect(fixture.protocol.deletedSessions).toEqual([]); + expect(fixture.protocol.discardedAttempts).toHaveLength(1); + expect(fixture.protocol.promotedAttempts).toHaveLength(0); + }); + + test("rejects duplicate session cookies before SDK restore", async () => { + const fixture = await authFixture(); + const { sessionId } = await completeLogin(fixture); + const cookie = `${sessionCookieName()}=${sessionId}`; + expect(await fixture.auth.authenticate(`${cookie}; ${cookie}`)).toBeUndefined(); + expect(fixture.protocol.restoreCalls).toEqual([]); + }); + + test("expires browser sessions and deletes only their matching local generation", async () => { + let now = 2_000_000; + const fixture = await authFixture({ now: () => now, sessionTtlMs: 60_000 }); + const { sessionId } = await completeLogin(fixture); + now += 60_001; + + expect(await fixture.auth.authenticate(`${sessionCookieName()}=${sessionId}`)).toBeUndefined(); + expect(fixture.protocol.revoked).toEqual([]); + expect(fixture.protocol.deletedSessions).toEqual([]); + expect(fixture.protocol.discardedPromotions).toEqual([1]); + expect(fixture.protocol.restoreCalls).toEqual([]); + }); + + test("deletes the matching local generation after restore failure", async () => { + const fixture = await authFixture(); + const { sessionId } = await completeLogin(fixture); + fixture.protocol.restoreFailure = true; + + expect(await fixture.auth.authenticate(`${sessionCookieName()}=${sessionId}`)).toBeUndefined(); + expect(fixture.protocol.restoreCalls).toEqual(["did:plc:allowed"]); + expect(fixture.protocol.revoked).toEqual([]); + expect(fixture.protocol.deletedSessions).toEqual([]); + expect(fixture.protocol.discardedPromotions).toEqual([1]); + expect(await fixture.auth.authenticate(`${sessionCookieName()}=${sessionId}`)).toBeUndefined(); + }); + + test("logout clears browser authority and only its exact local generation", async () => { + const fixture = await authFixture(); + const { sessionId } = await completeLogin(fixture); + const browser = await fixture.auth.authenticate(`${sessionCookieName()}=${sessionId}`); + + const cleared = await fixture.auth.logout(`${sessionCookieName()}=${sessionId}`, browser!.csrfToken); + expect(cleared).toContain("Max-Age=0"); + expect(fixture.protocol.revoked).toEqual([]); + expect(fixture.protocol.deletedSessions).toEqual([]); + expect(fixture.protocol.discardedPromotions).toEqual([1]); + expect(await fixture.auth.authenticate(`${sessionCookieName()}=${sessionId}`)).toBeUndefined(); + }); + + test("drops a non-allowlisted staged DID without remote revocation", async () => { + const fixture = await authFixture({ callbackDid: "did:plc:other" }); + const started = await fixture.auth.begin(); + const applicationState = cookieValue(started.setCookie, flowCookieName()); + + await expect(fixture.auth.finish( + callbackParams(fixture.protocol), + `${flowCookieName()}=${applicationState}`, + )).rejects.toThrow("not authorized"); + expect(fixture.protocol.revoked).toEqual([]); + expect(fixture.protocol.deletedSessions).toEqual([]); + expect(fixture.protocol.discardedAttempts).toHaveLength(1); + }); + + test("scopes deferred restore refresh writes to their initiating generation", async () => { + const { store } = await generationStoreFixture(); + const did = "did:plc:allowed"; + const generationOne = await stageAndPromote(store, "attempt-one", did, savedSession("one")); + let release!: () => void; + let entered!: () => void; + const gate = new Promise((resolve) => { release = resolve; }); + const started = new Promise((resolve) => { entered = resolve; }); + const oldRestore = store.restoreExactGeneration(did, generationOne, async () => { + expect(await store.get(did)).toEqual(savedSession("one")); + entered(); + await gate; + await store.set(did, savedSession("one-refreshed")); + }); + await started; + const generationTwo = await stageAndPromote(store, "attempt-two", did, savedSession("two")); + release(); + await expect(oldRestore).rejects.toThrow("changed during restore"); + + expect(await store.matchesPersistedGeneration(did, generationOne)).toBe(false); + expect(await store.matchesPersistedGeneration(did, generationTwo)).toBe(true); + expect(await store.get(did)).toEqual(savedSession("two")); + }); + + test("scopes deferred restore failure deletion to its initiating generation", async () => { + const { store } = await generationStoreFixture(); + const did = "did:plc:allowed"; + const generationOne = await stageAndPromote(store, "attempt-one", did, savedSession("one")); + let release!: () => void; + let entered!: () => void; + const gate = new Promise((resolve) => { release = resolve; }); + const started = new Promise((resolve) => { entered = resolve; }); + const oldRestore = store.runRestore(did, generationOne, async () => { + expect(await store.get(did)).toEqual(savedSession("one")); + entered(); + await gate; + await store.del(did); + throw new Error("synthetic generation-one refresh failure"); + }); + await started; + const generationTwo = await stageAndPromote(store, "attempt-two", did, savedSession("two")); + release(); + await expect(oldRestore).rejects.toThrow("generation-one refresh failure"); + + expect(await store.matchesPersistedGeneration(did, generationOne)).toBe(false); + expect(await store.matchesPersistedGeneration(did, generationTwo)).toBe(true); + expect(await store.get(did)).toEqual(savedSession("two")); + }); + + test("bounds retained callback quarantines and recovers after settlement or process recycle", async () => { + const { persistent, store } = await generationStoreFixture(2); + let releaseOne!: () => void; + let releaseTwo!: () => void; + const first = store.run("retained-one", () => new Promise((resolve) => { releaseOne = resolve; })); + const second = store.run("retained-two", () => new Promise((resolve) => { releaseTwo = resolve; })); + store.expire("retained-one"); + store.expire("retained-two"); + expect(store.quarantineStatus()).toEqual({ retained: 2, capacity: 2, recycleRequired: true }); + expect(() => store.run("refused", async () => undefined)).toThrow("recycle the proxy process"); + + releaseOne(); + await first; + await store.drop("retained-one"); + expect(store.quarantineStatus()).toEqual({ retained: 1, capacity: 2, recycleRequired: false }); + await store.run("accepted-after-settlement", async () => undefined); + await store.drop("accepted-after-settlement"); + releaseTwo(); + await second; + await store.drop("retained-two"); + + const restarted = new GenerationSessionStore(persistent, 2); + expect(restarted.quarantineStatus()).toEqual({ retained: 0, capacity: 2, recycleRequired: false }); + await restarted.run("accepted-after-recycle", async () => undefined); + await restarted.drop("accepted-after-recycle"); + }); + + test("fails closed when OAUTH_STORE_DIR escapes the systemd writable path", () => { + const home = "/tmp/thoughtstream-fixture-home"; + const expected = `${home}/.local/share/thoughtstream-inspector-auth`; + const base = { + HOME: home, + OAUTH_ENABLED: "1", + OAUTH_PUBLIC_ORIGIN: "https://thought.stream", + OAUTH_ALLOWED_DID: "did:plc:allowed", + OAUTH_EXPECTED_HANDLE: "cameron.stream", + OAUTH_STORE_KEY_B64: Buffer.alloc(32, 16).toString("base64"), + OAUTH_PRIVATE_JWK_B64: Buffer.from(JSON.stringify({ kty: "EC", d: "private-fixture" })).toString("base64"), + }; + expect(oauthConfigurationFromEnv({ ...base, OAUTH_STORE_DIR: expected }).storeDirectory).toBe(expected); + expect(() => oauthConfigurationFromEnv({ ...base, OAUTH_STORE_DIR: `${home}/custom` })) + .toThrow("systemd sandbox permits only"); + }); + + test("the installed official SDK generates protocol state separately from appState", async () => { + const key = await JoseKey.generate(["ES256"], "fixture-key"); + const states = new Map(); + const sessions = new Map(); + const stateStore = mapStore(states); + const sessionStore = mapStore(sessions); + const client = new NodeOAuthClient({ + clientMetadata: { + client_id: "https://thought.stream/oauth/client-metadata.json", + client_name: "ThoughtStream inspector", + client_uri: "https://thought.stream/", + redirect_uris: ["https://thought.stream/oauth/callback"], + grant_types: ["authorization_code", "refresh_token"], + scope: "atproto", + response_types: ["code"], + application_type: "web", + token_endpoint_auth_method: "private_key_jwt", + token_endpoint_auth_signing_alg: "ES256", + dpop_bound_access_tokens: true, + jwks_uri: "https://thought.stream/oauth/jwks.json", + }, + keyset: [key], + stateStore, + sessionStore, + requestLock: async (_name, operation) => operation(), + }); + const authorizationMetadata = { + issuer: "https://auth.example", + authorization_endpoint: "https://auth.example/authorize", + token_endpoint: "https://auth.example/token", + pushed_authorization_request_endpoint: "https://auth.example/par", + require_pushed_authorization_requests: true, + token_endpoint_auth_methods_supported: ["private_key_jwt"], + token_endpoint_auth_signing_alg_values_supported: ["ES256"], + dpop_signing_alg_values_supported: ["ES256"], + scopes_supported: ["atproto"], + response_types_supported: ["code"], + grant_types_supported: ["authorization_code", "refresh_token"], + code_challenge_methods_supported: ["S256"], + authorization_response_iss_parameter_supported: true, + }; + let pushedState: string | undefined; + (client.oauthResolver as unknown as { resolve: () => Promise }).resolve = async () => ({ + identityInfo: { did: "did:plc:allowed", handle: "cameron.stream" }, + metadata: authorizationMetadata, + }); + (client.serverFactory as unknown as { fromMetadata: () => Promise }).fromMetadata = async () => ({ + request: async (_endpoint: string, payload: { state?: string }) => { + pushedState = payload.state; + return { request_uri: "urn:ietf:params:oauth:request_uri:fixture", expires_in: 90 }; + }, + }); + const applicationState = "a".repeat(43); + const redirect = await client.authorize("did:plc:allowed", { state: applicationState, scope: "atproto" }); + const [protocolState, stored] = [...states.entries()][0]!; + + expect(protocolState).not.toBe(applicationState); + expect(stored.appState).toBe(applicationState); + expect(pushedState).toBe(protocolState); + expect(redirect.searchParams.get("request_uri")).toBe("urn:ietf:params:oauth:request_uri:fixture"); + }); + + test("official client metadata is exact, JWKS is public-only, and the store is singleton", async () => { + const root = await temporaryRoot("thoughtstream-oauth-client-"); + const key = await JoseKey.generate(["ES256"], "fixture-key"); + const configuration = { + publicOrigin: "https://thought.stream", + allowedDid: "did:plc:allowed", + expectedHandle: "cameron.stream", + storeDirectory: root, + storeKey: Buffer.alloc(32, 4), + privateJwk: { ...key.privateJwk! }, + }; + const auth = await createInspectorOAuthAuth(configuration); + + expect(auth.clientMetadata).toMatchObject({ + client_id: "https://thought.stream/oauth/client-metadata.json", + client_uri: "https://thought.stream/", + redirect_uris: ["https://thought.stream/oauth/callback"], + grant_types: ["authorization_code", "refresh_token"], + scope: "atproto", + response_types: ["code"], + token_endpoint_auth_method: "private_key_jwt", + token_endpoint_auth_signing_alg: "ES256", + dpop_bound_access_tokens: true, + jwks_uri: "https://thought.stream/oauth/jwks.json", + }); + expect(JSON.stringify(auth.jwks)).not.toContain('"d"'); + expect(JSON.stringify(auth.clientMetadata)).not.toContain('"d"'); + await expect(createInspectorOAuthAuth(configuration)).rejects.toThrow("already owned by process"); + await auth.close(); + const reacquired = await createInspectorOAuthAuth(configuration); + await reacquired.close(); + }); +}); + +interface FixtureOptions { + callbackDid?: string; + now?: () => number; + sessionTtlMs?: number; + callbackTimeoutMs?: number; +} + +async function authFixture(options: FixtureOptions = {}): Promise<{ + auth: InspectorOAuthAuth; + protocol: FakeProtocol; + flowStates: SecureJsonStore<{ createdAt: number }>; +}> { + const root = await temporaryRoot("thoughtstream-oauth-auth-"); + const key = Buffer.alloc(32, 3); + const now = options.now ?? Date.now; + const browserSessions = new SecureJsonStore({ + directory: root, + name: "browser", + key, + maxEntries: 8, + maxSerializedBytes: 32 * 1024, + }); + const flowStates = new SecureJsonStore<{ createdAt: number }>({ + directory: root, + name: "flow", + key, + maxEntries: 64, + maxSerializedBytes: 32 * 1024, + ttlMs: 15 * 60_000, + now, + }); + const protocol = new FakeProtocol(options.callbackDid ?? "did:plc:allowed"); + return { + auth: new InspectorOAuthAuth({ + protocol, + allowedDid: "did:plc:allowed", + expectedHandle: "cameron.stream", + browserSessions, + flowStates, + now, + ...(options.sessionTtlMs ? { sessionTtlMs: options.sessionTtlMs } : {}), + ...(options.callbackTimeoutMs ? { callbackTimeoutMs: options.callbackTimeoutMs } : {}), + }), + protocol, + flowStates, + }; +} + +class FakeProtocol implements OAuthProtocolClient { + readonly clientMetadata = { client_id: "https://thought.stream/oauth/client-metadata.json" }; + readonly jwks = { keys: [] }; + readonly protocolState = "sdk-protocol-state-1234567890abcdef"; + readonly authorizeCalls: Array<{ handle: string; applicationState: string; scope: string; signal: AbortSignal | undefined }> = []; + readonly callbackStates: string[] = []; + readonly restoreCalls: string[] = []; + readonly revoked: string[] = []; + readonly deletedSessions: string[] = []; + readonly discardedAttempts: string[] = []; + readonly expiredAttempts: string[] = []; + readonly discardedPromotions: number[] = []; + readonly promotedAttempts: string[] = []; + callbackCalls = 0; + restoreFailure = false; + revokeFailure = false; + callbackGate?: Promise; + private applicationState?: string; + private callbackUsed = false; + private readonly attemptDids = new Map(); + private currentGeneration: number | undefined; + private nextGeneration = 0; + + constructor(private readonly callbackDid: string) {} + + async authorize(handle: string, options: { state: string; scope: string; signal?: AbortSignal }): Promise { + this.applicationState = options.state; + this.authorizeCalls.push({ handle, applicationState: options.state, scope: options.scope, signal: options.signal }); + return new URL("https://pds.example/authorize?request_uri=urn:ietf:params:oauth:request_uri:fixture"); + } + + async callback(params: URLSearchParams, attemptId: string): Promise<{ session: { did: string }; state: string | null }> { + this.callbackCalls += 1; + const protocolState = params.get("state"); + this.callbackStates.push(protocolState ?? ""); + if (this.callbackUsed) throw new Error("protocol state replayed"); + if (protocolState !== this.protocolState) throw new Error("unknown protocol state"); + await this.callbackGate; + this.callbackUsed = true; + this.attemptDids.set(attemptId, this.callbackDid); + return { session: { did: this.callbackDid }, state: this.applicationState ?? null }; + } + + expireCallbackAttempt(attemptId: string): void { + this.expiredAttempts.push(attemptId); + } + + async dropCallbackAttempt(attemptId: string): Promise { + this.discardedAttempts.push(attemptId); + this.attemptDids.delete(attemptId); + } + + async promoteCallbackSession(attemptId: string, did: string, guard: () => boolean): Promise { + if (!guard() || this.attemptDids.get(attemptId) !== did) throw new Error("missing staged session"); + this.attemptDids.delete(attemptId); + this.promotedAttempts.push(attemptId); + this.currentGeneration = ++this.nextGeneration; + return this.currentGeneration; + } + + isCurrentGeneration(_did: string, generation: number): boolean { + return this.currentGeneration === generation; + } + + async discardPromotedSession(_did: string, generation: number): Promise { + this.discardedPromotions.push(generation); + if (this.currentGeneration === generation) this.currentGeneration = undefined; + } + + async restore(did: string, generation: number): Promise<{ did: string }> { + this.restoreCalls.push(did); + if (this.restoreFailure || this.currentGeneration !== generation) throw new Error("synthetic restore failure"); + return { did }; + } + + async revoke(did: string): Promise { + this.revoked.push(did); + if (this.revokeFailure) throw new Error("synthetic revoke failure"); + } + + async deleteSession(did: string): Promise { + this.deletedSessions.push(did); + this.currentGeneration = undefined; + } +} + +async function generationStoreFixture(maxRetainedAttempts = 8): Promise<{ + persistent: SecureJsonStore; + store: GenerationSessionStore; +}> { + const root = await temporaryRoot("thoughtstream-generation-store-"); + const persistent = new SecureJsonStore({ + directory: root, + name: "generation-sessions", + key: Buffer.alloc(32, 20), + maxEntries: 8, + maxSerializedBytes: 64 * 1024, + }); + return { persistent, store: new GenerationSessionStore(persistent, maxRetainedAttempts) }; +} + +async function stageAndPromote( + store: GenerationSessionStore, + attemptId: string, + did: string, + session: NodeSavedSession, +): Promise { + await store.run(attemptId, () => store.set(did, session)); + return store.promote(attemptId, did, () => true); +} + +function savedSession(label: string): NodeSavedSession { + return { fixture: label } as unknown as NodeSavedSession; +} + +async function authWithProtocol(protocol: OAuthProtocolClient, callbackTimeoutMs: number): Promise { + const root = await temporaryRoot("thoughtstream-oauth-watchdog-"); + const key = Buffer.alloc(32, 14); + return new InspectorOAuthAuth({ + protocol, + allowedDid: "did:plc:allowed", + expectedHandle: "cameron.stream", + browserSessions: new SecureJsonStore({ + directory: root, + name: "watchdog-browser", + key, + maxEntries: 8, + maxSerializedBytes: 32 * 1024, + }), + flowStates: new SecureJsonStore<{ createdAt: number }>({ + directory: root, + name: "watchdog-flow", + key, + maxEntries: 64, + maxSerializedBytes: 32 * 1024, + ttlMs: 15 * 60_000, + }), + callbackTimeoutMs, + }); +} + +class WatchdogProtocol implements OAuthProtocolClient { + readonly clientMetadata = { client_id: "https://thought.stream/oauth/client-metadata.json" }; + readonly jwks = { keys: [] }; + readonly promotedAttempts: string[] = []; + readonly discardedAttempts: string[] = []; + callbackCalls = 0; + lateDiscarded = false; + persistentDid: string | undefined; + currentGeneration: number | undefined; + remoteRevocations = 0; + private nextGeneration = 0; + private readonly expiredAttempts = new Set(); + private readonly applicationStates: string[] = []; + private readonly stagedAttempts = new Map(); + private readonly attemptIndexes = new Map(); + private callbackIndex = 0; + private lateResolve!: () => void; + private readonly latePromise = new Promise((resolve) => { this.lateResolve = resolve; }); + + constructor( + private readonly behaviors: Array<"never" | "late" | "success" | "hung-promotion">, + private readonly hangFirstDrop = false, + ) {} + + protocolState(index: number): string { + return `watchdog-protocol-state-0000000000000000-${index}`; + } + + resolveLate(): void { + this.lateResolve(); + } + + async authorize(_handle: string, options: { state: string; scope: string; signal?: AbortSignal }): Promise { + this.applicationStates.push(options.state); + return new URL("https://pds.example/authorize?request_uri=urn:ietf:params:oauth:request_uri:fixture"); + } + + async callback(params: URLSearchParams, attemptId: string): Promise<{ session: { did: string }; state: string | null }> { + const index = this.callbackIndex++; + this.attemptIndexes.set(attemptId, index); + this.callbackCalls += 1; + if (params.get("state") !== this.protocolState(index)) throw new Error("unexpected protocol state"); + const behavior = this.behaviors[index]; + if (behavior === "never") return new Promise(() => undefined); + if (behavior === "late") await this.latePromise; + this.stagedAttempts.set(attemptId, { did: "did:plc:allowed", index }); + return { session: { did: "did:plc:allowed" }, state: this.applicationStates[index] ?? null }; + } + + expireCallbackAttempt(attemptId: string): void { + this.expiredAttempts.add(attemptId); + } + + async dropCallbackAttempt(attemptId: string): Promise { + this.discardedAttempts.push(attemptId); + const index = this.attemptIndexes.get(attemptId); + if (index === 0 && this.hangFirstDrop) return new Promise(() => undefined); + const staged = this.stagedAttempts.get(attemptId); + this.stagedAttempts.delete(attemptId); + if (staged?.index === 0) this.lateDiscarded = true; + } + + async promoteCallbackSession(attemptId: string, did: string, guard: () => boolean): Promise { + const index = this.attemptIndexes.get(attemptId); + if (index !== undefined && this.behaviors[index] === "hung-promotion") return new Promise(() => undefined); + if (!guard() || this.expiredAttempts.has(attemptId) || this.stagedAttempts.get(attemptId)?.did !== did) { + throw new Error("missing staged session"); + } + this.stagedAttempts.delete(attemptId); + this.promotedAttempts.push(attemptId); + this.persistentDid = did; + this.currentGeneration = ++this.nextGeneration; + return this.currentGeneration; + } + + isCurrentGeneration(did: string, generation: number): boolean { + return this.persistentDid === did && this.currentGeneration === generation; + } + + async discardPromotedSession(did: string, generation: number): Promise { + if (this.persistentDid === did && this.currentGeneration === generation) { + this.persistentDid = undefined; + this.currentGeneration = undefined; + } + } + + async restore(did: string, generation: number): Promise<{ did: string }> { + if (this.persistentDid !== did || this.currentGeneration !== generation) throw new Error("missing persistent session"); + return { did }; + } + + async revoke(did: string): Promise { + this.remoteRevocations += 1; + if (this.persistentDid === did) this.persistentDid = undefined; + } + + async deleteSession(did: string): Promise { + if (this.persistentDid === did) { + this.persistentDid = undefined; + this.currentGeneration = undefined; + } + } +} + +async function waitFor(predicate: () => boolean): Promise { + for (let attempt = 0; attempt < 100; attempt += 1) { + if (predicate()) return; + await new Promise((resolve) => setTimeout(resolve, 5)); + } + throw new Error("Timed out waiting for condition"); +} + +function mapStore(map: Map): { + get(key: string): Promise; + set(key: string, value: T): Promise; + del(key: string): Promise; +} { + return { + get: async (key) => map.get(key), + set: async (key, value) => { map.set(key, value); }, + del: async (key) => { map.delete(key); }, + }; +} + +async function completeLogin(fixture: { auth: InspectorOAuthAuth; protocol: FakeProtocol }): Promise<{ sessionId: string }> { + const started = await fixture.auth.begin(); + const applicationState = cookieValue(started.setCookie, flowCookieName()); + const finished = await fixture.auth.finish( + callbackParams(fixture.protocol), + `${flowCookieName()}=${applicationState}`, + ); + const sessionCookie = finished.setCookies.find((value) => value.startsWith(`${sessionCookieName()}=`)); + if (!sessionCookie) throw new Error("Missing session cookie"); + return { sessionId: cookieValue(sessionCookie, sessionCookieName()) }; +} + +function callbackParams(protocol: FakeProtocol): URLSearchParams { + return new URLSearchParams({ state: protocol.protocolState, code: "fixture-code" }); +} + +function cookieValue(serialized: string, name: string): string { + const first = serialized.split(";", 1)[0]!; + const [cookieName, value] = first.split("=", 2); + if (cookieName !== name || !value) throw new Error(`Missing ${name} cookie`); + return value; +} + +async function temporaryRoot(prefix: string): Promise { + const root = await fs.mkdtemp(path.join(os.tmpdir(), prefix)); + roots.push(root); + return root; +} diff --git a/test/pi-runner.test.ts b/test/pi-runner.test.ts index 5f41fe6..9a15edd 100644 --- a/test/pi-runner.test.ts +++ b/test/pi-runner.test.ts @@ -1,17 +1,26 @@ import http from "node:http"; import path from "node:path"; +import os from "node:os"; +import fs from "node:fs/promises"; +import { createHash } from "node:crypto"; +import YAML from "yaml"; +import { canonicalJson } from "../src/core/json.js"; import { afterEach, describe, expect, test } from "vitest"; import { buildContextPacket } from "../src/agents/context.js"; +import { loadAgentDeclarations } from "../src/agents/declarations.js"; import { PiAgentRunner } from "../src/agents/pi.js"; import { staticProviderProfileResolver, type ProviderProfile } from "../src/agents/provider-profiles.js"; import type { ThoughtAgentDeclaration } from "../src/agents/types.js"; import type { ThoughtEvent } from "../src/events/types.js"; const servers: http.Server[] = []; +const temporaryRoots: string[] = []; const workerBundlePath = path.resolve("dist/sandbox/worker.cjs"); afterEach(async () => { delete process.env.THOUGHTSTREAM_TEST_API_KEY; + delete process.env.THOUGHTSTREAM_TINKER_ALLOWED_MODELS; + await Promise.all(temporaryRoots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); await Promise.all(servers.splice(0).map((server) => new Promise((resolve) => { server.closeAllConnections(); server.close(() => resolve()); @@ -68,7 +77,9 @@ describe("PiAgentRunner", () => { "provider.response", "pi.agent_start", "pi.agent_end", + "pi.message_updates_coalesced", ])); + expect(traces.map((trace) => trace.kind)).not.toContain("pi.message_update"); const serializedTraces = JSON.stringify(traces); expect(serializedTraces).not.toContain("fixture-secret"); expect(serializedTraces).not.toContain("fixture.md"); @@ -76,6 +87,130 @@ describe("PiAgentRunner", () => { expect(serializedTraces).toContain("redacted"); }, 15_000); + test("uses a process-local learned checkpoint while traces and output retain only public adapter identity", async () => { + const privateCheckpoint = "tinker://private-adapter-checkpoint-fixture"; + let receivedBody = ""; + const server = await startServer((request, response) => { + request.on("data", (part) => { receivedBody += part; }); + request.on("end", () => { + response.setHeader("x-tinker-checkpoint", privateCheckpoint); + respondWithOutput(response, validOutput("Adapted output")); + }); + }); + process.env.THOUGHTSTREAM_TEST_API_KEY = "fixture-secret"; + const declaration = await compiledAdapterDeclaration(privateCheckpoint); + const traces: Array<{ kind: string; data: unknown }> = []; + + const output = await fixtureRunner(server, { + allowedModels: [privateCheckpoint], + profileId: "tinker-default", + provider: "tinker", + }).run( + fixtureRunInput(declaration), + async (trace) => { traces.push(trace); }, + ); + + expect(JSON.parse(receivedBody)).toMatchObject({ model: privateCheckpoint }); + expect(output.model).toEqual({ + provider: "tinker", + id: "Qwen/Qwen3.5-4B", + revision: `sha256:${createHash("sha256").update(privateCheckpoint).digest("hex")}`, + }); + expect(JSON.stringify(traces)).not.toContain(privateCheckpoint); + expect(JSON.stringify(traces)).toContain("Qwen/Qwen3.5-4B"); + expect(JSON.stringify(output)).not.toContain(privateCheckpoint); + + receivedBody = ""; + await expect(fixtureRunner(server, { + allowedModels: [privateCheckpoint], + profileId: "tinker-default", + provider: "tinker", + }).run( + fixtureRunInput(structuredClone(declaration)), + async () => undefined, + )).rejects.toThrow("not compiled by the trusted declaration loader"); + expect(receivedBody).toBe(""); + }, 15_000); + + test("never persists a raw learned-checkpoint revision on adapter failure", async () => { + const privateCheckpoint = "tinker://private-adapter-failure-revision"; + const server = await startServer((_request, response) => { + response.setHeader("x-tinker-checkpoint", privateCheckpoint); + respondWithText(response, "not-json"); + }); + process.env.THOUGHTSTREAM_TEST_API_KEY = "fixture-secret"; + const declaration = await compiledAdapterDeclaration(privateCheckpoint); + + try { + await fixtureRunner(server, { + allowedModels: [privateCheckpoint], + profileId: "tinker-default", + provider: "tinker", + }).run(fixtureRunInput(declaration), async () => undefined); + throw new Error("Expected invalid adapter output"); + } catch (error) { + expect(error).toMatchObject({ + diagnostic: expect.objectContaining({ + checkpointRevision: `sha256:${createHash("sha256").update(privateCheckpoint).digest("hex")}`, + }), + }); + expect(JSON.stringify(error)).not.toContain(privateCheckpoint); + expect(String(error)).not.toContain(privateCheckpoint); + } + }, 15_000); + + test("does not begin adapter-backed provider egress until read-only prefetch is exhausted", async () => { + const privateCheckpoint = "tinker://prefetch-boundary-fixture"; + const server = await startServer((_request, response) => respondWithOutput(response, validOutput("After prefetch"))); + process.env.THOUGHTSTREAM_TEST_API_KEY = "fixture-secret"; + const declaration = await compiledAdapterDeclaration(privateCheckpoint, { atprotoPrefetch: true }); + let prefetchStarted!: () => void; + let releasePrefetch!: () => void; + const started = new Promise((resolve) => { prefetchStarted = resolve; }); + const released = new Promise((resolve) => { releasePrefetch = resolve; }); + let upstreamRequests = 0; + const base = fixtureRunInput(declaration); + const input = { + ...base, + event: { + ...base.event, + id: "evt_adapter_prefetch", + type: "stream.thought.source.atproto.commit", + source: "jetstream:test", + sourceKind: "jetstream" as const, + externalId: "did:plc:test/app.bsky.feed.post/one", + rootEventId: "evt_adapter_prefetch", + payload: { + atUri: "at://did:plc:test/app.bsky.feed.post/one", + cid: "bafyfixture", + collection: "app.bsky.feed.post", + rkey: "one", + operation: "create", + }, + }, + }; + const run = fixtureRunner(server, { + allowedModels: [privateCheckpoint], + profileId: "tinker-default", + provider: "tinker", + providerFetchImpl: async (...args) => { + upstreamRequests += 1; + return fetch(...args); + }, + fetchImpl: async () => { + prefetchStarted(); + await released; + return new Response("missing", { status: 404 }); + }, + }).run(input, async () => undefined); + + await started; + expect(upstreamRequests).toBe(0); + releasePrefetch(); + await expect(run).resolves.toMatchObject({ summary: "After prefetch" }); + expect(upstreamRequests).toBe(1); + }, 15_000); + test("uses the broker-authorized model identity instead of a provider-returned field", async () => { const untrustedModelField = "provider-secret-fragment"; const server = await startServer((_request, response) => { @@ -396,31 +531,121 @@ describe("PiAgentRunner", () => { function fixtureRunner( server: http.Server, - options: { fetchImpl?: typeof fetch; workerBundlePath?: string } = {}, + options: { + fetchImpl?: typeof fetch; + workerBundlePath?: string; + allowedModels?: string[]; + profileId?: string; + provider?: "tinker" | "openai-compatible"; + providerFetchImpl?: typeof fetch; + } = {}, ): PiAgentRunner { return new PiAgentRunner({ - providerProfiles: staticProviderProfileResolver([fixtureProfile(server)]), + providerProfiles: staticProviderProfileResolver([fixtureProfile( + server, + options.allowedModels, + options.profileId, + options.provider, + )]), workerBundlePath: options.workerBundlePath ?? workerBundlePath, ...(options.fetchImpl ? { fetchImpl: options.fetchImpl } : {}), + ...(options.providerFetchImpl ? { providerFetchImpl: options.providerFetchImpl } : {}), }); } -function fixtureProfile(server: http.Server): ProviderProfile { +function fixtureProfile( + server: http.Server, + allowedModels: string[] = ["fixture-model"], + id = "fixture", + provider: "tinker" | "openai-compatible" = "openai-compatible", +): ProviderProfile { return { - id: "fixture", - provider: "openai-compatible", + id, + provider, baseUrl: serverBaseUrl(server), route: "/chat/completions", apiKeyEnv: "THOUGHTSTREAM_TEST_API_KEY", - allowedModels: new Set(["fixture-model"]), + allowedModels: new Set(allowedModels), imageInputModels: new Set(), jsonObjectResponseFormat: true, + jsonSchemaResponseFormat: false, requestTimeoutMs: 5_000, maxRequestBytes: 2 * 1024 * 1024, maxResponseBytes: 1_000_000, }; } +async function compiledAdapterDeclaration( + privateCheckpoint: string, + options: { atprotoPrefetch?: boolean } = {}, +): Promise { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-pi-adapter-")); + temporaryRoots.push(root); + await fs.mkdir(path.join(root, "adapters", "releases"), { recursive: true }); + await fs.mkdir(path.join(root, "agents"), { recursive: true }); + await fs.mkdir(path.join(root, "prompts"), { recursive: true }); + const releaseManifest = [ + "schemaVersion: 1", + "id: fixture-adapter", + "version: 1", + "description: Fixture learned adapter", + "releasedAt: 2026-07-25T20:00:00.000Z", + "providerProfile: tinker-default", + "baseModel: Qwen/Qwen3.5-4B", + "checkpoint: { env: THOUGHTSTREAM_TEST_ADAPTER_MODEL }", + `dataset: { id: fixture-dataset, sha256: ${"c".repeat(64)} }`, + `evals: [{ id: fixture-eval, sha256: ${"d".repeat(64)} }]`, + "capabilities: [fixture-task]", + "privacyClass: private", + "exportClass: forbidden", + ].join("\n"); + await fs.writeFile(path.join(root, "adapters", "releases", "fixture.yaml"), `${releaseManifest}\n`); + const manifestSha256 = createHash("sha256").update(canonicalJson(YAML.parse(releaseManifest))).digest("hex"); + await fs.writeFile(path.join(root, "adapters", "deployment.yaml"), [ + "schemaVersion: 1", + "generation: 1", + "selections:", + " - declaration: { id: pi-test, version: 1 }", + ` release: { id: fixture-adapter, version: 1, manifestSha256: ${manifestSha256} }`, + " state: active", + "", + ].join("\n")); + await fs.writeFile(path.join(root, "prompts", "fixture.md"), "Read the fixture.\n"); + await fs.writeFile(path.join(root, "agents", "fixture.yaml"), [ + "id: pi-test", + "version: 1", + "description: Exercises the Pi adapter", + options.atprotoPrefetch + ? "subscribe: { types: [stream.thought.source.atproto.commit], sources: ['*'], privacy: [private] }" + : "subscribe: { types: [stream.thought.source.file.changed], sources: ['*'], privacy: [private] }", + "context: { maxEvents: 1, maxChars: 10000, strategy: single-event }", + "runner:", + " kind: pi", + " profile: tinker-default", + " adapter: { id: fixture-adapter, version: 1 }", + " maxOutputTokens: 500", + " timeoutMs: 5000", + "accounting:", + " leaseMs: 10000", + " reservation: { inputTokens: 5000, outputTokens: 500 }", + " limits: [{ window: hour, maxCalls: 10, maxInputTokens: 50000, maxOutputTokens: 5000 }]", + "prompt: prompts/fixture.md", + "emit: [stream.thought.derived.document.read]", + options.atprotoPrefetch + ? "policy: { tools: [atproto.fetch-markdown], externalActions: false }" + : "policy: { tools: [], externalActions: false }", + "enabled: true", + "", + ].join("\n")); + process.env.THOUGHTSTREAM_TINKER_ALLOWED_MODELS = `Qwen/Qwen3.5-4B,${privateCheckpoint}`; + const [declaration] = await loadAgentDeclarations(path.join(root, "agents"), { + THOUGHTSTREAM_TEST_ADAPTER_MODEL: privateCheckpoint, + THOUGHTSTREAM_TINKER_ALLOWED_MODELS: process.env.THOUGHTSTREAM_TINKER_ALLOWED_MODELS, + }); + if (!declaration) throw new Error("Missing compiled adapter declaration"); + return declaration; +} + function fixtureRunInput(declaration: ThoughtAgentDeclaration) { const event: ThoughtEvent = { id: "evt_failure_test", diff --git a/test/provider-broker.test.ts b/test/provider-broker.test.ts index 9c6360f..b87f82c 100644 --- a/test/provider-broker.test.ts +++ b/test/provider-broker.test.ts @@ -21,6 +21,28 @@ afterEach(async () => { }); describe("one-turn provider broker", () => { + test("keeps a rejected private model value out of broker errors", async () => { + const privateModel = "tinker://SYNTHETIC_PRIVATE_BROKER_FIXTURE"; + const upstream = http.createServer((_request, response) => response.end()); + servers.push(upstream); + await new Promise((resolve) => upstream.listen(0, "127.0.0.1", resolve)); + process.env.THOUGHTSTREAM_TEST_API_KEY = "fixture-secret"; + + try { + await startProviderBroker({ + runId: "run_private_model_rejection", + model: privateModel, + profile: fixtureProfile(upstream), + timeoutMs: 5_000, + maxOutputTokens: 100, + }); + throw new Error("Expected private model rejection"); + } catch (error) { + expect(String(error)).toContain("Requested model is not allowlisted"); + expect(String(error)).not.toContain(privateModel); + } + }); + test("injects authorization once, then rejects reuse of the run capability", async () => { let upstreamRequests = 0; let authorizationWasInjected = false; @@ -195,6 +217,100 @@ describe("one-turn provider broker", () => { expect(broker.usage()).toMatchObject({ requests: 1, responseBytes: 5 }); }); + test("atomically converts response reservations before delayed trace completion", async () => { + const upstream = http.createServer((_request, response) => response.end()); + servers.push(upstream); + await new Promise((resolve) => upstream.listen(0, "127.0.0.1", resolve)); + process.env.THOUGHTSTREAM_TEST_API_KEY = "fixture-secret"; + let firstResponseTraceStarted!: () => void; + let releaseFirstResponseTrace!: () => void; + const traceStarted = new Promise((resolve) => { firstResponseTraceStarted = resolve; }); + const traceRelease = new Promise((resolve) => { releaseFirstResponseTrace = resolve; }); + let responseTraceCount = 0; + let upstreamRequests = 0; + const broker = await startProviderBroker({ + runId: "run_broker_response_conversion", + model: "fixture-model", + profile: fixtureProfile(upstream), + timeoutMs: 5_000, + maxOutputTokens: 100, + maxRequests: 2, + maxTotalRequestBytes: 128 * 1024, + maxTotalResponseBytes: 10, + fetchImpl: async () => { + upstreamRequests += 1; + return new Response("12345", { status: 200 }); + }, + onTrace: async (kind) => { + if (kind !== "sandbox.provider.response") return; + responseTraceCount += 1; + if (responseTraceCount === 1) { + firstResponseTraceStarted(); + await traceRelease; + } + }, + }); + brokers.push(broker); + const request = fixtureRequest(broker, "run_broker_response_conversion"); + const first = exchange(broker.socketPath, request); + await traceStarted; + const second = await exchange(broker.socketPath, request); + expect(second).toMatchObject({ status: "completed", body: "12345" }); + expect(upstreamRequests).toBe(2); + expect(broker.usage()).toEqual({ + requests: 2, + requestBytes: Buffer.byteLength(request.body) * 2, + responseBytes: 10, + }); + releaseFirstResponseTrace(); + await expect(first).resolves.toMatchObject({ status: "completed", body: "12345" }); + }); + + test("atomically admits one concurrent socket against request-count and byte budgets", async () => { + const upstream = http.createServer((_request, response) => response.end()); + servers.push(upstream); + await new Promise((resolve) => upstream.listen(0, "127.0.0.1", resolve)); + process.env.THOUGHTSTREAM_TEST_API_KEY = "fixture-secret"; + let upstreamRequests = 0; + let releaseUpstream!: () => void; + const release = new Promise((resolve) => { releaseUpstream = resolve; }); + const profile = fixtureProfile(upstream); + const provisionalBody = JSON.stringify({ model: "fixture-model", stream: true, max_tokens: 100, messages: [] }); + const broker = await startProviderBroker({ + runId: "run_broker_atomic_admission", + model: "fixture-model", + profile, + timeoutMs: 5_000, + maxOutputTokens: 100, + maxRequests: 1, + maxTotalRequestBytes: Buffer.byteLength(provisionalBody), + maxTotalResponseBytes: 5, + fetchImpl: async () => { + upstreamRequests += 1; + await release; + return new Response("12345", { status: 200 }); + }, + }); + brokers.push(broker); + const request = fixtureRequest(broker, "run_broker_atomic_admission"); + const exchanges = Array.from({ length: 32 }, () => exchange(broker.socketPath, request)); + await vi.waitFor(() => expect(upstreamRequests).toBe(1)); + releaseUpstream(); + const responses = await Promise.all(exchanges); + + expect(responses.filter((response) => response.status === "completed")).toHaveLength(1); + expect(responses.filter((response) => response.status === "rejected")).toHaveLength(31); + expect(responses.filter((response) => response.status === "rejected").every((response) => ( + response.status === "rejected" && response.code === "turn-exhausted" + ))).toBe(true); + expect(upstreamRequests).toBe(1); + expect(broker.usage()).toEqual({ + requests: 1, + requestBytes: Buffer.byteLength(request.body), + responseBytes: 5, + }); + }); + test("rejects expired, forged, and authorization-bearing requests without egress", async () => { let upstreamRequests = 0; const upstream = http.createServer((_request, response) => { upstreamRequests += 1; response.end(); }); @@ -231,6 +347,7 @@ function fixtureProfile(server: http.Server): ProviderProfile { allowedModels: new Set(["fixture-model"]), imageInputModels: new Set(), jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: false, requestTimeoutMs: 5_000, maxRequestBytes: 64 * 1024, maxResponseBytes: 64 * 1024, diff --git a/test/provider-profiles.test.ts b/test/provider-profiles.test.ts index 07d527c..47dff95 100644 --- a/test/provider-profiles.test.ts +++ b/test/provider-profiles.test.ts @@ -14,6 +14,8 @@ describe("trusted provider profiles", () => { route: "/chat/completions", apiKeyEnv: "TINKER_API_KEY", imageInputModels: new Set(), + jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: false, }); expect(resolver.resolve("tinker-default", "openai/gpt-oss-120b")).toMatchObject({ provider: "tinker", @@ -40,4 +42,33 @@ describe("trusted provider profiles", () => { THOUGHTSTREAM_MODEL_ALLOWED_MODELS: "small", }).resolve("openai-compatible-default", "small")).toThrow("must use HTTPS"); }); + + test("keeps the JSON-enforcing OpenAI profile fixed and narrow", () => { + const resolver = createBuiltinProviderProfileResolver({}); + expect(resolver.resolve("openai-json-default", "gpt-4.1-mini")).toMatchObject({ + provider: "openai-compatible", + baseUrl: "https://api.openai.com/v1", + route: "/chat/completions", + apiKeyEnv: "OPENAI_API_KEY", + allowedModels: new Set(["gpt-4.1-mini"]), + jsonObjectResponseFormat: false, + jsonSchemaResponseFormat: true, + }); + expect(() => resolver.resolve("openai-json-default", "gpt-4.1")).toThrow("not allowlisted"); + }); + + test("keeps non-allowlisted image model values out of profile errors", () => { + const privateModel = "tinker://SYNTHETIC_PRIVATE_IMAGE_FIXTURE"; + try { + createBuiltinProviderProfileResolver({ + THOUGHTSTREAM_MODEL_BASE_URL: "https://models.example.test/v1", + THOUGHTSTREAM_MODEL_ALLOWED_MODELS: "small", + THOUGHTSTREAM_MODEL_IMAGE_MODELS: privateModel, + }).resolve("openai-compatible-default", "small"); + throw new Error("Expected image-model profile rejection"); + } catch (error) { + expect(String(error)).toContain("marks a non-allowlisted model as image-capable"); + expect(String(error)).not.toContain(privateModel); + } + }); }); diff --git a/test/rate-limit.test.ts b/test/rate-limit.test.ts new file mode 100644 index 0000000..20a0a23 --- /dev/null +++ b/test/rate-limit.test.ts @@ -0,0 +1,36 @@ +import { describe, expect, test } from "vitest"; +import { FixedWindowRateLimiter, createOAuthRouteRateLimiter } from "../src/web/rate-limit.js"; + +describe("OAuth process rate limits", () => { + test("bounds keys, rejects excess requests, and recovers after the window", () => { + let now = 1_000; + const limiter = new FixedWindowRateLimiter({ + maxRequests: 2, + windowMs: 10_000, + maxKeys: 2, + now: () => now, + }); + + expect(limiter.check("a").allowed).toBe(true); + expect(limiter.check("a").allowed).toBe(true); + expect(limiter.check("a")).toMatchObject({ allowed: false, retryAfterSeconds: 10 }); + expect(limiter.check("b").allowed).toBe(true); + expect(limiter.check("c").allowed).toBe(false); + now += 10_001; + expect(limiter.check("c").allowed).toBe(true); + }); + + test("applies both per-client and global OAuth route ceilings", () => { + const limiter = createOAuthRouteRateLimiter(() => 1_000); + for (let request = 0; request < 6; request += 1) { + expect(limiter.login("198.51.100.1").allowed).toBe(true); + } + expect(limiter.login("198.51.100.1").allowed).toBe(false); + expect(limiter.login("198.51.100.2").allowed).toBe(true); + + for (let request = 0; request < 12; request += 1) { + expect(limiter.callback("198.51.100.3").allowed).toBe(true); + } + expect(limiter.callback("198.51.100.3").allowed).toBe(false); + }); +}); diff --git a/test/repairs.test.ts b/test/repairs.test.ts index b221ee3..875149d 100644 --- a/test/repairs.test.ts +++ b/test/repairs.test.ts @@ -1,5 +1,6 @@ import fs from "node:fs/promises"; import { afterEach, describe, expect, test } from "vitest"; +import type { ModelAdapterIdentity } from "../src/adapters/model-adapters.js"; import { outputContractForDeclaration, outputContractIdentityJson } from "../src/agents/output-contracts.js"; import { CORRECTION_PROPOSAL_EVENT_TYPE, @@ -87,6 +88,12 @@ describe("append-only output repair", () => { repairRunId: expect.any(String), originalTriggerEventId: requests[0]!.payload.originalTriggerEventId, sourceRootEventId: requests[0]!.rootEventId, + originalModel: requests[0]!.payload.model, + repairModel: { + provider: "tinker", + id: "fixture-model", + executionAdapterRevision: "pi-openai-completions-v1", + }, structuredOutput: validOutput("Conservative repaired observation"), }, }); @@ -280,6 +287,16 @@ describe("append-only output repair", () => { authority: "accept", repairRunId: repairRun.id, judgmentEventId: accepted.id, + originalModel: { + provider: originalRun.provider, + id: originalRun.model, + executionAdapterRevision: originalRun.executionAdapterRevision, + }, + repairModel: { + provider: repairRun.provider, + id: repairRun.model, + executionAdapterRevision: repairRun.executionAdapterRevision, + }, structuredOutput: validOutput("Proposed repair"), }, }); @@ -325,6 +342,11 @@ describe("append-only output repair", () => { outputContract: outputContractIdentityJson(outputContractForDeclaration(original)), repair: true, }), + originalProvenance: expect.objectContaining({ + provider: originalRun.provider, + model: originalRun.model, + outputContract: outputContractIdentityJson(outputContractForDeclaration(original)), + }), }), ]); const serializedExamples = JSON.stringify(examples); @@ -333,6 +355,13 @@ describe("append-only output repair", () => { expect(serializedExamples).not.toContain("SECRET_DISALLOWED_TRACE_PAYLOAD"); expect(serializedExamples).not.toContain("A source item that needs a bounded observation"); + await store.upsertRun({ ...originalRun, modelAdapter: forbiddenAdapterIdentity(), privacy: "private" }); + expect(await projectTrainingExamples(store, { + includeSensitivePrivate: true, + includeRestrictedModelAdapters: true, + })).toEqual([]); + await store.upsertRun(originalRun); + const source = (await store.listEvents({ types: ["stream.thought.source.rss.item"] }))[0]!; const feedback = await store.appendEvent({ type: "stream.thought.source.telegram.reaction", @@ -580,3 +609,22 @@ function repairDeclaration(): ThoughtAgentDeclaration { function validOutput(summary: string) { return { summary, tags: ["repair"], importance: "normal" as const, confidence: 0.7 }; } + +function forbiddenAdapterIdentity(): ModelAdapterIdentity { + return { + id: "forbidden-original-adapter", + version: 1, + description: "Forbidden original repair fixture", + releasedAt: "2026-07-15T00:00:00.000Z", + manifestSha256: "a".repeat(64), + checkpointSelector: { kind: "env", reference: "THOUGHTSTREAM_FORBIDDEN_FIXTURE" }, + providerProfile: "tinker-default", + baseModel: "Qwen/Qwen3.5-4B", + dataset: { id: "forbidden-dataset", sha256: "b".repeat(64) }, + evals: [{ id: "forbidden-eval", sha256: "c".repeat(64) }], + capabilities: ["repair-fixture"], + privacyClass: "private", + exportClass: "forbidden", + binding: { checkpointReferenceSha256: "d".repeat(64) }, + }; +} diff --git a/test/review-web-capability.test.ts b/test/review-web-capability.test.ts new file mode 100644 index 0000000..4856d2d --- /dev/null +++ b/test/review-web-capability.test.ts @@ -0,0 +1,61 @@ +import { describe, expect, test } from "vitest"; +import { + REVIEW_NONCE_HEADER, + REVIEW_SIGNATURE_HEADER, + REVIEW_TIMESTAMP_HEADER, + ReviewCapabilityVerifier, + decodeReviewCapability, + signReviewRequest, +} from "../src/review/web-capability.js"; + +describe("Review loopback capability", () => { + test("binds method, path, body, time, and a one-time nonce", () => { + const key = Buffer.alloc(32, 17); + const now = 1_800_000_000_000; + const path = "/api/reviews/review%3Aitem/decisions"; + const body = Buffer.from('{"disposition":"skip"}', "utf8"); + const signed = signReviewRequest(key, { + method: "POST", + path, + body, + timestamp: now, + nonce: "A".repeat(32), + }); + const headers = { + [REVIEW_TIMESTAMP_HEADER]: signed.timestamp, + [REVIEW_NONCE_HEADER]: signed.nonce, + [REVIEW_SIGNATURE_HEADER]: signed.signature, + }; + const verifier = new ReviewCapabilityVerifier(key, { now: () => now }); + expect(verifier.verify(headers, "POST", path, body)).toBe(true); + expect(verifier.verify(headers, "POST", path, body)).toBe(false); + + const fresh = (suffix: string, at = now) => { + const signature = signReviewRequest(key, { + method: "POST", + path, + body, + timestamp: at, + nonce: suffix.repeat(32), + }); + return { + [REVIEW_TIMESTAMP_HEADER]: signature.timestamp, + [REVIEW_NONCE_HEADER]: signature.nonce, + [REVIEW_SIGNATURE_HEADER]: signature.signature, + }; + }; + expect(verifier.verify(fresh("B"), "GET", path, body)).toBe(false); + expect(verifier.verify(fresh("C"), "POST", `${path}/other`, body)).toBe(false); + expect(verifier.verify(fresh("D"), "POST", path, Buffer.from("{}"))).toBe(false); + expect(verifier.verify(fresh("E", now - 31_000), "POST", path, body)).toBe(false); + expect(verifier.verify({ ...fresh("F"), [REVIEW_SIGNATURE_HEADER]: "x".repeat(43) }, "POST", path, body)).toBe(false); + }); + + test("loads only canonical bounded key material", () => { + const encoded = Buffer.alloc(32, 9).toString("base64"); + expect(decodeReviewCapability(encoded)).toEqual(Buffer.alloc(32, 9)); + expect(decodeReviewCapability(undefined)).toBeUndefined(); + expect(() => decodeReviewCapability("not base64 !!!")).toThrow("canonical base64"); + expect(() => decodeReviewCapability(Buffer.alloc(8).toString("base64"))).toThrow("32 to 128"); + }); +}); diff --git a/test/review.test.ts b/test/review.test.ts new file mode 100644 index 0000000..c7bf23f --- /dev/null +++ b/test/review.test.ts @@ -0,0 +1,472 @@ +import fs from "node:fs/promises"; +import path from "node:path"; +import { afterEach, describe, expect, test } from "vitest"; +import { + REVIEW_RESPONSE_OUTPUT_CONTRACT, + canonicalStructuredOutput, + createOutputContractRegistry, + outputContractIdentityJson, + reviewResponseSummary, +} from "../src/agents/output-contracts.js"; +import type { JsonObject } from "../src/core/json.js"; +import type { PrivacyClass, ThoughtEvent } from "../src/events/types.js"; +import type { JazzThoughtStore } from "../src/jazz/store.js"; +import { + activeReviewDecisions, + appendReviewPrompt, + createReviewItem, + projectReviewQueue, + recordBrowserReviewDecision, +} from "../src/review/review.js"; +import { REVIEW_RESPONSE_EVENT_TYPE } from "../src/review/types.js"; +import { + REVIEW_NONCE_HEADER, + REVIEW_SIGNATURE_HEADER, + REVIEW_TIMESTAMP_HEADER, + signReviewRequest, +} from "../src/review/web-capability.js"; +import type { AgentRun } from "../src/store/types.js"; +import { projectTrainingExamples, writeTrainingJsonl } from "../src/training/judgments.js"; +import { startInspectorServer } from "../src/web/inspector.js"; +import { temporaryProject, testStore } from "./helpers.js"; + +const stores: JazzThoughtStore[] = []; +const roots: string[] = []; +const servers: import("node:http").Server[] = []; + +afterEach(async () => { + await Promise.all(servers.splice(0).map((server) => new Promise((resolve) => { + server.closeAllConnections(); + server.close(() => resolve()); + }))); + await Promise.all(stores.splice(0).map((store) => store.close())); + await Promise.all(roots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); +}); + +describe("ThoughtStream Review", () => { + test("freezes blinded same-trigger pairs and exports only active judgeable public decisions", async () => { + const { store, item } = await fixture(); + const unresolved = await projectReviewQueue(store); + expect(unresolved.counts).toMatchObject({ total: 1, unresolved: 1, decided: 0 }); + expect(unresolved.items[0]?.candidates.map((candidate) => Object.keys(candidate))).toEqual([ + ["label", "response"], + ["label", "response"], + ]); + expect(JSON.stringify(unresolved)).not.toContain("fixture/model-a"); + expect(JSON.stringify(unresolved)).not.toContain("run_a"); + + const underdetermined = await recordBrowserReviewDecision(store, item.id, { + disposition: "underdetermined", + confidence: "medium", + reasonCodes: ["missing-mechanism"], + responseTags: [], + notes: "The prompt lacks the critique that would make the mechanism change judgeable.", + trainingEligible: false, + submissionId: "submission-underdetermined-0001", + }); + expect((await projectTrainingExamples(store))).toEqual([]); + + const decidedQueue = await projectReviewQueue(store); + expect(decidedQueue.items[0]?.decision).toMatchObject({ + eventId: underdetermined.id, + disposition: "underdetermined", + trainingEligible: false, + externalExportEligible: false, + }); + expect(decidedQueue.items[0]?.candidates[0]?.provenance).toBeDefined(); + + const preferredLabel = decidedQueue.items[0]!.candidates[0].label; + const preferredResponse = decidedQueue.items[0]!.candidates[0].response; + const rejectedResponse = decidedQueue.items[0]!.candidates[1].response; + const preferred = await recordBrowserReviewDecision(store, item.id, { + disposition: "prefer", + preferredCandidate: preferredLabel, + preferenceStrength: "strong", + confidence: "high", + reasonCodes: ["mechanism-update"], + responseTags: ["specific"], + notes: "private note excluded from the dataset", + trainingEligible: true, + submissionId: "submission-preference-00000002", + supersedesDecisionEventId: underdetermined.id, + }); + const repeated = await recordBrowserReviewDecision(store, item.id, { + disposition: "prefer", + preferredCandidate: preferredLabel, + preferenceStrength: "strong", + confidence: "high", + reasonCodes: ["mechanism-update"], + responseTags: ["specific"], + notes: "private note excluded from the dataset", + trainingEligible: true, + submissionId: "submission-preference-00000002", + supersedesDecisionEventId: underdetermined.id, + }); + expect(repeated.id).toBe(preferred.id); + const active = await activeReviewDecisions(store); + expect(active.active.map((event) => event.id)).toEqual([preferred.id]); + expect(active.inactiveIds.has(underdetermined.id)).toBe(true); + + const examples = await projectTrainingExamples(store); + expect(examples).toHaveLength(1); + expect(examples[0]).toMatchObject({ + format: "thoughtstream.training-example.v4", + kind: "prefer", + judgment: { + criterion: "response-quality", + criterionVersion: 1, + preferenceStrength: "strong", + confidence: "high", + reasonCodes: ["mechanism-update"], + responseTags: ["specific"], + }, + input: { + prompt: "Given the critique and evidence, explain which mechanism should change.", + evidence: "The user critique says the retrieval gate was bypassed in the failed run.", + }, + chosen: { response: preferredResponse, summary: preferredResponse, confidence: 1 }, + rejected: { response: rejectedResponse, summary: rejectedResponse, confidence: 1 }, + review: { + campaignId: "review-pipeline-canary", + campaignVersion: 1, + disposition: "prefer", + candidateCount: 2, + }, + }); + const serialized = JSON.stringify(examples[0]); + expect(serialized).not.toContain("private note excluded"); + expect(serialized).not.toContain("submission-preference"); + expect(serialized).not.toContain(item.id); + expect(serialized).not.toContain(preferred.id); + + const output = path.join(roots[0]!, "review-dataset.jsonl"); + const manifest = await writeTrainingJsonl(output, examples); + expect(manifest).toMatchObject({ + format: "thoughtstream.training-dataset-manifest.v4", + examples: 1, + exampleFormats: { "thoughtstream.training-example.v4": 1 }, + reviewCampaigns: ["review-pipeline-canary@1"], + }); + }); + + test("records contract-valid corrections with both rejected candidates", async () => { + const { store, item } = await fixture(); + await expect(recordBrowserReviewDecision(store, item.id, { + disposition: "correct", + replacementResponse: "", + reasonCodes: [], + responseTags: [], + trainingEligible: true, + submissionId: "submission-invalid-correction", + })).rejects.toThrow(); + await recordBrowserReviewDecision(store, item.id, { + disposition: "correct", + replacementResponse: "The critique shows retrieval was bypassed, so the mechanism change is to enforce retrieval before generation and test that invariant.", + confidence: "high", + reasonCodes: ["mechanism-update"], + responseTags: ["specific"], + trainingEligible: true, + submissionId: "submission-valid-correction-01", + }); + const [example] = await projectTrainingExamples(store); + expect(example).toMatchObject({ + format: "thoughtstream.training-example.v4", + kind: "correct", + chosen: { + response: "The critique shows retrieval was bypassed, so the mechanism change is to enforce retrieval before generation and test that invariant.", + }, + additionalRejected: [expect.objectContaining({ response: "Candidate B proposes a vague tone adjustment." })], + }); + }); + + test("keeps private review useful but browser-ineligible for external training export", async () => { + const { store, item } = await fixture({ privacy: "private", externalExportEligible: false }); + const queue = await projectReviewQueue(store); + await recordBrowserReviewDecision(store, item.id, { + disposition: "prefer", + preferredCandidate: queue.items[0]!.candidates[0].label, + preferenceStrength: "slight", + reasonCodes: [], + responseTags: [], + trainingEligible: true, + submissionId: "submission-private-preference", + }); + const reviewed = await projectReviewQueue(store); + expect(reviewed.items[0]?.decision).toMatchObject({ + trainingEligible: true, + externalExportEligible: false, + }); + expect(await projectTrainingExamples(store, { includeSensitivePrivate: true })).toEqual([]); + await expect(appendReviewPrompt(store, { + externalId: "invalid-private-declassification", + privacy: "private", + payload: promptPayload(true), + })).rejects.toThrow("Only public-source"); + }); + + test("rejects cross-trigger pairs, undeclared agents, and forced labels on unjudgeable items", async () => { + const { store, prompt, runs } = await fixture({ materialize: false }); + const otherPrompt = await appendReviewPrompt(store, { + externalId: "other-prompt", + privacy: "public-source", + payload: promptPayload(true), + }); + const other = await candidate(store, otherPrompt, "agent-b", "run_other", "Other response."); + await expect(createReviewItem(store, { + promptEventId: prompt.id, + candidateRunIds: [runs[0].id, other.id], + })).rejects.toThrow("exact review prompt"); + const undeclared = await candidate(store, prompt, "agent-z", "run_undeclared", "Undeclared response."); + await expect(createReviewItem(store, { + promptEventId: prompt.id, + candidateRunIds: [runs[0].id, undeclared.id], + })).rejects.toThrow("not declared"); + const item = await createReviewItem(store, { + promptEventId: prompt.id, + candidateRunIds: [runs[0].id, runs[1].id], + }); + await expect(recordBrowserReviewDecision(store, item.id, { + disposition: "underdetermined", + reasonCodes: [], + responseTags: [], + trainingEligible: true, + submissionId: "submission-forced-label-invalid", + })).rejects.toThrow(); + await expect(recordBrowserReviewDecision(store, item.id, { + disposition: "prefer", + preferredCandidate: "A", + preferenceStrength: "strong", + reasonCodes: ["invented-reason"], + responseTags: [], + trainingEligible: false, + submissionId: "submission-unknown-reason", + })).rejects.toThrow("Unknown review reason code"); + }); + + test("accepts only a fresh body-bound loopback capability on the exact Review route", async () => { + const { store, item } = await fixture(); + const key = Buffer.alloc(32, 29); + const server = await startInspectorServer(store, { port: 0, reviewCapability: key }); + servers.push(server); + const address = server.address(); + if (!address || typeof address === "string") throw new Error("Missing inspector address"); + const base = `http://127.0.0.1:${address.port}`; + const pathName = `/api/reviews/${encodeURIComponent(item.id)}/decisions`; + const body = Buffer.from(JSON.stringify({ + disposition: "skip", + reasonCodes: [], + responseTags: [], + trainingEligible: false, + submissionId: "submission-loopback-canary-01", + }), "utf8"); + const unsigned = await fetch(`${base}${pathName}`, { + method: "POST", + headers: { "content-type": "application/json" }, + body, + }); + expect(unsigned.status).toBe(403); + expect((await activeReviewDecisions(store)).active).toEqual([]); + + const signed = signReviewRequest(key, { method: "POST", path: pathName, body }); + const headers = { + "content-type": "application/json", + [REVIEW_TIMESTAMP_HEADER]: signed.timestamp, + [REVIEW_NONCE_HEADER]: signed.nonce, + [REVIEW_SIGNATURE_HEADER]: signed.signature, + }; + const accepted = await fetch(`${base}${pathName}`, { method: "POST", headers, body }); + expect(accepted.status).toBe(201); + expect((await accepted.json()) as JsonObject).toMatchObject({ active: true }); + expect((await activeReviewDecisions(store)).active).toHaveLength(1); + const replay = await fetch(`${base}${pathName}`, { method: "POST", headers, body }); + expect(replay.status).toBe(403); + expect((await activeReviewDecisions(store)).active).toHaveLength(1); + }); + + test("fails queue, decision, and export closed when a frozen run receipt or output pointer changes", async () => { + const { store, item, runs } = await fixture(); + const queue = await projectReviewQueue(store); + await recordBrowserReviewDecision(store, item.id, { + disposition: "prefer", + preferredCandidate: queue.items[0]!.candidates[0].label, + preferenceStrength: "strong", + reasonCodes: [], + responseTags: [], + trainingEligible: true, + submissionId: "submission-before-run-tamper", + }); + await store.upsertRun({ ...runs[0], model: "fixture/tampered-model", updatedAt: "2026-07-26T23:00:00.000Z" }); + await expect(projectReviewQueue(store)).rejects.toThrow("receipt changed after materialization"); + await expect(projectTrainingExamples(store)).rejects.toThrow("receipt changed after materialization"); + await expect(recordBrowserReviewDecision(store, item.id, { + disposition: "skip", + reasonCodes: [], + responseTags: [], + trainingEligible: false, + submissionId: "submission-after-run-tamper", + })).rejects.toThrow("receipt changed after materialization"); + + const second = await fixture(); + await second.store.upsertRun({ + ...second.runs[0], + outputEventIds: [second.runs[1].outputEventIds[0]!], + updatedAt: "2026-07-26T23:00:00.000Z", + }); + await expect(projectReviewQueue(second.store)).rejects.toThrow("output pointer changed after materialization"); + }); + + test("rejects truncated, projected, enriched, or multi-event candidate context before materialization", async () => { + const { store, prompt, runs } = await fixture({ materialize: false }); + const create = () => createReviewItem(store, { + promptEventId: prompt.id, + candidateRunIds: [runs[0].id, runs[1].id], + }); + await store.upsertRun({ + ...runs[0], + contextManifest: { ...runs[0].contextManifest, truncated: true, sourceIncludedChars: 999 }, + updatedAt: "2026-07-26T23:00:00.000Z", + }); + await expect(create()).rejects.toThrow("truncated, projected, enriched, or action-capable"); + + await store.upsertRun({ + ...runs[0], + inputEventIds: [prompt.id, "event_extra"], + contextManifest: { + ...runs[0].contextManifest, + inputEventIds: [prompt.id, "event_extra"], + includedEventIds: [prompt.id, "event_extra"], + }, + updatedAt: "2026-07-26T23:00:01.000Z", + }); + await expect(create()).rejects.toThrow("exactly one prompt input"); + + await store.upsertRun({ + ...runs[0], + contextManifest: { ...runs[0].contextManifest, payloadFields: ["prompt"], tools: ["web.download-image"] }, + updatedAt: "2026-07-26T23:00:02.000Z", + }); + await expect(create()).rejects.toThrow("truncated, projected, enriched, or action-capable"); + }); +}); + +async function fixture(options: { + privacy?: PrivacyClass; + externalExportEligible?: boolean; + materialize?: boolean; +} = {}): Promise<{ + store: JazzThoughtStore; + prompt: ThoughtEvent; + runs: [AgentRun, AgentRun]; + item: ThoughtEvent; +}> { + const project = await temporaryProject("thoughtstream-review-"); + roots.push(project); + const store = testStore(project); + stores.push(store); + const privacy = options.privacy ?? "public-source"; + const prompt = await appendReviewPrompt(store, { + externalId: "prompt-1", + occurredAt: "2026-07-26T22:00:00.000Z", + privacy, + payload: promptPayload(options.externalExportEligible ?? true), + }); + const runs: [AgentRun, AgentRun] = [ + await candidate(store, prompt, "agent-a", "run_a", "Candidate A identifies the retrieval mechanism and proposes a concrete gate."), + await candidate(store, prompt, "agent-b", "run_b", "Candidate B proposes a vague tone adjustment."), + ]; + const item = options.materialize === false + ? ({ id: "not-materialized" } as ThoughtEvent) + : await createReviewItem(store, { promptEventId: prompt.id, candidateRunIds: [runs[0].id, runs[1].id] }); + return { store, prompt, runs, item }; +} + +function promptPayload(externalExportEligible: boolean): JsonObject { + return { + campaign: { id: "review-pipeline-canary", version: 1, label: "Review pipeline canary" }, + prompt: "Given the critique and evidence, explain which mechanism should change.", + evidence: "The user critique says the retrieval gate was bypassed in the failed run.", + criterion: { + id: "response-quality", + version: 1, + label: "Mechanism-changing response quality", + instructions: "Prefer the response that identifies the changed mechanism and preserves evidence boundaries.", + reasonCodes: ["missing-mechanism", "mechanism-update"], + responseTags: ["specific", "vague"], + }, + candidateAgentIds: ["agent-a", "agent-b"], + externalExportEligible, + }; +} + +async function candidate( + store: JazzThoughtStore, + prompt: ThoughtEvent, + agentId: string, + runId: string, + response: string, +): Promise { + const structuredOutput = canonicalStructuredOutput( + createOutputContractRegistry(), + REVIEW_RESPONSE_OUTPUT_CONTRACT.identity, + { response }, + ); + const output = (await store.appendEvent({ + type: REVIEW_RESPONSE_EVENT_TYPE, + schemaVersion: 1, + source: `agent:${agentId}`, + sourceKind: "agent", + externalId: runId, + idempotencyKey: `${runId}:output`, + occurredAt: "2026-07-26T22:00:01.000Z", + actor: agentId, + rootEventId: prompt.id, + parentEventId: prompt.id, + correlationId: prompt.correlationId, + privacy: prompt.privacy, + payload: { + runId, + executionKey: `execution:${runId}`, + inputEventId: prompt.id, + inputSourceSequence: prompt.sourceSequence, + summary: reviewResponseSummary(response), + outputContract: outputContractIdentityJson(REVIEW_RESPONSE_OUTPUT_CONTRACT.identity), + structuredOutput, + model: { provider: "fixture", id: `fixture/${agentId}` }, + }, + })).event; + const run: AgentRun = { + id: runId, + executionKey: `execution:${runId}`, + triggerEventId: prompt.id, + agentId, + agentVersion: 1, + status: "completed", + inputEventIds: [prompt.id], + outputEventIds: [output.id], + attempt: 1, + provider: "fixture", + model: `fixture/${agentId}`, + privacy: prompt.privacy, + promptHash: `prompt-hash-${agentId}`, + contextManifest: { + contextStrategy: "single-event", + inputEventIds: [prompt.id], + includedEventIds: [prompt.id], + omittedEventIds: [], + maxEvents: 1, + maxChars: 220_000, + sourceOriginalChars: 1_000, + sourceIncludedChars: 1_000, + truncated: false, + outputContract: outputContractIdentityJson(REVIEW_RESPONSE_OUTPUT_CONTRACT.identity), + tools: [], + externalActions: false, + }, + result: structuredOutput, + createdAt: "2026-07-26T22:00:00.000Z", + completedAt: "2026-07-26T22:00:02.000Z", + updatedAt: "2026-07-26T22:00:02.000Z", + }; + await store.upsertRun(run); + return run; +} diff --git a/test/runtime-failures.test.ts b/test/runtime-failures.test.ts index 968ef8e..dfe1a0c 100644 --- a/test/runtime-failures.test.ts +++ b/test/runtime-failures.test.ts @@ -155,7 +155,7 @@ describe("ThoughtAgentRuntime failures", () => { provider: "tinker", model: "Qwen/Qwen3.5-4B", checkpointRevision: "checkpoint-fixture", - adapterRevision: "deterministic-triage-v1", + executionAdapterRevision: "deterministic-triage-v1", }); const events = await store.listEvents(); const failedEvents = events.filter((event) => event.type === "stream.thought.agent.run.failed"); diff --git a/test/secure-store.test.ts b/test/secure-store.test.ts new file mode 100644 index 0000000..e97ad38 --- /dev/null +++ b/test/secure-store.test.ts @@ -0,0 +1,187 @@ +import fs from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { afterEach, describe, expect, test, vi } from "vitest"; +import { SecureJsonStore } from "../src/web/secure-store.js"; + +const roots: string[] = []; + +afterEach(async () => { + await Promise.all(roots.splice(0).map((root) => fs.rm(root, { recursive: true, force: true }))); +}); + +describe("encrypted OAuth store", () => { + test("persists authenticated ciphertext with owner-only modes and no plaintext", async () => { + const root = await temporaryRoot(); + const key = Buffer.alloc(32, 7); + const sentinel = "refresh-token-private-sentinel"; + const store = storeAt<{ token: string }>(root, "session", key); + + await store.set("did:plc:fixture", { token: sentinel }); + + const file = path.join(root, "session.enc.json"); + const raw = await fs.readFile(file, "utf8"); + expect(raw).not.toContain(sentinel); + expect(raw).not.toContain("did:plc:fixture"); + expect((await fs.stat(root)).mode & 0o777).toBe(0o700); + expect((await fs.stat(file)).mode & 0o777).toBe(0o600); + const restored = storeAt<{ token: string }>(root, "session", key); + expect(await restored.get("did:plc:fixture")).toEqual({ token: sentinel }); + + const wrongKey = storeAt<{ token: string }>(root, "session", Buffer.alloc(32, 8)); + await expect(wrongKey.initialize()).rejects.toThrow("authenticated or decoded"); + }); + + test("rejects direct malformed encrypted-envelope fixtures", async () => { + const root = await temporaryRoot(); + await fs.writeFile(path.join(root, "malformed.enc.json"), "{not-json\n", { mode: 0o600 }); + const malformed = storeAt<{ value: string }>(root, "malformed", Buffer.alloc(32, 17)); + await expect(malformed.initialize()).rejects.toThrow("envelope is invalid"); + + await fs.writeFile(path.join(root, "unsupported.enc.json"), JSON.stringify({ + version: 1, + algorithm: "plaintext", + iv: "", + ciphertext: "", + tag: "", + }), { mode: 0o600 }); + const unsupported = storeAt<{ value: string }>(root, "unsupported", Buffer.alloc(32, 17)); + await expect(unsupported.initialize()).rejects.toThrow("version is unsupported"); + }); + + test("rejects an oversized encrypted-envelope fixture before reading or decrypting it", async () => { + const root = await temporaryRoot(); + await fs.writeFile(path.join(root, "oversized.enc.json"), "x".repeat(6_000), { mode: 0o600 }); + const oversized = new SecureJsonStore<{ value: string }>({ + directory: root, + name: "oversized", + key: Buffer.alloc(32, 18), + maxEntries: 2, + maxSerializedBytes: 1_024, + }); + await expect(oversized.initialize()).rejects.toThrow("envelope exceeds its configured bound"); + }); + + test("expires short-lived records and take is one-time", async () => { + const root = await temporaryRoot(); + let now = 1_000_000; + const store = new SecureJsonStore<{ value: string }>({ + directory: root, + name: "state", + key: Buffer.alloc(32, 9), + maxEntries: 8, + maxSerializedBytes: 8 * 1024, + ttlMs: 1_000, + now: () => now, + }); + + await store.set("one", { value: "first" }); + expect(await store.take("one")).toEqual({ value: "first" }); + expect(await store.take("one")).toBeUndefined(); + await store.set("two", { value: "second" }); + now += 1_001; + expect(await store.getWithExpiration("two")).toMatchObject({ expired: true, value: { value: "second" } }); + expect(await store.get("two")).toBeUndefined(); + }); + + test("rejects entry-count exhaustion without corrupting accepted state", async () => { + const root = await temporaryRoot(); + const store = new SecureJsonStore<{ value: string }>({ + directory: root, + name: "bounded-count", + key: Buffer.alloc(32, 10), + maxEntries: 2, + maxSerializedBytes: 8 * 1024, + }); + + await store.set("a", { value: "one" }); + await store.set("b", { value: "two" }); + await expect(store.set("c", { value: "three" })).rejects.toThrow("entry limit"); + expect(await store.size()).toBe(2); + expect(await store.get("a")).toEqual({ value: "one" }); + expect(await store.get("c")).toBeUndefined(); + }); + + test("rechecks authority after temporary write and before atomic rename", async () => { + const root = await temporaryRoot(); + const store = storeAt<{ value: string }>(root, "guarded-write", Buffer.alloc(32, 19)); + await store.initialize(); + let authoritative = true; + const originalWriteFile = fs.writeFile.bind(fs); + const writeFile = vi.spyOn(fs, "writeFile").mockImplementation(async (...args: Parameters) => { + await originalWriteFile(...args); + authoritative = false; + }); + try { + await expect(store.set("late", { value: "must-not-commit" }, () => authoritative)) + .rejects.toThrow("lost authority"); + } finally { + writeFile.mockRestore(); + } + expect(await store.get("late")).toBeUndefined(); + const reopened = storeAt<{ value: string }>(root, "guarded-write", Buffer.alloc(32, 19)); + expect(await reopened.get("late")).toBeUndefined(); + }); + + test("has no fallible filesystem operation after the atomic rename commit point", async () => { + const root = await temporaryRoot(); + const key = Buffer.alloc(32, 15); + const store = storeAt<{ value: string }>(root, "rename-commit", key); + await store.initialize(); + const chmod = vi.spyOn(fs, "chmod").mockRejectedValue(new Error("post-rename chmod must not run")); + try { + await store.set("committed", { value: "disk-and-memory-agree" }); + expect(chmod).not.toHaveBeenCalled(); + } finally { + chmod.mockRestore(); + } + expect(await store.get("committed")).toEqual({ value: "disk-and-memory-agree" }); + const reopened = storeAt<{ value: string }>(root, "rename-commit", key); + expect(await reopened.get("committed")).toEqual({ value: "disk-and-memory-agree" }); + expect((await fs.stat(path.join(root, "rename-commit.enc.json"))).mode & 0o777).toBe(0o600); + }); + + test("atomically replaces prior browser-session entries", async () => { + const root = await temporaryRoot(); + const store = storeAt<{ value: string }>(root, "replace-all", Buffer.alloc(32, 13)); + await store.set("old-a", { value: "a" }); + await store.set("old-b", { value: "b" }); + await store.replaceAll("new", { value: "current" }); + expect(await store.size()).toBe(1); + expect(await store.get("old-a")).toBeUndefined(); + expect(await store.get("old-b")).toBeUndefined(); + expect(await store.get("new")).toEqual({ value: "current" }); + }); + + test("rejects serialized-byte exhaustion and rolls back the failed write", async () => { + const root = await temporaryRoot(); + const store = new SecureJsonStore<{ value: string }>({ + directory: root, + name: "bounded-bytes", + key: Buffer.alloc(32, 11), + maxEntries: 8, + maxSerializedBytes: 1_024, + }); + + await store.set("small", { value: "retained" }); + await expect(store.set("large", { value: "x".repeat(2_000) })).rejects.toThrow("byte limit"); + expect(await store.get("small")).toEqual({ value: "retained" }); + expect(await store.get("large")).toBeUndefined(); + }); +}); + +function storeAt(root: string, name: string, key: Buffer): SecureJsonStore { + return new SecureJsonStore({ + directory: root, + name, + key, + maxEntries: 16, + maxSerializedBytes: 64 * 1024, + }); +} + +async function temporaryRoot(): Promise { + const root = await fs.mkdtemp(path.join(os.tmpdir(), "thoughtstream-oauth-store-")); + roots.push(root); + return root; +} diff --git a/test/telegram-bot.test.ts b/test/telegram-bot.test.ts index 9cdf804..d77bb34 100644 --- a/test/telegram-bot.test.ts +++ b/test/telegram-bot.test.ts @@ -17,6 +17,7 @@ import { import { startTelegramWebhookServer } from "../src/connectors/telegram-webhook.js"; import { JetstreamConnector } from "../src/connectors/jetstream.js"; import type { JazzThoughtStore } from "../src/jazz/store.js"; +import { describeRunResult } from "../src/projections/activity.js"; import { projectTrainingExamples } from "../src/training/judgments.js"; import { temporaryProject, testDeclarationEnvironment, testStore } from "./helpers.js"; @@ -96,7 +97,8 @@ describe("TelegramBotConnector", () => { const ingress = await connector.ingest(store, parseTelegramBotUpdate(fixture.updates[0])); const trigger = ingress.events[0]!; const declarations = await loadAgentDeclarations(path.join(process.cwd(), "agents"), testDeclarationEnvironment); - const declaration = declarations.find((candidate) => candidate.id === "telegram-conversation"); + const loadedDeclaration = declarations.find((candidate) => candidate.id === "telegram-conversation"); + const declaration = loadedDeclaration ? structuredClone(loadedDeclaration) : undefined; if (!declaration) throw new Error("Missing Telegram conversation declaration"); declaration.initialReplay = "beginning"; const runner: AgentRunner = { @@ -111,6 +113,8 @@ describe("TelegramBotConnector", () => { const runtime = new ThoughtAgentRuntime(store, [runner]); const [processed] = await runtime.consumeBacklog([declaration]); const run = await store.getRun(processed!.runId); + expect(run).toBeDefined(); + expect(describeRunResult(run!)).toBe("Produced: A reaction target"); const dispatcher = new TelegramChannelDispatcher({ id: "telegram-dispatcher:telegram:thoughtstream-bot:123456789", client, @@ -424,7 +428,8 @@ describe("TelegramChannelDispatcher", () => { const ingress = await connector.ingest(store, parseTelegramBotUpdate(fixture.updates[0])); const trigger = ingress.events[0]!; const declarations = await loadAgentDeclarations(path.join(process.cwd(), "agents"), testDeclarationEnvironment); - const declaration = declarations.find((candidate) => candidate.id === "telegram-conversation"); + const loadedDeclaration = declarations.find((candidate) => candidate.id === "telegram-conversation"); + const declaration = loadedDeclaration ? structuredClone(loadedDeclaration) : undefined; if (!declaration) throw new Error("Missing Telegram conversation declaration"); declaration.initialReplay = "beginning"; const runner: AgentRunner = { diff --git a/test/web-deployment.test.ts b/test/web-deployment.test.ts new file mode 100644 index 0000000..125d34c --- /dev/null +++ b/test/web-deployment.test.ts @@ -0,0 +1,68 @@ +import fs from "node:fs/promises"; +import path from "node:path"; +import { describe, expect, test } from "vitest"; + +const root = process.cwd(); + +describe("public web deployment contract", () => { + test("rate-limits OAuth routes, strips query strings from access logs, and canonicalizes www before proxying", async () => { + const nginx = await fs.readFile(path.join(root, "deploy/nginx/thought.stream.conf"), "utf8"); + expect(nginx).toContain("limit_req_zone $binary_remote_addr zone=thoughtstream_oauth_login"); + expect(nginx).toContain("limit_req_zone $binary_remote_addr zone=thoughtstream_oauth_callback"); + expect(nginx).toContain("limit_req_zone $binary_remote_addr zone=thoughtstream_review"); + const logFormat = nginx.split("\n").find((line) => line.startsWith("log_format thoughtstream_no_query")); + expect(logFormat).toContain("$request_method $uri $server_protocol"); + expect(logFormat).not.toContain("$request_uri"); + expect(logFormat).not.toContain('"$request"'); + expect(logFormat).not.toContain("$http_referer"); + expect(logFormat).not.toContain("$http_user_agent"); + expect(nginx.match(/server_name thought\.stream www\.thought\.stream;([\s\S]*?)\n\}/)?.[1]).toContain("thoughtstream_no_query"); + + const callback = nginx.match(/location = \/oauth\/callback \{([\s\S]*?)\n \}/)?.[1]; + expect(callback).toContain("access_log off;"); + expect(callback).toContain("error_log /dev/null crit;"); + expect(callback).toContain("limit_req zone=thoughtstream_oauth_callback"); + const login = nginx.match(/location = \/oauth\/login \{([\s\S]*?)\n \}/)?.[1]; + expect(login).toContain("limit_req zone=thoughtstream_oauth_login"); + const logout = nginx.match(/location = \/oauth\/logout \{([\s\S]*?)\n \}/)?.[1]; + expect(logout).toContain("limit_except GET HEAD POST"); + const catchAll = nginx.match(/location \/ \{([\s\S]*?)\n \}/)?.[1]; + expect(catchAll).toContain("limit_except GET HEAD"); + expect(catchAll).not.toContain("GET HEAD POST"); + const review = nginx.match(/location ~ \^\/inspector\/api\/reviews\/\[\^\/\]\+\/decisions\$ \{([\s\S]*?)\n \}/)?.[1]; + expect(review).toContain("limit_req zone=thoughtstream_review"); + expect(review).toContain("limit_except POST"); + expect(review).toContain("client_max_body_size 100k"); + + const wwwServer = nginx.match(/server \{[\s\S]*?listen 443 ssl http2;[\s\S]*?server_name www\.thought\.stream;([\s\S]*?)\n\}/)?.[1]; + expect(wwwServer).toContain("return 308 https://thought.stream$request_uri;"); + expect(wwwServer).toContain("error_log /dev/null crit;"); + expect(wwwServer).not.toContain("proxy_pass"); + expect(nginx).toContain("server_name thought.stream;\n"); + }); + + test("keeps one singleton proxy process, Basic fallback explicit, and one shared Review capability file", async () => { + const unit = await fs.readFile(path.join(root, "deploy/systemd/thoughtstream-inspector-proxy.service"), "utf8"); + const inspectorUnit = await fs.readFile(path.join(root, "deploy/systemd/thoughtstream-inspector.service"), "utf8"); + expect(unit.match(/^ExecStart=/gm)).toHaveLength(1); + expect(unit).toContain("EnvironmentFile=-%h/.config/thoughtstream/credentials/inspector-oauth.env"); + expect(unit).toContain("EnvironmentFile=-%h/.config/thoughtstream/credentials/inspector-review.env"); + expect(inspectorUnit).toContain("EnvironmentFile=-%h/.config/thoughtstream/credentials/inspector-review.env"); + expect(unit).toContain("ReadWritePaths=-%h/.local/share/thoughtstream-inspector-auth"); + + const basicConfig = await fs.readFile(path.join(root, "scripts/configure-inspector-credentials.sh"), "utf8"); + const oauthConfig = await fs.readFile(path.join(root, "scripts/configure-inspector-oauth.ts"), "utf8"); + expect(basicConfig).toContain("PROXY_BASIC_FALLBACK_ENABLED=1"); + expect(oauthConfig).toContain('"PROXY_BASIC_FALLBACK_ENABLED=1"'); + expect(oauthConfig).toContain("THOUGHTSTREAM_OAUTH_STORE_DIR is unsupported"); + expect(oauthConfig).toContain('path.join(os.homedir(), ".local", "share", "thoughtstream-inspector-auth")'); + const reviewConfig = await fs.readFile(path.join(root, "scripts/configure-inspector-review.ts"), "utf8"); + expect(reviewConfig).toContain("randomBytes(32)"); + expect(reviewConfig).toContain("inspector-review.env"); + expect(reviewConfig).not.toContain("console.log"); + + const threatModel = await fs.readFile(path.join(root, "spec/web-auth.md"), "utf8"); + expect(threatModel).toContain("single process"); + expect(threatModel).toContain("Basic fallback"); + }); +});