A local-first event pipeline for independent agents, built on Jazz.
thought-stream spec testing.md
34 kB

Testing and acceptance #

Unit tests #

  • Canonical JSON and payload hashing.
  • Registry validation and unknown-type quarantine.
  • Source idempotency keys.
  • Event lineage construction.
  • Filesystem root containment and ignore rules.
  • Stable document identity and rename heuristics.
  • Diff size bounds and newline normalization.
  • Consumer YAML validation and deterministic Jazz query compilation.
  • Context budgets and explicit truncation receipts.
  • Canonical output-contract identity, definition hashing, and shared validation across original, repair, and human correction paths.
  • Repair eligibility allowlist and hard exclusion matrix for timeout/cancel, provider/rate-limit, sandbox/broker/process, credential/config/auth, incomplete evidence, and repair-origin runs.
  • Deterministic one-request identity across repeated reconciliation.
  • Secret redaction.
  • Operational incident normalization from connector failure/recovery, scheduler exhaustion, failed/blocked/abandoned runs, and Telegram delivery failure; raw error/source/provider/tool/credential sentinels remain absent from incident events.
  • Private incident-ledger path containment, mode 0600, bounded canonical JSONL, fsync-before-acknowledgement, and restart deduplication by deterministic incident id.
  • Incident Telegram alerts ignore normal source allowlists, deduplicate by fingerprint and cooldown, obey a separate destination window, render classifications only, and never recurse on Telegram delivery failure.
  • Cursor advancement only after durability.
  • Public web routing is a closed allowlist: every named public page renders from reviewed public files, traversal and unknown routes return local 404, and no public request reaches the inspector upstream.
  • Inspector forwarding is available only below /inspector, strips that prefix, permits only GET/HEAD, removes authorization/cookie/forwarding headers, and requires either a valid server-side OAuth browser session or explicitly enabled Basic fallback.
  • OAuth metadata and JWKS are public and exact; metadata declares HTTPS client_id/callback, authorization_code plus refresh_token, atproto scope, DPoP-bound tokens, private_key_jwt, and ES256. No private JWK is present in either response.
  • OAuth tests model SDK protocol state and browser-bound application state as distinct values and exercise the installed official SDK's real authorize path to prove its generated protocol-state store key/PAR value differs from stored appState. They prove the query protocol state reaches SDK callback while the SDK-returned application state matches the cookie, plus protocol replay, missing/duplicate/mismatched/expired application state, exact DID allowlisting, opaque and duplicate cookies, browser expiry, restore failure, flow-store failure after SDK success, generation-conditional local deletion on CSRF-checked logout, and generic failures. Watchdog tests cover never-resolving callback, hung promotion, hung non-timeout cleanup, and late callback settlement followed by a successful newer generation. They prove the serializer advances, authority guards reject late persistent writes, and exact-generation cleanup preserves newer browser authority. Deferred restore-success and restore-failure races promote generation 2 while generation 1 is in flight, then prove scoped refresh set/delete cannot mutate generation 2 and post-restore recheck denies the old browser. Quarantine tests accumulate to the hard ceiling, verify fail-closed operator-recycle refusal, then prove recovery after settlement and fresh-process construction. Tests make no claim about provider-side revocation from the unmodified SDK; they never contact a PDS or load live credentials.
  • Encrypted OAuth stores use direct malformed, unsupported-version, wrong-key, and oversized-envelope fixtures; retain no plaintext sentinel; enforce owner-only modes; expire state records; reject entry-count and serialized-byte exhaustion atomically; and replace files atomically. Direct commit-point tests prove no fallible chmod or other filesystem operation runs after rename and force authority loss after temporary write to prove the pre-rename guard removes the temporary file without committing, preventing disk/memory rollback or late-authority divergence. A live single-process owner lock prevents a second factory from opening the same store directory and can be reacquired after graceful release.
  • Process tests prove OAuth login/callback rate-limit rejection occurs before SDK invocation, browser disconnect aborts the SDK-supported authorization path, custom OAuth store paths fail before service startup because they fall outside the systemd writable path, and missing Basic configuration defaults off without decoding dormant credential values. Deployment tests prove nginx route limits, no-query access logging, callback access-log suppression, canonical www redirect before proxying, singleton systemd shape, and an optional separate Basic credential compartment.
  • A production OAuth acceptance gate separately proves a real external metadata/JWKS fetch, PKCE/PAR/DPoP authorization redirect, allowlisted callback, private inspector read, logout/local generation deletion, and one bounded Basic rollback read. Basic retirement then requires removal of its credential compartment, a proxy-only restart, generic rejection of the old credential with no challenge, and continued OAuth admission.

Jazz integration tests #

  • Real persistent local Jazz database in a temporary directory.
  • Deterministic Jazz row ids for events, document versions, consumer executions, traces, cursors, progress, and projections.
  • Identical event replay, conflicting same-id content detection, query ordering, and restart persistence.
  • Database-side filters, multi-column ordering, limits, offsets, and declared indexes for hot paths; no TypeScript post-filter/sort fallback.
  • Query subscription snapshot/delta behavior, per-source progress recovery, new-source discovery, and duplicate wakeup tolerance.
  • Direct-batch visibility and transaction accept/reject behavior at local durability.
  • The same transaction probes against an in-process Jazz server at edge durability.
  • Event-plus-source-sequence-plus-cursor atomicity under failure before commit, after commit, and while waiting for durability.
  • Consumer-output-plus-zero-to-two-agent-proposals-plus-terminal-evidence-plus-progress atomicity under the same failure points, including deterministic replay after ambiguous settlement.
  • Concurrent same-account inference reservation admits only the calls allowed by policy within one process, while separate agent accounts remain isolated. Aggregate enforcement across independent processes is intentionally best-effort rather than a globally serializable billing guarantee.
  • A consumer processes durable backlog when subscription wakeups are unavailable, proving that bounded reconciliation rather than process-local callback delivery is the recovery authority.
  • Inference reservation/settlement survives database restart. Actual token telemetry adjusts the charge; a tracked but unavailable cost dimension retains its conservative estimate; an untracked cost dimension remains absent from estimate, charged, and persisted JSON. Expired leases keep their tracked charge through the active window.
  • Cost-bearing policies remain backward compatible. Token-only policies become fully reported from complete input/output telemetry, reject any cost limit without a reservation, and preserve cost omission through fixed and true sliding rolling windows. The resident declaration keeps its 60,000 input / 2,000 output reservation estimate while exposing only very high emergency circuit breakers; tests treat its accounting as telemetry plus catastrophic-runaway protection, never as a budget or thrift mechanism.
  • Rolling limits are true sliding windows rather than first-call buckets.
  • Budget denial occurs before runner dispatch, writes one terminal blocked lifecycle result, emits no repair request, and persists no source body, prompt, model output, tool argument, or provider payload in accounting rows. onExhaustion: advance advances progress exactly once; defer leaves progress unchanged, records a limiting-window retry time, and creates no duplicate attempt before that time. Generic retry backoff is tested independently through the declaration-level retry policy.
  • Document version insertion and current projection update.
  • Independent producer processes can write disjoint source namespaces concurrently and synchronize through Jazz.
  • One producer restarted after an ambiguous failure replays without duplicate historical rows or skipped source sequence.
  • One consumer restarted after a nonterminal attempt preserves the attempt, replays its query safely, and reaches one accepted terminal transaction.
  • Permission tests prove historical insert-only behavior and sensitive-row isolation for backend, authorized, unauthorized, and anonymous sessions.
  • Mutation rejections are surfaced both through WriteHandle.wait and restart-safe mutation error reporting.
  • Projection rebuild from an empty projection table.
  • Repair request recovery after a crash between original terminal settlement and request append.
  • Interrupted repair recovery terminates without a second proposal generation.
  • Effective-output rebuild follows active accepted/corrected proposal, original valid output, unresolved failure, including supersession and retraction.

These are capability gates, not aspirational checks. An API named transaction, upsert, or permissions does not count as proof that the required settlement or authorization property holds.

Runner tests #

  • Pi-compatible deterministic provider streams text and structured output.
  • conversation-text is accepted only for a standard Pi Telegram conversation declaration, omits the provider JSON response-format request, wraps one bounded text part in deterministic observation metadata, and retains no rejected oversized text. A declaration may additionally expose only the fixed proposal tools in proposals.md; validated tool-only output receives one deterministic non-authoritative acknowledgment, while text-plus-tool preserves exact visible text. compaction-text is accepted only for the no-tool Telegram compactor, omits provider JSON formatting, and deterministically wraps one bounded plain-text summary in the canonical compaction contract. Every other declaration remains on strict JSON validation.
  • Trace events preserve order and execution identity.
  • Provider error, abort, malformed JSON, disallowed output, and timeout all produce correct terminal evidence.
  • Eligible invalid JSON and contract failures produce one sanitized request; every infrastructure/authorization/corrupt-evidence class produces none.
  • A repair declaration uses the mandatory fixture broker plus disposable sandbox path and emits one inert correction proposal.
  • A repair output validation failure is terminal, advances request progress, emits no proposal, and creates no repair-of-repair request.
  • Privacy tests seed a unique malformed-output sentinel in process-local provider output and prove it is absent from events, runs, traces, projections, Telegram candidates/receipts, and training export.
  • Training export includes accepted/corrected repairs only and emits only allowlisted trajectory/provenance fields.
  • Tinker provider configuration is tested without a real credential by inspecting the built model descriptor and request shape.
  • Proposal-tool tests cover native OpenAI tool serialization and parsing, tool_choice: auto, one provider request, strict TypeBox arguments, at most one call per fixed tool and two total, tool-only acknowledgment, text-plus-tool preservation, unknown/malformed/oversized/capability-escaping calls, tool-argument trace exclusion, and exact snapshot memory/evidence/delivered-output binding.
  • Proposal-core tests cover sensitive agent-proposed events with no publication/quality/export authority, atomic settlement and replay, correction non-effectiveness before a human decision, accept/edit/reject projection, stale/symlink/frontmatter/hash/version memory materialization, and content-dark failure receipts.
  • Private-training tests require an explicit destination and sensitive acknowledgment, never use stdout, include only active quality-eligible human judgments with exact private provenance, exclude unaccepted proposals, and prove the external exporter still excludes externally ineligible self-corrections.
  • Adapter manifests reject malformed ids/versions/digests, unknown fields, duplicate release identities, duplicate capabilities/evals, unknown provider profiles, missing or non-allowlisted public base/checkpoint resolution, ambiguous declaration selection, and adapter/model/tier combinations. Candidate/retired deployment entries remain unbound, and enabled declarations that name them fail compilation.
  • A trusted environment checkpoint resolves only during startup compilation and provider dispatch; declaration JSON, fingerprints, stored consumers, runs, contexts, outputs, events, traces, inspector responses, training examples, test failures, and generated manifests contain no private checkpoint value. The private binding lives in a module-local WeakMap. Returned release/binding objects are deeply frozen; mutation fails, and clones, rehydration, and manual forgery cannot recover the checkpoint.
  • Startup matrix tests cover missing/empty release catalogs, missing deployment catalog, malformed and duplicate releases, candidate/retired selection, missing or digest-mismatched release, duplicate declaration selection, unresolved/non-allowlisted checkpoint, expected catalog mismatch, and unauthorized process identity.
  • Release digest remains stable across deployment generations. Catalog digest changes with deployment selection, verifies its optional expected digest without self-reference, and appears with generation on declaration, run, output/failure evidence, inspector inventory, and training provenance. Declarations without learned adapters remain compatible.
  • No dynamic lifecycle machinery remains: no Jazz adapter registry, lifecycle events, authority file, recovery marker, owner election, registry lock/tombstone/reaper, dispatch lease, or operator recovery API. Activation/retirement documentation requires coordinated stop/install/restart plus PID/start-time, loaded-digest, and canary receipts.
  • Sensitive-adapter tests cover run evidence, success/failure output, repair request/proposal, judgment/retraction, delivery, effective-output projection, and training. Each path uses the shared privacy join.
  • Pairwise training tests gate privacy and export class on both primary and compared runs, reject forbidden pairs, require explicit restricted/private authority, and assert exact compared execution/model-adapter provenance in v3 examples and dataset manifests.
  • Review tests freeze exact same-trigger candidate runs; underdetermined, tie, malformed, and skip decisions never become preference examples; correction requires a contract-valid replacement; supersession leaves one active export label; and public prompt authority gates v4 input text. Private browser decisions cannot declassify data.
  • Review web tests prove Basic remains read-only; OAuth POST requires CSRF and a fresh one-time proxy signature; replay, stale signatures, unknown routes, oversized bodies, malformed decisions, and direct loopback POSTs fail without appending events.
  • Course web tests prove the repository course has eight ordered lessons and one typed workshop each; question ids are idempotent, stale revisions and divergent reuse fail, Basic remains read-only, OAuth forwarding requires CSRF and a separate fresh body-bound capability, signature replay fails, and the status API reveals an answer only after exact completed tutor lineage.
  • Workbench workflow tests (test/workbench-workflow.test.ts) run against temporary Jazz stores and prove: a working document created from an event keeps its origin reference across store reopen; operator edits require the exact head, are idempotent per request id, and fail closed with a stale-base error; context selection records exact event ids, payload hashes, version ids, and hashes in one immutable bounded plain-text snapshot; the fixture runner is deterministic and its diff is computed against the exact base; accepting once yields exactly one new version, one decision, and one non-exportable judgment while a duplicate submission creates nothing extra and a conflicting submission fails; reject leaves the document unchanged; an operator edit followed by accepting the older proposal fails with the stale error and the edit survives; a runner failure records a failed run and no proposal; and a second store client over the same project root sees the accepted head.
  • Workbench inspector tests (test/workbench-inspector.test.ts) prove every workbench write returns 405 without a Review verifier, unsigned and wrong-path signatures fail with 403 and no store effect, malformed bodies fail with 400, signed writes apply through the trusted workflow, the GET projections expose head, origin, exact selections, proposals with diff and stale flags, and versions, stale edits and stale accepts return 409 with the current head, decisions replay by submission id, and a consumed signature nonce cannot be replayed.
  • Proxy tests (test/authenticated-proxy.test.ts) prove every workbench write route rejects unauthenticated, Basic-only, missing-CSRF, wrong-CSRF, query-bearing, oversized, and non-JSON requests before any upstream contact, forwards one CSRF-checked OAuth write per route with a fresh body-bound Review signature and no browser credentials, and still serves workbench reads to Basic.
  • Inspector HTML tests compile the rendered inline JavaScript, contain the Learn destination and floating composer, and preserve fragment routes for exact lesson ids. Browser acceptance exercises every workshop, narrow and wide layouts, keyboard submission, disabled composer state, and one natural OAuth question through event, tutor run, output, and status readback.
  • Replay uses the exact public run binding but recompiles the private checkpoint through the trusted startup loader; it does not float to another release or trust serialized checkpoint authority.
  • Live Tinker sampling is an opt-in credentialed test and never runs in ordinary CI.
  • Letta Agent SDK declarations resolve the agent id from the named environment variable, reject non-Cloud backends and credential/base-URL selection, require single-event context and one or more bounded concrete sources, reject wildcard source patterns, and reject shared enabled agent ids.
  • Two ready source namespaces under one enabled Letta declaration execute through one shared agent scheduler key: the test runner observes maximum concurrency one while both per-source progress rows settle independently.
  • A public trigger processed by a resident declaration that also accepts sensitive events produces sensitive output and lifecycle events; current-trigger privacy may not downgrade persistent conversation state.
  • Resident ATProto batching is justified by semantic coherence and reduced redundant turns rather than cost savings; ATProto-object context keeps the existing Bluesky behavior: post URIs and like subjects use fixed public atproto.md and bsky.md endpoints in the trusted parent, with independently bounded protocol/social Markdown, observed-CID mismatch rejection, honest current-record-unverified labels, and per-service failure isolation.
  • One synthetic network.cosmik.collectionLink create produces one packet containing the exact original link record plus independently bounded atproto.md views for the link, referenced card, and referenced collection. Tests prove strong URI/CID preservation, one-view failure isolation, create-only ingress with no delete turn, defensive delete dereference skipping, immutable snapshot reuse, character bounds, and zero bsky.md calls. The tracked source filter excludes card, collection, and note-child ingestion so a save does not produce duplicate turns.
  • One immutable content-hashed Jazz document-version snapshot is written before the Cloud turn, survives projection rebuild semantics, and is reused across retries; model-visible compiled hashes, source hashes, and snapshot integrity are tested without persisting fetched bodies to Git or traces.
  • Conversation-text final instructions distinguish Telegram replies from private ATProto observations without repeating runtime character-count or JSON-format enforcement.
  • The Cloud adapter resumes the configured agent's main conversation in an SDK-managed sandbox, sends only the current event packet, closes the session, and maps terminal output through the canonical contract.
  • The Stream traces retain event type, counts, hashes, tool names, run ids, conversation id, duration, and allowlisted terminal metadata while excluding reasoning, assistant text, tool arguments/results, prompts, source bodies, SDK error detail, and credentials. Run-usage recovery reports complete token dimensions and does not manufacture zero or estimated OAuth dollar cost.
  • Repeated execution for one source event finds the deterministic turn marker in conversation history and recovers the existing assistant result without a second send(). A marker without assistant completion leaves progress unchanged. A delayed second history check runs before any attempt greater than one may send.
  • A Jetstream-rooted failed resident run can produce one content-dark operational alert while a successful run from the same source remains ineligible for normal Telegram delivery.
  • History pagination fails closed: a bounded window with older pages remaining is inconclusive, and hasMore without a usable cursor is a protocol failure. Neither case may send a new turn.
  • Timeout, Cloud sandbox expiry, protocol failure, terminal SDK failure, stream-without-result, invalid JSON, oversized conversation text, and history reconciliation ambiguity produce classified failures with the correct progress policy.
  • Cloud credit exhaustion is classified from process-local provider detail as letta-cloud-insufficient-credits, leaves progress unchanged, settles the inference reservation conservatively, and persists neither the provider detail nor source text.
  • A live Cloud canary is opt-in and credentialed. It must use a dedicated hidden agent, a disposable conversation or explicit canary agent main conversation, no private source payload, and no thought stream production activation. It records backend, package version, agent/conversation/run ids, terminal result, and sandbox lifecycle without printing credentials or raw internal reasoning.
  • Coil Public Knowledge tests scan a synthetic vault into a temporary Jazz database, preserve one documentId and conversation across renames and declaration-version upgrades, skip blocked/deferred/deleted paths before SDK session creation and accounting, bind one conversation per eligible document, recover a summary-marked remote conversation after a simulated pre-Jazz crash, reject duplicate remote markers, and verify unchanged scans make no model call.
  • The recommendation contract tests reject unknown revision targets, colliding new slugs, malformed candidate outlines, privacy-clear claims with blocked links, and source/catalog/hash mismatches. Accepted output settles one private recommendation, run lifecycle, conversation binding, accounting record, and source progress.
  • Local-memory Agent SDK tests verify backend=local, API-backed App Server selection, filesystemConfinement=memory, exact memory-root environment, empty skills/tools, strict permission mode, conversation-scoped model application, and absence of Coil/site paths or source text from durable traces. A live local canary is opt-in, uses Co's existing agent and one explicit synthetic or approved file conversation, and cannot activate the full Coil backlog.

Container-harness security gate #

The workspace-v1 profile is not activatable from a consumer declaration until an actual container runtime passes black-box tests. Command construction snapshots and mocked process exits are useful unit tests but are not containment proof.

  • The container sees only /workspace, /state, bounded tmpfs, its image filesystem, its own process namespace, and /broker/provider.sock. Canary reads of host home, parent repository paths, /proc/1/environ, credentials, container-runtime sockets, devices, and undeclared mounts fail.
  • Root filesystem writes, privilege escalation, Linux capabilities, setuid behavior, additional mounts, and container-runtime access fail. The process runs as the configured unprivileged uid/gid with no-new-privileges and the expected seccomp profile.
  • TCP, UDP, DNS, loopback services, link-local metadata addresses, and public Internet access fail. The provider Unix socket succeeds.
  • A workspace mutation persists only in the workspace lease. A state mutation persists only in the state lease. Fresh-container scratch files do not survive a second run.
  • CPU, memory, process-count, descriptor, output, wall-clock, total lease-byte, and inode limits terminate hostile fixtures with classified failures and no trusted-host fallback. Per-file limits alone do not satisfy the disk-exhaustion gate.
  • The launcher rejects missing runtimes, mutable/unresolved image references, symlink or out-of-root leases, wrong file types, duplicate mounts, oversized packets/results, trailing frames, run-id mismatch, nonzero exits, and timeout.
  • Broker tests cover capability forgery, run-id mismatch, model and route substitution, unauthorized headers, redirect rejection, response overflow, request-count exhaustion, cumulative request/response-byte exhaustion, expiry, replay, concurrent use, and opaque learned-adapter authority. Concurrent socket tests with maxRequests: 1 and one-request byte budgets prove serialized immediate-pre-egress reservation admits exactly one upstream request. Adapter tests prove no authority is acquired during prefetch, clock/state changes are observed at immediate pre-egress validation, upstream execution occurs inside the bound, the bound covers response consumption, live operations survive old lock age/process suspension, and retirement cannot cross the held provider request.
  • Pi coding-agent runs with extension/skill/template discovery disabled, only the profile's built-in tool allowlist, in-memory credentials/settings, and the broker placeholder key. A malicious .pi/extensions fixture is not executed.
  • A deterministic multi-turn provider fixture makes Pi call a workspace tool, observes the tool receipt, returns a final artifact, and proves that every provider turn crossed the broker. A second run resumes the selected session from /state; adapter/profile/workspace/model mismatch fails closed.
  • Sentinels placed in provider credentials, host environment, host-only files, tool arguments, shell output, provider bodies, and model reasoning are absent from durable run/trace/accounting/notification surfaces.

Passing this gate proves the named image digest, adapter revision, container runtime, and profile together. It does not bless arbitrary images, Pi extensions, runtime flags, or host platforms.

Connector contract tests #

Each connector has captured fixture responses and tests reconnect/cursor behavior without network access. Live tests are opt-in.

Jetstream fixtures include create, update, delete, non-commit, and filtered-collection messages. Tests preserve AT URIs/CIDs/records, retain deletes without a record, bind cursors to the filter revision, reject malformed batches without advancing the cursor, and absorb replay through the durable time_us cursor plus event idempotency.

The Jetstream CLI has a producer-only test that runs without an agent directory or model environment, persists source events and cursor evidence, reports zero consumer runs, and never loads or executes a declaration.

Live Jetstream transport tests use an injected in-process WebSocket fixture, never an official endpoint. They verify URL filters and rewind cursor construction, serial durable processing, overlap replay without duplicates or equal-timestamp loss, reconnect/backoff, bounded shutdown, and visibility of inserted events to a Jazz consumer subscription. Any production-network smoke test is opt-in and must use a temporary database plus narrow filters and hard runtime/message limits.

Telegram spool fixtures use normalized synthetic private messages, exact route identifiers, and opaque attachment references. Tests verify strict schema rejection, sensitive event classification, bounded restart-safe ingestion, incomplete trailing-line behavior, cursor advancement after durable records, and fail-closed detection when consumed spool content changes. They do not use a bot token, Telegram API, existing channel spool, or MessageChannel.

Telegram webhook and dispatcher tests use an in-process loopback receiver and synthetic Bot API updates. They verify exact path and method handling, constant-time secret-header admission, JSON/body bounds, allowlisted chat and reaction filtering, serial durable processing, deterministic replay, out-of-order lower update admission, migration from the former polling cursor revision, 2xx only after durable settlement, 5xx on store/projection failure, explicit registration with max_connections=1, and separation from dispatcher send authority. Dispatcher tests prove that only a recent running direct-reply run rooted in the exact allowlisted Telegram chat emits and refreshes one content-free typing action, terminal/stale/wrong-route runs do not, and chat-action failure does not block durable reply delivery. They never load a live token or webhook secret and never contact Telegram.

X webhook tests use an in-process loopback receiver, synthetic consumer secrets, and captured invented Activity envelopes. They verify exact CRC HMAC output, constant-time x-twitter-webhooks-signature admission over raw bytes, path-selected privacy lanes, event/user/tag allowlists, strict V1 post schemas, mutable metric/profile exclusion, body and deadline bounds, serial durable processing, duplicate and divergent event-id replay, out-of-order delivery, 200 only after durable settlement, and 5xx on store failure. Operator tests separate read-only status/plan from mutating register/apply/replay/delete commands and prove no command runs on receiver startup. No test loads a live X credential, user OAuth grant, account payload, or network endpoint. See x-webhook.md.

Agent-message tests use temporary Jazz databases and deterministic runners. They prove the exact Co → Stream route, strict source/actor/payload identity, idempotent exact replay and divergent-reuse rejection, bounded same-thread native-role reconstruction, exclusion of foreign/thread-crossing/failed outputs, retry-stable snapshots, typed response lineage, no Telegram candidate, and double opt-in private response display. They never forge Telegram updates, edit live context documents, or contact an external model.

Credential-compartment tests use generated synthetic values only. They prove Telegram ingress, X ingress, X management, consumer, dispatcher, and Jetstream outputs contain exactly their allowlisted variable names, no value is printed, malformed/duplicate assignments fail closed, Git/public destinations are refused, and generated directory/file modes are 0700/0600. Live credential files are never read by the test or implementation gate.

Fastmail tests use invented JMAP account, email, mailbox, address, preview, and attachment values. The captured-response tests verify Email/queryChanges + Email/changes + Email/get parsing, create/update/destroy observations, sensitive classification, full-body exclusion, retained state cursors, exact replay absorption, and fail-closed state-gap handling. Live-transport tests use a bounded in-process fetch double to prove same-origin HTTPS session validation, read-method allowlisting, first-start replay: now, serial pagination, response/page/id bounds, metadata-only fetches, replay-safe event identity, cursor-safe failure, and bounded cannotCalculateChanges resnapshot. Tests explicitly remove ambient Fastmail credentials and never contact Fastmail, read real mail, or mutate a mailbox.

First end-to-end acceptance #

  1. Start with an empty temporary Jazz database and fixture vault.
  2. Initial scan emits one add event and document version per accepted file.
  3. The deterministic document-structure consumer subscription receives Markdown adds and emits a typed structure event.
  4. Modify a linked heading; watcher emits one changed event with exact prior/current version ids and bounded diff.
  5. Agent emits a Charter re-evaluation recommendation linked to the changed event.
  6. Rename another document; identity remains stable when the heuristic is confident.
  7. Stop after a consumer attempt starts, restart that consumer, preserve the nonterminal attempt, and reach one valid terminal outcome without duplicate accepted outputs.
  8. Rebuild projections and compare hashes.
  9. Query the root activity endpoint and verify the full causal chain is visible.
  10. Verify producer and consumer processes communicate only through Jazz rows, queries, and subscriptions, with no central matcher queue or in-memory handoff required for recovery.

No milestone claim is valid without this receipt chain.