Recovery and evidence #
Web chat turn ownership #
One singleton proxy owns web turns independently of browser connections. Before any send, it fsyncs an admission containing the web conversation, random client request id/OTID, keyed payload digest, and dispatching state. Concurrent same-id/same-payload requests return the same receipt; differing payloads conflict. Receipts are never expired into resend permission. Bounded registry capacity fails closed rather than deleting idempotency evidence. SDK OTID is correlation, not exactly-once inference. Before a browser send, sessionStorage retains only its opaque web/request ids (never message text/context). Refresh reconciles that id against the server receipt; an absent receipt does not authorize a fresh-id resend. Re-entering the exact payload reuses the same request id. Missing browser storage fails admission before dispatch.
The controller consumes SDK events after the HTTP send response, maintains a bounded transient projection, and persists only status plus runtime run/session ids. Only a successful authoritative SDK result yields completed; an explicit interruption yields interrupted; stream closure, transport failure, restart, or missing terminal evidence yields unknown, not success or automatic retry. Uncertain conversation creation is likewise never replayed with the same request id. Restart keeps uncertain web sends blocked. Reconnecting owners can reconcile exact registered-conversation history and fresh device status; missing history or an idle device alone never proves completion. An unknown receipt remains visible and blocks another send until operator resolution; the owner may start a separate new web conversation. No synthetic inference is used for health checks.
Stop calls SDK abort(), persists cancellation-requested progress, and awaits terminal evidence; HTTP cancellation and SDK close() do not mean aborted. For recovered uncertain turns, Stop resumes only the registered web conversation and requests abort, without inventing a terminal outcome. Browser disconnect never aborts a server-owned turn. Initialization, history, send admission acknowledgement, and Stop have separate 60-second watchdogs; the SDK turn timeout is 30 minutes so a browser request deadline does not truncate ordinary long-running Co work. A watchdog is uncertainty, never proof that remote execution stopped. Browser reconnect fetches current history and server status rather than relying on SDK missed-event replay. Shutdown stops new admissions and closes transports; unfinished receipts remain unknown after restart.
Crash model #
Any producer or consumer process may stop between writes, during a source batch, during model streaming, or after the model responds but before outputs are durable. Recovery is built around single-owner process identities, source-local sequences, durable per-source consumer progress, deterministic row identity, and lifecycle evidence.
Connector recovery #
- Read the last durable source cursor.
- Resume or repoll from that position.
- Build deterministic event identities, assign source-local sequences, and validate candidates.
- In one Jazz transaction, write all new events, source sequence state, the external cursor, and cursor receipt.
- Wait for the transaction at the configured durability tier before acknowledging source progress.
Producer and consumer transactions keep canonical source/event keys separate from Jazz storage-object ids. Storage ids use a revisioned namespace, and ordinary reads prime rows by canonical key before transactions. This allows recovery from an older failed transaction that allocated an object id without ever linking a table row; canonical event identity and idempotency do not change.
At-least-once observation is expected. Duplicate canonical events are not. A replayed item addresses the same deterministic Jazz row; an already accepted identical row is unchanged. Each source id has one writer process, so ordinary recovery does not require a cross-worker claim.
For Jetstream reconnects, the transport requests a bounded interval before the last durable time_us. The connector offers messages from that overlap to Jazz and lets source idempotency absorb duplicates; it must not discard the overlap merely because a message is older than the stored high-water mark. Cursor writes remain monotone and reflect the greatest durably observed source timestamp, including ignored account or identity messages.
A live subscription processes messages serially. If local backpressure crosses its configured high-water mark, it closes the socket and reconnects from the durable cursor rather than dropping source messages and pretending the stream remained continuous.
For Telegram webhooks, the HTTP response is the source acknowledgement. The receiver returns success only after durable settlement. Telegram retries non-2xx responses, and a replay addresses the same deterministic update/message rows. The receiver serializes admitted deliveries and webhook registration requests one upstream connection, but correctness does not depend on arrival order: highestUpdateId is diagnostic only and cannot suppress a lower unseen update. A crash after durable settlement but before the HTTP response therefore produces an unchanged replay rather than duplicate canonical events.
An X webhook has no upstream cursor. Its local source sequence records durable receipt order, while X event_uuid and event idempotency absorb retries and explicit replay. The receiver admits earlier timestamps and replayed events without regressing state. X replay covers at most the preceding 24 hours and remains an explicit operator command over one webhook id and UTC interval; startup never infers or initiates replay. A read-only health process may report invalid webhook state or subscription drift but cannot repair it. Beyond 24 hours, surviving public posts may be recoverable through a separate REST reconciliation, while deletion completeness is unknowable. See x-webhook.md.
Fastmail recovery keeps the email-state and query-state cursors in one source cursor. First activation captures both current states without replay. A normal cycle reads bounded Email/changes and Email/queryChanges pages from those exact states, fetches metadata, and settles observations plus both final states together. A crash before settlement repeats the same changes; deterministic event identities absorb the replay. cannotCalculateChanges is the only automatic discontinuity repair: one bounded current-mailbox resnapshot emits metadata-only updated observations and advances both states atomically. Any other method error or exceeded page/byte bound leaves both prior states unchanged.
Deterministic batch recovery #
A batcher's durable per-source progress is a filtered-consumer high-water mark and is the authority after restart. It may cross raw nonmatching cursor/lifecycle rows, but settlement re-queries authoritative Jazz rows and requires every eligible matching event between the prior mark and selected terminal member exactly once in order. Each enabled declaration also holds a credential-dark, exclusive local runtime lock keyed by id/version and bound to boot id, PID, and kernel process-start identity; clean shutdown releases it and dead/rebooted/PID-reused owners are reclaimed. Synchronized or multi-host active-active batching remains unsupported until a distributed lease exists. Quiet-window and maximum-age eligibility are recomputed from canonical eligible matching-event timestamps; invalid future timestamps are clamped to the current cycle time, while wall-clock rollback remains a bounded delay risk rather than a monotonic-time guarantee; an in-memory timer is only a polling wakeup. Batch identity includes the declaration fingerprint and every ordered member event id. Batch append and progress settle in one Jazz transaction. A crash before settlement can replay the same deterministic candidate, whose idempotency absorbs duplication; progress may never advance beyond a member absent from that durable payload.
Live cutover uses replay: now: first stop resident consumption of raw ATProto, deploy the disabled batch declaration and resident declaration, then activate the batcher while the raw producer remains healthy. On its first cycle, before querying eligible work, the batcher atomically creates absent progress at the captured canonical raw source head; producer events appended after that captured head have larger sequences and remain pending. Only after verifying that head-aligned initialization should the batch-consuming resident be enabled. This avoids replay while retaining all commits that arrive after initialization. The previously budget-blocked source records remain canonical, but this cutover does not claim to recover or dispatch those four skipped resident observations.
Budget denial follows the declaration's explicit accounting.onExhaustion policy. advance atomically settles a terminal blocked run and consumer progress, preserving the original skip-on-denial behavior. defer atomically settles a terminal blocked attempt without progress, records the budget-window-derived retryAt, and creates a later deterministic attempt only after that time. The blocked attempt is terminal evidence for one provider-free attempt; the source event remains pending consumer work. Batching remains a semantic-coherence mechanism rather than an implicit cost-control queue.
Model-adapter startup recovery #
- A persisted declaration or run proves what previously executed; it cannot recreate the process-local checkpoint binding.
- Restart reloads the complete immutable release catalog and one deployment catalog, verifies optional expected catalog digest/process identity, resolves checkpoints, and recompiles declarations before provider egress.
- Candidate and retired entries create no runtime binding; any enabled declaration still naming one fails compilation. Missing, ambiguous, digest-mismatched, unresolved active, unallowlisted, or unauthorized selections fail startup. Recovery never floats to a newer release by capability similarity.
- Adapter activation and retirement use stop/install/restart. Operators verify that no old PID/process-start identity remains, every restarted process reports the expected loaded catalog digest, and a bounded conceptualizer canary passes.
- Mixed catalog generations are unsupported. If a process did not restart or reports the wrong digest, the deployment is incomplete and adapter-backed work stays disabled.
- There is no authority file, recovery marker, registry lock, owner election, dispatch lease, or operator mutation API to reconstruct after a crash.
Consumer recovery #
- Read the consumer declaration and per-source progress. A persisted run contains exact public-safe learned-adapter and catalog provenance, but a persisted or rehydrated declaration does not recreate the process-local checkpoint binding. Restart must recompile through the trusted startup loader; replay never silently floats to another release or binding.
- Query matching event rows after each source's last terminally handled sequence.
- Treat live subscription deltas as wakeups only. Re-query each installed source at the manifest's bounded reconciliation interval so writes committed by another local process cannot remain invisible merely because the persistent driver's cross-process callback did not fire.
- Inspect deterministic execution/lifecycle evidence before invoking external work.
- A nonterminal started attempt from the previous process is recorded as status-unknown or abandoned according to policy.
- Lifecycle settlement primes durable run, event, and progress rows before opening an edge transaction. Domain event identity remains stable, while consumer settlement uses revisioned Jazz storage-object ids so an object allocated but never linked by a failed older transaction cannot permanently poison terminal evidence.
- A retry creates a new attempt and preserves earlier traces and lifecycle evidence. Generic retryable runner failures use only the declaration's explicit
retry.initialDelayMs/retry.maxDelayMsexponential policy;accounting.onExhaustiongoverns budget denial only. An interrupted repair attempt is terminally abandoned and its request progress advances without a second proposal generation. - Accepted outputs, terminal evidence, and progress for terminally handled source sequences settle in one Jazz transaction. A deferred budget-blocked attempt deliberately settles run/lifecycle evidence without progress, leaving the source event pending.
- Startup and post-failure reconciliation deterministically append any missing eligible repair request. A crash after the original failure settles but before request append therefore heals to the same request row.
- One active process owns each consumer id/version. Redundant workers are a later explicit topology, not an implicit lease requirement.
- A model-backed attempt persists its run owner before reserving inference. The reservation is therefore recoverable even if the process dies before provider dispatch or terminal settlement.
- Reservation leases do not authorize automatic duplicate model calls. An expired unsettled lease becomes
expiredand keeps its conservative charge until its configured budget window elapses. Restart reconciliation terminally classifies the interrupted run and settles any still-reserved charge without inventing actual usage. - A denied reservation creates one terminal
blockedattempt without provider dispatch, repair generation, or dispatcher-visible failure noise. Underadvance, progress settles in the same transaction. Underdefer, progress remains unchanged and the failure diagnostic records a retry time derived from every limiting budget window. - A stateful Letta Cloud turn carries a deterministic marker derived from consumer id/version and source event id. Before sending, the adapter reads bounded main-conversation history. A marker followed by an assistant result is recoverable output; a marker without a result is still in-flight or ambiguous and may not be sent again blindly.
- The SDK currently does not expose a caller-owned
clientMessageIdonsend(). The marker/history protocol narrows the ambiguity window but is not a provider-native idempotency receipt. A second attempt performs a delayed second history check before any resend. If the marker remains present without an assistant result, source progress stays unchanged. If an earlier process died between transport send and durable marker visibility, exact recovery remains bounded by Cloud conversation-history consistency and must be named as such. - A multi-source resident keeps independent progress per source but serializes every source operation through one Letta-agent scheduler key. Restart reconciliation may enqueue ready sources in a different cross-source order than their wall-clock occurrence; it may not run two turns concurrently or advance either source before that turn's output and terminal evidence settle.
- Public ATProto object context is acquired before the resident turn and durably snapshotted as an immutable Jazz document version before any prompt can be sent. The snapshot key includes declaration fingerprint, source event id, and every source-specific target URI/CID: one Bluesky post target, or the Semble collection link plus referenced card and collection. Its stored text and canonical context manifest are integrity-checked on reuse and is not deleted by projection rebuild. A process death before snapshot persistence may refetch because no model turn exists. After persistence, every attempt and history-reconciliation recovery uses the exact same packet rather than mutable current network state.
- A subscribed-document Telegram context is likewise snapshotted before provider dispatch, but its deterministic identity uses declaration fingerprint plus trigger event id. The first attempt resolves each declared current projection to its exact immutable document version and records those ids and hashes beside the delivered-reply provenance. A retry loads the packet from the snapshot without consulting newer current projections. Missing required documents, inconsistent version evidence, non-text content, or a document set that exceeds its declared sub-budget fail before inference and leave source progress unchanged.
Dispatcher recovery #
- The first run for a channel writes one durable dispatcher activation event. This establishes the lower bound for notification candidates without flooding a new destination with historical activity.
- Completed consumer runs after activation are the durable accumulator. The dispatcher can stop without blocking producers or consumers; unclaimed candidates remain available when it returns.
- Each rendered batch gets a deterministic identity over dispatcher, destination, and included run ids.
- The dispatcher appends a
startedclaim before external delivery. Only the process that inserts that claim may call the Bot API. deliveredrecords the Telegram message id.failedpreserves the transport error. A crash afterstartedis status-unknown and is not blindly retried because Telegram does not provide an idempotency key forsendMessage; replay must be explicit to avoid duplicate human-visible sends.- Velocity accounting reads recent durable
deliveredreceipts, so restarting a dispatcher does not reset the channel's rate limit. - Typing presence is rebuilt from recent durable
runningruns on every dispatcher cycle and kept only in process-local refresh state. Restart may repeatsendChatActionbecause it is transient UI state, not a human-visible message. Stale runs beyond the bounded presence age are ignored; Bot API failure is suppressed until the next refresh interval and never changes run or delivery evidence.
Incident recovery #
Operational incident projection is replayable from immutable source/lifecycle/action evidence. Repeated projection addresses the same deterministic incident row. The private JSONL writer appends and fsyncs before remembering an incident id; restart reloads ids from the ledger. An interrupted append may be offered again, but the same incident id cannot become a different incident.
An incident Telegram dispatcher groups open incidents by normalized fingerprint. Its deterministic started claim remains the delivery ambiguity boundary: after a claim, neither a process restart nor an alert-delivery failure authorizes another Bot API call. Delivery failures remain ledger evidence and are excluded from recursive Telegram alerting.
Terminal evidence #
A run attempt is terminal only when one Jazz transaction has established its terminal run row, corresponding terminal event, every referenced accepted output, and configured durability. A source event is terminally handled only when that transaction also advances the relevant per-source consumer progress. Deferred budget-blocked attempts satisfy the first contract and deliberately not the second; the UI and recovery loop must keep those states distinct.
The UI flags contradictions rather than choosing whichever row looks friendlier.
Retry classes #
- Transient transport/rate limit: when the declaration has an explicit
retrypolicy and failure leaves progress unchanged, retry with bounded exponential backoff frominitialDelayMsthroughmaxDelayMs. This policy is independent of budget exhaustion. - Timeout/process loss: preserve the nonterminal attempt, classify it, and retry according to consumer policy.
- Invalid output: fail without blind retry; the deterministic repair coordinator appends exactly one versioned request only when
repairs.mdeligibility and evidence checks pass. - Unknown, missing, candidate, retired, ambiguous-capability, provider-mismatched, non-allowlisted, cloned, rehydrated, mutated-after-compilation, or unbound model adapter: reject declaration registration, run creation, and provider dispatch. Complete canonical compiler-bound identity is checked at every boundary. After read-only prefetch and local broker validation, invoke the opaque run-bound authority capability at the actual provider boundary. It revalidates the exact canonical active binding and holds the token-owned cross-process lock plus lease through the complete bounded provider request. Retirement cannot pass that bound; stale active processes fail before egress. A previously started run keeps its recorded adapter identity through recovery.
- Authorization/configuration: block until configuration changes.
- Letta Agent SDK
success: false: settle the attempt's conservative accounting charge, classify the remote failure, and leave source progress unchanged. A later retry first reconciles the deterministic turn marker before sending. - Per-document Letta conversation binding: derive one content-dark summary marker from declaration, agent, source, and stable document id; reuse one exact remote match, create only on zero matches, and fail closed on multiple matches. Jazz then inserts or verifies the mapping before the turn marker is reconciled.
- Inference budget exhaustion: terminally block the attempt before provider dispatch.
onExhaustion: advancecontinues from durable consumer progress;onExhaustion: deferleaves progress unchanged until the limiting windows permit a later attempt. - Permanent source deletion or malformed source: append an observation/error and advance only according to connector policy.
Rebuild #
thought stream rebuild drops only projection data, replays events, and verifies projection hashes. It does not delete events, document versions, source cursors, consumer lifecycle evidence, or traces. Effective output is rebuilt by original run id from active judgments over correction proposals, then original valid outputs, then unresolved failures.