Operational incidents #
Operational incidents are a content-dark projection over existing failure and recovery evidence. They make failures queryable, appendable to one private ledger, and eligible for a separate alert policy without turning source bodies, provider detail, or arbitrary error strings into another data sink.
Contract #
stream.thought.runtime.incident@1 is append-only and always sensitive. Its payload contains only bounded values from this allowlist:
incidentVersion:1incidentId: deterministic identity for one source failure, recovery, run terminal state, scheduler exhaustion, or delivery failurestate:openorrecoveredcategory: versioned incident categoryseverity:info,warning,error, orcriticalcomponent: bounded runtime component identifiercodeand optionalstage: allowlisted machine classificationsfingerprint: SHA-256 over normalized category, component, source, code, and stage; never over raw error textoccurredAt: the originating evidence timeretryable: whether ordinary reconciliation can attempt the underlying work againprogress:advanced,unchanged,unknown, ornot-applicable- optional bounded
source,agentId,agentVersion,attempt, and strong reference ids
The incident payload never contains Error.message, stack traces, provider or Telegram bodies, prompts, source payloads, model output or reasoning, tool arguments/results, routes, credentials, environment values, command lines, or hashes of credential-bearing raw text. The source failure row may predate this contract and contain more detail; the incident projector does not copy it.
Projection inputs #
The first slice normalizes:
stream.thought.connector.failedas an open connector warning;- a failed
stream.thought.connector.subscription.stoppedas an open terminal connector error; stream.thought.connector.recoveredas connector recovery evidence;stream.thought.agent.run.failed,.blocked, and.abandonedusing the terminal run, trigger, and consumer progress evidence;- exhausted
ConsumerScheduleroperations after their configured retries; stream.thought.action.telegram.send.failedas a delivery incident.
Projection is deterministic and idempotent. The first valid append-only incident row for an origin event is authoritative on later projection passes; current consumer progress cannot rewrite its retryable or progress classification. If an earlier failed attempt is first projected only after a later attempt under the same execution key exists, the earlier failure is classified as progress-unchanged and retryable even when the later attempt has since advanced source progress. It does not alter the originating event, retry count, consumer progress, model invocation, or delivery claim. A failed incident projection is an observability failure, not authority to replay an external action.
Private ledger #
When the manifest enables incidents, the incident service writes every normalized incident to one configured JSONL file below the runtime root. The writer:
- rejects absolute, escaping, and symlink paths;
- creates private parent directories and enforces mode
0600on the ledger; - emits canonical one-line JSON bounded to 4096 bytes;
- opens in append mode, writes one complete line, and calls
fsyncbefore acknowledging it; - reloads deterministic incident ids on restart so a completed append is not repeated;
- treats a malformed or oversized existing ledger as a startup failure rather than silently replacing it.
The ledger is a private operational document. It is not a source for model context, training export, public content, or ordinary inspector payload rendering.
Telegram incident alerts #
Incident alerts are a separate action policy from normal output notifications. They select normalized incident categories, not trigger-source allowlists. A Jetstream-rooted agent failure can therefore alert without making successful Bluesky observations Telegram-eligible.
The dispatcher does not suppress incidents by agent identity or retry disposition. Canary, test, and retryable failures all generate alerts. Every alert says what subsystem failed, whether the runtime will retry, and a short receipt hash for correlation. Technical component identifiers, source labels, classification codes, and stage names remain in the private ledger and inspector projection — they do not appear as raw field labels in the Telegram message.
Alert format #
Each alert is a short multi-line message:
- A header line:
The Stream · <short title>(with occurrence count when > 1). - A plain-language description of the failure.
- A retry line:
Will retryorWon't retry automatically, with attempt number when available. - A receipt line:
Receipt <short id>for correlation.
The alert does not contain raw component:, source:, classification:, retry:, or progress: field labels, agent ids, event ids, run ids, or tool names. The title and description use plain human language naming the affected subsystem and the nature of the failure.
Dispatch protocol #
For each destination, the incident dispatcher:
- establishes one durable activation cutoff;
- reads open incident events after that cutoff;
- groups by normalized fingerprint;
- excludes Telegram delivery failures to prevent recursive alert attempts;
- enforces a durable destination window and per-fingerprint cooldown;
- appends a deterministic
startedclaim before one Bot API call; - appends
deliveredor content-darkfailedevidence.
A claimed Telegram attempt is intentionally at-most-once: after the durable started claim, it is never retried automatically, even when no terminal receipt is observed. Telegram provides no idempotency key, so a blind retry could duplicate a notification whose first Bot API call succeeded but whose response was lost. Repeated incidents inside the cooldown remain in the ledger and may be summarized after the cooldown.
Process boundary #
This slice runs as a dedicated incident projection/ledger/alert process. It can record errors emitted by other running processes, but it cannot reliably report its own startup failure, a Jazz-open failure, or a host-level crash after it has stopped.
The external supervisor seam is explicit: a later systemd ExecStopPost/OnFailure helper should emit the same normalized contract from bounded unit metadata (SERVICE_RESULT, exit code/status, unit name, invocation identity), with explicit restart limits. It must not copy journal messages, environment, command lines, or stacks. This follow-up is required before claiming complete process-crash and malformed-startup coverage.