Output repair requests and correction proposals #
Boundary #
Output repair is an append-only proposal workflow for a narrow class of terminal model-output failures. It does not retry infrastructure, recover hidden reasoning, reinterpret quarantined text, or make a replacement effective without an existing judgment authority.
The deterministic coordinator uses repair policy stream.thought.agent.output-repair@1. Its identity includes a canonical SHA-256 hash over the policy definition. For each eligible original failed run and policy version it appends at most one stream.thought.agent.repair.requested@1 event. The idempotency key is derived from the original run id and complete policy identity, so crash recovery and repeated scans address the same event row.
Eligibility #
A run is eligible only when all of the following are true:
- it is a terminal failed run from a declaration whose role is
standard; - it has no output event and is not itself triggered by a repair request;
- it uses the currently supported
stream.thought.output.observation@1repair contract; - the original trigger event, failed lifecycle event, exact declaration version, prompt hash, declaration fingerprint, canonical output-contract identity, and sanitized validation evidence are all present and mutually consistent;
- its diagnostic is either invalid JSON, a canonical output-contract violation, or a semantic-validation failure whose stable rule id is explicitly allowlisted by the policy;
- its validation diagnostic contains only bounded counts, hashes, stable reason/rule ids, and sanitized contract issues.
Timeouts, cancellation, provider outages, rate limits, sandbox/broker/process failures, credential/configuration/authentication failures, incomplete or corrupt evidence, repair-origin runs, non-observation output contracts, and unallowlisted semantic failures are ineligible. The coordinator never reads malformed candidate text or quarantine data while deciding eligibility.
Canonical output contract #
The current repair workflow supports only stream.thought.output.observation@1. The output-contract registry binds its id and version to one canonical definition, validator, prompt description, and SHA-256 hash. Original execution, repair execution, human correct replacement validation, effective-output rebuilding, and training export resolve that same identity. A hash mismatch is an evidence failure, not a request to use a nearby schema. Other registered contracts, including conceptualization graphs, fail terminally without creating an unusable repair request until a contract-matched repair declaration exists.
Every run context manifest and accepted output event records the contract id, version, and hash. Invalid-output diagnostics record the same identity plus sanitized issue codes and paths. They never record rejected values or raw model text.
Repair agent #
A repair consumer is declared with role repair, selects the escalation model tier, and subscribes only to stream.thought.agent.repair.requested. It ships disabled. The trusted host must map THOUGHTSTREAM_TINKER_ESCALATION_MODEL to a larger or better-adapted Tinker model before enabling it; an enabled declaration fails closed when that mapping is absent. A disabled declaration may load without the mapping so unrelated services retain restart safety while activation is awaiting an explicit host decision. The repair consumer is still a Pi agent and therefore uses the mandatory disposable Bubblewrap worker and single-run provider broker. It has no external-action tools, trusted-host model fallback, or authority to mutate an original run.
The trusted parent constructs repair context from:
- immutable references to the original run, failed lifecycle event, trigger event, and source root;
- the exact original declaration metadata and repository prompt whose hashes match the failed run;
- the original provider/model identity, execution-adapter revision, learned-adapter identity, and startup catalog digest/generation;
- the original bounded source-event context regenerated from the immutable trigger event;
- the canonical output-contract definition and identity;
- bounded sanitized validation issues from the request.
The context excludes malformed candidate text, hidden reasoning, provider request/response bodies, prompt copies recovered from traces, arbitrary failure text, quarantine content, and trace payloads. Repair is fresh proposal generation from the original evidence, not substring salvage.
A repair request permits one model proposal generation. An ambiguous interrupted repair run is terminally abandoned and its request progress advances; restart does not generate a second proposal. Any repair validation failure writes only a sanitized terminal failed run and advances the request. No repair request is generated for that run, no model cascade occurs, and there is no repair-of-repair path.
Proposal and authority #
A valid repair appends stream.thought.derived.output.correction.proposed@1. The event links the original failed run, repair request, repair run, original trigger and source root, contract identity, separate original/repair model-adapter/catalog provenance, and one validated structured output. Its privacy is the join of request state and the repair declaration's learned-adapter privacy. It is inert by default.
Existing append-only judgments are the authority surface:
- active
acceptmakes the proposal output effective; - active
correctmakes its contract-validatedreplacementOutputeffective; rejectleaves the original failure unresolved;- supersession or retraction removes the affected judgment from the active set and rebuilding recomputes the result.
A correction judgment is rejected before recording unless its replacement validates against the original run's output contract. Originals, proposals, and judgments are never mutated.
Judgment sources are independent authorities. Supersession is explicit: a later judgment names the exact prior judgment it replaces, including when an exact Telegram correction follows a reaction judgment. Active leaves from different authorities may coexist. The effective-output projection chooses the latest contract-valid active authority by observation time with event id as a deterministic tie-breaker.
Effective-output projection #
The rebuildable projection is keyed by original run id and applies this order:
- latest contract-valid active
correctjudgment on the valid original run; - latest active accepted or corrected repair proposal;
- original valid output;
- unresolved failure.
A direct-original correction must reference the original run's exact output event and carry a replacement that canonicalizes against the run/output contract identity. Mismatched output pointers, missing replacements, unknown contracts, or schema-invalid replacements are ignored as corrupt historical authority rather than made effective. Its privacy joins the original source, run, output, and judgment. reject, accept, and prefer judgments on a valid original do not replace its effective output.
Projection rows contain only contract-validated structured output, joined privacy, provenance ids, and separate original/repair public model-adapter/catalog evidence. A direct correction row records status corrected, the original output event, correction judgment, authority correct, and canonical replacement. Rows can be dropped and reconstructed from runs and append-only events. Rebuilding after judgment supersession or retraction must produce the same result as uninterrupted processing.
Training boundary #
Repair trajectories enter external training export only when the repair proposal has an active accept or correct judgment with both quality and external-export eligibility. Default export also requires the joined original source/run, repair request/run/proposal, judgment, and chosen/rejected chain to be public-source; learned-adapter privacy/export policy is applied independently to every participating run. Sensitive/private repair material requires the separate declassification and destination gates. Rejected, unresolved, retracted, failed, and merely proposed repairs are excluded. Export contains validated chosen/rejected structured output, minimal source classification, allowlisted trajectory type/order, contract identity, and public model/adapter/catalog provenance. It excludes malformed candidate text, source bodies, prompts, provider bodies, reasoning, tool arguments, image bytes, arbitrary diagnostics, actor/route/external/correlation/idempotency identifiers, internal provenance ids, source hashes, trace-content hashes, private checkpoints, and quarantine material.