Consumer and agent system #
An agent is one model-backed kind of consumer. Projectors, bridges, rules, and deterministic transforms use the same Jazz subscription and progress contract without pretending to be model agents.
Declaration #
Agents are YAML files compiled into validated Jazz rows. A minimal declaration contains:
id: document-structure
version: 1
enabled: true
description: Identify document structure changes worth re-evaluating.
subscribe:
types:
- stream.thought.source.file.added
- stream.thought.source.file.changed
sources:
- filesystem:charter-fixture
context:
maxEvents: 8
maxChars: 24000
runner:
kind: deterministic
prompt: prompts/document-structure.md
emit:
- stream.thought.derived.document.structure
policy:
tools: []
externalActions: false
A Pi runner selects one exact learned Tinker release through adapter: { id, version }. Startup resolves it from the validated immutable release and deployment catalogs before the declaration becomes a Jazz row. A learned adapter is mutually exclusive with runner.model and runner.tier, must match the trusted provider profile, and receives a runtime binding only when its deployment entry is active. The deeply frozen declaration contains the public content-addressed release/binding and catalog identity. Only the resolved checkpoint value remains in module-private weak bindings keyed by the original compiled adapter/declaration objects. Cloning, rehydration, manual construction, or mutation cannot recreate dispatch authority. Declarations without a learned adapter remain backward compatible.
Every pi and letta-agent-sdk declaration also requires an accounting policy. It declares a conservative per-call reservation, a lease longer than the runner timeout, and one or more rolling/hour/day limits. Calls, input tokens, and output tokens are always reserved; micro-US-dollar cost reservation and limits are optional. A cost limit without a cost reservation is invalid because the runtime cannot enforce a dimension it did not reserve. A declaration that cannot admit one complete reservation in every tracked dimension is invalid. Deterministic consumers cannot declare inference accounting.
The retained Inkling-Small Telegram declaration reserves $0.05 per attempt because the OpenAI-compatible beta reports tokens but not provider cost. That estimate bounds one full 131,072-token Pi context plus the declared output at the current serverless list price, rather than assuming an ordinary short turn or a cache hit. Its 30-day rolling window admits at most $50 of those conservative reservations. Settlement replaces token estimates with reported usage but retains the $0.05 cost estimate, so the local dollar window is deliberately an upper-bound ledger rather than a provider invoice. The declaration and its reciprocal compactor are currently disabled; they remain a reversible implementation path, not the active direct Telegram owner.
A Pi declaration may select an explicit runner.thinkingLevel; omission preserves the off default. The trusted parent binds that value into the sandbox packet, and the worker passes it to Pi's agent state rather than imposing one process-wide setting. For Tinker/Qwen chat-template models, off serializes chat_template_kwargs.enable_thinking: false; every non-off level serializes enable_thinking: true while the model/provider decides the internal budget. The retained Pi Telegram declaration selects medium. This setting is declaration identity and therefore appears in the declaration fingerprint and run context, while hidden reasoning content remains excluded from durable traces.
The resident keeps measured token accounting and a conservative full-conversation reservation of 60,000 input / 2,000 output tokens per call, rather than estimating from the current packet size. Its tracked limits are emergency-only circuit breakers: 100 calls / 100,000,000 input / 10,000,000 output per rolling five minutes; 1,000 calls / 1,000,000,000 input / 100,000,000 output per hour; and 10,000 calls / 1,000,000,000 input / 1,000,000,000 output per day. Dollar cost is absent. This accounting is operational telemetry plus a catastrophic-runaway guard, not a budget or thrift mechanism; ordinary or extreme post, like, and Semble activity should not approach the ceilings.
Exactly one execution path owns direct Telegram replies. The active path is resident-letta-conversation@6: one persistent Letta Cloud agent main conversation receives the current webhook event as a bounded single trigger and owns its prior interaction internally. A direct turn includes projected text plus at most one current Telegram image artifact; the trusted parent revalidates the content-addressed file before passing it to the SDK as native multimodal content. Missing, malformed, changed, escaped, or otherwise invalid image evidence fails closed before the turn is sent. The persistent private SDK conversation may retain the validated multimodal message as remote history; local Jazz events, traces, incidents, and projections retain only the opaque artifact evidence. The Pi telegram-conversation@22 and reciprocal telegram-conversation-compactor@4 are disabled together, so one source event cannot produce competing Letta and Tinker replies. Switching paths is a coordinated declaration, credential, consumer, and dispatcher deployment; changing a model string alone is not a cutover.
The same resident-letta-conversation@6 also consumes batch:stream-activity. The deterministic ten-minute batch merges complete prefixes from several source namespaces into one ordered window and sets the batch privacy to its most-private member. Direct Stream Telegram messages are deliberately excluded from that batch because the resident has already received them in the same persistent conversation; non-conversation, reaction, and correction events may still join the window. Direct Telegram turns keep conversation-text output. Activity-window turns use strict JSON under the observation contract. Both routes serialize through the resident's one persistent main conversation.
An activity-window output may propose scarce human attention through the ordinary observation contract. Telegram eligibility requires the configured resident/source tuple, importance: high, the exact notify-cameron tag, and a recommendation targeting cameron-telegram with proposed action notify. The proposal stays inert until the dispatcher claims it. Direct Telegram replies use a disjoint source route and do not need the notification tuple. Ordinary window observations remain in Jazz.
Three disabled-by-default listener declarations add a layered social-attention path without replacing the resident. cameron-bluesky-listener@2 consumes the existing exact-member ATProto batch and receives canonical public ATProto expansion. Its explicit privacyFloor: private keeps persistent model derivations private even though every admitted source member is public. cameron-x-listener@1 consumes one deterministic most-private batch across the three X lanes while retaining each member's source identity. Both use separate persistent output-only Letta conversations on Luna, strict permissions, no skills, no tools, and no external actions. They emit ordinary private-or-stronger observations and have no dispatcher route.
cameron-social-listener@1 consumes only a deterministic batch of those two listeners' derived observations, not raw social records. Before context construction, the trusted parent requires each member to match exactly one completed run, trigger, agent source, output id, execution key, semantic result, and lineage. The validated receipt identities enter the context snapshot. The persistent Terra conversation may connect source-local observations across time. Its output settles to Jazz before any delivery decision. Only this agent and exact social-batch source may use the separate notification-proposal route; the dispatcher still requires the high-importance tuple and exact completed-output lineage. All three declarations and both new batches remain inert until their independent environment, identity, and manifest activation steps complete.
Every enabled main Agent SDK declaration must resolve to a distinct backend/agent identity. Multiple concrete sources inside one declaration still serialize through that declaration's one persistent conversation. Reusing one main agent across different declarations would collapse prompts and histories while looking like separate listeners, so declaration loading fails before any turn is admitted.
Listener identity provisioning is one target at a time and separate from activation. The repository command first derives a read-only confirmation from the exact declaration, prompt hash, model, and no-tool creation profile. Applying that confirmation creates at most one hidden Agent SDK identity, immediately persists owner-only resumable state before remote verification, and refuses declaration drift, duplicate concurrent provisioning, or a changed confirmation. Provisioning does not edit credentials, enable a declaration/batch, restart a service, or run the model.
Subscription #
The declaration compiles directly into a Jazz event query and live subscription. It may constrain event type, source, privacy class, address, typed payload fields, and bounded batching rules. The consumer process is both subscriber and runner; there is no separate matching service or queue.
A managed consumer maintains per-source progress keyed by consumer id/version and source id. It processes source sequences serially by default. Startup queries the bounded backlog after each source's progress, then follows live deltas. Query/subscription overlap is absorbed by deterministic execution identity.
The subscription declaration describes:
- exact event types or registered type-family prefixes;
- exact source ids and/or source kinds;
- accepted privacy classes;
- an optional destination/address;
- equality/range/set predicates over indexed envelope columns and promoted typed payload fields;
- initial replay policy for a source with no progress:
beginning,now, or an explicit source sequence; - bounded batch size, batch wait, and processing concurrency;
- behavior when the declaration version changes: preserve progress, replay selected sources, or start a new consumer identity.
Recovery compiles one bounded Jazz query per included source because each source has its own sequence space. Live delivery uses the corresponding narrow Jazz query subscription. The logical joint stream is the union of those source streams; no process has to manufacture a global queue first.
Context packet #
The packet contains exact ids and bounded rendered content:
- Trigger events.
- Referenced document versions or source records allowed by policy.
- Explicitly requested neighboring events.
- Agent prompt revision.
- Output schemas.
- Runtime model, execution-harness adapter revision, and optional learned model-adapter identity as separate fields.
- Privacy and tool capability statement.
The packet records omitted/truncated content. Silent truncation is forbidden.
Conversation reconstruction, compaction, selection, and composition #
Telegram conversation context has four explicit stages with different authority:
- Canonical history reconstruction rebuilds the same-chat conversation through the current inbound event from durable evidence. It includes every eligible user turn, a neutral
[image]placeholder for prior image-only turns, only assistant outputs with same-chat delivery receipts, exact validated effective corrections, and receipt-backed proposal call/result history. A corrected assistant turn retains both the exact originally delivered text and the validated effective replacement; only the replacement becomes model-facing content. Reconstruction does not apply declaration event/character bounds, assistant-history caps, repeated-fragment suppression, or model-facing prompt composition. Original events, runs, outputs, deliveries, corrections, and projections remain immutable. - Conversation compaction is an optional derived projection. A separate no-tool compactor clone consumes one immutable snapshot of an oldest history prefix, plus the previous boundary when present, and emits a sensitive typed boundary with exact coverage and input hashes. It never mutates reconstruction. See
compaction.md. - Model-context selection starts from reconstructed history and the latest valid boundary. Covered raw turns are replaced by the boundary; exact post-boundary turns then receive declaration
historyAgentIds, an explicit per-rolemaxEventscandidate window, the assistant-history cap, correction-derived repeated-fragment suppression, the combined event bound, andmaxChars. The per-role candidate window preserves v19 behavior that previously arose implicitly from limited source queries. Every omission or transformation is named in the context manifest; selection never changes canonical history. - Pi composition combines the selected typed boundary, native messages, the current user prompt and current-turn image artifacts, and trusted system/document context. It does not retrieve or reinterpret conversation history independently.
The reconstructor, compactor compiler, and selector are separately callable and tested. A declaration policy change may alter future model input without changing what the evidence-backed conversation history says happened.
Subscribed document context #
A Pi-backed Telegram conversation may add an exact trusted-document recipe under context.documents. Each subscription names one concrete filesystem source and an ordered list of normalized relative document paths. Wildcards, arbitrary queries, declaration-controlled roots, and a general-purpose template language are intentionally absent. required: true makes every named document fail closed when its current projection or immutable version evidence is unavailable. The declaration gives the document portion its own character ceiling inside the total context.maxChars; trusted document content is never silently truncated to make a turn fit.
The trusted parent resolves each current document projection to one immutable Jazz document version, verifies source, stable document identity, path, content type, byte count, and SHA-256, then renders the ordered versions into the system-role portion of the Pi request after the declaration-derived runtime authority block. These documents may supply operator-authored identity and continuity context. They cannot change the declaration's tools, external-action authority, provider capability, accounting policy, retry policy, or output contract; those controls remain runtime-owned.
The same system-role context begins with one small environment block: thought stream supplies one current event, selected earlier conversation, and operator-written identity/memory documents; the latest user message is the current task; earlier assistant messages are context rather than identity. Exact declaration, runner, provider, model, version, and continuity metadata remain snapshot-bound in the private context manifest for recovery and diagnostics but are not injected as model-facing identity prose. Prior delivered assistant replies remain available as conversation, including their mistakes.
The Telegram transcript is delivered to the provider as native role-separated messages, not a flattened JSON packet. Ordinary turns use user and assistant. A delivered run whose atomic completion receipt names durable proposal events created by The Stream is reconstructed as one assistant tool-call message, one matching tool-result message per call, then the assistant text actually delivered to Cameron. The call uses canonical proposal-event arguments and a stable synthetic call id. The normalized result explicitly says that proposal capture executed and created a durable inert proposal while human approval/application remains pending. This is semantic reconstruction, not a claim that the original provider-generated call id, literal result wording, or latest shorthand survived verbatim. Acknowledgment text without atomic proposal evidence remains text only.
Private agent-conversation context is a distinct native-role surface. It reconstructs one exact sender/recipient/thread from typed agent-message source and response events, uses the current source text once as the final prompt, and never imports Telegram history or exports its responses to the Telegram dispatcher. See agent-messages.md.
The post-training-course-tutor@1 declaration is a separate stateless private tutor. It subscribes only to stream.thought.source.course.question@1 from web-course:post-training-model-factory, uses one-event context with server-derived lesson text, emits one strict observation, and has no tools, external actions, proposals, persistent conversation, Telegram route, or public projection. Its output lineage must close over the exact question before the course status API returns an answer. See courses.md.
context.historyAgentIds may admit delivered replies from explicitly named prior conversation agents across declaration versions, allowing a Pi/Tinker declaration to inherit the visible channel history during a harness migration. Only completed runs with an actual same-chat Telegram delivery receipt are eligible. Undelivered output, a different chat or sender, unlisted agents, and proposal events not named by the exact completed-run receipt do not enter message history. Event ids and run ids stay in Jazz provenance except where snapshot-bound proposal arguments inherently name evidence or a correction target inside the native tool-call metadata; they never leak into visible assistant text. Prior assistant messages remain untrusted output that cannot override the current runtime authority block. The latest inbound user message is the current prompt; bounded prior messages precede it in the sandbox packet.
Image-only messages (empty text with exactly one stored, validated image attachment) are admitted as conversation triggers. Empty messages without a stored image use stream.thought.source.telegram.nonconversation; rejected images and non-image attachments remain source evidence outside the conversation subscription and cannot block later turns. The transcript renders a neutral [image] placeholder for both the current image turn and an eligible prior image turn, so the immediately following delivered reply retains its trigger and can re-enter history. The retry-stable snapshot carries a schema-validated opaque imageArtifacts reference array and its canonical SHA-256 (relative path, content SHA-256, MIME, byte count) for the current event only; prior turns retain only the placeholder and never replay image bytes. The trusted Pi parent resolves each current artifact reference beneath the artifact root and injects bounded base64 ImageContent into the sandbox packet before provider dispatch. Image-capable models (e.g. thinkingmachines/Inkling, thinkingmachines/Inkling-Small) are marked in the provider profile's imageInputModels set; text-only models receive no image parts.
The trusted system prompt names the current turn's resolved image count. With zero current image artifacts, the model must not claim visual access or reconstruct a photo from prior placeholders, assistant descriptions, or prose. Prior image descriptions remain ordinary untrusted text history.
An exact private /help command (or /help@<current-bot-username>) remains a durable Telegram source turn but bypasses provider and sandbox dispatch inside the trusted Pi parent. It emits one fixed versioned conversation observation with command, feedback, inspector, and documentation guidance, then follows the ordinary atomic run/output/progress and dispatcher delivery chain. The response version changes whenever fixed help content changes. The trace records only command name, response version, and zero provider requests. Other slash-shaped text remains ordinary conversation unless a separate fixed command contract owns it.
For every trigger, the complete system-document plus conversation packet is persisted as one immutable Jazz document version before provider dispatch. Its identity binds the declaration fingerprint and trigger event. The first attempt selects current document versions; every retry reuses and integrity-checks that exact snapshot even if a subscribed document changes later. A later Telegram event receives the newer document version. The model call remains stateless: durable events, document versions, receipts, context selection, and retry identity are the agent state.
The deployed context root is a separate producer. thoughtstream-agent-context.service watches the bounded private root in explicit --producer-only mode and owns filesystem:telegram-agent-context; the ordinary consumer service remains the sole model-runtime owner. The source service has no provider or channel credential compartment.
The Pi conceptualizer is the narrow non-resident user of atproto-batch. It binds the conceptualization output contract to stream.thought.derived.concept.graph, subscribes to one ATProto batch event type, uses no tools or external actions, and receives member-expanded context from the trusted parent. Other Pi declarations cannot opt into ATProto object expansion by configuration alone.
Model cells and agent harnesses #
The trusted consumer runtime owns context selection, provider authorization, persistence, output validation, accounting, concurrency, and retry policy. A model runtime or agent harness is never the event store and receives no external-action capability.
Ordinary observation and repair consumers use the observer-v1 model cell: Pi agent core runs with tools: [] in the existing disposable Bubblewrap sandbox. It has no workspace and a single-use broker capability permits one bounded provider request. This remains an inference-only profile; it is not evidence that a coding harness or arbitrary Pi extension is safe.
Provider JSON response formatting is a capability claim, not a request decoration. The built-in Tinker profile does not advertise it because observed Tinker routes accept response_format: {"type":"json_object"} without enforcing a JSON object. Strict JSON output on that profile therefore means explicit prompt instructions plus complete parent-side contract validation and rejection; it must not be described as constrained decoding. The separate openai-json-default profile fixes the OpenAI endpoint, credential name, and allowlisted model set and binds the conceptualization declaration to one trusted strict JSON Schema response format. The schema constrains object fields and enums at decoding time; parent validation still enforces cross-field graph invariants and remains authoritative.
Tool-using or persistent agent runtimes use the separate generic contract in harnesses.md. The first reference adapter is pi-coding@1 in the workspace-v1 container profile. Declarations select only an allowlisted adapter/profile and trusted model tier. They cannot supply an image, host mount, executable extension, provider URL, credential reference, broker budget, or container flag. Any container, broker, or lease setup failure leaves source progress unchanged for a later retry. Neither profile has an in-process or trusted-host fallback.
The letta-agent-sdk runner is a distinct stateful harness adapter. The letta-cloud-v1 profile uses the Agent SDK Cloud backend and a managed sandbox. The letta-local-memory-v1 profile uses an SDK-owned loopback App Server against an existing cloud-backed agent, while tools and skills are empty and the harness filesystem is fail-closed to that agent's memory root. Jazz remains authoritative for source events, consumer progress, lifecycle evidence, accepted outputs, and channel delivery. The Letta agent owns its conversational continuity and agent memory. thought stream must not rebuild a synthetic transcript and send it again on every turn.
An enabled letta-agent-sdk declaration names a trusted existing agent through an environment-variable reference and uses a bounded set of concrete source namespaces. It selects either the agent's main conversation or one conversation per stable filesystem document. The declaration may choose a fixed model handle, reasoning effort, permission mode, dreaming trigger, and final-response mode. Cloud declarations may choose sandbox TTL. Local-memory declarations name only the environment variable containing the agent memory root; they cannot select a general host path. No declaration can select an API-key variable, API base URL, websocket URL, sandbox image, arbitrary host root, or remote environment.
The stateful conversation topologies are intentionally narrow:
- one enabled declaration owns one Letta agent main conversation;
- one or more explicitly named concrete thought stream sources may feed that declaration;
- all source subscriptions owned by that declaration share one scheduler key derived from the Letta agent id, so only one SDK turn may execute against the main conversation at a time;
- the context strategy is
single-eventand the prompt contains only the current event packet plus trusted declaration instructions and a deterministic turn marker; - earlier Letta messages are available to the agent through its own conversation, but earlier thought stream messages are not copied into the new user turn;
- two enabled declarations may not share one Letta agent id, because concurrent sessions can race conversation state and MemFS.
- a
per-documentdeclaration accepts only filesystem add/change/rename events from one concrete source and usesdocumentIdas the conversation scope key; - Jazz and the remote conversation summary marker jointly recover the mapping, while the current path may change without changing conversation identity;
- all document conversations for one agent still share one scheduler operation key, so the default topology permits one Co turn at a time even though conversations are distinct.
The Coil Public Knowledge context strategy runs its default-deny policy before SDK session creation. It loads the exact event-bound document version from Jazz, verifies source/document/path/version/hash consistency, reads all public catalog metadata, and admits only a bounded deterministic set of relevant existing entry bodies as exact replacement targets. Blocked, deferred, deleted, unsupported, oversized, stale, or policy-ambiguous inputs never open a conversation. The model sees the source document, admitted public targets, and Co memory but receives no Jazz, vault browser, site writer, or publication tool. Its only tool is the controller-owned submit_public_knowledge_diff; the callback parses one bounded raw JSON string, validates the resulting object against stream.thought.output.public-knowledge-proposed-diff@1 and current context, captures it in memory, and performs no effect. The trusted parent binds a replacement to an admitted slug and base hash or proves a new slug absent, then settles the captured proposal rather than parsing the model's final acknowledgment. Full contract: public-knowledge.md.
Multi-source serialization is an execution-safety contract, not a globally canonical event order. Each source retains its own monotone sequence and independent consumer progress. When events from different sources become ready together, the resident sees them in the order their source operations enter the shared scheduler chain. That order is durable through each accepted turn and sufficient for one conversation owner; it must not be represented as a total order across the original source clocks.
A persistent conversation joins the information available across its accepted source classes. Every output and lifecycle event from a Letta SDK declaration therefore uses the most restrictive privacy class accepted by that declaration, even when the current trigger is public. A public ATProto event fed into a resident that also knows sensitive Telegram history cannot produce a public-source derived event.
For Cameron's resident stream, Telegram messages and the narrowly filtered jetstream:cameron-bluesky source may share the declaration. Derived ATProto batching exists for semantic coherence and to avoid redundant resident turns over one burst, not to save model cost. That Jetstream source admits Bluesky posts and likes plus network.cosmik.collectionLink; it does not separately admit network.cosmik.card, network.cosmik.collection, or note-child cards. The collection-link record is the one-per-save trigger, preventing a card write and its organizational link from producing duplicate resident turns. ATProto packets retain their strong atUri, CID, collection, operation, and bounded original record fields.
When context.atprotoObject is enabled, the trusted parent selects a source-specific compiler. Bluesky posts and likes keep the existing two-view behavior: atproto.md supplies repository/provenance Markdown and bsky.md supplies the social object, embeds, quotes, engagement, and thread links. A network.cosmik.collectionLink create instead produces three independently bounded atproto.md views for the exact link event, its strong-referenced network.cosmik.card, and its strong-referenced network.cosmik.collection; it never calls bsky.md. One failed dereference does not erase the other views or the original source record. Normal ingress admits collection-link creates only; updates and deletes produce no resident turn. If historical delete evidence reaches the compiler directly, it preserves the delete envelope and skips all stale dereferences.
Mutable Markdown does not prove which CID its rendering represents. Every view therefore remains labelled current-record-unverified unless stronger evidence exists. For Bluesky, the compiler compares the expected strong-reference CID to the current post CID reported by the public AppView; a mismatch discards both protocol and social renderings for that shared target. For Semble, any independently observed CID mismatch discards only that one mutable view. The original Jetstream record and its strong references remain the exact source evidence.
The complete bounded source-plus-enrichment packet is written once as an immutable content-hashed Jazz document version keyed by declaration fingerprint, event id, and every source-specific target URI/CID. Every retry reuses that exact snapshot and verifies its stored text hash. It survives ordinary projection rebuilds. Fetched source content is live runtime data and must never enter Git, traces, accounting, or a public projection.
The resident remains a capable Letta agent after the packet crosses the boundary. thought stream's security contract is to specify and snapshot the feed, label external content as data, withhold host/source credentials, prevent fetched or conversational content from entering Git or public projections, classify mixed-state outputs as sensitive, and grant no ATProto write authority. Prompt guidance reminds the resident not to expose private continuity, but the architecture does not split its conversation or cripple its normal sandbox solely because one trigger was public.
Stable resident instructions should be clear but small. Runtime-enforced response size and adapter format rules do not need to be narrated to the model on every turn. The final wrapper distinguishes only the semantic destination: a reply to Cameron for Telegram or a private internal observation for ATProto.
The deterministic turn marker is derived from declaration id/version and source event id. It is sent as trusted runtime metadata and never contains source text. Before sending a turn, the adapter checks bounded conversation history for that marker. If a matching assistant result already exists, the adapter recovers it instead of sending again. If the marker exists without a terminal assistant result, the adapter waits briefly and then leaves source progress unchanged. If the history window ends while older pages still exist, or the backend claims more pages without a valid cursor, marker absence is inconclusive and the adapter refuses to send. This is recovery evidence around the SDK's current lack of a caller-supplied idempotency key; it is not represented as provider-native exactly-once delivery.
An SDK terminal result with success: false settles the current inference reservation but does not advance source progress. Billing, authorization, rate-limit, and generic remote-agent failures are classified from process-local SDK detail into allowlisted codes; the raw detail is discarded. This may hold that declaration at its first failed event until configuration or service state changes, which is preferable to silently losing the event. A successful provider result whose final text fails the declared response contract remains a terminal invalid-output failure and follows the normal repair policy instead of blind provider retry.
letta-agent-sdk stream events are converted into metadata-only traces. Assistant and reasoning content, tool arguments, and tool results are represented by bounded counts and hashes. The terminal assistant value is kept process-local until it passes the declaration's response adapter and canonical output contract. Agent id, conversation id, SDK run ids, backend, package version, duration, stop reason, and bounded cost telemetry may be durable provenance.
Agent SDK 0.2.6 does not expose the Letta Code wire result's usage object. After receiving a bounded set of valid SDK run ids, the Cloud adapter therefore makes best-effort reads against the fixed GET /v1/runs/{run_id}/usage endpoint using the same process credential. Prompt and completion tokens are reported only when every run id has that dimension, so a partial lookup cannot silently undercount a multi-run turn. Lookup failure never changes a successful semantic result; it leaves that accounting dimension on the configured estimate. The token endpoint has no dollar-cost field. Dollar cost is absent unless the SDK terminal result supplies a finite, strictly positive totalCostUsd; zero is not converted into measured OAuth cost.
Before the broker or sandbox can dispatch a provider request, the trusted parent durably reserves that declaration's configured estimate against its Jazz-backed agent account. After either success or failure, the parent settles the reservation with reported usage for policy-tracked dimensions and retains configured estimates only for tracked dimensions that remain unavailable. Token-only OAuth policies therefore omit costMicrousd from reservation estimates and charges rather than serializing zero or invented dollars. Their usage status is reported when both tracked token dimensions are reported. A repair consumer performs an independent reservation under its own account; the original run never spends its repair budget implicitly.
Observer-cell tools are explicit declaration capabilities executed by the trusted parent before sandbox launch. The initial read-only set can dereference the current ATProto record through atproto.md and download only image URLs discovered in that record or its fetched Markdown. The runtime validates public destinations, follows a bounded redirect chain, caps response bytes and image count, and persists content-addressed image artifacts. A declaration without a tool name receives no corresponding evidence. Workspace-harness tools execute only inside the leased container boundary and are fixed by its adapter/profile pair.
For a learned adapter, startup compilation validates one exact active deployment selection and keeps its checkpoint in a process-local private binding. Read-only evidence prefetch completes before provider egress. The Pi runner resolves the checkpoint only from the original compiler-bound adapter object, then gives it to the trusted broker as the provider model. One serialized admission section rechecks expiry and atomically reserves request count plus cumulative request/response bytes before fetch; response reservation converts to actual usage in the same critical section. Durable broker traces render the public base model, learned-adapter identity, and catalog digest/generation, never the private checkpoint value.
The runner maps sandbox and Pi events to metadata-only trace chunks. Durable traces keep event type, role/model/stop metadata, content counts, and content hashes. Incremental pi.message_update callbacks are redundant with the final message receipt and may number in the hundreds, so the trusted parent replaces them with one pi.message_updates_coalesced count before Jazz persistence. They do not keep prompt text, provider thinking, final model text, tool arguments, provider bodies, image bytes, source bodies, or arbitrary provider errors. The provider response remains process-local long enough to validate the typed final part, then leaves no raw-content copy in the trace or lifecycle event stream.
Outputs #
The default strict-json output mode requires an agent's final assistant message to contain exactly one text part whose complete contents parse as one JSON value. Surrounding prose, Markdown fences, thinking tags in final text, multiple objects, trailing attempts, truncation, and substring recovery are rejected. The parsed value must be one strict object with:
summary: nonempty bounded text.tags: a bounded list of bounded strings.importance:low,normal, orhigh.confidence: a number from 0 through 1.recommendation: an optional strict{ target, reason, proposedAction }object.
The canonical registry identity for this shape is stream.thought.output.observation@1; its definition hash is recorded in each run context and output. Original model output, repair output, and human correction replacements resolve and validate through that same registry entry rather than duplicating nearby schemas.
conversation-text is a narrow serialization adapter for a standard Pi declaration using telegram-conversation context and stream.thought.derived.message.observation output. The model receives no JSON response-format request. Without proposal capability it must return exactly one nonempty text part of at most 4,096 characters. With the fixed proposal capability in proposals.md, it may instead return that text plus one to three validated proposal calls, or a proposal-only completion that the trusted parent maps to one fixed non-authoritative acknowledgment. A description-bearing /focus turn must contain exactly one focus proposal or it fails closed. The trusted parent treats visible text or the fixed acknowledgment as summary and constructs tags: ["conversation"], importance: "normal", and confidence: 0.5, then validates the resulting object against the same canonical contract. It does not extract substrings, strip fences, recover JSON, or expose proposal arguments as reply prose.
compaction-text is a separate Pi serialization adapter available only to the no-tool Telegram compactor role. The provider receives no JSON response-format request and must return exactly one nonempty plain-text part of at most 24,000 characters with no tool calls. The trusted parent derives the bounded preview and compatibility fields, validates the canonical compaction contract, and binds the resulting semantic object to the exact frozen-prefix plan. Model text never supplies route, coverage, snapshot, run, contract, or hash authority.
The Telegram context builder keeps prior turns in native role-shaped history and supplies the current inbound text only once, as the final user message passed to Pi. Output-format instructions remain system-owned. The runner must not append a synthetic user-role “required final answer” block after Cameron's message; harness prose in that position can be repeated as conversation and displaces the actual latest turn.
Ordinary reply language #
Ordinary conversation-text replies must be concise human conversation. Concision means content-proportional natural sentences without filler; it does not mean fragments, slogan-like line breaks, question-heading repetition, or announcing a style. Replies do not volunteer runtime versions, provider or model names, agent or event ids, tool names, receipt hashes, raw telemetry labels, or other internal terminology unless Cameron explicitly asks for diagnostics. The model should reply directly to the substance of Cameron's message, carry threads forward when useful, and state an opinion when it has one. When Cameron gives a simple behavioral correction, the model applies it instead of defending itself, reciting earlier instructions, asking Cameron to grade the correction, or explaining the correction back.
Unknown fields are rejected. Invalid output produces a classified diagnostic with counts, hashes, contract identity, and bounded issue codes/paths; it terminally fails the run and emits no derived event. Raw malformed text and provider thinking are not persisted. An eligible invalid-output failure may cause one deterministic append-only repair request under repairs.md; infrastructure and authorization failures may not.
The Cloud SDK adapter supports two response modes. strict-json applies the same complete-value parsing and canonical validation as the observer cell. conversation-text accepts one nonempty terminal assistant value, constructs the fixed observation metadata tags: ["conversation"], importance: "normal", and confidence: 0.5, then applies the same canonical contract. It does not persist rejected text, extract substrings, strip fences, or truncate an oversized answer into apparent success.
Agent declarations have an explicit standard or repair role. Repair declarations subscribe only to repair-request events, remain inside the same Pi sandbox/broker boundary, produce inert correction proposals, and cannot recursively repair repair-origin failures.
Completed output is not implicit approval. stream.thought.judgment.training-example@2 records explicit acceptance, rejection, correction, or pairwise preference with a criterion version plus independent qualityEligible and externalExportEligible fields. Quality eligibility means the judgment may inform private evaluation or adaptation; it is not publication or declassification authority. A judgment may name a feedback source event, delivery receipt, and superseded judgment. stream.thought.judgment.training-example.retracted removes an earlier judgment from the rebuildable projection without mutating history. A correction replacement must validate against the judged run's canonical output contract before the judgment is appended.
Target-bound Telegram correction feedback is a deterministic source/projector path, not an agent. stream.thought.source.telegram.correction is excluded from conversation declarations, creates no inference reservation or run, and carries no external-action authority. Its projector may replace only the display field defined by the supported output contract while preserving and recanonicalizing every other structured-output field.
Future Telegram transcript reconstruction uses a validated corrected effective-output projection in place of the original delivered assistant text. The context builder bulk-reads projection ids for the bounded delivered run set, requires projection version 2 with status: corrected and authority: correct, verifies exact run/output/root/contract/privacy lineage, verifies the referenced correction judgment and canonical replacement, and requires that judgment to predate the current inbound event. Invalid, stale, future, unresolved, repair, or contract-mismatched projections remain model-invisible and fall back to the original delivered output. The context manifest records correction authority and exact projection/judgment ids without placing those ids in provider messages. This changes conversational continuity; it does not mutate the original run or delivery evidence.
Telegram declarations may separately cap context.assistantHistoryMaxTurns, and the manifest records candidate, retained, and omitted assistant-turn ids when they do. The ordinary Stream declaration does not use this asymmetric cap. Retaining a large user history while dropping the replies that resolved those requests turns prior prompts into an apparent live backlog and is not coherent conversation context. Stream instead bounds user/assistant history together and relies on the recursive oldest-prefix summary plus exact recent tail. A future declaration may use the cap only when that asymmetry is intentional and tested against its actual history shape.
A valid correction also supplies narrow negative transcript evidence. The compiler compares the original and replacement summaries, extracts only corrected-away short sentence fragments of at most six words, and suppresses a fragment from assistant history only when it recurs in at least one other delivered assistant turn. Matching is exact and case-insensitive except for a corrected-away one-word sentence: after recurrence proves that word is a ritual, the compiler removes both later short assistant sentences beginning with it and remaining whole-word mentions from assistant history. Thus correcting away Slow. suppresses Slow again if you want. and later assistant rationalizations quoting Slow, without touching user text containing slow or slower. This targets repeated corrected-away rituals rather than arbitrary corrected prose. The manifest records the policy version, correction judgment ids, fragment hashes, affected assistant event ids, and removed occurrence count without exposing fragment text. Original runs, outputs, deliveries, corrections, and projections remain immutable. Invalid or divergent correction evidence supplies no suppression authority.
Default export includes only active judgments with both fields true and an entirely public-source source/output chain. Telegram reaction and correction projection writes quality-eligible, externally ineligible judgments. Legacy v1 exportEligible records remain readable but can authorize default export only for entirely public chains. Sensitive/private external eligibility requires explicit authorization at judgment creation and a second explicit private-export gate at dataset creation.
Legacy thoughtstream.training-example.v3 contains validated chosen/rejected structured output, a minimal source classification (type, schema version, source kind, joined privacy), sanitized context policy, trace type/order, output-contract identity, exact primary execution/model-adapter provenance, and for preferences a separate exact compared provenance and compared trace. Pairwise projection independently applies privacy and learned-adapter export policy to both runs. It excludes source payloads, actor/route/external/correlation/idempotency identifiers, event/run/delivery ids, source or trace-content hashes, trace timestamps, arbitrary context fields, prompts, provider content, thinking, tool arguments, quarantine, and legacy raw output. Review-derived examples use thoughtstream.training-example.v4 and may include only the exact preauthorized public review prompt/evidence plus campaign and bounded decision metadata described in review.md. thoughtstream.training-dataset-manifest.v4 contains the dataset content hash, count, kind distribution, participating model sets, mixed example-format counts, and Review campaign identities without judgment or event ids. Legacy v1/v2 judgment events remain readable but project into v3; Review decisions project into v4. Repair runs remain narrower: only active quality-eligible and externally eligible accept or correct judgments may export.
Lifecycle #
Model-backed consumer attempts use:
received → started → completed | failed | blocked | status-unknown | abandoned
blocked is the terminal attempt outcome for a denied pre-dispatch inference reservation. It emits no provider request, no repair request, and no normal or failure notification. accounting.onExhaustion: advance settles consumer progress with the blocked attempt. defer leaves progress unchanged and records a budget-window-derived retry time. Generic unchanged-progress runner failures use the separate declaration-level retry policy; changing budget exhaustion behavior does not silently change provider or sandbox retry timing.
Lifecycle events are authoritative evidence. An optional execution row materializes the current state for inspection; it is not a queue claim. A started attempt that survives a process restart without terminal evidence becomes status-unknown or abandoned according to consumer policy before a new attempt begins. Repair consumers are stricter: an interrupted repair attempt is terminally abandoned and advances request progress, because one repair request authorizes only one proposal generation.
Every started, completed, failed, blocked, or abandoned model run binds executionAdapterRevision and, when selected, the exact modelAdapter identity. The output event and context manifest carry the same learned-adapter identity; replay compares that identity rather than re-resolving a later active release.
completed means output events, terminal evidence, execution state, and consumed source progress settled together at the configured durability tier. A process exiting after model response but before that transaction settles is not completed and may repeat the external call on recovery. The earlier nonterminal attempt remains visible.
Escalation #
An agent may emit an escalation event addressed to another consumer and naming the evidence that justifies it. The addressed consumer's Jazz subscription receives it like any other event. Agents do not call one another recursively inside an execution. Output repair follows the same event-mediated rule and the additional single-proposal constraints in repairs.md.