Runtime control protocol v1 #
Implementation authority: klbr-runtime/src/protocol.rs and the Python runtime, checked
against tests/runtime-fixtures/envelopes.json. This protocol is local to the starter,
not a claimed compatible implementation of Prime or the old klbr IPC protocol.
Transport and bounds #
One dedicated Unix-domain stream socket per worker, located in a private temporary directory created by the host. A random startup token binds the expected subprocess handshake. This is distinct from stdout/stderr, and is not a network-accessible service. Processes run with same-user trust; the token and framing are not an OS sandbox.
Each frame is four unsigned big-endian bytes giving the UTF-8 JSON payload length, followed by that many bytes. Length must be 1..1,048,576. All outer fields are required:
{"version":1,"session_id":"root","generation":"g-1","message":{"kind":"shutdown"}}
Identifiers are 1..128 ASCII letters/digits/underscore/dot/hyphen. Unknown envelope and typed
message fields fail. Version, generation and session mismatches fail. NaN/Infinity are not
valid wire numbers. Python's frame decoder rejects duplicate JSON object keys; Rust derives
reject duplicate typed fields. Arbitrary nested serde_json::Value payload maps are not yet
a duplicate-key equivalence guarantee between parsers; strengthen the shared decoder before
accepting third-party wire producers. This is a trusted local protocol, not a claim of a
fully fuzzed hostile-input parser.
On EOF, bad framing, timed-out/abandoned Rust exchange, or identity mismatch, invalidate the connection/generation. Never attempt to resume a new frame after a partial payload. Output limits: cell source <=262,144 bytes; returned text <=65,536 bytes across <=256 records; error text <=65,536 bytes. The Python runtime currently returns output on completion, not as a continuous stream. Excess is counted and discarded. Raw/native stdout/stderr are continuously drained by the Rust worker into separate bounded 8 KiB diagnostic tails.
Message vocabulary #
hello: token, role (workbench or hooks), revision (null for workbench), and slots
(empty for workbench; exactly attention.plan, delivery.review for this hook worker).
execute: operation id and code, allowed in workbench only. One foreground operation
per worker. Namespace and future flags persist. The final expression has a displayed repr;
_ retains its last non-None value. Evaluation is arbitrary CPython, not sandboxed Python.
invoke: operation id, exact slot, typed domain input, hook role only. A handler is an
importable source entrypoint, not a pickled closure. The workbench is never made to execute
an authoritative send-review hook behind the cell waiting for that review.
host_request: request id, execution_id, and request. Implemented request variants:
Local speech and the turn boundary:
{"method":"local.send","text":"hello"}
{"method":"runtime.wait","seconds":null,"reason":"waiting for input"}
{"method":"hooks.list"}
{"method":"hooks.describe","slot":"delivery.review"}
{"method":"runtime.inspect"}
{"method":"runtime.instructions"}
{"method":"skills.index"}
{"method":"skills.load","id":"workspace"}
Source work and policy:
{"method":"workspace.info"}
{"method":"workspace.run","argv":["cargo","check"],"cwd":".","env":{},"timeout_ms":120000,"max_output_bytes":750000}
{"method":"evidence.resolve","refs":["event:42"]}
{"method":"behavior.activate","path":"/abs/releases/<sha256>","expected_epoch":3}
{"method":"maintenance.originate","purpose":"...","reason":"...","evidence_refs":[]}
{"method":"maintenance.assess","work_item_id":"...","outcome":"no-change"}
{"method":"orientation.stage","candidate_id":"c","text":"...","target":"...","scope":{"kind":"global"}}
{"method":"orientation.current","scope":{"kind":"global"}}
{"method":"orientation.adopt","candidate_id":"c","operation_id":"op-1"}
{"method":"frontier.select"}
Peers and remote namespaces:
{"method":"peer.open","mode":"independent","evidence":[]}
{"method":"peer.send","peer_id":"p","request_id":"req-1","content":"...","evidence":[]}
{"method":"peer.result","request_id":"req-1","wait_ms":0}
{"method":"peer.submit_result","request_id":"req-1","result":{}}
{"method":"peer.cancel","request_id":"req-1"}
{"method":"peer.get","peer_id":"p"}
{"method":"peer.list"}
{"method":"peer.stop","peer_id":"p","drain":true}
{"method":"remote.ensure","target":"build-host"}
{"method":"remote.call","target":"build-host","namespace_id":"n","namespace_generation":"g","operation_id":"op","operation":"execute","payload":{}}
{"method":"remote.cancel","namespace_id":"n","namespace_generation":"g","operation_id":"op"}
{"method":"remote.close","namespace_id":"n","namespace_generation":"g"}
peer.open also accepts agents.open, and peer.send also accepts peer.message.
host_reply repeats request/execution identity and contains either
{"kind":"ok","value":...} or {"kind":"error","code":"...","message":"..."}.
The worker host enforces a maximum of 128 requests per foreground cell. Pending replies
are correlated by identity, not arrival order. The Python reader keeps running while a cell
awaits an RPC. Late replies after cancellation are discarded rather than satisfying another
cell's request.
completed repeats the operation id with one of:
{"kind":"cell","status":"ok","output":[{"stream":"result","text":"42"}],"truncated_bytes":0,"error":null}
{"kind":"hook","value":{"action":"allow"}}
{"kind":"failed","error":"worker-level failure description"}
Cell statuses: ok, error, yielded, interrupted. Output streams: stdout, stderr,
result. Worker death has no invented completion; the host creates a failed/interrupted
record from its own observation. A worker cannot claim a committed yield without the host
having accepted the wait.
interrupt targets one operation ID. Python supports cooperative task cancellation; the
current Rust management API conservatively terminates the generation instead. shutdown
requests clean worker exit. Rust retains the final hard-stop authority when code blocks or
catches cancellation. Only the direct child is currently managed, not all descendants.
The instruction header and its revisions #
The head of a request is the text this transcript pinned, not the live contents of the runtime's
instruction sources. prompt.header records the pin; prompt.revision records a change:
{"type":"header_pin","digest":"<sha256>","text":"...","reason":"session_start|fold|fold_covered_revision","pinned_at_ms":0}
{"type":"header_revision","digest":"<sha256 of the current text>","change":{"form":"diff"|"full", ...},"sources":[{"name":"...","digest":"...","text":"..."}],"actor":{"origin":"session|host","execution_id":"...","step":1,"path":"...","reason":"..."},"observed_at_ms":0}
Three rules, in the order they apply:
- No pin yet: the transcript pins the current text (
session_start) and runs under it. - The text changed and this transcript has (or can take) the change as an observation: the head
keeps its bytes and the observation renders at the sequence where it was noticed. A change a
session makes through a tool call is attributed to that execution; a change with nothing in the
transcript to attribute it to is reported as
host, with what is known and no invented actor. The observation is written once and replayed like any other record. - A fold moves the head. Installing a frontier rebuilds the prefix anyway, so the current text
costs nothing extra there, and observations the frontier covers can no longer be rendered -
a head whose only support was one of those must move (
fold_covered_revision) rather than silently drop an instruction the model is running under.
A revision is carried as verified patches against the newest text the transcript can reconstruct, and as the whole text when a patch does not reproduce that text exactly:
{"digest":"<sha256 of the result>","change":{"form":"diff","base_digest":"<sha256>","patches":[
{"source":"contract","start_line":99,"old_lines":["\n"],"new_lines":["A revision ...\n","\n"]}]},
"sources":[{"name":"contract","digest":"...","text":"..."}],"actor":{"origin":"host","reason":"..."}}
Verification is not optional: a patch that does not reconstruct its text is replaced by the
whole text, so a reader never reconstructs an instruction text nobody wrote. runtime.instructions
resolves the chain on the host for a reader that wants the exact current text instead of applying
patches by hand; a link that does not reconstruct stops the chain at the last text that did.
Patches are per source while the two source sets line up by name. When they do not - a head pinned
before sources were recorded, a source that appeared or vanished - the change is patched over the
joined texts under the source name instructions instead. A rename then reads as the edit it is
rather than as a rewrite of everything under both names, and a head whose parts are unknown is one
honestly anonymous source rather than a claim about parts nobody recorded.
The rendered observation names the mechanism, because two instruction texts in one request are otherwise ambiguous:
[2026-09-13T19:54:53Z · instructions changed for this runtime · changed outside this transcript]
the header of this transcript is the revision it started under; the text below is current from here on.
<current instruction text or the verified patch that reaches it>
Why it is shaped this way: rewriting a head invalidates every block after it, and it rewrites what the transcript claims to have been running under all along. Appending costs the observation's own tokens. A pin move costs the prefix, which is why it happens at folds and nowhere else.
What the request costs: cache reporting #
Usage fields are read from every shape the routes we talk to actually send, and never derived:
prompt_tokens_details.cached_tokens and cache_write_tokens (also under
input_tokens_details), or DeepSeek's prompt_cache_hit_tokens / prompt_cache_miss_tokens.
cached_prompt_tokens, cache_write_tokens and uncached_prompt_tokens stay absent when the
route said nothing about them: an absent measurement and a reported zero are different records,
and prompt_tokens - cached_prompt_tokens is our arithmetic, not the route's.
Which cache fields a request may carry is the route's declared capability (see DOGFOOD.md);
prompt_cache_retention and prompt_cache_options are stripped from route options and only
re-added by a profile that says the route accepts them.
Framing what the model reads #
An input is assembled into the request with a frame in front of the sender's own words, derived from the durable event on every assembly:
[2026-09-13T18:59:00Z · operator]
do the thing
The label is the event's source, or peer request <id> from <session> when the payload carries a
peer request identity - the recipient has to be able to answer it, and the model never sees the
sibling fields. Inputs whose event has no usable time say time unknown rather than rendering
epoch zero as 1970. The stored content is exactly what arrived, so a frame is never part of a
durable record and never has to be parsed back out of one.
Tool results carry the same kind of line: the instant the observation was made, plus how long the cell ran when one ran at all.
[2026-09-13T21:17:22Z · ran 420ms]
"It took 420ms" is a fact about the past and stays true on replay. A relative rendering such as
"four minutes ago" is deliberately absent from stored content: the same block is replayed later
and would then assert something false. Currency is stated by the host clock instead -
runtime.inspect returns now and now_ms - so a reader compares two instants rather than
trusting a stale phrase. now uses the same Z shape as these stamps.
Scope, waiting and receipts #
Host effects require the current live execution plus matching generation. Background tasks inherit an origin for diagnostics, but that scope expires when the foreground cell ends. After terminal wait, catching the interpreter sentinel does not restore scope authority. The host also rejects further requests from a yielded execution.
Wait seconds are null or finite >0 and <=600. The host records the deadline/state and checks
for already-pending input in one transaction. No suspended Python continuation is persisted.
Wakeups start a future model/cell boundary; code below await runtime.wait(...) does not
run normally. The caller must not mistake ordinary asyncio.sleep for this durable wait.
Local send receipt is effect_id, status (confirmed or held), event_id (nullable),
and reason (nullable). Confirmed means a durable operator-lane event exists. Held is not
approval, queue-success or proof anyone received the message. There is no Discord transport
in this slice. A retry with the same execution/request identity returns the existing receipt;
a deliberately new request can send identical text again.
Skills #
A skill is a directory holding SKILL.md whose first non-empty line is # <id>, where <id>
equals the directory name. There is no other field: a header beside the body would be a second,
staler definition of it. The body is the trail to the state the id names - the moves, the
receipts, the refusals - not a signature list, because the signatures are already in the
installed code.
Two roots form one namespace and their order is not precedence. Builtin skills come from
KLBR_SKILLS_DIR, defaulting to <workspace>/skills, and are allowed to be absent. Managed
skills come from <data>/skills, which init/serve create so an agent can write one without
inventing the directory. An id defined in both roots is refused rather than resolved, as is a
title that disagrees with its directory, a non-UTF-8 or oversize file, or a symlinked directory.
Refused ids are reported with their reason by skills.index, and by the assembled contract, so
an agent that writes a skill learns why it was not offered.
The model-visible index is ids only. The contract message carries the ids, and the host
skills.load request carries one body; a body that differs from the digest the host just indexed
is refused rather than half-read. Reading grants nothing: a skill cannot widen the delivery review
and cannot outrank the contract.
A local spawn also materializes the whole validated set beside the control socket and passes it
as --skills <path>. The worker reads it before it connects, so the payload file cannot outlive
the handshake, makes the ids importable as klbr.skills.<id> (a synthetic module whose
documentation is the body), and echoes the set's revision in hello. That echo is what
kernel.reset records, so the durable log names the instruction set a generation actually
loaded. A body written during a turn is therefore offered on the next generation, not mid-turn.
Implemented decision slots #
attention.plan input: positive event_id, host-classified source (operator, discord.dm,
discord.mention, discord.ambient), nonnegative receipt timestamp. Result: wake with priority,
defer with bounded seconds, or ignore with reason. Operator attention cannot be suppressed
or forged from an external source by returning a priority. This is a policy API awaiting
authenticated ingress/model-coordinator integration.
delivery.review input: effect_id, destination=operator, bounded message text. Result:
allow or hold with a reason. Any invalid/unavailable required review holds the send.
No host-effect scope is given to a hook. Additional domain hooks should be added alongside
the owning implementation, not as no-op placeholders.
The host's two-second invocation budget includes queueing. One bounded actor serves a hook worker; a hung callback can delay other calls until timeout invalidates that generation. There is no separate per-handler scheduling or automatic hook restart yet.
Source revisions and activation #
A manifest contains version 1 and exactly the two implemented entrypoints:
{"version":1,"hooks":{"attention.plan":"klbr_hooks.defaults:attention","delivery.review":"klbr_hooks.defaults:delivery"}}
Handlers receive (request, context) and return the corresponding SDK decision dataclass,
synchronously or asynchronously. Startup validates the import/call signature, not semantic
behavior. Imports themselves are trusted executable code and can fail or block.
Source identity hashes sorted POSIX relative filenames and exact bytes. For each file feed
SHA-256: u64BE name-byte-length, UTF-8 name, u64BE content-length, content. Source snapshots
allow .py, .json, .md, maximum 128 files/2 MiB, no symlinks. Run source tests with -B
or put caches outside the draft. The shared behavior.json fixture fixes the initial digest.
Database epoch is activation authority; a path/symlink is not a second active pointer. Candidate handshake precedes the compare-and-swap switch. Each executing cell pins its policy object, so later hooks in that cell use the same revision even during activation. A hook-only change does not replace the workbench. A broader environment change is not yet implemented.