Clayterm Renderer Specification #
Version: 0.1 (draft) Status: Current-state specification. Normative for the rendering contract. Descriptive for settling surfaces.
1. Purpose #
Clayterm is a terminal rendering engine. It accepts a declarative description of a terminal UI layout, performs layout computation and cell-level diffing internally, and returns ANSI escape byte sequences suitable for direct write to a terminal output stream.
This specification defines Clayterm's current-state rendering contract: its architectural model, its invariants, its stable public API surface, and its intentional boundaries. It is written to allow future feature work to extend the project without destabilizing the core.
This specification does not attempt to define areas of Clayterm that are still settling. Where the project has working but evolving surfaces — including the pointer event model and certain wrapper types — those are described in Section 12 as current implementation rather than normative contract.
Input parsing is specified separately in the Clayterm Input Specification.
Transitions are specified separately in the Clayterm Transitions Specification.
2. Scope #
In scope (normative) #
- The rendering pipeline and its architectural commitments
- The frame-snapshot rendering model
- The stable public rendering API
- The directive model and core helpers
- Element identity and frame semantics
- Boundary responsibilities (what Clayterm owns and what it does not)
- Capability-gated emission (Section 7.6; the capability layer itself is specified in the Terminfo Specification)
In scope (non-normative, descriptive) #
- Current implementation surfaces that are settling but not yet stable enough to freeze (Section 12)
- Implementation notes that aid understanding but do not define contract (Section 13)
Out of scope #
- Internal C code organization, function names, or file structure
- WASM memory layout or compilation details beyond behavioral requirements
- Performance targets or benchmark methodology
- Packaging, CI, or distribution workflow details
- Higher-level UI framework concerns (e.g., component lifecycle, reconciliation)
- Demo applications
- The crankterm project or any specific framework built on Clayterm
- Input parsing (see Clayterm Input Specification)
- Terminfo parsing and capability probing (see Terminfo Specification); this specification consumes the capability struct, it does not define it
3. Terminology #
Frame. A single, complete rendering pass. Each frame begins with the caller providing directives and ends with the renderer returning ANSI bytes. Frames are independent; the renderer carries no UI tree state between them.
Directive (op). A plain object that declares one element of the UI tree for a single frame. Directives are typed by an identifier field and carry layout, styling, and content properties. The set of directives for a frame is ordered and forms an implicit tree via open/close pairing.
Directive array. An ordered array of directives constituting a complete frame description. The array is the input to the rendering transaction.
Render transaction. The process of accepting a directive array, performing layout, walking render commands, diffing against the previous frame's cell buffer, and producing ANSI byte output. A render transaction is a single, synchronous operation from the caller's perspective.
ANSI bytes. A byte sequence of UTF-8–encoded ANSI escape codes and text content that, when written to a terminal file descriptor, produces the visual output described by the frame's directives. ANSI bytes include cursor-positioning sequences, SGR (Select Graphic Rendition) attribute sequences, and UTF-8 text.
Renderer core. The WASM module and its TS entry points that together implement the render transaction. The renderer core owns layout computation, render-command walking, cell-buffer diffing, and ANSI byte generation.
Caller. Any code that invokes Clayterm's public API to produce terminal output. The caller owns terminal setup, IO, input handling, and application lifecycle.
Higher-level framework. A component model, reconciler, or application framework built on top of Clayterm. Examples include crankterm. Clayterm has no dependency on any higher-level framework, and this specification does not constrain their design.
Term. An instance of the Clayterm renderer, bound to specific terminal dimensions. A Term is the object through which the caller performs render transactions.
4. Architectural Model #
This section is normative.
4.1 Pipeline #
Clayterm implements a rendering pipeline with the following stages:
-
Directive acceptance. The caller provides a complete directive array representing the desired UI state for a single frame.
-
Transfer. The renderer transfers the frame description into the WASM module. The transfer mechanism is an implementation detail. The normative requirement is that the transfer occurs as part of a single render transaction; the caller does not interact with the transfer mechanism directly.
-
Render transaction. The WASM module processes the frame description. Internally, it drives a layout engine to compute element positions and sizes, walks the resulting render commands to populate a cell buffer, and diffs the cell buffer against the previous frame.
-
Output generation. For each cell that differs from the previous frame, the renderer emits ANSI escape sequences (cursor positioning, color attributes, and text) into an output buffer.
-
Output retrieval. The caller reads the ANSI byte output.
4.2 Single-transaction rendering #
A frame MUST be rendered in a single transaction that crosses the TS→WASM boundary once. The caller provides the complete directive array, invokes the render transaction, and reads the output. There are no intermediate callbacks, yields, or partial results.
4.3 Frame-snapshot model #
Each render transaction operates on a complete, self-contained snapshot of the UI. The renderer MUST NOT maintain an internal component tree or UI state across frames. The only state the renderer retains between frames is the cell buffer used for diffing, which is an implementation detail of output minimization and not observable to the caller except through reduced output size.
4.4 Double-buffered diffing #
The renderer maintains two cell buffers: a front buffer (the previously rendered frame) and a back buffer (the frame being rendered). After populating the back buffer from the current frame's render commands, the renderer compares it against the front buffer and emits ANSI bytes only for cells that differ. Changed cells are then copied from the back buffer to the front buffer so that both buffers are identical at the end of the transaction. This mechanism is internal to the renderer and not directly observable to the caller.
5. Contract Layer Boundary #
This section is normative.
This specification defines the architectural rendering contract: the commitments that make Clayterm what it is and that callers and framework authors can depend on.
This specification does not define the following as normative:
-
The internal transfer encoding. The mechanism by which directives are serialized for the WASM module — its byte format, opcode structure, and field encoding — is an implementation detail. The normative commitment is that the transfer happens within a single render transaction; the encoding is described in Section 12.1 as current implementation surface.
-
Validation or error semantics. How the renderer responds to invalid input (malformed directive arrays, unbalanced open/close pairs) is not yet specified as contract. Section 9.1 defines what constitutes valid input. Behavior for invalid input is currently unspecified.
-
The complete set of directive properties. The existence of the core directive constructors (
open,close,text) and the core sizing helpers (grow,fixed,fit) is normative. The full set of properties accepted by these constructors — which layout fields, which styling options, which configuration groups are available — is current implementation surface described in Section 12.2. New property groups have been added over time and more may follow. -
The return type wrapper of
render(). The commitment thatrender()produces ANSI bytes accessible as aUint8Arrayis normative. The wrapper type around those bytes is current implementation surface described in Section 12.3.
Future readers should not treat current implementation surface as identical to the contract boundary.
6. Core Invariants #
This section is normative.
INV-1. Zero IO. The renderer MUST NOT perform any terminal input or output. It MUST NOT write to stdout, read from stdin, open file descriptors, or interact with the terminal device. The renderer produces bytes; the caller writes them.
INV-2. Single transaction per frame. Each frame MUST be rendered in a single transaction that crosses the TS→WASM boundary once. The caller provides the complete frame as a directive array and receives ANSI bytes in return.
INV-3. Frame-snapshot independence. The renderer MUST NOT require the caller
to maintain or provide state across frames beyond calling render() on the same
Term instance. Each directive array fully describes its frame.
INV-4. ANSI byte output. The output of a render transaction MUST be a byte sequence of valid UTF-8–encoded ANSI escape codes that is directly writable to a terminal output stream without further transformation or encoding.
INV-5. Layout/render/diff ownership. The renderer owns the layout computation, render-command walk, cell-buffer diffing, and ANSI byte generation stages. The caller MUST NOT need to perform any of these operations.
INV-6. Internal lifecycle symmetry. The renderer's internal layout lifecycle (begin-layout and end-layout calls to the underlying layout engine) MUST be symmetric: both calls occur within the same render transaction, in the same function scope.
INV-7. Separation of concerns. The rendering concern and the input-parsing
concern MUST remain independent. Neither MUST depend on the other's state,
types, or API surface. They MAY share a compiled WASM binary for loading
efficiency, but this is an implementation convenience, not an architectural
coupling. Both MAY consume the shared capability layer defined in the
Terminfo Specification; the renderer reads capabilities and
the input parser surfaces probe responses as CapabilityEvent values, but
neither observes the other beyond the capability facts it carries.
7. Rendering Contract #
This section is normative.
7.1 Inputs #
The rendering transaction accepts:
- A directive array: an ordered array of directive objects constituting a complete frame. The array MUST contain balanced open/close pairs forming a valid tree structure.
The directive array is the sole required input to a render transaction.
7.2 Rendering transaction #
When the caller invokes a render transaction:
- The renderer accepts the directive array and transfers the frame description into the WASM module.
- The WASM module processes the frame: it computes layout, walks render commands, populates the cell buffer, diffs against the previous frame, and writes ANSI bytes for changed cells.
- Control returns to the caller with the ANSI byte output available.
The render transaction is synchronous from the caller's perspective once invoked. It MUST NOT yield, suspend, or require callbacks during execution.
7.3 Output #
The render transaction produces ANSI bytes as a Uint8Array. These bytes:
- MUST be valid UTF-8
- MUST consist of ANSI escape sequences (CSI, SGR) and text content
- MUST be directly writable to a terminal file descriptor to produce the described visual output
- In cursor update mode, MUST represent only the cells that changed since the previous frame (on a Term instance that has rendered at least one prior frame)
- In line mode, MUST represent all cells in the frame as newline-separated rows
The output reflects the complete visual state of the frame. The caller SHOULD write the output to the terminal without modification.
The output Uint8Array may be a view over renderer-owned memory. It is valid
until the next render() or update() call on the same Term instance, at which
point the buffer may be reused or detached (§7.7). Callers who need to retain
the output beyond that point MUST copy it.
7.4 Lifecycle #
A Term instance is created for specific terminal dimensions. The caller provides width and height at creation time.
Creation of a Term is asynchronous because it may involve WASM module preparation. A Term instance MAY be used for any number of render transactions. The Term retains its cell buffers across frames for diffing purposes.
A Term instance's dimensions MAY be changed after creation through the update transaction (§7.7). A resized Term remains valid: it continues to accept render transactions at the new dimensions. Creating a new Term is NOT required to handle terminal resize.
7.5 Clip semantics #
An element whose props include a clip group declares a clip region: a
rectangular bound on the cells its descendants are permitted to write. Cells
produced by descendants that fall outside this region MUST be suppressed from
the output. The clip region is determined by the element's computed layout box
and the axes selected by the clip group (horizontal, vertical, or both).
Clip regions stack. When clip elements nest:
- The effective clip region of an element MUST be the intersection of its own declared region with the effective clip region of its nearest clipping ancestor, if any.
- When the renderer finishes processing a clip element's subtree, it MUST restore the effective clip region of that element's clipping ancestor. Later siblings drawn within an ancestor clip MUST therefore remain bounded by that ancestor.
- A
clipelement whose declared region is fully outside its ancestor's effective region produces an empty effective region; descendants of that element MUST NOT contribute any cells to the output.
The renderer MAY impose an implementation-defined limit on the depth of clip regions it can track. The limit itself is not normatively bounded. When a frame nests clip regions more deeply than the renderer can track:
- All clip regions whose entry the renderer successfully tracked MUST continue to be honored for the remainder of the frame, including for siblings drawn after the over-deep subtree closes. The renderer MUST maintain push/pop symmetry so that exiting an untracked clip does not disturb any ancestor's effective region.
- Content drawn inside an untracked clip region MUST remain bounded by the deepest successfully-tracked ancestor clip region. The untracked region's own additional restriction MAY be lost.
- The renderer MUST surface the condition via the render result's error channel (see §12.3) before returning, so the caller can detect that some clipping was not applied.
7.6 Hardware cursor visibility and positioning #
A text() directive MAY declare a caret property whose value is a
non-negative integer code-point offset into the directive's content. 0 means
"before the first code point"; [...content].length means "after the last."
Offsets in between sit immediately before the cell where the corresponding code
point would render.
The cell where a declared caret sits is the cell at which the code point at the
caret's offset would be drawn given the layout engine's text wrapping. For an
offset N:
- If
N < [...content].length, the caret's cell is the display position of theN-th code point (zero-indexed) within the rendered text. - If
N == [...content].length, the caret's cell is one display position past the last rendered code point: on the same wrapped line if there is room, or at the start of the next line if the layout wraps at the end. - If
N > [...content].lengthorN < 0, behavior is unspecified; callers must keep offsets within bounds.
When the layout engine consumes whitespace at a wrap boundary — dropping it from the rendered text so that neither wrapped line contains a display cell for those code points — the caret has no display position of its own for any offset that falls in that dropped run. In that case the caret's cell is the origin of the following wrapped line (the display position of the first code point on that line). This rule keeps the caret on-screen at the point where subsequent input will land, rather than orphaning it past the end of the previous line.
When content is empty, the only in-range offset is 0. In that case the
caret's cell is the text element's origin — the cell at which the first code
point of content would be drawn if content were a single space. This is a
rendering commitment implied by the presence of a caret declaration, and the
renderer must synthesize whatever minimal geometry is required to honor it. How
this outcome is achieved is implementation-defined.
Display position accounts for code points wider than one cell (CJK, fullwidth forms, some emoji): such code points occupy two cells, and the caret sits at the first cell of the pair. Cell widths are determined by the same measurement the renderer uses to lay out the text itself.
The renderer manages the terminal's hardware cursor based on the presence of
caret: declarations:
-
When the current frame contains one or more
caret:declarations, the renderedoutputMUST, when written to the terminal, leave the hardware cursor visible at the cell where the first declared caret (in directive order) sits. -
When the current frame contains no
caret:declarations:- If the previous frame contained one, the rendered
outputMUST leave the hardware cursor hidden. - Otherwise, the rendered
outputMUST NOT include cursor-positioning or cursor-visibility bytes; the caller's cursor state is preserved.
- If the previous frame contained one, the rendered
The byte-level path the renderer takes to satisfy these outcomes is implementation-defined. The renderer MAY hide the cursor before cell writes to prevent flicker, MAY omit such a hide for in-place edits where flicker is acceptable, MAY emit only a CUP when the caret has moved within an already-visible state, MAY use whatever escape-sequence path it judges appropriate. Only the post-frame cursor state above is normative.
If more than one caret: declaration is present in a frame, the renderer SHOULD
use the first in directive order for the hardware cursor. Behavior for
additional declarations is intentionally unspecified, leaving room for a future
multi-cursor extension without breaking this contract.
This responsibility is limited to the hardware cursor's position and visibility. Cursor shape and blink rate remain caller-managed.
7.7 Update transaction #
The update transaction changes a Term instance's dimensions or capability state in place. Like the render transaction (§7.2), it is synchronous: it MUST NOT yield, suspend, or require callbacks during execution.
Inputs. The update transaction accepts an ordered array of InputEvent
values (see Input Specification §5), discriminated by type:
- A
ResizeEvent({ type: "resize"; width: number; height: number }) — a resize to the given character-cell dimensions. Both MUST be positive integers; the transaction MUST throw otherwise. - A
CapabilityEvent(see Terminfo Specification §6.3) — a capability value delivered by the input parser from a probe response. - Any other
InputEvent— a no-op step. It changes no state and contributes no bytes.
The Term folds each event in array order. The returned bytes are the concatenation of each fold's output. An empty array is a no-op.
A resize event whose target dimensions equal the Term's current dimensions MUST be a no-op for that step.
Resize semantics. A non-no-op resize step:
- Reallocates renderer state for the new dimensions within the Term's existing WASM instance and linear memory, growing the memory if required. The renderer MUST NOT create a new WASM module or instance.
- Discards all diff state. The next render transaction MUST emit the frame as a complete redraw, so that no cell laid out at the previous dimensions can linger on screen.
- Discards pointer interaction state (hover and press tracking), since cell coordinates under the pointer have changed. No synthetic pointer events are emitted.
Capability semantics. A CapabilityEvent step folds the event's key and
value into the Term's private RuntimeCapabilities. The foundation update
transaction emits bytes or alters renderer output as a consequence of a
capability event only where a capability section of this specification says so
(§7.9).
Return value. update() returns a Uint8Array of bytes to write to the
terminal immediately. An empty array is valid when the update changes no
rendered state (TINV-5). The caller writes the bytes to its output stream before
the next render.
Memory. WASM linear memory can only grow. An update to smaller dimensions retains the high-water-mark allocation; memory is not reclaimed until the Term itself is discarded. This is accepted behavior, not a defect.
Output invalidation. Growing linear memory detaches existing buffer views.
Output Uint8Arrays returned by render transactions prior to a resize update
MUST NOT be used after it (this strengthens the validity window in §7.3: output
is valid until the next render() or update() call).
7.8 Capability consumption #
The renderer consumes the read-only capability snapshot maintained by
Term.update(). A protocol that changes emitted bytes MUST define its gating
capability, output, and invalidation rules in its own section of this
specification. Pointer shape (§7.9) is the first.
Color encoding, synchronized-output wrapping, Kitty keyboard mode setup, and Kitty graphics emission are deferred to their respective follow-up PRs. The renderer continues to emit its existing hardcoded ANSI output until each is specified.
7.9 Pointer shape #
An open() directive MAY declare a pointerShape property naming the mouse
pointer shape to show while the pointer is over the element. Values are the
kitty pointer shape names (the CSS cursor keywords): default, none,
context-menu, help, pointer, progress, wait, cell, crosshair,
text, vertical-text, alias, copy, move, no-drop, not-allowed,
grab, grabbing, e-resize, n-resize, ne-resize, nw-resize,
s-resize, se-resize, sw-resize, w-resize, ew-resize, ns-resize,
nesw-resize, nwse-resize, zoom-in, zoom-out. Any other value MUST be
treated as absent and MUST NOT reach the output. The property does not affect
layout or cell output.
The gate is RuntimeCapabilities.pointerShape (see
Terminfo Specification §6.3). While it is false, the
renderer MUST NOT read pointerShape properties and MUST NOT emit OSC 22.
Resolved shape. While the capability is true, each render resolves one
shape:
defaultwhen the render options carry nopointer.- Otherwise, the
pointerShapeof the last declaring element in hit-test order (the order behindpointerenterevents, §12.4), ordefaultwhen none declares one. Ancestors precede descendants, so the innermost element wins. A floating element in"capture"mode hides the elements beneath it; one in"passthrough"mode lets them decide.
Elements inside a snapshot() participate as direct directives would.
Resolution and output are identical in line mode (§8.2.2).
Output. The Term tracks the shape it last emitted, starting at default.
When the resolved shape differs, the render MUST append ESC ] 22 ; <shape> ST
to output after the frame's cell bytes and record the new shape. Otherwise it
MUST NOT emit OSC 22.
Restore. OSC 22 set cannot pop, so the Term restores by setting default:
- A render that resolves
defaultemits it per Output above. - An update step that changes
pointerShapefromtruetofalsewhile the emitted shape is notdefaultMUST returnESC ] 22 ; default STand recorddefault. - A resize step leaves the emitted shape unchanged.
To restore before exit, the caller renders a frame without pointer, or passes
{ type: "capability", key: "pointer-shape", value: false } to update(), and
writes the result.
8. Public Rendering API #
This section is normative. Only items with high confidence of stability are included. See Section 5 for what this section does and does not freeze.
8.1 Term creation #
createTerm(options: {
width: number;
height: number;
terminfo?: TerminalInfo;
}): Promise<Term>
Creates a new Term instance bound to the specified terminal dimensions. The
returned promise resolves when the renderer is ready. The width and height
parameters specify the terminal dimensions in character cells.
The optional terminfo value (from detectTerminal(); see
Terminfo Specification §10.1) initializes the Term's private
RuntimeCapabilities from the static Capabilities it carries. When omitted,
the Term uses the §7.1 baseline.
8.2 Render invocation #
term.render(ops: Op[], options?: RenderOptions): <result containing ANSI bytes as Uint8Array>
Accepts an ordered array of directive objects and performs a render transaction
as defined in Section 7. Returns the ANSI byte output as a Uint8Array.
The optional options parameter controls the rendering mode. See Section 8.2.1
and 8.2.2 for the two available modes.
The return type is specified here only to the extent that the ANSI bytes MUST be
accessible as a Uint8Array. The precise shape of the return value — whether it
is a bare Uint8Array, a wrapper object, or a structure carrying additional
data — is part of the current implementation surface described in Section 12.3
and is not locked down by this specification.
8.2.1 Cursor update mode (default) #
When mode is omitted, the renderer operates in cursor update mode. Output
consists of ANSI bytes with absolute CUP (\x1b[row;colH) cursor positioning
for each changed cell. Only cells that differ from the previous frame are
emitted, making this efficient for full-screen UIs where most of the screen is
static between frames.
The optional row parameter specifies a 1-based row offset for CUP positioning.
This allows the caller to render into a region of the terminal starting at a row
other than the top. The offset is applied to all emitted cursor positions. When
omitted, it defaults to 1.
8.2.2 Line mode #
When mode is "line", the renderer emits all cells as newline-separated rows
without CUP positioning. Every cell is written regardless of whether it changed
since the previous frame. The front buffer is updated so that a subsequent
cursor update mode render can diff efficiently.
Line mode is intended for inline region rendering where the caller manages cursor positioning externally and the output must work in pipes or non-alternate screen contexts.
8.3 Directive constructors #
Directives are created using constructor functions that return plain objects. The caller assembles these into an array. This pattern — functions returning plain directives, composed into arrays — is normative. A builder, fluent, or mutation-based API is explicitly rejected.
8.3.1 open #
open(id: string, props?): OpenElement
Creates an element-open directive. The id parameter provides an identity for
the element within the frame, used by the underlying layout engine for element
tracking and hit-testing. IDs MUST be unique within a frame; passing duplicate
IDs is undefined behavior. The optional props parameter carries configuration
for layout, styling, and behavior.
Elements opened with open() MUST be closed with a corresponding close()
directive later in the same directive array.
The set of properties accepted by props is part of the current implementation
surface described in Section 12.2. This specification defines the existence and
signature of open() normatively but does not freeze the complete property
surface, which has been extended incrementally and may continue to grow.
8.3.2 close #
close(): CloseElement
Creates an element-close directive. Each close() MUST correspond to a
preceding open().
8.3.3 text #
text(content: string, props?): Text
Creates a text directive. The content parameter provides the text string to
render. The optional props parameter carries text styling configuration.
Text directives MUST appear between a matching open/close pair.
When props.bg is provided, the renderer MUST apply that background color only
to cells occupied by glyphs emitted by that text directive. It MUST NOT fill
trailing cells or other cells in the text element's bounding rectangle that are
not occupied by emitted glyphs. When props.bg is omitted, text rendering MUST
NOT override the background already present in each glyph cell; element
backgrounds established by open({ bg }) remain in effect, and the terminal
default remains in effect where no element background applies.
Grapheme cluster preservation. When content contains grapheme clusters — a
base codepoint followed by one or more combining marks (Unicode codepoints with
wcwidth ≤ 0) — the renderer MUST preserve and emit the full cluster. Combining
marks MUST NOT be silently dropped. Combining marks do not advance the cursor
position; they are attached to the preceding base codepoint's cell and emitted
immediately after the base codepoint's bytes in the output stream. A cell
occupied by a grapheme cluster with combining marks MUST be treated as changed
(and thus emitted) when any mark in the cluster changes between frames.
The set of styling properties accepted by props is part of the current
implementation surface and may be extended.
8.3.4 snapshot #
snapshot(ops: Op[]): Op
Creates a snapshot by pre-packing the given directive array into its transfer
encoding. The returned value is an Op and can appear anywhere in a directive
array where the original ops would have appeared. The internal representation is
opaque.
When the renderer encounters a snapshot during transfer, it copies the
pre-packed bytes directly into the command buffer without re-encoding. The
snapshot's ops MUST be structurally balanced (every open matched by a
close).
Snapshots enable higher-level frameworks to implement dirty tracking: a component whose inputs have not changed can reuse a previously created snapshot, avoiding the cost of re-packing its subtree each frame.
8.4 Sizing helpers #
These functions produce sizing-axis values for use in element layout configuration:
grow(): SizingAxis
The element expands to fill available space in the parent along this axis.
fixed(value: number): SizingAxis
The element has a fixed size of value cells along this axis.
fit(min?: number, max?: number): SizingAxis
The element sizes to fit its content, optionally constrained by minimum and maximum bounds.
8.5 Color helper #
rgba(r: number, g: number, b: number, a?: number): number
Packs color channel values (each 0–255) into a single 32-bit integer in ARGB format. Alpha defaults to 255 (fully opaque). The returned value is used wherever the directive model expects a color.
8.6 Term update #
term.update(events: readonly InputEvent[]): Uint8Array
Performs an update transaction as defined in §7.7. update() is the universal
sink for both resize and capability change.
Events. A resize step is a ResizeEvent
({ type: "resize", width, height }). A capability step is any
CapabilityEvent value (see Terminfo Specification §6.3).
Every other InputEvent is accepted and ignored, so the full events array
from input.scan() can be passed without filtering. Events are folded in array
order.
Return value. update() always returns a Uint8Array. Write it to the
terminal immediately when non-empty. Do not wait for the next render(). An
empty array means the update changed no rendered state.
Host usage. Events from input.scan() pass through directly. Resizes
observed outside the input stream (e.g. SIGWINCH) are passed as a constructed
ResizeEvent:
const { events } = input.scan(bytes);
const out = term.update(events);
if (out.length) stdout.write(out);
// on SIGWINCH
term.update([{ type: "resize", width: cols(), height: rows() }]);
A resize to the Term's current dimensions is a no-op for that step.
9. Directive Model #
This section is normative.
9.1 Directive-array pattern #
The rendering input is an ordered array of directive objects. Each directive is a plain JavaScript/TypeScript object created by a constructor function (Section 8.3). Directives are not classes, do not carry methods, and do not participate in a prototype chain. They MAY be spread, composed, stored, or inspected as ordinary objects.
The array is processed in order. Open and close directives form an implicit tree. The renderer processes them sequentially.
A directive array with unbalanced open/close pairs, or with close directives that do not match a preceding open, is invalid input. Callers SHOULD validate directive arrays before rendering. The renderer's behavior when given an invalid directive array is unspecified by this specification.
A snapshot is semantically equivalent to splicing its source ops into the array at the snapshot's position. The renderer MUST produce identical layout and output regardless of whether ops are provided directly or via a snapshot.
9.2 Transfer to the WASM module #
As part of the render transaction, the directive array is transferred into a form that the WASM module can process. This transfer is handled internally by the renderer and is not an operation the caller performs or observes. The transfer mechanism is an implementation detail described in Section 12.1.
If a frame exceeds transfer-buffer capacity while packing string content, the
renderer MUST throw a descriptive RangeError that identifies the condition as
a transfer-buffer, frame-capacity, or packing overflow. The renderer MUST NOT
expose only the raw host-level TypedArray message "offset is out of bounds"
for this condition. The error message SHOULD direct callers to render a smaller
visible slice or reduce frame content.
9.3 Directive identity #
Each element directive carries an id provided by the caller via open().
Element IDs MUST be unique within a frame. The renderer uses the ID directly as
the element's identity for the layout engine. Passing duplicate IDs within a
single frame is undefined behavior.
10. Identity and Frame Semantics #
This section is normative.
10.1 Frame completeness #
A directive array provided to render() MUST represent a complete frame. The
renderer does not support incremental updates, partial frames, or delta
descriptions. Every frame fully specifies the desired UI state.
10.2 Directive ordering #
Directives MUST be provided in depth-first tree order. An open() directive
begins an element; its children (including nested open/close pairs and text
directives) follow in order; a close() directive ends the element. The
renderer processes directives in the order they appear in the array.
10.3 Element identity within a frame #
Within a single frame, each element MUST have a unique identity for the layout engine. As specified in Section 9.3, element IDs MUST be unique within a frame. Passing duplicate IDs is undefined behavior.
10.4 No cross-frame identity #
The renderer does not track element identity across frames. An element with id "sidebar" in frame N and an element with id "sidebar" in frame N+1 are not related from the renderer's perspective. Cross-frame identity, if needed, is the responsibility of a higher-level framework.
11. Boundaries and Non-Responsibilities #
This section is normative.
11.1 The renderer does not perform IO #
The renderer MUST NOT write to any output stream. The renderer MUST NOT read from any input stream. The renderer produces bytes; the caller decides when and how to write them. This enables the renderer to operate in any environment where WebAssembly is available, including browsers, server-side runtimes, and embedded contexts.
11.2 The renderer does not manage terminal state #
The renderer MUST NOT emit escape sequences for any of the following terminal-management operations:
- Entering or leaving the alternate screen buffer
- Setting the cursor shape or blink state
- Enabling or disabling mouse reporting
- Enabling or disabling keyboard protocol modes (e.g., Kitty progressive enhancement)
- Enabling or disabling raw mode or similar terminal disciplines
These are the caller's responsibility. The renderer's output contains only the
escape sequences needed to render the frame content (cursor positioning for cell
writes, SGR attributes for styling, and UTF-8 text) and, when a caret
declaration is present, the cursor-positioning and cursor-visibility sequences
specified in §7.6.
Capability-specific output modes are not terminal-state management owned by the foundation renderer. Each capability section of §7 defines its own state boundary. The mouse pointer shape (§7.9) is distinct from the cursor shape above, which remains caller-managed.
11.3 The renderer does not own application lifecycle #
The renderer MUST NOT maintain a run loop, event loop, timer, or subscription mechanism. It does not schedule frames. It does not manage component state. It renders when asked and returns. The decision of when to render is entirely the caller's.
11.4 The renderer does not own input parsing #
Input parsing (keyboard events, mouse events, escape sequence decoding) is an independent concern specified separately in the Clayterm Input Specification. The renderer MUST NOT depend on input-parsing state, types, or API.
However, pointer hit detection does require the render loop to participate. The
caller may pass the current pointer position as part of render options, and the
renderer returns the ids of every element the pointer is over. This is how the
PointerEvent[] array in the render result is populated. See Section 12.4 for
the current pointer event surface.
11.5 The renderer does not own higher-level framework concerns #
The renderer MUST NOT implement or depend on:
- Component models or component lifecycles
- Reconciliation or diffing of directive trees (the renderer diffs cells, not trees)
- State management or reactivity
- Event propagation through a component hierarchy
These are the domain of higher-level frameworks built on Clayterm.
12. Current Surface That Remains Elastic #
This entire section is non-normative. It describes the current implementation surface to aid consumers and future spec authors. The shapes described here are real, working, and in many cases deliberately designed, but they do not yet meet the stability threshold for normative specification. They MAY change in future versions without constituting a breaking change to the normative core defined above.
12.1 Transfer encoding (command protocol) #
The renderer currently serializes directives into a flat byte buffer using a
command protocol based on fixed-width Uint32 words. Each directive is encoded
as an opcode word followed by directive-specific data. Element-open directives
use a property mask to indicate which optional field groups (layout, border,
corner radius, clip, floating, scroll) are present, followed by the data for
each indicated group. Strings are encoded as length-prefixed UTF-8 byte
sequences within the word stream. Floats are stored as bit-reinterpreted
Uint32 values.
This encoding has been extended incrementally (floating, clip, and scroll groups were added after the initial protocol) but has never been restructured. It is likely to remain stable in structure while continuing to grow. However, specific opcode values, mask definitions, and field layouts are implementation details and are not locked down by this specification.
12.2 Directive property groups #
The open() constructor currently accepts the following property groups in its
props parameter:
layout— sizing (width and height, specified via sizing helpers), padding (per-side), alignment (alignX:"left"|"center"|"right";alignY:"top"|"center"|"bottom", defaulting to left/top when omitted), direction (top-to-bottom or left-to-right), and gapborder— per-side border configuration. Each side field (top,right,bottom,left) accepts either a scalar width or a structured object{ width, color?, bg? }. The sharedcolorfield is required and is the fallback foreground for every side; the optional sharedbgfield is the fallback border-cell background for every sidecornerRadius— per-corner radius values, producing rounded box-drawing charactersclip— Declares the element as a clip region (see §7.5). Currently acceptshorizontal: booleanandvertical: booleanaxis selectors. Originally added for scroll containers; nesting and standalone use are supported.floating— floating-element configuration (offset, expansion, parent reference, attach target, structured attach points, pointer capture mode, clip target, z-index)scroll— scroll container configurationpointerShape— mouse pointer shape while hovered (§7.9); not transferred to the WASM module
The floating object shape is:
floating?: {
x?: number;
y?: number;
expand?: { width?: number; height?: number };
parent?: number;
attachTo?: "none" | "parent" | "element" | "root";
attachPoints?: {
element?:
| "left-top"
| "left-center"
| "left-bottom"
| "center-top"
| "center-center"
| "center-bottom"
| "right-top"
| "right-center"
| "right-bottom";
parent?:
| "left-top"
| "left-center"
| "left-bottom"
| "center-top"
| "center-center"
| "center-bottom"
| "right-top"
| "right-center"
| "right-bottom";
};
pointerCaptureMode?: "capture" | "passthrough";
clipTo?: "none" | "attached-parent";
/** signed 16-bit integer */
zIndex?: number;
}
The floating object configures Clay floating layout behavior. x and y
provide the floating offset. expand expands the floating bounds. parent
identifies the target element when attachTo is "element". attachTo selects
whether the element is attached to no target, its parent, an element, or the
layout root. attachPoints.element describes the anchor on the floating
element, and attachPoints.parent describes the anchor on the attached target.
pointerCaptureMode controls whether the floating element captures pointer
input or lets it pass through, clipTo controls inherited clipping, and
zIndex controls floating order and is transferred as a signed 16-bit integer.
The text() constructor accepts: color, bg, fontSize, letterSpacing,
lineHeight, and attribute flags (bold, italic, underline,
strikethrough).
These property groups represent the current implementation surface. New groups and fields have been added incrementally and more may follow.
Border sides. Each border side is declared independently as either a scalar
width (top: 1) or a structured object (top: { width: 1, color?, bg? }). The
two forms are equivalent when the object form provides only width. A side is
enabled when its resolved width is greater than zero; an omitted side or a side
with width 0 MUST NOT be drawn. Scalar side declarations MUST keep their
pre-existing behavior.
Border side colors (fallback resolution). Side attributes resolve in a CSS-like shorthand/longhand fashion before rendering:
- A structured side with
colorMUST render with that foreground color. A scalar side, or a structured side that omitscolor, MUST fall back to the sharedborder.color. The sharedcolorremains required. - A structured side with
bgMUST render border cells of that side with that background color. A scalar side, or a structured side that omitsbg, MUST fall back to the sharedborder.bgwhen it is provided. - When neither the side nor the shared border provides
bg, border rendering MUST NOT override the background already present in each border cell of that side; element backgrounds established byopen({ bg })remain in effect, and the terminal default remains in effect where no element background applies.
Fallback resolution is performed on the TypeScript side before the frame is transferred; the WASM renderer consumes explicit per-side attributes and does not implement the public fallback rules.
Independent sides and corners. Each enabled side renders as a straight edge
(─ for horizontal sides, │ for vertical sides). A corner glyph MUST be
rendered only when both adjacent sides for that corner are enabled; when either
adjacent side is absent, the present side continues straight through the
endpoint with no corner glyph. A left-only border is therefore a plain vertical
line, and a top-plus-bottom border is two plain horizontal rules.
Corner styling approximation. A terminal cell carries a single glyph,
foreground, and background, so CSS-style diagonally split corners cannot be
represented. When corners are rendered: top corners (┌, ┐, and their rounded
variants) MUST use the resolved attributes of the top side, and bottom corners
(└, ┘, and their rounded variants) MUST use the resolved attributes of the
bottom side. Left and right side attributes apply to vertical edge cells
excluding joined corner cells. Per-side attributes affect only the styling of
corner cells; corner glyph shape selection (including rounded corners via
cornerRadius) is unchanged.
Border width and layout interaction. The renderer automatically reserves
space for each enabled border side. For each side, the effective padding passed
to the layout engine is userPadding + borderWidth; the WASM renderer applies
this adjustment when decoding the element declaration. Border glyphs are drawn
at the same positions as before; the change is purely in how much layout space
Clay allocates for the element.
Semantics of the additive rule:
- No explicit padding. The border width itself becomes the effective padding, so content is placed immediately inside the border.
- Explicit user padding. Adds breathing room beyond the border edge.
padding: 1withborder: 1places content 2 cells from the element edge — 1 for the border glyph, 1 for the padding inset. - Prior workaround pattern (
padding == borderWidth). These callers now receive double-reservation (effective = 2 × borderWidth). This is a breaking change: remove the workaround padding to restore the original visual.
12.3 Render return type #
The render() method currently returns a RenderResult object shaped as
{ output: Uint8Array, events: PointerEvent[], info: RenderInfo, errors: ClayError[] }.
The output field is the ANSI byte output specified normatively in Section 7.3
and Section 8.2.
The events field contains pointer events (enter, leave, click) derived from
the underlying layout engine's element hit-testing. This field was added during
a pointer-events feature implementation. The pointer event model is functional
but has acknowledged gaps (no modifier keys on click events) and its interaction
protocol (passing pointer state via render options, then reading events from the
return value) was arrived at through iteration rather than upfront design.
The info field implements RenderInfo, a read-only lookup keyed by element id
(the id parameter passed to open()):
interface RenderInfo {
get(id: string): ElementInfo | undefined;
}
interface ElementInfo {
bounds: BoundingBox;
}
interface BoundingBox {
x: number;
y: number;
width: number;
height: number;
}
Each ElementInfo provides post-layout metadata. The bounds field is the
element's computed bounding box in character cells, as determined by the layout
engine after the render transaction completes. x and y are zero-indexed from
the top-left corner of the layout root.
Querying an element with an empty-string id or an id not present in the frame
returns undefined.
The errors field contains any errors reported by the Clay layout engine during
the most recent render() call. Each error is a ClayError object with:
type: a string identifying the error category. The following types are defined. Most mirror Clay's error taxonomy;"CLIP_DEPTH_EXCEEDED"and"COMBINING_MARKS_EXCEEDED"are Clayterm-specific."TEXT_MEASUREMENT_FUNCTION_NOT_PROVIDED""ARENA_CAPACITY_EXCEEDED""ELEMENTS_CAPACITY_EXCEEDED""TEXT_MEASUREMENT_CAPACITY_EXCEEDED""DUPLICATE_ID""FLOATING_CONTAINER_PARENT_NOT_FOUND""PERCENTAGE_OVER_1""INTERNAL_ERROR""UNBALANCED_OPEN_CLOSE""CLIP_DEPTH_EXCEEDED"— A frame nested clip regions more deeply than the renderer could track. See §7.5 for the guarantees that still hold in this case. ThemessageSHOULD identify the renderer's tracking limit."COMBINING_MARKS_EXCEEDED"— A frame attached more combining marks to a single cell than the cell can store. The excess marks are truncated (see §13, Cell representation). Reported at most once per frame, on the first truncation. ThemessageSHOULD identify the per-cell limit.
message: a human-readable string describing the error in detail.
Errors are collected per-render; each call to render() returns only the errors
from that invocation. The array is empty when no errors occurred.
The return type of render() has changed twice since the project's inception
(string, then Uint8Array, then RenderResult). While the ANSI bytes
commitment (Section 7.3) is stable, the wrapper shape around those bytes is not.
Future versions may restructure the return type.
12.4 Pointer event model #
Clayterm currently supports pointer hit-testing via the underlying layout
engine's element-identification mechanism. The caller passes pointer state
({ x, y, down }) as part of render options, and the renderer returns pointer
events as part of the render result:
pointerenter— the pointer has entered an element's bounding boxpointerleave— the pointer has left an element's bounding boxpointerclick— a pointer-up occurred over an element that was also under the pointer at pointer-down
This surface is functional but should not be treated as stable contract. The calling convention was discovered through iteration, the event model has acknowledged gaps, and the approach may evolve.
12.5 Validation and packing #
validate(ops) — A public API function that checks a directive array for
structural errors (unbalanced open/close pairs, invalid field types). Exported
and used in tests.
pack(ops, mem, offset) — An internal function that serializes a directive
array into the transfer encoding described in Section 12.1. Currently exported
but not public API; its exposure is incidental to the module structure.
13. Implementation Notes #
This section is non-normative. These notes describe current implementation details that aid understanding but do not define contract. They may change without notice.
WASM module structure. The renderer is implemented in C and compiled to WebAssembly as a single module. The module contains both rendering and input-parsing functionality; they share a binary but maintain independent state.
WASM loading. The WASM binary is inlined as a base64-encoded string in a generated module and instantiated per Term or Input with fresh memory.
Memory layout. WASM linear memory is initialized with 256 pages (16MB). The renderer state struct and the transfer buffer are allocated in WASM linear memory. The specific layout is an implementation detail.
Capability state. RuntimeCapabilities is currently held in TypeScript and
folded by update(); the WASM renderer state carries no capability fields
because no renderer output depends on them yet (§7.8). When a
capability-consuming feature lands, the authoritative state is expected to move
into WASM linear memory alongside the renderer state, so that render() reads
it without per-frame transfer; term.capabilities remains a frozen snapshot
decoded on access.
Layout engine. The underlying layout engine is Clay, included as a dependency. Clay provides flexbox-like layout computation with support for fixed, grow, and fit sizing; padding; alignment; direction; gap; floating elements; clip regions; and scroll containers.
Text measurement. Text width measurement uses wcwidth-based character
width computation, supporting ASCII, CJK wide characters, and other Unicode
codepoints. Combining marks (codepoints with wcwidth ≤ 0) contribute zero to
measured width; they attach to the preceding base codepoint's cell and are not
counted as separate cells in layout. Measurement and rendering MUST agree: if
the measurer ignores a combining mark for width purposes, the renderer MUST
still attach and emit it.
Cell representation. Each cell in the buffer stores a grapheme cluster — a
base Unicode codepoint plus up to 8 combining-mark codepoints — together with a
foreground color (packed ARGB with attribute flags in the high byte) and a
background color. The combining-mark slots are zero-terminated; a cell with no
combining marks stores zero in every slot. When a text string produces more than
8 combining marks for a single base codepoint, the excess marks are truncated
from the end (marks 1–8 are kept, marks 9+ are discarded), ensuring that the
first and most semantically significant marks always survive. The first
truncation in a frame is reported as a "COMBINING_MARKS_EXCEEDED" error
(§12.3); later truncations in the same frame are not reported again. Cell
comparison for diffing considers combining marks: a cell is considered changed
when any combining mark differs from the front buffer, not only when the base
codepoint or color attributes differ.
Border junction resolution. When bordered elements share edges, the renderer accumulates per-cell direction bitmasks and resolves them to correct box-drawing junction glyphs in a post-render pass.
Clip stack. Section 7.5 requires the effective clip region of a nested
clip element to be the intersection of its declared region with its clipping
ancestor's effective region. The underlying layout engine (Clay) emits
per-clip-element bounding boxes that are not pre-intersected with any ancestor's
clip, so the renderer maintains an internal stack of effective clip rectangles:
it pushes the intersected rect on each clip-region entry and pops on exit. The
stack capacity is a small fixed value sufficient for realistic UIs; depth beyond
that is handled per §7.5 (prior clips honored, the over-deep level coalesced
into its deepest tracked ancestor, and a "CLIP_DEPTH_EXCEEDED" error
surfaced).
Upstream Clay may eventually flatten nested clip emission so renderers only need
single-rect handling; see
nicbarker/clay#466 (the
underlying issue),
nicbarker/clay#485 (in-flight
Clay-side fix), and
nicbarker/clay#87 (renderer
guidance). When upgrading Clay, check whether a single clip element now produces
multiple SCISSOR_START/SCISSOR_END pairs across its lifetime (one per
nesting transition rather than just an outer pair); if so, the renderer-side
stack can be removed and replaced with a single rect storing Clay's bounding box
directly.
14. Deferred / Future Areas #
This section is non-normative. These topics are explicitly excluded from this specification. Their omission is intentional, not an oversight.
Scroll container API. The underlying layout engine supports scroll containers. No TypeScript-side API exists for providing scroll state to the renderer.
CSI helper for terminal setup. No helper exists for generating paired apply/rollback byte arrays for terminal mode configuration.
Browser-specific adapter. The renderer's zero-IO architecture makes browser portability possible. No adapter exists.
betweenChildren border support. The underlying layout engine supports
this. It is not exposed in the directive model.
Open Decisions Intentionally Left Out of This Spec #
The following decisions are open. This specification omits them deliberately. Future readers should not interpret their absence as oversight or implicit resolution.
-
What is the normative return type of
render()? This specification commits to ANSI bytes asUint8Arraybut does not lock down the wrapper type. The currentRenderResultshape may evolve. -
Is pointer event detection part of the rendering contract? The current implementation returns pointer events from
render(). This specification does not include pointer events in the normative core. Whether pointer detection is intrinsic to the renderer or should be a separate concern is unresolved. -
Is
pack()public API?pack()is currently exported but is an internal implementation detail, not public API.validate()is public API. -
What are the specific transfer encoding details? The encoding structure is described in Section 12.1 as current implementation surface. Locking down opcode values would constrain future extensions unnecessarily.
-
What is the complete set of directive properties? The property groups available in
open()andtext()are described in Section 12.2 as current implementation surface. They have been extended incrementally and will continue to grow. -
What are the validation and error semantics? How the renderer responds to invalid input is unspecified. Callers SHOULD validate, but the validation model is not yet settled enough to define normatively.