WIP experiment with 95% vibe coding
tiny-workshop docs design ipc-protocol.md
27 kB
Markdown
at main

IPC Protocol Design #

Component: IPC Protocol (stdin/stdout JSON-RPC) Status: Design Related: Frontend Architecture, Backend Adapter


Overview #

The IPC protocol defines the message format and lifecycle for communication between the Fyrox frontend (Rust) and the backend adapter (Bun/TypeScript). It uses JSON-RPC over stdin/stdout for cross-platform, language-agnostic communication.

Design Goals #

  1. Language-agnostic: Works with any language that can read/write stdin/stdout
  2. Simple: JSON lines, no complex framing or multiplexing
  3. Type-safe: Strongly-typed messages on both sides (TypeScript discriminated unions, Rust enums)
  4. Versionable: Protocol version negotiation at startup
  5. Streamable: Supports incremental updates (text deltas, tool progress)
  6. Bidirectional: Both frontend and backend can initiate requests

Transport Layer #

Why stdin/stdout over Unix sockets or HTTP? #

Consideration stdin/stdout Unix Sockets HTTP
Cross-platform ✓ (universal) ✗ (Windows named pipes differ) ✓
Process coupling Child process Separate process Separate process
Setup complexity None Socket file management Port allocation, CORS
Language support Universal Library-dependent Universal
Multiplexing No (sequential JSON lines) Yes Yes
Debugging Easy (tee stderr) Harder (inspect socket) Easy (curl)

Decision: stdin/stdout provides the simplest cross-platform process communication with zero setup overhead. The frontend spawns the backend as a child process, giving automatic lifecycle management (backend dies when frontend exits).

Framing #

Each message is a single JSON object on one line, terminated by \n:

{"type":"init","version":1,"working_dir":"/foo","config":{...}}\n
{"type":"session_started","session_id":"abc123","model":"claude-sonnet-4.5","tools":[...]}\n
{"type":"assistant_text_delta","content":"Hello! "}\n
{"type":"assistant_text_delta","content":"How can I help?"}\n

Rationale:

  • Line-buffered: Both Rust (BufReader::lines()) and Node.js (readline) have built-in line-based reading
  • No escaping: JSON already handles special characters (newlines become \n in strings)
  • Human-readable: Easy to debug with tee or log files
  • Streaming-friendly: Sender can flush after each message for low latency

Limitation: No binary data. Images and files are referenced by path, not transmitted inline.


Protocol Version Negotiation #

Version Number #

PROTOCOL_VERSION = 1

Version is a single integer (not semantic versioning) because:

  • Protocol is internal to this project (not a public API)
  • Frontend and backend are released together
  • Breaking changes require frontend + backend update

Negotiation Flow #

Frontend                     Backend
   │                            │
   │─── init (version: 1) ─────►│
   │                            │
   │                            ├─ Check version
   │                            │
   │◄── session_started ────────│  (version OK)
   │                            │
   │     OR                     │
   │                            │
   │◄── error (PROTOCOL_MISMATCH)│  (version mismatch)
   │                            │
   │─── shutdown ──────────────►│  (close and show error)

Rationale: Version is part of init message, so backend can reject incompatible frontends immediately.

Version Compatibility #

Frontend Backend Compatible? Behavior
1 1 ✓ Normal operation
1 2 ✗ Backend sends PROTOCOL_MISMATCH error
2 1 ✗ Backend sends PROTOCOL_MISMATCH error

Future: If we need backwards compatibility, backend could support multiple versions:

if (msg.version >= 1 && msg.version <= 2) {
  // Adapt behavior based on version
}

Message Types #

Naming Convention #

  • Frontend → Backend: Imperative (command) or noun (data)
    • init, user_message, cancel, shutdown
  • Backend → Frontend: Past tense (event) or noun (data)
    • session_started, assistant_text_delta, tool_start, error

Message Categories #

Category Direction Purpose
Lifecycle Frontend → Backend init, shutdown
Lifecycle Backend → Frontend session_started, session_ended, compaction_occurred
Conversation Frontend → Backend user_message
Conversation Backend → Frontend assistant_text_delta, assistant_message_complete
Tools Backend → Frontend tool_start, tool_end
Room Backend → Frontend room_action, room_query
Room Frontend → Backend room_query_response, room_action_response
Control Frontend → Backend cancel, ping, mode_change
Control Backend → Frontend pong, error
Auth Backend → Frontend auth_status

Message Schemas #

Frontend → Backend Messages #

init #

Start a new session or resume an existing one.

{
  type: 'init';
  version: number;                  // Protocol version (currently 1)
  working_dir: string;              // Absolute path to project directory
  api_key?: string;                 // z.AI API key (optional)
  resume_session_id?: string;       // SDK session ID to resume (optional)
}

When sent: On app startup.

Response: session_started or error

Note: Future extensions planned but not yet implemented:

  • resume?: string - SDK session ID to resume
  • initial_prompt?: string - First user message
  • config?: { model?: string; permission_mode?: ... } - SDK configuration

user_message #

Send a user message to the agent.

{
  type: 'user_message';
  content: string;                  // User's message text
  attachments?: Array<{             // Optional file attachments
    type: 'image' | 'file';
    path: string;                   // Absolute path to file
    mime?: string;                  // MIME type (optional, inferred if missing)
  }>;
}

When sent: User presses Enter in chat input or sends a file.

Response: Stream of assistant_text_delta, tool_start, tool_end, assistant_message_complete

Open Issue: Does SDK support mid-session user messages, or must we restart query() with resume?


cancel #

Cancel the current generation.

{
  type: 'cancel';
}

When sent: User presses Esc during streaming response.

Response: session_ended with reason 'cancelled'


ping #

Heartbeat to check backend responsiveness.

{
  type: 'ping';
}

When sent: Every 5 seconds (from frontend).

Response: pong (backend should respond within 10 seconds)


mode_change #

Switch between auto and plan modes.

{
  type: 'mode_change';
  mode: 'auto' | 'plan';
}

When sent: User presses Shift+Tab.

Response: None (backend updates SDK permission mode)


room_query_response #

Response to a room_query from backend.

{
  type: 'room_query_response';
  request_id: string;               // Matches room_query.request_id
  description: string;              // Text description of room state
}

When sent: Frontend responds to backend's room tool query.

Response: None (backend resolves pending promise)


room_action_response #

Response to a room_action from backend. Provides actual success/failure status of the action so the agent receives accurate feedback.

{
  type: 'room_action_response';
  request_id: string;               // Matches room_action.request_id
  success: boolean;                 // Whether the action succeeded
  message: string;                  // Human-readable result or error message
}

When sent: Frontend responds to backend's room action after executing it.

Response: None (backend resolves pending promise and returns result to MCP tool)

Example success: { success: true, message: "Created yellow note 'TODO' on Desk" }

Example failure: { success: false, message: "Mantelpiece is full" }


shutdown #

Gracefully shut down backend.

{
  type: 'shutdown';
}

When sent: App is closing.

Response: Backend exits (no response message)


Backend → Frontend Messages #

session_started #

Session initialization complete.

{
  type: 'session_started';
  session_id: string;               // SDK session ID (UUID)
  model?: string;                   // Model being used (optional)
  tools?: string[];                 // Available tools (optional, defaults to [])
}

When sent: After init, once SDK session is ready.

Source: Transformed from SDKSystemMessage with subtype: 'init'.

Note: model and tools are optional for backwards compatibility. Older backends may not send these fields.


assistant_text_delta #

Incremental text from assistant.

{
  type: 'assistant_text_delta';
  content: string;                  // Text fragment
}

When sent: As SDK streams assistant message.

Source: Extracted from SDKAssistantMessage text blocks.

Rationale: Streaming allows frontend to render text incrementally (better perceived performance).


assistant_message_complete #

Assistant message fully received.

{
  type: 'assistant_message_complete';
  content: string;                  // Full text (concatenated deltas)
}

When sent: When SDK message has stop_reason.

Source: SDKAssistantMessage with stop_reason present.

Use case: Frontend triggers "turn from camera" animation, logs complete message.


tool_start #

Tool execution began.

{
  type: 'tool_start';
  id: string;                       // Tool use ID (for matching with tool_end)
  tool: string;                     // Tool name (e.g., "Read", "Bash")
  input: Record<string, unknown>;   // Tool arguments
}

When sent: When SDK sends assistant message with tool_use block.

Source: Extracted from SDKAssistantMessage tool_use content blocks.


tool_end #

Tool execution completed.

{
  type: 'tool_end';
  id: string;                       // Matches tool_start.id
  output: unknown;                  // Tool result (structure depends on tool)
  success: boolean;                 // true if tool succeeded, false if error
}

When sent: When SDK sends user message with tool_result block.

Source: Extracted from SDKUserMessage tool_result content blocks.


subagent_spawn #

Subagent spawned.

{
  type: 'subagent_spawn';
  id: string;                       // Subagent ID
  task: string;                     // Task description
}

When sent: When SDK indicates subagent creation.

Source: TBD (needs SDK documentation on subagent detection)

Open Issue: How are subagents represented in SDK messages?


subagent_message #

Message from subagent.

{
  type: 'subagent_message';
  id: string;                       // Subagent ID
  content: string;                  // Message text
}

When sent: Subagent sends output.

Source: TBD


subagent_end #

Subagent completed.

{
  type: 'subagent_end';
  id: string;                       // Subagent ID
  success: boolean;                 // true if task succeeded
}

When sent: Subagent finishes.

Source: TBD


room_action #

Agent manipulated room state.

{
  type: 'room_action';
  action: RoomAction;               // See Room System design doc
}

// Action types:
type RoomAction =
  | { type: 'create_note'; title: string; content: string; surface: string; color?: NoteColor }
  | { type: 'remove_note'; title: string; surface: string }
  | { type: 'move_note'; title: string; from_surface: string; to_surface: string }
  | { type: 'clear_room' };

type NoteColor = 'yellow' | 'pink' | 'blue' | 'green';

When sent: MCP room tool is invoked.

Source: Emitted from room-tools.ts MCP server.


room_query #

Backend requests room state description.

{
  type: 'room_query';
  request_id: string;               // UUID for response matching
  query: 'describe' | 'examine';    // Query type
  target?: string;                  // For 'describe': surface ID (optional)
                                    // For 'examine': note title (required)
}

When sent: MCP room tool (look_around or examine) is invoked.

Query types:

  • describe: Returns room overview or specific surface description. If target is omitted, describes entire room. If target is a surface ID (e.g., "desk"), describes that surface.
  • examine: Returns detailed note information. target must be the exact note title. Returns note content, color, and location.

Response: Frontend sends room_query_response with same request_id.


session_ended #

Session terminated.

{
  type: 'session_ended';
  reason: 'success' | 'error_max_turns' | 'error_during_execution' | 'error_max_budget_usd' | 'cancelled';
}

When sent: SDK query completes or errors.

Source: Transformed from SDKResultMessage.


compaction_occurred #

Conversation was compacted (context pruning).

{
  type: 'compaction_occurred';
  old_session_id: string;           // Previous SDK session ID
  new_session_id: string;           // New SDK session ID (after compaction)
  trigger: 'manual' | 'auto';       // How compaction was triggered
}

When sent: SDK prunes conversation history.

Source: Transformed from SDKCompactBoundaryMessage.

Frontend action: Update ProjectState.active_session, log compaction event.


pong #

Response to ping.

{
  type: 'pong';
}

When sent: Immediately after receiving ping.

Use case: Frontend tracks last_pong timestamp to detect hung backend.


error #

Error occurred.

{
  type: 'error';
  code: string;                     // Error code (e.g., "SDK_ERROR", "PROTOCOL_MISMATCH")
  message: string;                  // Human-readable error description
  recoverable: boolean;             // Can frontend retry/continue?
}

When sent: Any error condition.

Error codes:

  • PROTOCOL_MISMATCH: Version incompatibility
  • SDK_ERROR: SDK threw exception
  • PARSE_ERROR: Failed to parse frontend message
  • AUTH_FAILED: Authentication error
  • NOT_AUTHENTICATED: Backend not initialized with auth
  • TOKEN_REFRESH_FAILED: OAuth token refresh failed
  • SDK_NOT_FOUND: Claude Code CLI not installed

auth_status #

Authentication status notification.

{
  type: 'auth_status';
  authenticated: boolean;           // Currently authenticated?
  error?: string;                   // Error message (only when authenticated=false)
}

When sent:

  • After init to indicate authentication result
  • If no API key provided: authenticated: false with error "No API key provided"
  • If API key is invalid: authenticated: false with error from z.AI

Frontend action:

  • If authenticated=false: Show error message, prompt for API key

Message Flow Examples #

New Session Flow #

Frontend                                Backend
   │                                       │
   │─── init ──────────────────────────────►│
   │    { version: 1,                      │
   │      working_dir: "/foo",             │
   │      initial_prompt: "hello" }        │
   │                                       │
   │◄── session_started ────────────────────│
   │    { session_id: "abc123",            │
   │      model: "claude-sonnet-4.5",      │
   │      tools: [...] }                   │
   │                                       │
   │◄── assistant_text_delta ───────────────│
   │    { content: "Hello! " }             │
   │                                       │
   │◄── assistant_text_delta ───────────────│
   │    { content: "How can I help?" }     │
   │                                       │
   │◄── assistant_message_complete ─────────│
   │    { content: "Hello! How can I help?"}│

Tool Execution Flow #

Frontend                                Backend
   │                                       │
   │◄── tool_start ─────────────────────────│
   │    { id: "t1",                        │
   │      tool: "Read",                    │
   │      input: { file_path: "/foo" } }   │
   │                                       │
   │    [SDK executes Read tool]           │
   │                                       │
   │◄── tool_end ───────────────────────────│
   │    { id: "t1",                        │
   │      output: { content: "..." },      │
   │      success: true }                  │
   │                                       │
   │◄── assistant_text_delta ───────────────│
   │    { content: "The file contains..." }│

Room Query Flow #

Frontend                                Backend
   │                                       │
   │◄── room_query ─────────────────────────│
   │    { request_id: "req1",              │
   │      query: "describe" }              │
   │                                       │
   │    [Frontend generates description]   │
   │                                       │
   │─── room_query_response ────────────────►│
   │    { request_id: "req1",              │
   │      description: "The desk has 3     │
   │        notes and 2 image stacks..." } │
   │                                       │
   │    [Backend returns to MCP tool]      │
   │                                       │
   │◄── tool_end ───────────────────────────│
   │    { id: "t2",                        │
   │      output: { content: [...] },      │
   │      success: true }                  │

Crash Recovery Flow #

Frontend                                Backend
   │                                       │
   │◄── tool_start ─────────────────────────│
   │    { id: "t1", tool: "Bash", ... }    │
   │                                       │
   │    [Backend process crashes]          │
   │                                       X
   │
   │    [Frontend detects EOF on stdout]
   │    [Frontend restarts backend]
   │                                       │
   │─── init ──────────────────────────────►│
   │    { resume: "abc123", ... }          │
   │                                       │
   │◄── session_started ────────────────────│
   │    { session_id: "abc123", ... }      │
   │                                       │
   │    [SDK resumes from last checkpoint] │
   │                                       │
   │◄── tool_end ───────────────────────────│
   │    { id: "t1", ... }                  │

Error Handling #

Error Response Format #

All errors use the error message type:

{
  type: 'error';
  code: string;
  message: string;
  recoverable: boolean;
}

Error Code Catalog #

Code Meaning Recoverable Frontend Action
PROTOCOL_MISMATCH Version incompatibility No Show error dialog, exit
SDK_NOT_FOUND Claude Code CLI not installed No Show installation instructions, exit
SDK_ERROR SDK threw exception Yes Show error, allow retry
PARSE_ERROR Malformed frontend message Yes Log warning, continue
AUTH_FAILED Authentication error Yes Show error, prompt for API key
NOT_AUTHENTICATED Backend not initialized with auth Yes Send init message with api_key
PERMISSION_DENIED User denied permission prompt Yes Continue (SDK handles)

Performance Characteristics #

Message Size #

Message Type Typical Size Max Size
init 200 bytes 1 KB
user_message 500 bytes 100 KB
assistant_text_delta 50 bytes 1 KB
tool_start 300 bytes 10 KB
tool_end 1 KB 1 MB (large tool outputs)
room_action 500 bytes 10 KB
room_query_response 2 KB 50 KB

Rationale: Most messages are small (<1 KB). Large messages (tool outputs) are rare and handled with streaming.

Latency #

Metric Target Measurement
Message serialization <1ms serde_json::to_string()
Message deserialization <1ms JSON.parse()
Stdout write <1ms console.log() + flush
Stdin read <1ms readline event
End-to-end (frontend → backend → frontend) <10ms Round-trip ping/pong

Throughput #

  • Target: 1000 messages/second
  • Bottleneck: JSON serialization/deserialization (CPU-bound)
  • Mitigation: Keep messages small, batch where possible (e.g., coalesce text deltas)

Extensibility #

Adding New Messages #

  1. Add type to FrontendMessage or BackendMessage union
  2. Increment protocol version if breaking (optional fields don't require version bump)
  3. Update handler in frontend/backend
  4. Add tests

Example: Adding a pause_generation message:

// Frontend message
{
  type: 'pause_generation';  // New type
}

// Backend response
{
  type: 'generation_paused';
  can_resume: boolean;
}

Optional Fields #

Optional fields can be added without breaking compatibility:

// v1
{
  type: 'user_message';
  content: string;
}

// v2 (backwards compatible)
{
  type: 'user_message';
  content: string;
  context?: string;  // New optional field
}

Old backends ignore unknown fields, new backends handle gracefully.


Testing Strategy #

Unit Tests #

#[test]
fn test_serialize_init_message() {
    let msg = FrontendMessage::Init {
        version: 1,
        working_dir: "/foo".into(),
        resume: None,
        initial_prompt: Some("hello".into()),
        config: SdkConfig {
            model: None,
            permission_mode: PermissionMode::Default,
        },
    };

    let json = serde_json::to_string(&msg).unwrap();
    assert_eq!(
        json,
        r#"{"type":"init","version":1,"working_dir":"/foo","initial_prompt":"hello","config":{"permission_mode":"default"}}"#
    );
}

#[test]
fn test_deserialize_session_started() {
    // With all fields
    let json = r#"{"type":"session_started","session_id":"abc123","model":"claude-sonnet-4.5","tools":["Read","Write"]}"#;
    let msg: BackendMessage = serde_json::from_str(json).unwrap();

    match msg {
        BackendMessage::SessionStarted { session_id, model, tools } => {
            assert_eq!(session_id, "abc123");
            assert_eq!(model, Some("claude-sonnet-4.5".to_string()));
            assert_eq!(tools, vec!["Read", "Write"]);
        }
        _ => panic!("Wrong message type"),
    }
}

#[test]
fn test_deserialize_session_started_minimal() {
    // Backwards compatible: only session_id required
    let json = r#"{"type":"session_started","session_id":"abc123"}"#;
    let msg: BackendMessage = serde_json::from_str(json).unwrap();

    match msg {
        BackendMessage::SessionStarted { session_id, model, tools } => {
            assert_eq!(session_id, "abc123");
            assert_eq!(model, None);
            assert!(tools.is_empty());
        }
        _ => panic!("Wrong message type"),
    }
}

Integration Tests #

#[tokio::test]
async fn test_ping_pong() {
    let mut backend = spawn_backend().await;

    // Send ping
    backend.send(FrontendMessage::Ping).await.unwrap();

    // Expect pong
    let msg = backend.recv().await.unwrap();
    assert!(matches!(msg, BackendMessage::Pong));
}

#[tokio::test]
async fn test_protocol_mismatch() {
    let mut backend = spawn_backend().await;

    // Send init with wrong version
    backend.send(FrontendMessage::Init {
        version: 999,
        // ...
    }).await.unwrap();

    // Expect error
    let msg = backend.recv().await.unwrap();
    match msg {
        BackendMessage::Error { code, recoverable, .. } => {
            assert_eq!(code, "PROTOCOL_MISMATCH");
            assert!(!recoverable);
        }
        _ => panic!("Expected error message"),
    }
}

Open Questions #

  1. Message batching: Should we support sending multiple messages in one line (JSON array)?

    • Pro: Reduces syscall overhead for burst sends
    • Con: More complex parsing, line buffering breaks
    • Recommendation: Not needed unless profiling shows bottleneck
  2. Binary data: How do we handle large binary data (e.g., images from tools)?

    • Option A: Write to temp file, send path in message (current approach)
    • Option B: Base64 encode and embed in message
    • Recommendation: Use temp files to keep messages small
  3. Message ordering guarantees: Are messages guaranteed to arrive in order?

    • Answer: Yes, stdin/stdout are sequential streams
    • But: Async processing in frontend may reorder handling
    • Mitigation: Frontend processes messages sequentially from channel
  4. Backpressure: What if frontend can't keep up with backend message rate?

    • Current: Frontend's mpsc channel has bounded capacity (100 messages)
    • Behavior: Backend blocks on stdout write if frontend is slow
    • Acceptable?: Yes, backend is I/O bound anyway (waiting for SDK)