This repository has no description

Personal runtime #

Each installation addresses the PersonalAgent binding/class and the stable instance personal. Keep these identities and the Worker name across upgrades. Wrangler's declarative SQLite exports provisions the namespace; do not add legacy tagged migrations beside it. onStart initializes the application-owned flarebot_runtime singleton with schema version, installation ID, owner subject and creation time. An incompatible schema or changed installation/owner identity fails closed. Future schema changes need explicit forward migrations. SDK tables, state, socket hibernation and alarm dispatch remain owned by agents.

PersonalState is a small public projection containing schemaVersion and createdAt. Native setState persists and broadcasts it. Raw client state updates are rejected before persistence or broadcast; application writes will use narrow validated callables. Credentials, owner identity and raw configuration must not be added to this projection.

Request and session boundary #

The content-free Octane shell and static assets are public. Every /agents/* and /api/* request authenticates at Worker ingress, before the native router can resolve an instance, return protocol state, handle HTTP, or accept a socket. The exact root and /status routes are allowed. Registered conversation facets also allow /agents/personal-agent/personal/sub/conversation/<uuid> (WebSocket) and its /get-messages endpoint (GET). Worker ingress validates the full path, then awaits parent and child readiness before routing. The parent independently checks the exact class, existing native registry entry and leaf path before forwarding. Unknown IDs, top-level child routes and deeper facets cannot create or recreate conversations. Private HTTP responses are no-store.

The __Host-flarebot-session cookie is HMAC-SHA256 signed with the customer's FLAREBOT_SESSION_SECRET. Verification checks owner subject, installation ID, exact runtime audience, issuance and expiry, and rejects duplicate cookies. HTTP uses the configured runtime origin. WebSocket handshakes and mutations require its exact Origin; cross-origin reads carrying Origin are also rejected. Responses containing status or authentication errors are no-store.

worker/session.ts provides server-only createOwnerSession and verification helpers for the later OAuth ownership bridge. Issuance produces a Secure, HttpOnly, SameSite=Lax, Path=/ cookie lasting at most eight hours. There is no public issuance endpoint or arbitrary-owner developer login. Before calling issuance in production, the bridge must authenticate the owner and consume its one-time assertion/challenge. Signing keys never enter frontend props or state. Socket attachments store only expiry metadata; native SDK schedules close sockets at session expiry with code 4001, and close cleanup cancels their schedules. Conversation facets use the same lifecycle while retaining authentication at Worker ingress.

Native client integration #

Use the framework-independent AgentClient from agents/client in Octane:

const client = new AgentClient({
  host: location.host,
  agent: "PersonalAgent",
  name: "personal",
  onStateUpdate: (state) => {
    /* publish to Octane state */
  },
});
await client.ready;
const status = await client.call("getStatus");
// Teardown only detaches this client; it does not delete durable state.
client.close();

The browser sends the same-origin cookie automatically. The SDK handles protocol identity, state snapshots and reconnect backoff; ready resets on disconnect. Mount a client only after login, and return to login after an expired session. Do not create a custom WebSocket hub or mirror SDK persistence/protocol tables. See native routing and the installed agents/docs/state.md and readonly-connections.md references.

Validation #

Run pnpm build:release, then pnpm test:runtime. The test exercises the actual bundled Worker in local workerd with a native AgentClient: rejected HTTP/socket credentials, ownership/audience/origin/path isolation, two clients, state injection, RPC restrictions, socket expiry, reconnect and a full runtime restart using the same filesystem persistence and installation Worker name. Test-only cookie issuance needs no external OAuth registration and adds no production route. pnpm test:deployment separately checks release integrity and internal native Agent instantiation. Local tests do not certify a deployed OAuth bridge.

Think conversation harness #

Conversation extends Think<Env> is an internal native child facet of PersonalAgent. Each child has its own SQLite transcript, workspace, queued turns, buffered streams and recovery fibers. The parent remains the one user-facing personal agent. Never switch Think.session to multiplex conversations: Think owns one session/cache/stream per instance. Native Session persistence owns the transcript; the application does not duplicate it in parent tables.

Authenticated owner clients call the parent using native AgentClient.call:

const conversation = await client.call("createConversation", ["Research"]);
const conversations = await client.call("listConversations");
await client.call("renameConversation", [conversation.id, "Release research"]);
await client.call("deleteConversation", [conversation.id]);

Creation generates the UUID on the server and defaults to New conversation. Names must be strings of 1–120 UTF-16 code units before trimming, contain no control characters, and remain nonempty after trimming. Create, rename and list return only { id, name, createdAt, updatedAt }; timestamps are ISO strings. updatedAt tracks display metadata changes, not message activity. Lists sort by that timestamp descending, then ID. Invalid IDs/names fail; renaming an unknown or deleted conversation fails, while deleting an absent valid ID is idempotent. Native facet resolution methods remain internal, and client state injection is still rejected. UI list refresh and chat selection belong to the sidebar issue.

The parent stores only display metadata in flarebot_conversations. Both an active metadata row and a native registry entry are required to route a child. The first metadata migration adopts earlier harness facets; subsequent starts never infer active conversations from unrecognized registry entries. Creation persists a pending row before native initialization. Deletion marks the row inaccessible before awaiting native teardown, closes its bridged sockets with 4004 (Conversation deleted), and uses deleteSubAgent to abort work, wipe child storage, and clean native schedules. Concurrent deletes share the same teardown. Failed or interrupted creation/deletion remains inaccessible and startup retries cleanup. A client may also retry deletion explicitly. Readiness checks recheck metadata after initialization yields to a concurrent deletion.

Agents 0.22 bridges facet sockets through the parent. Its constructor installs onMessage/onClose wrappers that resolve child stubs before invoking subclass hooks. The parent therefore wraps those installed handlers in its constructor to reject stale frames and closes before they can recreate a deleted facet. Ordinary root traffic and active conversation frames retain SDK dispatch. Native registry inspection in the test fixture verifies deletion after close traffic and a full restart, beyond simply hiding a child from the metadata list.

The initial model is @cf/meta/llama-3.3-70b-instruct-fp8-fast, using the native AI binding and workers-ai-provider. No provider key is needed. The turn loop is limited to eight steps and 4096 output tokens per model call. The 60-second stream inactivity watchdog enters native bounded recovery (three attempts without progress, with the SDK's finite recovery work/time budgets). Think persists partial output, handles explicit cancellation, reconstructs interrupted turns on wake and reports model errors through its native chat protocol. A subsequent user turn can proceed after an error. Recovery does not make arbitrary external tool side effects exactly-once; future mutating tools need the native action/idempotency boundary.

MCP auto-tools, dynamic extensions and Think workspace Bash are disabled. beforeTurn limits active tools to explicit getTools() and getActions() entries: Think's automatically assembled workspace and client tools are not offered to the model. The deterministic model and echo tool exist only in the test Worker entry. Do not add environment flags selecting a fake production model. Editable instructions, tool settings and the Octane chat UI arrive in their own issues.

Model and provider settings #

The customer installation owns one shared configuration in the parent's private flarebot_model_settings SQLite table. Authenticated owner clients use these native parent callables (the settings UI comes separately):

const catalog = await client.call("getModelCatalog");
const settings = await client.call("getModelSettings");
await client.call("setProviderKey", ["anthropic", apiKey]);
await client.call("updateModelSettings", [
  {
    provider: "anthropic",
    model: catalog.anthropic[0],
  },
]);
// Replace a key with another setProviderKey call; null removes it.
await client.call("setProviderKey", ["anthropic", null]);

The public DTO contains only { configuration: { provider, model }, credentials: { anthropic: "configured" | "missing" } }. configured means a key is stored, not that provider access has been verified. The catalog is an explicit allowlist of text/tool models; arbitrary slugs, endpoint overrides and extra fields are rejected without changing working settings. Keys must be 20–512 printable ASCII characters without whitespace. Inputs and responses must not be logged by the future settings client. Native ingress ownership, origin and socket expiry checks protect these callables; the internal credential-reading RPC is not browser callable.

Each beforeTurn fetches one coherent configuration/credential snapshot through native parentAgent(PersonalAgent), then supplies a model override to Think. All steps of that turn retain that model; the next turn (including a recovered turn) reads the current settings. Existing and newly created conversations share the configuration. Keys never enter native state, child persistent configuration, props, URLs, transcripts or control-plane records. The child holds a redacted Secret wrapper and reveals it only to the provider constructor.

Workers AI uses workers-ai-provider@4.0.0 and the customer's AI binding. Anthropic BYOK uses the official @ai-sdk/anthropic@4.0.49 adapter at its fixed https://api.anthropic.com/v1/messages endpoint within Think's native loop. The customer must supply their own Anthropic API key with model access and billing. There is no AI Gateway or deployment API-token prerequisite and no added manifest resource for FLA9. A plain external Think model slug uses unified billing and is not treated as BYOK. The adapter contract was checked against its exact published source (createAnthropic, key header and endpoint) and Think 0.17's model override.

Provider keys are AES-256-GCM encrypted with a fresh 96-bit IV on every write. HKDF-SHA256 derives a domain-separated key from the installation's FLAREBOT_SESSION_SECRET, using the installation ID as salt; the ID is also authenticated as additional data. Only ciphertext is stored, in customer-owned parent SQLite. Preserve the existing session secret and namespace on upgrades. Rotating the session secret invalidates existing encrypted provider keys; re-enter or remove them after rotation. Failed decryption returns safe replacement guidance. Removing a key clears the active value, but platform backup/PITR retention still applies; revoke a compromised key at Anthropic as well.

The model middleware removes provider request/response diagnostics, raw stream events, warnings and stream metadata, and replaces thrown/stream errors before Think can log or persist provider response bodies. Authentication and rate-limit errors have bounded actionable messages; cancellation keeps AbortError semantics. Missing credentials fail the turn with settings guidance. No silent fallback to another provider occurs. Native stream, tool, cancellation and recovery behavior remain owned by Think.

pnpm test:providers verifies encryption, strict validation and the actual Anthropic adapter's endpoint/key use with mocked network responses, including credential-echoing failures and raw/failed/aborted streams. pnpm test:think also checks live authenticated settings RPC across independent/new conversations, invalid edits, replacement/removal, native error redaction, full restart with unchanged ciphertext, and secret-free snapshots, history, protocol and logs. These local gates make no live provider request and do not verify account billing.

Model references: Workers AI Scout, Anthropic model IDs, and Think lifecycle hooks.

The framework-independent browser stack is AgentClient (agents/client) plus WebSocketChatTransport (agents/chat/transport) and an Octane adapter for AI SDK AbstractChat. Remove the leading slash when passing a pathname as AgentClient.basePath: PartySocket itself adds it. Keep the slash for HTTP history requests. Register native resume handlers before connecting; transport resumption requires forwarding CF_AGENT_STREAM_RESUMING, ...RESUME_NONE and ...STREAM_PENDING to its matching handlers. cancelActiveServerTurn() explicitly stops work; closing a tab/socket merely detaches and lets durable work continue. The UI issue owns hydration, cross-tab snapshots and conversation-switch races.

pnpm test:think exercises the actual production parent/facet classes with a fixture-only AI SDK model in workerd. It checks incremental streaming, concurrent isolated conversations, native server tool results, partial cancellation, model errors followed by a successful turn, native stream reattachment, full runtime restart during a turn, unchanged sibling history, routing rejection and facet socket expiry. pnpm test:runtime still tests the packaged production Worker. It also covers authenticated lifecycle RPC, malformed names/IDs, native registry removal during live deletion, concurrent deletes/history reads, durable metadata, and cleanup after injected deletion failure and interrupted creation. Local deterministic inference does not certify a live Workers AI call.

References: Think configuration, default model and the installed @cloudflare/think/docs/sub-agents.md. Published 0.17 source is the authority when examples describe older signatures. Before debugging runtime or client behavior, symptom-match bug lessons.

Personal agent instructions #

PersonalAgent stores an optional instructions override and modification time in its private flarebot_instructions SQLite table. Defaults describe Flarebot as the one personal assistant and require honest tool, memory, and task reporting without claiming unavailable capabilities. A null override means use the shipped defaults; reset clears the override, so later default improvements can apply on upgrades. Custom text is preserved exactly, accepts tabs and line breaks, rejects other control characters and blank/non-string values, and is limited to 16,000 UTF-16 code units. Invalid writes leave working settings intact.

The authenticated owner callables are getInstructionSettings, updateInstructions(text), and resetInstructions. They return effective text, shipped defaults, customization status, and modification time. Text is never part of public SSR HTML or the generic native agent state. It stays in the customer's account and is supplied to their selected inference provider as a system prompt. Instructions do not grant permissions: Worker ownership/origin checks and the server tool allowlist still apply independently of any customized prompt.

Every Conversation.beforeTurn reads readInstructions() through internal parent RPC and returns Think 0.17's complete instructions override alongside the model. This replaces the assembled frozen fallback; appending to ctx.system would risk keeping obsolete instructions. Existing and new conversations see edits on their next turn, including recovery turns. An already running turn keeps its snapshot. Session history, streaming, and persistence remain native Think behavior. Future memory/tool context must be composed deliberately into this complete instructions override rather than assuming frozen context blocks are appended.

/settings provides an Octane/Kumo form to load, edit, save, and reset the personal agent's instructions. The small createOwnerClient seam preflights authenticated status before opening a native AgentClient, bounds readiness and RPC waits, and closes/aborts when the route unmounts. Late replies cannot replace a later mount's state. Save errors retain the draft; successful reset restores actual defaults. Concurrent owner edits use last successful write; reopen Settings to load changes made in another tab. This route requires an owner session; OAuth sign-in is a later issue and no public test login endpoint is provided.

pnpm test:think verifies the exact system messages received by fixture inference for defaults, edits, resets, existing/new conversations and full restart, alongside invalid write rejection. pnpm exec playwright install chromium then pnpm test:settings runs a real Chromium browser against the packaged production Worker: SSR, signed-out preflight, loading, save/reload, failed save retention, reset/reload, route navigation and a 375px mobile viewport. CI installs Chromium and runs this gate. Tests create owner cookies only in their browser fixture.

Unified tool activity #

Conversation inherits ActivityThink, a small adapter around Think's public beforeToolCall, afterToolCall, onChunk, onChatResponse, and lifecycle hooks. The framework-independent contract is shared/tool-activity.ts. Join native tool parts and activity by conversation ID plus native toolCallId; render generic cards from the contract without importing tool implementations. Native Think continues to own tool inputs/results, transcript persistence, stream replay, backpressure, abort, and recovery. Activity metadata is not a second transcript or an execution/idempotency ledger.

Each customer-owned conversation stores only safe normalized metadata in flarebot_tool_activity: kind, summaries, status, server observation timestamps, optional progress, and observed execution count. Native Agent state publishes at most 32 records, prioritizing active calls then newest calls. Older metadata is retained with its original timestamps; the owner-only native callable listToolActivities(before?) returns up to 100 records with an exclusive nextCursor (null at the end). Cursors order immutable creation sequence, so progress updates do not shift pagination. Both metadata and state stay behind existing owner/origin/socket-expiry gates. Client state injection is rejected. A UI mounts this state/history alongside the native conversation transport and must invalidate old asynchronous reads when switching conversations.

New tools register server-owned ToolActivityDescriptors through getToolActivityDescriptors(). Select explicit safe fields in formatters; never serialize raw inputs, results, URLs with credentials/query strings, commands, headers, or errors. Unknown tools use a generic label. Text is limited to 160 characters and control characters become spaces. These bounds do not redact secrets from arbitrary formatter output: allowlisting belongs to the descriptor. Tool-specific outcome classifies scalar envelopes (for example fetch ok:false or paused Code Mode); native action error envelopes are always failures even when the SDK delivered them as successful results. applicationToolNames() enables only explicit getTools() and getActions() names, including action name overrides. Workspace, MCP and client tools remain disabled. Production registers explicit memory tools/actions; deterministic test tools are fixture-only.

Server executions move pending → running → succeeded/failed/cancelled. Pending means input is arriving or an explicitly classified result awaits continuation; running marks the server execution hook, not a claim that an external side effect has started. Provider-owned calls can have zero observed server attempts and no startedAt. Native preliminary generator results update progress; scalar tools and actions may call reportToolProgress(toolCallId, value). The descriptor selects safe progress text/counts. The first progress report is immediate, subsequent reports are coalesced to four writes per second, and the latest pending value is flushed with the terminal update. Progress does not reset Think's native stream-stall watchdog; slow tools must set an appropriate bounded turn timeout.

Abort listeners persist cancellation before native stream draining stops. The cancelled summary deliberately means the conversation stopped waiting; browser or shell tools must separately stop their external resources. Late completions cannot overwrite cancelled or known failed/succeeded records. The response hook only reconciles IDs in its own native message, including partial input on abort, so it cannot cancel a newer turn. Isolate restart marks unfinished observations failed/interrupted with an explicitly unknown outcome. Native recovery may then execute the same call again: only this interrupted outcome can reopen on a real server start, incrementing attempts, or resolve from an authoritative native result. createdAt/first startedAt survive; interruptedAt records the most recent interruption. Native recovery owns replay; this does not guarantee exactly-once external work.

Both native clear paths call the documented protected resetTurnState seam. The activity adapter synchronously removes all presentation metadata there, leaving ID-only tombstones to reject already-running callbacks; subsequent turns keep their new rows. Conversation deletion wipes the whole native facet, including tombstones. Do not override lifecycle hooks without composing the activity superclass, or use resetTurnState as a lightweight stop: use native cancelChat/cancelAllChats for cancellation that preserves history.

pnpm test:activities runs real native tools/actions in local workerd with owner cookies, WebSocketChatTransport, native preliminary results, parallel execution, errors and semantic failure envelopes, progress coalescing, late completion after cancel, detach/reconnect, paginated retention beyond the live window, full restart and same-ID recovery, and clear during an uncooperative tool. It checks private metadata access, client injection rejection, bounded safe summaries, and native transcript/activity agreement. No live provider or external execution is used.

Explicit memory #

PersonalAgent owns one private flarebot_memories table shared by all native conversation facets. SQLite FTS5 (flarebot_memories_search) indexes its active facts; application triggers update that index in the same statement as each write. These are application tables accessed through native Agent.sql, never SDK-owned tables. No vector pipeline, conversation harvesting, external memory service entitlement or extra agent identity is required.

The authenticated Settings owner connection exposes listMemories(), addMemory(content), updateMemory(id, content, version) and deleteMemory(id, version). Facts contain only { id, content, version, createdAt, updatedAt }. Limits are 200 active facts, 1,000 UTF-16 code units per fact, and strict lowercase UUID-shaped server IDs. Blank content, disallowed controls and invalid versions fail before writes. Updates compare the version, so concurrent edits fail instead of overwriting changes. Deleting an absent valid ID is idempotent. A stale edit never inserts a replacement fact. Settings keeps failed drafts, provides explicit reload, and locks reload/edit/mutation operations against one another. The existing owner preflight, native socket and lifecycle are shared with the instructions editor; no fact enters public SSR HTML or native generic state.

Think registers remember, updateMemory and forget as native actions, and recall as a native read-only tool. All four have static, bounded memory activity descriptors; fact contents are present only in intentional private tool results and model context, never generic activity labels. Actions validate the current conversation in the parent before its synchronous write and check native abort signals around preparation. Clearing a turn invalidates pending preparation. A write already dispatched to the parent may complete before cancellation; cancellation does not promise to undo a committed fact.

For remember, a deterministic ID is derived from the native conversation name and tool-call ID. That ID makes the parent mutation idempotent even if its reply is lost before Think settles the native action ledger. Deletion clears content and the FTS entry but retains an ID/version/timestamp tombstone, so later retries cannot recreate it. These small tombstones persist beyond the active-fact limit. Native settled action replay may still contain its historical result. Update version checks prevent duplicate or stale mutation; an ambiguous lost update reply can require a fresh recall/reload to confirm the outcome. Native action replay and parent writes are separate commits, not an exactly-once transaction.

At every beforeTurn, the current last user message supplies a bounded search query alongside fresh custom instructions and model configuration. Up to 24 quoted Unicode word tokens (common English stop words removed) use FTS OR and BM25 ranking with English stemming. Queries are limited to 1,000 characters and results to eight facts, each at most 1,000 characters; punctuation/operators are never interpolated as SQL or raw FTS syntax. This is lexical relevance, so the model can call recall with better search words when needed. JSON-encoded facts are appended to the complete effective instructions override with an explicit untrusted-data boundary. They grant no authorization or additional tools. Current edits and deletions affect future retrieval in existing and new conversations, including after restart. Historical transcripts, native action results and platform backups may still contain old fact text.

pnpm test:memory exercises actual native model/tool/action execution, shared relevance, current instruction composition, versioned update/delete, lost-reply replay, cancellation, late update after deletion, FTS escaping, limits, owner boundaries, state privacy and full namespace restart. pnpm test:settings also covers packaged-Worker memory CRUD, persistence, failure retention, mobile layout and SSR privacy. No live inference is used by these fixture tests.

References: Think actions and native SQLite storage. The installed Agents 0.22 AgentSearchProvider provides set/search but no public list/delete; app-owned native SQLite permits the complete inspect/edit/delete contract without mutating its internal storage.

Web research #

Conversation.getTools() composes web_search and read_url with memory tools. Search uses the existing AI binding with openai/gpt-5.4-mini and native Responses API web_search through AI Gateway default, independently of the conversation's selected model. There is no Bing/DuckDuckGo scraping, browser acquisition, or search-specific API key. Cloudflare manages upstream credentials; model and search usage require AI Gateway credits (or its configured default BYOK). See Cloudflare's web-search contract.

Queries are capped at 500 characters and returned sources at five. Each request forces web search, permits one hosted tool call, uses low search context and low reasoning effort, and caps output at 2,000 tokens. There are no application retries or scraper fallbacks. Gateway caching and content logging are disabled for this request, and Responses storage is disabled. A single 60-second deadline covers the binding request and response body; cancellation is forwarded to the binding and cancels body consumption. Upstream work already performed may still be billed. Responses are limited to 256,000 bytes; generated summaries to 6,000 characters. Billing, access, rate-limit, malformed/incomplete response and timeout failures return safe errors, never empty success. An empty source list is accepted only when a completed search explicitly supplies an empty provider source list.

read_url uses native Think createFetchTools for GET, redirect validation, download limits and cancellation. Download limit: 256,000 bytes; returned page text: 16,000 UTF-16 code units; deadline: 15 seconds. HTMLRewriter parses bounded HTML without executing JavaScript, excludes scripts/styles/navigation/forms, and entities decodes assembled text (HTMLRewriter preserves entity spelling). Markdown/plain text are preserved; binary, PDF and other unsupported types fail explicitly. This is general page-text extraction, not a semantic article reader. Private/local hostname and literal checks apply on native redirects; initial and citation URLs additionally reject all IP literals and credentials. These checks are not DNS-resolution or rebinding protection for arbitrary public names.

Think 0.17.0 needs the committed one-line pnpm patch in patches/@cloudflare__think@0.17.0.patch: awaiting finalizeResponse keeps its request timer and caller abort listener alive through the body read. Do not remove this patch on an SDK upgrade until the slow-body timeout/cancel regression passes. No custom downloader, transport, or inference-stream override is introduced.

Source records in ordinary native tool results contain stable URL-derived IDs, title, requested/final URL, fetched-at timestamp, source kind (search or page), content and truncation. Search sources come only from provider citation annotations and search-action source metadata, with validated/deduplicated URLs. Search source content is empty: the Responses API does not supply per-source page excerpts. Its separately labeled generatedSummary is model synthesis, not a verbatim excerpt or proof that Flarebot read a page. Follow-up read_url calls supply page evidence. No publication date is inferred. The per-turn instruction asks for Markdown citations to exact final URLs and treats source text as untrusted evidence. Native Think history owns persistence and reconnect; its current stream omits source UI parts, so the later citation UI should consume these tool-output records. Activity state contains static labels and classified outcomes, never URLs, queries or page text.

pnpm test:web runs native fixture-model tool calls in real local workerd, HTML/entity/redirect/size/type/error/slow-body checks, provider response fixtures, search request budgets, citation validation, summary bounds, billing/access errors, request/body cancellation, the actual search deadline, shared browser lifecycle races, safe activities and reconnect/full-restart source persistence. CI uses no paid inference and does not depend on an external search frontend.

FLAREBOT_WEB_LIVE_SMOKE=1 node --test tests/web-search-live.test.mjs opts into one paid search using the developer's Wrangler authentication. It starts a local Worker with a remote AI binding through unstable_startWorker, without deploying. The live smoke returned three official Cloudflare documentation sources on September 6, 2026. It does not verify another installation's credits or model access.

Rendered browser research #

browser_read({url, waitForSelector?}) complements read_url when evidence needs JavaScript. Each invocation acquires its own Browser Run session through an internal parent RPC using the public agents/browser create/connect/delete primitives and native CDP command handling. The model can choose returned links for subsequent independent reads. There is no shared login/session state or model-authored execution code. The existing BROWSER binding is sufficient; no new deployment resource or runtime class is required.

The tool waits for load and a short text/title stability window (at least one second after load), or for an optional CSS selector. These are bounded readiness heuristics, not a promise that every SPA has finished loading. Use a selector or retry if evidence still contains loading placeholders. One absolute 30-second budget covers acquisition, navigation, waiting and extraction. A fresh five-second budget closes the remote session; metadata acknowledgement is separately bounded. Cancellation stops waiting and attempts immediate owned-session deletion. Late create/connect replies cannot start navigation after cancellation or native clear.

A narrow observer on the native binding WebSocket handles CDP request events; CdpSession still owns command correlation and timeouts. HTTP(S) requests and redirects are checked using the existing public hostname/literal policy before continuation. Additional targets are paused and closed. This is not DNS rebinding protection. Extraction runs a fixed host-authored expression in an isolated world; selectors are JSON-serialized data. Final source provenance comes from CDP's main frame URL, never an untrusted canonical tag. Results contain up to 16,000 text characters, a 240-character title and 20 validated follow-up links. Arbitrary HTML, page errors, logs and browser session IDs are not broadcast as activity metadata. Sources persist in native Think tool output with sourceKind: "browser".

Only browser_read uses these cleanup hooks; AI-backed web_search creates no browser lease. The parent owns acquisition and keeps its continuation alive with native waitUntil: deleting a child facet destroys that child's continuations, so a late creation reply must be recorded and cleaned by the surviving parent. The child receives only its own session ID and controls its commands; no shared current browser or transferable AbortSignal crosses the native RPC boundary. The parent's private flarebot_browser_leases records known session IDs before connection, with the original absolute expiry. A native Agent schedule attempts closure at expiry; it never extends the execution deadline. Successful deletion removes the record and corresponding schedules. Failed scheduled cleanup retries at most three times with five-second request bounds; an exhausted private record remains unresolved. Startup reconciles prior unexhausted leases immediately. Conversation deletion closes its known browsers before native facet teardown; parent records and cleanup schedules survive deletion, including late acquisition. No cleanup callback resolves a deleted child. Native clear invalidates the call's generation, and normal tool cancellation still closes only its own browser.

An isolate interruption runs no JavaScript finally. Native schedules may run late, and a lost creation response can hide a remotely accepted session ID. Browser Run's 60-second inactivity expiry is the final backstop, not a claim of an exact remote hard shutdown time. Local deadline enforcement, attempted cancellation and confirmed remote deletion are distinct. The generic activity view reports safe browser progress and terminal status; uncertain cleanup returns a structured cleanup_failed result instead of claiming success.

pnpm test:browser runs native Think fixture inference and real local Wrangler Chromium. It checks delayed JavaScript evidence absent from the plain reader, follow-up links, final provenance, bounded text, selector and HTTP failures, private redirects, progress, cancellation, separate concurrent sessions, timeout, late acquisition during deletion, cleanup retries/native schedules and restart reconciliation. Chromium also stops on local runtime shutdown; the restart test checks persisted cleanup intent and idempotent deletion, not remote service survival. This gate makes no Cloudflare account deployment or live provider call.

Temporary shell #

The registered shell tool uses @cloudflare/sandbox@0.12.9 and the matching immutable native Linux image. Each invocation gets a fresh customer-owned container with Bash, Node.js and Bun. A command can write scripts and input data with heredocs, then process those files within the same invocation. Python is not included in this image. The workspace is not a project checkout, file browser or artifact store. No customer Worker/provider secrets are injected and no container preview routes, tunnels, mounts or secret bridge are configured. Commands can access the public network; external side effects are not rolled back or made exactly-once by cancellation or Think recovery.

The input command is at most 32 KiB UTF-8, the default absolute lifetime is 30 seconds (including cold start), and the caller may request 1–45 seconds. Combined retained stdout/stderr is at most 32 KiB. Reaching the output limit triggers container destruction, as do completion, nonzero exit, cancellation, lifetime expiry and conversation deletion. Four active or unresolved workspaces are allowed per installation. The native lite instance type bounds container CPU/RAM/disk; the application does not pretend those are per-process quotas. Files and background processes are temporary and cannot be accessed by later invocations. Durable chat and explicit memory remain in the native agent stores.

The parent persists an opaque, server-generated lease before external work and owns native launch settlement outside the deletable conversation facet. The Sandbox subclass adds only an absolute expiry and durable one-use/cancellation record around native execution. This protects late launches when the parent restarts; Sandbox remains the native Containers implementation and owns its own waitUntil. Parent native Agent schedules retry cleanup; startup reconciles old leases rather than attaching recovered Think calls to old workspaces. Confirmed closure removes the lease. A failed or slow destroy returns cleanup: "pending", keeps the private record/capacity reservation, and retries without claiming that the process has stopped. If execution succeeded its exit code/output stay available and the activity summary states cleanup is pending. Platform outages can delay cleanup; native one-minute inactivity sleep is an additional backstop, not a hard execution deadline. Preserve the Sandbox class and namespace on upgrade.

execTemporary delegates to native execStream; the child consumes native SSE with parseSSEStream and emits ordinary AI SDK async-generator preliminary tool results. These bounded cumulative stdout/stderr values travel and persist through Think's native tool-output parts. Generic activity state/SQLite retain only safe labels, byte counts and status, never commands or stdout/stderr. Future chat UI expands the native tool parts alongside activity; no second transcript, custom WebSocket protocol or shell-log database is introduced.

Run pnpm test:shell with Docker running after pnpm build:release. Its fake model invokes the real registered tool through native Think/AgentClient in local Wrangler and uses the pinned Linux image. This is local container execution; it does not certify customer-account Containers provisioning or live inference. The other runtime fixtures keep Wrangler's default container execution disabled while retaining the actual release export/configuration contract.

References: stable streaming, lifecycle, local development.

Scheduled tasks and durable execution #

The personal parent owns task definitions and a small history projection in flarebot_tasks / flarebot_task_runs. An owner can createTask(input), getTask(id), listTasks(), updateTask(id, expectedVersion, input), deleteTask(id, expectedVersion), runTaskNow(id, requestId) and listTaskRuns(id, { limit?, before? }) over the existing authenticated native Agent client. Task creation takes a stable UUID id, an existing active conversationId, name, instructions, schedule and boolean enabled. Updates replace the four editable fields and use the returned version; the conversation target is immutable. Names allow 120 characters, instructions 8,000 and the installation holds at most 100 active task definitions. Task content is private RPC data, absent from public SSR and native state broadcasts.

Schedules are either { kind: "once", at: "2096-02-29T09:30:00Z" } or { kind: "cron", expression: "0 9 * * 1", timezone: "UTC" }. One-offs require a valid future whole-second instant with Z or an explicit offset; they normalize to UTC. Invalid calendar dates, fractional seconds, zone-less times and unknown offsets (-00:00) are rejected. Cron accepts five numeric fields with stars, lists, ranges and steps on stars/ranges, at most 120 characters. Seconds, nicknames, named weekdays/months and non-UTC zones are outside this contract. The public cron-schedule@6.0.0 parser also used by Agents calculates future occurrences in the Worker. Its environment-local Date arithmetic runs in UTC there; this helper must not be moved into browser code or presented as local wall-clock recurrence. Whitespace, numeric leading zeroes and explicit one-off offsets are normalized before checking creation identity.

nextRunAt comes from the actual native schedule's epoch-second time, read through getScheduleById. Native schedule rows are the binding registry. A missing binding is reconciled from task intent; if arming remains unavailable, the response contains nextRunAt: null and safe schedulingError: "schedule_unavailable". Disabled tasks and consumed one-offs have no next occurrence. An enabled overdue one-off retains its intended instant for later execution reconciliation. previousRun means the latest stored actual occurrence, ordered by creation time then ID; an earlier cron calendar match is never fabricated as history. History pages use the same descending keyset order, default to 25 and allow at most 100 entries. No transcript or raw model/tool errors are copied into history; the run links to its conversation and deterministic native submission ID.

beginTaskRun and projectTaskRun are internal, non-callable integration seams. Scheduled identities include task ID, version and intended instant; manual identities use task ID and a stable owner request UUID. Persist a placeholder before dispatch and use native submission inspection/status hooks to project the result. A late pending receipt cannot overwrite running/completed history, and unknown/deleted/mismatched reports cannot create rows. An uncertain dispatch failure can still be repaired by later native acceptance. A manual run does not consume a scheduled one-off. Once consumed, name/instruction edits preserve that state; rearming needs a new future instant. Dispatch gates the current definition and existing native conversation before and after child RPCs. Native onSubmissionStatus projects durable results, and rechecks pending/running tasks before Think applies queued messages. Owner reads and a bounded native maintenance schedule repair missed observer reports using inspectSubmission. beginTaskRun returns a prior occurrence on replay even after its task version or enabled state changes. The dispatcher must independently recheck the current task version, enabled state and conversation lifecycle before submitting that replay.

Creation keeps a hash of the normalized initial request, so retries after edits return the current task while conflicting reuse fails. Deletes retain only a content-free task ID tombstone and privately preserve unresolved run/conversation/ submission IDs until native cancellation resolves. Native storage transactionSync groups occurrence/one-off consumption and deletion writes, so a failed later SQL statement rolls the whole mutation back. Terminal history is erased; other history payloads are stripped with set-based SQL. Deleting an absent ID reserves it against an in-flight creation. Task deletion preserves the shared conversation and its Think transcript. Conversation deletion tombstones referencing tasks synchronously before its first cleanup await; retries/restarts cannot resurrect tasks or lazily recreate the deleted facet. Disable/edit/delete cancels stale native schedules and targets only the task's pending/running submission IDs with cancelSubmission, preserving unrelated conversation turns. Confirmed native facet deletion erases its private cleanup references.

pnpm test:tasks runs real workerd owner RPCs, UTC calculations, strict input and CAS/replay checks, deterministic mixed-source history pagination, full Worker restart persistence and deletion failure/recovery. Its separate fixture entry seeds internal occurrence projections solely to test the model; this is not unattended execution evidence. Native schedule and Think tables remain untouched.

Agent.schedule(Date | UTC cron) runs without a browser connection. Cron skips missed occurrences and advances after dispatch returns; there is no backfill. Native due callbacks are serialized, so a long-running native drain can delay another alarm; due times are intended wake times rather than latency guarantees. Each callback submits one ordinary user message to its existing Conversation Think facet, retaining the normal model/provider, instructions, memory, tools, transcript, streaming and durable recovery. Acceptance is queued work, not a completion report. runTaskNow persists its UUID-keyed occurrence and native one-shot dispatch before replying; retries return the same occurrence, even after completion. A manual run can execute a disabled task and never consumes/shifts its scheduled occurrence. Editing/disable after that request invalidates the old run.

Native callback retries cover acceptance (three attempts). A placeholder older than ten seconds whose submission is absent can receive one further native recovery batch, reusing the same submission ID; a private boolean bounds that batch. This closes the parent crash gap before acceptance without a replacement queue. Native terminal inference errors remain visible as turn_failed; they are not replayed with fresh submission IDs. Think's provider retries and chatRecovery remain authoritative. If a stopped accepted turn lacks sufficient native continuation evidence after applying its prompt, Think seals it as a safe terminal error; task history preserves that outcome and the original submission identity rather than replaying possible tool side effects. Subsequent cron occurrences continue normally. A recurring thirty-second native reconciliation callback pages through unresolved projections with a five-second processing budget. Reads target the requested task (or a small page for lists), and parent startup arms reconciliation without awaiting child initialization, avoiding parent/child recovery cycles. Once no future task or unresolved cleanup remains, the maintenance schedule is removed.

pnpm test:execution uses real workerd Agents alarms and Think submissions with a fixture model, including no-client one-offs, restart before deadline without requests until after due, two real minute-based UTC cron occurrences, acceptance reply loss, missed status reports, manual dedupe, stale callback and queued/running cancellation, inference errors, and restart between placeholder and acceptance. Fixture routes and faults are separate test exports and absent from the production release. This verifies native local execution, not a live external provider or a deployed customer account. The task UI and conversational schedule creation are separate issues.