Bug lessons #
Durable, evidence-backed lessons from debugging sessions in this repo. Symptom-match new bug reports against these entries before theorising.
2026-09-06 — Public Worker-to-Worker requests require explicit routing #
- Affected area: Publisher readiness fetch and customer-to-publisher login bridge; both Wrangler compatibility flag lists.
- Symptom signature: Worker and Sandbox provisioning completed, but verification failed immediately. The signed customer health endpoint passed from an external client, while the same request from a cloud Worker returned HTTP 404 with Cloudflare error
1042. - Root cause: Neither Worker enabled
global_fetch_strictly_public. Native fetch could not use the other Worker's public workers.dev endpoint. This routing flag is not implied by the compatibility date ornodejs_compat. - Resolution: Enable the flag on both publisher and customer deployments. Publish the changed customer contract as
0.1.0-dev.2, retaining the originaldev.1archive and an explicit forward-upgrade edge rather than modifying its immutable bytes. - Regression signal: The same signed cloud-Worker probe returned error
1042before the flag and full readiness HTTP 200 after it. Configuration/artifact tests require the new release's flag and retain support for the exact legacy flag list needed to recover and upgrade existing installations. - Prevention rule: Test public Worker-to-Worker calls from a deployed Worker in each direction. A successful external HTTP request or local service fixture does not certify Cloudflare's edge routing; maintain explicit routing requirements in both deployment contracts.
2026-09-06 — Native fetch rejected the deployment client's receiver #
- Affected area:
control-plane/deployment-api.ts, account preparation and uploaded Worker verification. - Symptom signature: Installation stayed at “Preparing your account”; the native Workflow exhausted four immediate
temporarily_unavailableattempts inresolve customer origin, before deploying resources. - Root cause: The adapter stored ambient Workers
fetchon the client and invokedthis.network(...). Nativefetchrejected theDeploymentAPIreceiver with “Illegal invocation” before sending a request; the adapter sanitized that exception. Arrow-function provider fixtures did not enforce the native receiver constraint. - Resolution: Call the injected transport as a standalone function in both the JSON request and multipart content verification paths. Preserve the fixed endpoints, credentials, bounds and error classifications.
- Regression signal:
pnpm test:deployment-networkexercises the actual adapter and native Workersfetch, replacing only outbound service responses. Before the fix, account preparation fails without reaching the fixture endpoint; the corrected transport resolves the origin and verifies uploaded content. - Prevention rule: Tests for an injected platform primitive must preserve its native calling convention. Exercise default transports in the target runtime as well as arrow-function substitutes; receiver-sensitive APIs cannot safely be invoked through arbitrary owning objects.
2026-09-06 — Cloudflare callback scope metadata rejected valid sign-ins #
- Affected area:
control-plane/http.ts, OAuth callback query validation. - Symptom signature: Every real Cloudflare sign-in returned
/connect?error=oauth_invalid_callbackand “This sign-in link is invalid or expired. Connect to Cloudflare again.” even with a fresh transaction. - Root cause: Cloudflare returned
code,scope, andstate; the callback allowlist rejectedscopebefore claiming the transaction or exchanging the code. Both the direct and browser fixtures had omitted this provider field. - Resolution: Accept a single callback
scopeparameter as non-authoritative metadata. Continue validating granted scopes from the token exchange and retaining state, cookie, issuer, duplicate-query, expiry and one-use checks. - Regression signal:
pnpm test:oauthreproduces the exact error before the fix and passes afterwards through the native Worker/DO callback and Chromium navigation. It also rejects duplicate scope parameters and a token response missing permissions despite a complete callback scope list. - Prevention rule: Model the real provider's callback shape in both HTTP and browser fixtures; distinguish provider metadata from verified token grants. Testing OAuth endpoints separately does not validate the application's complete callback path.
2026-09-06 — Local runtime had no owner login path #
- Affected area:
worker/bridge.ts, local customer-runtime setup - Symptom signature: Opening
http://localhost:8787/auth/loginreturnedOwner authentication or installation verification failed.and the shell remained disconnected when using.dev.vars.example. - Root cause: The development template intentionally omitted production bridge keys, while the only owner-session issuer required a configured control-plane bridge.
- Resolution: An explicit development runtime on a loopback origin issues its local owner session from
/auth/login; non-loopback and production installations retain the signed OAuth bridge. - Regression signal:
pnpm test:runtimeasserts that the actual packaged local Worker redirects/auth/login, sets the hardened owner cookie, and accepts it on the private status route. - Prevention rule: Every documented local startup path that exposes authenticated UI must include a bounded way to establish its development identity.
2026-09-05 — Sidebar SSR flash "window is not defined" #
- Affected area:
octane-kumoSidebar.Provider→useIsMobile(packages/octane-kumo/src/components/sidebar/sidebar.tsx); surfaced in flarebot'sRootLayout, which rendersSidebar.Provider. - Symptom signature: First paint shows the root error boundary
("Something went wrong!
window is not defined"), replaced by the real shell after hydration. The SSR HTML of every route contains the error text. - Root cause:
useIsMobile'suseSyncExternalStoresnapshot called unguardedwindow.matchMedia. Octane's serveruseSyncExternalStorefalls back togetSnapshot()when the compiled call carries no slot arg — and the server hook-slot transform wraps the nesteduseIsMobile()call site withwithSlotbut does not inject slots into its body (verified indist/server/entry.js: 3-arg call, no slot; the client bundle for the same source has the slot). So SSR threwReferenceError: window is not defined. - Resolution: Guarded
getSnapshotwithtypeof window === "undefined" ? false : window.matchMedia(query).matches, matchinggetServerSnapshot. Fixed inoctane-kumo, consumed via thelink:dependency. - Regression signal:
packages/octane-kumo/tests/sidebar-ssr.test.tsx— SSR-rendersSidebar.Providerwithwindowstubbed toundefined; fails pre-fix, passes post-fix. Manual loop: startpnpm dev, thencurl -s http://localhost:5176/about | grep -c "window is not defined"must print0. - Prevention rule: Every
useSyncExternalStoregetSnapshot— and any render-path browser-global access — in Octane-ported components must be SSR-safe (typeof window/documentguards). Never rely ongetServerSnapshotalone; the server transform may drop it.
2026-09-05 — Native facet client never becomes ready #
- Affected area:
AgentClient.basePath/ PartySocket URL construction. - Symptom signature: Authenticated facet history works, but
client.readyremains pending with no identity frames and WebSocket handshakes return 404. - Root cause: Passing
/agents/...asbasePathgeneratesws://host//agents/...; PartySocket prepends its own slash. - Resolution: Pass the authenticated pathname with its first slash removed
to
basePath; preserve the slash for HTTP requests. - Regression signal:
pnpm test:thinkawaits bounded native facet readiness and completes streamed turns usingbasePath: pathFor(id).slice(1). - Prevention rule: Treat a PartySocket base path separately from an absolute HTTP pathname. Do not loosen the server's exact route guard for malformed URLs.
2026-09-06 — Tool activity recovery and clear use different native paths #
- Affected area:
worker/tool-activity.ts, Think tool hooks and native clear. - Symptom signature: An interrupted call recovers with the same tool-call ID,
but a blanket terminal-state guard leaves its activity failed while the native
transcript succeeds. A public
clearMessagesoverride alone also misses WebSocket clear because Think invokes its own clear handler. - Root cause: An interrupted observation has an unknown outcome, unlike a
known tool failure or explicit cancellation. Think 0.17's two clear paths both
call the documented protected
resetTurnState, but WebSocket clear does not dispatch through publicclearMessages. - Resolution: Permit only interrupted observations to reopen on authoritative execution/result evidence, preserve first timestamps and count observed attempts. Strip activity presentation metadata synchronously at the native reset seam, retaining ID-only tombstones against late callbacks.
- Regression signal:
pnpm test:activitiesrestarts real workerd during a tool, reissues that exact native call ID and compares transcript/activity success; it also clears native chat while an abort-ignoring tool is running and checks that a later turn survives without restored old summaries. - Prevention rule: Distinguish known terminal outcomes from interrupted observations, and verify both native HTTP/RPC and WebSocket lifecycle paths before choosing an override. Tests must actually reissue the same call ID to exercise replay, rather than merely letting the recovered model finish text.
2026-09-06 — Think fetch timeout stopped after response headers #
- Affected area: Think 0.17.0
dist/tools/fetch.js,executeRequest. - Symptom signature: A server returns headers and an initial body chunk, then stalls; native timeout and caller cancellation no longer interrupt the body.
- Root cause: Returning
finalizeResponse(...)without awaiting it exits the surrounding try/finally immediately, clearing the request timer and removing caller abort forwarding beforereadCappedfinishes. - Resolution: Version-pinned pnpm patch adds
awaitat that return. Native limits, redirect filtering and download code remain authoritative. - Regression signal:
pnpm test:webfeeds a controlled slow body to the actual native fetch tool and verifies timeout plus underlying signal abort, then stops a native Thinkread_urlinvocation during body consumption. - Prevention rule: When cleanup releases cancellation or resource ownership, await asynchronous body processing before leaving the protected scope. Retest slow bodies before removing the patch on a Think upgrade.
2026-09-06 — Browser acquisition outlived its conversation facet #
- Affected area: native facet deletion and Browser Run session acquisition.
- Symptom signature: Deleting a conversation while browser creation was
awaiting its response left the actual remote session open. The child's late
continuation logged
Facet was deleted; capturing a parent RPC stub alone did not keep that continuation alive. - Root cause:
deleteSubAgentdestroys child execution and its pending continuations. A child-ownedfinallyorwaitUntilcannot guarantee cleanup after the facet itself is destroyed. - Resolution: The surviving parent owns the native create request and its
waitUntil, records the returned ID before connection, and rechecks whether the conversation still exists. A late result for a deleted conversation is closed directly. The child still owns only its session's browser commands. - Regression signal:
pnpm test:browserdelays a real local Chromium create reply, deletes the conversation, and probes that exact session for HTTP 404. - Prevention rule: Own external acquisition in a lifetime that survives its caller's deletion. Cleanup intent must survive alongside that owner; remote creation whose ID is lost still requires an honest service-expiry fallback.
2026-09-06 — Stable Sandbox session buffering defeated output limits #
- Affected area: Sandbox 0.12.9
getSandboxand shell streaming. - Symptom signature: Direct subclass
execStreambuffered an unterminated output line until the deadline; explicitexitended the SDK's persistent session without an ordinary command completion event. - Root cause:
getSandbox(..., { enableDefaultSession: false })implements stateless execution in its public helper wrapper. Callingthis.execStreaminside a subclass bypasses that wrapper and uses a persistent shell session. - Resolution: The one-use Sandbox lifecycle wrapper delegates through the public stateless helper, verifies its own namespace identity, and preserves native SSE byte chunks and terminal exit codes.
- Regression signal:
pnpm test:shellchecks nonzero exit and floods real Docker stdout without newlines; the byte cap must destroy the container before its time deadline. It also checks partial-output cancellation. - Prevention rule: Test exact streaming semantics with unterminated output, explicit exit and silence. Wrapper options are not necessarily stored runtime configuration; retain the SDK helper that implements them.
2026-09-06 — Native one-shot cleanup retry deduplication #
- Affected area: Agent native scheduled resource cleanup callbacks.
- Symptom signature: A failed cleanup schedules an idempotent retry with identical type, callback and payload; no later retry occurs.
- Root cause: Native scheduling deduplicates against the currently executing one-shot row, then deletes that row after the callback returns.
- Resolution: Cleanup retries create a new native one-shot schedule without self-deduplication; confirmed cleanup cancels remaining matching schedules.
- Regression signal:
pnpm test:shellinjects four consecutive destroy failures and waits for actual native scheduled cleanup and stopped container. - Prevention rule: Test successive failures, not only the first retry, and account for scheduler row advancement when rearming inside a callback.
2026-09-06 — Local public container images still trigger Wrangler auth #
- Affected area: Credentials-free native Worker and Docker tests.
- Symptom signature: CI fails before starting a local Worker with an account or token error, although local tests pass with a developer login.
- Root cause: Wrangler 4.128 normalizes container image references and fills registry API authentication even when local container execution is disabled.
- Resolution: The shared test fixture supplies a synthetic account/token and redirects Cloudflare API requests to closed loopback. Optional registry login fails locally; Docker pulls the pinned public image directly. Production deployment configuration stays account-neutral.
- Regression signal: Worker and real Docker shell tests pass with an isolated Wrangler config directory and no real Cloudflare credentials or API access.
- Prevention rule: Verify native local tests without developer authentication; disabling remote execution does not necessarily disable config-time auth.
2026-09-06 — Chat UI artifacts and failed-case cleanup were host-dependent #
- Affected area:
tests/chat-ui.test.mjsscreenshot capture and native client lifetime. - Symptom signature: Unprivileged Linux reports
EACCES: permission denied, mkdir '/private'in the rich-message case; the later history case times out locating a desktop sidebar link. A failed clear case can leave the process alive after TAP has reported its assertion. - Root cause: Screenshots used a developer's macOS path. Its failure skipped viewport restoration, contaminating later cases. A native leaf client closed only on success kept reconnecting after failed assertions; held HTTP gates and route handlers also lacked failure cleanup.
- Resolution: Store screenshots beneath the test-owned temporary directory
(or explicit
FLAREBOT_CHAT_SCREENSHOTS), restore viewport and release case resources infinally, and register bounded, idempotent suite cleanup. Use TAP output so the original assertion is visible immediately. - Regression signal:
pnpm test:chat-uiin an unprivileged Linux container exercises the same screenshot and desktop-navigation sequence. Injecting an assertion after the clear case opens its client and holds HTTP made the old harness hit an external 25-second timeout; the corrected harness reports the intentional failure and exits nonzero normally in about three seconds. A separate pending-body injection exits after its 10-second parent timeout, confirming cleanup is independent of the suspended test body. - Prevention rule: Derive test artifact paths from
tmpdir()or an explicit caller path. Register resource cleanup before awaiting readiness, and verify failed assertions release native reconnecting clients as well as browsers.
2026-09-06 — Completed resume observer overlaid authoritative history #
- Affected area:
ConversationSessionnative resume ownership and terminal history. - Symptom signature: After full Worker restart, native history contains a partial assistant and one separate continuation, but the connected UI also appends the continuation to the earlier partial. Native IDs match while text remains duplicated, even after waiting.
- Root cause: An unsolicited fallback observer could become transport-owned during resume. Owned terminal frames bypassed the broadcast state transition, leaving its accumulator observing after the stream finished. The final fresh HTTP history was then overlaid with that obsolete accumulator.
- Resolution: Retire the observer only when its stream ID matches the completed native request, following the SDK's own owned-response handling. Preserve native partial and continuation rows and let final history replace their text without an obsolete overlay.
- Regression signal:
pnpm test:chat-uicompares each rendered message's ID and text parts against native history at Connected after a real Worker restart, in addition to requiring the continuation text exactly once. The Linux failure reproduced with fresh history revision unchanged and an obsolete observer; the native mock had made only one continuation call. - Prevention rule: When streaming ownership changes, terminal cleanup must retire both ownership paths for that request. Compare authoritative message content as well as IDs, and never clear a different active stream's observer.
2026-09-06 — Unrouting raced an intercepted history response #
- Affected area: Clear and stale-selection HTTP gates in
tests/chat-ui.test.mjs. - Symptom signature: Linux CI fails Clear with
route.fulfill: Route is already handled!; parent cancellation then produces closed-page cleanup errors. The same gate can pass on another runner. - Root cause: UI readiness did not mean every intercepted history handler
had finished. A later non-aborted handler was still awaiting
route.fetchwhenpage.unroutedisabled interception; its subsequent fulfillment raced Chromium's handling of that request. - Resolution: Both held-history cases use the public
page.unrouteAll({ behavior: "wait" })within the existing deadline. Active handlers finish before interception is disabled; route errors are not ignored on the success path and stale-history assertions remain intact. - Regression signal: The actual Clear case in an unprivileged Linux container
reproduced the exact error with delayed fulfillment. Changing only route
teardown to await handlers made that same narrowed case pass and exit normally.
pnpm test:chat-uiretains both held-history acceptance cases. - Prevention rule: Await routing work itself before removing interception. A rendered ready state and release of a gate do not prove its asynchronous callback has finished.
2026-09-06 — Shell reconnect selector matched the conversation control #
- Affected area: Session-expiry recovery in
tests/app-shell.test.mjs. - Symptom signature: Playwright reports two matching Reconnect buttons after both shell and conversation connections learn that the session expired.
- Root cause: The shell test used a page-wide action selector. Timing could expose either one or both independently owned reconnect controls.
- Resolution: Scope shell actions to
.connection-notice. The signed-session expiry case establishes both connections, awaits their native expiry, then verifies only shell Reconnect restores both. Cookie removal remains a separate new-request authorization check: it does not revoke an accepted native socket. - Regression signal:
pnpm test:app-shellreproduced the original strict-mode failure. Preserving the chat socket while delivering offline/online events also reproduced the mistaken expectation that cookie removal must close chat. The real signed-expiry regression passes without relying on network disconnect timing. - Prevention rule: Scope repeated actions to their component, and trigger the actual authorization event before expecting an established socket to expire.
2026-09-06 — Held task page exceeded the native RPC deadline #
- Affected area: Older-history repair case in
tests/tasks-ui.test.mjs. - Symptom signature:
Captured older page appended: 25, with the UI showingCould not load older runs. Try again.on slower CI runs. - Root cause: The test withheld a real older-page response until 29 native executions finished. That wait could exceed the browser client's 10-second RPC deadline, so the correctly captured four-row page arrived after its request failed.
- Resolution: Pause the browser clock only during that deliberate response
hold, leaving Worker execution in real time, and resume in
finally. Assert the captured task, older cursor, four rows and active native status before release. - Regression signal: The real native case with a 12-second execution delay reproduced the exact failure and passed with controlled browser time. The test retains that delay, all 29-row append and completed-status reconciliation checks.
- Prevention rule: Deliberate transport holds must control the client's deadline independently of slow server work when the assertion concerns stale data, not timeout.
2026-09-06 — Conversation socket outlived the browser's network loss #
- Affected area:
ConversationSessioninsrc/runtime/conversation-session.ts; surfaced in the session-expiry recovery case oftests/app-shell.test.mjs. - Symptom signature: After
setOffline(true), cleared cookies andsetOffline(false), the shell reachesSign-in requiredwhile the conversation heading still readsConnectedand never showsSign in to this installation to view this conversation. - Root cause: The shell drops its socket on the browser's
offlineevent, but the conversation only reacted to a WebSocketclose. Chromium's offline emulation (and a real loss on some networks) does not reliably close an established socket, so the conversation kept trusting a stale connection and never re-read history, which is where the 401 would have been observed. - Resolution:
ConversationSessionnow listens tooffline/online:offlinedetaches immediately and shows the reconnecting notice;onlinereconnects (queued if an aborted connection is still unwinding). Terminal states are left alone. - Regression signal: CI at head 54d813e kept the conversation
Connectedafter a brief offline toggle and cookie removal.pnpm test:chat-uicovers offline stream recovery;pnpm test:app-shellseparately verifies automatic recovery after both connections observe actual signed-session expiry. - Prevention rule: Every independently owned connection must observe the same browser network signals; a live socket is not evidence that the session is valid.
2026-09-06 — Native pending task status rejected by a test assertion #
- Affected area: Captured older-page assertion in
tests/tasks-ui.test.mjs. - Symptom signature:
The captured real older page includes active runsfails after the correct four-row older page is captured during execution. - Root cause: The assertion used
queued, but the task domain calls an acknowledged, unfinished submissionpending. Faster native acknowledgement changed the captured rows fromdispatchingto validpending. - Resolution: Check the actual active statuses:
dispatching,pending, andrunning; retain the cursor, row-count and eventual completion checks. - Regression signal: CI rejected the captured older page; a local native
capture confirmed real
pendingrows. The full task UI gate retains its active page assertion and all 29-row append/completion checks with the documented value. - Prevention rule: Read protocol and domain status values from their source; similar natural-language descriptions are not interchangeable enum values.
2026-09-06 — Referrer policy changed OAuth form Origin #
- Affected area: Public OAuth onboarding forms and exact-Origin CSRF checks.
- Symptom signature: Chromium submitted the real Connect form with
Origin: null, so the Worker correctly rejected the request with HTTP 403. - Root cause: Applying
Referrer-Policy: no-referrerto the public document also affected the Origin header on navigation form POSTs. - Resolution: Public onboarding uses
same-origin, which retains same-origin form verification without disclosing referrers to Cloudflare. Callback and API responses retainno-referrer; exact-Origin checks remain unchanged. - Regression signal:
pnpm test:oauthexercises actual Chromium form POSTs, redirect-chain provider interception, callback cookies, account selection and logout against the real Worker and native Durable Objects. - Prevention rule: Test browser navigation forms as well as direct HTTP API calls. Never weaken Origin validation to accommodate a referrer-policy mistake.
2026-09-06 — Native Workflow error messages are not a result protocol #
- Affected area:
control-plane/installation-workflow.ts - Symptom signature: Native
health_failedandresource_conflictstep failures becametemporarily_unavailablein installation metadata, despite the correct callback error appearing in local logs. - Root cause: The Workflow/RPC boundary changes nonretryable error messages; matching an exact application enum against the wrapped message loses the original classification.
- Resolution: Nonretryable callback failures return a strict safe result containing the enum. The Workflow interprets that persisted result outside the callback. Only transient failures throw a sanitized error for native retries.
- Regression signal:
pnpm test:orchestratorchecks exact durable failure codes through real local Workflow execution and confirms a failed boot never assigns an installed release. - Prevention rule: Persist declared, nonsecret step results for domain failures. Do not depend on native exception message formatting as an application protocol.
2026-09-06 — Deployment fixtures must exercise the runtime configuration consumer #
- Affected area:
control-plane/deployment-config.ts, customer installation configuration and the orchestrator fixture. - Symptom signature: The deployment adapter included
bridge.issuer, while the finished customer schema accepted onlykeyIdandpublicKey. Independent orchestrator and login fixtures passed, but the production customer loader rejected the actual uploaded variables. - Root cause: The producer retained an earlier interface proposal and the orchestration fixture verified multipart metadata without invoking its real consumer.
- Resolution: Deployment configuration now passes through the production installation parser after origin resolution. The native provider fixture also feeds the exact uploaded variables to
loadCustomerConfigandloadCustomerSecrets, using a fixed HTTPS publisher origin. - Regression signal:
pnpm test:orchestratorvalidates the actual multipart bindings with the production customer loaders before accepting a Worker upload. - Prevention rule: Independently valid subsystem fixtures do not establish integration. Exercise the real consumer against the exact serialized producer output, including production origin and secret-binding requirements.
2026-09-06 — Upgrade preflight rejected native Worker metadata #
- Affected area:
control-plane/deployment-api.ts, upgrade observation and metadata preservation. - Symptom signature: A Ready installation rejects its first upgrade with
resource_conflictbefore any upload, despite unchanged code, configuration and namespace identities. - Root cause: The provider fixture omitted default
tags, versionannotations, andscript_runtime.assets/containers. The strict observation allowlist therefore classified ordinary Cloudflare upload results as resource drift. - Resolution: Preserve Worker tags and writable message/tag annotations, excluding only read-only
workers/triggered_byprovenance. Validate the exact asset-routing defaults and installation-owned container mapping produced by Flarebot's upload. Unknown metadata and changed routing/ownership still fail closed. - Regression signal:
pnpm test:orchestratormodels the native response fields, checks preserved customer metadata through upgrade/recovery, and rejects unsupported changes before customer mutation. The read-only liveDeploymentAPI.baselinereproduction must also pass against the original installation. - Prevention rule: Build upgrade fixtures from the full native upload/read response shape. Separate writable metadata, server-generated provenance and deployment-owned settings explicitly; never fix an allowlist mismatch by dropping all unfamiliar fields.
2026-09-06 — Status-read deadline aborted valid upgrade commands #
- Affected area:
control-plane/ui/installation-status.ts, installation command requests. - Symptom signature: Continuing a saved upgrade displays “Flarebot could not be reached” while retaining the previous Ready release; ordinary status reads still succeed.
- Root cause: Commands shared the 15-second polling deadline even though upgrade/recovery performs synchronous Cloudflare resource and code verification before acknowledgement. A native read-only preflight against the deployed artifacts took 15.2 seconds. A genuine accepted upgrade response delayed 16 seconds reproduced the exact browser alert; the original user's failed POST was not captured.
- Resolution: Allow installation commands 120 seconds while retaining the 15-second read deadline, explicit cancellation, frozen request identity and replay behavior.
- Regression signal:
pnpm test:installation-statusholds the real saved-upgrade 202 response for 16 seconds, rejects a false network alert, and verifies the exact request and final installed release. Existing deliberate network-loss and recovery cases remain covered. - Prevention rule: Budget acknowledgement time for the synchronous work an endpoint performs. A status-read timeout is not automatically suitable for a command that verifies remote resources before starting durable background work.
2026-09-06 — Live missing-Workflow errors blocked startup recovery #
- Affected area:
control-plane/start-installation.ts, recovery after the registry saves an operation but before its Workflow is created. - Symptom signature: Installation stays Updating / Preparing with no matching native Workflow; continuing the saved recovery request returns HTTP 503 before any Worker upload.
- Root cause: Local native
Workflow.getthrowsinstance.not_found, while the deployed binding throws(instance.not_found) Instance not found. Recovery recognized only the local message and its RPC prefix, so the live absence signal becametemporarily_unavailable. - Resolution: Recognize the observed deployed absence message alongside the existing local forms. All other lookup/termination failures remain unavailable and cannot authorize replacement execution.
- Regression signal:
pnpm test:orchestratorexercises the real recovery endpoint through a missing native instance with the deployed error serialization and verifies that unknown lookup errors leave the saved operation and provider resources unchanged. A protected read-only live probe confirmed both the missing execution and exact remote error shape. - Prevention rule: Native service error serialization can differ between local and deployed runtimes. Capture the actual remote contract at a failing boundary and cover its known representation explicitly; never interpret every lookup failure as absence.
2026-09-06 — Script PUT rejected version-pinned inheritance #
- Affected area:
control-plane/deployment-api.ts, upgraded Worker upload and source-version checks. - Symptom signature: Upgrade upload records its intent, fails with
temporarily_unavailable, then stops withrecovery_requiredwhile dev.1 still serves. The live PUT returns HTTP 400 / code 10057: inheritedversion_idaccepts only the literallatest. - Root cause: The fixture accepted concrete version UUIDs from the shared API schema, but the deployed script PUT endpoint rejects them for every inherited binding.
- Resolution: Use strict inheritance with
version_id: "latest". Require the newest uploaded version to equal the checked active source during preflight and again immediately before PUT, including undeployed uploads. Preserve post-upload fingerprints; concurrent external uploads remain a documented non-atomic boundary. - Regression signal:
pnpm test:orchestratorrejects UUID inheritance, checks latest-version reads around staging, and blocks both preexisting and late newer uploads before PUT. The guarded real dev.2 upload changed from HTTP 400 to success with the same saved configuration and resources. - Prevention rule: Exercise the actual provider endpoint when its behavior differs from a shared schema. Do not replace a pinned source with
latestwithout checking all uploaded versions, and do not describe the resulting check/write sequence as atomic.
Upgrade operation defaults belong only at storage boundaries #
The native upgrade gate exposed a baseline-erasure bug: using a defaulted Zod storage schema with .partial() for intent mutations inserted upgrade: null and rolloutId: null into otherwise unrelated Worker-intent changes. The upload and health completed, but request replay and failed-upgrade retry lost their immutable source baseline. Use an explicit strict mutation schema with no storage defaults, excluding the immutable upgrade baseline entirely. Native replay/retry tests and an unknown-field intent rejection protect this boundary.
A ready installation is not evidence that an unsubmitted upgrade succeeded. The real browser test aborted the upgrade request before the registry saw it, then reloaded against the still-ready old installation. Reusing the initial-install ready cleanup erased the frozen upgrade target. Only clear an upgrade intent from a ready read when the installed identity matches that exact pending target; otherwise retain it for explicit replay.
2026-09-06 — Preserve request provenance in native fixture failures #
- Affected area: JSON response reads in
tests/orchestrator.test.mjs. - Symptom signature: Linux CI's upgrade process-restart case failed with
Unexpected token 'E', "Error: Net"... is not valid JSON; the Undici-only stack did not identify the request or HTTP status. - Root cause: Bare response
.json()failures discarded request provenance. The underlying intermittent transport/restart cause remains unconfirmed; focused local/Linux runs and the full Linux suite passed under instrumentation. - Resolution: Fixture JSON reads now report method, pathname, status and content type, without response bodies, cookies or request payloads. No delay, retry or weakened preservation assertion was added.
- Regression signal:
pnpm test:orchestratorretains the real process-stop/restart and once-only upgrade PUT assertions; subsequent CI failures will identify the exact HTTP boundary. - Prevention rule: Preserve safe request context when parsing fixture responses so a local transport failure is distinguishable from a production reconciliation failure before choosing a fix.
2026-09-06 — Consume accepted responses before native restart tests #
- Affected area: The upgrade process-restart case in
tests/orchestrator.test.mjs. - Symptom signature: With Linux x64, Node 24.14.0 and
CI=true, the firstPOST /__test__/provider/inspectafter an accepted upgrade returned HTTP 500 with a non-JSON response. This happened beforeworker.stop(), after the preceding upgrade cases passed. - Root cause: The fixture checked only the upgrade response's 202 headers and left its JSON body unread before polling. Native phase probes showed that the failed inspection never entered the fixture Worker, while the provider continued serving other requests. Consuming the accepted body removed the failure; reversing only that change restored it. The precise internal Wrangler connection failure mechanism was not established.
- Resolution: Consume and validate the accepted response's
installation.status === "updating"before inspection. Preserve the held upload, actual process restart, stable-state assertions and exactly-two-total-Worker-PUT assertion; add no retry or delay. - Regression signal:
CI=true pnpm test:orchestratorin Linux x64 Node 24.14.0: the original full suite failed, the response assertion passed all 17 tests, and reversing the assertion reproduced the same inspection failure. Run this CI-mode gate when changing the fixture HTTP lifecycle. - Prevention rule: Complete and validate HTTP response bodies before advancing a native lifecycle test. Do not mistake a headers-only acknowledgement for a completed fixture exchange, or attribute a pre-stop transport failure to restart readiness.
2026-09-06 — Hold observed native phases instead of racing UI polling #
- Affected area:
tests/installation-status.test.mjsand the native provider fixture intests/fixtures/orchestrator-worker.ts. - Symptom signature: Linux CI timed out waiting for the active “Verifying installation” progress step after observing “Provisioning your Sandbox”. The real installation UI polls every 2 seconds.
- Root cause: The fixture delayed each provider operation for only 1.8 seconds, so a valid native phase could start and finish between UI reads. Sequential assertions that every transient phase appears depended on polling alignment; a longer assertion timeout cannot recover a phase that already completed.
- Resolution: Test-only latches hold the Containers create request and health result separately, keyed by installation resource name and phase. The browser observes each active phase, confirms matching native installation metadata with no installed release yet, then explicitly releases its provider operation. Each latch has a 45-second failure deadline, and
finallyreleases both even when an assertion fails. The old timing delays are removed; production polling and Workflow transitions are unchanged. - Regression signal:
CI=true pnpm test:installation-statusexercises the actual built browser UI and native Workflow with deterministic provisioning/verifying observations, then requires Ready. The captured CI failure at the former verifying-step wait is preserved in the task handoff. - Prevention rule: When a browser test must observe an intermediate asynchronous phase, hold the corresponding test provider operation until that observation. Do not assume a fixed sleep exceeds every polling interval, scheduling delay or browser round trip.
2026-09-06 — Select one format from mixed Workers AI stream events #
- Affected area:
workers-ai-provider@4.0.0streaming adapter and shell tool calls. - Symptom signature: A single
echo Hello Worldrequest produced eight failed shell calls with interleaved JSON such as{"command": "{"command": "echoecho Hello Hello World"} World"}, and streamed prose repeated every token. - Root cause: Live Workers AI events contained both native top-level fields and their OpenAI-compatible equivalents. The provider processed both representations, emitting every text and tool-argument fragment twice.
- Resolution: The pinned provider patch gives native
responseandtool_callsprecedence within a mixed event, while retaining the OpenAI-compatible path when native fields are absent. - Regression signal:
pnpm test:providersfeeds mixed-format SSE events through the public Workers AI adapter and requires one text fragment and one valid tool input. A live local smoke call must execute one shell invocation with stdoutHello Worldand unduplicated final prose. - Prevention rule: Treat alternate wire representations within one provider event as mutually exclusive. Test the adapter with the exact mixed event shape returned by live inference.
2026-09-06 — Numeric zero is tool input, not a finalization sentinel #
- Affected area:
workers-ai-provider@4.0.0streaming tool-call assembly and quote-heavy shell scripts. - Symptom signature: A Fibonacci command arrived as
{}with an unterminated JSON error, or reached Bash truncated at.slice(. The remaining command was printed as pseudoshell(...)text. - Root cause: Workers AI emitted the
0in.slice(0, ...)as a numeric argument fragment. The provider used!argsto detect the end of a tool call, so numeric zero closed the call before the remaining JSON arrived. - Resolution: Only
null,undefined, and an empty string finalize an argument stream. The shell schema and instructions also direct multiline JavaScript through a single-quoted heredoc and require structured retries. - Regression signal:
pnpm test:providerssends a mixed-format stream with a numeric-zero fragment and requires the completeprintf 0input. The live Fibonacci request executes one heredoc command and returns all ten rows. - Prevention rule: Never use truthiness to classify streamed protocol values; valid argument fragments can be
0,false, or an empty-looking scalar.
2026-09-06 — Explicit URL reads need first-step tool selection #
- Affected area: Think turn assembly and the
read_url/browser_readtools. - Symptom signature: Flarebot promised to read a URL, then printed text such as
[read_url(url="https://…")]without creating any tool activity or page evidence. - Root cause: The model had all application tools under automatic selection, so an explicit URL request could be completed as prose instead of a structured call. A turn-wide forced choice was also incorrect because it repeated the reader on every agentic step.
- Resolution: Explicit URL-reading intent, including contextual questions such as “what's this telling me?”, now selects the appropriate reader through Think's native
beforeStephook on step zero only. Later steps can consume the result and answer. Web instructions reject pseudo-call syntax, and a short retry can recover the intended reader from the previous assistant response. - Regression signal:
pnpm test:thinkcovers direct reads, contextual link questions, rendered reads, pseudo-call retries, non-reading URL text, and continuation behavior. Live local requests for the reported Hacker News and Cloudflare documentation URLs each show exactly one successfulRead webpageactivity followed by sourced Markdown. - Prevention rule: When a request requires a tool, enforce it at the first model step. Do not use turn-wide tool choice for an agentic loop that must answer after receiving the result.
2026-09-06 — Unified catalog models retain their declared wire format #
- Affected area:
worker/model-provider.ts, Cloudflare unified AI catalog models - Symptom signature: Selecting
thinkingmachines/inkling-256ksucceeded in Settings, but sending an ordinary message returned a model request error or completed without response text. - Root cause: The model was added to the Workers AI allowlist and passed to the generic Workers AI adapter, whose stream mapper expects native/OpenAI-compatible events. Cloudflare exposes Inkling only through the Anthropic Messages request and response format, so its content events were not understood.
- Resolution: Inkling now uses the official Anthropic AI SDK converter and parser while transport remains the native Cloudflare
AI.runbinding with the default AI Gateway and session affinity. - Regression signal:
pnpm test:providersfeeds an Anthropic Messages event stream through the configured Inkling model and requires its text delta. - Prevention rule: Before allowlisting a unified catalog slug, route it by the request format declared in Cloudflare's model catalog and test that exact streaming event shape through the production model factory.
2026-09-06 — Cross-isolate expirations need validation margin #
- Affected area: Production owner-login challenge creation across the customer Worker and
PersonalAgentDurable Object. - Symptom signature:
/auth/loginreturned the generic installation-verification 403 even though all production variables, the secret, owner identity, bridge key, and Durable Object namespace were correct. - Root cause: The Worker issued
expiresAtat the store's exact ten-minute maximum. The Durable Object validated the timestamp against its own request clock, so small cross-isolate clock skew could reject the challenge as too far in the future. - Resolution: Login challenges use a nine-minute lifetime while the store retains the ten-minute absolute validation ceiling. The public response remains generic and no request or credential data is logged.
- Regression signal: The production login route creates a native
PersonalAgentchallenge and redirects to the configured control plane after a code update; the bridge integration gate retains expiry and replay validation. - Prevention rule: Do not issue distributed expiry values at the validator's exact upper boundary. Reserve explicit time for request transit and clock skew.
2026-09-06 — Third-party AI Gateway balance failures need an actionable boundary error #
- Affected area:
worker/model-provider.ts, Cloudflare unified AI catalog billing - Symptom signature: Retrying
thinkingmachines/inkling-256kin the local conversation returned “The response could not be completed” even after its Anthropic Messages adapter was installed. - Root cause: A direct call through the same remote
AIbinding returned HTTP 402, code 2021: the account's default AI Gateway had insufficient credits. The model boundary collapsed 402 into the generic Workers AI failure. - Resolution: HTTP 402 is sanitized to an actionable insufficient-balance error. Actual inference remains unavailable until the account adds AI Gateway credits (or configures a supported BYOK route).
- Regression signal:
pnpm test:providersinjects a 402 response through the configured Inkling model and requires the bounded AI Gateway balance error; the direct remote binding repro remains HTTP 402 until billing changes. - Prevention rule: Before debugging a third-party model's request or stream codec, probe its native Cloudflare binding status. Preserve actionable authentication, rate-limit, and payment categories while discarding provider response bodies.
2026-09-06 — Local AI configuration leaked into release artifacts #
- Affected area:
scripts/build-release.mjs, deployment fixtures, and the production artifact validator. - Symptom signature: CI rejected the public-fetch compatibility fixture with
artifact_unavailable; installation tests could not accept the generated release. - Root cause: Adding
ai.remote: truefor local inference copied a development-only option intodeployment.json. The installer correctly accepts only the AI binding name. Existing packaging tests checked that name but never loaded the complete release through the installer validator. - Resolution: Package only the AI binding name, retain the local remote-inference setting, and validate the actual packaged bytes through
loadArtifact. Compatibility fixtures must derive their deployment configuration from the packaged artifact, not local Wrangler settings. - Regression signal:
pnpm test:deploymentincludes a production artifact-validator test that failed on the original package and passes after rebuilding. The public-fetch fixture also rejects the original source-derived configuration and accepts the packaged one. - Prevention rule: Project development configuration into the explicit installation contract. Exercise the production artifact consumer against the exact packaged bytes; do not weaken its schema to accept local-only settings.
2026-09-06 — Optional providers need explicit provisioned enablement #
- Affected area:
control-plane/installation-workflow.ts,control-plane/deployment-api.ts, and OAuth capability coverage. - Symptom signature: OpenCode was selectable and its key could be saved, but the owner had never configured the required AI Gateway custom provider. The installation had nevertheless reached ready.
- Root cause: The runtime assumed an enabled account-level
opencode-goroute, while install/upgrade provisioned only Workers and Containers. Gateway setup existed solely as a manual documentation prerequisite. - Resolution: Install/upgrade automatically reconcile only
default. Explicit OpenCode enablement checks owner, grant and account membership, provisions the trusted route, and sends a signed receipt to persist runtime enablement. Key/model writes and inference cannot bypass that gate. Live rollout still requires publisher OAuth configuration and customer authorization; successful live inference with the user's key has not yet been verified. - Regression signal:
pnpm test:deployment-networkcovers idempotent setup and conflicting URLs.pnpm test:orchestratorrequires automatic default gateway setup without optional providers.pnpm test:bridgeexercises two-site enablement, reconnect, lost replies, failed acknowledgment and durable enablement;pnpm test:settingscovers disabled model/key controls. - Prevention rule: Separate always-on infrastructure from opt-in provider setup. A stored key is not proof that a provider is enabled. Validate credential destinations and require a verified provisioning acknowledgment before exposing inference.
2026-09-06 — Model failure exports discarded the diagnostic status #
- Affected area:
worker/model-provider.ts,worker/conversation.ts, and the diagnostic schema. - Symptom signature: A dev.4 export recorded an OpenCode Go
gpt-5.6-lunaattempt failing after 582 ms, but contained onlystatus: error; the conversation displayed a generic failed-response banner. - Root cause: The model boundary recorded terminal status without the provider HTTP status or failure stage, then sanitized the error. The original OpenCode failure remains unclassified; this finding concerns the lost diagnostic evidence, not its upstream cause.
- Resolution: Local diagnostics retain only integer HTTP error statuses (400–599, otherwise null) and an allowlisted request/stream failure stage. No messages, bodies, headers, credentials, or prompts are retained. Production deployment and a fresh failing attempt are still required to diagnose the original failure.
- Regression signal:
node --test --test-concurrency=1 tests/model-provider.test.mjs tests/diagnostics.test.mjs tests/diagnostics-native.test.mjscovers thrown and streamed failures, rejects arbitrary diagnostic fields, and verifies a native exported 401 with the request stage. - Prevention rule: Preserve bounded protocol evidence before sanitization. A generic error flag cannot distinguish authentication, billing, routing, or malformed-response failures; never infer one from latency alone.
2026-09-06 — Reachable local control plane is not configured provider setup #
- Affected area: Local control-plane configuration,
control-plane/config.ts,control-plane/installation-metadata.ts, and OpenCode Settings enablement. - Symptom signature: Local
/connectrendered successfully, but authorization showed "The Flarebot publisher needs to verify deployment permissions for this release" and the runtime's OpenCode key input was disabled. - Root cause: The local control plane was started with a placeholder OAuth client and no capability manifest. The local runtime also lacked a bridge key and provider receipt. The installation registry accepts deployed
workers.devorigins, not localhost; merely starting both servers cannot establish a full local installation workflow. - Resolution: Registered a separate development OAuth client, verified its real grant/UserInfo/account access, and restarted the control plane with its validated configuration. For this debugging session only, verified the existing enabled gateway route and issued a localhost-bound receipt using an independent development signing key. Production gates and registry records were not changed. Full localhost install/upgrade support remains separate work.
- Regression signal: Live local
POST /auth/startchanged from 503 to 303;/auth/provider-enabledaccepted the development receipt with 200. A Playwright check selected OpenCode Go, typed into the API key field, and asserted that Save key became enabled without saving a credential or sending inference. - Prevention rule: Verify authorization and provider enablement, not just page availability, before claiming a local setup is usable. Keep development keys separate, verify the destination before issuing a receipt, and never represent local bootstrap as a completed deployed installation.
2026-09-06 — Missing local Sandbox image is not a model failure #
- Affected area: Long-running local Wrangler runtime and Docker Sandbox execution.
- Symptom signature: OpenRouter streamed text and emitted a shell call successfully, but the tool failed after about 12 seconds. Wrangler logged
No such image available named cloudflare-dev/sandbox:9c19c745andContainer failed to start. - Root cause: The running dev runtime referenced a generated Docker image tag that no longer existed.
docker image inspectindependently confirmed its absence; the cause of removal was not established. - Resolution: Restarted local Wrangler to prepare a fresh Sandbox image. A new OpenRouter turn then executed the harmless printf command successfully, with the expected marker in the actual tool output. No provider or production configuration changed.
- Regression signal: Live diagnostics changed from shell
failedtosucceeded(one attempt, about 1.3 seconds). Both the model's tool-selection and follow-up streaming calls completed. The temporary API key was removed and the previous model restored afterward. - Prevention rule: Verify actual tool status and output rather than an assistant's mention of expected output. For local container startup errors, check the exact referenced Docker image before changing model routing, tool schemas, or timeouts.
2026-09-06 — Publisher release bump preserved an obsolete permission attestation #
- Affected area: Publisher deployment configuration,
control-plane/config.ts, and release verification. - Symptom signature: After deploying dev.6,
/connectreturned HTTP 200 but/api/connectionreturnedoauth_capability_unavailable: "The Flarebot publisher needs to verify deployment permissions for this release." - Root cause:
keep_varspreservedFLAREBOT_CONTROL_PLANE.oauthCapabilities.artifactVersionas dev.5. The configuration gate correctly requires it to match the deployed package version. Static-page and bundle checks missed that gate; OAuth fixtures generated matching versions automatically. - Resolution: Reviewed the unchanged required capability contract and updated only the production attestation's artifact version to dev.6. No client registration, OAuth scopes, secrets, or personal Worker changes were needed.
- Regression signal: The actual configuration function rejects the saved dev.5 attestation and accepts the same configuration with dev.6.
pnpm test:oauthcovers wrong-release rejection; release smoke checks must exercise live/api/connectionand/auth/start, not just/connect. - Prevention rule: Review and update the publisher permission attestation on every release bump. Preserve unrelated bindings, but do not mistake a preserved version-bound attestation for a valid new-release configuration. Require unauthenticated connection HTTP 401
not_connectedand authorization-start HTTP 303 before declaring onboarding ready.
2026-09-06 — Local AI smoke tests need the remote-binding dev API #
- Affected area:
tests/web-search-live.test.mjs, Wrangler local AI bindings. - Symptom signature: The live search smoke returned
search_unavailableimmediately underunstable_dev({ local: true }), despiteai.remote: truein its config. - Root cause: The older dev API did not establish the remote binding connection required by the native AI wrapper. Local fixture credentials also deliberately cannot access Cloudflare.
- Resolution: Run the opt-in paid smoke separately from
tests/fixtures/config.mjs, usingunstable_startWorkerwith a remote AI binding and withoutdev.remote: false. - Regression signal:
FLAREBOT_WEB_LIVE_SMOKE=1 node --test tests/web-search-live.test.mjschanged from an immediate failure to three real provider sources in about 12 seconds, without changing the search request. - Prevention rule: Verify that a live inference harness establishes remote bindings before diagnosing model access or request shape. Keep paid smoke authentication separate from credentials-free fixtures.
2026-09-07 — Reasoning signatures are protocol state, not diagnostics #
- Affected area:
worker/model-provider.ts, Anthropic-compatible reasoning and tool follow-ups. - Symptom signature: Enabling adaptive reasoning while stripping all stream
providerMetadataremoves the signature carried by an empty reasoning delta. The next SDK-generated tool-follow-up request loses the signed thinking block; redacted thinking is likewise lost. - Root cause: The privacy boundary treated every provider metadata field as optional diagnostics, but the Anthropic adapter uses
anthropic.signatureandanthropic.redactedDatato reconstruct required conversation blocks. - Resolution: Preserve only those string fields on reasoning chunks and generated reasoning parts. Continue removing other provider metadata, raw events, request/response envelopes, and unsafe errors.
- Regression signal:
pnpm test:providersexercises a native Anthropic SDK two-step tool round-trip with empty signed thinking and redacted thinking, and separately checks that unrelated metadata is stripped. - Prevention rule: When enabling reasoning, verify the actual follow-up request—not just the initial effort payload. Preserve required opaque protocol state and empty signature-bearing deltas without treating them as user-visible diagnostics.
2026-09-07 — Queued submission metadata is not active turn metadata #
- Affected area:
worker/conversation.ts, Think 0.17 scheduled model selection. - Symptom signature: A scheduled task in a chat with High effort used High even after the installation default became Low; the submission inspection correctly contained Low.
- Root cause:
submitMessages(..., { metadata })stores queue-ledger metadata but does not stamp the user message's reservedturnMetadata.activeTurnMetadatatherefore did not identify the scheduled turn, andbeforeTurnfell back to the last interactive request body. - Resolution: Also carry the scheduled model snapshot in application-owned metadata on the submitted user message and resolve that snapshot ahead of the interactive body. Keep queue metadata for submission inspection and task lifecycle handling.
- Regression signal:
FLAREBOT_EXECUTION_CASE='scheduled turns ignore' node --test --test-reporter=tap tests/execution.test.mjsfailed with High versus Low, then passed with the scheduled tool follow-up using Low, the chat override remaining High, and a subsequent interactive message using High. - Prevention rule: Test each native submission path's actual hook-visible data. Do not assume queue inspection metadata is automatically available through per-turn APIs.
2026-09-07 — Catalog removal broke persisted model settings and conversation startup #
- Affected area:
shared/model-providers.ts,PersonalAgent.getModelSettings, andConversationSession.connect. - Symptom signature: Settings cannot load model settings; conversation startup repeatedly reports connection loss after successfully opening its socket.
- Root cause: dev.6 offered OpenRouter
openai/gpt-5-mini. Removing it from the dev.7 catalog made the strict parser reject existing saved settings. Conversation startup awaits the same settings read and classifies its RPC failure as a reconnectable connection failure. Seeding the dev.6 selection in native SQLite reproduced the loop; the reported live diagnostics did not expose the stored selection, so the live cause remains unconfirmed. - Resolution: Retain the previously supported model and its provider-default effort behavior. Do not rewrite the owner's model, credentials, or conversations.
- Regression signal:
FLAREBOT_CHAT_CASE='previous release model' node --test --test-reporter=tap tests/chat-ui.test.mjsseeds the old selection directly in native storage and checks conversation readiness, unchanged settings, and the real Settings UI. - Prevention rule: A selectable-model catalog is also a persisted-data contract. Preserve previously valid selections or provide an explicit migration before removing an entry; exercise reads through upgraded storage, not only fresh installations.
2026-09-07 — Completed search response contained an unfinished search item #
- Affected area:
worker/web-search.ts, native AI Gateway Responses validation. - Symptom signature:
web_searchreturnsinvalid_search_responsefor ordinary monitor queries despite HTTP 200. - Root cause: One authorized replay of
1440p 120Hz OLED monitor USB-C Macreturned overallstatus: completed, a completed search with 22 source entries, a second search still markedsearchingwith three source entries, and a completed message with four citations. The request specifiedmax_tool_calls: 1, but two search items were returned. Flarebot rejects the entire response when any search item is not completed. The provider's reason for the inconsistent item state is unknown. - Resolution: Select only completed search items for source-list extraction while retaining completed-message citations and overall completion validation. Require at least one completed search and reject an empty result when another search remains unfinished. Do not relabel unfinished calls as completed or retry paid requests automatically.
- Regression signal: A single live replay and an offline Miniflare replay reproduced the rejection.
pnpm test:webexercises the mixed response through native tool execution and history, requires the completed source and summary, excludes the unfinished source, and rejects unfinished-only and mixed-empty responses. The mixed-response assertion failed before the fix. - Prevention rule: Validate response-level and item-level completion separately. Retain bounded failure classifications so mixed completion states can be distinguished from missing sources or invalid JSON without exporting queries or response content.
2026-09-07 — Custom Domain IDs differ from zone IDs #
- Affected area:
control-plane/domain-api.ts,control-plane/domain-metadata.ts,control-plane/installation-registry.ts. - Symptom signature: Cloudflare attaches the hostname and serves the app, but domain setup reports failure and
/auth/loginreturns 403 with "Owner authentication or installation verification failed." - Root cause: The live Workers Domains API returned a 40-character hexadecimal domain ID. Both the API verifier and persisted domain schema assumed 32 characters, rejecting the attached mapping before runtime authorization and HTTPS verification. Replaying that identifier shape through the native provisioning fixture reproduces the rejection; the authenticated live workflow record was not available for inspection.
- Resolution: Use a dedicated domain ID schema accepting the existing 32-character and observed 40-character formats. Keep account/zone IDs and ownership checks strict. Deployment and an authenticated retry are needed to reconcile the existing attachment; do not delete or reattach it.
- Regression signal:
node --test tests/domain-provisioning.test.mjs tests/domain-workflow.test.mjsexercises 40-character IDs through API reconciliation, SQLite persistence, activation and removal. The reconciliation test failed before the fix. - Prevention rule: Do not reuse account/zone identifier formats for other provider resources. Exercise provider-observed identifiers through both response validation and durable storage, including recovery after an accepted write.
2026-09-07 — Settings anchor jumps raced responsive navigation measurement #
- Affected area:
src/components/settings-navigation.tsx. - Symptom signature: Keyboard navigation immediately after resizing to 1024px put the Memory heading under the sticky navigation. At 390px the jump could leave the active-section indicator on the preceding section.
- Root cause: The native anchor jump used a scroll margin last measured for the previous layout. Deferred resize measurement and responsive layout changes then changed the offset without realigning the target.
- Resolution: Measure on link activation and realign the target on the next animation frame with the current navigation height, retaining native hash navigation and focus.
- Regression signal:
pnpm test:settingsfailed its heading-visibility assertion before the fix and passes responsive keyboard navigation, manual scrolling and active-section checks afterward. - Prevention rule: When native scrolling depends on dynamically measured sticky content, synchronize the measurement with navigation and test activation across breakpoint changes, not only settled layouts.
2026-09-08 — Resource navigation tests left later chat cases on Settings #
- Affected area:
tests/chat-ui.test.mjs, PR #45 browser CI. - Symptom signature: The memory banner test passes, then the reload/offline and pending-stream tests time out finding the
Messagetextbox. - Root cause: The new memory-link case ends on Settings. The next two cases share its browser page and send messages without opening a conversation. Running the resource case alone missed this dependency.
- Resolution: Explicitly open the fixture conversation at the start of each affected chat case.
- Regression signal: The full
CI=true pnpm test:chat-uirun reproduces the missing-composer timeout before the fix; run the complete sequence to verify navigation isolation. - Prevention rule: Each browser case must establish its starting route. After adding a navigation test, run the full containing suite as well as focused cases.
2026-09-07 — Default Wrangler target was not the live installation #
- Affected area: Manual deployment using
wrangler.jsonc. - Symptom signature: A successful
wrangler deployreportedflarebot.nbedd2.workers.dev, whose homepage returned HTTP 503 with invalid configuration, while the owner's existing app still worked. - Root cause: The default
flarebotWorker was mistaken for the installed customer Worker. Both its previous and newly deployed versions lacked installation configuration bindings. The custom domain actually mapped to a separateflarebot-<installationId>Worker. - Resolution: Confirmed
https://bot.nathanbeddoe.com/and the publisher's/connectreturned HTTP 200. Neither the installed customer Worker nor publisher was updated by this deployment; the intended release target still needs to be resolved. - Regression signal: Cloudflare domain mapping identifies the installed Worker;
curl https://flarebot.nbedd2.workers.dev/still returns the configuration error, whilecurl https://bot.nathanbeddoe.com/returns HTTP 200. - Prevention rule: Resolve the live domain, Worker identity, and release workflow before deploying. A default Wrangler target and successful upload do not establish that the user's installation was updated. Compare prior bindings before attributing a post-deploy error to changed configuration.
2026-09-08 — Publisher origin migration left customer bridge pins behind #
- Affected area: Publisher
publicOrigin, customerFLAREBOT_INSTALLATION.controlPlaneOrigin, andcontrol-plane/bridge.ts/worker/bridge.ts. - Symptom signature: After moving the publisher to
control.flarebot.app, an existing installation's login redirects to the old publisher hostname and receives HTTP 403bridge_denied. - Root cause: The publisher's exact-origin gate now requires the new hostname, while the deployed customer still pins the old origin for login redirects, code exchange, and assertion issuer verification. Keeping the old workers.dev endpoint enabled does not make it an accepted bridge origin.
- Resolution: Migrated the existing customer's pinned origin to
https://control.flarebot.appwith authorization, preserving its other installation fields, bindings and runtime metadata. No browser-only redirect or weakened issuer validation was added. - Regression signal: Live
curlrequests to/auth/loginon both customer hostnames originally returned 303 to the old publisher, followed by 403bridge_denied. After migration both return 303 to the new bridge and then 303 into OAuth with the new callback. Full authenticated callback verification remains outstanding. - Prevention rule: Treat a publisher origin change as a migration of every installed customer's trust configuration, not just DNS and OAuth. Verify customer login, server-side exchange, and issuer checks as well as publisher onboarding before declaring the migration complete.
2026-09-08 — Worker settings patch dropped container runtime metadata #
- Affected area: Cloudflare Workers script settings PATCH for a container-backed customer Worker.
- Symptom signature: A successful binding-only settings update retained bindings and exports but removed
resources.script_runtime.containersfrom the active version. - Root cause: The settings PATCH schema does not support
containers; supplying the previous container mapping there did not preserve it. Adding an export'scontainerreference was rejected because that endpoint did not declare the referenced container. - Resolution: Re-uploaded the existing deployed text bundle through script PUT with the original exports and container mapping, inherited bindings and retained assets. No local rebuild or container application change was performed.
- Regression signal: Compared original and final active-version resources after normalizing binding order and the intended installation-origin change; all other version resources matched, including the Sandbox container mapping. Both live login redirect checks passed afterward.
- Prevention rule: For container-backed Workers, verify active-version container metadata after configuration changes. Use a metadata-preserving script upload when settings PATCH cannot represent the original runtime configuration; do not equate HTTP 200 with preservation.