This repository has no description
flarebot docs bug-lessons.md
50 kB

Bug lessons #

Durable, evidence-backed lessons from debugging sessions in this repo. Symptom-match new bug reports against these entries before theorising.

2026-09-06 — Public Worker-to-Worker requests require explicit routing #

  • Affected area: Publisher readiness fetch and customer-to-publisher login bridge; both Wrangler compatibility flag lists.
  • Symptom signature: Worker and Sandbox provisioning completed, but verification failed immediately. The signed customer health endpoint passed from an external client, while the same request from a cloud Worker returned HTTP 404 with Cloudflare error 1042.
  • Root cause: Neither Worker enabled global_fetch_strictly_public. Native fetch could not use the other Worker's public workers.dev endpoint. This routing flag is not implied by the compatibility date or nodejs_compat.
  • Resolution: Enable the flag on both publisher and customer deployments. Publish the changed customer contract as 0.1.0-dev.2, retaining the original dev.1 archive and an explicit forward-upgrade edge rather than modifying its immutable bytes.
  • Regression signal: The same signed cloud-Worker probe returned error 1042 before the flag and full readiness HTTP 200 after it. Configuration/artifact tests require the new release's flag and retain support for the exact legacy flag list needed to recover and upgrade existing installations.
  • Prevention rule: Test public Worker-to-Worker calls from a deployed Worker in each direction. A successful external HTTP request or local service fixture does not certify Cloudflare's edge routing; maintain explicit routing requirements in both deployment contracts.

2026-09-06 — Native fetch rejected the deployment client's receiver #

  • Affected area: control-plane/deployment-api.ts, account preparation and uploaded Worker verification.
  • Symptom signature: Installation stayed at “Preparing your account”; the native Workflow exhausted four immediate temporarily_unavailable attempts in resolve customer origin, before deploying resources.
  • Root cause: The adapter stored ambient Workers fetch on the client and invoked this.network(...). Native fetch rejected the DeploymentAPI receiver with “Illegal invocation” before sending a request; the adapter sanitized that exception. Arrow-function provider fixtures did not enforce the native receiver constraint.
  • Resolution: Call the injected transport as a standalone function in both the JSON request and multipart content verification paths. Preserve the fixed endpoints, credentials, bounds and error classifications.
  • Regression signal: pnpm test:deployment-network exercises the actual adapter and native Workers fetch, replacing only outbound service responses. Before the fix, account preparation fails without reaching the fixture endpoint; the corrected transport resolves the origin and verifies uploaded content.
  • Prevention rule: Tests for an injected platform primitive must preserve its native calling convention. Exercise default transports in the target runtime as well as arrow-function substitutes; receiver-sensitive APIs cannot safely be invoked through arbitrary owning objects.

2026-09-06 — Cloudflare callback scope metadata rejected valid sign-ins #

  • Affected area: control-plane/http.ts, OAuth callback query validation.
  • Symptom signature: Every real Cloudflare sign-in returned /connect?error=oauth_invalid_callback and “This sign-in link is invalid or expired. Connect to Cloudflare again.” even with a fresh transaction.
  • Root cause: Cloudflare returned code, scope, and state; the callback allowlist rejected scope before claiming the transaction or exchanging the code. Both the direct and browser fixtures had omitted this provider field.
  • Resolution: Accept a single callback scope parameter as non-authoritative metadata. Continue validating granted scopes from the token exchange and retaining state, cookie, issuer, duplicate-query, expiry and one-use checks.
  • Regression signal: pnpm test:oauth reproduces the exact error before the fix and passes afterwards through the native Worker/DO callback and Chromium navigation. It also rejects duplicate scope parameters and a token response missing permissions despite a complete callback scope list.
  • Prevention rule: Model the real provider's callback shape in both HTTP and browser fixtures; distinguish provider metadata from verified token grants. Testing OAuth endpoints separately does not validate the application's complete callback path.

2026-09-06 — Local runtime had no owner login path #

  • Affected area: worker/bridge.ts, local customer-runtime setup
  • Symptom signature: Opening http://localhost:8787/auth/login returned Owner authentication or installation verification failed. and the shell remained disconnected when using .dev.vars.example.
  • Root cause: The development template intentionally omitted production bridge keys, while the only owner-session issuer required a configured control-plane bridge.
  • Resolution: An explicit development runtime on a loopback origin issues its local owner session from /auth/login; non-loopback and production installations retain the signed OAuth bridge.
  • Regression signal: pnpm test:runtime asserts that the actual packaged local Worker redirects /auth/login, sets the hardened owner cookie, and accepts it on the private status route.
  • Prevention rule: Every documented local startup path that exposes authenticated UI must include a bounded way to establish its development identity.

2026-09-05 — Sidebar SSR flash "window is not defined" #

  • Affected area: octane-kumo Sidebar.Provider → useIsMobile (packages/octane-kumo/src/components/sidebar/sidebar.tsx); surfaced in flarebot's RootLayout, which renders Sidebar.Provider.
  • Symptom signature: First paint shows the root error boundary ("Something went wrong! window is not defined"), replaced by the real shell after hydration. The SSR HTML of every route contains the error text.
  • Root cause: useIsMobile's useSyncExternalStore snapshot called unguarded window.matchMedia. Octane's server useSyncExternalStore falls back to getSnapshot() when the compiled call carries no slot arg — and the server hook-slot transform wraps the nested useIsMobile() call site with withSlot but does not inject slots into its body (verified in dist/server/entry.js: 3-arg call, no slot; the client bundle for the same source has the slot). So SSR threw ReferenceError: window is not defined.
  • Resolution: Guarded getSnapshot with typeof window === "undefined" ? false : window.matchMedia(query).matches, matching getServerSnapshot. Fixed in octane-kumo, consumed via the link: dependency.
  • Regression signal: packages/octane-kumo/tests/sidebar-ssr.test.tsx — SSR-renders Sidebar.Provider with window stubbed to undefined; fails pre-fix, passes post-fix. Manual loop: start pnpm dev, then curl -s http://localhost:5176/about | grep -c "window is not defined" must print 0.
  • Prevention rule: Every useSyncExternalStore getSnapshot — and any render-path browser-global access — in Octane-ported components must be SSR-safe (typeof window/document guards). Never rely on getServerSnapshot alone; the server transform may drop it.

2026-09-05 — Native facet client never becomes ready #

  • Affected area: AgentClient.basePath / PartySocket URL construction.
  • Symptom signature: Authenticated facet history works, but client.ready remains pending with no identity frames and WebSocket handshakes return 404.
  • Root cause: Passing /agents/... as basePath generates ws://host//agents/...; PartySocket prepends its own slash.
  • Resolution: Pass the authenticated pathname with its first slash removed to basePath; preserve the slash for HTTP requests.
  • Regression signal: pnpm test:think awaits bounded native facet readiness and completes streamed turns using basePath: pathFor(id).slice(1).
  • Prevention rule: Treat a PartySocket base path separately from an absolute HTTP pathname. Do not loosen the server's exact route guard for malformed URLs.

2026-09-06 — Tool activity recovery and clear use different native paths #

  • Affected area: worker/tool-activity.ts, Think tool hooks and native clear.
  • Symptom signature: An interrupted call recovers with the same tool-call ID, but a blanket terminal-state guard leaves its activity failed while the native transcript succeeds. A public clearMessages override alone also misses WebSocket clear because Think invokes its own clear handler.
  • Root cause: An interrupted observation has an unknown outcome, unlike a known tool failure or explicit cancellation. Think 0.17's two clear paths both call the documented protected resetTurnState, but WebSocket clear does not dispatch through public clearMessages.
  • Resolution: Permit only interrupted observations to reopen on authoritative execution/result evidence, preserve first timestamps and count observed attempts. Strip activity presentation metadata synchronously at the native reset seam, retaining ID-only tombstones against late callbacks.
  • Regression signal: pnpm test:activities restarts real workerd during a tool, reissues that exact native call ID and compares transcript/activity success; it also clears native chat while an abort-ignoring tool is running and checks that a later turn survives without restored old summaries.
  • Prevention rule: Distinguish known terminal outcomes from interrupted observations, and verify both native HTTP/RPC and WebSocket lifecycle paths before choosing an override. Tests must actually reissue the same call ID to exercise replay, rather than merely letting the recovered model finish text.

2026-09-06 — Think fetch timeout stopped after response headers #

  • Affected area: Think 0.17.0 dist/tools/fetch.js, executeRequest.
  • Symptom signature: A server returns headers and an initial body chunk, then stalls; native timeout and caller cancellation no longer interrupt the body.
  • Root cause: Returning finalizeResponse(...) without awaiting it exits the surrounding try/finally immediately, clearing the request timer and removing caller abort forwarding before readCapped finishes.
  • Resolution: Version-pinned pnpm patch adds await at that return. Native limits, redirect filtering and download code remain authoritative.
  • Regression signal: pnpm test:web feeds a controlled slow body to the actual native fetch tool and verifies timeout plus underlying signal abort, then stops a native Think read_url invocation during body consumption.
  • Prevention rule: When cleanup releases cancellation or resource ownership, await asynchronous body processing before leaving the protected scope. Retest slow bodies before removing the patch on a Think upgrade.

2026-09-06 — Browser acquisition outlived its conversation facet #

  • Affected area: native facet deletion and Browser Run session acquisition.
  • Symptom signature: Deleting a conversation while browser creation was awaiting its response left the actual remote session open. The child's late continuation logged Facet was deleted; capturing a parent RPC stub alone did not keep that continuation alive.
  • Root cause: deleteSubAgent destroys child execution and its pending continuations. A child-owned finally or waitUntil cannot guarantee cleanup after the facet itself is destroyed.
  • Resolution: The surviving parent owns the native create request and its waitUntil, records the returned ID before connection, and rechecks whether the conversation still exists. A late result for a deleted conversation is closed directly. The child still owns only its session's browser commands.
  • Regression signal: pnpm test:browser delays a real local Chromium create reply, deletes the conversation, and probes that exact session for HTTP 404.
  • Prevention rule: Own external acquisition in a lifetime that survives its caller's deletion. Cleanup intent must survive alongside that owner; remote creation whose ID is lost still requires an honest service-expiry fallback.

2026-09-06 — Stable Sandbox session buffering defeated output limits #

  • Affected area: Sandbox 0.12.9 getSandbox and shell streaming.
  • Symptom signature: Direct subclass execStream buffered an unterminated output line until the deadline; explicit exit ended the SDK's persistent session without an ordinary command completion event.
  • Root cause: getSandbox(..., { enableDefaultSession: false }) implements stateless execution in its public helper wrapper. Calling this.execStream inside a subclass bypasses that wrapper and uses a persistent shell session.
  • Resolution: The one-use Sandbox lifecycle wrapper delegates through the public stateless helper, verifies its own namespace identity, and preserves native SSE byte chunks and terminal exit codes.
  • Regression signal: pnpm test:shell checks nonzero exit and floods real Docker stdout without newlines; the byte cap must destroy the container before its time deadline. It also checks partial-output cancellation.
  • Prevention rule: Test exact streaming semantics with unterminated output, explicit exit and silence. Wrapper options are not necessarily stored runtime configuration; retain the SDK helper that implements them.

2026-09-06 — Native one-shot cleanup retry deduplication #

  • Affected area: Agent native scheduled resource cleanup callbacks.
  • Symptom signature: A failed cleanup schedules an idempotent retry with identical type, callback and payload; no later retry occurs.
  • Root cause: Native scheduling deduplicates against the currently executing one-shot row, then deletes that row after the callback returns.
  • Resolution: Cleanup retries create a new native one-shot schedule without self-deduplication; confirmed cleanup cancels remaining matching schedules.
  • Regression signal: pnpm test:shell injects four consecutive destroy failures and waits for actual native scheduled cleanup and stopped container.
  • Prevention rule: Test successive failures, not only the first retry, and account for scheduler row advancement when rearming inside a callback.

2026-09-06 — Local public container images still trigger Wrangler auth #

  • Affected area: Credentials-free native Worker and Docker tests.
  • Symptom signature: CI fails before starting a local Worker with an account or token error, although local tests pass with a developer login.
  • Root cause: Wrangler 4.128 normalizes container image references and fills registry API authentication even when local container execution is disabled.
  • Resolution: The shared test fixture supplies a synthetic account/token and redirects Cloudflare API requests to closed loopback. Optional registry login fails locally; Docker pulls the pinned public image directly. Production deployment configuration stays account-neutral.
  • Regression signal: Worker and real Docker shell tests pass with an isolated Wrangler config directory and no real Cloudflare credentials or API access.
  • Prevention rule: Verify native local tests without developer authentication; disabling remote execution does not necessarily disable config-time auth.

2026-09-06 — Chat UI artifacts and failed-case cleanup were host-dependent #

  • Affected area: tests/chat-ui.test.mjs screenshot capture and native client lifetime.
  • Symptom signature: Unprivileged Linux reports EACCES: permission denied, mkdir '/private' in the rich-message case; the later history case times out locating a desktop sidebar link. A failed clear case can leave the process alive after TAP has reported its assertion.
  • Root cause: Screenshots used a developer's macOS path. Its failure skipped viewport restoration, contaminating later cases. A native leaf client closed only on success kept reconnecting after failed assertions; held HTTP gates and route handlers also lacked failure cleanup.
  • Resolution: Store screenshots beneath the test-owned temporary directory (or explicit FLAREBOT_CHAT_SCREENSHOTS), restore viewport and release case resources in finally, and register bounded, idempotent suite cleanup. Use TAP output so the original assertion is visible immediately.
  • Regression signal: pnpm test:chat-ui in an unprivileged Linux container exercises the same screenshot and desktop-navigation sequence. Injecting an assertion after the clear case opens its client and holds HTTP made the old harness hit an external 25-second timeout; the corrected harness reports the intentional failure and exits nonzero normally in about three seconds. A separate pending-body injection exits after its 10-second parent timeout, confirming cleanup is independent of the suspended test body.
  • Prevention rule: Derive test artifact paths from tmpdir() or an explicit caller path. Register resource cleanup before awaiting readiness, and verify failed assertions release native reconnecting clients as well as browsers.

2026-09-06 — Completed resume observer overlaid authoritative history #

  • Affected area: ConversationSession native resume ownership and terminal history.
  • Symptom signature: After full Worker restart, native history contains a partial assistant and one separate continuation, but the connected UI also appends the continuation to the earlier partial. Native IDs match while text remains duplicated, even after waiting.
  • Root cause: An unsolicited fallback observer could become transport-owned during resume. Owned terminal frames bypassed the broadcast state transition, leaving its accumulator observing after the stream finished. The final fresh HTTP history was then overlaid with that obsolete accumulator.
  • Resolution: Retire the observer only when its stream ID matches the completed native request, following the SDK's own owned-response handling. Preserve native partial and continuation rows and let final history replace their text without an obsolete overlay.
  • Regression signal: pnpm test:chat-ui compares each rendered message's ID and text parts against native history at Connected after a real Worker restart, in addition to requiring the continuation text exactly once. The Linux failure reproduced with fresh history revision unchanged and an obsolete observer; the native mock had made only one continuation call.
  • Prevention rule: When streaming ownership changes, terminal cleanup must retire both ownership paths for that request. Compare authoritative message content as well as IDs, and never clear a different active stream's observer.

2026-09-06 — Unrouting raced an intercepted history response #

  • Affected area: Clear and stale-selection HTTP gates in tests/chat-ui.test.mjs.
  • Symptom signature: Linux CI fails Clear with route.fulfill: Route is already handled!; parent cancellation then produces closed-page cleanup errors. The same gate can pass on another runner.
  • Root cause: UI readiness did not mean every intercepted history handler had finished. A later non-aborted handler was still awaiting route.fetch when page.unroute disabled interception; its subsequent fulfillment raced Chromium's handling of that request.
  • Resolution: Both held-history cases use the public page.unrouteAll({ behavior: "wait" }) within the existing deadline. Active handlers finish before interception is disabled; route errors are not ignored on the success path and stale-history assertions remain intact.
  • Regression signal: The actual Clear case in an unprivileged Linux container reproduced the exact error with delayed fulfillment. Changing only route teardown to await handlers made that same narrowed case pass and exit normally. pnpm test:chat-ui retains both held-history acceptance cases.
  • Prevention rule: Await routing work itself before removing interception. A rendered ready state and release of a gate do not prove its asynchronous callback has finished.

2026-09-06 — Shell reconnect selector matched the conversation control #

  • Affected area: Session-expiry recovery in tests/app-shell.test.mjs.
  • Symptom signature: Playwright reports two matching Reconnect buttons after both shell and conversation connections learn that the session expired.
  • Root cause: The shell test used a page-wide action selector. Timing could expose either one or both independently owned reconnect controls.
  • Resolution: Scope shell actions to .connection-notice. The signed-session expiry case establishes both connections, awaits their native expiry, then verifies only shell Reconnect restores both. Cookie removal remains a separate new-request authorization check: it does not revoke an accepted native socket.
  • Regression signal: pnpm test:app-shell reproduced the original strict-mode failure. Preserving the chat socket while delivering offline/online events also reproduced the mistaken expectation that cookie removal must close chat. The real signed-expiry regression passes without relying on network disconnect timing.
  • Prevention rule: Scope repeated actions to their component, and trigger the actual authorization event before expecting an established socket to expire.

2026-09-06 — Held task page exceeded the native RPC deadline #

  • Affected area: Older-history repair case in tests/tasks-ui.test.mjs.
  • Symptom signature: Captured older page appended: 25, with the UI showing Could not load older runs. Try again. on slower CI runs.
  • Root cause: The test withheld a real older-page response until 29 native executions finished. That wait could exceed the browser client's 10-second RPC deadline, so the correctly captured four-row page arrived after its request failed.
  • Resolution: Pause the browser clock only during that deliberate response hold, leaving Worker execution in real time, and resume in finally. Assert the captured task, older cursor, four rows and active native status before release.
  • Regression signal: The real native case with a 12-second execution delay reproduced the exact failure and passed with controlled browser time. The test retains that delay, all 29-row append and completed-status reconciliation checks.
  • Prevention rule: Deliberate transport holds must control the client's deadline independently of slow server work when the assertion concerns stale data, not timeout.

2026-09-06 — Conversation socket outlived the browser's network loss #

  • Affected area: ConversationSession in src/runtime/conversation-session.ts; surfaced in the session-expiry recovery case of tests/app-shell.test.mjs.
  • Symptom signature: After setOffline(true), cleared cookies and setOffline(false), the shell reaches Sign-in required while the conversation heading still reads Connected and never shows Sign in to this installation to view this conversation.
  • Root cause: The shell drops its socket on the browser's offline event, but the conversation only reacted to a WebSocket close. Chromium's offline emulation (and a real loss on some networks) does not reliably close an established socket, so the conversation kept trusting a stale connection and never re-read history, which is where the 401 would have been observed.
  • Resolution: ConversationSession now listens to offline/online: offline detaches immediately and shows the reconnecting notice; online reconnects (queued if an aborted connection is still unwinding). Terminal states are left alone.
  • Regression signal: CI at head 54d813e kept the conversation Connected after a brief offline toggle and cookie removal. pnpm test:chat-ui covers offline stream recovery; pnpm test:app-shell separately verifies automatic recovery after both connections observe actual signed-session expiry.
  • Prevention rule: Every independently owned connection must observe the same browser network signals; a live socket is not evidence that the session is valid.

2026-09-06 — Native pending task status rejected by a test assertion #

  • Affected area: Captured older-page assertion in tests/tasks-ui.test.mjs.
  • Symptom signature: The captured real older page includes active runs fails after the correct four-row older page is captured during execution.
  • Root cause: The assertion used queued, but the task domain calls an acknowledged, unfinished submission pending. Faster native acknowledgement changed the captured rows from dispatching to valid pending.
  • Resolution: Check the actual active statuses: dispatching, pending, and running; retain the cursor, row-count and eventual completion checks.
  • Regression signal: CI rejected the captured older page; a local native capture confirmed real pending rows. The full task UI gate retains its active page assertion and all 29-row append/completion checks with the documented value.
  • Prevention rule: Read protocol and domain status values from their source; similar natural-language descriptions are not interchangeable enum values.

2026-09-06 — Referrer policy changed OAuth form Origin #

  • Affected area: Public OAuth onboarding forms and exact-Origin CSRF checks.
  • Symptom signature: Chromium submitted the real Connect form with Origin: null, so the Worker correctly rejected the request with HTTP 403.
  • Root cause: Applying Referrer-Policy: no-referrer to the public document also affected the Origin header on navigation form POSTs.
  • Resolution: Public onboarding uses same-origin, which retains same-origin form verification without disclosing referrers to Cloudflare. Callback and API responses retain no-referrer; exact-Origin checks remain unchanged.
  • Regression signal: pnpm test:oauth exercises actual Chromium form POSTs, redirect-chain provider interception, callback cookies, account selection and logout against the real Worker and native Durable Objects.
  • Prevention rule: Test browser navigation forms as well as direct HTTP API calls. Never weaken Origin validation to accommodate a referrer-policy mistake.

2026-09-06 — Native Workflow error messages are not a result protocol #

  • Affected area: control-plane/installation-workflow.ts
  • Symptom signature: Native health_failed and resource_conflict step failures became temporarily_unavailable in installation metadata, despite the correct callback error appearing in local logs.
  • Root cause: The Workflow/RPC boundary changes nonretryable error messages; matching an exact application enum against the wrapped message loses the original classification.
  • Resolution: Nonretryable callback failures return a strict safe result containing the enum. The Workflow interprets that persisted result outside the callback. Only transient failures throw a sanitized error for native retries.
  • Regression signal: pnpm test:orchestrator checks exact durable failure codes through real local Workflow execution and confirms a failed boot never assigns an installed release.
  • Prevention rule: Persist declared, nonsecret step results for domain failures. Do not depend on native exception message formatting as an application protocol.

2026-09-06 — Deployment fixtures must exercise the runtime configuration consumer #

  • Affected area: control-plane/deployment-config.ts, customer installation configuration and the orchestrator fixture.
  • Symptom signature: The deployment adapter included bridge.issuer, while the finished customer schema accepted only keyId and publicKey. Independent orchestrator and login fixtures passed, but the production customer loader rejected the actual uploaded variables.
  • Root cause: The producer retained an earlier interface proposal and the orchestration fixture verified multipart metadata without invoking its real consumer.
  • Resolution: Deployment configuration now passes through the production installation parser after origin resolution. The native provider fixture also feeds the exact uploaded variables to loadCustomerConfig and loadCustomerSecrets, using a fixed HTTPS publisher origin.
  • Regression signal: pnpm test:orchestrator validates the actual multipart bindings with the production customer loaders before accepting a Worker upload.
  • Prevention rule: Independently valid subsystem fixtures do not establish integration. Exercise the real consumer against the exact serialized producer output, including production origin and secret-binding requirements.

2026-09-06 — Upgrade preflight rejected native Worker metadata #

  • Affected area: control-plane/deployment-api.ts, upgrade observation and metadata preservation.
  • Symptom signature: A Ready installation rejects its first upgrade with resource_conflict before any upload, despite unchanged code, configuration and namespace identities.
  • Root cause: The provider fixture omitted default tags, version annotations, and script_runtime.assets/containers. The strict observation allowlist therefore classified ordinary Cloudflare upload results as resource drift.
  • Resolution: Preserve Worker tags and writable message/tag annotations, excluding only read-only workers/triggered_by provenance. Validate the exact asset-routing defaults and installation-owned container mapping produced by Flarebot's upload. Unknown metadata and changed routing/ownership still fail closed.
  • Regression signal: pnpm test:orchestrator models the native response fields, checks preserved customer metadata through upgrade/recovery, and rejects unsupported changes before customer mutation. The read-only live DeploymentAPI.baseline reproduction must also pass against the original installation.
  • Prevention rule: Build upgrade fixtures from the full native upload/read response shape. Separate writable metadata, server-generated provenance and deployment-owned settings explicitly; never fix an allowlist mismatch by dropping all unfamiliar fields.

2026-09-06 — Status-read deadline aborted valid upgrade commands #

  • Affected area: control-plane/ui/installation-status.ts, installation command requests.
  • Symptom signature: Continuing a saved upgrade displays “Flarebot could not be reached” while retaining the previous Ready release; ordinary status reads still succeed.
  • Root cause: Commands shared the 15-second polling deadline even though upgrade/recovery performs synchronous Cloudflare resource and code verification before acknowledgement. A native read-only preflight against the deployed artifacts took 15.2 seconds. A genuine accepted upgrade response delayed 16 seconds reproduced the exact browser alert; the original user's failed POST was not captured.
  • Resolution: Allow installation commands 120 seconds while retaining the 15-second read deadline, explicit cancellation, frozen request identity and replay behavior.
  • Regression signal: pnpm test:installation-status holds the real saved-upgrade 202 response for 16 seconds, rejects a false network alert, and verifies the exact request and final installed release. Existing deliberate network-loss and recovery cases remain covered.
  • Prevention rule: Budget acknowledgement time for the synchronous work an endpoint performs. A status-read timeout is not automatically suitable for a command that verifies remote resources before starting durable background work.

2026-09-06 — Live missing-Workflow errors blocked startup recovery #

  • Affected area: control-plane/start-installation.ts, recovery after the registry saves an operation but before its Workflow is created.
  • Symptom signature: Installation stays Updating / Preparing with no matching native Workflow; continuing the saved recovery request returns HTTP 503 before any Worker upload.
  • Root cause: Local native Workflow.get throws instance.not_found, while the deployed binding throws (instance.not_found) Instance not found. Recovery recognized only the local message and its RPC prefix, so the live absence signal became temporarily_unavailable.
  • Resolution: Recognize the observed deployed absence message alongside the existing local forms. All other lookup/termination failures remain unavailable and cannot authorize replacement execution.
  • Regression signal: pnpm test:orchestrator exercises the real recovery endpoint through a missing native instance with the deployed error serialization and verifies that unknown lookup errors leave the saved operation and provider resources unchanged. A protected read-only live probe confirmed both the missing execution and exact remote error shape.
  • Prevention rule: Native service error serialization can differ between local and deployed runtimes. Capture the actual remote contract at a failing boundary and cover its known representation explicitly; never interpret every lookup failure as absence.

2026-09-06 — Script PUT rejected version-pinned inheritance #

  • Affected area: control-plane/deployment-api.ts, upgraded Worker upload and source-version checks.
  • Symptom signature: Upgrade upload records its intent, fails with temporarily_unavailable, then stops with recovery_required while dev.1 still serves. The live PUT returns HTTP 400 / code 10057: inherited version_id accepts only the literal latest.
  • Root cause: The fixture accepted concrete version UUIDs from the shared API schema, but the deployed script PUT endpoint rejects them for every inherited binding.
  • Resolution: Use strict inheritance with version_id: "latest". Require the newest uploaded version to equal the checked active source during preflight and again immediately before PUT, including undeployed uploads. Preserve post-upload fingerprints; concurrent external uploads remain a documented non-atomic boundary.
  • Regression signal: pnpm test:orchestrator rejects UUID inheritance, checks latest-version reads around staging, and blocks both preexisting and late newer uploads before PUT. The guarded real dev.2 upload changed from HTTP 400 to success with the same saved configuration and resources.
  • Prevention rule: Exercise the actual provider endpoint when its behavior differs from a shared schema. Do not replace a pinned source with latest without checking all uploaded versions, and do not describe the resulting check/write sequence as atomic.

Upgrade operation defaults belong only at storage boundaries #

The native upgrade gate exposed a baseline-erasure bug: using a defaulted Zod storage schema with .partial() for intent mutations inserted upgrade: null and rolloutId: null into otherwise unrelated Worker-intent changes. The upload and health completed, but request replay and failed-upgrade retry lost their immutable source baseline. Use an explicit strict mutation schema with no storage defaults, excluding the immutable upgrade baseline entirely. Native replay/retry tests and an unknown-field intent rejection protect this boundary.

A ready installation is not evidence that an unsubmitted upgrade succeeded. The real browser test aborted the upgrade request before the registry saw it, then reloaded against the still-ready old installation. Reusing the initial-install ready cleanup erased the frozen upgrade target. Only clear an upgrade intent from a ready read when the installed identity matches that exact pending target; otherwise retain it for explicit replay.

2026-09-06 — Preserve request provenance in native fixture failures #

  • Affected area: JSON response reads in tests/orchestrator.test.mjs.
  • Symptom signature: Linux CI's upgrade process-restart case failed with Unexpected token 'E', "Error: Net"... is not valid JSON; the Undici-only stack did not identify the request or HTTP status.
  • Root cause: Bare response .json() failures discarded request provenance. The underlying intermittent transport/restart cause remains unconfirmed; focused local/Linux runs and the full Linux suite passed under instrumentation.
  • Resolution: Fixture JSON reads now report method, pathname, status and content type, without response bodies, cookies or request payloads. No delay, retry or weakened preservation assertion was added.
  • Regression signal: pnpm test:orchestrator retains the real process-stop/restart and once-only upgrade PUT assertions; subsequent CI failures will identify the exact HTTP boundary.
  • Prevention rule: Preserve safe request context when parsing fixture responses so a local transport failure is distinguishable from a production reconciliation failure before choosing a fix.

2026-09-06 — Consume accepted responses before native restart tests #

  • Affected area: The upgrade process-restart case in tests/orchestrator.test.mjs.
  • Symptom signature: With Linux x64, Node 24.14.0 and CI=true, the first POST /__test__/provider/inspect after an accepted upgrade returned HTTP 500 with a non-JSON response. This happened before worker.stop(), after the preceding upgrade cases passed.
  • Root cause: The fixture checked only the upgrade response's 202 headers and left its JSON body unread before polling. Native phase probes showed that the failed inspection never entered the fixture Worker, while the provider continued serving other requests. Consuming the accepted body removed the failure; reversing only that change restored it. The precise internal Wrangler connection failure mechanism was not established.
  • Resolution: Consume and validate the accepted response's installation.status === "updating" before inspection. Preserve the held upload, actual process restart, stable-state assertions and exactly-two-total-Worker-PUT assertion; add no retry or delay.
  • Regression signal: CI=true pnpm test:orchestrator in Linux x64 Node 24.14.0: the original full suite failed, the response assertion passed all 17 tests, and reversing the assertion reproduced the same inspection failure. Run this CI-mode gate when changing the fixture HTTP lifecycle.
  • Prevention rule: Complete and validate HTTP response bodies before advancing a native lifecycle test. Do not mistake a headers-only acknowledgement for a completed fixture exchange, or attribute a pre-stop transport failure to restart readiness.

2026-09-06 — Hold observed native phases instead of racing UI polling #

  • Affected area: tests/installation-status.test.mjs and the native provider fixture in tests/fixtures/orchestrator-worker.ts.
  • Symptom signature: Linux CI timed out waiting for the active “Verifying installation” progress step after observing “Provisioning your Sandbox”. The real installation UI polls every 2 seconds.
  • Root cause: The fixture delayed each provider operation for only 1.8 seconds, so a valid native phase could start and finish between UI reads. Sequential assertions that every transient phase appears depended on polling alignment; a longer assertion timeout cannot recover a phase that already completed.
  • Resolution: Test-only latches hold the Containers create request and health result separately, keyed by installation resource name and phase. The browser observes each active phase, confirms matching native installation metadata with no installed release yet, then explicitly releases its provider operation. Each latch has a 45-second failure deadline, and finally releases both even when an assertion fails. The old timing delays are removed; production polling and Workflow transitions are unchanged.
  • Regression signal: CI=true pnpm test:installation-status exercises the actual built browser UI and native Workflow with deterministic provisioning/verifying observations, then requires Ready. The captured CI failure at the former verifying-step wait is preserved in the task handoff.
  • Prevention rule: When a browser test must observe an intermediate asynchronous phase, hold the corresponding test provider operation until that observation. Do not assume a fixed sleep exceeds every polling interval, scheduling delay or browser round trip.

2026-09-06 — Select one format from mixed Workers AI stream events #

  • Affected area: workers-ai-provider@4.0.0 streaming adapter and shell tool calls.
  • Symptom signature: A single echo Hello World request produced eight failed shell calls with interleaved JSON such as {"command": "{"command": "echoecho Hello Hello World"} World"}, and streamed prose repeated every token.
  • Root cause: Live Workers AI events contained both native top-level fields and their OpenAI-compatible equivalents. The provider processed both representations, emitting every text and tool-argument fragment twice.
  • Resolution: The pinned provider patch gives native response and tool_calls precedence within a mixed event, while retaining the OpenAI-compatible path when native fields are absent.
  • Regression signal: pnpm test:providers feeds mixed-format SSE events through the public Workers AI adapter and requires one text fragment and one valid tool input. A live local smoke call must execute one shell invocation with stdout Hello World and unduplicated final prose.
  • Prevention rule: Treat alternate wire representations within one provider event as mutually exclusive. Test the adapter with the exact mixed event shape returned by live inference.

2026-09-06 — Numeric zero is tool input, not a finalization sentinel #

  • Affected area: workers-ai-provider@4.0.0 streaming tool-call assembly and quote-heavy shell scripts.
  • Symptom signature: A Fibonacci command arrived as {} with an unterminated JSON error, or reached Bash truncated at .slice(. The remaining command was printed as pseudo shell(...) text.
  • Root cause: Workers AI emitted the 0 in .slice(0, ...) as a numeric argument fragment. The provider used !args to detect the end of a tool call, so numeric zero closed the call before the remaining JSON arrived.
  • Resolution: Only null, undefined, and an empty string finalize an argument stream. The shell schema and instructions also direct multiline JavaScript through a single-quoted heredoc and require structured retries.
  • Regression signal: pnpm test:providers sends a mixed-format stream with a numeric-zero fragment and requires the complete printf 0 input. The live Fibonacci request executes one heredoc command and returns all ten rows.
  • Prevention rule: Never use truthiness to classify streamed protocol values; valid argument fragments can be 0, false, or an empty-looking scalar.

2026-09-06 — Explicit URL reads need first-step tool selection #

  • Affected area: Think turn assembly and the read_url/browser_read tools.
  • Symptom signature: Flarebot promised to read a URL, then printed text such as [read_url(url="https://…")] without creating any tool activity or page evidence.
  • Root cause: The model had all application tools under automatic selection, so an explicit URL request could be completed as prose instead of a structured call. A turn-wide forced choice was also incorrect because it repeated the reader on every agentic step.
  • Resolution: Explicit URL-reading intent, including contextual questions such as “what's this telling me?”, now selects the appropriate reader through Think's native beforeStep hook on step zero only. Later steps can consume the result and answer. Web instructions reject pseudo-call syntax, and a short retry can recover the intended reader from the previous assistant response.
  • Regression signal: pnpm test:think covers direct reads, contextual link questions, rendered reads, pseudo-call retries, non-reading URL text, and continuation behavior. Live local requests for the reported Hacker News and Cloudflare documentation URLs each show exactly one successful Read webpage activity followed by sourced Markdown.
  • Prevention rule: When a request requires a tool, enforce it at the first model step. Do not use turn-wide tool choice for an agentic loop that must answer after receiving the result.

2026-09-06 — Unified catalog models retain their declared wire format #

  • Affected area: worker/model-provider.ts, Cloudflare unified AI catalog models
  • Symptom signature: Selecting thinkingmachines/inkling-256k succeeded in Settings, but sending an ordinary message returned a model request error or completed without response text.
  • Root cause: The model was added to the Workers AI allowlist and passed to the generic Workers AI adapter, whose stream mapper expects native/OpenAI-compatible events. Cloudflare exposes Inkling only through the Anthropic Messages request and response format, so its content events were not understood.
  • Resolution: Inkling now uses the official Anthropic AI SDK converter and parser while transport remains the native Cloudflare AI.run binding with the default AI Gateway and session affinity.
  • Regression signal: pnpm test:providers feeds an Anthropic Messages event stream through the configured Inkling model and requires its text delta.
  • Prevention rule: Before allowlisting a unified catalog slug, route it by the request format declared in Cloudflare's model catalog and test that exact streaming event shape through the production model factory.

2026-09-06 — Cross-isolate expirations need validation margin #

  • Affected area: Production owner-login challenge creation across the customer Worker and PersonalAgent Durable Object.
  • Symptom signature: /auth/login returned the generic installation-verification 403 even though all production variables, the secret, owner identity, bridge key, and Durable Object namespace were correct.
  • Root cause: The Worker issued expiresAt at the store's exact ten-minute maximum. The Durable Object validated the timestamp against its own request clock, so small cross-isolate clock skew could reject the challenge as too far in the future.
  • Resolution: Login challenges use a nine-minute lifetime while the store retains the ten-minute absolute validation ceiling. The public response remains generic and no request or credential data is logged.
  • Regression signal: The production login route creates a native PersonalAgent challenge and redirects to the configured control plane after a code update; the bridge integration gate retains expiry and replay validation.
  • Prevention rule: Do not issue distributed expiry values at the validator's exact upper boundary. Reserve explicit time for request transit and clock skew.

2026-09-06 — Third-party AI Gateway balance failures need an actionable boundary error #

  • Affected area: worker/model-provider.ts, Cloudflare unified AI catalog billing
  • Symptom signature: Retrying thinkingmachines/inkling-256k in the local conversation returned “The response could not be completed” even after its Anthropic Messages adapter was installed.
  • Root cause: A direct call through the same remote AI binding returned HTTP 402, code 2021: the account's default AI Gateway had insufficient credits. The model boundary collapsed 402 into the generic Workers AI failure.
  • Resolution: HTTP 402 is sanitized to an actionable insufficient-balance error. Actual inference remains unavailable until the account adds AI Gateway credits (or configures a supported BYOK route).
  • Regression signal: pnpm test:providers injects a 402 response through the configured Inkling model and requires the bounded AI Gateway balance error; the direct remote binding repro remains HTTP 402 until billing changes.
  • Prevention rule: Before debugging a third-party model's request or stream codec, probe its native Cloudflare binding status. Preserve actionable authentication, rate-limit, and payment categories while discarding provider response bodies.

2026-09-06 — Local AI configuration leaked into release artifacts #

  • Affected area: scripts/build-release.mjs, deployment fixtures, and the production artifact validator.
  • Symptom signature: CI rejected the public-fetch compatibility fixture with artifact_unavailable; installation tests could not accept the generated release.
  • Root cause: Adding ai.remote: true for local inference copied a development-only option into deployment.json. The installer correctly accepts only the AI binding name. Existing packaging tests checked that name but never loaded the complete release through the installer validator.
  • Resolution: Package only the AI binding name, retain the local remote-inference setting, and validate the actual packaged bytes through loadArtifact. Compatibility fixtures must derive their deployment configuration from the packaged artifact, not local Wrangler settings.
  • Regression signal: pnpm test:deployment includes a production artifact-validator test that failed on the original package and passes after rebuilding. The public-fetch fixture also rejects the original source-derived configuration and accepts the packaged one.
  • Prevention rule: Project development configuration into the explicit installation contract. Exercise the production artifact consumer against the exact packaged bytes; do not weaken its schema to accept local-only settings.