This repository has no description
flarebot docs agent-diagnostics.md
6.8 kB

Agent diagnostics #

Settings → Diagnostics downloads a JSON support file through the authenticated owner connection. Choose runtime only, or runtime plus one existing conversation. The download stays on the owner's device; Flarebot does not upload it to support or the control plane. Conversation names appear in the selector but are excluded from the file. Unknown/deleting conversations cannot be exported or recreated by an export. Public HTML contains no diagnostic data.

Collection uses each native Agents instance's public observability receiver, Think onStepEnd, and the existing application tool/task status projection seams. Each parent/facet has an application-owned SQLite history. There is no additional scheduler, delivery queue, stream consumer, trace platform, or transcript reader. The SDK's generic observability receiver remains connected as an independent sink. Both receivers contain observer failures. Runtime Cloudflare observability remains disabled by default, and Think storeMessages/storeTools remain false. This support file's allowlist is not a general sanitizer for native logs/traces; changing deployment logging settings is a separate operator decision.

Contents and interpretation #

The versioned file contains application/SDK versions, export time, scope, limits, availability, and retained events with immutable per-store sequence numbers. Only known enums, bounded finite scalars, allowlisted model/tool names and SHA-256 correlation IDs are persisted/exported. IDs may originate with clients/providers, so their raw values never enter diagnostic rows. No prompt, response, reasoning, custom instructions, memory, tool summaries/progress text, commands, outputs, URLs, file paths, provider errors, headers, close reasons, arbitrary metadata, credentials, environment configuration or installation addresses are exported. Rows are validated again during export; malformed/unknown-schema rows are omitted and counted, not serialized. IDs are correlation hints, not authorization tokens.

  • Turns: admitted native start/finish events, request correlation, known trigger, status and measured execution duration. A start without a finish has an unknown outcome; it is never implicitly successful. The history is best effort, not a complete audit log, and may miss pre-admission rejection or isolate-death events.
  • Model steps: normalized AI SDK v7 token usage, model response/step/first-output timing when available. Step timing can include tool work. A completed step is one usage observation; attempt events never add a second token charge. Workers AI's adapter supplies zero defaults for absent usage, so zero-valued counts are conservatively unavailable (null). Missing/nonfinite fields remain null. Billed cost is unavailable (null); no prices or charges are estimated.
  • Model attempts: invocation and one terminal status/duration per underlying provider invocation, including thrown errors, stream failures and cancellation. Stream attempts finish when consumption finishes/fails/cancels, not at response headers. They use separate attempt IDs; no native request correlation is claimed. The existing stream demand, raw-chunk filtering and cancellation are preserved.
  • Tools: status/attempt transitions from the existing activity lifecycle, including native action result errors and interrupted execution. Progress chunks do not each create an event. No tool argument/result presentation fields enter the observer. A generic/blocked tool result does not prove an external side effect occurred.
  • Tasks: scheduled/manual occurrence status from authoritative application projection, including receipt and native inspection reconciliation. Dispatching and dispatch errors remain distinct from inference completion. claimedAt is native submission claim, which can precede queue admission. claimToCompletionMs includes possible queue wait; use native turn/model timing for actual execution. Export does not reconcile tasks or cause them to run. Selected task events are included only with their conversation's scope.
  • Connections: native transport connect/disconnect with hashed connection ID, numeric close code and duration when a retained start exists. A connect event describes transport connection, not an independent credential validation.

Retention and lifecycle #

Each runtime/conversation stores at most 500 events, 256 KiB of serialized event payloads, and seven days of history; the first applicable limit removes oldest rows. Each event is at most 2 KiB. Pruning runs on writes/exports without a timer. The metadata tables and SQLite bookkeeping have small additional storage overhead. The complete compact JSON export is capped at 600 KiB. It contains at most two stores and never fans out across all conversations. Rows use descending immutable sequence order. prunedRows is cumulative for the whole store, including scopes not selected; omittedRows counts invalid retained records in the selected scope. These are not claims about missing observations before collection or during failure.

Clearing messages clears that conversation's local diagnostic history and advances a generation gate so late model/turn/recovery callbacks cannot restore it. Existing native activity tombstones continue to block old tool callbacks. Parent task history remains governed by task lifecycle; deleting a task purges its diagnostic rows, and deleting a conversation purges associated parent events before native facet teardown. Retention never deletes SDK submission/action idempotency records, activity tombstones or task lifecycle rows. A facet deletion uses the existing native lifecycle. Failed diagnostic writes cannot fail real turns/tools/tasks; a failed diagnostic read reports the scope unavailable.

Validation #

pnpm test:diagnostics checks strict adversarial projections, zero-default usage, SQL age/row/byte limits, immutable ordering, persisted reopen, clear generations, native receivers and storage failure isolation. Its workerd test uses actual native authenticated connections, two model steps plus a tool, task alarms with all clients disconnected, persisted process restart, clear/cancellation, deletion, and nonempty content/credential/error sentinels in both JSON and diagnostic rows. Inference is replaced only by the test entry's model; no provider bypass exists in production. pnpm test:providers checks attempt terminal statuses and stream backpressure/cancellation. pnpm test:settings downloads the actual production Worker's JSON through Chromium and covers the existing settings expiry, reconnect, mobile and private-HTML boundaries. Execution/Think/activity regression suites continue to own their respective native scheduling/recovery semantics.