This repository has no description
flarebot docs attachments.md
3.9 kB

Image and PDF attachments #

The composer accepts up to four PNG, JPEG, WebP or PDF files per message, including messages containing only attachments. Images are limited to 3.5 MiB and PDFs to 20 MiB. Each conversation currently allows 100 uploaded files. Audio and video are deferred.

Uploads use authenticated, same-origin HTTP requests independently of the chat WebSocket. The server detects the format from the bytes and allocates an opaque ID. Think's conversation workspace stores originals, with native R2 spillover; messages retain attachment metadata, not base64 payloads. Download and preview routes verify the owner session and active conversation. Removing an unsent attachment deletes its workspace directory. Saved-message attachments are retained until the conversation is deleted, which also removes their R2 objects. Interrupted or abandoned uploads can remain until conversation deletion.

read_attachment is available for files referenced in the conversation. It delegates image and PDF output to Think's native workspace reader. Anthropic models use native image/PDF input. Llama 4 Scout and Qwen 3.8 use images and extracted PDF text. OpenRouter's GPT-5.6 Luna and GPT-5 Mini routes use images and native PDFs; GLM-5.3 Flash and MiMo V2.5 use images and extracted PDF text. Other catalog models use extracted PDF text. MODEL_INPUTS records capabilities per model, with exhaustive coverage of the selectable catalog; provider identity alone does not establish vision or native PDF support.

New attachments require a structured read on the first model step. Explicit follow-up references such as “The attached image?” also require a read. Later steps can answer normally after receiving the result. Uploading a file already authorizes reading it; the assistant must not ask again or print pseudo calls.

Chat-completions providers do not deliver native multimodal tool results. A provider middleware keeps the textual tool result and supplies media through the provider's supported user-message path. Request-body tests cover both streaming and generation. This uses the existing SDK rather than a replacement transport.

PDFs above Think's 3.5 MiB inline limit, or sent to models without native PDF support, are converted with AI.toMarkdown() when read. Extracted Markdown is cached in the workspace and read in bounded line ranges. Extraction failures retain the original and can be retried. Extraction does not establish scanned page or chart contents; the agent receives that limitation explicitly. Scanned large PDFs should be split into smaller PDFs for a native visual model.

The installer provisions flarebot-attachments-<installationId>, adds the ATTACHMENTS binding on upgrades, and verifies its identity before marking a deployment ready. New artifacts require a reviewed attachment-storage OAuth capability and a matching grant; existing control-plane configurations remain parseable but must be updated before installing this artifact. No real scope identifier is invented or automatically added to the registered OAuth client.

Validation: pnpm test:attachments, pnpm test:providers, and the attachment cases in pnpm test:chat-ui exercise native SQLite/R2, format/size limits, extraction caching, compact tool output, provider serialization, authorized downloads, conversation isolation, reload and cleanup. Provider requests and Markdown conversion use fixtures. The optional live check FLAREBOT_ATTACHMENTS_LIVE=1 node --test tests/attachments-live.test.mjs uses the real remote AI binding with Scout and Qwen: it requires a structured read and checks both a shape/color and text visible only in the uploaded image.

References: Think workspace tools, Workers AI Markdown conversion.