Image and PDF attachments #
The composer accepts up to four PNG, JPEG, WebP or PDF files per message, including messages containing only attachments. Images are limited to 3.5 MiB and PDFs to 20 MiB. Each conversation currently allows 100 uploaded files. Audio and video are deferred.
Uploads use authenticated, same-origin HTTP requests independently of the chat WebSocket. The server detects the format from the bytes and allocates an opaque ID. Think's conversation workspace stores originals, with native R2 spillover; messages retain attachment metadata, not base64 payloads. Download and preview routes verify the owner session and active conversation. Removing an unsent attachment deletes its workspace directory. Saved-message attachments are retained until the conversation is deleted, which also removes their R2 objects. Interrupted or abandoned uploads can remain until conversation deletion.
read_attachment is available for files referenced in the conversation. It
delegates image and PDF output to Think's native workspace reader. Anthropic
models use native image/PDF input. Llama 4 Scout and Qwen 3.8 use images and
extracted PDF text. OpenRouter's GPT-5.6 Luna and GPT-5 Mini routes use images
and native PDFs; GLM-5.3 Flash and MiMo V2.5 use images and extracted PDF text.
Other catalog models use extracted PDF text. MODEL_INPUTS records capabilities
per model, with exhaustive coverage of the selectable catalog; provider identity
alone does not establish vision or native PDF support.
New attachments require a structured read on the first model step. Explicit follow-up references such as “The attached image?” also require a read. Later steps can answer normally after receiving the result. Uploading a file already authorizes reading it; the assistant must not ask again or print pseudo calls.
Chat-completions providers do not deliver native multimodal tool results. A provider middleware keeps the textual tool result and supplies media through the provider's supported user-message path. Request-body tests cover both streaming and generation. This uses the existing SDK rather than a replacement transport.
PDFs above Think's 3.5 MiB inline limit, or sent to models without native PDF
support, are converted with AI.toMarkdown() when read. Extracted Markdown is
cached in the workspace and read in bounded line ranges. Extraction failures
retain the original and can be retried. Extraction does not establish scanned
page or chart contents; the agent receives that limitation explicitly. Scanned
large PDFs should be split into smaller PDFs for a native visual model.
The installer provisions flarebot-attachments-<installationId>, adds the
ATTACHMENTS binding on upgrades, and verifies its identity before marking a
deployment ready. New artifacts require a reviewed attachment-storage OAuth
capability and a matching grant; existing control-plane configurations remain
parseable but must be updated before installing this artifact. No real scope
identifier is invented or automatically added to the registered OAuth client.
Validation: pnpm test:attachments, pnpm test:providers, and the attachment
cases in pnpm test:chat-ui exercise native SQLite/R2, format/size limits,
extraction caching, compact tool output, provider serialization, authorized
downloads, conversation isolation, reload and cleanup. Provider requests and
Markdown conversion use fixtures. The optional live check
FLAREBOT_ATTACHMENTS_LIVE=1 node --test tests/attachments-live.test.mjs uses
the real remote AI binding with Scout and Qwen: it requires a structured read
and checks both a shape/color and text visible only in the uploaded image.
References: Think workspace tools, Workers AI Markdown conversion.