Provider-neutral LLM requests, transcripts and responses, I/O-free
README.md

llm #

Provider-neutral values for one turn with a language model, an OpenAI-compatible chat-completions client for Eio, and a fireworks command that chats with the open-weight models Fireworks AI serves and rewrites named sections of a Markdown document.

llm holds checked, inert values: transcripts, requests with their options and tools, streamed events, terminal responses, token usage and classified errors, each with a JSON codec. It performs no I/O. A client is an injected interpreter, so a test or another transport supplies its own.

llm-eio interprets requests against any server that speaks the OpenAI Chat Completions protocol, over requests-eio. It streams the reply as server-sent events, retries a failed request within a bound (honouring the server's Retry-After), re-runs a broken stream only while no output has been shown, and classifies failures by HTTP status and the error body. Llm_eio.Fireworks points it at Fireworks AI and lists the models an account serves.

Security #

Sending text to Fireworks publishes it to a third party. The fireworks command refuses a file (a prompt, a system prompt, rewriting instructions or a document) that lives in a git repository whose remote belongs to a private owner, including a local clone of one, unless --allow-private is given.

The API key is sent only as an Authorization: Bearer header and never printed: -v names where the key came from, not the key. A key file that its group or others may read is refused, with the chmod that fixes it.

Install #

opam install llm-eio

Usage #

Library #

A request is built and checked without a network:

# let model = Llm_eio.Fireworks.model "accounts/fireworks/models/glm-4p6" ;;
val model : Llm.Model.t = <abstr>
# let transcript = Llm.Transcript.of_list_exn [ Llm.Message.user_text "Hello" ] ;;
val transcript : Llm.Transcript.t = <abstr>
# let request = Llm.Request.v_exn ~model transcript ;;
val request : Llm.Request.t = <abstr>
# Llm.Model.id (Llm.Request.model request) ;;
- : string = "accounts/fireworks/models/glm-4p6"

Sending it with the Eio client:

let reply request =
  Eio_main.run @@ fun env ->
  Eio.Switch.run @@ fun sw ->
  let clock = Eio.Stdenv.clock env in
  let session = Requests_eio.v ~sw ~clock (Eio.Stdenv.net env) in
  let credential =
    Llm_eio.Fireworks.Credential.api_key (Sys.getenv "FIREWORKS_API_KEY")
  in
  let client = Llm_eio.Fireworks.client ~session ~clock ~credential () in
  match Llm.Client.response client request with
  | Ok r -> print_endline (Llm.Response.text r)
  | Error e -> prerr_endline (Llm.Error.message e)

Llm.Client.response ~on_event sees the reply's text as it streams, the token usage, and each retry before it is made.

Command line #

The key is read from --key-file PATH, then from FIREWORKS_API_KEY, then from $XDG_CONFIG_HOME/fireworks/token (by default ~/.config/fireworks/token), which must have mode 0600. A model is named by its resource name or, for the serverless catalogue, by its last segment.

$ fireworks chat -m kimi-k2-instruct "Name three moons of Jupiter."
$ fireworks chat -m glm-4p6 --system-file style.md -f draft.md --json
$ git log -1 --format=%B | fireworks chat -s "Shorten this commit message."
$ fireworks models
$ fireworks rewrite -m qwen3-235b-a22b --instructions style.md \
    --in README.md --section Overview --section Usage --out README.new.md

chat prints the reply as it streams, or the whole response with --json, and the token usage on standard error. Fireworks publishes no per-model price table, so no cost is estimated.

rewrite sends each named section, heading and body, with the instructions, and nothing else of the document. It writes a copy in which only those sections' bodies changed, and refuses the result when a reply alters a fenced code block of its section byte for byte, adds a heading at the section's level or higher, or leads to any change outside the named sections.

--base-url URL sends the requests to another OpenAI-compatible server.

Tests #

The codec tests decode the examples of the Fireworks API reference, copied verbatim. The client is tested against a local fake of the API, the rewrite against an injected client. Interop traces of the live API are not recorded: run FIREWORKS_API_KEY=... REGEN=1 dune build @traces with a key to record them.

License #

ISC. See LICENSE.md.