llm #
Provider-neutral values for one turn with a language model, an
OpenAI-compatible chat-completions client for Eio, and a fireworks command
that chats with the open-weight models Fireworks AI serves and rewrites named
sections of a Markdown document.
llm holds checked, inert values: transcripts, requests with their options
and tools, streamed events, terminal responses, token usage and classified
errors, each with a JSON codec. It performs no I/O. A client is an injected
interpreter, so a test or another transport supplies its own.
llm-eio interprets requests against any server that speaks the
OpenAI Chat Completions
protocol, over requests-eio. It streams the reply as server-sent events,
retries a failed request within a bound (honouring the server's
Retry-After), re-runs a broken stream only while no output has been shown,
and classifies failures by HTTP status and the error body. Llm_eio.Fireworks
points it at Fireworks AI
and lists the models an account serves.
Security #
Sending text to Fireworks publishes it to a third party. The fireworks
command refuses a file (a prompt, a system prompt, rewriting instructions or a
document) that lives in a git repository whose remote belongs to a private
owner, including a local clone of one, unless --allow-private is given.
The API key is sent only as an Authorization: Bearer header and never
printed: -v names where the key came from, not the key. A key file that its
group or others may read is refused, with the chmod that fixes it.
Install #
opam install llm-eio
Usage #
Library #
A request is built and checked without a network:
# let model = Llm_eio.Fireworks.model "accounts/fireworks/models/glm-4p6" ;;
val model : Llm.Model.t = <abstr>
# let transcript = Llm.Transcript.of_list_exn [ Llm.Message.user_text "Hello" ] ;;
val transcript : Llm.Transcript.t = <abstr>
# let request = Llm.Request.v_exn ~model transcript ;;
val request : Llm.Request.t = <abstr>
# Llm.Model.id (Llm.Request.model request) ;;
- : string = "accounts/fireworks/models/glm-4p6"
Sending it with the Eio client:
let reply request =
Eio_main.run @@ fun env ->
Eio.Switch.run @@ fun sw ->
let clock = Eio.Stdenv.clock env in
let session = Requests_eio.v ~sw ~clock (Eio.Stdenv.net env) in
let credential =
Llm_eio.Fireworks.Credential.api_key (Sys.getenv "FIREWORKS_API_KEY")
in
let client = Llm_eio.Fireworks.client ~session ~clock ~credential () in
match Llm.Client.response client request with
| Ok r -> print_endline (Llm.Response.text r)
| Error e -> prerr_endline (Llm.Error.message e)
Llm.Client.response ~on_event sees the reply's text as it streams, the
token usage, and each retry before it is made.
Command line #
The key is read from --key-file PATH, then from FIREWORKS_API_KEY, then
from $XDG_CONFIG_HOME/fireworks/token (by default
~/.config/fireworks/token), which must have mode 0600. A model is named by
its resource name or, for the serverless catalogue, by its last segment.
$ fireworks chat -m kimi-k2-instruct "Name three moons of Jupiter."
$ fireworks chat -m glm-4p6 --system-file style.md -f draft.md --json
$ git log -1 --format=%B | fireworks chat -s "Shorten this commit message."
$ fireworks models
$ fireworks rewrite -m qwen3-235b-a22b --instructions style.md \
--in README.md --section Overview --section Usage --out README.new.md
chat prints the reply as it streams, or the whole response with --json,
and the token usage on standard error. Fireworks publishes no per-model price
table, so no cost is estimated.
rewrite sends each named section, heading and body, with the instructions,
and nothing else of the document. It writes a copy in which only those
sections' bodies changed, and refuses the result when a reply alters a fenced
code block of its section byte for byte, adds a heading at the section's level
or higher, or leads to any change outside the named sections.
--base-url URL sends the requests to another OpenAI-compatible server.
Tests #
The codec tests decode the examples of the Fireworks API reference, copied
verbatim. The client is tested against a local fake of the API, the rewrite
against an injected client. Interop traces of the live API are not recorded:
run FIREWORKS_API_KEY=... REGEN=1 dune build @traces with a key to record
them.
License #
ISC. See LICENSE.md.