llm-bridge #
llm-bridge is a self-hosted inference gateway. It exposes OpenAI and Anthropic APIs, discovers the models available through configured providers, and routes each request to the matching backend.
It supports:
- Anthropic, Codex, and Google Antigravity subscription accounts through OAuth
- OpenAI-compatible API providers and providers from the
pi-airegistry - OpenAI Chat Completions and Responses APIs
- Anthropic Messages and token-counting APIs
- image generation and editing, audio transcription, and WebRTC Realtime call creation
- Exa-compatible search backed by Codex
- multiple provider accounts with per-key routing and conversation affinity
- usage accounting, quota-aware account selection, conversation compaction, and an admin UI
Requirements #
- Node.js 22.12 or newer
- npm 10 or newer
Node 22.12 is required for the built-in node:sqlite module.
Quick start #
Install dependencies:
npm ci
Create config.toml with at least one provider:
host = "127.0.0.1"
port = 4040
[providers.openrouter]
base_url = "https://openrouter.ai/api/v1"
api_key_file = "/run/secrets/openrouter-key"
Start the bridge:
npm start -- --config ./config.toml
The server listens at http://127.0.0.1:4040. Without a keys file, it accepts unauthenticated requests only when bound to a loopback address.
Check the server and list its discovered models:
curl http://127.0.0.1:4040/health
curl http://127.0.0.1:4040/v1/models
Send an OpenAI Chat Completions request:
curl http://127.0.0.1:4040/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"model": "MODEL_ID",
"messages": [{"role": "user", "content": "Hello"}]
}'
Use the bridge with an OpenAI client by setting its base URL to http://127.0.0.1:4040/v1. Set an Anthropic client's base URL to http://127.0.0.1:4040; it will call /v1/messages.
Providers #
Providers are configured under [providers.<name>]:
[providers.openai]
api_key_file = "/run/secrets/openai-key"
base_url = "https://api.openai.com/v1"
[providers.openrouter]
api_key_file = "/run/secrets/openrouter-key"
base_url = "https://openrouter.ai/api/v1"
Each provider accepts these fields:
| Field | Purpose |
|---|---|
api_key |
API key stored directly in the config |
api_key_file |
File containing the API key |
base_url |
Provider API base URL |
protocol |
Generic API protocol: openai, anthropic, or responses |
models |
Optional upstream model IDs used only when discovery yields no usable catalog |
auth_path |
OAuth credential file for a subscription provider |
model |
Default Codex model |
reasoning_effort |
Default Codex reasoning effort |
Use api_key_file instead of api_key for deployed services. Generic providers with base_url and a key discover their upstream model catalog automatically; models is an optional fallback, not a required list. For providers supported by pi-ai, the bridge registers models when the provider has credentials in the config or environment.
Generic API-key providers in the web UI #
Open Accounts → Add generic. Enter a unique provider ID, select OpenAI Chat Completions, Anthropic Messages, or OpenAI Responses, then supply the API base URL and API key. Leave the optional upstream model IDs blank to use discovery, or enter fallback IDs (one per line or comma-separated). An API key is still required; use a dummy value such as local for local servers that do not require authentication. Use an API root, not the completion endpoint: OpenAI-compatible URLs generally end in /v1; Anthropic accepts the root or a trailing /v1. Reverse-proxy path prefixes are preserved.
Discovery uses bounded requests in richest-first order: LiteLLM /model/info, LM Studio /api/v1/models, Ollama /api/tags enriched with /api/show, then the wire-base /models endpoint (Anthropic root URLs use /v1/models). It stops at the first usable chat catalog rather than merging catalogs. Available metadata populates model names, context windows, output limits, reasoning and vision capabilities, and costs. Explicit fallback IDs are used only when every probe is unavailable, empty, or unusable.
Models become available immediately as provider-id/upstream-model-id, with the same catalog metadata available through the API and admin UI. The bridge can expose them through all three client API formats using its existing translations. Generic providers are separate from OAuth accounts and their routing pools. Their catalogs refresh periodically, including providers added through the UI after startup.
The UI can list and remove saved generic providers. API keys are write-only in the admin API, and providers persist in generic-providers.json under the bridge state directory with mode 0600. Explicit config entries take precedence over saved entries and cannot be removed from the UI. To change a UI-managed provider, remove it and add it again.
The same protocol selection and optional fallback are available in config:
[providers.my_gateway]
protocol = "responses"
base_url = "https://gateway.example/v1"
api_key_file = "/run/secrets/gateway-key"
# Optional: used only if discovery finds no usable catalog.
# models = ["upstream-model-id"]
For a local unauthenticated server, set its API root as base_url and use api_key = "local" instead of api_key_file; omit models to discover its catalog.
The complete root configuration is:
host = "127.0.0.1"
port = 4040
codex_fast_max_output_cost_per_mtok = 20
trace_retention_days = 7
[providers.openrouter]
base_url = "https://openrouter.ai/api/v1"
api_key_file = "/run/secrets/openrouter-key"
The corresponding environment variables are:
| Variable | Purpose |
|---|---|
BRIDGE_HOST |
Listen address; default 127.0.0.1 |
BRIDGE_PORT |
Listen port; default 4040 |
BRIDGE_PROVIDERS |
JSON provider object when no providers are present in TOML |
BRIDGE_STATE_DIR |
Persistent state directory |
BRIDGE_TRACE_RETENTION_DAYS |
Stored trace lifetime in days; default 7; fractional days are supported |
CODEX_CLIENT_VERSION |
Pin the Codex wire version and skip automatic version lookup |
PI_AI_ANTIGRAVITY_VERSION |
Pin the Antigravity hub wire version and skip release-manifest lookup |
CODEX_FAST_MAX_OUTPUT_COST_PER_MTOK |
Maximum output price for advertised Codex :nitro aliases |
BRIDGE_MODEL_REFRESH_MS |
Subscription and generic model refresh interval in milliseconds; default 900000 (15 minutes), including generic providers added after startup |
TOML settings take precedence over environment variables.
API keys #
Pass --keys to require bridge authentication:
npm start -- --config ./config.toml --keys ./keys.json
keys.json may contain one object or an array:
[
{
"key": "11111111-2222-4333-8444-555555555555",
"name": "operator",
"allowedModels": ["*"],
"caps": ["admin"]
}
]
Newly generated keys use sk-bridget- followed by 37 securely random, unpadded base64url characters (A-Z, a-z, 0-9, -, _): 48 characters total with 222 random bits. These keys are case-sensitive.
Existing UUID keys remain supported, including the optional sk-bridget-uuid-<uuid> spelling. UUIDs remain case-insensitive; bare and prefixed spellings authenticate as the same key and retain the same permissions, routing, and usage history. No migration is required, and aliases cannot be registered as separate keys.
allowedModels contains case-insensitive globs with * and ?. The admin capability grants access to the admin API and UI. Model access and admin access are separate; allowedModels: ["*"] does not grant admin.
Clients can authenticate with either header:
Authorization: Bearer 11111111-2222-4333-8444-555555555555
x-api-key: 11111111-2222-4333-8444-555555555555
A non-loopback bind requires at least one API key. The bridge refuses to start as an unauthenticated public proxy.
The admin UI can create, edit, and delete keys when the file passed to --keys is a writable regular file. Leave the key field blank to generate a new 48-character key, or supply a key in any supported format. The creation response shows the secret once; later listings only show masked keys. Key writes are atomic and use mode 0600.
Repeatable HTTP verification: node --import tsx --test src/api-keys.e2e.test.ts. The scenario checks generation, UUID aliases, case-sensitive secrets, routing/usage identity, restart persistence, and revocation, and writes a secret-free report to /tmp/llm-bridge-api-key-artifacts/report.json.
Subscription accounts #
Open http://127.0.0.1:4040/admin. When API keys are enabled, sign in with a key that has the admin capability. Add one or more accounts for:
- Claude — Anthropic authorization-code flow
- Codex — OpenAI device-code flow
- Antigravity — Google authorization-code flow
The UI stores each account separately, refreshes its credentials, and updates the model catalog after login. It also manages per-key account pools, routing policies, compaction, API keys, model metadata, usage, upstream errors, and stored traces.
Claude OAuth follows Oh My Pi 18.6.2's authentication and request fingerprint: token exchange/refresh use api.anthropic.com, and Messages, Chat Completions, Responses, token counting, and model-driven compaction share the final-body billing checksum and tool-name translation. Existing credential files remain supported; new logins also retain account and organization identity. PI_AI_CLAUDE_CODE_VERSION pins the wire version; otherwise a version-too-old rejection can update it for one retry.
An Anthropic account exposes the bundled Claude catalog immediately. Live /v1/models discovery enriches it; a failed listing does not hide the bundled models. Previously discovered IDs are cached in anthropic-models.json under the state directory. Catalog presence is not proof of account entitlement or a valid unexpired grant: inference still reports upstream authorization errors.
Repeatable controlled-upstream verification: node --import tsx --test src/anthropic-wire.e2e.test.ts src/anthropic-compaction.e2e.test.ts. Captured requests and independently checked billing attestations are written under /tmp/llm-bridge-anthropic-artifacts/. These scenarios use synthetic credentials, not live Anthropic inference.
For Codex accounts, the remote compaction strategy uses Codex Responses compaction v2 when a configured context or idle trigger fires. It sends a compaction-only beta feature header, keeps the native compaction item and recent user turns for subsequent requests, and preserves explicit service_tier (priority or fast); :nitro models select priority when no tier is specified. Other providers do not support this remote strategy. Failed remote compaction returns an error instead of silently switching strategies. A repeatable real-server scenario writes its captured request and header artifact to /tmp/llm-bridge-codex-artifacts/remote-compaction.json via node --import tsx --test src/codex-remote-compaction.e2e.test.ts.
Stored traces #
Open Stored traces in /admin to browse requests, preview their complete trace, filter by API key, or download conversation as json. Both the page and downloads require admin access. HTTP and Responses WebSocket requests can retain their original body, response bytes, headers, provider attempts and events, routing, usage, errors, and compaction details. Authentication and cookie headers are redacted. Binary content is preserved; binary responses use base64 in JSON exports. Trace capture is disabled globally by default; enable Store traces in the admin settings to begin storing new requests.
Traces live in native tables in usage.sqlite. Scalar values have native types; objects and arrays use shared references. Large text and binary values are Brotli-compressed. Conversation prefixes and nested array prefixes are shared and extended, so later turns, retries, and forks do not copy earlier conversations. Garbage collection removes expired requests and unreferenced content while keeping prefixes still used by live traces. Existing token-usage accounting and routing state remain independent of trace retention.
Retention defaults to 7 days from receipt of each request. Set trace_retention_days in TOML, BRIDGE_TRACE_RETENTION_DAYS, or the retention form in Stored traces. The UI setting persists in SQLite, takes precedence on restart, and recalculates expiry for existing traces. Shortening retention can immediately collect older traces. Collection runs at startup, once per minute, and before listing or downloading traces. Durations must be greater than zero and at most 3650 days.
Trace capture requires both the global Store traces setting and the individual API key's store requests setting. The global setting defaults off and is managed through GET /admin/api/settings and PUT /admin/api/settings with { "storeTraces": true }. Each API key otherwise stores requests by default; turn off store requests on its API-key card, or set "storeRequests": false in its keys-file entry, to further restrict capture for that key. Open access (when no API key is required) is also subject to the global switch. Disabling global capture or a key's setting prevents new traces but does not delete existing traces; they remain available until their retention expiry.
Repeatable verification: node --import tsx --test src/trace-retention*.test.* exercises the real HTTP server and WebUI. It writes a JSON export and rendered HTML to /tmp/llm-bridge-retention-artifacts/.
A Codex account adds its live Codex model catalog. Models with a live Fast tier receive a :nitro alias when their known doubled output price is within the configured cap. Known models, including GPT-6 Astra, Luna, and Sol, report standard prices and doubled :nitro prices (including long-context tiers) in /v1/models and the admin catalog. Plain gpt-6 and future variants not yet in the pinned pricing catalog also receive the alias when the live catalog reports Fast, but remain unpriced until pricing is available. A live Codex account can expose these media aliases:
codex-imageandgpt-image-2codex-transcribeand supported OpenAI transcription model IDs
At startup, the bridge checks npm for the latest stable Codex CLI version. If it is newer than the version stored in the bridge, the first real Codex inference request tries it. When that attempt is rejected and the same request succeeds with the stored version, the bridge uses the stored version and refreshes its model catalog. If both attempts fail, the newer version remains unjudged. A failed npm lookup uses the stored version. Set CODEX_CLIENT_VERSION to pin a version and skip the lookup.
An Antigravity account adds its live Gemini catalog and logical model-family aliases. Chat reasoning_effort, Messages output_config.effort, and Responses reasoning.effort select the family's upstream model and thinking configuration. Gemini Flash 3.6+ selects advertised low/medium/high siblings; tiered-only catalogs retain the advertised -tiered ID with the requested thinking level. Older families use their curated token budgets. Known wire profiles set output caps and model labels; unknown models do not inherit another model's label.
Inference uses the Antigravity hub agent envelope, shared conversation identifiers and advancing steps, and validated tool calling by default. The hub version is discovered from the release manifest with a five-second timeout and a 2.8.0 fallback. PI_AI_ANTIGRAVITY_VERSION pins it; PI_AI_ANTIGRAVITY_CL, PI_AI_ANTIGRAVITY_OS, and PI_AI_ANTIGRAVITY_ARCH override the captured client metadata. OAuth login and refresh are unchanged.
Repeatable wire verification: node --import tsx --test src/antigravity-wire.e2e.test.ts exercises Chat, Messages, and Responses against a controlled upstream and writes captured requests to /tmp/llm-bridge-antigravity-artifacts/wire-parity.json. This checks wire contracts without using real credentials; it is not a live Google inference test.
Existing credential files can be configured directly:
[providers.anthropic-oauth]
auth_path = "/path/to/pi-ai/auth.json"
[providers.codex]
auth_path = "/path/to/codex/auth.json"
If no Codex path is configured, the bridge also checks CODEX_AUTH_PATH and ~/.codex/auth.json.
Inference APIs #
The bridge exposes these client-facing routes:
| Method | Route | API |
|---|---|---|
GET |
/health |
Health and model count |
GET |
/v1/models |
Models available to the current key |
POST |
/v1/chat/completions |
OpenAI Chat Completions |
POST |
/v1/responses |
OpenAI Responses |
POST |
/v1/messages |
Anthropic Messages |
POST |
/v1/messages/count_tokens |
Anthropic token counting |
GET |
/v1/usage?model=MODEL_ID |
Normalized provider usage |
POST |
/v1/images/generations |
OpenAI image generation |
POST |
/v1/images/edits |
OpenAI image editing |
POST |
/v1/audio/transcriptions |
OpenAI audio transcription |
POST |
/v1/realtime/calls |
OpenAI API-key WebRTC call creation |
POST |
/v1/realtime/transcription_sessions |
OpenAI transcription session creation |
POST |
/search |
Exa-compatible search |
Chat Completions, Responses, and Messages requests can use any discovered model. The bridge translates between API formats when the selected backend uses a different protocol. Streaming and tool calls are supported across the main text APIs.
Chat Completions also accepts raw Responses configuration controls interleaved
with messages on GPT-6 Codex routes:
{ "type": "configuration_update", "reasoning": { "effort": "high" } }
Keep the request-level reasoning_effort at its original value, append the
control before the next user message, and replay prior controls in their original
positions. The bridge forwards controls without treating them as text messages.
Malformed or adjacent controls, GPT-5.6, and non-Codex routes are rejected.
Automatic threshold/transition compaction rejects histories containing these
controls; idle compaction skips them. Start a fresh context to compact/rewrite
such a history. /v1/models exposes supported_reasoning_levels and
default_reasoning_level when the provider advertises them.
Exa-compatible search #
POST /search uses a connected Codex account and accepts the common Exa Search request shape:
curl http://127.0.0.1:4040/search \
-H 'content-type: application/json' \
-H 'x-api-key: YOUR_BRIDGE_KEY' \
-d '{
"query": "latest work on coding agents",
"type": "auto",
"numResults": 5,
"includeDomains": ["github.com"],
"contents": {"text": true},
"context": true
}'
The optional model field selects an allowed Codex model. Without it, the bridge selects an available Codex reasoning model.
Account routing #
Each API key can use a separate pool of Anthropic, Codex, and Antigravity accounts. Four selection policies are available:
| Policy | Behavior |
|---|---|
fill-first |
Use the first available account until its quota is exhausted |
lowest-usage |
Select the account with the lowest known quota usage |
round-robin |
Rotate new conversations across available accounts |
specific-plan-first |
Prefer a selected account, then fall back to the pool |
Conversations stay on their selected account. Send a stable identifier on every turn for explicit affinity:
x-llm-bridge-conversation-id: conversation-123
The Responses API conversation field and OpenAI prompt_cache_key are also recognized. Without an explicit identifier, the bridge derives one from the opening request content.
When a conversation moves to another account, the bridge carries portable context across the account boundary. Per-key compaction supports three strategies:
remote— provider-native compactionmodel-driven— a model-generated summary plus recent turnssnapcompact— local SnapCompact frames plus recent turns
Compaction can also run at a configured context percentage or after an idle interval. Conversation state is retained for the configured number of days.
State and backups #
The state directory is selected in this order:
- The first non-empty component of
STATE_DIRECTORY BRIDGE_STATE_DIR$HOME/.local/share/llm-bridge
It contains:
<state_dir>/accounts.json
<state_dir>/accounts/<account-uuid>.json
<state_dir>/usage.sqlite
Account files contain OAuth credentials. usage.sqlite contains token usage, routing state, conversation snapshots, and recent upstream errors. The directory uses mode 0700; credential and database files use mode 0600.
Back up the complete state directory with the service stopped. Include usage.sqlite-wal and usage.sqlite-shm if present. The state directory contains API credentials and conversation data.
NixOS #
The flake exports nixosModules.default and nixosModules.llm-bridge:
{
inputs.llm-bridge.url = "git+https://tangled.org/okami.mom/llm-bridge";
outputs = { self, nixpkgs, llm-bridge, ... }: {
nixosConfigurations.host = nixpkgs.lib.nixosSystem {
modules = [
llm-bridge.nixosModules.default
{
services.llm-bridge = {
enable = true;
keysFile = "/run/secrets/llm-bridge-keys.json";
settings = {
host = "0.0.0.0";
port = 4040;
providers.openrouter = {
base_url = "https://openrouter.ai/api/v1";
api_key_file = "/run/secrets/openrouter-key";
};
};
};
}
];
};
};
}
The module runs the service as llm-bridge and stores state in /var/lib/llm-bridge.
Development #
Run the backend and Vite admin UI together:
npm run dev -- --config ./config.toml --keys ./keys.json
The backend listens on port 4040 and the Vite UI is available at http://localhost:5173/admin/.
Run the admin UI with in-memory development data and no backend:
npm run mock
Run checks:
npm run typecheck
npm test
Repeatable controlled-upstream generic discovery verification: node --import tsx --test src/generic-model-discovery.e2e.test.ts. Catalog parity artifacts are written to /tmp/llm-bridge-generic-discovery-artifacts/catalog-parity.json; these scenarios do not require live provider credentials.
Format with Prettier:
npm run format
Check formatting without writing:
npm run format:check
Prettier formats JavaScript, TypeScript, React (.tsx/.jsx), CSS, HTML, JSON, and Markdown using the repository's existing no-semicolon, double-quote, 2-space style. Nix files are not covered. Existing files have not been bulk reformatted, so format:check may initially report formatting drift.
A pre-commit hook enforces formatting on staged files. Husky installs it through the prepare script when you run npm ci or npm install, and lint-staged runs prettier --write on the supported staged files. Staged formatting is rewritten in place and unstaged changes are preserved; formatting errors block the commit. The repository-wide check may still report preexisting drift.
Build the admin UI:
npm run build:admin