Bring Your Own Client
Use any OpenAI- or Anthropic-compatible client or SDK with Tokamak
Tokamak exposes OpenAI Chat Completions and Responses routes and Anthropic-compatible /v1/messages. Point a compatible client at your Tokamak endpoint and use a model exposed by that deployment. API dialect and model/provider capabilities determine which client features work.
This is the flagship way to use Tokamak: keep your code, swap the base URL and the key.
What you get
- Unified model access — reach every model your deployment exposes through one endpoint
- No per-provider API keys — one Tokamak key in front of every upstream credential
- Usage tracking — routed inference is metered by the server; usage coverage and early refusals have the limits described in Usage and spending
Setup
Set two environment variables in your application:
OPENAI_BASE_URL=https://api.tokamak.sh/v1
OPENAI_API_KEY=sk_live_your_tokamak_api_keyThen use your client as normal:
from openai import OpenAI
client = OpenAI(
base_url="https://api.tokamak.sh/v1",
api_key="sk_live_your_tokamak_api_key"
)
response = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "Hello"}]
)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.tokamak.sh/v1",
apiKey: "sk_live_your_tokamak_api_key",
});
const response = await client.chat.completions.create({
model: "claude-sonnet-4-6",
messages: [{ role: "user", content: "Hello" }],
});Anthropic SDK
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.tokamak.sh",
api_key="sk_live_your_tokamak_api_key"
)
message = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}]
)Tokamak accepts the key as Authorization: Bearer sk_live_... or as an X-API-Key header, so both SDK conventions work.
Get an API key
tokamak authOr create a key from the dashboard, or via the API keys endpoint.
List available models
To see which models are available on your Tokamak instance:
curl https://api.tokamak.sh/v1/models \
-H "Authorization: Bearer sk_live_..."Pass the model ID from this list as the model field in your completions.
Tool calling
Tokamak forwards tool/function definitions to the provider and returns the model's tool calls to your client. Your code executes the tool and sends the result back on the next request. The router does not run a server-side tool loop.
On the Responses API this means Tokamak returns function_call items, and you reply with function_call_output items plus previous_response_id.
On the Anthropic Messages dialect, server tools (web_search_20250305, web_fetch_20250910, and the rest) are forwarded to the provider exactly as declared — Tokamak does not intercept or execute them.
Current limitations
When routing through Tokamak, the following OpenAI features behave differently:
- Assistants API — not supported
- Batch API — not supported
These are routing constraints, not limitations of the underlying models.
Launch coding clients
After tokamak auth, run tokamak launch <tool> --model <catalog-id>.
Supported tool names are claude, opencode, deepseek-harness,
codex-app, codex, copilot, droid, oh-my-pi, jan-agent, and
claude-cowork; legacy pi is also retained. The selected model must support
the client's API dialect and required tools, streaming, images or reasoning.
The launcher preserves native client permissions, extensions and sessions.
Hosted/account features retain the client's own requirements.
Plain tokamak launch claude uses Claude's default model selection through
Tokamak. Explicit --model selects catalog routing. --native uses local Claude
authentication and is outside Tokamak metering. Extra client options follow --.
Claude uses a private per-run settings overlay; OpenCode uses an inline overlay;
Codex CLI keeps the user's real home and instructions. Desktop adapters persist
configuration: tokamak launch codex-app --restore or
tokamak launch claude-cowork --restore restores prior files offline, refusing
conflicting edits. Fully quit/reopen the desktop app yourself to load changes.
Cowork applies a Claude-3p profile, which is what switches current Claude Desktop into its third-party deployment, and records the deployment mode in both regular Claude and Claude-3p configuration.
Run tokamak launch claude-cowork without --model: its native picker discovers
eligible Claude models from the configured gateway's authenticated /v1/models.
Tokamak removes the old fixed list and context preference. An empty eligible
catalog needs a Claude-capable server route; no fallback model is forced. Use
tokamak launch claude-cowork --check for offline diagnostics. Existing tasks may
retain their saved model; select a discovered model or start a new task.
Use tokamak launch <tool> --model <catalog-id> --check to inspect local setup
without network, installation, writes or execution. It does not certify model,
server or client compatibility. Droid installs from its official npm package when missing. Oh My Pi, Jan Agent and desktop apps must already be installed. See the complete client guides for setup and native-feature limits.
Routed inference uses Tokamak authentication, usage limits and metering. Optional client telemetry is separate nonbilling activity. The Tools analytics tab shows attribution coverage, including unknown clients; Copilot inference attribution can remain unknown even when its OTLP activity is labelled. See Usage and spending and the usage API. No client label changes access or pricing.
Client installation is separate from runtime compatibility. The Windows validation matrix records startup checks and API/account blockers. A temporary configuration must remain valid for the whole client runtime, including background processes; do not substitute another provider identity to make a configuration appear supported.