DocsAPI Reference
Integrations

Bring Your Own Client

Use any OpenAI- or Anthropic-compatible client or SDK with Tokamak


Tokamak exposes OpenAI Chat Completions and Responses routes and Anthropic-compatible /v1/messages. Point a compatible client at your Tokamak endpoint and use a model exposed by that deployment. API dialect and model/provider capabilities determine which client features work.

This is the flagship way to use Tokamak: keep your code, swap the base URL and the key.

What you get

  • Unified model access — reach every model your deployment exposes through one endpoint
  • No per-provider API keys — one Tokamak key in front of every upstream credential
  • Usage tracking — routed inference is metered by the server; usage coverage and early refusals have the limits described in Usage and spending

Setup

Set two environment variables in your application:

OPENAI_BASE_URL=https://api.tokamak.sh/v1
OPENAI_API_KEY=sk_live_your_tokamak_api_key

Then use your client as normal:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.tokamak.sh/v1",
    api_key="sk_live_your_tokamak_api_key"
)

response = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content": "Hello"}]
)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.tokamak.sh/v1",
  apiKey: "sk_live_your_tokamak_api_key",
});

const response = await client.chat.completions.create({
  model: "claude-sonnet-4-6",
  messages: [{ role: "user", content: "Hello" }],
});

Anthropic SDK

from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.tokamak.sh",
    api_key="sk_live_your_tokamak_api_key"
)

message = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}]
)

Tokamak accepts the key as Authorization: Bearer sk_live_... or as an X-API-Key header, so both SDK conventions work.

Get an API key

tokamak auth

Or create a key from the dashboard, or via the API keys endpoint.

List available models

To see which models are available on your Tokamak instance:

curl https://api.tokamak.sh/v1/models \
  -H "Authorization: Bearer sk_live_..."

Pass the model ID from this list as the model field in your completions.

Tool calling

Tokamak forwards tool/function definitions to the provider and returns the model's tool calls to your client. Your code executes the tool and sends the result back on the next request. The router does not run a server-side tool loop.

On the Responses API this means Tokamak returns function_call items, and you reply with function_call_output items plus previous_response_id.

On the Anthropic Messages dialect, server tools (web_search_20250305, web_fetch_20250910, and the rest) are forwarded to the provider exactly as declared — Tokamak does not intercept or execute them.

Current limitations

When routing through Tokamak, the following OpenAI features behave differently:

  • Assistants API — not supported
  • Batch API — not supported

These are routing constraints, not limitations of the underlying models.

Launch coding clients

After tokamak auth, run tokamak launch <tool> --model <catalog-id>. Supported tool names are claude, opencode, deepseek-harness, codex-app, codex, copilot, droid, oh-my-pi, jan-agent, and claude-cowork; legacy pi is also retained. The selected model must support the client's API dialect and required tools, streaming, images or reasoning. The launcher preserves native client permissions, extensions and sessions. Hosted/account features retain the client's own requirements.

Plain tokamak launch claude uses Claude's default model selection through Tokamak. Explicit --model selects catalog routing. --native uses local Claude authentication and is outside Tokamak metering. Extra client options follow --. Claude uses a private per-run settings overlay; OpenCode uses an inline overlay; Codex CLI keeps the user's real home and instructions. Desktop adapters persist configuration: tokamak launch codex-app --restore or tokamak launch claude-cowork --restore restores prior files offline, refusing conflicting edits. Fully quit/reopen the desktop app yourself to load changes. Cowork applies a Claude-3p profile, which is what switches current Claude Desktop into its third-party deployment, and records the deployment mode in both regular Claude and Claude-3p configuration. Run tokamak launch claude-cowork without --model: its native picker discovers eligible Claude models from the configured gateway's authenticated /v1/models. Tokamak removes the old fixed list and context preference. An empty eligible catalog needs a Claude-capable server route; no fallback model is forced. Use tokamak launch claude-cowork --check for offline diagnostics. Existing tasks may retain their saved model; select a discovered model or start a new task.

Use tokamak launch <tool> --model <catalog-id> --check to inspect local setup without network, installation, writes or execution. It does not certify model, server or client compatibility. Droid installs from its official npm package when missing. Oh My Pi, Jan Agent and desktop apps must already be installed. See the complete client guides for setup and native-feature limits.

Routed inference uses Tokamak authentication, usage limits and metering. Optional client telemetry is separate nonbilling activity. The Tools analytics tab shows attribution coverage, including unknown clients; Copilot inference attribution can remain unknown even when its OTLP activity is labelled. See Usage and spending and the usage API. No client label changes access or pricing.

Client installation is separate from runtime compatibility. The Windows validation matrix records startup checks and API/account blockers. A temporary configuration must remain valid for the whole client runtime, including background processes; do not substitute another provider identity to make a configuration appear supported.

On this page