Core Concepts
Providers, models, dialects, and usage — the mental model behind the router
Tokamak is built around one idea: every LLM call in your organization should go through a single door. This page explains what sits behind that door.
Providers and models
A provider is an upstream LLM endpoint plus the credential used to reach it — Anthropic, OpenAI, an OpenAI-compatible gateway, or your own inference server. Providers are configured once by an administrator and their credentials never leave the server.
A provider model is one model offered by one provider. Provider models are grouped into a model catalog, and the subset your deployment chooses to publish is the set of exposed models.
GET /v1/models returns the exposed models your account can reach. The model field of any completion request is resolved against that catalog, which is how a client can switch upstreams without changing anything but a string.
Administrators manage this surface through /v1/admin/providers and /v1/admin/models, gated by the admin.models.manage permission.
One model ID is virtual: tokamak/auto. A request that names it is served by a real catalog model, chosen per request from a routing policy of tiers, and from then on it is a request for that model. See Auto routing.
API dialects
Tokamak accepts three request shapes and forwards each one to whatever provider serves the requested model.
OpenAI Chat Completions — POST /v1/chat/completions. The broadest-compatibility path: any OpenAI SDK works unchanged. Supports streaming via Server-Sent Events.
Anthropic Messages — POST /v1/messages. For the Anthropic SDK and clients built against it. Anthropic's own server tools (web_search_20250305, web_fetch_20250910, and the rest) are forwarded to the provider exactly as declared — Tokamak does not intercept or execute them.
OpenAI Responses — POST /v1/responses, plus GET/DELETE /v1/responses/{response_id} and the /cancel and /input_items sub-resources (the legacy /items spelling still works as an alias). The provider that served the original request owns the state the dialect implies — including previous_response_id chaining — so a follow-up turn, or a retrieve/delete/cancel by id, is forwarded to that same provider. Tokamak keeps only a routing index from response id to provider, not the response itself.
Tools are executed by the client
On the Responses API, Tokamak returns function_call items to your client, exactly as the provider sent them. Your client runs the tool and sends the result back as a function_call_output item on the next request, referencing previous_response_id. The router does not run a server-side tool loop — it never executes your functions for you.
Request fidelity
Every request on the three dialects above is forwarded to the selected provider byte-for-byte — no translation, no rewriting, no injected defaults — with exactly two edits: model is replaced by the provider's own model id, and a streaming Chat Completions body gets stream_options.include_usage set so usage can be metered. The provider's response — JSON or SSE — is relayed unchanged apart from SSE keep-alive comments on a silent stream (and, when a provider sends nothing for 30 seconds, a stream Tokamak starts itself; see Streaming), and a usage scanner reads token counts out of it as it streams past, so metering works without parsing or rewriting the payload. The settled generation id comes back in the X-Tokamak-Execution-Id response header rather than a body field.
Use Tokamak's router for auth, key management, routing and accounting; the wire behaviour on the way to the provider is the provider's own.
Authentication and scopes
Credential class and permissions determine which endpoints each credential can use:
- API keys — sent as
X-API-Key: <key>orAuthorization: Bearer sk_live_.... For scripts, CI, agents, and applications. - OIDC / JWT sessions — for interactive use in the web app.
Authorization resolves over three scopes: platform, organization and team. Roles are held at one of those scopes; each role carries a permission set that the gateway checks on every request. Platform standing does not reach into an organization — that wall is enforced, not conventional — while an organization's admins do reach its teams, through a short enumerated list rather than a general rule. Usage limits independently match platform, organization, team and user policies. Organization-created team/member budgets stay anchored to that organization, and calendar-month windows use UTC. Team attribution comes from the credential, not from changing an analytics selector. See Access Control.
Organization and payer
Browser context identifies the active workspace. Existing API keys retain their billing organization and optional default team. Personal and shared organizations use the same billing machinery. Read Organizations and teams.
Usage and cost
The usage pipeline records routed activity, including: input tokens, output tokens, the model, and the principal that made the call. Tokens are captured from the provider response as it streams past. Telemetry that coding tools push over OTLP is separate, nonbilling activity for Insights and the Tools view; it never adds tokens or cost.
Pricing is stored per provider model, so Tokamak computes a cost for each record. Those records roll up into analytics and feed usage limits. Reported usage cost, remaining budget and available credit answer different questions. Wallet charges depend on billing activation and settlement; see Usage and spending.
tokamak usage
tokamak usage --period weekSee Usage.
Next steps
- Quick Start — make your first routed request
- Coding Agents — point Claude Code, Codex, opencode, or Pi at the router
- Architecture — how gateway, core, auth and billing fit together
- Self-hosting — run Tokamak yourself