Auto routing (tokamak/auto)
Send model tokamak/auto and let Tokamak pick the cheapest model that answers each request well. Request, response, decision record and routing report.
tokamak/auto is a virtual model. When your organization has auto routing enabled, a request that names it is served by a real model from your organization's routing policy, chosen per request:
- Tokamak reads the latest message you wrote, plus up to two earlier ones. Tool results, assistant turns, system prompts and the context blocks coding agents add are left out.
- A TypeSafe Jev classifier picks the cheapest of the policy's tiers that would answer it well.
- If the classifier is unsure, the next stronger tier serves.
- If it is unavailable, a fallback tier serves: the strongest, unless your administrator chose another.
- When the policy names kinds of work, the same call also names the kind of work.
- When the policy uses a follow-up setting, the same call also judges whether a follow-up corrects the previous answer.
- A model of that tier serves the request. Models marked strong at the request's kind of work are tried first, and the policy can prefer the family of the API you call. The model must be permitted for you and large enough for the conversation.
From then on the request is exactly a request for that model. Its limits, pricing and usage record are that model's, and the provider's response reaches you unchanged.
This page is the API contract. For the concepts, the settings and how to read the report, see the Auto routing guide: how it works, configure it, read the report.
Use it
It works on every text endpoint: /v1/chat/completions, /v1/messages (including count_tokens) and /v1/responses.
curl -sS https://api.tokamak.sh/v1/chat/completions \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
-H "Content-Type: application/json" \
-D - \
-d '{"model": "tokamak/auto", "max_tokens": 512,
"messages": [{"role": "user", "content": "Rename usr to user in: func f(usr *User) {}"}]}'The response header names the model that served the request:
X-Tokamak-Routed-Model: anthropic/claude-haiku-4-5With the Tokamak CLI, pin it for a coding agent:
tokamak launch claude --model tokamak/autoHow conversations are handled:
- Tool loops. Every request of one turn, including the agent's tool-calling follow-ups, stays on the same model. The next message you type can land on a different one.
- Long conversations. If your administrator keeps long conversations on their model, a turn past a size they set that needs the same tier or a cheaper one stays where the conversation's prompt cache is (
sticky). - Responses continuations. A Responses request that continues with
previous_response_idstays on the model that produced that response. - Corrections. A follow-up that says the previous answer was wrong ("that didn't work", "try again") can start one tier above the tier that answered (
reason: "correction_escalated"). - Recognizing a conversation. Tokamak uses the session ID your client sends:
X-Session-Idand similar headers, Claude Code'smetadata.user_id, or Codex'sprompt_cache_key. Failing that, it uses the conversation's first message, when that is at least 24 characters.
Candidate models are chosen by your administrator. Make sure they can serve the API your client uses: a Claude Code session (/v1/messages) needs candidates served through the Messages API. A request routed to a model that cannot serve that API fails the same way as naming that model directly.
Availability
GET /v1/models lists tokamak/auto (display name Tokamak Auto, last in the list) only when your organization can use it. Its context_length and maximum output are the smallest among the candidate models, so clients that size their requests from the listing stay safe whichever model serves.
| Situation | Response |
|---|---|
| Auto routing is not enabled for your organization | 404, the same as any unknown model |
| No candidate model is available right now, or the router cannot run (for example, it cannot read the access policies) | 503 |
| Every candidate is excluded by your organization's access policy | 403 |
See what was chosen
Request listings under /v1/usage/*/requests show the model that served each request. Auto-routed requests also carry model_route:
"model_route": {
"requested": "tokamak/auto",
"tier": "balanced",
"classified_tier": "fast",
"reason": "correction_escalated",
"confidence": 1.0,
"task": "math",
"correction": true,
"previous_model": "openai/gpt-5.4-mini",
"previous_tier": "fast",
"baseline_model": "openai/gpt-6-astra",
"baseline_cost_usd": "0.0121"
}| Field | Meaning |
|---|---|
requested | Always tokamak/auto |
tier | The tier that served. The request's model is the model that served |
classified_tier | The classifier's own choice, before any escalation |
reason | Why that tier (table below) |
confidence | The classifier's confidence in its tier |
cached | The classification was reused from an earlier request of the same human turn |
task | The kind of work the classifier named, when the policy names kinds |
picked_by | strength or native_api when a strength or the API-family preference put this model ahead of the tier's first servable one; absent when the tier's order decided |
correction | The request corrects the previous turn's answer |
previous_model, previous_tier | The route of the conversation's previous turn, when the router remembered it |
sticky | The conversation stayed on its previous model for its prompt cache, rather than moving to another model of the same or a cheaper tier |
baseline_model, baseline_cost_usd | The strongest tier's model this request could have used (its own model when the strongest tier served it), and the same tokens priced at it. baseline_cost_usd is present when the baseline has a catalog price and the request's own cost is a real price: a catalog price, or on a billed request its full charge. A request your organization's own provider key served has no baseline_cost_usd |
reason | Tier used |
|---|---|
classified | The classifier's choice, at or above the policy's confidence floor |
low_confidence_escalated | One tier above the classifier's choice |
correction_escalated | One tier above the tier that gave the corrected answer |
classifier_failed | The fallback tier: error, timeout, or an answer that is not a tier |
no_request_text | The fallback tier: no human text in the body |
classifier_skipped | The fallback tier: a token count with no classified turn yet |
continuation | The model that produced the previous response |
Request listings accept three filters:
routed=true: only auto-routed requests;route_tier=<tier>: one tier;route_reason=<reason>: one reason.
Routing report
GET …/analytics/routing sums the decisions of a scope and window. Windows are up to 92 days; granularity is hour or day, by default hourly up to three days.
| Audience | Path |
|---|---|
| Your requests | /v1/usage/me/analytics/routing |
| A team | /v1/usage/teams/{teamId}/analytics/routing |
| Your organization (member view) | /v1/usage/active-org/analytics/routing |
| Your organization (administrators) | /v1/admin/active-org/usage/analytics/routing |
| The platform (platform administrators) | /v1/admin/analytics/routing |
Parameters are start_date, end_date, granularity, and the audience's usual filters (model, provider, api_key_id, user_id, team_id, organization_id).
| Field | Contents |
|---|---|
totals | The routed requests' metrics (below), plus all_requests and routed_share |
totals.classifier | calls, input_tokens, p50_ms, p95_ms, cost_usd (null when the classifier has no catalog price), fallback_rate, cached_share, escalated_share |
totals.sessions | follow_ups, corrections, correction_rate, remembered, escalated, sticky, sticky_share. Follow-ups, corrections and escalations count each turn once |
series | Per bucket and tier: requests, spend_usd, baseline_cost_usd, baseline_spend_usd |
by_tier, by_model, model_remainder | Metrics per tier (with rank, 0 = cheapest) and per served model (top 25; the rest summed in model_remainder) |
by_task, task_models | Metrics per kind of work ("" = none named), and per kind × served model with share of the kind's requests |
picked_by | Counts of "" (tier order), strength and native_api |
reasons, confidence | Counts per reason; ten confidence buckets |
rejections | Passed-over candidates by model, tier and reason |
classifiers, policies | Calls per classifier model; every policy version seen, with first and last use |
by_previous_tier, by_previous_model | follow_ups, corrections and correction_rate after each tier or model that gave the previous answer |
recent_corrections | The 20 latest corrected follow-ups: at, request_id, previous_model, previous_tier, model, tier, task, probability, escalated. null on the team, member and platform audiences, which list no individual requests |
Every metrics object has:
- volume and cost:
requests,tokens,spend_usd; - savings inputs:
baseline_requests,baseline_spend_usd,baseline_cost_usd; - outcome counts:
errors,truncated,cached,escalated,fallbacks,task_applied; - latency and confidence:
p50_latency_ms,p95_latency_ms,p50_ttft_ms,avg_confidence; - derived figures:
share,cost_per_request,savings_usd,savings_pct,error_rate,truncation_rate.
Savings compare baseline_cost_usd with baseline_spend_usd over the baseline_requests whose both prices are known.
Data sent to the classifier
To classify a request, Tokamak sends the classifier the text described in step 1: your latest message and up to two earlier ones, trimmed to a few thousand characters. It also sends the policy's tier and kind-of-work descriptions. It sends nothing else from the request: no system prompt, tool results, attachments, keys, identities or session IDs.
The classifier is TypeSafe's Jev (api.typesafe.ai). TypeSafe states that Jev is not trained on customer requests; see TypeSafe data handling. If that does not fit your data policy, name a model directly instead of tokamak/auto. See Under the hood for the exact questions and what Tokamak keeps.
Routing policies are set by platform administrators, per organization or as a platform default; see Configure auto routing.