How auto routing works
The five steps between a request that names tokamak/auto and the model that answers it, and how follow-ups, agent tool loops and long conversations are handled.
A request that names tokamak/auto goes through five steps before it reaches a provider. After the fifth, it is an ordinary request for the model the router chose.
1. Read the human turns
The router extracts the text a person wrote and nothing else:
- Chat Completions and Messages: the
userentries ofmessages, using only theirtextparts. - Responses: the
inputstring, or themessageitems of theinputarray. - Left out: tool results, images, documents, assistant turns and system prompts. So are the blocks coding agents inject into your messages, such as Claude Code's
<system-reminder>and Codex's<environment_context>and<user_instructions>. A message made only of such blocks does not count as something you wrote.
The latest human turn is kept up to 4,000 characters; a longer one keeps about its first 3,000 and its last 1,000 characters. Up to two earlier human turns ride along, 600 characters each.
The earlier turns matter. Without them, a follow-up such as "yes, do it" after a hard question was classified as the cheapest tier at 0.95 confidence; with two earlier turns it was routed correctly in every case measured.
A request with no human text at all takes the policy's fallback tier. An example is a Responses request whose input is only a tool result.
2. Classify in one call
The router makes one call to the policy's classifier, TypeSafe's Jev System One model (jev/jev-latest by default). The call goes through the Jev provider configured in Tokamak, with that provider's credential and capacity. Jev does not write text: it answers typed questions about a piece of text with calibrated probabilities. Up to three questions travel in that one call:
| Question | Asked | Answer |
|---|---|---|
| Tier: which tier is the cheapest one that would answer the request well? | Always. The options are the policy's tiers, each described by the administrator's text | A tier, a probability for every tier and a confidence |
| Kind of work: what kind of work does the request mainly ask for? | When the policy names kinds of work (coding, math, writing, …) | A kind, with probabilities and a confidence |
| Correction: does the request say an earlier answer was wrong, incomplete or did not work? | On a follow-up (a request with earlier human turns), when the policy uses either follow-up setting | A probability; 0.5 or more counts as a correction |
The exact wording and the evidence behind it are in Under the hood.
One answer per human turn. The answer is cached for 15 minutes, keyed on the policy revision and the extracted text. Every tool-calling request that follows the same human message reuses it, so an agent turn stays on one model and pays for one classification. Only the request that made the call records the classifier's tokens and latency. The ones that reuse it are marked cached.
Token counting never classifies. count_tokens reuses an existing answer for the same turns, or takes the fallback tier (classifier_skipped). It generates nothing to justify a classifier call.
3. Choose the tier
The tier comes from the classifier's answer, adjusted only upward:
- The classifier's tier is the starting point.
- If its confidence is below the policy's confidence floor (0.5 by default), the next stronger tier serves instead. Uncertainty escalates; it never downgrades.
- If the request corrects the previous answer and the policy escalates on correction, it starts at least one tier above the tier that gave that answer. When that answer came from the strongest tier, it stays on the strongest. This never lowers a tier the classifier already put higher.
Every decision records the reason:
reason | Tier used |
|---|---|
classified | The classifier's choice, at or above the floor |
low_confidence_escalated | One tier above the classifier's choice |
correction_escalated | One tier above the tier that gave the corrected answer |
classifier_failed | The fallback tier: error, timeout, or an answer that is not a tier |
no_request_text | The fallback tier: no human text to judge |
classifier_skipped | The fallback tier: a token count with no earlier classification of its turns |
continuation | The model that produced the previous response (a Responses request with previous_response_id) |
The fallback tier is the strongest tier unless the administrator chose another. When the router cannot show that a cheaper model is good enough, it does not guess.
4. Pick the model
A tier can list several candidate models. The request tries them in this order:
- Strengths. When the kind of work's confidence reaches the policy's kind-of-work floor (0.5 by default), models marked strong at that kind go first. The decision's
picked_byisstrengthwhen that moved another model ahead of the tier's first. - API family. With prefer the request's API family on, a Messages request (Claude Code) tries
anthropic/*models first and a Responses request triesopenai/*models first. Chat Completions keeps the listed order, because nearly every family also serves it.picked_byisnative_apiwhen this decided.
The walk starts at the chosen tier and moves to stronger tiers, then to cheaper ones. The first candidate that passes every check serves. A candidate that fails a check is recorded as passed over:
| Passed over as | When |
|---|---|
unavailable | No active provider serves the model |
not_generative | It is a classifier model, not one that writes answers |
access_policy | An access policy that blocks for the organization, team, user or API key refuses it. An observe-mode policy lets it through and flags the decision |
context_window | The request's estimated size is larger than the model's context window |
other_provider | A continuing Responses conversation is held by another provider |
The size estimate is deliberately rough. It is a quarter of the request's bytes, which overstates real tokens. So a candidate refused only for its window is kept as a last resort instead of failing the request (context_window_ignored).
When nothing can serve, the request fails:
403when every candidate was refused by an access policy;503when none is available.
The router also answers 503 if it cannot load the access policies. It will not dispatch a model it cannot show is permitted.
Stickiness comes last. With the policy's keep long conversations on their model setting, the router checks whether the walk moved a long conversation to another model of the same tier or a cheaper one. If it did, and the previous turn's model still passes every check, that model serves instead. The decision is marked sticky. See Conversations and sessions.
5. Serve and record
The chosen model's ID replaces tokamak/auto before the request is resolved. Everything after that behaves as if you had named the model:
- admission checks and credit;
- usage limits, including a limit scoped to that model;
- pricing and the provider forward;
- the usage record, whose model is the chosen one.
Nothing is recorded under tokamak/auto. The provider's response reaches you unchanged, plus one header:
X-Tokamak-Routed-Model: openai/gpt-5.4-miniThe decision is stored with the request as model_route:
- the tier, the classifier's own choice, the reason, the confidence and probabilities;
- the kind of work and why a model was picked;
- the correction answer and the previous turn's route;
- stickiness;
- the candidates passed over;
- the policy revision.
It also names a baseline model: the request's own model when the strongest tier served it, otherwise the first model of the strongest tier this request could have used (active, permitted for the caller and large enough). The usage record prices the same tokens at that model (baseline_cost_usd). That comparison is what the report's savings are made of. The fields are listed in the API reference.
Conversations and sessions
Agent tool loops stay on one model. Tool calls do not add human turns, so every request of one agent turn reuses the turn's classification and lands on the same model. The next message you type is classified again and can land elsewhere, unless the conversation is long enough to stay put (below).
Responses continuations stay where the conversation is. A Responses request with previous_response_id continues a conversation the provider holds. It stays on the model that produced that response while the policy still lists it and it passes every check (continuation). Otherwise it is classified as usual but limited to candidates on the provider that holds the conversation. Only when that provider serves none of them does it route without that limit (continuation_unpinned).
The router remembers each conversation's previous turn. Correction escalation and stickiness need to know which model answered last. The router keeps, per conversation, the latest routed turn and the one before it:
- How a conversation is recognized:
- the session ID the client attaches:
X-Session-Idand similar headers, the session inside Claude Code'smetadata.user_id, or Codex'sprompt_cache_key; - failing that, the conversation's first message, when it is at least 24 characters. A greeting such as "hi" opens too many conversations to name one.
- the session ID the client attaches:
- Scope: always one caller (organization, user and API key), so two people never share a conversation.
- Where it lives: in memory on each server, for up to two hours, and never in the database. A conversation whose requests reach another server simply has no previous turn there for that turn. That costs the escalation or the stickiness, never correctness.
- What does not count as a turn: token counts and console previews.
Escalate on correction. "That didn't work", "try again", or "you missed one" after a cheap model's answer would otherwise be judged on those few words, often straight back onto the cheap tier. With the setting on, such a follow-up starts one tier above the tier that answered. The agent's tool calls in that turn keep the same previous turn, so they do not escalate again. The decision records correction, its probability and the previous model.
Keep long conversations on their model. A provider's prompt cache makes continuing a long conversation on the same model much cheaper than starting it over on another. With many providers a cache read costs about a tenth of fresh input. When the request's estimated size reaches the policy's threshold, which new policies set at 20,000 tokens, a turn that would move to another model of the same tier or a cheaper one stays on the previous model instead.
A turn that needs a stronger tier still moves up. A turn that corrects the previous answer never stays: the person just said that model was wrong. sticky is recorded only when it changed the model.
A session, step by step
Here is how one policy handled real requests on a test deployment with the live classifier. The model names are examples. The policy:
- Tiers: fast (
gpt-5.4-mini), balanced (gpt-5.4), and frontier (gpt-6-astrathenclaude-fable-5). - Strengths: Astra is marked for architecture, math and research; Fable for coding, front-end work and security.
- Settings: the native-API preference, escalate on correction, and stickiness from 20,000 tokens are all on.
| The person sends | The classifier says | Served by | Why |
|---|---|---|---|
| "Design the architecture for a multi-region, active-active API gateway…", in a conversation of about 28,000 tokens | frontier 0.99, architecture | gpt-6-astra | classified; Astra is first and strong at architecture |
| The agent's tool calls for that turn | (reused) | gpt-6-astra | The same human turn, cached |
| "Great, thanks. Now turn the three approaches into a bullet list for the slide deck." | fast 0.82, writing; correction 0.04 | gpt-6-astra | sticky: the long conversation keeps its cache |
| The same follow-up in a short conversation | fast 0.82 | gpt-5.4-mini | classified |
| "Now prove that your token-bucket reconciliation converges under partitions…", back in the long conversation | frontier 0.99, math | gpt-6-astra | classified; the walk chose Astra by itself, so it is not marked sticky |
| "Convert 72 degrees Fahrenheit to Celsius.", in a new conversation | fast 1.00 | gpt-5.4-mini | classified |
| "That's wrong. Recalculate it." | fast 1.00; correction 0.98 | gpt-5.4 | correction_escalated: one tier above the answer it corrects |
| "Thanks, that's right now. Also convert 100 degrees Fahrenheit." | fast 1.00; correction 0.10 | gpt-5.4-mini | classified |
| "Refactor our 3,000-line payment reconciliation module into clean architecture layers…" | frontier, coding | claude-fable-5 | picked_by: strength: Fable is marked for coding |
| From Claude Code: "This Go service leaks goroutines only in production…" | frontier 0.97, debugging | claude-fable-5 | picked_by: native_api: no model is marked for debugging, and Claude Code speaks Messages |
| The same request over Chat Completions | frontier 0.97, debugging | gpt-6-astra | The tier's listed order |
When something goes wrong
| What happens | What the router does |
|---|---|
| The classifier times out, errors or answers something that is not a tier | Serves the fallback tier. After a timeout or an error it skips that policy's classifier for 30 seconds, so an outage costs one timeout rather than one per request |
| The caller goes away during classification | Not counted as a classifier failure |
| No candidate can serve | 503, or 403 when every candidate was refused by an access policy |
| The chosen model cannot speak the client's API | The provider's error, exactly as if you had named that model. Choose candidates your clients' APIs can reach |
| The organization has no enabled policy | 404, the same answer as any unknown model; tokamak/auto is not listed |
Policy changes reach every server within 15 seconds, and apply at once on the server that saved them. Access-policy changes also apply within 15 seconds. A candidate model or provider switched on or off can take up to 30 seconds to reach the other servers' candidate checks.
Auto routing
Name one model, tokamak/auto. Tokamak sends each request to the cheapest model that can answer it well, learns from follow-ups that say an answer was wrong, and shows what it saved.
Configure auto routing
Set up the policy behind tokamak/auto in the admin console. Choose tiers and models, the classifier, kinds of work and strengths, and follow-up settings, and try them before your users do.