Under the hood
The classifier behind tokamak/auto, the exact questions it is asked, the evaluations behind the defaults, what it costs, what leaves your deployment, and why the router works the way it does.
The classifier
Auto routing is judged by TypeSafe's Jev, a System One model. Jev does not generate text. It reads a piece of text, the state, and answers typed questions about it with calibrated probabilities:
- a Choice picks one option and gives a probability for each;
- a Noul gives the probability that a statement is true.
Several questions can be asked in one call, which reads the state once. That is why the router asks everything it needs in a single call per human turn.
The call is an ordinary request through the Jev provider configured in Tokamak, with the same provider credential, capacity limits and rate-limit cooldown as any request. You can call Jev yourself on the System One endpoint.
The questions
Here is what the router sends for the follow-up "That's wrong. Recalculate it." The tier and kind-of-work descriptions are abbreviated:
{
"model": "jev-latest",
"state": {
"request": "That's wrong. Recalculate it.",
"earlier_requests": ["Convert 72 degrees Fahrenheit to Celsius."]
},
"questions": {
"tier": {
"type": "choice",
"instructions": "Which model tier is the cheapest one that would answer the `request` well? The `request` may be a short follow-up; judge it in the context of the `earlier_requests` from the same conversation.",
"criteria": {
"fast": "A small, fast model is enough: greetings, thanks, short factual lookups, …",
"balanced": "A capable mid-size model is needed: ordinary coding tasks such as …",
"frontier": "The strongest model is needed: deep multi-step reasoning, …"
}
},
"task": {
"type": "choice",
"instructions": "What kind of work does the `request` mainly ask for? Judge the work the user wants done, not the subject they mention. The `request` may be a short follow-up; judge it in the context of the `earlier_requests` from the same conversation.",
"criteria": { "chat": "…", "coding": "…", "math": "…", "…": "…" }
},
"correction": {
"type": "noul",
"instructions": "Is the `request` a complaint, correction or retry because an earlier answer in this conversation (see `earlier_requests`) was wrong, incomplete or did not work?",
"criteria": {
"true": "The user says the previous answer failed, was wrong, misunderstood them, or asks to redo it",
"false": "The user continues, accepts, extends or asks something new"
}
}
}
}And the answers, abbreviated:
{
"model": "jev-1.13.0",
"answers": {
"tier": { "type": "choice", "choice": "fast", "confidence": 1.0,
"probabilities": { "fast": 1.0, "balanced": 0.0, "frontier": 0.0 } },
"task": { "type": "choice", "choice": "math", "confidence": 0.97 },
"correction": { "type": "noul", "noul": 0.98 }
},
"usage": { "input_tokens": 958 }
}The router took fast from the tier answer and saw the correction. The previous answer had come from the fast tier, so it started on balanced.
The wording carries weight.
- "The cheapest one that would answer the request well" is what makes the classifier pick the lower of two adequate tiers.
- The follow-up clause makes a bare "continue" inherit the difficulty of the conversation it continues.
- "The work the user wants done, not the subject they mention" keeps a request that mentions React but asks for a SQL query classified as SQL work.
- The tiers' own descriptions are the administrator's, and they are what the tier answer is judged against.
The evidence
The defaults were chosen on labelled evaluations against the live classifier. The labels are one person's, so treat the figures as indicative, not as a benchmark.
| Evaluation | Set | Result |
|---|---|---|
| Tier question and confidence floor | 34 labelled prompts across the three default tiers | All 34 routed correctly with the 0.5 floor. A 0.3 floor missed the one hard prompt the classifier was unsure about; 0.7 sent three prompts a tier too high |
| Earlier turns | Bare follow-ups such as "yes, do it" | Classified as the cheapest tier at 0.95 confidence on their own; routed correctly with two earlier turns |
| Kind of work, in the same call | The 34 prompts plus 12 covering front-end work, writing, research, security and math | 45 of 46 kinds correct; the miss named "Sort these words alphabetically" other. Lowest confidence 0.67. Asking it changed 1 of 46 tier answers, one already at 0.28 confidence |
| Correction, alone | 18 labelled follow-ups to one coding request | 18 of 18 correct. Corrections scored 0.94 to 0.98, including "try again", "This doesn't compile" and a Vietnamese "still broken, fix it". Continuations such as "thanks, that worked" and "now add tests" scored 0.02 to 0.23 |
| All three questions together | The 18 follow-ups, and the 46 prompts sent as follow-ups | Corrections 18 of 18. No new request flagged as a correction. Every kind of work unchanged. Tier answers unchanged on 44 of 46, and both changes moved to the labelled tier |
Cost and speed
- One call per human turn. Every tool-loop request of the turn reuses the answer, and
count_tokensnever calls. - Input tokens: about 500 to 700 for the tier question alone. With kinds of work and the correction question, about 850 to 950. Output is free at TypeSafe's list price.
- Price: at the list price of 42 µUSD per 1,000 input tokens, that is roughly $0.00002 to $0.00004 per human turn.
- Latency: about 250 to 300 ms median when warm. Adding the second and third questions did not measurably change it.
The call is not recorded as a usage row of its own. Its tokens and latency are recorded on the decision of the request that made it, and the report prices them at the classifier's catalog rate.
Caches and memory
| What | Kept for | Why | Where |
|---|---|---|---|
| The classifier's answer | 15 minutes, per policy revision and human turns | Every request of one turn gets the same answer, and the same model | Each server |
| Each conversation's previous route | 2 hours, up to 16,384 conversations total and 1,024 per organization (or per authenticated user/API-key pair for organization-less requests) | Correction escalation and stickiness | Each server, in memory only |
| The routing policy that governs an organization | 15 seconds | Policy changes reach every server within 15 seconds | Each server |
| A caller's access policies | 15 seconds | Access changes reach the router within 15 seconds | Each server |
| Classifier backoff after a failure | 30 seconds, per policy | An outage costs one timeout, not one per request | Each server |
| Which active provider serves each candidate model | 30 seconds | A model or provider switched on or off reaches the other servers' candidate checks within 30 seconds; the server that made the change sees it at once | Each server |
Each server keeps its own. A conversation whose requests land on different servers may be classified again, which costs one extra call, or may have no previous turn on that server. That costs that turn's escalation or stickiness and nothing else.
What leaves your deployment
For each classification, Tokamak sends TypeSafe (api.typesafe.ai) the extracted text:
- Sent: your latest message, up to 4,000 characters, and up to two earlier messages, 600 characters each.
- Also sent: the policy's tier and kind-of-work descriptions, and the fixed question text above.
- Never sent: system prompts, tool results, attachments, images, the model's answers, API keys, user or organization identities, or session IDs.
TypeSafe states that Jev is not trained on customer requests and offers zero data retention to enterprise customers; see TypeSafe data handling. Whether that fits is a decision for each deployment to make before enabling a policy. People who cannot send text to the classifier can always name a model directly.
What Tokamak keeps
-
The decision. It is stored with the request's usage record as
model_route:- the tier, the reason, the classifier's answers and version;
- the kind of work and the correction probability;
- the previous turn's model;
- the candidates passed over;
- the policy version;
- the baseline model and cost.
It holds no request text.
-
The conversation memory. It holds a hash of the conversation's key and the models and tiers of its last two turns, in memory, for up to two hours. It is not written to the database.
-
The answer cache. It holds the classifier's answers keyed by a hash of the text, not the text itself.
Why it works this way
- Classify the ask, not the transcript. Tool output, system prompts and the blocks agents inject describe the conversation, not what the person asked for. They are large and identical from request to request, and a classifier's accuracy falls as its input fills with detail the question does not need.
- Doubt goes up, never down. Low confidence escalates a tier; a failure takes the fallback, which is the strongest tier by default. When the router cannot show that a cheaper model is good enough, it does not guess.
- One call, every question. The state is read once, and more questions cost only their own tokens.
- One answer per turn. Agents make many requests per human message. Reusing the answer keeps a turn on one model and keeps the classifier's cost per turn, not per request.
- Strengths are claims; the report is the evidence. The router does what the policy says. The kind of work × model table shows whether the claim holds.
- Stickiness only sideways or down. Keeping a cache is worth a slightly pricier model; it is never worth a weaker one. A correction always gets to move.
- Ask the correction question only where it is used. Policies without a follow-up setting send exactly the call they sent before.
- Remember in memory, not in a database. The previous turn is a hint that steers one decision. Losing it costs that decision's refinement, never correctness, so it needs no new store to protect.
- Record every decision with its request. Every figure in the report can be traced back to the requests behind it.
Limits
- No learned selection. The tiers, their order and the strengths are the only inputs. The router does not learn from outcomes by itself; the report shows you the outcomes so you can adjust the policy.
- The router does not know which APIs a model serves. It checks that a candidate is servable, not that it speaks the request's API. Candidates must suit your clients.
- Text only. Images and attachments are not seen, so a request whose difficulty lies in an image is judged by its words.
- The memory is per server. A conversation spread across many servers gets fewer escalations and less stickiness.
- Savings are an estimate. They price the same tokens at the strongest tier's model; they do not say what that model would have answered.
- It is not a budget. Usage limits still cap spend, and they apply to the model that served.
Read the routing report
What the Auto routing report shows — where auto-routed requests went, what they saved against the strongest tier, how each tier, model and kind of work fared, and how often a follow-up corrected an answer — and how to act on it.
Bring your own key
Route your organization's Anthropic and OpenAI traffic through your own provider accounts. Tokamak still meters every request at a price you set, and charges nothing for what your keys serve.