DocsAPI Reference
Auto routing

Under the hood

The classifier behind tokamak/auto, the exact questions it is asked, the evaluations behind the defaults, what it costs, what leaves your deployment, and why the router works the way it does.


The classifier

Auto routing is judged by TypeSafe's Jev, a System One model. Jev does not generate text. It reads a piece of text, the state, and answers typed questions about it with calibrated probabilities:

  • a Choice picks one option and gives a probability for each;
  • a Noul gives the probability that a statement is true.

Several questions can be asked in one call, which reads the state once. That is why the router asks everything it needs in a single call per human turn.

The call is an ordinary request through the Jev provider configured in Tokamak, with the same provider credential, capacity limits and rate-limit cooldown as any request. You can call Jev yourself on the System One endpoint.

The questions

Here is what the router sends for the follow-up "That's wrong. Recalculate it." The tier and kind-of-work descriptions are abbreviated:

{
  "model": "jev-latest",
  "state": {
    "request": "That's wrong. Recalculate it.",
    "earlier_requests": ["Convert 72 degrees Fahrenheit to Celsius."]
  },
  "questions": {
    "tier": {
      "type": "choice",
      "instructions": "Which model tier is the cheapest one that would answer the `request` well? The `request` may be a short follow-up; judge it in the context of the `earlier_requests` from the same conversation.",
      "criteria": {
        "fast": "A small, fast model is enough: greetings, thanks, short factual lookups, …",
        "balanced": "A capable mid-size model is needed: ordinary coding tasks such as …",
        "frontier": "The strongest model is needed: deep multi-step reasoning, …"
      }
    },
    "task": {
      "type": "choice",
      "instructions": "What kind of work does the `request` mainly ask for? Judge the work the user wants done, not the subject they mention. The `request` may be a short follow-up; judge it in the context of the `earlier_requests` from the same conversation.",
      "criteria": { "chat": "…", "coding": "…", "math": "…", "…": "…" }
    },
    "correction": {
      "type": "noul",
      "instructions": "Is the `request` a complaint, correction or retry because an earlier answer in this conversation (see `earlier_requests`) was wrong, incomplete or did not work?",
      "criteria": {
        "true": "The user says the previous answer failed, was wrong, misunderstood them, or asks to redo it",
        "false": "The user continues, accepts, extends or asks something new"
      }
    }
  }
}

And the answers, abbreviated:

{
  "model": "jev-1.13.0",
  "answers": {
    "tier": { "type": "choice", "choice": "fast", "confidence": 1.0,
              "probabilities": { "fast": 1.0, "balanced": 0.0, "frontier": 0.0 } },
    "task": { "type": "choice", "choice": "math", "confidence": 0.97 },
    "correction": { "type": "noul", "noul": 0.98 }
  },
  "usage": { "input_tokens": 958 }
}

The router took fast from the tier answer and saw the correction. The previous answer had come from the fast tier, so it started on balanced.

The wording carries weight.

  • "The cheapest one that would answer the request well" is what makes the classifier pick the lower of two adequate tiers.
  • The follow-up clause makes a bare "continue" inherit the difficulty of the conversation it continues.
  • "The work the user wants done, not the subject they mention" keeps a request that mentions React but asks for a SQL query classified as SQL work.
  • The tiers' own descriptions are the administrator's, and they are what the tier answer is judged against.

The evidence

The defaults were chosen on labelled evaluations against the live classifier. The labels are one person's, so treat the figures as indicative, not as a benchmark.

EvaluationSetResult
Tier question and confidence floor34 labelled prompts across the three default tiersAll 34 routed correctly with the 0.5 floor. A 0.3 floor missed the one hard prompt the classifier was unsure about; 0.7 sent three prompts a tier too high
Earlier turnsBare follow-ups such as "yes, do it"Classified as the cheapest tier at 0.95 confidence on their own; routed correctly with two earlier turns
Kind of work, in the same callThe 34 prompts plus 12 covering front-end work, writing, research, security and math45 of 46 kinds correct; the miss named "Sort these words alphabetically" other. Lowest confidence 0.67. Asking it changed 1 of 46 tier answers, one already at 0.28 confidence
Correction, alone18 labelled follow-ups to one coding request18 of 18 correct. Corrections scored 0.94 to 0.98, including "try again", "This doesn't compile" and a Vietnamese "still broken, fix it". Continuations such as "thanks, that worked" and "now add tests" scored 0.02 to 0.23
All three questions togetherThe 18 follow-ups, and the 46 prompts sent as follow-upsCorrections 18 of 18. No new request flagged as a correction. Every kind of work unchanged. Tier answers unchanged on 44 of 46, and both changes moved to the labelled tier

Cost and speed

  • One call per human turn. Every tool-loop request of the turn reuses the answer, and count_tokens never calls.
  • Input tokens: about 500 to 700 for the tier question alone. With kinds of work and the correction question, about 850 to 950. Output is free at TypeSafe's list price.
  • Price: at the list price of 42 µUSD per 1,000 input tokens, that is roughly $0.00002 to $0.00004 per human turn.
  • Latency: about 250 to 300 ms median when warm. Adding the second and third questions did not measurably change it.

The call is not recorded as a usage row of its own. Its tokens and latency are recorded on the decision of the request that made it, and the report prices them at the classifier's catalog rate.

Caches and memory

WhatKept forWhyWhere
The classifier's answer15 minutes, per policy revision and human turnsEvery request of one turn gets the same answer, and the same modelEach server
Each conversation's previous route2 hours, up to 16,384 conversations total and 1,024 per organization (or per authenticated user/API-key pair for organization-less requests)Correction escalation and stickinessEach server, in memory only
The routing policy that governs an organization15 secondsPolicy changes reach every server within 15 secondsEach server
A caller's access policies15 secondsAccess changes reach the router within 15 secondsEach server
Classifier backoff after a failure30 seconds, per policyAn outage costs one timeout, not one per requestEach server
Which active provider serves each candidate model30 secondsA model or provider switched on or off reaches the other servers' candidate checks within 30 seconds; the server that made the change sees it at onceEach server

Each server keeps its own. A conversation whose requests land on different servers may be classified again, which costs one extra call, or may have no previous turn on that server. That costs that turn's escalation or stickiness and nothing else.

What leaves your deployment

For each classification, Tokamak sends TypeSafe (api.typesafe.ai) the extracted text:

  • Sent: your latest message, up to 4,000 characters, and up to two earlier messages, 600 characters each.
  • Also sent: the policy's tier and kind-of-work descriptions, and the fixed question text above.
  • Never sent: system prompts, tool results, attachments, images, the model's answers, API keys, user or organization identities, or session IDs.

TypeSafe states that Jev is not trained on customer requests and offers zero data retention to enterprise customers; see TypeSafe data handling. Whether that fits is a decision for each deployment to make before enabling a policy. People who cannot send text to the classifier can always name a model directly.

What Tokamak keeps

  • The decision. It is stored with the request's usage record as model_route:

    • the tier, the reason, the classifier's answers and version;
    • the kind of work and the correction probability;
    • the previous turn's model;
    • the candidates passed over;
    • the policy version;
    • the baseline model and cost.

    It holds no request text.

  • The conversation memory. It holds a hash of the conversation's key and the models and tiers of its last two turns, in memory, for up to two hours. It is not written to the database.

  • The answer cache. It holds the classifier's answers keyed by a hash of the text, not the text itself.

Why it works this way

  • Classify the ask, not the transcript. Tool output, system prompts and the blocks agents inject describe the conversation, not what the person asked for. They are large and identical from request to request, and a classifier's accuracy falls as its input fills with detail the question does not need.
  • Doubt goes up, never down. Low confidence escalates a tier; a failure takes the fallback, which is the strongest tier by default. When the router cannot show that a cheaper model is good enough, it does not guess.
  • One call, every question. The state is read once, and more questions cost only their own tokens.
  • One answer per turn. Agents make many requests per human message. Reusing the answer keeps a turn on one model and keeps the classifier's cost per turn, not per request.
  • Strengths are claims; the report is the evidence. The router does what the policy says. The kind of work × model table shows whether the claim holds.
  • Stickiness only sideways or down. Keeping a cache is worth a slightly pricier model; it is never worth a weaker one. A correction always gets to move.
  • Ask the correction question only where it is used. Policies without a follow-up setting send exactly the call they sent before.
  • Remember in memory, not in a database. The previous turn is a hint that steers one decision. Losing it costs that decision's refinement, never correctness, so it needs no new store to protect.
  • Record every decision with its request. Every figure in the report can be traced back to the requests behind it.

Limits

  • No learned selection. The tiers, their order and the strengths are the only inputs. The router does not learn from outcomes by itself; the report shows you the outcomes so you can adjust the policy.
  • The router does not know which APIs a model serves. It checks that a candidate is servable, not that it speaks the request's API. Candidates must suit your clients.
  • Text only. Images and attachments are not seen, so a request whose difficulty lies in an image is judged by its words.
  • The memory is per server. A conversation spread across many servers gets fewer escalations and less stickiness.
  • Savings are an estimate. They price the same tokens at the strongest tier's model; they do not say what that model would have answered.
  • It is not a budget. Usage limits still cap spend, and they apply to the model that served.

On this page