DocsAPI Reference
Auto routing

Auto routing

Name one model, tokamak/auto. Tokamak sends each request to the cheapest model that can answer it well, learns from follow-ups that say an answer was wrong, and shows what it saved.


tokamak/auto is a model you can name like any other. It is a virtual model: behind it is a routing policy your platform administrator sets up. The policy has tiers of real models, ordered from the cheapest to the strongest. For every request, Tokamak asks a small classifier which tier is the cheapest one that would answer it well. A model from that tier then serves the request, exactly as if you had named it.

Why use it

Most requests in a working session do not need the strongest model. In a coding session, "rename this variable", "yes, do it" and "convert this JSON to YAML" sit between the few requests that do need it, such as "design the failover for this service" or "find why this leaks goroutines". Naming one strong model for everything pays frontier prices for all of it. Naming a cheap model gets the hard requests wrong. Auto routing makes that choice per request.

  • You pay frontier prices only for frontier work. Easy requests go to a small, fast model and hard ones to the strongest. The routing report compares what you spent with what the same tokens would have cost on the strongest tier.
  • Doubt costs a little extra, never a worse answer.
    • When the classifier is unsure, the next stronger tier serves.
    • When the classifier is unavailable, the fallback tier serves; that is the strongest tier unless your administrator chose another.
    • When a follow-up says the previous answer was wrong, it can start one tier higher.
    • A tier cheaper than the classifier's choice serves only when no model of that tier or a stronger one can serve at all.
  • Within a tier, the model that fits the work.
    • A tier can hold several models, each marked as better at some kinds of work, for example one for math and research and another for coding and front-end work.
    • A tier can also prefer the model family whose API the client speaks, such as Claude models for Claude Code.
  • Agent sessions stay coherent and cache-friendly.
    • Every tool-calling request of one agent turn stays on one model.
    • A long conversation can stay on the model that already holds its prompt cache instead of hopping to a cheaper one and paying for the whole context again.
  • One model ID for every client. It works on Chat Completions, Messages (including count_tokens) and Responses, so it fits an SDK, Claude Code or Codex alike.
  • Every decision is visible.
    • The response names the model that served it.
    • Every request records why that model was chosen.
    • The report shows where traffic went, what it cost, how each tier and model fared, and how often a follow-up corrected an answer.
  • It is cheap to decide. One classifier call per human turn, about 300 ms and a few thousandths of a cent at list price. Every tool-loop request after the first reuses it.

Who does what

You areYou doRead
A developer or a coding-agent userSend model: "tokamak/auto" and read which model served youAuto routing API
A platform administratorCreate the routing policy: tiers, models, kinds of work, follow-up settingsConfigure auto routing
An organization administrator or analystCheck what routing saved, where it went and whether answers held upRead the routing report
Anyone who wants the detailsSee the classifier, the questions it is asked, the evidence and the data it seesHow it works, Under the hood

Try it

  1. A platform administrator enables a policy in the admin console under Models › Auto router. Until one governs your organization, tokamak/auto answers 404 like any unknown model and is not listed. See Configure auto routing.

  2. Name the model.

    curl -sS https://api.tokamak.sh/v1/chat/completions \
      -H "Authorization: Bearer $TOKAMAK_API_KEY" \
      -H "Content-Type: application/json" \
      -D - \
      -d '{"model": "tokamak/auto",
           "messages": [{"role": "user", "content": "Convert 72 degrees Fahrenheit to Celsius."}]}'

    For a coding agent, launch it with tokamak launch claude --model tokamak/auto, or with the equivalent command for your tool.

  3. See what served you. The response carries X-Tokamak-Routed-Model. The request appears in Analytics › Requests with an auto badge, and the Auto routing tab of Analytics (Advanced view) sums all of them.

Use your deployment's gateway address in place of api.tokamak.sh.

Good to know

  • The classifier reads text you wrote. It sees the text of your latest message, including anything you pasted into it, and of up to two messages before it. It never sees tool results, attachments, images, system prompts or the model's answers. It is TypeSafe's Jev. If your data policy does not allow that, name a model directly. See Under the hood.
  • Candidates must speak your client's API. Tokamak does not translate between APIs. A Claude Code session needs candidates served through the Messages API. Your administrator chooses the candidates.
  • The served model's rules apply. Usage limits, pricing and the usage record are those of the model that served the request, not of tokamak/auto.
  • Your organization's own provider keys still apply. The router chooses among models a Tokamak provider serves. When your organization's key maps the chosen model, the key serves it, exactly as if you had named that model. A model only your key serves is never a candidate, and requests your key serves are left out of the routing report's savings.

On this page