DocsAPI Reference
Auto routing

Configure auto routing

Set up the policy behind tokamak/auto in the admin console. Choose tiers and models, the classifier, kinds of work and strengths, and follow-up settings, and try them before your users do.


Platform administrators manage routing policies in the admin console under Models › Auto router, or through the admin API. Doing either needs the admin.models.manage permission.

  • The platform default governs every organization that has no policy of its own. Until a platform default or an organization's override is enabled, tokamak/auto answers 404 for that organization and is not listed.
  • An organization override replaces the default entirely for one organization. It can narrow the candidates, change the bar, or switch tokamak/auto off for that organization alone. Removing the override hands the organization back to the platform default.
The platform default policy in the admin console: status and settings badges above three tiers, the frontier tier listing two models with their strengths
The platform default policy on a test deployment. The badges summarize the settings; each tier lists its models in order, with the kinds of work each is marked strong at. Model names are examples. Open full size ↗

Before you start

  • A classifier is connected. Add TypeSafe as a provider of kind jev, so that jev/jev-latest is in the catalog; see System One. A synced Jev model gets TypeSafe's list price, 42 µUSD per 1,000 input tokens, unless the provider states its own; keep it priced so the report can price classifications.
  • The candidate models are in the catalog, active and priced. Savings in the report are computed only where the strongest tier's model has a catalog price and the request's own cost is known: the served model's catalog price or, on a billed request, its full charge.
  • The candidates speak your clients' APIs. Tokamak forwards each API as it is and does not translate. Claude Code needs candidates served through the Messages API, and Codex needs the Responses API. A request routed to a model that cannot answer its API fails exactly as it would if the client had named that model.
  • Access policies are in place. The router never sends a request to a model that an access policy which blocks refuses the caller. A candidate the caller cannot reach is skipped. Only tokamak/auto reads access policies: a request that names a model directly is not checked against them, and they have no console page or API yet.
  • You know who goes first. Many deployments start with one organization's override as a pilot, watch its report, then create the platform default.

1. Create the policy

Choose Create on the platform default, or pick an organization from Add override…. The editor opens with a starting draft:

  • the three evaluated tiers (fast, balanced, frontier) with their descriptions;
  • the ten kinds of work;
  • the API-family preference;
  • both follow-up settings.

An organization's override starts from the platform default when there is one. The draft does not serve anything until you switch on Serve tokamak/auto and save.

2. Tiers and models

Tiers are the heart of the policy. They are ordered cheapest first, and that order is also the escalation order.

RuleLimit
Tiers per policy2 to 6
Tier nameUp to 32 characters: lowercase letters, digits, -, _
Tier descriptionRequired, up to 600 characters
Models per tier1 to 8 catalog IDs; not a classifier model and not tokamak/auto itself. A model may sit in more than one tier

The description is all the classifier knows about a tier. The classifier is asked which tier is the cheapest one that would answer the request well, and it judges against your descriptions. Some guidance:

  • Describe the work each tier is enough for, with concrete examples.
  • Keep neighbouring tiers clearly apart.
  • Start from the defaults, which routed all 34 prompts of the evaluation correctly:
TierDefault description
fastA small, fast model is enough: greetings, thanks, short factual lookups, one-line commands, typo fixes, format conversion, trivial renames or edits.
balancedA capable mid-size model is needed: ordinary coding tasks such as writing a function with tests, fixing a clear bug, writing a query, explaining or refactoring a small piece of code, a short comparison.
frontierThe strongest model is needed: deep multi-step reasoning, system or architecture design with trade-offs, subtle concurrency or distributed-systems bugs, security review, proofs, large migrations or refactors.

List models in order of preference. The first model that can serve does serve, unless a strength or the API preference puts another first. List fallbacks after it: a model that is down, refused by an access policy, or too small for the request yields to the next.

The tier section of the policy editor: three tiers with descriptions and models; under the frontier tier, strength chips per model with architecture, math and research marked for one model and coding, frontend and security for the other
Tiers in the editor. Strength chips appear under a tier that has two or more models to choose between. Model names are examples. Open full size ↗

3. Classifier settings

SettingRangeDefaultWhat it does
ClassifierA catalog model served by a jev providerjev/jev-latestThe System One model that answers the questions
Confidence floor0 to 10.5Below it, the classified tier is escalated one step
Classifier timeout100 to 10,000 ms1,500 msPast it, the fallback tier serves
Fallback tierAny tierThe strongestServes requests the classifier could not judge

The confidence floor trades cost for safety. In the evaluation, a floor of 0.5 routed every prompt correctly. At 0.3 it missed the one hard prompt the classifier had been unsure about, and at 0.7 it sent three prompts to a stronger tier than they needed. A warm classifier call takes about 250 to 300 ms, so the default timeout leaves room. If the report shows a noticeable fallback rate, raise the timeout before anything else.

4. Kinds of work and strengths

A tier with several models can send different work to different models. For example, the frontier tier can send math and research to one model and coding, front-end work and security to another.

  1. Switch on Route by kind of work. The classifier then also names the kind of work, in the same call. The ten default kinds are chat, coding, frontend, debugging, architecture, security, math, writing, research, and other as a catch-all. Edit their descriptions or add your own: 2 to 12 kinds, each with a name like a tier's and a description of up to 300 characters.
  2. Set the kind-of-work confidence floor (0.5 by default). Below it the kind is recorded but does not reorder the tier.
  3. Under each tier with two or more models, mark which kinds each model is strong at. For a request of a marked kind, the models strong at it go first, in their listed order.

The editor warns when a strength cannot change anything. An example is when only the tier's first model is marked and the API-family preference is off: that model goes first anyway.

Set strengths from evidence, not reputation. The routing report's kind of work × model table shows how each kind fared on each model. Watch it for a while before marking strengths, and revisit the marks when it changes. Asking the kind of work adds about 300 input tokens to each classification.

5. Prefer the request's API family

With Prefer the request's own API family on, a tier's candidates are ordered by the API the request arrived on:

  • Messages requests (Claude Code) try anthropic/* models first.
  • Responses requests (Codex) try openai/* models first.
  • Chat Completions keeps the listed order, because nearly every family serves it too.

It matters when a tier mixes vendors, so that each client meets the family it was built around. A strength outranks it.

6. Follow-up turns

The policy editor's settings: serve switch, classifier, confidence floor, timeout, fallback tier, API-family preference, the Follow-up turns section with two switches and a token threshold, and the kinds of work
The policy settings. Follow-up turns holds the two settings this section describes. Open full size ↗

Escalate when a follow-up says the answer was wrong. Examples are "That didn't work", "try again" and "you missed one". The follow-up then starts one tier above the tier that gave the answer. When that answer already came from the strongest tier, it stays there. With the setting off, a correction does not change the tier.

Keep long conversations on their model. Set the threshold in estimated input tokens. When a turn reaches it and would move to another model of the same tier or a cheaper one, it stays on the model that answered the previous turn instead, keeping that model's prompt cache. The trade-off:

  • With many providers a cache read costs about a tenth of fresh input. So staying is cheaper whenever the cheaper model's input price is above a tenth of the current model's.
  • At around 20,000 tokens that is almost always true, which is why new policies start there.
  • A turn that needs a stronger tier still moves up, and a correction never stays put.

Either setting also turns on the correction question. On every follow-up the classifier then judges whether it corrects the previous answer, which costs about 90 more input tokens. The report starts counting corrections against the route that gave the answer. With both settings off, which is how every policy created before these settings existed stays, the classifier call is exactly what it was. Both depend on the router knowing the conversation's previous turn; see Conversations and sessions.

7. Try a prompt

Try a prompt, below the policies, runs the saved policy of the platform or an organization against a sample conversation. It shows the decision a real request would get. Nothing is forwarded to a model, and no cache or conversation memory is touched.

FieldUse
PolicyThe platform default, or an organization's override
Arrives onChat Completions, Messages (Claude Code) or Responses (Codex); what the API-family preference orders by
Earlier promptsUp to five earlier messages of the conversation, separated by a blank line
Earlier prompts answered byThe model that answered them, as a conversation would remember it: shows a correction's escalation
Conversation sizeTokens to assume: shows stickiness and the context-window check
PromptThe message to route

The result shows:

  • the model and tier;
  • the reason;
  • the classifier's confidence and its probability for each tier;
  • the kind of work, and why the model was picked;
  • whether the prompt corrects the previous answer and whether the conversation stayed on its model;
  • the classifier's version, latency and input tokens;
  • the strongest-tier alternative the report uses as its baseline;
  • every candidate passed over, and why.
Try a prompt: the earlier prompt asks to convert 72 degrees Fahrenheit, the previous answer came from the fast tier's model, and the prompt says the answer is wrong. The decision routes to the balanced tier with the reason corrects the previous answer, escalated
A correction after the cheap model's answer. The classifier judges the words fast with full confidence, but the follow-up corrects the fast tier's answer, so it starts on balanced. Model names are examples. Open full size ↗
Try a prompt: a large refactoring request arriving on the Messages API is routed to the frontier tier's second model, with the note that it was served by the tier's model strong at coding
A frontier coding request goes to the tier's model marked strong at coding rather than the tier's first model. Model names are examples. Open full size ↗

Previews always call the classifier, even while a live policy is backing off after a failure. That makes them a quick way to check whether the classifier has recovered.

8. Switch it on and roll out

Switch on Serve tokamak/auto and save. Each save gives the policy a new version.

  • It applies at once on the server that saved it, and within 15 seconds everywhere.
  • The routing report lists every version seen in its window, so you can tell which figures belong to which policy.

For organizations it governs:

  • GET /v1/models lists tokamak/auto last, as Tokamak Auto.
  • Its context length and maximum output are the smallest among the candidates, so clients that size their requests from the listing stay safe whichever model serves.

To take tokamak/auto away from one organization only, give it an override with Serve tokamak/auto off.

Policy reference

The "If omitted" column is also what a policy saved before a field existed has. The last column is what the console's new-policy draft starts with.

FieldMeaningRangeIf omittedConsole draft
enabledServe tokamak/autoon/offoffoff
classifier_modelThe System One modela jev catalog modeljev/jev-latestjev/jev-latest
min_confidenceTier confidence floor0–10.50.5
classifier_timeout_msClassifier deadline100–10,0001,5001,500
fallback_tierTier for unjudged requestsa tier name; empty is the strongestthe strongestthe strongest
tiers[]name, description, models[], optional strengths2–6 tiersrequiredfast, balanced, frontier
tiers[].strengthsModel → kinds of work it goes first forkinds from tasksnonenone
tasks[]Kinds of work: name, descriptionnone, or 2–12nonethe ten defaults
task_min_confidenceKind-of-work floor0–10.50.5
prefer_native_apiAPI-family preferenceon/offoffon
escalate_on_correctionEscalate a correctionon/offoffon
sticky_min_context_tokensStickiness threshold; 0 is off0–2,000,000020,000

A save is validated before it is stored:

  • the classifier must be a jev model;
  • every candidate must be a catalog model that is not one;
  • tier and kind names must be unique;
  • a strength must name a model of its tier and a kind of the policy.

A rejected save answers 400 with the field that is wrong.

Admin API

The console uses these endpoints on the gateway. They need a platform session with admin.models.manage:

Method and pathDoes
GET /v1/admin/models/routerEvery policy, plus the defaults a new policy starts from
PUT /v1/admin/models/router/platformCreate or replace the platform default
PUT /v1/admin/models/router/organizations/{organizationId}Create or replace one organization's override
DELETE /v1/admin/models/router/organizations/{organizationId}Remove an override
POST /v1/admin/models/router/previewTry a prompt. Send policy to preview an unsaved draft; also accepts earlier_prompts, dialect, previous_model and context_tokens
{
  "enabled": true,
  "classifier_model": "jev/jev-latest",
  "min_confidence": 0.5,
  "classifier_timeout_ms": 1500,
  "tiers": [
    { "name": "fast", "description": "A small, fast model is enough: greetings, short lookups, trivial edits.",
      "models": ["openai/gpt-5.4-mini"] },
    { "name": "balanced", "description": "A capable mid-size model is needed: ordinary coding tasks.",
      "models": ["openai/gpt-5.4"] },
    { "name": "frontier", "description": "The strongest model is needed: deep reasoning, design, subtle bugs, proofs.",
      "models": ["openai/gpt-6-astra", "anthropic/claude-fable-5"],
      "strengths": { "openai/gpt-6-astra": ["math", "research", "architecture"],
                     "anthropic/claude-fable-5": ["coding", "frontend", "security"] } }
  ],
  "tasks": [
    { "name": "coding", "description": "Writing, editing, explaining or testing ordinary code." },
    { "name": "math", "description": "Mathematics, proofs, derivations or algorithm complexity." }
  ],
  "task_min_confidence": 0.5,
  "prefer_native_api": true,
  "escalate_on_correction": true,
  "sticky_min_context_tokens": 20000
}

The model names are examples; use IDs from your catalog. The tasks list is shortened here, and a real policy should use the ten defaults returned under defaults.tasks.

Tune with the report

The routing report is how you find out whether the policy does what you meant. Common signals:

You seeIt usually meansTry
A high correction rate after the cheapest tierThe fast tier gets work it cannot doNarrow its description, or switch on escalation on correction
Many requests escalated for low confidence between the same two tiersTheir descriptions overlapMake the difference between them concrete
A noticeable classifier fallback rateTimeouts, or an outageRaise the timeout; check the Jev provider
One candidate passed over again and againIt is down, refused by an access policy, or too smallFix the provider or the access policy, or reorder the tier
One model of a tier doing clearly better at a kind of workA strength worth markingMark it, then watch the table
Savings shown for only some requestsSome served or baseline models have no catalog price, or some requests were capped at their credit ceiling, only observed by billing, or served by your own provider keyPrice the models; the other requests are left out by design
Long agent sessions hopping between modelsNo stickinessSwitch on Keep long conversations on their model

On this page