DocsAPI Reference
Auto routing

Read the routing report

What the Auto routing report shows — where auto-routed requests went, what they saved against the strongest tier, how each tier, model and kind of work fared, and how often a follow-up corrected an answer — and how to act on it.


Every auto-routed request carries the router's decision. The routing report sums those decisions with each request's outcome: cost, errors, truncated answers and latency. It answers three questions:

  • Is auto routing saving money?
  • Is it sending work to models that can do it?
  • Is the classifier healthy?

Where to find it

WhereWhoScope
Customer app › Analytics › Auto routing (Advanced view)Anyone who can see AnalyticsYour own requests, a team, or the organization. The audience switch works as on every other Analytics tab
Admin console › Models › Auto router › Routing qualityPlatform administratorsAll organizations, or one. The Usage analytics link opens the full report
APISee belowThe same report as JSON

The report covers up to 92 days and buckets its series by hour for windows up to three days, by day beyond. Where your audience can see individual requests, click a tier or a reason in the report to open Requests narrowed to it. The request detail then says why that request went where it did.

Totals

The top of the Auto routing report: six figures for auto-routed requests, savings against the strongest tier, auto-routed spend, classifier latency and fallback rate, low-confidence escalations and errors, above a stacked daily chart of requests per tier
The top of the report over seven days of test traffic. Figures come from local test requests, not a customer workload. Open full size ↗
FigureWhat it says
Auto-routed requestsHow many requests named tokamak/auto, and their share of all requests in the scope
Saved vs strongest tierThe strongest-tier cost minus what those requests actually cost, and the percentage saved. See How savings are computed
Auto-routed spendWhat the auto-routed requests cost, and the cost per request
ClassifierThe classifier's median latency, and how often it fell back because it failed
Escalated for low confidenceRequests sent one tier up because the classifier was unsure
ErrorsThe error rate, and how many answers were cut off by their output limit (truncated)

Tier mix over time stacks requests per tier, cheapest at the bottom. A step in the mix often lines up with a policy revision, which the report lists at the end.

By tier and by model

Two tables, each with requests, share, spend, cost per request, savings, error and truncation rates, and median and 95th-percentile latency:

  • By tier lists tiers cheapest first, with the classifier's mean confidence. Click a tier to list its requests.
  • By model lists the models that served, busiest first. After the top 25, the rest are summed into one row.

Read them against each other. A cheap tier with a high error or truncation rate is doing work it should not get. A strong tier taking most requests means the savings are small; check whether its description is too broad.

By kind of work

The kind of work table: rows for coding, math, writing, debugging, architecture and more; columns for served models; each cell shows the request count and error or truncation rate, shaded by the model's share of that kind; beside it, a bar list of why each model served within its tier
Kinds of work against the models that served them, from test traffic. The darker the cell, the larger that model's share of the kind. Model names are examples. Open full size ↗

When the policy names kinds of work, each row is a kind and each column a served model:

  • Each cell shows how many requests of that kind the model took, and their error and truncation rates.
  • The shade is the model's share of the kind.
  • Under each kind, the row says how often the classifier was confident enough for strengths to apply.

This table is the evidence for setting strengths. If one model of a tier takes a kind's requests with fewer errors or truncated answers than the other, mark it strong at that kind; see Configure › Kinds of work and strengths.

Why this model within its tier counts how the model was chosen inside its tier:

  • Tier order: the tier's first servable model.
  • Strength: a model strong at the request's kind went first.
  • Native API: the API-family preference decided.

Follow-up turns

Corrections by the tier that answered before: follow-ups, corrections and a correction rate for the fast, balanced and frontier tiers, with a correction rate per model below; beside it, totals for follow-ups, corrections, previous turn known, escalations on correction and conversations kept on their model
Corrections charged to the tier and model that gave the answer being corrected. On this test deployment the traffic script sent more complaints after short prompts, so the rates show the shape of the view, not a measurement. Model names are examples. Open full size ↗

With either follow-up setting on, the classifier judges on every follow-up whether it corrects the previous answer. Each correction is charged to the route that gave that answer, not the one that served the complaint. Each turn counts once, however many tool calls it made.

  • Corrections by the tier that answered before. Follow-ups, corrections and the correction rate after each tier, with the rate after each model below.
    • A clearly higher rate after the cheapest tier is the signal of under-routing: that tier is being sent work it cannot do.
    • A rate that is high everywhere is more likely the work itself, or the users.
  • Follow-up turns:
    • how many turns were follow-ups, and how many of them were corrections;
    • how many had a remembered previous turn (each server remembers its own conversations, so a conversation spread across servers is remembered less);
    • how many corrections were escalated;
    • how many requests stayed on their model for its prompt cache.
  • Latest corrections. The 20 most recent corrected follow-ups, each with the route that gave the answer, the route that served the correction, the kind of work and the classifier's probability. Those escalated by the policy are marked, and Show escalated corrections opens them in Requests. This list appears only where the report's audience can already see individual requests: your own requests, and the organization for its administrators. Team, member and platform views show the totals only.

Why requests went where they did

Reasons counts each decision reason:

  • classified;
  • escalated for low confidence;
  • escalated on a correction;
  • fallback after the classifier failed;
  • no request text;
  • not classified, for a token count;
  • continuation.

Click a reason to list its requests. Reason codes are defined in How it works.

Classifier confidence is a histogram of how sure the classifier was of its tier. A lot of mass just below the confidence floor means many requests are paying for the next tier up. Below the histogram:

  • the number of classifier calls and their input tokens;
  • median and 95th-percentile latency;
  • how often a turn reused an earlier answer;
  • the classifier's cost at its catalog price.

Passed-over candidates lists the models requests skipped on the way to the one that served, and why:

  • not servable;
  • refused by an access policy;
  • context window too small;
  • the conversation is held by another provider.

A frequent refusal usually means the policy lists a model its users cannot reach.

Policy revisions in this window lists every policy version the requests were routed under, with when it was first and last used. A change of policy changes what the figures above measure.

How savings are computed

For every auto-routed request, the router names a baseline model. That is the model the request would have used had every request gone to the strongest tier: the request's own model if it was served there, otherwise the first model in the strongest tier the request could have used. That model must be active, allowed for the caller and large enough for the request. The usage record prices the request's own tokens at that model.

Saved is that baseline cost minus what the same requests actually cost. It is summed only over requests where both prices are known: the baseline's catalog price, and the request's own catalog price or, when billing charged it, its full charge. A request billing only observed (shadow mode) or whose charge was capped at its credit ceiling is left out, and so is a request your organization's own provider key served: its cost is your internal price, so counting it would report your discount as a routing saving. The savings figure says how many requests were counted.

It is a list-price estimate of the alternative, the same tokens on a stronger model. It is not what that model would have produced, and it does not include the classifier's own cost, which the classifier figures report separately.

Acting on it

You seeTry
High correction rate after the cheapest tierNarrow the fast tier's description, or switch on escalation on correction
Many low-confidence escalationsMake neighbouring tier descriptions more distinct
Classifier fallback rate above a few percentRaise the classifier timeout; check the Jev provider
High error or truncation on one modelMove it later in its tier, or out of it
One model doing clearly better at a kind of workMark it strong at that kind
Savings on only some requestsGive the served and baseline models catalog prices
A frequent passed-over candidateFix its provider or access policy, or remove it

API

The report is GET …/analytics/routing on every Analytics audience:

AudiencePath
Your own requests/v1/usage/me/analytics/routing
A team/v1/usage/teams/{teamId}/analytics/routing
The organization, member view/v1/usage/active-org/analytics/routing
The organization, administrator view/v1/admin/active-org/usage/analytics/routing
The platform/v1/admin/analytics/routing (one organization with organization_id)

It accepts start_date, end_date, granularity=hour|day, and the audience's usual filters (model, provider, api_key_id, user_id, team_id, organization_id). Request listings accept:

  • routed=true: only auto-routed requests;
  • route_tier: one tier;
  • route_reason: one reason.

The response fields are listed in the Auto routing API reference.

On this page