Error Codes
HTTP status codes and error response shapes returned by the Tokamak API
Tokamak returns different error envelopes depending on which surface produced the error. There are four, plus the admission refusals described below, and the difference matters if you are parsing them.
Platform envelope
Most /v1/* endpoints — media, usage, admin, models — return:
{
"code": "5e1d3524-929e-4c7a-9bb7-0a8b74fa6f10",
"error": "Media object not found",
"message": "no media object with that id",
"request_id": "req_abc123"
}| Field | Meaning |
|---|---|
code | An opaque correlation identifier for the specific call site that raised the error. It is not a stable symbolic error code — do not branch on it. Quote it in support requests. |
error | Human-readable summary |
message | Optional additional detail |
request_id | Present when a request ID was assigned; use it to find the trace |
Branch on the HTTP status, not on code.
Anthropic envelope
POST /v1/messages returns the Anthropic-style error shape so the Anthropic SDK can parse it:
{
"type": "error",
"error": {
"type": "api_error",
"message": "upstream request failed"
}
}The complete set of error.type values:
error.type | HTTP status | Meaning |
|---|---|---|
invalid_request_error | 400 | Malformed request |
request_too_large | 413 | Payload exceeds the size limit |
authentication_error | 401 | Credential missing, expired, or invalid |
permission_error | 403 | Authenticated but not permitted |
not_found_error | 404 | Resource or model does not exist |
rate_limit_error | 429 | Rate limit exceeded |
api_error | 5xx | Internal failure (including a failed model lookup), a temporarily unavailable dependency such as tokamak/auto routing, or an unrecognized upstream error |
billing_error | 402 | Insufficient credit; only when /v1/messages refusals are switched to the Anthropic dialect (see Admission refusals) |
byok_key_unavailable | 503 | Your organization's own provider key could not serve the request, and fallback to Tokamak is off. See Bring your own key |
Responses envelope
/v1/responses and its sub-resources return the OpenAI error shape:
{"error": {"message": "response not found", "type": "invalid_request_error", "code": "response_not_found", "param": null}}error.type follows the status: authentication_error (401), permission_error (403), rate_limit_error (429), server_error (5xx), otherwise invalid_request_error. A WebSocket upgrade on GET /v1/responses gets 426 with code websocket_not_supported; send POST /v1/responses over HTTPS instead.
Admission refusals
Before a generation reaches a provider, Tokamak can refuse it for credit, a usage budget or traffic capacity. On POST /v1/chat/completions, /v1/messages and /v1/responses these refusals use their own envelope, whatever the route's other errors look like. By default error is the code as a string, and message sits beside it:
{
"error": "insufficient_credit",
"message": "Insufficient credit for this request. Add funds and retry; see /v1/billing/credits for the current balance.",
"billing": {"code": "insufficient_credit", "retryable": false, "balance_url": "/v1/billing/credits", "top_up_url": "/v1/billing/deposits"}
}An operator can switch a route to its own dialect so Claude Code, Codex and the OpenAI and Anthropic SDKs can read the refusal. Then error becomes that dialect's error object and carries the code in error.code. The billing, result, retryable and retry_after_seconds fields stay where they were. On /v1/messages:
{
"type": "error",
"error": {"type": "billing_error", "message": "Insufficient credit for this request. ...", "code": "insufficient_credit"},
"billing": {"code": "insufficient_credit", "retryable": false, "balance_url": "/v1/billing/credits", "top_up_url": "/v1/billing/deposits"}
}On /v1/chat/completions and /v1/responses:
{
"error": {"message": "Insufficient credit for this request. ...", "type": "billing_error", "param": null, "code": "insufficient_credit"},
"billing": {"code": "insufficient_credit", "retryable": false, "balance_url": "/v1/billing/credits", "top_up_url": "/v1/billing/deposits"}
}The status code and Retry-After header are the same in both forms. To read the code in either form, use billing.code for credit refusals and error.code (dialect) or error (default) otherwise. Codes: insufficient_credit (402), billing_frozen (403), not_safely_billable (422), billing_unavailable and principal_provisioning (503), usage_limit_exceeded (429), rate_limit, concurrency_limit and upstream_cooldown (429), traffic_token_bound_required (422) and traffic_capacity_unavailable (503).
Bring your own key
When your organization's own provider key cannot serve a request and fallback to Tokamak is off, the generation routes answer 503 in their own envelope:
| Route | Where the code is |
|---|---|
POST /v1/chat/completions | error.type and error.code are both byok_key_unavailable |
POST /v1/messages, /v1/messages/count_tokens | error.type is byok_key_unavailable, in the Anthropic envelope |
POST /v1/responses and the stored-response routes | error.code is byok_key_unavailable; error.type is server_error |
The provider-key management routes (/v1/admin/active-org/provider-credentials, byok-templates, model-prices, byok/*, and the console's /v1/admin/organizations/{id}/byok) answer {"code": …, "error": …}, where code is a stable code you can branch on: invalid_request (400), byok_disabled (403), byok_kind_not_allowed (403), not_found (404), key_rejected (422), byok_encryption_unconfigured (503) and internal_error (500). See When a key fails.
Guardrails
When your organization's guardrail blocks a request, every generation route (and /v1/messages/count_tokens) answers 400 in its own envelope with error.type invalid_request_error and error.code guardrail_blocked, plus a top-level guardrail object naming the policy, rule, entity and stage. The message never contains the matched value. The request was not sent to a provider and is not charged.
If the policy could not be loaded and the organization chose to refuse requests in that case, the answer is 503 with error.type api_error and error.code guardrail_unavailable: an outage, not a content refusal.
Every screened response carries X-Tokamak-Guardrail (none, flagged, masked, cloaked or blocked) and X-Tokamak-Guardrail-Count.
Gateway envelope
Errors raised by the gateway before a request reaches core or auth are plain strings:
{"error": "permission denied"}| Status | Body | Cause |
|---|---|---|
401 | {"error":"authentication required"} | No credential presented |
401 | {"error":"unauthorized"} | Credential rejected |
400 | {"error":"authorization scope required"} | The route needs a scope the request did not resolve |
403 | {"error":"permission denied"} | Authenticated, but the required permission is missing |
502 | {"error":"authorization unavailable"} | The auth service could not be reached |
502 | {"error":"delegation token unavailable"} | The gateway could not mint an internal delegation token |
HTTP status codes
Internally, errors carry a type that maps onto a status code:
| Internal type | Status |
|---|---|
VALIDATION | 400 Bad Request |
UNAUTHORIZED | 401 Unauthorized |
FORBIDDEN | 403 Forbidden |
NOT_FOUND | 404 Not Found |
CONFLICT | 409 Conflict |
NOT_IMPLEMENTED | 501 Not Implemented |
EXTERNAL | 502 Bad Gateway |
INTERNAL, DATABASE_ERROR, TOO_MANY_RECORDS | 500 Internal Server Error |
These type names are internal and are not serialized into the response body — only the status code and the human-readable text reach the client.
Note that validation failures return 400. A 422 from Tokamak means it would not admit the request: an output length it can't bound for a billed organization (not_safely_billable) or under a token-per-minute limit (traffic_token_bound_required); see Admission refusals. A provider's own validation error can also arrive as 422.
Common errors
| Scenario | Status |
|---|---|
| No credential header | 401 |
| Access token expired | 401 |
| API key revoked or expired | 401 |
| Caller lacks the route's required permission | 403 |
| Unknown media or model ID | 404 |
| Removing the last Owner of a scope | 409, with error code owner_floor |
| Request body or file exceeds the size limit | 413 |
Billed org's chat completion has no max_tokens and the model has no max_completion_tokens cap | 422 |
| Rate limit exceeded | 429 |
WebSocket upgrade on GET /v1/responses | 426 |
| Upstream provider failure | 502 |