DocsAPI Reference
Endpoints

Chat, Responses and Messages

Generate model output using the API format your application supports.


Choose the endpoint supported by your model. All three use your Tokamak inference key.

EndpointFormatInput
POST /v1/chat/completionsChat Completionsmessages
POST /v1/responsesResponsesinput
POST /v1/messagesAnthropic Messagesmessages

Use the exact model ID from Models. Support for tools, images, reasoning and other optional parameters depends on the model and provider.

Chat Completions

curl -sS https://api.tokamak.sh/v1/chat/completions \
  -H "Authorization: Bearer $TOKAMAK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_CHAT_MODEL_ID",
    "messages": [{"role": "user", "content": "Say hello."}],
    "max_tokens": 64
  }'

Replace the placeholder with a Chat Completions-compatible model. For a text response, read choices[0].message.content.

FieldPurpose
modelPublic model ID
messagesConversation messages
max_tokensOutput limit for models supporting this field
streamOptional; set true for incremental output

Some models require max_completion_tokens instead of max_tokens. Use the limit field supported by the selected model.

Responses

curl -sS https://api.tokamak.sh/v1/responses \
  -H "Authorization: Bearer $TOKAMAK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "YOUR_RESPONSES_MODEL_ID",
    "input": "Say hello.",
    "max_output_tokens": 64
  }'

Use a Responses-compatible model. The native response contains output items in output; handle their types when extracting text or tool calls.

Saved-response operations are available when the upstream provider supports them:

MethodPath
GET/v1/responses/{response_id}
DELETE/v1/responses/{response_id}
POST/v1/responses/{response_id}/cancel
GET/v1/responses/{response_id}/input_items

Use the provider's response ID for these operations and keep the same authorized context.

Responses over WebSocket are not served. A WebSocket upgrade on GET /v1/responses returns 426 with code websocket_not_supported; use POST /v1/responses over HTTPS (in Codex, supports_websockets = false for this provider).

Anthropic Messages

curl -sS https://api.tokamak.sh/v1/messages \
  -H "Authorization: Bearer $TOKAMAK_API_KEY" \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "YOUR_MESSAGES_MODEL_ID",
    "max_tokens": 64,
    "messages": [{"role": "user", "content": "Say hello."}]
  }'

Use a Messages-compatible model. Read the native content blocks in the response. A text block has type: "text" and a text value.

When used, system instructions belong in the top-level system field. Token counting is available through POST /v1/messages/count_tokens where supported by the provider.

Tools and media

Use the selected endpoint's native tool and media schemas. Your application validates and executes application tool calls, then sends the results back to the model. Tokamak does not execute your application's tools.

See Images and files for media input guidance.

Streaming and errors

Set stream: true to receive the chosen endpoint's event stream. See Streaming.

Check the HTTP status and native error body. Save X-Tokamak-Execution-Id from response headers when present if you need to inspect usage. See Errors and debugging for failed requests.

On this page