Chat, Responses and Messages
Generate model output using the API format your application supports.
Choose the endpoint supported by your model. All three use your Tokamak inference key.
| Endpoint | Format | Input |
|---|---|---|
POST /v1/chat/completions | Chat Completions | messages |
POST /v1/responses | Responses | input |
POST /v1/messages | Anthropic Messages | messages |
Use the exact model ID from Models. Support for tools, images, reasoning and other optional parameters depends on the model and provider.
Chat Completions
curl -sS https://api.tokamak.sh/v1/chat/completions \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_CHAT_MODEL_ID",
"messages": [{"role": "user", "content": "Say hello."}],
"max_tokens": 64
}'Replace the placeholder with a Chat Completions-compatible model. For a text response, read choices[0].message.content.
| Field | Purpose |
|---|---|
model | Public model ID |
messages | Conversation messages |
max_tokens | Output limit for models supporting this field |
stream | Optional; set true for incremental output |
Some models require max_completion_tokens instead of max_tokens. Use the limit field supported by the selected model.
Responses
curl -sS https://api.tokamak.sh/v1/responses \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_RESPONSES_MODEL_ID",
"input": "Say hello.",
"max_output_tokens": 64
}'Use a Responses-compatible model. The native response contains output items in output; handle their types when extracting text or tool calls.
Saved-response operations are available when the upstream provider supports them:
| Method | Path |
|---|---|
| GET | /v1/responses/{response_id} |
| DELETE | /v1/responses/{response_id} |
| POST | /v1/responses/{response_id}/cancel |
| GET | /v1/responses/{response_id}/input_items |
Use the provider's response ID for these operations and keep the same authorized context.
Responses over WebSocket are not served. A WebSocket upgrade on GET /v1/responses returns 426 with code websocket_not_supported; use POST /v1/responses over HTTPS (in Codex, supports_websockets = false for this provider).
Anthropic Messages
curl -sS https://api.tokamak.sh/v1/messages \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
-H "Content-Type: application/json" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "YOUR_MESSAGES_MODEL_ID",
"max_tokens": 64,
"messages": [{"role": "user", "content": "Say hello."}]
}'Use a Messages-compatible model. Read the native content blocks in the response. A text block has type: "text" and a text value.
When used, system instructions belong in the top-level system field. Token counting is available through POST /v1/messages/count_tokens where supported by the provider.
Tools and media
Use the selected endpoint's native tool and media schemas. Your application validates and executes application tool calls, then sends the results back to the model. Tokamak does not execute your application's tools.
See Images and files for media input guidance.
Streaming and errors
Set stream: true to receive the chosen endpoint's event stream. See Streaming.
Check the HTTP status and native error body. Save X-Tokamak-Execution-Id from response headers when present if you need to inspect usage. See Errors and debugging for failed requests.