Streaming
Display model output as it arrives.
Set stream: true to receive incremental output. Use the stream format for your selected endpoint.
Chat Completions example
Replace YOUR_MODEL_ID with a compatible model from the catalog:
curl -sS -N https://api.tokamak.sh/v1/chat/completions \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_MODEL_ID",
"messages": [{"role": "user", "content": "Explain API routing briefly."}],
"max_tokens": 128,
"stream": true
}'-N disables curl buffering so chunks appear as they arrive.
Read the stream
| Endpoint | What to handle |
|---|---|
| Chat Completions | Native completion chunks, including text and tool-call deltas |
| Responses | Named Responses events and completed, failed or incomplete states |
| Anthropic Messages | Message and content-block events, including errors |
Use an SSE parser or a client compatible with the chosen protocol. One network read can contain part of an event or several events; do not assume every chunk is a complete JSON object.
For Chat Completions text, read the content deltas in choices. Compatible providers commonly terminate with data: [DONE]. Other endpoints have their own terminal events.
Long requests
Reasoning models can think for minutes before they write anything. To keep the connection open through that silence, Tokamak writes an SSE comment to the stream after 15 seconds without output:
: keep-aliveA line that starts with : is a comment under the SSE standard, and SSE parsers skip it, including the OpenAI and Anthropic SDKs, Claude Code and Codex. If you parse the stream yourself, ignore these lines. They carry no data and do not count as output or usage. Keep-alives are written only into an uncompressed stream: if your client sends Accept-Encoding and the provider compresses its stream, none are added, so send Accept-Encoding: identity on long reasoning requests whose connection is cut while the model thinks.
Some providers send nothing at all, not even response headers, until their first token. If a streaming request gets no answer from the provider within 30 seconds, Tokamak starts the stream itself: it answers 200 with Content-Type: text/event-stream and sends keep-alive comments until the provider answers. From then on:
- A successful answer streams as usual.
- A provider error arrives as one error event in the endpoint's own format, and the stream ends. The event is the same one the provider would send for an error in the middle of a stream:
| Endpoint | Error event |
|---|---|
| Chat Completions | data: {"error":{"message":...,"type":...}} |
| Anthropic Messages | event: error with {"type":"error","error":{"type":...,"message":...}} |
| Responses | event: response.failed with response.error.code and message |
So check for an error event in a stream even when the status was 200. Errors that arrive within the first 30 seconds, including most rate limits (429) and authentication failures, still come back as their real HTTP status.
Stream anything that may take more than about 100 seconds. A non-streaming request sends nothing until the whole answer is ready, and the network in front of the API closes a connection that stays silent that long. The answer is still generated and charged, but your client never receives it.
| Request | Limit |
|---|---|
| Non-streaming | About 100 seconds from request to the complete answer |
| Streaming, waiting for the provider's first answer | Tokamak starts the stream after 30 seconds, so the edge never cuts it; the silence limit below still applies |
| Streaming, silence from the model | 5 minutes without any output from the provider, before or after its first token, ends the stream as failed |
| Streaming, total | 60 minutes in Tokamak (STREAM_TIMEOUT); 15 minutes through api.tokamak.sh, the load balancer's limit |
Handle errors and cancellation
- Check HTTP status before parsing the stream.
- Handle error events after the stream starts.
- Treat an unexpected connection close as potentially incomplete.
- Cancel the underlying HTTP request when the user stops generation.
A successful HTTP status does not prove the stream completed. Retrying can create another generation and charge.
Usage
Usage can arrive near the end of a stream. Tokamak requests usage reporting from the provider for streaming Chat Completions.
Capture X-Tokamak-Execution-Id from response headers when present. Use generation lookup if your client missed final usage or needs billing status. Missing stream usage does not mean the request was free.