Usage
Check your usage and inspect the cost of an individual model request.
Use Analytics in the workspace for an overview, or read usage through the API.
Your usage summary
GET /v1/usage/me
curl -sS https://api.tokamak.sh/v1/usage/me \
-H "Authorization: Bearer $TOKAMAK_API_KEY"This reads your usage across credentials, not only the key used for this request.
For recent recorded requests, use GET /v1/usage/me/requests. For daily totals, use GET /v1/usage/me/daily.
Inspect one generation
Capture the X-Tokamak-Execution-Id response header from an inference request, then set it as EXECUTION_ID:
curl -sS --get https://api.tokamak.sh/v1/generation \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
--data-urlencode "id=$EXECUTION_ID"The response contains one record in data.
| Field | Meaning |
|---|---|
id | Tokamak execution UUID |
model | Requested model, when recorded |
streamed | Whether the request streamed |
native_tokens_prompt | Recorded input tokens, when available |
native_tokens_completion | Recorded output tokens, when available |
billing_status | final, pending, not_billed or cancelled |
total_cost | Customer charge, when known |
currency | Currency, when available |
Use a credential for the same organization and payer context. A foreign or unavailable execution is returned as not found. The provider's response ID is not the Tokamak execution UUID.
Read costs correctly
A missing cost or token count means it is unavailable, not zero. Billing can remain pending after a response completes.
Generation money fields are exact JSON numbers. Use lossless decimal parsing when reconciling charges. Analytics money fields use decimal strings.
If returned, rated_cost describes measured usage cost and authorized_amount describes the admission ceiling. Neither is a substitute for the final customer charge.
Correlate requests
You can send X-Client-Request-Id with your own short identifier on inference requests. Search it with:
curl -sS --get https://api.tokamak.sh/v1/generations \
-H "Authorization: Bearer $TOKAMAK_API_KEY" \
--data-urlencode "client_request_id=my-app-request-001"The response always contains an array in data. A correlation value can match multiple executions; it does not make retries idempotent.
Tag requests
Send X-Tokamak-Tag on inference requests to label them, for example by experiment or batch: X-Tokamak-Tag: exp-42, task:pick-and-place. Repeat the header or separate tags with commas. Tags are lower-cased, must match ^[a-z0-9][a-z0-9._:/@-]{0,63}$, and at most 16 are kept per request; a malformed tag is dropped without failing the request. Tags are not sent to the provider. In Analytics, group by Tag: a request with two tags counts under both, so tag rows do not add up to the total. Tags are labels only and never affect routing, pricing, limits or access.
Limits and failures
GET /v1/usage/limits/current reads current usage-limit status. Usage limits and wallet credit are separate checks.
Requests rejected before metering may have no usage record. See Errors and debugging when an expected record is missing.