DocsAPI Reference

Guardrails

Screen what your organization sends to models. Find secrets, personal data and denylisted content in prompts, tool results and answers, and flag, mask, cloak or block it, with no client changes.


Guardrails are content policies your organization applies to its own inference traffic. A policy holds rules that find secrets (API keys, tokens, passwords, private keys), personal data (emails, phone numbers, cards, bank and ID numbers) and words or patterns you choose. Each match is flagged, masked, cloaked or blocked before the request reaches a model, and optionally on the way back too.

Screening runs inside Tokamak on every route (/v1/chat/completions, /v1/messages, /v1/responses, /v1/messages/count_tokens), in every turn, including tool results such as a .env file an agent just read. Nothing is sent to another service to classify it. Clients do not change: a coding agent keeps working, and only what it would have leaked changes.

Three things to know before you start:

  • It is an organization opt-in. Guardrails are off until a platform administrator switches the feature on for your organization. Then your owners and admins (org.admin) write the policies.
  • It never touches billing. A match is not billing evidence and grants nothing. A blocked request is never sent to a provider and is not charged. Cost stays in Analytics.
  • No header can switch it off. A policy is chosen by the API key and the organization, never by anything a caller sends, including X-Tokamak-Tag.

1. Switch it on (platform administrator)

In the platform console open Organizations → your organization → Features and turn on Guardrails. The Guardrails tab on the same sheet shows counts only. See Guardrails for administrators.

2. Create a policy

In the app open Settings → Guardrails. You need org.admin.

The Guardrails policies tab: a 7-day match bar with flag 27, mask 24, cloak 32 and block 17, the default policy acme-default and a second policy ci-strict, and eight presets below
Policies. The default policy applies to every key without its own policy. Presets are a starting point. Open full size ↗

Start from a preset or New guardrail. A policy is a list of rules that run in order:

RuleFindsActions
Secrets223 secret formats (AWS, GitHub, Stripe, Slack, OpenAI, Anthropic, private keys, …) from the gitleaks ruleset, passwords in connection strings and Tokamak keysflag · mask · block
PIIemail, phone, US and Vietnamese phone, card number (Luhn + issuer), IBAN (checksum), US SSN, Vietnamese CCCD (needs a context word), public IP, MAC, crypto wallet, and your own custom entitiesflag · mask · cloak · block, per entity
Keywordwords you list, optionally whole words onlyflag · mask · block
Regexup to 100 RE2 patternsflag · mask · block
Size caprequests longer than a number of charactersblock
Prompt injectioninstruction overrides, role-tag spoofing and hidden Unicode in tool output and the newest user turnflag · block
The policy editor for acme-default: name, description, and four rules in order: secrets, pii, rival and injection
A policy. Saving makes a new version; the history keeps every version and can revert to one. Open full size ↗

Click a rule to edit it. For PII each entity can have its own action, so cards can block while emails cloak.

The rule drawer for the pii rule: applies to both requests and responses, default action cloak, card number and US SSN set to block, Vietnam CCCD set to mask
Per-entity actions. Validated entities (cards, IBANs, SSNs, CCCD) only match numbers that pass their checks. Open full size ↗

What each action does

ActionThe provider receivesYour client receives
Flagthe request unchangedthe answer unchanged; the match is recorded
Maskthe match replaced with a tag, e.g. [AWS_ACCESS_TOKEN]the answer as the model wrote it
Cloaka realistic stand-in, e.g. [email protected], a test card number, an IBAN with valid check digitsthe answer with your real values put back, in text and in tool-call arguments
Blocknothing400 guardrail_blocked in the route's own error format; not charged

A blocked request looks like this. The message names the rule and where it matched, never the value:

{"type":"error","error":{"type":"invalid_request_error","code":"guardrail_blocked",
 "message":"Request blocked by guardrail \"acme-default\" (rule \"pii\": credit_card in messages[0].content)."},
 "guardrail":{"policy":"acme-default","rule":"pii","entity":"credit_card","stage":"input"}}

Every screened response carries X-Tokamak-Guardrail: none|flagged|masked|cloaked|blocked and X-Tokamak-Guardrail-Count. On a streamed answer the headers describe the request only, because they are sent before the answer starts.

Cloaked stand-ins are derived from a secret, your organization and a key version, so the same value always gets the same stand-in within a conversation. Nothing is stored and prompt caching keeps working.

Start in Flag. Run a new rule in flag for a few days, check Matches for false positives, then switch it to mask, cloak or block.

3. Test before you save

Test runs a saved policy against text or a request body in any of the three API formats. Nothing is sent to a provider, recorded or charged.

The Test tab: an agent prompt containing an email, a phone number, a .env with a database URL, AWS and GitHub keys, and a CSV row with an IBAN; the decision is forward with changes, 5 masked and 4 cloaked, with each match, its score, action and replacement
The sandbox shows each match, its score and what it becomes. Here secrets are masked and customer data is cloaked. Open full size ↗

4. Give a key its own policy

By default every key and every signed-in member uses the default policy. Under API keys, assign a stricter policy to a CI or release key.

The API keys tab: the ci-agent key uses the ci-strict policy; every other key uses acme-default
A key's own policy replaces the default. A key set to a disabled policy gets no screening; it never falls back. Open full size ↗

5. Review matches

Matches lists every match with its policy, rule, entity, action, stage, JSON path, key and member. Mark a match as a false positive to tune the rule, or Export CSV.

The Matches tab: 100 matches and 11 blocked requests in the last 7 days, top rule pii email, and a table of matches with time, policy, rule, action, stage, path, key and member
Matches show where something matched, not what. The matched text is kept only if you turn on Log matched content, and secrets never are. Open full size ↗

6. Settings

The Settings tab: if screening is unavailable let requests through or refuse them; log matched content off; keep matches 30 days; scan up to 4 MiB; cloak key ck_v1 with a rotate button
Organization settings for guardrails. Open full size ↗
SettingDefaultWhat it means
If screening is unavailableLet requests throughIf a policy cannot be loaded the request is forwarded unscreened and counted. Refuse answers 503 guardrail_unavailable instead, an outage rather than a content refusal.
Log matched contentOffKeep a short excerpt (≤ 256 characters) of each new match for review. Secrets are never kept.
Keep matches for30 daysMatch records are deleted after this.
Scan up to4 MiBNewest content is scanned first. A larger request is forwarded and marked partly scanned, or blocked if you choose.
Cloak key—Rotate it to change every stand-in. Expect one prompt-cache miss per active conversation.

Answers

When a policy has rules that apply to responses, or the request was cloaked, Tokamak also screens the answer:

  • A non-streaming answer is checked whole, then sent.
  • A streamed answer is checked as it streams. A short stretch of text is held back so a match split across chunks is caught before any of it leaves. A tool call is released when it is complete.
  • A block on the answer ends the stream the way the API already does (stop_reason: "refusal", finish_reason: "content_filter", or response.incomplete). Text already sent cannot be recalled.
  • An answer block is charged for the tokens the model produced. Only request blocks are free.

Limits

  • Detection is pattern-based. Names and free-text addresses without a label are not found.
  • Prompt-injection detection is a heuristic for flagging, not a guarantee.
  • Thinking and reasoning blocks and images are never scanned or changed.
  • Changes reach every server within 30 seconds.

On this page