Browse documentation

Guardrails

Deploy pre-call and post-call guardrails in nRouter to prevent prompt injections, enforce content safety policies, and redact sensitive PII data in real time.

Last updated

Guardrails inspect requests before they reach a model (pre-call) and responses before they reach your user (post-call). They can block, redact, warn, or simply log — letting you enforce safety and compliance without changing your application code.

Guardrail types

TypeWhat it does
Presidio PIIDetects and anonymizes personally identifiable information using Microsoft Presidio
RegexFilters content matching patterns you define
KeywordBlocks content containing words on a blocklist
Prompt InjectionDetects and blocks attempts to hijack the model's instructions
CustomCalls your own webhook to make the decision

Modes and actions

Each guardrail runs in a mode — pre-call (before the LLM request) or post-call (on the response) — and takes an action when it triggers:

ActionEffect
BlockReject the request before it reaches the provider — it costs zero credits and never appears in the provider's logs
RedactStrip the sensitive content and continue
WarnAllow the request through, but attach a warning header
LogAllow through and record the event only

Scope hierarchy

Guardrails apply at three scopes, combined per request — key > team > org:

  • Organization — a master kill-switch plus org-wide guardrails that apply to every key.
  • Team — guardrails scoped to a team, applying to every key on that team.
  • Key — guardrails assigned to specific virtual keys.

Scopes resolve by specificity, not union: for each guardrail, the assignment at the narrowest scope that mentions it decides, and that row alone. A key-level assignment overrides its team's and org's rather than adding to them. Guardrails that no assignment mentions are untouched, so different guardrails on the same request can be decided at different scopes.

In practice today every assignment the dashboard creates is an enabled one, and removing an assignment deletes the row rather than disabling it — so a narrower scope currently tightens or re-points protection and does not remove it. The distinction matters if that ever changes: because the narrowest row decides alone, a disabled assignment at the winning scope would switch that guardrail off for the key even though the broader org row still exists. It would not be overruled by the wider rule.

Manage org guardrails on the main Guardrails page and per-key assignments under Guardrails → Keys.

Adding a guardrail to a single request

The scopes above are your standing configuration. A single call can add guardrails on top of whatever that configuration resolves to, by naming them in the JSON request body:

{
  "model": "gpt-5.4-mini",
  "messages": [{ "role": "user", "content": "Draft a reply to this customer." }],
  "nrouter_guardrails": ["pii-strict", "3f2a8c14-9b7e-4d21-8f66-1a2b3c4d5e6f"]
}

Live. nrouter_guardrails is available on api.nrouter.ai as of the 2026-09-18 gateway release.

Each entry is a guardrail ID or name your organization owns. The field is deliberately one-directional:

  • Add-only, with no exceptions. A request can add protection. It can never remove or weaken a guardrail your key, team, or organization configured, and it can never remove or downgrade the platform safety floor that runs on every request whatever your settings say. There is no per-request way to switch a rule off — turning one off is an organization-level change made on this page.
  • Added, not substituted. Scope resolution (key > team > org) runs exactly as described above and decides your configured rules; the requested ones are appended to the result, and the platform floor is appended after that.
  • Your organization's guardrails only, and pre-call ones in this version. A guardrail belonging to another organization is not addressable from your key.
  • At most eight entries.
  • An unknown entry is a refusal, not a silent skip: 400 guardrail_not_found. A guardrail that does not exist anywhere and one that belongs to another organization return byte-identical bodies, so the refusal cannot be used to probe what exists outside your organization.
  • If your organization has guardrails switched off, a request cannot switch them back on for itself — the same guardrail_not_found refusal applies. A call may add to a plane you left on; it may not re-open one you turned off. The platform floor still runs regardless.

An added guardrail that blocks behaves like any other block: 400 with x-nr-guardrails: blocked, no provider call, zero credits, and a row in Guardrails → Logs naming the rule that fired. The added rules also take part in the response-cache key, so a reply produced under an added guardrail is never replayed to a call that did not ask for it.

The full per-request field reference, alongside the routing and caching fields, is in Per-Request Options.

Reading the posture header

Every response from a wire that runs a pre-call guardrail chain carries x-nr-guardrails, a single token saying what the chain did. It is the one place a client can see a control run without opening the dashboard:

x-nr-guardrailsWhat it tells you
noneNo guardrail applied to this request. A legitimate state if you configured nothing — and the alarm if you did.
monitorA chain ran, but nothing in it could have refused: every rule was in Warn or Log mode. It observed; it did not protect.
passAn enforcing chain inspected the whole request and allowed it.
redactedAn enforcing rule rewrote part of the prompt before it was sent, and the request then served normally. The model did not see what you sent, so an answer that quotes your prompt back is quoting the rewritten one.
partialAn enforcing chain ran but some content was not inspected. It never means something was acted on — that is redacted. It is the ordinary answer on the two audio upload wires, not an anomaly there.
blockedAn enforcing chain refused the request. 400, no provider call, zero credits.
unavailableAn enforcing chain could not run, so the request was refused without being judged (503). Retry it; do not rewrite the prompt, because nothing objected to its content.

Two things to get right when you wire this up:

  • Match the value exactly, and case-sensitively. An unrecognized token is unknown, never a guess at the nearest one.
  • Absence is not none. A response with no x-nr-guardrails header at all is making no guardrail claim — it never reached the pre-call chain, as with an authentication refusal or a /v1/models call. The explicit claim that nothing applied is the token none.

The header reports posture only, by design. It never carries a policy name, a policy ID, a rule count, or — for partial — which part went uninspected. A rule count moves when a policy moves, so a caller watching it could map your controls without ever tripping one, and naming the uninspected part would hand an evader the route around it. Which rule fired is in Guardrails → Logs, where it is yours to read and nobody else's.

Testing before you ship

Expand any guardrail to find a Test tab: paste sample input and see exactly which action fires and how long it took. This lets you tune a rule against realistic content before it touches live traffic.

Versioning and rollback

Every change to a guardrail is snapshotted. The Versions tab lists each update, and you can roll back to any prior version with one click — useful if a tightened rule starts blocking legitimate requests.

Templates

The Templates gallery offers one-click setups for common patterns — PII redaction, jailbreak/prompt-injection detection, and language blocklists — so you can stand up sensible defaults quickly, then customize.

Guardrail logs

The Guardrails → Logs view records every guardrail evaluation:

ColumnMeaning
TimeWhen the guardrail ran
GuardrailWhich rule evaluated the request
ModePre-call or post-call
ActionBlocked, Redacted, Allowed, Logged, or Error
LatencyHow long the check took (ms)
Request IDCorrelate with the request in observability logs

Filter by guardrail or action, or search by name/request ID.

Next steps

FAQ

Do I need to change my application code to use guardrails?

No. Guardrails run in the request path automatically once configured — they inspect requests before they reach a model and responses before they reach your user, with no SDK or code change on your side.

What kinds of content can a guardrail check for?

There are five types: Presidio PII detection/anonymization, Regex pattern filtering, Keyword blocklists, Prompt Injection detection, and Custom — which calls your own webhook to make the decision. You can combine multiple guardrails across the same request.

What's the difference between the Block, Redact, Warn, and Log actions?

Block rejects the request before it reaches the provider, Redact strips the sensitive content and lets the request continue, Warn allows it through but attaches a warning header, and Log allows it through and only records the event. Each guardrail runs in either pre-call (before the model) or post-call (on the response) mode.

If a guardrail blocks a request, am I still charged credits for it?

No. A blocked request is rejected before it reaches the provider, so it costs zero credits and never appears in the provider's logs. Redact, Warn, and Log all let the request proceed, so those calls are billed normally.

How do organization, team, and key guardrails combine on a single request?

The narrowest scope that mentions a guardrail decides it, in the order key > team > org > org default. A key assignment overrides team, team overrides org, and a disable at the winning scope turns that guardrail off — it is not re-enabled by a broader assignment still existing. Guardrails no assignment mentions are unaffected, so different guardrails on one request can be decided at different scopes.

Can a single request add a guardrail without changing my configuration?

Yes. Put nrouter_guardrails in the JSON request body with up to eight guardrail IDs or names your organization owns, and they run on that call in addition to whatever your key, team, and organization scopes resolve to. Nothing about your saved configuration changes, and the next call without the field is unaffected.

Can a request turn a guardrail off for one call?

No, and that is the design. nrouter_guardrails is add-only: it can add a guardrail your organization owns and can never remove one your key, team, or organization configured, nor remove or weaken the platform safety floor. Switching a rule off is an organization-level change made on the Guardrails page, by an owner or admin.

What happens if I name a guardrail that doesn't exist?

The request is refused with 400 guardrail_not_found. A guardrail that exists nowhere and one belonging to another organization return byte-identical bodies, so the error reveals nothing about what exists outside your organization. The same refusal applies if your organization has guardrails switched off — a request may add to a plane you left on, but cannot re-open one you turned off.

How do I tell what the guardrails did on a particular request?

Read the x-nr-guardrails response header. It carries one of seven tokens: none (nothing applied), monitor (a chain ran but nothing in it could refuse), pass (an enforcing chain inspected everything and allowed it), redacted (an enforcing rule rewrote part of the prompt before it was sent), partial (an enforcing chain ran but some content went uninspected), blocked (refused), and unavailable (the chain could not run, so the request was refused without being judged — retry it). Match the value exactly and case-sensitively. No header at all is not the same as none: it means the response makes no guardrail claim.

Who on my team can create or edit guardrails?

Creating, editing, deleting, and rolling back guardrails requires the owner or admin role. Any org member can view guardrails and their version history, and any member can use the Test tab to try a rule against sample input.

Can I test a guardrail before it affects live traffic?

Yes. Expand any guardrail and open the Test tab, paste sample input, and you'll see exactly which action fires and how long the check took — so you can tune a rule against realistic content before it touches live traffic.

What happens if I tighten a rule and it starts blocking legitimate requests?

Every change to a guardrail is snapshotted. Open the Versions tab, and you can roll back to any prior version with one click to immediately restore the earlier behavior.

Is there a fast way to stand up common guardrails?

Yes. The Templates gallery offers one-click setups for common patterns — PII redaction, jailbreak/prompt-injection detection, and language blocklists — that you can apply as sensible defaults and then customize.

How do I see what my guardrails did, and can I trace it back to a specific request?

The Guardrails → Logs view records every evaluation with its time, which rule ran, the mode, the action taken (Blocked, Redacted, Allowed, Logged, or Error), and the latency in milliseconds. You can filter by guardrail or action, or use the Request ID to correlate an evaluation with the full request in your observability logs.

Was this page helpful?