Chat Completions
Create streaming and non-streaming chat completions using nRouter. OpenAI-compatible request format with automated fallbacks, guardrails, and cost headers.
Last updated
The /v1/chat/completions endpoint is nRouter's primary text wire. It serves standard OpenAI- and Anthropic-compatible chat requests across six provider clouds—OpenAI, Anthropic, AWS Bedrock, Google Vertex AI, Azure Foundry, and Alibaba US—behind a single unified API key and a single bill.
POST https://api.nrouter.ai/v1/chat/completionsDynamic latency and cost routing with automatic fallback chains across clouds.
Content inspection blocks injection before the credit hold. Zero tokens spent.
All models billed at raw provider list prices. No hidden surcharge.
Every response headers includes exact USD spend and edge-measured latency.
SDK Code Examples
Every request requires your virtual key prefixed with sk-nrouter-. You can use our official branded SDKs, native community packages, or point any standard OpenAI client library at https://api.nrouter.ai/v1.
import { nRouter } from "@nrouter_ai/sdk";
// Reads NROUTER_API_KEY from environment
const client = new nRouter();
// 1. Standard Chat Completion
const completion = await client.chat.completions.create({
model: "gpt-5.4-mini",
messages: [
{ role: "system", content: "You are a senior systems architect." },
{ role: "user", content: "Explain how prompt-injection defense operates before credit reservation." }
],
max_completion_tokens: 1024,
temperature: 0.7,
});
console.log(completion.choices[0].message.content);
console.log(`Cost: $${client.lastResponse?.cost ?? "unpriced"}`);
console.log(`Latency: ${client.lastResponse?.latencyMs}ms`);
console.log(`Request ID: ${client.lastResponse?.requestId}`);
// 2. Real-Time Streaming
const stream = await client.chat.completions.create({
model: "claude-sonnet-4-5-20250929",
messages: [{ role: "user", content: "List 3 advantages of an LLM gateway." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Streaming & Server-Sent Events (SSE)
When "stream": true is passed, the gateway immediately establishes an HTTP chunked transfer connection and streams Server-Sent Events (SSE) formatted as data: {...} payloads.
data: {"id":"chatcmpl-98a","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-98a","choices":[{"index":0,"delta":{"content":" world"},"finish_reason":null}]}
data: {"id":"chatcmpl-98a","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]Stream Telemetry & Cost Accounting
- Latency Header (
x-nr-latency-ms): On a stream, this header measures Time-To-First-Byte (TTFB) from edge arrival until the initial response headers are flushed—not the total generation time of the stream. - Settlement: Spend is computed and committed atomically at end-of-stream (EOS) from the tokens actually emitted, so a streamed request is billed on real usage rather than an estimate.
Router Aliases vs. Pinned Models
nRouter decouples client applications from rigid vendor model strings. What you place in the "model" field determines the routing strategy:
| Parameter Value | Resolution Type | Behavior |
|---|---|---|
fastest-claude | Router Alias | Evaluates real-time cloud latency, routes to the lowest-latency healthy endpoint, and falls back if degraded. |
nrouter/auto | Smart Router | Evaluates prompt complexity, context size, and price ceilings to select the optimal model automatically. |
gpt-5.4-mini | Concrete Model | Bypasses router fallback chains and pins the request directly to the specified model. |
Request Options
Three optional JSON body fields change routing, guardrails, and caching for one call, without touching your organization's configuration. They are nRouter control fields: the Rust gateway reads them and removes them from the body before forwarding, so no provider ever receives a key beginning with nrouter_. Any other nrouter_* key is refused with 400 rather than passed on.
| Field | Type | Limit | Behavior |
|---|---|---|---|
nrouter_fallbacks | string[] | 1–4 entries | Models tried in order if the primary is refused admission. Replaces your organization's fallback list for this call; the primary stays whatever is in model. Every target must be routable by your key or the call is refused 400 fallback_not_allowed. Failover only on 429, 503, Anthropic 529, or a connect-phase failure. Still one credit reservation, one rate-limit slot, and at most 3 provider calls. |
nrouter_guardrails | string[] | 1–8 entries | Guardrail IDs or names your organization owns, added to this call's pre-call chain. Add-only: it never removes a key, team, or organization rule and never weakens the platform safety floor. An unknown or foreign entry is refused 400 guardrail_not_found with a byte-identical body. |
nrouter_cache | boolean | — | false skips the response cache in both directions — not answered from it, not written into it — and the response carries x-nr-response-cache: bypass. |
Live
All three fields, and the x-nr-routing and x-nr-attempts response headers below, are live on api.nrouter.ai as of the 2026-09-18 gateway release.
curl -X POST https://api.nrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $NROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.4-mini",
"messages": [
{ "role": "user", "content": "Summarize this incident report in three bullets." }
],
"nrouter_fallbacks": ["claude-sonnet-4-5", "gemini-2.5-pro"],
"nrouter_guardrails": ["pii-strict"],
"nrouter_cache": false
}'The same three fields work from every nRouter SDK; SDKs that model the request body strictly expose them through an extra_body-style escape hatch.
Refusal codes
Every refusal carries a machine-readable code alongside its message, so a client can branch on it rather than matching prose:
code | HTTP | When |
|---|---|---|
fallback_not_allowed | 400 | A nrouter_fallbacks target your key cannot route to, or a fallback list on a request shape that does not accept one. |
guardrail_not_found | 400 | A nrouter_guardrails entry your organization does not own or does not have enabled. |
input_too_large | 400 | The prompt exceeds the model's declared input ceiling. |
max_output_tokens_too_large | 400 | The requested output length exceeds the model's declared maximum. |
guardrail_blocked | 400 | A guardrail in the resolved chain refused the request before it reached a provider. |
service_unavailable | 503 | A required dependency was unavailable. When it carries x-nr-guardrails: unavailable, an enforcing guardrail could not run and the request was refused without being judged — retry it rather than rewriting the prompt. |
A request refused at any of these points never reaches a provider and costs nothing.
Full narrative reference: Per-Request Options.
Headers Reference
Request Headers (Inbound)
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
x-nr-compress | string | No | on requests prompt compression, off skips it for this request. Compression is already on by default for qualified models, so this header is how a single request opts out. |
x-nr-tags | string | No | Comma-separated key-value metadata tags for cost attribution (e.g. env=prod,tenant=acme). |
x-nr-mcp-server | string | No | MCP routing target when issuing requests to model context servers. |
Response Headers (Outbound)
Every response contains accurate telemetry stamped at the gateway edge:
| Header | Type | Emitted | Description |
|---|---|---|---|
x-nr-request-id | string | Always | Unique correlation identifier for tracing and support lookup. |
x-nr-latency-ms | integer | Always | Gateway edge turnaround time in milliseconds (TTFB for streams). |
x-nr-request-cost | float | When priced | Exact USD cost of the inference call calculated from provider token rates. Absent — never 0 — when the request could not be priced. |
x-nr-cost-status | string | Every served response | exact when the call was priced, or unpriced when it could not be. It is what tells an absent cost apart from a free one — no enabled model is free. |
x-nr-model | string | Success | The exact physical model that served the request (auditing alias resolution). |
x-nr-routing | string | Provider answered | direct when the first chain entry answered, or fallback:<n> where n is the 0-based chain index of the entry that answered — so the first fallback is fallback:1, the second model in the chain. Absent on cache hits and refusals. |
x-nr-attempts | integer | Provider answered | Provider calls this request made, counting retries and failovers alike (≥ 1). Absent on cache hits and refusals. |
x-nr-guardrails | string | Guarded | Guardrail evaluation posture: none, monitor, pass, redacted (an enforcing rule rewrote part of the prompt before it was sent), partial (some content was not inspected), blocked, or unavailable. |
x-nr-limit-source | string | 429 / 402 | Which ceiling refused: key, plan, team, user, budget, plan_window_h8, plan_window_day, plan_window_week, capacity, plan_allowance_exhausted, or plan_required. |
x-nr-auth-reason | string | 401 | Machine-readable reason a virtual key was refused: unauthorized, key_blocked, key_expired, key_route_not_allowed, key_model_not_allowed, key_ip_not_allowed, key_network_policy_invalid, or auth_backend_unavailable. |
x-nr-budget-warning | string | Warning | Emitted when a soft budget threshold has been crossed (e.g. org soft_budget 80.00/100.00). |
x-nr-response-cache | string | Cached | hit (served from cache), miss (called provider), or bypass (opted out). |
x-nr-response-cache-age | integer | Cache hit | Age of the cache entry in seconds, so a two-second-old replay is distinguishable from a much older one. |
x-nr-compression | string | Compression | applied, not_requested, off, or skipped. |
Prompt Compression
Prompt compression is on by default. On a qualified model, a long prompt is shortened before the model reads it, which means the prompt the model reads is rewritten and the answer can differ from the one your original prompt would have produced. You are billed for what the provider bills — the shorter prompt.
It only runs on models nRouter has qualified for compression, and only on parts of the prompt that are long enough to be worth rewriting; every other request is forwarded unchanged. Your latest turn and any system or developer instructions are never compressed.
Every response tells you what happened on x-nr-compression:
| Value | Meaning |
|---|---|
applied | The prompt was compressed and the shorter prompt was sent. |
skipped | Compression was entitled but did not run or did not help — the original prompt was sent. |
not_requested | Your organization is set to opt-in and this request did not ask. |
off | Compression is switched off for this organization, key, or request. |
Three ways to opt out, from widest to narrowest:
- Per organization — set prompt compression to
Offin Settings → Privacy. An Owner or Organization Admin can change it. - Per request — send
x-nr-compress: off. - Per message or block — add
"nrouter_compress": falseto that message. This is a per-message field, not a request-level one; to switch a whole request off, use the header.
A virtual key can also be set never to compress, whatever its organization says.
Guardrails & Prompt Injection Defense
Every inference request passes through nRouter's 4-phase preflight gate chain before touching an upstream provider:
- Phase 1: In-Memory Validation: Validates virtual key hash, model ACLs, and organization status.
- Phase 2: Rate Limits & Ceilings: Verifies tenant RPM and TPM sliding window ceilings.
- Phase 3: Content Inspection: Scores the prompt for prompt injection and for the seven moderation categories, concurrently. If a score crosses its threshold, the gateway halts with HTTP 400 (
x-nr-guardrails: blocked). - Phase 4: Credit Reservation: Only runs after Phase 3 passes. $0 is spent and $0 is held on blocked injection attacks.
{
"error": {
"type": "guardrail_blocked",
"message": "Request blocked by pre-call prompt injection safety policy."
}
}API Reference
Complete nRouter API reference: explore OpenAI and Anthropic compatible endpoints for chat completions, embeddings, audio, and intelligent model routing.
Completions POST
Create raw text completions using legacy prompt models with streaming support, token usage tracking, automated failover, and multi-tenant budget controls.