Messages
Create conversational turns using the native Anthropic Messages API wire format across Claude and cloud providers with prompt caching and cost tracking.
Last updated
The /v1/messages endpoint provides direct compatibility with the native Anthropic Messages API wire format. It allows applications built for Anthropic Claude models to route seamlessly through nRouter without changing client SDKs or message payload structures. nRouter dynamically routes requests across direct Anthropic APIs, AWS Bedrock, and Google Vertex AI with automated failover, prompt caching telemetry, and zero-markup pricing.
POST https://api.nrouter.ai/v1/messagesDirect support for content blocks, system prompts, thinking mode, and tool use.
Automatic failover from Anthropic direct to AWS Bedrock and Google Vertex AI.
Full telemetry on cache-read and cache-write tokens passed directly from providers.
Opus usage billed at Opus rates; Sonnet usage billed at Sonnet rates. Zero markup.
Architectural Role & Lifecycle
When a client application submits a request to /v1/messages, nRouter executes a 4-phase preflight lifecycle before dispatching to upstream providers:
- Phase 1: In-Memory Key Auth & ACLs: Verifies virtual key hash (
sk-nrouter-...), tenant organization state, and ensures entitlement for Claude models. - Phase 2: Rate Limits & Ceilings: Validates requests-per-minute (RPM) and tokens-per-minute (TPM) sliding window quotas.
- Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input message payload and system prompt against safety classifiers. If prompt injection is detected, the request is halted with HTTP 400 (
x-nr-guardrails: blocked). - Phase 4: Credit Reservation: Estimates token volume from prompt length plus
max_tokensand reserves the credit envelope. Blocked requests incur zero credit hold and zero spend. - Multi-Cloud Dispatch & Stream Processing: Sends the request to the lowest-latency active provider cloud. If Anthropic direct returns 529 or 429, nRouter automatically shifts execution to Bedrock or Vertex AI. For streams, Server-Sent Events (SSE) format is preserved, and exact token spend is settled atomically at end-of-stream.
Request Parameters
The request body must be a JSON object complying with the Anthropic Messages specification:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | — | The model that will complete your prompt (e.g. claude-sonnet-4-5-20250929, claude-opus-4-20250514, claude-3-5-haiku-20241022). |
messages | array | Yes | — | Input messages. Each item is an object with role (user or assistant) and content (string or array of content blocks). |
max_tokens | integer | Yes | — | The maximum number of tokens to generate before stopping. |
system | string or array | No | — | System prompt specifying high-level context and behavior. Can be a string or an array of text blocks supporting cache controls. |
metadata | object | No | — | An object describing metadata about the request (e.g. user_id). |
stop_sequences | string[] | No | — | Custom text sequences that will cause the model to stop generating tokens. |
stream | boolean | No | false | Whether to incrementally stream the response using Server-Sent Events. |
temperature | number | No | 1.0 | Amount of randomness injected into the response (range 0.0 to 1.0). |
top_p | number | No | — | Nucleus sampling cutoff threshold. |
top_k | integer | No | — | Only sample from the top K options for each subsequent token. |
tools | array | No | — | Definitions of tools that the model may use. |
tool_choice | object | No | {"type": "auto"} | How the model should use the provided tools (auto, any, tool). |
thinking | object | No | — | Configuration for extended thinking mode (e.g. {"type": "enabled", "budget_tokens": 2048}). |
nrouter_fallbacks | string[] | No | — | 1–4 fallback models to try if primary provider admission fails. |
nrouter_guardrails | string[] | No | — | 1–8 organization-defined guardrail policy IDs to enforce on this call. |
nrouter_cache | boolean | No | true | When false, bypasses nRouter's response cache for this call. |
Response Payloads
Standard Messages Response
{
"id": "msg_01XFDUDYJgAACzvnptvVoYEE",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "nRouter provides seamless multi-cloud routing across Claude models with automated failover."
}
],
"model": "claude-sonnet-4-5-20250929",
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 24,
"output_tokens": 19,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}Headers Reference
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
anthropic-version | string | No | Compatibility header: 2023-06-01. |
x-nr-tags | string | No | Billing tags for cost attribution (e.g. service=agent,env=prod). |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique UUID correlation identifier for tracing. |
x-nr-latency-ms | integer | Gateway edge turnaround time in milliseconds (TTFB for streams). |
x-nr-request-cost | float | Exact USD cost of the inference call calculated from provider rates. |
x-nr-cost-status | string | exact when priced or unpriced if pending rate metadata. |
x-nr-model | string | Upstream physical model that served the request. |
x-nr-routing | string | Routing chain outcome: direct or fallback:<n>. |
x-nr-attempts | integer | Provider calls made for this request. |
x-nr-guardrails | string | Guardrail evaluation outcome: none, monitor, pass, or blocked. |
x-nr-cache-read-tokens | integer | Provider prompt cache-read tokens; emitted when nonzero. |
x-nr-cache-write-tokens | integer | Provider prompt cache-write tokens; emitted when nonzero. |
x-nr-input-tokens | integer | Billed input tokens. |
x-nr-output-tokens | integer | Billed output tokens. |
x-nr-total-tokens | integer | Total token count. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const response = await client.messages.create({
model: "claude-sonnet-4-5-20250929",
max_tokens: 1024,
system: "You are an expert distributed systems engineer.",
messages: [
{ role: "user", content: "Explain how multi-cloud failover ensures 99.99% gateway availability." }
],
});
console.log(response.content[0].text);
console.log(`Cost: \$${client.lastResponse?.cost}`);Error Codes & Failure Modes
{
"error": {
"type": "gateway_error",
"message": "max_tokens: field required",
"code": "invalid_request"
}
}| HTTP Status | Error Code | Cause | Recommended Action |
|---|---|---|---|
| 400 Bad Request | invalid_request | Missing required max_tokens or messages field, or invalid content block structure. | Ensure all required fields conform to Anthropic Messages schema. |
| 400 Bad Request | guardrail_blocked | Preflight inspection detected prompt injection or safety policy violation. | Review prompt content; blocked requests incur $0 hold and $0 spend. |
| 401 Unauthorized | invalid_api_key | Virtual key missing, expired, or invalid. | Check Authorization: Bearer sk-nrouter-... key in dashboard. |
| 402 Payment Required | insufficient_credits | Organization balance is insufficient to reserve token hold. | Add credit balance via dashboard; minimum top-up is $5. |
| 404 Not Found | model_not_found | Model identifier not found or unauthorized for tenant. | Query GET /v1/models for active Claude model identifiers. |
| 429 Too Many Requests | rate_limit_exceeded | RPM or TPM ceiling crossed for virtual key. | Check x-nr-limit-source header and apply exponential backoff. |
| 500 / 503 Provider Error | service_unavailable | Upstream provider outage or connection timeout. | nRouter automatically initiates multi-cloud fallback (e.g. Anthropic ↔ Bedrock). |
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/messages" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-4-5-20250929", "max_tokens": 1024, "messages": [ { "role": "user", "content": "Explain multi-tenant credit reservation." } ] }'{ "id": "msg_01", "type": "message", "role": "assistant", "content": [ { "type": "text", "text": "Hello!" } ], "usage": { "input_tokens": 9, "output_tokens": 3 }}