Browse documentation
API ReferenceMessages POST

Messages

Create conversational turns using the native Anthropic Messages API wire format across Claude and cloud providers with prompt caching and cost tracking.

Last updated

The /v1/messages endpoint provides direct compatibility with the native Anthropic Messages API wire format. It allows applications built for Anthropic Claude models to route seamlessly through nRouter without changing client SDKs or message payload structures. nRouter dynamically routes requests across direct Anthropic APIs, AWS Bedrock, and Google Vertex AI with automated failover, prompt caching telemetry, and zero-markup pricing.

POST https://api.nrouter.ai/v1/messages
Native Wire
Anthropic Spec

Direct support for content blocks, system prompts, thinking mode, and tool use.

Multi-Cloud Failover
Bedrock & Vertex

Automatic failover from Anthropic direct to AWS Bedrock and Google Vertex AI.

Prompt Caching
Up to 90% Savings

Full telemetry on cache-read and cache-write tokens passed directly from providers.

Pricing Model
Raw List Price

Opus usage billed at Opus rates; Sonnet usage billed at Sonnet rates. Zero markup.


Architectural Role & Lifecycle

When a client application submits a request to /v1/messages, nRouter executes a 4-phase preflight lifecycle before dispatching to upstream providers:

  1. Phase 1: In-Memory Key Auth & ACLs: Verifies virtual key hash (sk-nrouter-...), tenant organization state, and ensures entitlement for Claude models.
  2. Phase 2: Rate Limits & Ceilings: Validates requests-per-minute (RPM) and tokens-per-minute (TPM) sliding window quotas.
  3. Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input message payload and system prompt against safety classifiers. If prompt injection is detected, the request is halted with HTTP 400 (x-nr-guardrails: blocked).
  4. Phase 4: Credit Reservation: Estimates token volume from prompt length plus max_tokens and reserves the credit envelope. Blocked requests incur zero credit hold and zero spend.
  5. Multi-Cloud Dispatch & Stream Processing: Sends the request to the lowest-latency active provider cloud. If Anthropic direct returns 529 or 429, nRouter automatically shifts execution to Bedrock or Vertex AI. For streams, Server-Sent Events (SSE) format is preserved, and exact token spend is settled atomically at end-of-stream.

Request Parameters

The request body must be a JSON object complying with the Anthropic Messages specification:

ParameterTypeRequiredDefaultDescription
modelstringYes—The model that will complete your prompt (e.g. claude-sonnet-4-5-20250929, claude-opus-4-20250514, claude-3-5-haiku-20241022).
messagesarrayYes—Input messages. Each item is an object with role (user or assistant) and content (string or array of content blocks).
max_tokensintegerYes—The maximum number of tokens to generate before stopping.
systemstring or arrayNo—System prompt specifying high-level context and behavior. Can be a string or an array of text blocks supporting cache controls.
metadataobjectNo—An object describing metadata about the request (e.g. user_id).
stop_sequencesstring[]No—Custom text sequences that will cause the model to stop generating tokens.
streambooleanNofalseWhether to incrementally stream the response using Server-Sent Events.
temperaturenumberNo1.0Amount of randomness injected into the response (range 0.0 to 1.0).
top_pnumberNo—Nucleus sampling cutoff threshold.
top_kintegerNo—Only sample from the top K options for each subsequent token.
toolsarrayNo—Definitions of tools that the model may use.
tool_choiceobjectNo{"type": "auto"}How the model should use the provided tools (auto, any, tool).
thinkingobjectNo—Configuration for extended thinking mode (e.g. {"type": "enabled", "budget_tokens": 2048}).
nrouter_fallbacksstring[]No—1–4 fallback models to try if primary provider admission fails.
nrouter_guardrailsstring[]No—1–8 organization-defined guardrail policy IDs to enforce on this call.
nrouter_cachebooleanNotrueWhen false, bypasses nRouter's response cache for this call.

Response Payloads

Standard Messages Response

{
  "id": "msg_01XFDUDYJgAACzvnptvVoYEE",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "nRouter provides seamless multi-cloud routing across Claude models with automated failover."
    }
  ],
  "model": "claude-sonnet-4-5-20250929",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 24,
    "output_tokens": 19,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}

Headers Reference

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be application/json.
anthropic-versionstringNoCompatibility header: 2023-06-01.
x-nr-tagsstringNoBilling tags for cost attribution (e.g. service=agent,env=prod).

Outbound Response Headers

HeaderTypeDescription
x-nr-request-idstringUnique UUID correlation identifier for tracing.
x-nr-latency-msintegerGateway edge turnaround time in milliseconds (TTFB for streams).
x-nr-request-costfloatExact USD cost of the inference call calculated from provider rates.
x-nr-cost-statusstringexact when priced or unpriced if pending rate metadata.
x-nr-modelstringUpstream physical model that served the request.
x-nr-routingstringRouting chain outcome: direct or fallback:<n>.
x-nr-attemptsintegerProvider calls made for this request.
x-nr-guardrailsstringGuardrail evaluation outcome: none, monitor, pass, or blocked.
x-nr-cache-read-tokensintegerProvider prompt cache-read tokens; emitted when nonzero.
x-nr-cache-write-tokensintegerProvider prompt cache-write tokens; emitted when nonzero.
x-nr-input-tokensintegerBilled input tokens.
x-nr-output-tokensintegerBilled output tokens.
x-nr-total-tokensintegerTotal token count.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const response = await client.messages.create({
  model: "claude-sonnet-4-5-20250929",
  max_tokens: 1024,
  system: "You are an expert distributed systems engineer.",
  messages: [
    { role: "user", content: "Explain how multi-cloud failover ensures 99.99% gateway availability." }
  ],
});

console.log(response.content[0].text);
console.log(`Cost: \$${client.lastResponse?.cost}`);

Error Codes & Failure Modes

{
  "error": {
    "type": "gateway_error",
    "message": "max_tokens: field required",
    "code": "invalid_request"
  }
}
HTTP StatusError CodeCauseRecommended Action
400 Bad Requestinvalid_requestMissing required max_tokens or messages field, or invalid content block structure.Ensure all required fields conform to Anthropic Messages schema.
400 Bad Requestguardrail_blockedPreflight inspection detected prompt injection or safety policy violation.Review prompt content; blocked requests incur $0 hold and $0 spend.
401 Unauthorizedinvalid_api_keyVirtual key missing, expired, or invalid.Check Authorization: Bearer sk-nrouter-... key in dashboard.
402 Payment Requiredinsufficient_creditsOrganization balance is insufficient to reserve token hold.Add credit balance via dashboard; minimum top-up is $5.
404 Not Foundmodel_not_foundModel identifier not found or unauthorized for tenant.Query GET /v1/models for active Claude model identifiers.
429 Too Many Requestsrate_limit_exceededRPM or TPM ceiling crossed for virtual key.Check x-nr-limit-source header and apply exponential backoff.
500 / 503 Provider Errorservice_unavailableUpstream provider outage or connection timeout.nRouter automatically initiates multi-cloud fallback (e.g. Anthropic ↔ Bedrock).
POST
/v1/messages

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Header Parameters

x-nr-compress?string

Request prompt compression: on to compress eligible prompts, off to skip

Value in

  • "on"
  • "off"
x-nr-tags?string

Custom spend and attribution tags (comma-separated key=value pairs)

x-nr-mcp-server?string

Target MCP server ID when routing MCP tool calls or prompts through the gateway

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/messages" \  -H "Content-Type: application/json" \  -d '{    "model": "claude-sonnet-4-5-20250929",    "max_tokens": 1024,    "messages": [      {        "role": "user",        "content": "Explain multi-tenant credit reservation."      }    ]  }'
{  "id": "msg_01",  "type": "message",  "role": "assistant",  "content": [    {      "type": "text",      "text": "Hello!"    }  ],  "usage": {    "input_tokens": 9,    "output_tokens": 3  }}
Was this page helpful?