Browse documentation
API ReferenceResponses POST

Responses

Create model responses using the OpenAI Responses API format with integrated tool calling, agentic multi-turn loops, streaming, and exact list pricing.

Last updated

The /v1/responses endpoint provides native support for the OpenAI Responses API wire format. Built for autonomous agent loops, multi-step tool execution, and complex reasoning pipelines, this endpoint unifies next-generation response structures behind nRouter's multi-cloud gateway with automatic failover, content guardrails, and zero-markup pricing.

POST https://api.nrouter.ai/v1/responses
Responses API
Agent-First Wire

Native support for input-output structures, tool executions, and multi-turn state.

Security Shield
$0 Injection Spend

Content inspection blocks prompt injections before credit hold. Zero tokens spent.

Transparent Pricing
Raw List Price

Exact provider pass-through token pricing with zero hidden surcharges.

Telemetry
Unified Tracing

Every response reports turnaround latency, x-nr-request-cost, and routing headers.


Architectural Role & Lifecycle

When handling requests on /v1/responses, nRouter enforces its robust 4-phase preflight lifecycle:

  1. Phase 1: In-Memory Key Auth & ACLs: Verifies virtual key hash (sk-nrouter-...) and confirms tenant organization permissions for the selected model.
  2. Phase 2: Sliding-Window Rate Limiting: Enforces tenant-level and key-level RPM and TPM ceilings.
  3. Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input and instructions fields against safety classifiers. If prompt injection is detected, the request is halted with HTTP 400 (x-nr-guardrails: blocked).
  4. Phase 4: Credit Reservation: Estimates token consumption based on input tokens and max_output_tokens and places an atomic credit reservation hold. Blocked requests cost $0.
  5. Execution & Settlement: Dispatches the call to the upstream provider cloud. If an outage occurs, automatic fallback routing runs. Tokens are counted and spend is settled atomically at end-of-response at the exact list price of the served model.

Request Parameters

The request body must be a JSON object:

ParameterTypeRequiredDefaultDescription
modelstringYes—Model ID to generate responses (e.g. gpt-5.4-mini, gpt-4o).
inputstring or arrayYes—The input content to generate a response for. Can be a text string or an array of input items.
instructionsstringNo—System-level guidance or developer instructions directing the model's behavior and tone.
toolsarrayNo—An array of tools the model may call (e.g. custom functions, code interpreter, web search).
tool_choicestring or objectNoautoSpecifies whether and how tools should be invoked (auto, required, none, or a specific function).
temperaturenumberNo1.0Sampling temperature between 0.0 and 2.0.
max_output_tokensintegerNo—The maximum number of tokens to generate in the response.
streambooleanNofalseWhen true, response deltas are streamed via Server-Sent Events (SSE).
nrouter_fallbacksstring[]No—1–4 fallback models to try if the primary model admission fails.
nrouter_guardrailsstring[]No—1–8 organization-defined guardrail policy IDs to enforce on this call.
nrouter_cachebooleanNotrueWhen false, bypasses nRouter's response cache for this call.

Response Payloads

Standard Response Structure

{
  "id": "resp_01j78qwed928374",
  "object": "response",
  "status": "completed",
  "model": "gpt-5.4-mini",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "nRouter simplifies enterprise AI infrastructure by providing unified routing, governance, and billing."
        }
      ]
    }
  ],
  "usage": {
    "input_tokens": 16,
    "output_tokens": 20,
    "total_tokens": 36
  }
}

Headers Reference

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be application/json.
x-nr-tagsstringNoMetadata tags for FinOps attribution (e.g. agent=researcher,env=prod).
x-nr-compressstringNoSet to off to bypass prompt compression for this request.

Outbound Response Headers

HeaderTypeDescription
x-nr-request-idstringUnique correlation UUID assigned to this request.
x-nr-latency-msintegerGateway edge turnaround time in milliseconds (TTFB for streams).
x-nr-request-costfloatExact USD cost of the response calculated from provider token rates.
x-nr-cost-statusstringexact when priced or unpriced if rate metadata is pending.
x-nr-modelstringUpstream physical model that served the response.
x-nr-routingstringRouting outcome: direct or fallback:<n>.
x-nr-attemptsintegerProvider calls made for this request.
x-nr-guardrailsstringGuardrail evaluation outcome: none, monitor, pass, or blocked.
x-nr-response-cachestringCache outcome: hit, miss, or bypass.
x-nr-input-tokensintegerBilled input token count.
x-nr-output-tokensintegerBilled output token count.
x-nr-total-tokensintegerTotal token count.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const response = await client.responses.create({
  model: "gpt-5.4-mini",
  instructions: "You are an autonomous incident response assistant.",
  input: "Summarize the root cause of the database connection pool timeout.",
});

console.log(response.output[0].content[0].text);
console.log(`Cost: \$${client.lastResponse?.cost}`);

Error Codes & Failure Modes

{
  "error": {
    "type": "gateway_error",
    "message": "input: field required",
    "code": "invalid_request"
  }
}
HTTP StatusError CodeCauseRecommended Action
400 Bad Requestinvalid_requestMissing required input or model fields, or invalid tool schemas.Verify request body structure against Responses API specification.
400 Bad Requestguardrail_blockedPreflight inspection detected prompt injection or safety policy violation.Review input content against safety policies; $0 is charged.
401 Unauthorizedinvalid_api_keyVirtual key missing, expired, or invalid.Check Authorization: Bearer sk-nrouter-... key in dashboard.
402 Payment Requiredinsufficient_creditsOrganization balance is insufficient to reserve token hold.Add credit balance via dashboard; minimum top-up is $5.
404 Not Foundmodel_not_foundModel identifier not found or unauthorized for tenant.Query GET /v1/models to confirm active provider catalog entitlements.
429 Too Many Requestsrate_limit_exceededVirtual key or organization RPM/TPM ceiling exceeded.Check x-nr-limit-source header and apply exponential backoff.
500 / 503 Provider Errorservice_unavailableUpstream provider outage or connection timeout.nRouter automatically initiates fallback routing if alternative models are configured.
POST
/v1/responses

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Header Parameters

x-nr-compress?string

Request prompt compression: on to compress eligible prompts, off to skip

Value in

  • "on"
  • "off"
x-nr-tags?string

Custom spend and attribution tags (comma-separated key=value pairs)

x-nr-mcp-server?string

Target MCP server ID when routing MCP tool calls or prompts through the gateway

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/responses" \  -H "Content-Type: application/json" \  -d '{    "model": "gpt-5.4-mini",    "input": "Explain how prompt injection defense operates before credit reservation."  }'
{  "id": "resp_01",  "object": "response",  "status": "completed",  "output": [    {      "type": "message",      "role": "assistant",      "content": [        {          "type": "output_text",          "text": "Hello!"        }      ]    }  ]}
Was this page helpful?