Responses
Create model responses using the OpenAI Responses API format with integrated tool calling, agentic multi-turn loops, streaming, and exact list pricing.
Last updated
The /v1/responses endpoint provides native support for the OpenAI Responses API wire format. Built for autonomous agent loops, multi-step tool execution, and complex reasoning pipelines, this endpoint unifies next-generation response structures behind nRouter's multi-cloud gateway with automatic failover, content guardrails, and zero-markup pricing.
POST https://api.nrouter.ai/v1/responsesNative support for input-output structures, tool executions, and multi-turn state.
Content inspection blocks prompt injections before credit hold. Zero tokens spent.
Exact provider pass-through token pricing with zero hidden surcharges.
Every response reports turnaround latency, x-nr-request-cost, and routing headers.
Architectural Role & Lifecycle
When handling requests on /v1/responses, nRouter enforces its robust 4-phase preflight lifecycle:
- Phase 1: In-Memory Key Auth & ACLs: Verifies virtual key hash (
sk-nrouter-...) and confirms tenant organization permissions for the selected model. - Phase 2: Sliding-Window Rate Limiting: Enforces tenant-level and key-level RPM and TPM ceilings.
- Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the
inputandinstructionsfields against safety classifiers. If prompt injection is detected, the request is halted with HTTP 400 (x-nr-guardrails: blocked). - Phase 4: Credit Reservation: Estimates token consumption based on input tokens and
max_output_tokensand places an atomic credit reservation hold. Blocked requests cost $0. - Execution & Settlement: Dispatches the call to the upstream provider cloud. If an outage occurs, automatic fallback routing runs. Tokens are counted and spend is settled atomically at end-of-response at the exact list price of the served model.
Request Parameters
The request body must be a JSON object:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | — | Model ID to generate responses (e.g. gpt-5.4-mini, gpt-4o). |
input | string or array | Yes | — | The input content to generate a response for. Can be a text string or an array of input items. |
instructions | string | No | — | System-level guidance or developer instructions directing the model's behavior and tone. |
tools | array | No | — | An array of tools the model may call (e.g. custom functions, code interpreter, web search). |
tool_choice | string or object | No | auto | Specifies whether and how tools should be invoked (auto, required, none, or a specific function). |
temperature | number | No | 1.0 | Sampling temperature between 0.0 and 2.0. |
max_output_tokens | integer | No | — | The maximum number of tokens to generate in the response. |
stream | boolean | No | false | When true, response deltas are streamed via Server-Sent Events (SSE). |
nrouter_fallbacks | string[] | No | — | 1–4 fallback models to try if the primary model admission fails. |
nrouter_guardrails | string[] | No | — | 1–8 organization-defined guardrail policy IDs to enforce on this call. |
nrouter_cache | boolean | No | true | When false, bypasses nRouter's response cache for this call. |
Response Payloads
Standard Response Structure
{
"id": "resp_01j78qwed928374",
"object": "response",
"status": "completed",
"model": "gpt-5.4-mini",
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "nRouter simplifies enterprise AI infrastructure by providing unified routing, governance, and billing."
}
]
}
],
"usage": {
"input_tokens": 16,
"output_tokens": 20,
"total_tokens": 36
}
}Headers Reference
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
x-nr-tags | string | No | Metadata tags for FinOps attribution (e.g. agent=researcher,env=prod). |
x-nr-compress | string | No | Set to off to bypass prompt compression for this request. |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique correlation UUID assigned to this request. |
x-nr-latency-ms | integer | Gateway edge turnaround time in milliseconds (TTFB for streams). |
x-nr-request-cost | float | Exact USD cost of the response calculated from provider token rates. |
x-nr-cost-status | string | exact when priced or unpriced if rate metadata is pending. |
x-nr-model | string | Upstream physical model that served the response. |
x-nr-routing | string | Routing outcome: direct or fallback:<n>. |
x-nr-attempts | integer | Provider calls made for this request. |
x-nr-guardrails | string | Guardrail evaluation outcome: none, monitor, pass, or blocked. |
x-nr-response-cache | string | Cache outcome: hit, miss, or bypass. |
x-nr-input-tokens | integer | Billed input token count. |
x-nr-output-tokens | integer | Billed output token count. |
x-nr-total-tokens | integer | Total token count. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const response = await client.responses.create({
model: "gpt-5.4-mini",
instructions: "You are an autonomous incident response assistant.",
input: "Summarize the root cause of the database connection pool timeout.",
});
console.log(response.output[0].content[0].text);
console.log(`Cost: \$${client.lastResponse?.cost}`);Error Codes & Failure Modes
{
"error": {
"type": "gateway_error",
"message": "input: field required",
"code": "invalid_request"
}
}| HTTP Status | Error Code | Cause | Recommended Action |
|---|---|---|---|
| 400 Bad Request | invalid_request | Missing required input or model fields, or invalid tool schemas. | Verify request body structure against Responses API specification. |
| 400 Bad Request | guardrail_blocked | Preflight inspection detected prompt injection or safety policy violation. | Review input content against safety policies; $0 is charged. |
| 401 Unauthorized | invalid_api_key | Virtual key missing, expired, or invalid. | Check Authorization: Bearer sk-nrouter-... key in dashboard. |
| 402 Payment Required | insufficient_credits | Organization balance is insufficient to reserve token hold. | Add credit balance via dashboard; minimum top-up is $5. |
| 404 Not Found | model_not_found | Model identifier not found or unauthorized for tenant. | Query GET /v1/models to confirm active provider catalog entitlements. |
| 429 Too Many Requests | rate_limit_exceeded | Virtual key or organization RPM/TPM ceiling exceeded. | Check x-nr-limit-source header and apply exponential backoff. |
| 500 / 503 Provider Error | service_unavailable | Upstream provider outage or connection timeout. | nRouter automatically initiates fallback routing if alternative models are configured. |
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/responses" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.4-mini", "input": "Explain how prompt injection defense operates before credit reservation." }'{ "id": "resp_01", "object": "response", "status": "completed", "output": [ { "type": "message", "role": "assistant", "content": [ { "type": "output_text", "text": "Hello!" } ] } ]}