Completions
Create raw text completions using legacy prompt models with streaming support, token usage tracking, automated failover, and multi-tenant budget controls.
Last updated
The /v1/completions endpoint provides raw text completion for base models, instruction-tuned foundation models, and legacy completion workloads across all major provider clouds. While modern conversational workloads predominantly use Chat Completions, the Completions API remains essential for code synthesis, text infilling, unstructured generation, and deterministic prompt completion.
POST https://api.nrouter.ai/v1/completionsDirect string or array-of-strings prompt evaluation without message turn overhead.
Tenant-isolated caching returns repeat completions instantly with zero provider spend.
Phase 3 inspection detects prompt injections before credit reservation.
Exact provider pass-through pricing per token with zero markup.
Architectural Role & Lifecycle
When a prompt is dispatched to /v1/completions, nRouter runs the request through its unified preflight engine:
- Phase 1: In-Memory Validation: Verifies the virtual key hash (
sk-nrouter-...), tenant organization membership, and model access policies. - Phase 2: Rate Limits & Token Ceilings: Evaluates requests-per-minute (RPM) and tokens-per-minute (TPM) sliding windows.
- Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input prompt against safety classifiers. If content violates policy, the request is halted with HTTP 400 (
x-nr-guardrails: blocked). - Phase 4: Credit Reservation: Calculates the upper-bound token cost (
input_tokens + max_tokens) and places an atomic hold on organization credits. Blocked requests never incur holds or spend. - Execution & Streaming Settlement: Dispatches the prompt to the selected provider or serves the response from the tenant cache (
x-nr-response-cache: hit). Streaming requests commit exact spend upon receiving the final chunk.
Request Parameters
The request body must be a JSON object:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | — | Model ID to generate completions (e.g. gpt-5.4-mini, gpt-3.5-turbo-instruct). |
prompt | string or string[] | Yes | — | The prompt or prompts to generate completions for. Can be a single string or an array of strings. |
max_tokens | integer | No | 16 | The maximum number of tokens to generate in the completion. |
temperature | number | No | 1.0 | Sampling temperature between 0.0 and 2.0. Higher values yield more creative outputs. |
top_p | number | No | 1.0 | Nucleus sampling parameter. Alternative to temperature sampling. |
n | integer | No | 1 | Number of completions to generate for each prompt. |
stream | boolean | No | false | When true, tokens are streamed via Server-Sent Events (SSE). |
logprobs | integer | No | null | Include the log probabilities on the logprobs most likely tokens, up to 5. |
echo | boolean | No | false | Echo back the prompt in addition to the completion. |
stop | string or string[] | No | null | Up to 4 sequences where the API will stop generating further tokens. |
presence_penalty | number | No | 0 | Number between -2.0 and 2.0 penalizing new tokens based on presence in text. |
frequency_penalty | number | No | 0 | Number between -2.0 and 2.0 penalizing new tokens based on existing frequency. |
nrouter_cache | boolean | No | true | When false, bypasses nRouter's response cache for this call. |
nrouter_prompt_template_id | string | No | — | Tenant-scoped managed prompt template UUID. |
nrouter_prompt_variables | object | No | — | Key-value variables interpolated into the managed prompt template. |
Request & Response Headers
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
x-nr-tags | string | No | Metadata tags for FinOps attribution (e.g. workload=batch,team=ml). |
x-nr-compress | string | No | Set to off to disable prompt compression for this request. |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique trace ID for the request. |
x-nr-latency-ms | integer | Gateway turnaround time in milliseconds (TTFB for streams). |
x-nr-request-cost | float | Exact USD cost of the completion calculated from provider token rates. |
x-nr-cost-status | string | exact when priced or unpriced if pending rate metadata. |
x-nr-model | string | Physical model that served the completion. |
x-nr-routing | string | Routing chain outcome: direct or fallback:<n>. |
x-nr-attempts | integer | Total provider attempts made. |
x-nr-guardrails | string | Guardrail evaluation outcome: none, monitor, pass, or blocked. |
x-nr-response-cache | string | Cache outcome: hit, miss, or bypass. |
x-nr-input-tokens | integer | Billed input token count. |
x-nr-output-tokens | integer | Billed output token count. |
x-nr-total-tokens | integer | Total tokens consumed. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const completion = await client.completions.create({
model: "gpt-5.4-mini",
prompt: "Write a high-performance Rust function to calculate SHA-256 hashes.",
max_tokens: 512,
temperature: 0.2,
});
console.log(completion.choices[0].text);
console.log(`Tokens: ${completion.usage.total_tokens}`);
console.log(`Cost: \$${client.lastResponse?.cost}`);Error Handling & Status Codes
{
"error": {
"type": "gateway_error",
"message": "The requested max_tokens exceeds model maximum context ceiling.",
"code": "max_output_tokens_too_large"
}
}| HTTP Status | Error Code | Root Cause | Remediation |
|---|---|---|---|
| 400 Bad Request | invalid_request | Malformed JSON, missing prompt, or invalid token bounds. | Verify request schema and ensure prompt is non-empty. |
| 400 Bad Request | guardrail_blocked | Preflight inspection detected prompt injection or policy violation. | Review prompt text against safety guidelines; $0 is charged. |
| 401 Unauthorized | invalid_api_key | Virtual key missing, expired, or invalid. | Verify sk-nrouter-... key in dashboard settings. |
| 402 Payment Required | insufficient_credits | Tenant credit balance insufficient for reservation hold. | Recharge balance in dashboard; minimum top-up is $5. |
| 404 Not Found | model_not_found | Model identifier not found or unauthorized for tenant. | Query GET /v1/models for active model identifiers. |
| 429 Too Many Requests | rate_limit_exceeded | RPM or TPM ceiling crossed for virtual key. | Check x-nr-limit-source and back off request rate. |
| 500 / 503 Provider Error | service_unavailable | Upstream provider outage or connection timeout. | Enable fallback chains via nrouter_fallbacks parameter. |
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/completions" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.4-mini", "prompt": "Write a haiku about cloud infrastructure.", "max_tokens": 64 }'{ "id": "cmpl-abc123", "object": "text_completion", "model": "gpt-4o-mini", "choices": [ { "text": " there!", "index": 0, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 2, "completion_tokens": 2, "total_tokens": 4 }}