Browse documentation
API ReferenceCompletions POST

Completions

Create raw text completions using legacy prompt models with streaming support, token usage tracking, automated failover, and multi-tenant budget controls.

Last updated

The /v1/completions endpoint provides raw text completion for base models, instruction-tuned foundation models, and legacy completion workloads across all major provider clouds. While modern conversational workloads predominantly use Chat Completions, the Completions API remains essential for code synthesis, text infilling, unstructured generation, and deterministic prompt completion.

POST https://api.nrouter.ai/v1/completions
Raw Prompting
Unstructured Input

Direct string or array-of-strings prompt evaluation without message turn overhead.

Response Cache
Sub-Millisecond Hits

Tenant-isolated caching returns repeat completions instantly with zero provider spend.

Security Shield
$0 Injection Spend

Phase 3 inspection detects prompt injections before credit reservation.

Pricing Model
Flat List Price

Exact provider pass-through pricing per token with zero markup.


Architectural Role & Lifecycle

When a prompt is dispatched to /v1/completions, nRouter runs the request through its unified preflight engine:

  1. Phase 1: In-Memory Validation: Verifies the virtual key hash (sk-nrouter-...), tenant organization membership, and model access policies.
  2. Phase 2: Rate Limits & Token Ceilings: Evaluates requests-per-minute (RPM) and tokens-per-minute (TPM) sliding windows.
  3. Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input prompt against safety classifiers. If content violates policy, the request is halted with HTTP 400 (x-nr-guardrails: blocked).
  4. Phase 4: Credit Reservation: Calculates the upper-bound token cost (input_tokens + max_tokens) and places an atomic hold on organization credits. Blocked requests never incur holds or spend.
  5. Execution & Streaming Settlement: Dispatches the prompt to the selected provider or serves the response from the tenant cache (x-nr-response-cache: hit). Streaming requests commit exact spend upon receiving the final chunk.

Request Parameters

The request body must be a JSON object:

ParameterTypeRequiredDefaultDescription
modelstringYes—Model ID to generate completions (e.g. gpt-5.4-mini, gpt-3.5-turbo-instruct).
promptstring or string[]Yes—The prompt or prompts to generate completions for. Can be a single string or an array of strings.
max_tokensintegerNo16The maximum number of tokens to generate in the completion.
temperaturenumberNo1.0Sampling temperature between 0.0 and 2.0. Higher values yield more creative outputs.
top_pnumberNo1.0Nucleus sampling parameter. Alternative to temperature sampling.
nintegerNo1Number of completions to generate for each prompt.
streambooleanNofalseWhen true, tokens are streamed via Server-Sent Events (SSE).
logprobsintegerNonullInclude the log probabilities on the logprobs most likely tokens, up to 5.
echobooleanNofalseEcho back the prompt in addition to the completion.
stopstring or string[]NonullUp to 4 sequences where the API will stop generating further tokens.
presence_penaltynumberNo0Number between -2.0 and 2.0 penalizing new tokens based on presence in text.
frequency_penaltynumberNo0Number between -2.0 and 2.0 penalizing new tokens based on existing frequency.
nrouter_cachebooleanNotrueWhen false, bypasses nRouter's response cache for this call.
nrouter_prompt_template_idstringNo—Tenant-scoped managed prompt template UUID.
nrouter_prompt_variablesobjectNo—Key-value variables interpolated into the managed prompt template.

Request & Response Headers

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be application/json.
x-nr-tagsstringNoMetadata tags for FinOps attribution (e.g. workload=batch,team=ml).
x-nr-compressstringNoSet to off to disable prompt compression for this request.

Outbound Response Headers

HeaderTypeDescription
x-nr-request-idstringUnique trace ID for the request.
x-nr-latency-msintegerGateway turnaround time in milliseconds (TTFB for streams).
x-nr-request-costfloatExact USD cost of the completion calculated from provider token rates.
x-nr-cost-statusstringexact when priced or unpriced if pending rate metadata.
x-nr-modelstringPhysical model that served the completion.
x-nr-routingstringRouting chain outcome: direct or fallback:<n>.
x-nr-attemptsintegerTotal provider attempts made.
x-nr-guardrailsstringGuardrail evaluation outcome: none, monitor, pass, or blocked.
x-nr-response-cachestringCache outcome: hit, miss, or bypass.
x-nr-input-tokensintegerBilled input token count.
x-nr-output-tokensintegerBilled output token count.
x-nr-total-tokensintegerTotal tokens consumed.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const completion = await client.completions.create({
  model: "gpt-5.4-mini",
  prompt: "Write a high-performance Rust function to calculate SHA-256 hashes.",
  max_tokens: 512,
  temperature: 0.2,
});

console.log(completion.choices[0].text);
console.log(`Tokens: ${completion.usage.total_tokens}`);
console.log(`Cost: \$${client.lastResponse?.cost}`);

Error Handling & Status Codes

{
  "error": {
    "type": "gateway_error",
    "message": "The requested max_tokens exceeds model maximum context ceiling.",
    "code": "max_output_tokens_too_large"
  }
}
HTTP StatusError CodeRoot CauseRemediation
400 Bad Requestinvalid_requestMalformed JSON, missing prompt, or invalid token bounds.Verify request schema and ensure prompt is non-empty.
400 Bad Requestguardrail_blockedPreflight inspection detected prompt injection or policy violation.Review prompt text against safety guidelines; $0 is charged.
401 Unauthorizedinvalid_api_keyVirtual key missing, expired, or invalid.Verify sk-nrouter-... key in dashboard settings.
402 Payment Requiredinsufficient_creditsTenant credit balance insufficient for reservation hold.Recharge balance in dashboard; minimum top-up is $5.
404 Not Foundmodel_not_foundModel identifier not found or unauthorized for tenant.Query GET /v1/models for active model identifiers.
429 Too Many Requestsrate_limit_exceededRPM or TPM ceiling crossed for virtual key.Check x-nr-limit-source and back off request rate.
500 / 503 Provider Errorservice_unavailableUpstream provider outage or connection timeout.Enable fallback chains via nrouter_fallbacks parameter.
POST
/v1/completions

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Header Parameters

x-nr-compress?string

Request prompt compression: on to compress eligible prompts, off to skip

Value in

  • "on"
  • "off"
x-nr-tags?string

Custom spend and attribution tags (comma-separated key=value pairs)

x-nr-mcp-server?string

Target MCP server ID when routing MCP tool calls or prompts through the gateway

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/completions" \  -H "Content-Type: application/json" \  -d '{    "model": "gpt-5.4-mini",    "prompt": "Write a haiku about cloud infrastructure.",    "max_tokens": 64  }'
{  "id": "cmpl-abc123",  "object": "text_completion",  "model": "gpt-4o-mini",  "choices": [    {      "text": " there!",      "index": 0,      "finish_reason": "stop"    }  ],  "usage": {    "prompt_tokens": 2,    "completion_tokens": 2,    "total_tokens": 4  }}
Was this page helpful?