Audio Speech
Synthesize natural speech audio from input text using OpenAI-compatible text-to-speech models with customizable voices, multiple codecs, and list pricing.
Last updated
The /v1/audio/speech endpoint provides low-latency, streaming text-to-speech (TTS) synthesis across top-tier speech synthesis providers through a unified OpenAI-compatible API wire. By routing through nRouter, your infrastructure gains multi-provider failover, pre-call content moderation, transparent character-level billing at raw provider list prices, and comprehensive edge telemetry.
POST https://api.nrouter.ai/v1/audio/speechStream MP3, Opus, AAC, FLAC, WAV, and uncompressed PCM directly to client applications.
Preflight content inspection inspects text before credit hold. Blocked synthesis costs $0.
Exact provider pass-through pricing per 1,000 characters with zero hidden surcharges.
Every response includes x-nr-request-cost, edge latency, and correlation IDs.
Architectural Role & Lifecycle
When a client application submits text to /v1/audio/speech, nRouter executes a four-phase preflight lifecycle before emitting audio bytes:
- Phase 1: In-Memory Authentication & ACLs: Verifies the virtual key hash (
sk-nrouter-...), tenant organization status, and access permissions for the requested speech synthesis model. - Phase 2: Sliding-Window Rate Limiting: Enforces tenant-level and key-level requests-per-minute (RPM) and characters-per-minute ceilings without external cache bottlenecks.
- Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input text against safety policies. If flagged content or prompt injections are detected, the request is halted with HTTP 400 (
x-nr-guardrails: blocked). - Phase 4: Character-Level Credit Reservation: Estimates the character volume and reserves credits against the tenant organization balance. If the preflight checks fail at Phase 3, zero credits are deducted or held.
- Egress & Audio Streaming: Forwards the request to the upstream voice provider. As binary audio frames stream back, nRouter flushes the byte stream directly to the client while calculating exact cost settlement from actual characters rendered.
Request Parameters
The request body must be a JSON object containing the required synthesis parameters:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | — | Model ID for speech synthesis (e.g. tts-1, tts-1-hd). |
input | string | Yes | — | The text to generate audio for. Maximum length is 4,096 characters per request. Text exceeding this limit is rejected with HTTP 400. |
voice | string | Yes | — | The voice persona to use for synthesis. Supported options include alloy, echo, fable, onyx, nova, and shimmer. |
response_format | string | No | mp3 | The output audio format. Supported formats: mp3, opus, aac, flac, wav, and pcm. |
speed | number | No | 1.0 | The playback speed multiplier for synthesized speech. Accepts values from 0.25 (slowest) to 4.0 (fastest). |
Request & Response Headers
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
x-nr-tags | string | No | Optional billing tags for cost attribution (e.g. team=support,env=prod). |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
Content-Type | string | The MIME type corresponding to the requested audio format (e.g. audio/mpeg, audio/opus, audio/aac). |
x-nr-request-id | string | Unique UUID correlation identifier stamped on every request for distributed tracing and support auditing. |
x-nr-latency-ms | integer | Gateway edge turnaround time in milliseconds (Time-To-First-Byte for streaming audio). |
x-nr-request-cost | float | Exact USD cost of the audio generation calculated from upstream character rates. |
x-nr-cost-status | string | Set to exact when priced from provider rates or unpriced if rate metadata is pending. |
x-nr-model | string | The physical upstream model that served the synthesis. |
x-nr-routing | string | Routing chain outcome: direct for primary provider or fallback:<n> if failover occurred. |
x-nr-attempts | integer | Number of upstream provider attempts made (retries and failovers). |
x-nr-guardrails | string | Preflight safety outcome: none, monitor, pass, or blocked. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
import * as fs from "node:fs";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const response = await client.audio.speech.create({
model: "tts-1",
voice: "alloy",
input: "Welcome to nRouter, the multi-cloud enterprise AI routing gateway.",
response_format: "mp3",
});
const buffer = Buffer.from(await response.arrayBuffer());
await fs.promises.writeFile("welcome.mp3", buffer);
console.log("Audio saved successfully to welcome.mp3");Audio Codec & Format Specifications
nRouter supports six distinct audio encodings to accommodate varied client latency, fidelity, and bandwidth requirements:
| Format | Content-Type | Typical Bitrate | Best Used For |
|---|---|---|---|
mp3 | audio/mpeg | 128 kbps | Universal client playback, consumer mobile apps, web browsers. |
opus | audio/opus | 24–64 kbps | Ultra-low latency streaming, VoIP applications, WebRTC voice pipelines. |
aac | audio/aac | 128 kbps | High-efficiency audio compression for iOS applications and Apple ecosystems. |
flac | audio/flac | Lossless | Studio-grade archival audio, professional production, audiophile post-processing. |
wav | audio/wav | Uncompressed | Low-latency audio decoders requiring instant playback without container decompression. |
pcm | audio/pcm | 24kHz 16-bit | Raw audio samples in host-endian format for digital signal processing (DSP). |
Error Handling & Status Codes
Every error response emitted by nRouter returns a structured JSON payload accompanied by machine-readable error headers:
{
"error": {
"type": "gateway_error",
"message": "Input text exceeds maximum allowed length of 4096 characters.",
"code": "input_too_large"
}
}| HTTP Status | Error Type | Trigger Conditions | Mitigation Strategy |
|---|---|---|---|
| 400 Bad Request | invalid_request | Missing required parameters (model, input, voice), invalid speed value, or text exceeding 4096 characters. | Validate payload against schema constraints before issuing the request. |
| 400 Bad Request | guardrail_blocked | Input text violated tenant content moderation or prompt injection thresholds during Phase 3. | Review input content against safety policies; no credits are held or deducted. |
| 401 Unauthorized | invalid_api_key | Virtual key is missing, malformed, expired, or rejected by tenancy network policies. | Verify the Authorization: Bearer sk-nrouter-... header in the organization dashboard. |
| 402 Payment Required | insufficient_credits | Organization prepaid credit balance is exhausted or lower than the character reservation hold. | Add credit balance via the dashboard billing interface or enable auto-recharge. |
| 404 Not Found | model_not_found | Requested speech model is unrecognized or not enabled for the tenant organization. | Check available models via GET /v1/models to confirm active provider catalog entitlements. |
| 429 Too Many Requests | rate_limit_exceeded | Requests-per-minute (RPM) limit has been exceeded on the virtual key or tenant account. | Implement exponential backoff retry algorithms; check x-nr-limit-source header. |
| 500 / 503 Gateway Error | service_unavailable | Upstream voice synthesis provider is experiencing an outage or connection timeout. | nRouter automatically initiates fallback routing when configured; retry with backoff. |
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/audio/speech" \ -H "Content-Type: application/json" \ -d '{ "model": "string", "input": "string", "voice": "string" }'"audio/mpeg bytes (rendered as an <audio> player)"