Browse documentation
API ReferenceAudio Speech POST

Audio Speech

Synthesize natural speech audio from input text using OpenAI-compatible text-to-speech models with customizable voices, multiple codecs, and list pricing.

Last updated

The /v1/audio/speech endpoint provides low-latency, streaming text-to-speech (TTS) synthesis across top-tier speech synthesis providers through a unified OpenAI-compatible API wire. By routing through nRouter, your infrastructure gains multi-provider failover, pre-call content moderation, transparent character-level billing at raw provider list prices, and comprehensive edge telemetry.

POST https://api.nrouter.ai/v1/audio/speech
Audio Formats
6 Codecs Supported

Stream MP3, Opus, AAC, FLAC, WAV, and uncompressed PCM directly to client applications.

Security Shield
$0 Injection Spend

Preflight content inspection inspects text before credit hold. Blocked synthesis costs $0.

Transparent Pricing
Raw List Price

Exact provider pass-through pricing per 1,000 characters with zero hidden surcharges.

Edge Telemetry
Per-Request Tracking

Every response includes x-nr-request-cost, edge latency, and correlation IDs.


Architectural Role & Lifecycle

When a client application submits text to /v1/audio/speech, nRouter executes a four-phase preflight lifecycle before emitting audio bytes:

  1. Phase 1: In-Memory Authentication & ACLs: Verifies the virtual key hash (sk-nrouter-...), tenant organization status, and access permissions for the requested speech synthesis model.
  2. Phase 2: Sliding-Window Rate Limiting: Enforces tenant-level and key-level requests-per-minute (RPM) and characters-per-minute ceilings without external cache bottlenecks.
  3. Phase 3: Content Moderation & Prompt Injection Scoring: Evaluates the input text against safety policies. If flagged content or prompt injections are detected, the request is halted with HTTP 400 (x-nr-guardrails: blocked).
  4. Phase 4: Character-Level Credit Reservation: Estimates the character volume and reserves credits against the tenant organization balance. If the preflight checks fail at Phase 3, zero credits are deducted or held.
  5. Egress & Audio Streaming: Forwards the request to the upstream voice provider. As binary audio frames stream back, nRouter flushes the byte stream directly to the client while calculating exact cost settlement from actual characters rendered.

Request Parameters

The request body must be a JSON object containing the required synthesis parameters:

ParameterTypeRequiredDefaultDescription
modelstringYes—Model ID for speech synthesis (e.g. tts-1, tts-1-hd).
inputstringYes—The text to generate audio for. Maximum length is 4,096 characters per request. Text exceeding this limit is rejected with HTTP 400.
voicestringYes—The voice persona to use for synthesis. Supported options include alloy, echo, fable, onyx, nova, and shimmer.
response_formatstringNomp3The output audio format. Supported formats: mp3, opus, aac, flac, wav, and pcm.
speednumberNo1.0The playback speed multiplier for synthesized speech. Accepts values from 0.25 (slowest) to 4.0 (fastest).

Request & Response Headers

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be application/json.
x-nr-tagsstringNoOptional billing tags for cost attribution (e.g. team=support,env=prod).

Outbound Response Headers

HeaderTypeDescription
Content-TypestringThe MIME type corresponding to the requested audio format (e.g. audio/mpeg, audio/opus, audio/aac).
x-nr-request-idstringUnique UUID correlation identifier stamped on every request for distributed tracing and support auditing.
x-nr-latency-msintegerGateway edge turnaround time in milliseconds (Time-To-First-Byte for streaming audio).
x-nr-request-costfloatExact USD cost of the audio generation calculated from upstream character rates.
x-nr-cost-statusstringSet to exact when priced from provider rates or unpriced if rate metadata is pending.
x-nr-modelstringThe physical upstream model that served the synthesis.
x-nr-routingstringRouting chain outcome: direct for primary provider or fallback:<n> if failover occurred.
x-nr-attemptsintegerNumber of upstream provider attempts made (retries and failovers).
x-nr-guardrailsstringPreflight safety outcome: none, monitor, pass, or blocked.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";
import * as fs from "node:fs";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const response = await client.audio.speech.create({
  model: "tts-1",
  voice: "alloy",
  input: "Welcome to nRouter, the multi-cloud enterprise AI routing gateway.",
  response_format: "mp3",
});

const buffer = Buffer.from(await response.arrayBuffer());
await fs.promises.writeFile("welcome.mp3", buffer);
console.log("Audio saved successfully to welcome.mp3");

Audio Codec & Format Specifications

nRouter supports six distinct audio encodings to accommodate varied client latency, fidelity, and bandwidth requirements:

FormatContent-TypeTypical BitrateBest Used For
mp3audio/mpeg128 kbpsUniversal client playback, consumer mobile apps, web browsers.
opusaudio/opus24–64 kbpsUltra-low latency streaming, VoIP applications, WebRTC voice pipelines.
aacaudio/aac128 kbpsHigh-efficiency audio compression for iOS applications and Apple ecosystems.
flacaudio/flacLosslessStudio-grade archival audio, professional production, audiophile post-processing.
wavaudio/wavUncompressedLow-latency audio decoders requiring instant playback without container decompression.
pcmaudio/pcm24kHz 16-bitRaw audio samples in host-endian format for digital signal processing (DSP).

Error Handling & Status Codes

Every error response emitted by nRouter returns a structured JSON payload accompanied by machine-readable error headers:

{
  "error": {
    "type": "gateway_error",
    "message": "Input text exceeds maximum allowed length of 4096 characters.",
    "code": "input_too_large"
  }
}
HTTP StatusError TypeTrigger ConditionsMitigation Strategy
400 Bad Requestinvalid_requestMissing required parameters (model, input, voice), invalid speed value, or text exceeding 4096 characters.Validate payload against schema constraints before issuing the request.
400 Bad Requestguardrail_blockedInput text violated tenant content moderation or prompt injection thresholds during Phase 3.Review input content against safety policies; no credits are held or deducted.
401 Unauthorizedinvalid_api_keyVirtual key is missing, malformed, expired, or rejected by tenancy network policies.Verify the Authorization: Bearer sk-nrouter-... header in the organization dashboard.
402 Payment Requiredinsufficient_creditsOrganization prepaid credit balance is exhausted or lower than the character reservation hold.Add credit balance via the dashboard billing interface or enable auto-recharge.
404 Not Foundmodel_not_foundRequested speech model is unrecognized or not enabled for the tenant organization.Check available models via GET /v1/models to confirm active provider catalog entitlements.
429 Too Many Requestsrate_limit_exceededRequests-per-minute (RPM) limit has been exceeded on the virtual key or tenant account.Implement exponential backoff retry algorithms; check x-nr-limit-source header.
500 / 503 Gateway Errorservice_unavailableUpstream voice synthesis provider is experiencing an outage or connection timeout.nRouter automatically initiates fallback routing when configured; retry with backoff.
POST
/v1/audio/speech

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Header Parameters

x-nr-compress?string

Request prompt compression: on to compress eligible prompts, off to skip

Value in

  • "on"
  • "off"
x-nr-tags?string

Custom spend and attribution tags (comma-separated key=value pairs)

x-nr-mcp-server?string

Target MCP server ID when routing MCP tool calls or prompts through the gateway

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/audio/speech" \  -H "Content-Type: application/json" \  -d '{    "model": "string",    "input": "string",    "voice": "string"  }'
"audio/mpeg bytes (rendered as an <audio> player)"
Was this page helpful?