Audio Translations
Translate foreign language spoken audio directly into fluent English text across Whisper models with sub-second failover and exact list-price settlement.
Last updated
The /v1/audio/translations endpoint translates spoken audio from over 90 source languages directly into English text using state-of-the-art ASR models like Whisper. By wrapping translation workloads within nRouter's unified gateway, engineering teams gain multi-provider failover, strict multi-tenant budget ceilings, and zero-markup provider pass-through pricing.
POST https://api.nrouter.ai/v1/audio/translationsDirect audio-to-English translation without requiring intermediate transcription steps.
Supports streaming uploads for MP3, WAV, M4A, FLAC, OGG, and WebM containers.
Billed at exact provider duration rates with zero markup or hidden platform fees.
Every response reports exact cost, processing latency, and routing path.
Architectural Role & Lifecycle
Translating spoken audio requires reliable stream processing and automatic fallback when upstream cloud providers experience transient degradation:
- Phase 1: In-Memory Key Auth & ACLs: Verifies virtual key hash (
sk-nrouter-...) and checks tenant permissions for the target audio translation model. - Phase 2: Tenant Rate Limits & Budget Check: Validates the organization's sliding-window RPM quotas and verifies balance availability.
- Phase 3: Stream Inspection & Prompt Guardrails: Validates the audio container structure and inspects any guiding prompt for prompt injections. If safety rules trigger, the request is halted with HTTP 400 (
x-nr-guardrails: blocked). - Phase 4: Duration Credit Reservation: Places an atomic credit hold based on estimated audio duration. If the request was halted in preflight, $0 is deducted or held.
- Provider Execution & Final Settlement: Sends the multipart stream to the upstream provider. Upon receiving the translated English transcript, nRouter calculates the exact duration in seconds, settles against the provider list price, and emits telemetry headers.
Request Parameters
The request body must be submitted as multipart/form-data:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
file | binary | Yes | — | The audio file to translate. Accepted containers: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm. Maximum size: 25 MB. |
model | string | Yes | — | Model ID for translation (e.g. whisper-1). |
prompt | string | No | — | An optional text prompt to guide the model's translation style or provide context for proper names, terminology, and formatting. |
response_format | string | No | json | The output format of the translation. Supported values: json, text, srt, verbose_json, or vtt. |
temperature | number | No | 0 | Sampling temperature between 0.0 and 1.0. Higher values lead to more varied translations. |
Response Payloads
Standard JSON (response_format: "json")
{
"text": "Hello, how are you today? We are testing multi-cloud voice routing."
}Verbose JSON (response_format: "verbose_json")
{
"task": "translate",
"language": "spanish",
"duration": 5.48,
"text": "Hello, how are you today? We are testing multi-cloud voice routing.",
"segments": [
{
"id": 0,
"seek": 0,
"start": 0.0,
"end": 5.48,
"text": " Hello, how are you today? We are testing multi-cloud voice routing.",
"tokens": [50364, 2425, 11, 703, 389, 345, 1909, 30, 775, 389, 4268, 5978, 12, 11624, 4678, 26978, 13, 50638],
"temperature": 0.0,
"avg_logprob": -0.198,
"compression_ratio": 1.08,
"no_speech_prob": 0.002
}
]
}Headers Reference
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be multipart/form-data; boundary=.... |
x-nr-tags | string | No | Metadata tags for FinOps attribution (e.g. service=localization,region=us). |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique correlation UUID assigned to this translation request. |
x-nr-latency-ms | integer | Edge turnaround time in milliseconds. |
x-nr-request-cost | float | Exact USD cost of the translation calculated from raw provider duration rates. |
x-nr-cost-status | string | exact when priced from provider rates or unpriced if rate metadata is pending. |
x-nr-model | string | Physical model that served the translation. |
x-nr-routing | string | Routing chain outcome: direct or fallback:<n>. |
x-nr-attempts | integer | Number of upstream provider attempts made. |
x-nr-guardrails | string | Preflight safety outcome: none, monitor, pass, or blocked. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
import * as fs from "node:fs";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const fileStream = fs.createReadStream("spanish_audio.mp3");
const translation = await client.audio.translations.create({
file: fileStream,
model: "whisper-1",
response_format: "verbose_json",
});
console.log("English Translation:", translation.text);
console.log(`Original Audio Duration: ${translation.duration}s`);Error Codes & Failure Handling
Errors return standard JSON error structures with descriptive error messages:
{
"error": {
"type": "gateway_error",
"message": "Unsupported file format. Please upload an MP3, MP4, MPEG, MPGA, M4A, WAV, or WEBM file.",
"code": "invalid_request"
}
}| HTTP Status | Error Code | Cause | Recommended Action |
|---|---|---|---|
| 400 Bad Request | invalid_request | Missing required file or model fields, or corrupted container headers. | Verify the audio file format and multipart form structure. |
| 400 Bad Request | guardrail_blocked | Optional prompt steering text triggered content safety or injection rules. | Sanitize prompt text; $0 is spent or held on blocked requests. |
| 401 Unauthorized | invalid_api_key | Virtual key missing, expired, or lacking translation scope. | Ensure virtual key starts with sk-nrouter- and has active status. |
| 402 Payment Required | insufficient_credits | Organization balance is insufficient to reserve duration credits. | Add funds via dashboard billing; minimum credit deposit is $5. |
| 404 Not Found | model_not_found | Target translation model not found or disabled in catalog. | Verify model availability via GET /v1/models. |
| 413 Payload Too Large | payload_too_large | File size exceeds the 25 MB gateway limit. | Compress or segment the audio file before issuing request. |
| 429 Too Many Requests | rate_limit_exceeded | Account RPM or TPM ceiling reached. | Inspect x-nr-limit-source and retry using exponential backoff. |
| 500 / 503 Provider Error | service_unavailable | Upstream translation provider failure or timeout. | Gateway automatically tries fallback providers if configured. |
Audio Translation Architecture & Billing Mechanics
The audio translation pipeline combines acoustic feature extraction, multi-lingual phoneme recognition, and neural machine translation into a single unified step.
Duration-Based Billing Precision
Unlike LLM completions which charge per token, audio translation is priced per second of parsed audio duration:
- Phase 4 Credit Reservation: The gateway inspects container headers (e.g. ID3 tags, RIFF chunks) during initial upload to estimate duration and place a credit hold.
- Exact Duration Settlement: Spend settles at exact list prices rounded to the nearest millisecond upon completion.
- Zero Cost on Rejections: If a file upload is rejected by edge WAF or guardrails (e.g. unsupported container or file size ceiling), $0 is reserved or charged.
Best Practices for Multilingual Ingestion
- Chunking Long Recordings: For recordings exceeding 25 MB or 15 minutes, split audio files into continuous segments at natural pauses using tools like FFmpeg or WebAudio.
- Steering Context with Prompts: Provide relevant terminology, company names, or industry acronyms in the optional
promptparameter to anchor the translation vocabulary. - Format Selection: Uploading compressed formats such as MP3 or Opus encoded M4A drastically cuts network transmission latency compared to raw PCM WAV files without sacrificing translation accuracy.
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
multipart/form-data
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/audio/translations" \ -F model="string" \ -F file="string"{ "text": "Hello, how are you?"}