Audio Transcriptions
Transcribe spoken audio into accurate text across Whisper models with automatic language detection, timestamping, and transparent duration-based billing.
Last updated
The /v1/audio/transcriptions endpoint provides high-accuracy automatic speech recognition (ASR/STT) by proxying requests to state-of-the-art transcription models like Whisper through nRouter's unified gateway. The endpoint accepts multipart file uploads, provides automatic provider failover, and guarantees strict per-second billing at raw provider list prices.
POST https://api.nrouter.ai/v1/audio/transcriptionsAccepts MP3, MP4, M4A, WAV, WEBM, FLAC, and OGG containers seamlessly.
Extract word-level and segment-level timestamps for captions and alignment.
Exact provider pass-through pricing calculated per second of processed audio.
Every response returns execution latency, request IDs, and exact USD spend.
Architectural Role & Lifecycle
Audio transcription workloads require streaming binary uploads and resilient execution for prolonged audio processing tasks. nRouter orchestrates transcription requests through a structured preflight pipeline:
- Phase 1: Key Authentication & Tenancy Verification: Validates the caller's virtual key (
sk-nrouter-...) against the tenant organization and verifies model entitlement for ASR engines. - Phase 2: Concurrency & Rate Limit Gating: Validates that the tenant organization has active RPM and concurrent job capacity before accepting the multipart payload stream.
- Phase 3: Payload Inspection & Safety Screening: Verifies audio container integrity, size ceilings (25 MB maximum per upload), and applies prompt injection checks to the optional prompt steering parameter.
- Phase 4: Duration Credit Reservation: Estimates audio duration from stream headers and places an atomic hold on organization credits. If the request is rejected during preflight, zero credits are deducted or held.
- Upstream Execution & Settlement: Dispatches the audio stream to the designated provider. If an upstream outage occurs, nRouter executes automatic retry or fallback routing. Upon completion, exact audio duration in seconds is settled against the provider's published list price.
Request Parameters
Requests must be sent as multipart/form-data:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
file | binary | Yes | — | The audio file object to transcribe. Supported container formats include flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, and webm. Maximum size is 25 MB. |
model | string | Yes | — | Model ID for transcription (e.g. whisper-1). |
language | string | No | — | The language of the input audio in ISO-639-1 format (e.g. en, es, de, fr, ja). Providing a hint improves transcription accuracy and reduces processing latency. |
prompt | string | No | — | An optional text prompt to guide the model's style or specify domain-specific jargon, brand names, or uncommon technical terminology. |
response_format | string | No | json | The format of the transcript output. Options: json, text, srt, verbose_json, or vtt. |
temperature | number | No | 0 | The sampling temperature between 0.0 and 1.0. Lower values produce more deterministic transcriptions. |
timestamp_granularities[] | array | No | ["segment"] | The timestamp granularities to populate for verbose_json. Options: word, segment, or both ["word", "segment"]. |
Response Formats & Payload Schemas
Standard JSON Response (response_format: "json")
{
"text": "Welcome to nRouter, the multi-cloud enterprise AI routing gateway."
}Verbose JSON Response (response_format: "verbose_json")
When detailed timing information is requested, verbose_json returns full metadata including segments, tokens, and word-level timestamps:
{
"task": "transcribe",
"language": "english",
"duration": 4.82,
"text": "Welcome to nRouter, the multi-cloud enterprise AI routing gateway.",
"segments": [
{
"id": 0,
"seek": 0,
"start": 0.0,
"end": 4.82,
"text": " Welcome to nRouter, the multi-cloud enterprise AI routing gateway.",
"tokens": [50364, 7062, 281, 412, 17820, 11, 264, 4833, 12, 11624, 7762, 9552, 26978, 13, 50605],
"temperature": 0.0,
"avg_logprob": -0.1824,
"compression_ratio": 1.12,
"no_speech_prob": 0.0014
}
],
"words": [
{ "word": "Welcome", "start": 0.0, "end": 0.52 },
{ "word": "to", "start": 0.54, "end": 0.72 },
{ "word": "nRouter", "start": 0.74, "end": 1.34 }
]
}Headers Reference
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be multipart/form-data; boundary=.... |
x-nr-tags | string | No | Billing tags for cost attribution (e.g. team=transcription,client=mobile). |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique correlation UUID assigned to the request. |
x-nr-latency-ms | integer | Gateway edge turnaround time in milliseconds. |
x-nr-request-cost | float | Exact USD cost of the transcription calculated from audio duration rates. |
x-nr-cost-status | string | exact when cost was calculated or unpriced if pending rate settlement. |
x-nr-model | string | Upstream physical model that performed the transcription. |
x-nr-routing | string | Routing chain outcome: direct or fallback:<n>. |
x-nr-attempts | integer | Number of provider attempts made to complete the request. |
x-nr-guardrails | string | Guardrail evaluation state: none, monitor, pass, or blocked. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
import * as fs from "node:fs";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const fileStream = fs.createReadStream("interview.mp3");
const transcription = await client.audio.transcriptions.create({
file: fileStream,
model: "whisper-1",
language: "en",
response_format: "verbose_json",
timestamp_granularities: ["word", "segment"],
});
console.log("Transcribed Text:", transcription.text);
console.log(`Audio Duration: ${transcription.duration}s`);Error Handling & Common Failure Modes
Every error response returns a standard JSON error object with a corresponding HTTP status code:
{
"error": {
"type": "gateway_error",
"message": "Maximum content size of 25MB exceeded.",
"code": "payload_too_large"
}
}| HTTP Status | Error Code | Trigger Condition | Recommended Resolution |
|---|---|---|---|
| 400 Bad Request | invalid_request | Missing file or model parameter, unsupported audio codec, or corrupt container headers. | Verify the audio file plays locally and is encoded in an accepted format (mp3, wav, m4a). |
| 400 Bad Request | guardrail_blocked | Prompt guidance text violated safety policies or triggered prompt-injection defenses. | Review input prompt text; blocked calls incur $0 cost. |
| 401 Unauthorized | invalid_api_key | Virtual key is invalid, revoked, or lacks access permissions for this route. | Inspect the Authorization: Bearer sk-nrouter-... key configuration. |
| 402 Payment Required | insufficient_credits | Organization balance is insufficient to reserve duration credits. | Top up balance in the billing dashboard; minimum credit purchase is $5. |
| 404 Not Found | model_not_found | Specified transcription model is unrecognized or unavailable. | Consult the model catalog via GET /v1/models for active transcription engines. |
| 413 Payload Too Large | payload_too_large | Audio file exceeds the 25 MB file size limit. | Compress the audio stream or split longer files into chunks prior to upload. |
| 429 Too Many Requests | rate_limit_exceeded | Account or key RPM ceiling exceeded. | Check x-nr-limit-source header and apply exponential backoff. |
| 500 / 503 Provider Error | service_unavailable | Upstream transcription provider failure or connection timeout. | nRouter automatically initiates fallback routing if alternative models are configured. |
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
multipart/form-data
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/audio/transcriptions" \ -F model="string" \ -F file="string"{ "text": "The transcribed speech."}