Browse documentation

Audio Transcriptions

Transcribe spoken audio into accurate text across Whisper models with automatic language detection, timestamping, and transparent duration-based billing.

Last updated

The /v1/audio/transcriptions endpoint provides high-accuracy automatic speech recognition (ASR/STT) by proxying requests to state-of-the-art transcription models like Whisper through nRouter's unified gateway. The endpoint accepts multipart file uploads, provides automatic provider failover, and guarantees strict per-second billing at raw provider list prices.

POST https://api.nrouter.ai/v1/audio/transcriptions
Input Formats
Up to 25 MB

Accepts MP3, MP4, M4A, WAV, WEBM, FLAC, and OGG containers seamlessly.

Timestamps
Word & Segment

Extract word-level and segment-level timestamps for captions and alignment.

Pricing Model
Zero Markup

Exact provider pass-through pricing calculated per second of processed audio.

Edge Observability
Full Telemetry

Every response returns execution latency, request IDs, and exact USD spend.


Architectural Role & Lifecycle

Audio transcription workloads require streaming binary uploads and resilient execution for prolonged audio processing tasks. nRouter orchestrates transcription requests through a structured preflight pipeline:

  1. Phase 1: Key Authentication & Tenancy Verification: Validates the caller's virtual key (sk-nrouter-...) against the tenant organization and verifies model entitlement for ASR engines.
  2. Phase 2: Concurrency & Rate Limit Gating: Validates that the tenant organization has active RPM and concurrent job capacity before accepting the multipart payload stream.
  3. Phase 3: Payload Inspection & Safety Screening: Verifies audio container integrity, size ceilings (25 MB maximum per upload), and applies prompt injection checks to the optional prompt steering parameter.
  4. Phase 4: Duration Credit Reservation: Estimates audio duration from stream headers and places an atomic hold on organization credits. If the request is rejected during preflight, zero credits are deducted or held.
  5. Upstream Execution & Settlement: Dispatches the audio stream to the designated provider. If an upstream outage occurs, nRouter executes automatic retry or fallback routing. Upon completion, exact audio duration in seconds is settled against the provider's published list price.

Request Parameters

Requests must be sent as multipart/form-data:

ParameterTypeRequiredDefaultDescription
filebinaryYes—The audio file object to transcribe. Supported container formats include flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, and webm. Maximum size is 25 MB.
modelstringYes—Model ID for transcription (e.g. whisper-1).
languagestringNo—The language of the input audio in ISO-639-1 format (e.g. en, es, de, fr, ja). Providing a hint improves transcription accuracy and reduces processing latency.
promptstringNo—An optional text prompt to guide the model's style or specify domain-specific jargon, brand names, or uncommon technical terminology.
response_formatstringNojsonThe format of the transcript output. Options: json, text, srt, verbose_json, or vtt.
temperaturenumberNo0The sampling temperature between 0.0 and 1.0. Lower values produce more deterministic transcriptions.
timestamp_granularities[]arrayNo["segment"]The timestamp granularities to populate for verbose_json. Options: word, segment, or both ["word", "segment"].

Response Formats & Payload Schemas

Standard JSON Response (response_format: "json")

{
  "text": "Welcome to nRouter, the multi-cloud enterprise AI routing gateway."
}

Verbose JSON Response (response_format: "verbose_json")

When detailed timing information is requested, verbose_json returns full metadata including segments, tokens, and word-level timestamps:

{
  "task": "transcribe",
  "language": "english",
  "duration": 4.82,
  "text": "Welcome to nRouter, the multi-cloud enterprise AI routing gateway.",
  "segments": [
    {
      "id": 0,
      "seek": 0,
      "start": 0.0,
      "end": 4.82,
      "text": " Welcome to nRouter, the multi-cloud enterprise AI routing gateway.",
      "tokens": [50364, 7062, 281, 412, 17820, 11, 264, 4833, 12, 11624, 7762, 9552, 26978, 13, 50605],
      "temperature": 0.0,
      "avg_logprob": -0.1824,
      "compression_ratio": 1.12,
      "no_speech_prob": 0.0014
    }
  ],
  "words": [
    { "word": "Welcome", "start": 0.0, "end": 0.52 },
    { "word": "to", "start": 0.54, "end": 0.72 },
    { "word": "nRouter", "start": 0.74, "end": 1.34 }
  ]
}

Headers Reference

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be multipart/form-data; boundary=....
x-nr-tagsstringNoBilling tags for cost attribution (e.g. team=transcription,client=mobile).

Outbound Response Headers

HeaderTypeDescription
x-nr-request-idstringUnique correlation UUID assigned to the request.
x-nr-latency-msintegerGateway edge turnaround time in milliseconds.
x-nr-request-costfloatExact USD cost of the transcription calculated from audio duration rates.
x-nr-cost-statusstringexact when cost was calculated or unpriced if pending rate settlement.
x-nr-modelstringUpstream physical model that performed the transcription.
x-nr-routingstringRouting chain outcome: direct or fallback:<n>.
x-nr-attemptsintegerNumber of provider attempts made to complete the request.
x-nr-guardrailsstringGuardrail evaluation state: none, monitor, pass, or blocked.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";
import * as fs from "node:fs";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const fileStream = fs.createReadStream("interview.mp3");

const transcription = await client.audio.transcriptions.create({
  file: fileStream,
  model: "whisper-1",
  language: "en",
  response_format: "verbose_json",
  timestamp_granularities: ["word", "segment"],
});

console.log("Transcribed Text:", transcription.text);
console.log(`Audio Duration: ${transcription.duration}s`);

Error Handling & Common Failure Modes

Every error response returns a standard JSON error object with a corresponding HTTP status code:

{
  "error": {
    "type": "gateway_error",
    "message": "Maximum content size of 25MB exceeded.",
    "code": "payload_too_large"
  }
}
HTTP StatusError CodeTrigger ConditionRecommended Resolution
400 Bad Requestinvalid_requestMissing file or model parameter, unsupported audio codec, or corrupt container headers.Verify the audio file plays locally and is encoded in an accepted format (mp3, wav, m4a).
400 Bad Requestguardrail_blockedPrompt guidance text violated safety policies or triggered prompt-injection defenses.Review input prompt text; blocked calls incur $0 cost.
401 Unauthorizedinvalid_api_keyVirtual key is invalid, revoked, or lacks access permissions for this route.Inspect the Authorization: Bearer sk-nrouter-... key configuration.
402 Payment Requiredinsufficient_creditsOrganization balance is insufficient to reserve duration credits.Top up balance in the billing dashboard; minimum credit purchase is $5.
404 Not Foundmodel_not_foundSpecified transcription model is unrecognized or unavailable.Consult the model catalog via GET /v1/models for active transcription engines.
413 Payload Too Largepayload_too_largeAudio file exceeds the 25 MB file size limit.Compress the audio stream or split longer files into chunks prior to upload.
429 Too Many Requestsrate_limit_exceededAccount or key RPM ceiling exceeded.Check x-nr-limit-source header and apply exponential backoff.
500 / 503 Provider Errorservice_unavailableUpstream transcription provider failure or connection timeout.nRouter automatically initiates fallback routing if alternative models are configured.
POST
/v1/audio/transcriptions

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Header Parameters

x-nr-compress?string

Request prompt compression: on to compress eligible prompts, off to skip

Value in

  • "on"
  • "off"
x-nr-tags?string

Custom spend and attribution tags (comma-separated key=value pairs)

x-nr-mcp-server?string

Target MCP server ID when routing MCP tool calls or prompts through the gateway

Request Body

multipart/form-data

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/audio/transcriptions" \  -F model="string" \  -F file="string"
{  "text": "The transcribed speech."}
Was this page helpful?