Models / GPT TranscribeID: openai/gpt-transcribe
by OpenAI·via OpenAI

GPT Transcribe

OpenAI GPT Transcribe is a high-performance audio transcription foundation model with automated multi-cloud failover and guardrails on nRouter.

Maker

OpenAI

Served On

OpenAI

Pricing

Rates are read live

Pricing for gpt-transcribe is fetched from the live catalogue on load, so it is never served from a cache that could outlive a repricing.

Pricing Transparency & Zero-Markup Guarantee

nRouter strictly operates with a Zero per-token markup guarantee. All inference requests for GPT Transcribe are billed at exact upstream provider list prices with zero per-token margin, zero routing fees, and zero hidden platform overhead.

Exact List Price

Token consumption is settled at the published upstream rate of the exact served model and provider tier that executed your call.

100% Cache Savings

Upstream prompt caching discounts (up to 50%–90% reduction on cached prompt tokens) pass through directly to your balance with zero added fee.

Live Cost Header

Every response includes the authoritative x-nr-request-cost header for instant accounting and FinOps reconciliation.

Specifications

Context window

—

Max output

—

Modality

audio transcription

Context window, compared

gpt-transcribeUnknownUnknowngpt-4o-mini-transcribe-2025-12-15gpt-4o-mini-transcribe-2025-12-15 — 16K tokens16Kwhisper-1UnknownUnknowngpt-4o-transcribe-diarizegpt-4o-transcribe-diarize — 16K tokens16Kgpt-4o-transcribegpt-4o-transcribe — 16K tokens16Kgpt-4o-mini-transcribe-2025-03-20gpt-4o-mini-transcribe-2025-03-20 — 16K tokens16Kgpt-4o-mini-transcribegpt-4o-mini-transcribe — 16K tokens16K

Bars are scaled to 16K tokens. A model with no published context window reads Unknown, never zero.

Vision
Function calling
System prompt

Context Limits & Caching Economics

Context Window Boundaries

With an effective context ceiling of — tokens and a maximum output generation ceiling of — tokens, gpt-transcribe is engineered to support long-document analysis, repository code exploration, and complex multi-turn dialog without token truncation.

Prompt Caching Economics

Upstream prefix caching stores static context (system prompts, tool definitions, documentation corpora) in memory. Cached token reads receive up to 50%–90% cost discounts from upstream providers, passed through directly to your nRouter balance with zero added fee.

Ideal Production Workloads
  • ✦Autonomous Tool Use: Structured JSON schema validation and multi-step function calling loops.
  • ✦Code Intelligence: Multi-file codebase refactoring, unit test generation, and complex algorithmic reasoning.
  • ✦High-Density RAG: In-context information extraction across large token windows without needle loss.
  • ✦Multimodal Ingestion: Visual documentation analysis, image parsing, and structured asset extraction.

Multi-Provider Redundancy

nRouter routes traffic for GPT Transcribe across multiple underlying cloud hyperscalers to guarantee continuous zero-downtime availability and eliminate vendor lock-in:

Cross-Cloud Fleet

Traffic for openai/gpt-transcribe routes dynamically across Microsoft Azure, AWS Bedrock, Google Cloud Vertex AI, and direct provider endpoints without code changes.

Automatic Zero-Downtime Failover

If an upstream cloud provider reports rate limits (HTTP 429), regional degradation, or an outage, nRouter automatically retries on an alternate cloud backend in <1ms.

Unified Single-Key Control

A single nRouter virtual key provides unified access with enforced spending budgets, model allowlists, and end-to-end SOC 2 compliant audit logging.

Availability

Not enough data yet

We have not collected enough health probes for this model to publish an uptime or latency figure. A percentage from a handful of samples is fabricated precision, so none is shown until the sample count supports one.

More from OpenAI