Models / Voxtral Mini 3B (Audio)ID: bedrock/mistral.voxtral-mini-3b-2507
by Mistral AI·via AWS Bedrock

Voxtral Mini 3B (Audio)

Voxtral Mini 3B (Audio) is a chat model made by Mistral AI, served on nRouter.

Maker

Mistral AI

Served On

AWS Bedrock

Start building with voxtral-mini-3b-2507

1

Get a key

2

Set your API key

export NROUTER_API_KEY="<YOUR_API_KEY>"
3

Configure Claude Code

npm install -g @anthropic-ai/claude-code
# Point Claude Code to nRouter gateway
export ANTHROPIC_BASE_URL="https://api.nrouter.ai"
export ANTHROPIC_API_KEY="$NROUTER_API_KEY"

# Launch with bedrock/mistral.voxtral-mini-3b-2507
claude --model bedrock/mistral.voxtral-mini-3b-2507

Pricing

Rates are read live

Pricing for voxtral-mini-3b-2507 is fetched from the live catalogue on load, so it is never served from a cache that could outlive a repricing.

Pricing Transparency & Zero-Markup Guarantee

nRouter strictly operates with a Zero per-token markup guarantee. All inference requests for Voxtral Mini 3B (Audio) are billed at exact upstream provider list prices with zero per-token margin, zero routing fees, and zero hidden platform overhead.

Exact List Price

Token consumption is settled at the published upstream rate of the exact served model and provider tier that executed your call.

100% Cache Savings

Upstream prompt caching discounts (up to 50%–90% reduction on cached prompt tokens) pass through directly to your balance with zero added fee.

Live Cost Header

Every response includes the authoritative x-nr-request-cost header for instant accounting and FinOps reconciliation.

Specifications

Context window

—

Max output

—

Modality

chat

Context window, compared

voxtral-mini-3b-2507UnknownUnknownqwen-mt-turboqwen-mt-turbo — 131.1K tokens131.1Kglm-5.3-primeUnknownUnknownqwen-plus-latestqwen-plus-latest — 131.1K tokens131.1Kgpt-4o-mini-realtime-preview-2024-12-17gpt-4o-mini-realtime-preview-2024-12-17 — 128K tokens128Kqwen3-235b-a22b-2507UnknownUnknownqwen3.7-max-2026-05-17UnknownUnknown

Bars are scaled to 131.1K tokens. A model with no published context window reads Unknown, never zero.

Vision
Function calling
System prompt

Context Limits & Caching Economics

Context Window Boundaries

With an effective context ceiling of — tokens and a maximum output generation ceiling of — tokens, voxtral-mini-3b-2507 is engineered to support long-document analysis, repository code exploration, and complex multi-turn dialog without token truncation.

Prompt Caching Economics

Upstream prefix caching stores static context (system prompts, tool definitions, documentation corpora) in memory. Cached token reads receive up to 50%–90% cost discounts from upstream providers, passed through directly to your nRouter balance with zero added fee.

Ideal Production Workloads
  • ✦Autonomous Tool Use: Structured JSON schema validation and multi-step function calling loops.
  • ✦Code Intelligence: Multi-file codebase refactoring, unit test generation, and complex algorithmic reasoning.
  • ✦High-Density RAG: In-context information extraction across large token windows without needle loss.
  • ✦Multimodal Ingestion: Visual documentation analysis, image parsing, and structured asset extraction.

Provider-Reported Capabilities

What the hosting provider reports for this model, shown as it was returned. nRouter adds nothing to this report and infers nothing that is missing from it.

Model ID modelId

mistral.voxtral-mini-3b-2507

Converse converse
Max Tokens Default maxTokensDefault

null

Max Tokens Maximum maxTokensMaximum

32768

Reasoning Supported reasoningSupported

null

System Role Supported systemRoleSupported
Stop Sequences Default stopSequencesDefault
User Image Types Supported userImageTypesSupported
User Video Types Supported userVideoTypesSupported
User Document Types Supported userDocumentTypesSupported
Model Name modelName

Voxtral Mini 3B 2507

Description description

null

Model Family modelFamily

Voxtral

Provider Name providerName

Mistral AI

Batch Supported batchSupported
Base Model Supported baseModelSupported
Cross Region Supported crossRegionSupported
Custom Model Supported customModelSupported
Model Lifecycle modelLifecycle
Status status

ACTIVE

Legacy Time legacyTime

null

End Of Life Time endOfLifeTime

null

Public Extended Access Time publicExtendedAccessTime

null

Input Modalities inputModalities
SPEECHTEXT
Output Modalities outputModalities
TEXT
Features Supported featuresSupported
Count Tokens countTokens
Prompt Caching promptCaching
Batch Inference batchInference
In Region inRegion
Cross Region crossRegion
Prompt Optimization promptOptimization
Provisioned Throughput provisionedThroughput
Intelligent Prompt Routing intelligentPromptRouting
Console IDEMetadata consoleIDEMetadata

null

Guardrails Supported guardrailsSupported
Explicit Prompt Caching explicitPromptCaching
Is Supported isSupported
Inference APIs Supported inferenceAPIsSupported
Converse converse
Sync sync
Streaming streaming
Async Invoke asyncInvoke
Invoke Model invokeModel
Sync sync
Response Streaming responseStreaming
Bidirectional Streaming bidirectionalStreaming
Open AI Responses openAiResponses
Open AI Chat Completions openAiChatCompletions
Customizations Supported customizationsSupported
Inference Types Supported inferenceTypesSupported
ON_DEMAND
Intelligent Prompt Routing intelligentPromptRouting
Is Supported isSupported
Response Streaming Supported responseStreamingSupported
Latency Optimization Supported latencyOptimizationSupported

Multi-Provider Redundancy

nRouter routes traffic for Voxtral Mini 3B (Audio) across multiple underlying cloud hyperscalers to guarantee continuous zero-downtime availability and eliminate vendor lock-in:

Cross-Cloud Fleet

Traffic for bedrock/mistral.voxtral-mini-3b-2507 routes dynamically across Microsoft Azure, AWS Bedrock, Google Cloud Vertex AI, and direct provider endpoints without code changes.

Automatic Zero-Downtime Failover

If an upstream cloud provider reports rate limits (HTTP 429), regional degradation, or an outage, nRouter automatically retries on an alternate cloud backend in <1ms.

Unified Single-Key Control

A single nRouter virtual key provides unified access with enforced spending budgets, model allowlists, and end-to-end SOC 2 compliant audit logging.

Availability

Not enough data yet

We have not collected enough health probes for this model to publish an uptime or latency figure. A percentage from a handful of samples is fabricated precision, so none is shown until the sample count supports one.

Interactive Playground

Send a real test request to voxtral-mini-3b-2507 using your virtual key.

Test it

Prompt
cURL
curl https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"bedrock/mistral.voxtral-mini-3b-2507","messages":[{"role":"user","content":"Hello!"}]}'

More from Mistral AI