Use Case · Voice & Realtime

Voice assistants that answer fast — every turn.

A voice app lives or dies on latency. nRouter steers each conversational turn to the quickest healthy endpoint, keeps the line alive with transparent failover, and logs every turn for observability.

voice-turn · conversation 4c1d

One turn of a live conversation

Route strategylatency-based
Endpoint chosenquickest healthy
Streamingtoken-by-token
Gateway overhead~95 ms p50
Mid-call failovertransparent
Turn loggedp50 / p99
low-latencyfailover-safeobservable
Voice latency overhead
<1 ms

p50 added proxy latency in native Rust

Turn failover SLA
99.99%

Mid-conversation fallback with zero dropped calls

Cost reduction
40–60%

Via latency-aware flash tier steering

Realtime catalog
169+

models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic

Why nRouter for voice

Speed, resilience, and a turn-level trail

Voice is the least forgiving LLM workload. Every turn is on a clock and a caller is listening. nRouter gives that turn fast routing, transparent failover, and observability.

Latency-based routing

Silence on the line is the failure mode. Each conversational turn is steered to the quickest healthy endpoint using live latency signal. The gateway adds ~95 ms p50; LLM time dominates.

Failover mid-conversation

A provider blip during a live call cannot be a dead-air moment. The fallback chain retries transparently. The turn completes and the caller hears a response, not an error.

Turn-level observability

Every turn lands in the request log with model, latency, tokens, and cost — p50 / p99 per model surfaces a slow tail before callers feel it.

Model catalog for realtime

Pick the model with the latency profile you need, swap as faster models ship. No SDK change. Streaming runs with zero hot-path buffering.

How it works

A conversational turn, end to end

nRouter is the LLM hop between transcription and synthesis. The turn routes for latency, streams back token-by-token, and lands in the log with p50 / p99 metrics.

Voice turn flow

  1. Audio Turn Ingress

    speech-to-text layer

    Whisper / Deepgram transcriptions dispatched to nRouter.

  2. Realtime Gateway Auth

    :4000 · In-Memory RLS

    Sub-millisecond auth, concurrent channel budget allocation, and rate limits.

  3. Low-Latency Guardrail

    Streaming Safety Filter

    Non-blocking toxicity and PII checks without inducing speech jitter.

  4. Smart Latency Router

    p95 Latency Optimizer

    Live latency steering routes turns to the fastest endpoint (40–60% ROI).

  5. Streaming Providers

    Claude · GPT-4o · Gemini

    99.99% multi-provider failover with zero-buffered SSE stream to TTS.

Your speech-to-text and text-to-speech layers stay yours. nRouter gives the reasoning turn in between low-latency routing, failover, and a logged record. 169+ models to choose from.

The code

Stream the reasoning turn

A voice turn is a streaming chat request: transcript in, tokens out to your speech layer. These snippets come from the same SDK examples the playground uses; enable streaming and tokens flow as they generate.

Installpip install openai
1# Cache: enabled (org default). Pass nrouter_cache: false to skip.
2from openai import OpenAI
3import os
4
5client = OpenAI(
6 api_key=os.environ["NROUTER_API_KEY"],
7 base_url="https://api.nrouter.ai/v1",
8)
9
10response = client.chat.completions.create(
11 model="gpt-5.4-mini",
12 temperature=1,
13 max_completion_tokens=1024,
14 messages=[
15 {"role": "user", "content": "Hello! What models do you support?"},
16 ],
17 extra_body={
18 # "nrouter_cache": False, # Uncomment to skip cache
19 },
20)
21
22print(response.choices[0].message.content)

Set stream: true — nRouter streams the token stream natively with no hot-path buffering.

FAQ

Common voice & realtime questions

How does nRouter keep latency low for voice applications?

Routing decisions happen in-memory and add less than 1 millisecond (<1ms) at p50. LLM inference time dominates. Latency-based routing continuously monitors provider responsiveness and steers each conversational turn to the quickest healthy endpoint, and streaming responses run with zero buffering so tokens reach your TTS speech layer instantly.

What happens to a live voice call if a provider degrades or rate-limits?

nRouter’s circuit breaker executes transparent mid-conversation failover. If the primary foundation model returns a 5xx error or latency exceeds your SLA threshold, the request retries the backup provider in milliseconds. The caller continues speaking without experiencing an abrupt hangup or error tone.

Can inline guardrails operate without inducing audio stutter or jitter?

Yes. For voice workloads, guardrail filters run in an asynchronous stream-forwarding mode that inspects incoming transcripts and outgoing token chunks in parallel. Malicious prompts or offensive completions are deflected without delaying time-to-first-token (TTFT).

Can I monitor p50 and p99 latency for every conversational turn?

Yes. Every turn records end-to-end latency, time-to-first-token, provider processing time, token counts, and cost in the request log. Real-time metrics allow your voice engineering team to diagnose latency degradation before callers notice delays.

Does nRouter handle speech-to-text and text-to-speech directly?

nRouter provides the high-performance LLM intelligence layer between your STT (e.g. Whisper, Deepgram) and TTS (e.g. ElevenLabs, Cartesia) services. By focusing on sub-millisecond LLM routing, budget enforcement, and failover, nRouter keeps live conversational pipelines resilient.

Fast turns, no dead air

Build a voice assistant that never leaves the caller waiting

Latency-based routing, transparent failover, and turn-level observability. All unlocked on every plan.

nRouter is the LLM hop. Pair it with any speech-to-text and text-to-speech stack.

Explore other use cases

Production AI workflows powered by nRouter