Use Cases

Build it on one managed gateway.

RAG pipelines, AI agents, support bots, coding assistants, document processors, voice apps. They all need routing, guardrails, cost control, and observability. nRouter gives you all four behind one key across 169+ models.

nrouter · router console
gateway active · 6 workloads
Incoming RequestAuth Verified

Task: "Decompose SQL schema migration and run verification tests"

Autonomous Agents
TTFT 138 ms
Loop guard · Budget ceiling: $5.00/run
Routed Modelclaude-sonnet-4-5
Request Cost$0.00220
Multi-turn chain → Auto-failover ready (0.01% drop)0 BYOK keys
One nRouter KeyBudget GuardZero Server BYOK
99.99% Routing SLA
One key
169+

models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic

Cost reduction
40–75%

Via tier routing & semantic caching

Gateway overhead
95 ms

p50 added latency — LLM time dominates

Platform fee
4%

Flat fee on credits — no feature paywalls

Pick your workload

Six use cases, one gateway

Each page below is tailored to how that workload actually runs: which nRouter capabilities matter, and how the request flow looks end to end.

Workload 01

RAG pipelines

Retrieval-augmented generation, one key for embed + chat

A RAG pipeline calls two model families — embeddings to index and query, a chat model to synthesize the answer. Route both through one endpoint, cache the repeats, and track cost per pipeline stage.

  • Embed + chat routing
  • Response caching
  • Per-pipeline cost & usage
Explore rag pipelines
Workload 02

AI agents

Autonomous agents that stay up and stay in budget

Agents fan out dozens of LLM calls per task. Auto-failover keeps a long run alive when a provider degrades, per-agent budgets cap spend, and every tool call lands in the request log for audit.

  • Fallback chains
  • Per-agent budgets
  • Tool-call observability
Explore ai agents
Workload 03

Customer support bots

Support bots with PII redaction and injection defense built in

Support traffic carries real customer data and hostile inputs. Guardrails redact PII and block prompt injection on every request; versioned prompt templates keep tone consistent across the fleet.

  • PII redaction
  • Injection defense
  • Versioned prompt templates
Explore customer support bots
Workload 04

Code generation

Coding assistants with model choice and per-seat budgets

Coding tools mix fast autocomplete with deep reasoning. Pick the model per task from the catalog, route latency-sensitive calls to the quickest endpoint, and cap spend with per-key, per-seat budgets.

  • Model catalog choice
  • Latency routing
  • Per-seat budgets
Explore code generation
Workload 05

Document processing

Extract, classify, and summarize documents at volume

Document workloads run vision-capable models over scans and PDFs in large batches. Per-key model access keeps a batch on the vision-capable set, and cost & usage tracking attributes spend per batch job.

  • Vision-capable models
  • Batch-friendly routing
  • Per-job cost & usage
Explore document processing
Workload 06

Voice & realtime AI

Low-latency routing for voice and realtime assistants

Voice apps live or die on latency. Latency-based routing steers each turn to the quickest healthy endpoint, fallbacks keep the conversation alive, and every turn is logged for observability.

  • Latency-based routing
  • Turn-level observability
  • Model catalog
Explore voice & realtime ai
The common foundation

What every use case shares

The workloads differ, but the gateway underneath is the same. These three guarantees hold whether you ship a RAG bot or a voice agent.

One key, no provider config

Every use case below starts the same way: one nRouter key, the OpenAI-compatible endpoint, and the model catalog. No provider accounts, no key vault, no per-model SDK.

Guardrails and budgets, always on

PII redaction, prompt-injection detection, and per-org / per-team / per-key budgets apply to every workload: RAG, agents, or voice. You opt out, never in.

Cost and logs you can attribute

nRouter reports the real cost of every call; nRouter logs the strategy, model, latency, and result. Slice spend by pipeline, agent, batch job, or seat.

Enterprise Architecture

Universal workload execution, end to end

Whether powering autonomous agent loops, high-volume RAG embeddings, voice streams, or IDE completions, every request passes through an identical audited gateway chain.

Enterprise request lifecycle: client → gateway → guardrail filter → smart router → model provider

  1. Client Workload

    RAG · Agents · Voice · Apps

    Standard OpenAI-compatible SDK calls from any application.

  2. Unified Gateway

    :4000 · In-Memory RLS

    Sub-millisecond auth, per-key budgets, rate limits, and tenancy.

  3. Guardrail Filter

    PII · Secret · Safety Scan

    Inline PII redaction and prompt-injection defense before inference.

  4. Smart Router

    Cost · Latency · Intent

    40–75% compute savings steering tasks to optimal model tiers.

  5. Model Providers

    OpenAI · Anthropic · Azure · Alibaba

    99.99% multi-provider failover with zero-markup list pricing.

The Rust proxy adds less than 95 ms at p50. Cache hits return in under 15ms with 0% token spend, while multi-provider failover guarantees 99.99% operational uptime.

Unified Integration

One client setup for any AI workload

Point your standard OpenAI-compatible client to nRouter. Configure guardrail profiles, virtual key budgets, and fallback chains via standard headers without custom SDKs.

Installpip install openai
1# Cache: enabled (org default). Pass nrouter_cache: false to skip.
2from openai import OpenAI
3import os
4
5client = OpenAI(
6 api_key=os.environ["NROUTER_API_KEY"],
7 base_url="https://api.nrouter.ai/v1",
8)
9
10response = client.chat.completions.create(
11 model="gpt-5.4-mini",
12 temperature=1,
13 max_completion_tokens=1024,
14 messages=[
15 {"role": "user", "content": "Hello! What models do you support?"},
16 ],
17 extra_body={
18 # "nrouter_cache": False, # Uncomment to skip cache
19 },
20)
21
22print(response.choices[0].message.content)

One key reaches every model across leading providers — no per-provider infrastructure or key vault sprawl.

FAQ

Workload & gateway questions, answered

How does nRouter support multiple AI workloads behind a single key?

nRouter provides a unified OpenAI-compatible endpoint that serves chat completions, streaming, embeddings, vision, and audio models across leading foundation providers. You configure your client once with an sk-nrouter-* key, and your entire stack (agents, RAG pipelines, support bots, and coding tools) can access any supported model without separate provider accounts or SDK swaps.

How does smart tier routing reduce operating costs across use cases?

Smart tier routing analyzes prompt intent, context length, and required capabilities to dynamically route requests to the most cost-effective model that satisfies the task. Routine queries, data extractions, and classification tasks route to high-speed micro-models (saving 40% to 75%), while complex reasoning tasks automatically escalate to frontier models.

What happens if a foundation model provider experiences an outage during production traffic?

nRouter monitors upstream provider health in real time with continuous health checks. If an upstream provider returns 429 rate-limit errors, 5xx server errors, or degrades in latency, nRouter automatically fails over to the next configured fallback model or deployment in milliseconds without dropping the connection or failing the user request.

How do inline guardrails protect production applications from data leaks and prompt injections?

Every request passes through an inline guardrail inspection layer before payload egress. It detects and sanitizes sensitive data (PII, SSNs, credit cards, credentials) and blocks adversarial jailbreaks or prompt-injection exploits before they reach foundation models. Redacted events and blocked attempts are captured in tamper-evident security audit logs.

Can I enforce distinct budgets and rate limits per workload, team, or agent?

Yes. Virtual keys can be scoped by organization, team, project, or autonomous agent. You can assign hard or soft budget envelopes (with automatic HTTP 402 rejection when depleted) and configure distinct requests-per-minute (RPM) and tokens-per-minute (TPM) limits to prevent any single workload from exhausting shared account capacity.

Does nRouter add noticeable latency to streaming or interactive voice applications?

No. The native Rust gateway proxy overhead is less than 1 millisecond at p50. Streaming responses are delivered token-by-token directly from upstream providers with zero hot-path buffering, making nRouter ideal for latency-critical voice bots, IDE completions, and realtime conversational agents.

One key. One bill. Every workload.

Start with the use case that fits your build

Sign up, paste your virtual key, change the base URL. Routing, guardrails, budgets, and observability are unlocked on every plan.