One key, no provider config
Every use case below starts the same way: one nRouter key, the OpenAI-compatible endpoint, and the model catalog. No provider accounts, no key vault, no per-model SDK.
RAG pipelines, AI agents, support bots, coding assistants, document processors, voice apps. They all need routing, guardrails, cost control, and observability. nRouter gives you all four behind one key across 169+ models.
Task: "Decompose SQL schema migration and run verification tests"
models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic
Via tier routing & semantic caching
p50 added latency — LLM time dominates
Flat fee on credits — no feature paywalls
Each page below is tailored to how that workload actually runs: which nRouter capabilities matter, and how the request flow looks end to end.
RAG pipelines
A RAG pipeline calls two model families — embeddings to index and query, a chat model to synthesize the answer. Route both through one endpoint, cache the repeats, and track cost per pipeline stage.
AI agents
Agents fan out dozens of LLM calls per task. Auto-failover keeps a long run alive when a provider degrades, per-agent budgets cap spend, and every tool call lands in the request log for audit.
Customer support bots
Support traffic carries real customer data and hostile inputs. Guardrails redact PII and block prompt injection on every request; versioned prompt templates keep tone consistent across the fleet.
Code generation
Coding tools mix fast autocomplete with deep reasoning. Pick the model per task from the catalog, route latency-sensitive calls to the quickest endpoint, and cap spend with per-key, per-seat budgets.
Document processing
Document workloads run vision-capable models over scans and PDFs in large batches. Per-key model access keeps a batch on the vision-capable set, and cost & usage tracking attributes spend per batch job.
Voice & realtime AI
Voice apps live or die on latency. Latency-based routing steers each turn to the quickest healthy endpoint, fallbacks keep the conversation alive, and every turn is logged for observability.
The workloads differ, but the gateway underneath is the same. These three guarantees hold whether you ship a RAG bot or a voice agent.
Every use case below starts the same way: one nRouter key, the OpenAI-compatible endpoint, and the model catalog. No provider accounts, no key vault, no per-model SDK.
PII redaction, prompt-injection detection, and per-org / per-team / per-key budgets apply to every workload: RAG, agents, or voice. You opt out, never in.
nRouter reports the real cost of every call; nRouter logs the strategy, model, latency, and result. Slice spend by pipeline, agent, batch job, or seat.
Whether powering autonomous agent loops, high-volume RAG embeddings, voice streams, or IDE completions, every request passes through an identical audited gateway chain.
Enterprise request lifecycle: client → gateway → guardrail filter → smart router → model provider
Client Workload
RAG · Agents · Voice · Apps
Standard OpenAI-compatible SDK calls from any application.
Unified Gateway
:4000 · In-Memory RLS
Sub-millisecond auth, per-key budgets, rate limits, and tenancy.
Guardrail Filter
PII · Secret · Safety Scan
Inline PII redaction and prompt-injection defense before inference.
Smart Router
Cost · Latency · Intent
40–75% compute savings steering tasks to optimal model tiers.
Model Providers
OpenAI · Anthropic · Azure · Alibaba
99.99% multi-provider failover with zero-markup list pricing.
The Rust proxy adds less than 95 ms at p50. Cache hits return in under 15ms with 0% token spend, while multi-provider failover guarantees 99.99% operational uptime.
Point your standard OpenAI-compatible client to nRouter. Configure guardrail profiles, virtual key budgets, and fallback chains via standard headers without custom SDKs.
pip install openai| 1 | # Cache: enabled (org default). Pass nrouter_cache: false to skip. |
| 2 | from openai import OpenAI |
| 3 | import os |
| 4 | |
| 5 | client = OpenAI( |
| 6 | api_key=os.environ["NROUTER_API_KEY"], |
| 7 | base_url="https://api.nrouter.ai/v1", |
| 8 | ) |
| 9 | |
| 10 | response = client.chat.completions.create( |
| 11 | model="gpt-5.4-mini", |
| 12 | temperature=1, |
| 13 | max_completion_tokens=1024, |
| 14 | messages=[ |
| 15 | {"role": "user", "content": "Hello! What models do you support?"}, |
| 16 | ], |
| 17 | extra_body={ |
| 18 | # "nrouter_cache": False, # Uncomment to skip cache |
| 19 | }, |
| 20 | ) |
| 21 | |
| 22 | print(response.choices[0].message.content) |
One key reaches every model across leading providers — no per-provider infrastructure or key vault sprawl.
nRouter provides a unified OpenAI-compatible endpoint that serves chat completions, streaming, embeddings, vision, and audio models across leading foundation providers. You configure your client once with an sk-nrouter-* key, and your entire stack (agents, RAG pipelines, support bots, and coding tools) can access any supported model without separate provider accounts or SDK swaps.
Smart tier routing analyzes prompt intent, context length, and required capabilities to dynamically route requests to the most cost-effective model that satisfies the task. Routine queries, data extractions, and classification tasks route to high-speed micro-models (saving 40% to 75%), while complex reasoning tasks automatically escalate to frontier models.
nRouter monitors upstream provider health in real time with continuous health checks. If an upstream provider returns 429 rate-limit errors, 5xx server errors, or degrades in latency, nRouter automatically fails over to the next configured fallback model or deployment in milliseconds without dropping the connection or failing the user request.
Every request passes through an inline guardrail inspection layer before payload egress. It detects and sanitizes sensitive data (PII, SSNs, credit cards, credentials) and blocks adversarial jailbreaks or prompt-injection exploits before they reach foundation models. Redacted events and blocked attempts are captured in tamper-evident security audit logs.
Yes. Virtual keys can be scoped by organization, team, project, or autonomous agent. You can assign hard or soft budget envelopes (with automatic HTTP 402 rejection when depleted) and configure distinct requests-per-minute (RPM) and tokens-per-minute (TPM) limits to prevent any single workload from exhausting shared account capacity.
No. The native Rust gateway proxy overhead is less than 1 millisecond at p50. Streaming responses are delivered token-by-token directly from upstream providers with zero hot-path buffering, making nRouter ideal for latency-critical voice bots, IDE completions, and realtime conversational agents.
One key. One bill. Every workload.
Sign up, paste your virtual key, change the base URL. Routing, guardrails, budgets, and observability are unlocked on every plan.