Auto-failover for long runs
Agents fan out dozens of calls per task. A provider hiccup at step 11 triggers the next fallback link transparently. Step 12 never knows, and the failover is logged.
An autonomous agent fans out dozens of LLM calls per task. nRouter is the agent control plane. It keeps a long run alive when a provider degrades, caps each agent with a per-key budget, and logs every tool call for audit.
One agent, one key, one budget
Via adaptive step-level model tiering
Continuous failover through upstream failures
p50 added proxy latency in native Rust
Hard budget caps prevent infinite loops
Agents are unpredictable by design. The gateway makes them dependable: a run that survives a provider blip, a budget it cannot exceed, and a log of everything it did.
Agents fan out dozens of calls per task. A provider hiccup at step 11 triggers the next fallback link transparently. Step 12 never knows, and the failover is logged.
One virtual key per agent with its own budget. A runaway reasoning loop hits a hard 402 ceiling instead of burning credits; per-key RPM/TPM throttles a misbehaving agent.
Planning, tool selection, reflection: every LLM call lands in the request log with model, latency, cost, and result. Filter by the agent’s key to replay a whole run.
Strong reasoning model for planning, fast model for tool-argument extraction. Pick per call from one catalog, no provider account, no SDK swap.
Each agent carries its own virtual key. Every planning and tool-selection step is a separate LLM call — routed, failover-protected, budget-checked, and logged.
Agent run flow
Agent Turn Ingress
LangChain / CrewAI / AutoGen
Autonomous planning and tool-calling step dispatches.
Gateway & Agent Auth
:4000 · In-Memory RLS
Per-agent virtual keys, rate limits, and hard budget envelopes.
Tool & Safety Guardrail
Action & PII Defense
Prevents tool injection exploits and redacts credentials before execution.
Smart Step Router
Task & Cost Optimizer
Micro-models for tool args; reasoning models for planning (40–70% ROI).
Multi-Model Providers
Anthropic · OpenAI · Bedrock
99.99% multi-provider failover keeping hours-long agent loops intact.
The budget check is a reserve-then-settle: credits are reserved before the call and settled at the real cost after. If an agent is out of budget, it gets a 402 — never a surprise overspend.
nRouter speaks the OpenAI API, so any agent framework that targets the OpenAI SDK works with a base-URL and key change. These snippets are generated from the same SDK examples the playground uses. Give each agent its own key and the per-key budget does the rest.
pip install openai| 1 | # Cache: enabled (org default). Pass nrouter_cache: false to skip. |
| 2 | from openai import OpenAI |
| 3 | import os |
| 4 | |
| 5 | client = OpenAI( |
| 6 | api_key=os.environ["NROUTER_API_KEY"], |
| 7 | base_url="https://api.nrouter.ai/v1", |
| 8 | ) |
| 9 | |
| 10 | response = client.chat.completions.create( |
| 11 | model="gpt-5.4-mini", |
| 12 | temperature=1, |
| 13 | max_completion_tokens=1024, |
| 14 | messages=[ |
| 15 | {"role": "user", "content": "Hello! What models do you support?"}, |
| 16 | ], |
| 17 | extra_body={ |
| 18 | # "nrouter_cache": False, # Uncomment to skip cache |
| 19 | }, |
| 20 | ) |
| 21 | |
| 22 | print(response.choices[0].message.content) |
Issue one virtual key per agent — spend, rate limits, and the request log all scope to that key.
Autonomous agents fan out dozens or hundreds of sequential LLM calls per goal: planning, tool evaluation, schema reflection, and final synthesis. A single provider rate limit or timeout aborts the entire multi-step run. nRouter provides multi-provider failover, hard spending limits, and an immutable trace ledger to make agents reliable, cost-bounded, and auditable.
Each autonomous agent runs with an isolated virtual key assigned a strict dollar ceiling. The moment an agent’s cumulative spend meets its cap, the gateway rejects subsequent calls with HTTP 402 Payment Required. Database-level credit holds prevent balances from ever going negative.
Not every step in an agent trajectory requires an expensive frontier reasoning model. nRouter enables agents to route JSON extraction, tool argument parsing, and interim checks to high-speed micro-models (saving 40% to 70%), reserving frontier models (like Claude 3.5 Sonnet or o1) for high-level goal decomposition.
As agents execute tools and retrieve untrusted external inputs (web scrapers, SQL queries, APIs), inline guardrails detect and neutralize indirect prompt injections before they can hijack the agent’s execution loop. System prompts, API keys, and sensitive database schemas are shielded from prompt exfiltration.
Yes. Because nRouter provides a 100% OpenAI-compatible endpoint, any agent framework (LangChain, LangGraph, CrewAI, AutoGen, LlamaIndex, OpenAI Swarm) works by setting the baseURL to https://api.nrouter.ai/v1 and supplying your nRouter key. Function calling, streaming, and tool execution are supported natively.
AI is becoming autonomous; control must become automatic.
Auto-failover, per-agent budgets, and a full tool-call log — unlocked on every plan, behind one nRouter key.
Explore other use cases