Use Case · AI Agents

Agents that stay up and stay in budget.

An autonomous agent fans out dozens of LLM calls per task. nRouter is the agent control plane. It keeps a long run alive when a provider degrades, caps each agent with a per-key budget, and logs every tool call for audit.

agent-run · key sk-nrouter-7f2a

One agent, one key, one budget

Steps this run24 calls
Fallback usedstep 11
Run statuscompleted
Budget$2.40 / $5.00
Tool calls logged24 / 24
Runaway protection402 ceiling
per-agent keyfailover-safeaudited
Agent cost savings
40–70%

Via adaptive step-level model tiering

Multi-hour uptime
99.99%

Continuous failover through upstream failures

Gateway overhead
95 ms

p50 added proxy latency in native Rust

Runaway protection
HTTP 402

Hard budget caps prevent infinite loops

Why nRouter for agents

Reliability, spend control, and a trail

Agents are unpredictable by design. The gateway makes them dependable: a run that survives a provider blip, a budget it cannot exceed, and a log of everything it did.

Auto-failover for long runs

Agents fan out dozens of calls per task. A provider hiccup at step 11 triggers the next fallback link transparently. Step 12 never knows, and the failover is logged.

Per-agent budgets

One virtual key per agent with its own budget. A runaway reasoning loop hits a hard 402 ceiling instead of burning credits; per-key RPM/TPM throttles a misbehaving agent.

Every tool call audited

Planning, tool selection, reflection: every LLM call lands in the request log with model, latency, cost, and result. Filter by the agent’s key to replay a whole run.

Model choice per step

Strong reasoning model for planning, fast model for tool-argument extraction. Pick per call from one catalog, no provider account, no SDK swap.

How it works

An agent run, end to end

Each agent carries its own virtual key. Every planning and tool-selection step is a separate LLM call — routed, failover-protected, budget-checked, and logged.

Agent run flow

  1. Agent Turn Ingress

    LangChain / CrewAI / AutoGen

    Autonomous planning and tool-calling step dispatches.

  2. Gateway & Agent Auth

    :4000 · In-Memory RLS

    Per-agent virtual keys, rate limits, and hard budget envelopes.

  3. Tool & Safety Guardrail

    Action & PII Defense

    Prevents tool injection exploits and redacts credentials before execution.

  4. Smart Step Router

    Task & Cost Optimizer

    Micro-models for tool args; reasoning models for planning (40–70% ROI).

  5. Multi-Model Providers

    Anthropic · OpenAI · Bedrock

    99.99% multi-provider failover keeping hours-long agent loops intact.

The budget check is a reserve-then-settle: credits are reserved before the call and settled at the real cost after. If an agent is out of budget, it gets a 402 — never a surprise overspend.

The code

Point your agent framework at one endpoint

nRouter speaks the OpenAI API, so any agent framework that targets the OpenAI SDK works with a base-URL and key change. These snippets are generated from the same SDK examples the playground uses. Give each agent its own key and the per-key budget does the rest.

Installpip install openai
1# Cache: enabled (org default). Pass nrouter_cache: false to skip.
2from openai import OpenAI
3import os
4
5client = OpenAI(
6 api_key=os.environ["NROUTER_API_KEY"],
7 base_url="https://api.nrouter.ai/v1",
8)
9
10response = client.chat.completions.create(
11 model="gpt-5.4-mini",
12 temperature=1,
13 max_completion_tokens=1024,
14 messages=[
15 {"role": "user", "content": "Hello! What models do you support?"},
16 ],
17 extra_body={
18 # "nrouter_cache": False, # Uncomment to skip cache
19 },
20)
21
22print(response.choices[0].message.content)

Issue one virtual key per agent — spend, rate limits, and the request log all scope to that key.

FAQ

Common agent questions

Why do autonomous AI agents require an LLM gateway?

Autonomous agents fan out dozens or hundreds of sequential LLM calls per goal: planning, tool evaluation, schema reflection, and final synthesis. A single provider rate limit or timeout aborts the entire multi-step run. nRouter provides multi-provider failover, hard spending limits, and an immutable trace ledger to make agents reliable, cost-bounded, and auditable.

How does nRouter prevent runaway spend during infinite agent loops?

Each autonomous agent runs with an isolated virtual key assigned a strict dollar ceiling. The moment an agent’s cumulative spend meets its cap, the gateway rejects subsequent calls with HTTP 402 Payment Required. Database-level credit holds prevent balances from ever going negative.

How does step-level model tiering reduce autonomous agent costs?

Not every step in an agent trajectory requires an expensive frontier reasoning model. nRouter enables agents to route JSON extraction, tool argument parsing, and interim checks to high-speed micro-models (saving 40% to 70%), reserving frontier models (like Claude 3.5 Sonnet or o1) for high-level goal decomposition.

How do inline guardrails protect agents against prompt injection and malicious tools?

As agents execute tools and retrieve untrusted external inputs (web scrapers, SQL queries, APIs), inline guardrails detect and neutralize indirect prompt injections before they can hijack the agent’s execution loop. System prompts, API keys, and sensitive database schemas are shielded from prompt exfiltration.

AI is becoming autonomous; control must become automatic.

Run agents on a gateway built for long runs

Auto-failover, per-agent budgets, and a full tool-call log — unlocked on every plan, behind one nRouter key.