Enterprise LLM Gateway
The High-Performance Gateway
for Production LLMs.
A high-throughput, drop-in OpenAI-compatible Rust gateway. Route across Live models with sub-millisecond proxy overhead, automated multi-cloud failover, Cortex mTLS safety preflight, and atomic credit reservation.
Rust-native proxy · Sub-1ms routing overhead · 5 providers live · Flat 4% fee
- <1msGateway routing overheadRust async core (P99)
- LiveModels liveAlibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic
- 99.99%Gateway availabilityMulti-cloud automated failover
- 4%Platform feeFlat fee on credits, added on top
One OpenAI-compatible endpoint. Sub-millisecond routing.
Replace dozens of fragmented vendor client libraries with a single, ultra-low-latency Rust gateway. Switch between models dynamically or leverage intent-based routing with zero code changes.
- Drop-in /v1/chat/completions wire protocol for Live models across 5 providers
- Sub-1ms proxy overhead powered by an asynchronous native Rust core
- Intent-aware routing (nrouter/auto) dynamically selects optimal cost/quality winner
Automated failover and circuit breaking across providers.
Never drop a user request due to upstream outages, rate limits (429), or degraded latency. The gateway continuously health-probes provider backends and automatically fails over in milliseconds.
- Instant circuit breaking on upstream HTTP 5xx errors or latency spikes
- Configurable fallback chains (e.g. Anthropic Direct → AWS Bedrock → Azure Foundry)
- Transparent client retries with zero connection hangs and preserved streaming SSE
Cortex mTLS inspection: $0 held, $0 spent on attacks.
Every request is evaluated through an air-gapped mTLS 1.3 sidecar (nrouter-cortex) before token egress. Block prompt injection, jailbreaks, and sensitive PII with zero egress leakage.
- Air-gapped mTLS 1.3 sidecar inspection: PII redaction and prompt injection classification
- Strict preflight sequencing: requests blocked during Phase 3 incur exactly $0 in spend
- Configurable guardrail policies scoped to organization, team, or individual virtual key
Atomic reserve-and-settle with zero financial overruns.
Prevent balance overdrafts before provider invocation. Phase 4 credit reservation atomically holds funds, then settles exact usage at the served model's flat list price down to the cent.
- Two-tier funding: debits subscription plan allowance first, then overflows to top-up credits
- Downwards-only reservation adjustments for smart routing — guarantees zero overcharges
- Precision telemetry headers on every response: x-nr-request-id, x-nr-request-cost, x-nr-latency-ms
The 5-phase gateway preflight chain.
Before a single token egresses to external providers, every incoming inference call undergoes strict sequential validation to guarantee tenant isolation, security, and credit safety.
Gate 0: Safety Hold
Kill-switch & account hold
Phase 1: WAF & Auth
Key hash & rate slots
Phase 2: Model ACL
Context ceilings & balance
Phase 3: Cortex mTLS
Injection & PII score
Phase 4: Credit Hold
Two-tier atomic hold
Phase 5: Egress & Settle
Exact list price settle
Gateway developer tooling & control plane.
The high-performance Rust data plane integrates directly with management tooling to monitor, configure, and inspect production LLM workloads in real time.
Get started
Route your first LLM request.
In less than two minutes.
Point your existing OpenAI SDK client to api.nrouter.ai/v1. Access Live models across all top providers with a single virtual key and built-in enterprise resilience.
Pay as you go from $5 · Flat 4% fee · Zero token markup