
nRouter Sets New AI Gateway Speed Benchmark with Sub-Millisecond p99 Overhead Powered by Pure Rust Engine
Independent stress testing confirms nRouter's asynchronous Rust gateway streams tokens at wire speed with sub-millisecond p99 latency overhead, outperforming legacy Python and Node.js proxies.
SUNNYVALE, CA, July 24, 2026 — nRouter, the enterprise LLM gateway platform built by CloudAct Inc., today released independent benchmark results demonstrating that its native Rust gateway core introduces under 1 millisecond of p99 proxy latency overhead across high-concurrency streaming workloads.
As AI software transitions from asynchronous background processing to real-time agentic swarms, code autocomplete, and voice conversation loops, proxy latency has emerged as the critical bottleneck in enterprise user experience. Traditional gateways built on interpreted runtimes frequently introduce 30ms to 120ms of jitter, connection queuing, and token buffering delays before the first byte reaches the end user.
Bare-Metal Tokio Architecture
Engineered from the ground up in native Rust, nRouter's data plane eliminates runtime garbage collection pauses and memory copying overhead:
- Zero-Copy Streaming Pipeline: Server-Sent Events (SSE) and raw HTTP/2 chunked responses are proxied at wire speed directly to the client without intermediate payload buffering.
- Asynchronous Tokio Kernel: Leverages epoll/kqueue non-blocking I/O event loops capable of saturating 10Gbps network interfaces while maintaining flat latency curves.
- Sub-Millisecond p99 Overhead: Under load tests executing 100,000 concurrent streaming connections, nRouter measured an average latency overhead of 0.62ms and a p99 ceiling of 0.89ms.
Real-World Impact for Autonomous AI Agents
In multi-turn agentic workflows where an orchestrator issues dozens of chained reasoning requests to evaluate tools and synthesize context, cumulative gateway overhead in legacy systems routinely degrades responsiveness by several seconds. nRouter's wire-speed streaming ensures that time-to-first-token (TTFT) is bounded strictly by foundation model inference speeds, not routing infrastructure.
Transparent Model Health & Circuit Breaking
The Rust core continuously monitors provider socket health, TLS handshake durations, and upstream HTTP status codes. When an upstream provider experiences transient degradation or rate limiting (HTTP 429/503), nRouter's internal circuit breakers trigger sub-millisecond failover to designated backup models without disconnecting the client application.
Commentary from Engineering Leadership
"When we founded nRouter, we made the architectural decision to build our entire data plane in Rust rather than taking the common shortcut of wrapping Python or Node frameworks," said Rama Surasani, Founder and CEO of nRouter. "Today's benchmark results validate that choice. For latency-sensitive enterprise applications, every millisecond counts, and our gateway proves that enterprise governance does not have to come at the expense of speed."
Getting Started
Engineering teams can benchmark nRouter's proxy latency directly using our open-source client SDKs or standard OpenAI-compatible client libraries. Complete technical documentation and benchmark methodologies are available at nrouter.ai/docs.
For media inquiries: sales@nrouter.ai. For technical partnerships: sales@nrouter.ai.
About nRouter
nRouter is a managed LLM gateway for teams that want one key, one bill, and zero provider configuration. Customers buy credits, call any major model through a single endpoint, and get cost tracking, routing, and safety controls built in. nRouter is built and operated by CloudAct Inc..
More from the press room

Enterprise Adoption Surges as nRouter Reaches SOC 2 Type II In-Progress Milestone and Launches Dedicated VPC Deployments
Accelerating enterprise adoption across fintech and digital health with institutional security, Postgres Row-Level Security isolation, and private VPC deployment options.

nRouter Expands Gateway to 100+ Multimodal Models
Unified enterprise inference across Google Vertex AI, Anthropic Claude, OpenAI, AWS Bedrock, Meta Llama, and DeepSeek with sub-millisecond provider failover and multimodal support.

nRouter Abolishes Broker Markups with Flat List-Price Settlement and FinOps FOCUS 1.4 Tokenomics
Eliminating the 15% broker markup tax: nRouter establishes 0% per-token markup, settling at exact upstream provider list prices with automated 3-way FinOps FOCUS 1.4 reconciliation.