13 Evaluated CategoriesArena ELO · Verified Production Telemetry

Gateway Model Ladderboard

Live LLM intelligence, speed, and cost evaluations across 13 categories. Every model is servable instantly through api.nrouter.ai/v1/* with 0% token markup and automatic multi-cloud failover.

Data Plane Invariants

Engineered for Enterprise Production Scale

Ladderboard metrics backed by carrier-grade gateway infrastructure. nRouter unifies every top foundation model under one resilient interface.

Sub-150ms Gateway TTFT

Global edge proxying with direct fiber peering to AWS, Azure, and Google Cloud datacenters, adding less than 2ms of network overhead.

99.99% Multi-Cloud Failover

Instant zero-code failover across Azure OpenAI, AWS Bedrock, GCP Vertex AI, and direct provider wires to eliminate 503 outages and queue spikes.

0% Token Markup (Rule #28)

Direct upstream list prices. 100% of prompt caching discounts (-50% to -90%) passed through directly with zero hidden surcharges.

11-Intent Smart Routing

Gateway analyzes incoming prompt complexity and intent (chitchat vs heavy reasoning), cutting enterprise spend by up to 78% via nrouter/auto.

Methodology & Architecture

Frequently Asked Questions

Key technical details on evaluation synthesis, failover engineering, and billing transparency.

How does the nRouter ladderboard synthesize Arena ELO with production metrics?

The ladderboard combines crowdsourced double-blind pairwise Arena ELO scores with empirical production telemetry flowing through our high-performance Rust gateway: measured time-to-first-token (TTFT), actual provider rate-limit frequencies, streaming token throughput, and verified task benchmarks (SWE-bench, HumanEval, MMLU).

How does multi-cloud failover prevent P99 latency spikes?

When an individual cloud region experiences GPU queue accumulation or HTTP 429 throttling, nRouter detects the delay in sub-milliseconds and automatically shifts subsequent requests to alternate provider wires (e.g. from Azure OpenAI to Bedrock or Direct), preserving sub-second response times.

How do prompt caching discounts affect effective model costs?

On models supporting prompt prefix caching (Anthropic Claude, OpenAI GPT-4o, Google Gemini), cached token reads reduce input costs by 50% to 90%. nRouter passes 100% of these savings directly to your account with zero per-token markup, lowering effective blended costs dramatically.

How can developers automate model selection using nRouter smart aliases?

Instead of hardcoding individual model IDs, applications can target smart aliases like nrouter/auto. The gateway pre-scores incoming prompt intent and routes simple queries to ultra-fast commodity models ($0.10/1M) and complex multi-step reasoning to frontier models ($3.00/1M).

Route to any model on the ladderboard in under 60 seconds

One universal API key (sk-nrouter-*) gives your applications instant access to all models above with OpenAI-compatible endpoints, zero token markups, and automatic multi-cloud failover.