13 Evaluated CategoriesArena ELO · Verified Production Telemetry

Gateway Model Leaderboard

Live LLM intelligence, speed, and cost evaluations across 13 categories. Every model is servable instantly through api.nrouter.ai/v1/* with 0% token markup and automatic multi-cloud failover.

Data Plane Invariants

Engineered for Enterprise Production Scale

Leaderboard metrics backed by carrier-grade gateway infrastructure. nRouter unifies every top foundation model under one resilient interface.

Sub-150ms Gateway TTFT

Global edge proxying with direct fiber peering to AWS, Azure, and Google Cloud datacenters, adding less than 2ms of network overhead.

99.99% Multi-Cloud Failover

Instant zero-code failover across Azure OpenAI, AWS Bedrock, GCP Vertex AI, and direct provider wires to eliminate 503 outages and queue spikes.

0% Token Markup (Rule #28)

Direct upstream list prices. 100% of prompt caching discounts (-50% to -90%) passed through directly with zero hidden surcharges.

11-Intent Smart Routing

Gateway analyzes incoming prompt complexity and intent (chitchat vs heavy reasoning), cutting enterprise spend by up to 78% via nrouter/auto.

Methodology & Architecture

Frequently Asked Questions

Key technical details on evaluation synthesis, failover engineering, and billing transparency.

Why is this called "Arena" and what does the word "Arena" mean?

The word "Arena" originates from the LMSYS Chatbot Arena (lmarena.ai), the foundational open research project by researchers from UC Berkeley, UCSD, and CMU. In an arena evaluation, two AI models compete blindly side-by-side (like gladiators in a digital arena) responding to real-world user prompts. Human evaluators judge the winning response without knowing model identities (double-blind). Using Bradley-Terry statistical models, pairwise battle outcomes are computed into Elo ratings—the same metric used in tournament chess—along with 95% bootstrap confidence intervals. nRouter displays official Chatbot Arena Elo ratings for conversational capability while combining them with our live gateway metrics: Time-To-First-Token (TTFT), token throughput, multi-cloud failover, and exact 0% markup list pricing.

Where is the leaderboard data sourced from?

All model rankings, Arena Elo ratings, and pairwise battle metrics are sourced directly from the official LMSYS Chatbot Arena (lmarena-ai/leaderboard-dataset hosted on Hugging Face) via verified Datasets Server APIs with official token authentication. Real-world runtime telemetry—including edge Time-To-First-Token (TTFT), streaming throughput, cloud failover rate-limits, and prompt cache hit rates—is gathered from live production requests routed through the nRouter Rust gateway with 0% token markup.

How are the Arena Elo scores and 95% Confidence Intervals calculated?

Arena Elo ratings are derived from crowdsourced double-blind, randomized A/B human evaluation battles. The 95% confidence intervals (±CI or [rating_lower, rating_upper]) are computed using 1,000 bootstrap resamplings on the pairwise battle matrix. Overlapping confidence intervals indicate that two models are statistically tied, meaning neither model can be asserted as definitively superior without additional battle volume.

How does the nRouter leaderboard synthesize Arena ELO with production metrics?

The leaderboard combines crowdsourced double-blind pairwise Arena ELO scores with empirical production telemetry flowing through our high-performance Rust gateway: measured time-to-first-token (TTFT), actual provider rate-limit frequencies, streaming token throughput, and verified task benchmarks (SWE-bench, HumanEval, MMLU).

How does nRouter offer zero per-token markup on all leaderboard models?

In strict adherence to Rule #28, nRouter charges the exact list price of the served provider (e.g. OpenAI, Anthropic, AWS Bedrock, GCP Vertex AI, TypeSafe AI). When upstream providers offer prompt caching discounts (-50% to -90%), 100% of these savings are passed through directly to your account. There are zero token surcharges, zero hidden markups, and all costs reconcile transparently in your 22-column spend ledger.

How does multi-cloud failover prevent P99 latency spikes?

When an individual cloud region experiences GPU queue accumulation or HTTP 429 throttling, nRouter detects the delay in sub-milliseconds and automatically shifts subsequent requests to alternate provider wires (e.g. from Azure OpenAI to Bedrock or Direct), preserving sub-second response times.

How do prompt caching discounts affect effective model costs?

On models supporting prompt prefix caching (Anthropic Claude, OpenAI GPT-4o, Google Gemini), cached token reads reduce input costs by 50% to 90%. nRouter passes 100% of these savings directly to your account with zero per-token markup, lowering effective blended costs dramatically.

How can developers automate model selection using nRouter smart aliases?

Instead of hardcoding individual model IDs, applications can target smart aliases like nrouter/auto. The gateway pre-scores incoming prompt intent and routes simple queries to ultra-fast commodity models ($0.10/1M) and complex multi-step reasoning to frontier models ($3.00/1M).

How can developers route to any model on this leaderboard in 60 seconds?

Any model listed on the leaderboard can be called immediately through our universal OpenAI-compatible endpoint (https://api.nrouter.ai/v1/chat/completions) using a single sk-nrouter-* API key. Simply configure the endpoint base URL and your authorization key in your existing OpenAI, LangChain, Vercel AI SDK, or LlamaIndex application.

Route to any model on the leaderboard in under 60 seconds

One universal API key (sk-nrouter-*) gives your applications instant access to all models above with OpenAI-compatible endpoints, zero token markups, and automatic multi-cloud failover.