Empirical Evaluations & LatencyVerified Technical Reports

Model Benchmarks

Empirical evaluations across MMLU, HumanEval, GPQA, MATH, and SWE-bench with measured latency percentiles and zero-markup pricing.

Methodology

How a nRouter benchmark is run

A benchmark is only worth printing if someone else can reproduce it. Every study we publish follows the same four rules, and ships the config so you can run it yourself.

1

Isolate the gateway

Run against the zero-cost mock provider so provider latency is constant. Whatever moves is gateway overhead. Nothing else.

2

Warm, then sample

Discard cold-start requests, then collect a large fixed sample at steady state. Cold starts are reported separately, never blended into p50.

3

Report the tail

Publish p50, p95, and p99. Never just the average. A gateway that looks fast at p50 and stalls at p99 is a slow gateway.

4

Show the config

Every run ships its config: region, instance size, guardrails enabled, concurrency, sample size, and date. A benchmark you cannot reproduce is marketing.

How results will be published
  • A dated report, not a moving number. Each study is timestamped and versioned. Older reports stay live so you can see the trend, not just today’s headline.

  • Full run configuration. Region, instance size, concurrency, guardrails enabled, sample size, and provider mix — everything needed to reproduce the run.

  • p50, p95, and p99 — always the tail. Every latency claim ships all three percentiles. A single average is not a benchmark; it is a hope.

  • Failover and error rates alongside latency. Reliability and speed are reported together. A gateway that is fast only when nothing fails is not actually fast.

Want the methodology in detail, or to run a benchmark against your own workload before committing? Email sales@nrouter.ai.

Uptime

Availability is a number we sign

The 99.9% uptime SLA is a contractual commitment in the legal SLA, not a figure we picked because it looked good on a slide. The gateway runs on managed autoscaling container infrastructure and managed Supabase Postgres.

99.9%

Committed, monitored, and public — component health and incident history are never behind a login.

FAQ

Frequently asked questions

Understanding synthetic evaluation metrics, LMSYS Chatbot Arena human preference scores, and live production telemetry.

How do synthetic benchmarks differ from the LMSYS Chatbot Arena?

Synthetic benchmarks (such as MMLU, HumanEval, MATH, GPQA, and SWE-bench Verified) measure an AI model against fixed, objective ground-truth datasets with automated scoring. In contrast, the LMSYS Chatbot Arena evaluates models through double-blind pairwise battles where human users judge which model gave the better subjective answer. nRouter combines both: reproducible synthetic accuracy metrics on this Benchmarks page, and crowdsourced Arena Elo ratings on our Leaderboard page, combined with real-world latency, availability, and cost data.

Why does nRouter evaluate latency and failover alongside raw benchmark scores?

A model with a top-tier MMLU or Arena score is ineffective in production if its Time-To-First-Token (TTFT) exceeds 5 seconds, or if its cloud provider frequently returns HTTP 429 rate-limit errors during peak traffic. nRouter measures and optimizes the entire inference chain: benchmark intelligence, P95/P99 latency, multi-cloud failover uptime, and exact list-price costs with 0% token markup.

Can I route between high-accuracy and low-cost models automatically?

Yes. Through nRouter smart router aliases such as nrouter/auto, incoming requests are scored by preflight intent. Lightweight tasks (chitchat, simple classification) route to lightning-fast commodity models, while complex tasks (multi-step coding, deep reasoning) route to frontier models, reducing blended inference costs by up to 78% without compromising output quality.

Measure it yourself

The fastest benchmark is your own traffic

Point a real workload at nRouter and watch the overhead in your own dashboard. No fabricated leaderboard to trust. Just the gateway, your requests, and the numbers.