Sub-15ms LLM Gateway Routing with Warm HTTP/2 Pools
How connection pooling, in-memory preflight gating, and multiplexed HTTP/2 streams eliminate cold TCP/TLS handshakes and keep routing overhead under 15ms.
Static routing to a single frontier foundation model is financially inefficient and operationally fragile. Intelligent LLM routing dynamically matches incoming prompt complexity, modality requirements, and service level objectives with the optimal downstream model and provider deployment. By classifying prompt intent in real time, intelligent routing architectures slash token expenditure by up to 70% while improving response latency and uptime.
Routing simple chitchat, classification, or extraction queries to costly frontier models (like GPT-4o or Claude 3.5 Sonnet) unnecessarily inflates token spend when smaller, faster models (such as GPT-4o-mini or Gemini 2.0 Flash) provide equivalent accuracy.
Single-provider architectures suffer downtime during provider incidents or sudden tier exhaustion. Systems need dynamic fallback chains that automatically switch deployments within milliseconds.
Balancing cost constraints against strict time-to-first-token SLAs requires continuous evaluation of provider queue depths, historical response latencies, and regional geographic proximity.
Intelligent dispatchers must inspect token counts, multimodal attachments (images, audio, video), and tool schemas to filter only candidate models capable of serving the exact request specification.
11-intent categorization engine distinguishing light tasks from complex multi-step reasoning.
Configurable tier cascades with per-org retry budgets, timeout thresholds, and automated cooldowns.
Evaluates context ceilings, tool-calling support, and vision capabilities before candidate ranking.
nRouter implements Strategy::Intent to categorize incoming prompts across 11 distinct intent categories—from light tasks (chitchat, simple QA, summarization) to heavy computational tasks (code synthesis, mathematical reasoning, agentic planning). During Preflight Phase 3, nRouter calculates intent scoring alongside moderation, resolving the winning model tier before credit reservation. If a primary provider degrades or encounters rate limits, nRouter transparently executes configured fallback chains across alternative providers without dropping client streams. All models are billed at their exact upstream list prices with zero per-token markup.
Examine our engineering guides on smart routing algorithms, fallback chain patterns, and multi-model cost reduction below.

How connection pooling, in-memory preflight gating, and multiplexed HTTP/2 streams eliminate cold TCP/TLS handshakes and keep routing overhead under 15ms.

Point a router alias at a candidate set with the Latency strategy, keep the automatic cross-provider failover, tune the two settings that decide how fast a bad provider is abandoned, and prove the effect on your own p95.

Routing is the one cost lever you can pull from a dashboard. Point an alias at a set of models, choose cost or latency or weighted, and change what a request costs without touching a line of application code.

Request count and dollar cost tell different stories, and the gap between them is where the savings are. Here are the four shapes an overlay of cost and usage produces, which one to chase first, and what makes the numbers trustworthy enough to act on.

An LLM gateway is one endpoint in front of every model provider that owns six cross-cutting jobs — auth, cost, limits, safety, observability and failover. What it does, what happens to a request inside it, and when you need one.

Quality is a property of a task, not of a model. Here is how to inventory your traffic by task, point a router alias at a candidate set, prove each downgrade with an A/B test, and read the saving off the Cost vs Usage report.

A coin flip on every request is not an experiment — it is noise with a dashboard. Here is how deterministic hash-based assignment gives each user a stable variant for the life of a test, why the experiment id belongs in the hash, and what the gateway refuses to let a caller override.

How an ordered LLM fallback chain reroutes failing requests to healthy providers, advances only on retryable errors, and bills spend for exactly one hop.

Vendor-neutral buyer's-guide decision tree across the three durable LLM routing-intelligence shapes — benchmark-anchored (Unify-style), ML-classifier (NotDiamond-style), and operator-controlled (nRouter-style). Three questions, one shape, one product. Pick the failure mode your team is best equipped to own.

Not Diamond returns a model recommendation and charges $0.05 per million tokens routed — you still hold every provider key and make the call yourself. What that leaves you to build, and how deterministic A/B tests compare to a trained router when you have to reproduce a decision.

nRouter vs Unify AI for teams weighing benchmark-driven model arbitration against operator-pinned routing. Why a vendor leaderboard is not your eval, and how to compute cost-per-passing-answer from your own requests.

Eight axes that actually matter when picking an LLM gateway in 2026. Shortlist matrix across OpenRouter, Portkey, Helicone, nRouter. Decision tree by buyer profile, 90-minute evaluation.

Route each agent step to the model that fits its complexity, survive mid-chain provider outages, and capture one settled cost per run with an LLM gateway.