← All posts
Buyer's Guide

LLM gateway buyer's guide 2026: routing, guardrails, evals, prompt management

Eight axes that actually matter when picking an LLM gateway in 2026. Shortlist matrix across OpenRouter, Portkey, Helicone, nRouter. Decision tree by buyer profile, 90-minute evaluation.

LLM gateway buyer's guide 2026: routing, guardrails, evals, prompt management

The Direct Answer: Selecting an LLM gateway in 2026 comes down to eight core evaluation criteria: intelligent routing/failover, inline guardrails (PII/jailbreak), 4-tier budget ceilings, evals, prompt management, full-stack observability, provider economics (0% token markup vs reseller inflation), and enterprise compliance (SOC2/RBAC). nRouter provides every governance and security feature out of the box on all tiers without enterprise paywalls, charging a transparent 4% platform fee on credits.

LLM Gateway Buyer's Guide 2026: 8 Evaluation Axes Across Top Gateways

Figure 1: Enterprise Gateway Evaluation Framework — The eight architectural pillars determining production longevity, security posture, and true total cost of ownership.

If you are evaluating an LLM gateway in 2026, the OpenAI / Anthropic / Vertex / Bedrock direct path has stopped scaling — on cost, attribution, guardrails, or the velocity of swapping models without redeploying half your stack. This guide names the eight evaluation axes that matter and applies them across the gateways most teams shortlist, with vendor pricing pages cited by URL with a 2026-05-16 audit timestamp.

We are unambiguous about where nRouter sits — that is the wedge: every gateway feature you would otherwise pay an enterprise contract for is included on every nRouter plan (Pay as you go, Starter, Pro, Max, Enterprise), with one platform fee on every self-serve plan (4% of the credits at purchase, charged on top, no minimum fee) and no plan varying the feature set.

Five-minute path: eight axes, shortlist matrix, decision tree.


What an LLM gateway actually is (and isn't)

An LLM gateway is a single API endpoint between your application and N upstream model providers. It does three things: routes requests by your rules (cost, latency, fallback, A/B), observes them with attribution metadata (team, customer, feature), and governs them with budgets, guardrails (PII / jailbreak / regex), rate limits, and prompt versioning.

A gateway is not a vector DB, a RAG framework (LangChain, LlamaIndex), a fine-tuning platform, or an evaluation harness in isolation. It is the chokepoint where governance, cost, and routing decisions live regardless of upstream model. Swapping model A for model B is now a quarterly event; re-plumbing each application does not scale. Anthropic, OpenAI, Azure OpenAI, and Vertex publish broadly compatible chat APIs, but per-request shape, key rotation, and quota stories differ — exactly the wedge a gateway closes.


The eight evaluation axes

These are the eight axes that determine whether a gateway will hold up in production for the next 18 months. We have ordered them by the rate at which buyers we talk to discover they got them wrong post-switch.

1. Routing & fallback (table stakes)

Expect: model alias resolution, weighted A/B splits, failover chains by latency or 5xx, per-request model override, streaming. The gateway should expose an OpenAI-shaped /chat/completions and an Anthropic-shaped /messages endpoint so existing SDK code works with a 2-line change. Ask: show me a UI that shifts 5% of gpt-5.5 traffic to claude-sonnet-4-5-20250929 without a redeploy.

2. Guardrails (PII / jailbreak / regex)

PII redaction, jailbreak detection, custom regex denylists, prompt-injection mitigation. Often the first feature gated behind a paid tier — note where each vendor places this on its pricing page. Ask: can I write a custom regex guardrail today on the tier I am evaluating, without a sales call?

3. Per-team / per-customer budgets + virtual keys

Scoped API keys (per team, per environment, per customer) with hard or soft monthly caps; alert at 80%, enforced cutoff at 100%. Critical for multi-tenant SaaS — without it, one runaway tenant eats the month's budget. Ask: can I issue 50 virtual keys, each with a $200/month cap, and see real-time spend per key?

4. Eval pipelines

Run a fixed prompt set across N models on a schedule, compare output quality (LLM-as-judge, exact match, semantic similarity), gate upgrades on regressions. Often enterprise-only across competitors. Ask: show me an eval comparing gpt-5.5 vs claude-sonnet-4-5-20250929 vs claude-haiku-4-5 on my own dataset, with last-run delta visible to engineers.

5. Prompt management & versioning

Centralized prompt registry, versioned, with diffs, environment promotion (dev → staging → prod), and rollback without a code change. Ask: can a non-engineer (PM, CS) edit a production prompt safely, with audit trail and one-click rollback?

6. Observability & attribution

Per-request, per-team, per-customer cost and latency rollups; export to your warehouse (BigQuery, Snowflake, S3) or to LangSmith / Helicone / your own SIEM. Ask: show me $/customer for the last 30 days, broken down by feature.

7. Cost controls + provider economics

The lever buyers most underestimate. A gateway's aggregated volume can justify reserved capacity and dedicated throughput commitments with the model providers that no single mid-market customer can sustain alone. Annual reservations on Azure PTU are documented at up to 70% off PAYG; monthly up to 30%. Full math: the cost teardown post. Ask: do you pool reservations across customers, and how does that show up on my invoice?

8. Compliance posture (RLS, RBAC, residency, audit)

Multi-tenant isolation at the database (Postgres RLS, not app-layer only), SSO/SAML, regional residency (US / EU / customer-private), per-request audit log, SOC 2 / GDPR. Ask: show me the RLS policy that prevents tenant A from reading tenant B's prompt history; show me the residency knob.


The 2026 shortlist matrix

Four gateways come up most in 2026 buyer conversations: OpenRouter, Portkey, Helicone, and nRouter. All have public pricing pages; the rows below are sourced from those pages, with audit timestamps in the Sources section. We are explicit about where features sit on the pricing ladder — that is the whole point of this guide.

AxisOpenRouterPortkeyHeliconenRouter
1. Routing & fallbackRouting-first productYes, Pro+Limited; observability-firstYes, every tier
2. Guardrails (PII/jailbreak/regex)Not offered as a product featurePro / EnterprisePro / EnterpriseIncluded, every plan
3. Per-team / per-customer budgetsNot offeredEnterpriseNot offeredIncluded, every plan
4. Eval pipelinesNot offeredEnterprisePro / EnterpriseIncluded, every plan
5. Prompt management / versioningNot offeredPro / EnterpriseLimitedIncluded, every plan
6. Observability & attributionSpend dashboard onlyYes, Pro+Yes, primary productYes, every tier
7. 0% platform fee on a self-serve plan5.5% credit feeAnnual contract / salesAnnual contract / salesNo — 4% of the credits on every plan, no minimum fee
8. Reserved provider capacityNot documentedNot documentedNot at the gateway layerYes, post-$10k ARR

OpenRouter, Portkey, and Helicone are trademarks of their respective owners. nRouter is not affiliated with or endorsed by any of these vendors. Every row is sourced from each vendor's public pricing or documentation page on 2026-05-16; if any has changed, email us and we'll update.

The pattern across the matrix is the same one we wrote up in the OpenRouter alternative comparison: governance features (rows 2–5) are gated by every competitor and included on every nRouter plan. The competing wedge is feature-gating — ours is platform-fee-gating, with every feature on every plan.


Decision tree: which gateway for which buyer

Three buyer profiles cover roughly 90% of evaluations we see. Map yourself to one and use the decision below as a starting point — then validate with your own list of must-haves.

Buyer A — Indie / hobbyist / single-developer prototype

LLM spend under $500/month, no compliance, no team. Any of these works; nRouter Pay as you go (platform fee of 4% of the credits, no subscription) is a strict superset of OpenRouter (5.5% credit fee) — same OpenAI shape, lower fee, with the option to grow into per-team budgets without re-platforming. Stop reading; go to app.nrouter.ai/signup.

Buyer B — Mid-market SaaS adding AI (ICP 2)

$5k–$50k/month spend, 50–500 engineers, compliance asking for per-team / per-customer attribution and budget predictability. nRouter (4% platform fee on every plan, all features on every plan) fits: a subscription from $20/mo adds a monthly nrouter/auto allowance and higher rate limits, not a lower fee, while keeping the same feature set as every other plan. Portkey and Helicone solve observability, but their Pro / Enterprise gating on guardrails, evals, and budgets requires multi-year contract negotiation nRouter does not.

Buyer C — Platform / infra eng at AI-native startup (ICP 3)

$10k–$200k+/month, multiple product lines, multiple regions; internal platform team that wants per-team budgets, guardrails, and reservation economics without running its own infrastructure. nRouter is the managed path: every governance feature on every plan, plus the reservation-pooling lever. That lever is one no self-hosted gateway can replicate without aggregating customers across companies — a gateway-provider economic, not a feature flag. Steady spend → reservations: nRouter Pro / Enterprise.


The 90-minute evaluation

You do not need a 2-week vendor cycle. Run this:

  1. (15 min) Mark the eight axes above as MUST, SHOULD, or DEFER. Be honest — SOC 2 is a MUST for healthcare; for a 4-person startup it is a DEFER.
  2. (20 min) For each gateway, find the tier where every MUST first appears (pricing page or docs). If unclear inside 5 minutes, mark "unknown" and assume worst case.
  3. (20 min) For each gateway, write annual cost at the lowest tier covering all MUSTs (12 × monthly minimum + per-seat + platform-fee-on-$X). That is the real annual cost in your scenario, not the headline.
  4. (15 min) Run a 5-prompt smoke test on the two cheapest gateways that cover the MUSTs. Same prompt, model, temperature. Compare latency p50/p95, error rate, and any guardrail wired up in the trial.
  5. (20 min) Read the What this is not section of the cost teardown. Decide whether the 60% / 70% reservation lever is actually available to you (annualized spend > $10k/month, willing to prepay) or whether you should stay PAYG. Either answer is fine; the wrong answer is pretending you are somewhere on the curve you aren't.

If you want help running this on your own numbers, that is what the 30-minute walk-through on /community is for.


Switching cost: the two-line diff

The migration into nRouter (or any OpenAI-shaped gateway) is two lines for an existing OpenAI integration:

# Before — direct to OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

# After — through nRouter
client = OpenAI(
    api_key=os.environ["NROUTER_API_KEY"],
    base_url="https://api.nrouter.ai/v1",
)

The Anthropic SDK accepts the same shape via its base_url parameter. Typical scoping: half a day for one engineer plus 24–72 hours of dual-run. Add it to your spreadsheet alongside the annual platform fee in step 3 above.


What this is not

A few buyers should not switch to a gateway at all in 2026:

  • Single-provider committed accounts with negotiated direct-deal discounts that beat reservation pricing — rare above $1M/yr but possible.
  • FedRAMP High / IL5+ workloads where the gateway provider is not on the relevant authorization boundary. Direct endpoints inside the boundary, gateway outside.
  • Sub-$500/month spend with zero compliance load — any of the five works; pick the cleanest signup and move on.

We would rather you skip the switch than switch on a math error.


How nRouter is different in one paragraph

Every other gateway in the matrix uses feature gating as its monetization wedge: guardrails on Pro, evals on Enterprise, budgets on call-sales. We use platform-fee gating instead — the same complete feature set on every plan (Pay as you go, Starter $20/mo, Pro $50/mo, Max $200/mo, Enterprise custom), with one fee we add on top of provider cost on every self-serve plan (4% of the credits at purchase, no minimum fee). We then commit to reserved provider capacity once aggregated revenue clears the provider's commitment minimum. That is the entire wedge, and it holds on the $5 account exactly as it holds on the enterprise contract — because if it stops being true, we have failed.


Frequently asked questions

How do enterprise buyers evaluate an LLM gateway platform?

Enterprise procurement teams evaluate LLM gateways across six core criteria: multi-model access through a single OpenAI-compatible contract, provider failover and intelligent routing, zero-markup token pricing (flat list price pass-through), enterprise security (SSO, SAML, SCIM, RBAC), append-only audit logs with write-time PII redaction, and flexible deployment models (managed, dedicated tenancy, or customer VPC).

Does an LLM gateway charge markups on upstream model tokens?

Traditional aggregators often mark up token rates by 5% to 20% over upstream provider prices. nRouter charges zero per-token markup: all usage settles at the exact published list price of the serving provider (OpenAI, Anthropic, Google Vertex, AWS Bedrock). A transparent 4% platform fee is applied to credit purchases, keeping unit economics predictable.

What enterprise security and compliance features does nRouter support?

nRouter includes enterprise-grade guardrails, prompt-injection defense, write-time PII redaction, role-based access control (RBAC), virtual key scoping with per-key and per-team budget ceilings, and immutable audit logs. SOC 2 Type II compliance is actively in progress. Organizations requiring custom security reviews or BAAs can engage through nRouter Enterprise.

Can enterprise teams deploy nRouter in their own VPC or BYO-cloud?

Yes. In addition to managed multi-tenant and dedicated cloud tenancy, enterprise customers can deploy nRouter within their own cloud VPC (AWS, Azure, GCP) or bring their own provider enterprise agreements. On-premise and air-gapped deployments are supported for scoped technical pilots.


Try it on your own numbers

Load the $5 minimum in credits, with the platform fee on top. Enough to route 5–10 production prompts on Pay as you go, exercise the guardrails + per-team budgets UI on your real traffic, and decide whether a subscription's monthly nrouter/auto allowance is worth it before you sign anything.

→ Get started at app.nrouter.ai/signup — Pay as you go from $5, with the platform fee on top. No subscription. Mid-market SaaS or larger? Bring your last 90 days of LLM invoices to a 30-min walk-through (book through /community) and we'll redo the cost teardown against your actual spend. For dedicated cloud tenancy, VPC deployments, or procurement documentation, explore nRouter Enterprise or browse nRouter.ai.


See also

Sources

Vendor pricing/docs pages, verified 2026-05-16:

Share
Written by nRouter teamEngineering, product, and company posts from the nRouter team — code-first, cost-honest, no vendor-marketing fluff.