Solutions

One managed gateway,
tuned to your industry’s rules.

Every team ships on the same gateway: smart routing, built-in guardrails, real-time cost & usage tracking, audit logs. What changes by industry is which controls matter most. Pick yours.

No BYOK · one bill · every control on every plan

  • 169Models liveAlibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic
  • 8Residency regionsUS + EU GA, more on request
  • 99.9%Uptime SLASame SLA on every tier
  • 0Features gatedEvery control on every plan
Compliance & Controls

Enterprise governance across every vertical.

RequirementFinanceHealthLegalRetailGovEdu
Inline PII/PHI Redaction
Zero-Content Data Policy——
Sovereign Data Residency (US/EU)
Per-Matter / Team Budget Caps—
SOC 2 CC7.2 Immutable Audit Logs
Dedicated Tenancy / VPC Link——
Universal Architecture

Cross-industry enterprise LLM gateway architecture

A unified data and control plane that enforces multi-tenancy, inline compliance guardrails, cost-optimized tier routing, and zero-retention provider egress across every workload.

Enterprise request lifecycle: client → gateway → guardrail filter → smart router → model provider

  1. Enterprise Apps & Agents

    Internal Apps / SaaS / AI Agents

    Standard OpenAI-compatible API client endpoints

  2. Unified Rust Gateway

    Sub-millisecond Ingress & RLS Auth

    Key verification, tenant budget isolation, and RPM/TPM enforcement

  3. Universal Guardrail Filter

    PII/PHI Redaction & Threat Defense

    In-flight prompt injection detection and data privacy masking

  4. Multi-Model Smart Router

    Cost / Latency / Intent Routing

    Dynamic fallback and intent-based model tiering

  5. Foundation Model Providers

    OpenAI / Anthropic / Bedrock / Vertex

    Direct zero-retention provider inference with pass-through pricing

40–75%

Cost Reduction

Through semantic caching and smart tier routing

99.99%

Uptime Reliability

Continuous multi-provider failover protection

<1ms

Engine Overhead

Sub-millisecond proxy latency in native Rust

100%

Spend Predictability

Hard-cap budget ceilings returning HTTP 402

Stage 01 — Application Ingress

1. Unified Application Ingress

Internal corporate applications, customer-facing workflows, autonomous coding agents, and departmental microservices connect to one standard OpenAI-compatible API endpoint across HTTP and streaming WebSocket protocols.

Stage 02 — Gateway Preflight & Budgets

2. Sub-Millisecond Gateway & Budget Enforcement

The Rust-powered inference engine verifies organization and virtual key permissions in under 1 millisecond. Real-time token buckets enforce rate limits while strict credit reservations prevent account overdrafts before upstream dispatch.

Stage 03 — Universal Compliance Guardrail

3. Universal Guardrail & Compliance Filter

Every prompt is scrubbed by deterministic pattern engines and semantic classifiers to strip sensitive PII, PHI, financial account identifiers, and API keys. Prompt-injection heuristics halt adversarial payloads before provider egress.

Stage 04 — Dynamic Model Routing

4. Multi-Model Smart Router

Traffic is dynamically assigned to the optimal foundation model based on prompt complexity, latency requirements, and cost profiles. High-throughput queries route to fast micro-tier models, saving 40% to 75% on compute bills.

Stage 05 — Zero-Retention Settlement

5. Zero-Retention Provider Egress & Settlement

Requests terminate at provider endpoints under enterprise zero-retention agreements. Spend is settled down to the micro-cent against exact provider response token usage without hidden token markups.

Unified Integration

Connect any application stack with standard OpenAI SDKs

Drop in nRouter as an API proxy in minutes. Set guardrail profiles, sovereign residency constraints, and fallback model chains via standard client headers.

Installpip install openai
1# Cache: enabled (org default). Pass nrouter_cache: false to skip.
2from openai import OpenAI
3import os
4
5client = OpenAI(
6 api_key=os.environ["NROUTER_API_KEY"],
7 base_url="https://api.nrouter.ai/v1",
8)
9
10response = client.chat.completions.create(
11 model="gpt-5.4-mini",
12 temperature=1,
13 max_completion_tokens=1024,
14 messages=[
15 {"role": "user", "content": "Hello! What models do you support?"},
16 ],
17 extra_body={
18 # "nrouter_cache": False, # Uncomment to skip cache
19 },
20)
21
22print(response.choices[0].message.content)

Header Configuration & Policy Spec

import OpenAI from 'openai';

// Universal enterprise client configuration
const nrouter = new OpenAI({
  apiKey: process.env.NROUTER_ENTERPRISE_KEY,
  baseURL: 'https://api.nrouter.ai/v1',
  defaultHeaders: {
    'x-nr-guardrail-profile': 'enterprise-strict',
    'x-nr-residency': 'us-sovereign', // or 'eu-central'
    'x-nr-zero-retention': 'true',
  }
});

// Smart Router automatically balances cost, latency, and capability
const response = await nrouter.chat.completions.create({
  model: 'nrouter/auto', // Intelligently switches between Haiku/Flash and Sonnet/GPT-4o
  messages: [
    { role: 'system', content: 'Enterprise assistant with zero-trust data boundaries.' },
    { role: 'user', content: 'Synthesize quarterly compliance findings and flag material risks.' }
  ],
  extra_body: {
    fallback_models: ['anthropic/claude-3-5-sonnet', 'openai/gpt-4o', 'google/gemini-2.0-flash'],
    max_cost_per_request_usd: 0.05
  }
});

console.log(response.choices[0].message.content);

Autonomous Multi-Provider Failover

Eliminate single points of failure with millisecond health-check circuits that instantly route around provider outages and rate-limit storms.

Multi-Tenant RBAC & PostgreSQL RLS

Hierarchical organizations, teams, and members governed by Row-Level Security guarantee that no team or cost center can view or leak another tenant’s data.

Zero-Markup List Price Transparency

Never pay arbitrary token markups. You pay exactly the underlying model provider’s published list price plus a transparent flat 4% platform fee on credits.

Common foundation

Smart routing + failover

One OpenAI-compatible endpoint; automatic retry on a backup model when a provider degrades.

Guardrails on every request

PII/PHI redaction, prompt-injection detection, and secret scanning, inline on every plan.

Budgets at three scopes

Hard and soft caps per org, team, and key. A tripped hard cap returns 402 — never a negative balance.

Audit + observability

Append-only audit trail with actor, IP, and diff; request logs export to Langfuse, Datadog, or S3.

Need the deployment, contract, and residency depth? See Enterprise — dedicated tenancy, VPC, SSO, and procurement-ready legal artifacts.

Industry solutions questions, answered

How does nRouter support regulated enterprise industries?

Every industry inherits the same audited gateway core: inline PII and PHI redaction, multi-tenant isolation enforced via database Row Level Security, append-only SOC 2 CC7.2 audit logs, and hardware encryption at rest and in transit. What changes by vertical is which policy toggles and residency constraints are applied.

Can we enforce data residency for EU, US, or specific sovereign regions?

Yes. Enterprise and Pro plans can pin data residency to US or EU regions, guaranteeing that prompt traffic, telemetry, and cached responses never transit outside approved sovereign jurisdictions.

How do guardrails prevent PII and PHI leakage across external model providers?

nRouter guardrails execute inline before any outbound request reaches OpenAI, Anthropic, or Google. Detectors scan for social security numbers, medical record identifiers, credit cards, and API secrets, redacting or rejecting violations before payloads leave our security boundary.

How does unified billing work across multiple AI providers?

You receive one consolidated monthly invoice from nRouter regardless of how many providers (Azure, Vertex AI, AWS Bedrock, OpenAI, Anthropic) your teams consume. Pre-funded credit reservations prevent surprise overdrafts, and per-key ceilings enforce hard caps.

What are the latency overheads and SLA guarantees of the gateway?

The Rust proxy adds less than 1 millisecond of processing latency to non-cached requests. When response cache hits occur, latency drops to sub-15ms.

Enterprise customers receive a guaranteed 99.99% availability SLA backed by contractual remedies.

How does multi-provider failover work during upstream cloud outages?

If an upstream foundation model provider degrades or throws 5xx errors, nRouter detects the threshold in under 200 milliseconds and transparently redirects incoming requests to your configured fallback provider without returning errors to end users.

Solutions · scoped to your stack

Don’t see your exact use case? We will map it.

Tell us how your team uses LLMs and which rules you answer to. We will walk through the routing, guardrail, budget, and residency setup that fits — no over-promised timelines.

No invented certifications · no claims we cannot back