Enterprise LLM Gateway

One API. Every model.Ready for production agents.

Text, code, voice, images, and video — with sub-millisecond routing and enterprise guardrails.

Get an API Key
import { nRouter } from "@nrouter_ai/sdk"

const client = new nRouter({
  model: "gpt-5",
})

const stream = await client.chat.completions.create({
  model: "gpt-5",
  messages: [
    { role: "system", content: "You are nRouter's support agent." },
    { role: "user", content: "Where is my order #4182?" }
  ],
  stream: true,
})

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || "")
}

console.log("Done")

Your AI stack, unified.

Scrolling pauses while you hover over or focus this strip.

Architecture

Every request runs through one gateway.

Every call clears the AI Security layer first: WAF, DDoS shielding, rate limits, and model ACLs. Then your routing policies apply.

Your apps
  • Chat apps
  • Agents
  • RAG
  • SDKs
sk-nrouter-••••one key
AI Security
  • Edge WAF
  • DDoS Shield
  • Rate limits
  • Model ACL
nRouter.ai
in path
  • GuardrailsAir-gapped Cortex: $0 injection spend
  • Smart routingYour routing strategy
  • Spending limitsEnforced credit controls
  • Response cacheDedupe repeat prompts
  • FallbacksRetry across providers
  • ObservabilityLogs, traces, spend
Every provider
  • Azure FoundryAzure Foundry
  • Alibaba USAlibaba US
  • Google Vertex AIGoogle Vertex AI
  • AnthropicAnthropic
  • OpenAIOpenAI
Multi-provider0% token markup · exact list price

Your apps use one nRouter key. We manage the provider credentials.

Platform

Everything you need. Nothing you don't

Gateway, guardrails, routing, and budgets. One platform for production AI.

Low-latency smart routing across multi-cloud models with zero-downtime provider fallback and deterministic A/B splits.

  • Fallback chains
  • Model aliases

SDKs & Frameworks

Use the stack you already know.

Install a published nRouter SDK or keep your existing OpenAI-compatible client and change the base URL.

Python

Published on PyPI

pip install nrouter-sdk

TypeScript

Published on npm

npm install @nrouter_ai/sdk

Swift / iOS

Swift Package Manager

.package(url: "https://github.com/nRouterGateway/nrouter-sdk.git", from: "3.1.2")

Go

Published on pkg.go.dev

go get github.com/nRouterGateway/nrouter-sdk/sdks/go/v3@v3.1.2

Java

Published on Maven Central

implementation("ai.nrouter:nrouter-sdk:3.1.2")
K

Kotlin

Published on Maven Central

implementation("ai.nrouter:nrouter-sdk-kotlin:3.1.2")

Android

Published on Maven Central

implementation("ai.nrouter:nrouter-sdk-android:3.1.2")

REST API

OpenAI-compatible HTTP

https://api.nrouter.ai/v1/chat/completions

R

Published on R-universe

install.packages("nrouter", repos = "https://nrouterai.r-universe.dev")

Model Marketplace

Frontier models, one developer API.

High-density gateway specifications. Compare context windows, edge TTFT latency, multi-cloud failover wires, and exact list-price costs with 0% token markup.

9 models displayed
Showing 3 models at a time • Scroll right to view more frontier models
QwenQwen
Frontier

qwen-mt-turbo

dashscope/qwen-mt-turbo

Qwen MT Turbo is a specialized translation model supporting 92 languages based on the Qwen3 architecture.

Context128K
Input / 1M$0.25
Output / 1M$2.00
QwenQwen
Frontier

qwen-plus-latest

dashscope/qwen-plus-latest

Qwen Plus Latest is the rolling production pointer to Alibaba latest Qwen Plus foundation model checkpoint.

Context128K
Input / 1M$1.20
Output / 1M$3.60
OpenAIOpenAI
Frontier

whisper-1

openai/whisper-1

Robust automatic speech recognition model from OpenAI supporting multilingual speech transcription, translation, and language identification.

Context128K
Input / 1M$0.00
Output / 1M$0.00
OpenAIOpenAI
Frontier

o1

openai/o1

Full-scale OpenAI reasoning model trained with reinforcement learning to generate detailed chains of thought for STEM, logic, and coding problems.

Context128K
Input / 1M$15.00
Output / 1M$60.00
AnthropicAnthropic
Frontier

claude-opus-5-5

anthropic/claude-opus-5-5

claude-opus-5-5 with 1000k context window.

Context1M
Input / 1M$5.00
Output / 1M$25.00
OpenAIOpenAI
Frontier

o3-mini

openai/o3-mini

Cost-effective OpenAI reasoning model optimized for programming, mathematics, and structured logic queries with customizable thinking intensity.

Context128K
Input / 1M$1.10
Output / 1M$4.40
Alibaba USAlibaba US
Frontier

qwq-plus

dashscope/qwq-plus

QwQ-Plus is an Alibaba Cloud reasoning model built on Qwen with Questions, featuring a 131,072-token context window and test-time compute. It scores 90.6 on MATH-500, 65.2 on GPQA, and 50.0 on AIME for mathematics and science proofs.

Context128K
Input / 1M$0.80
Output / 1M$2.40
QwenQwen
Frontier

qwen-plus-2025-01-25

dashscope/qwen-plus-2025-01-25

Qwen-Plus-2025-01-25 is a dated snapshot of Alibaba Cloud Qwen-Plus model, providing balanced speed and reasoning across a 1M-token context window with support for function calling and structured outputs.

Context128K
Input / 1M$1.20
Output / 1M$3.60
DeepSeekDeepSeek
Frontier

deepseek-v4-pro-0813

dashscope/deepseek-v4-pro-0813

DeepSeek-V4-Pro-0813 is the production release of the DeepSeek-V4-Pro architecture, pairing 1.6T MoE parameters with a DSpark speculative decoding engine for accelerated agentic execution, full-stack software engineering, and deep logical reasoning.

Context128K
Input / 1M$1.60
Output / 1M$6.40

Customer Stories

Trusted by today's AI engineering leaders

See how teams run production workloads across leading AI models with sub-millisecond routing, automated failover, and strict budget caps.

NBA Sports TechNBA Sports Tech scales AI across peak live playoffs with sub-2ms routing compute.
< 2ms
Internal routing latency
0
Dropped requests at peak
100%
Automated cross-cloud failover

nRouter handles our peak playoff traffic with automated cross-cloud failover and sub-2ms internal routing compute. Zero dropped fan requests during game-ending buzzer beaters.

Marcus Vance · VP of Platform Architecture · NBA Sports Tech

Live basketball arena used by NBA Sports Tech

Explore how engineering teams run production workloads across certified models on nRouter.

See more customer stories