Pick the model per task
Fast model for inline autocomplete, strong reasoning model for refactors and test generation. Set the model per request from one catalog, no SDK swap.
A coding tool mixes fast autocomplete with deep reasoning. nRouter lets you pick the model per task from one catalog, route latency-sensitive calls to the quickest endpoint, and cap spend per developer seat.
One catalog, the right model per task
models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic
Via smart tier routing for completions
p50 added latency in native Rust
Multi-provider failover across clouds
A coding assistant has two competing needs (speed for completion, depth for reasoning) and cost that scales with every seat. nRouter handles all three behind one key.
Fast model for inline autocomplete, strong reasoning model for refactors and test generation. Set the model per request from one catalog, no SDK swap.
A developer is waiting on every keystroke. Latency-based routing steers each request to the model with the lowest recent p95 for your org; streaming runs natively token-by-token.
One virtual key per developer with a hard 402 ceiling. Cost-per-seat is a number, not a guess, and a per-key RPM/TPM cap stops one seat hogging the org’s shared plan capacity.
When a suggestion is wrong or slow, filter the request log by the developer’s key: model, latency, tokens, and real cost per call. A/B test two models on real traffic.
Each developer carries a seat key. Completion and reasoning requests route to the model the task needs, stream back token-by-token, and land in the log attributed to that seat.
Code-assistant request flow
Editor Request
POST /v1/chat/completions
Keystroke completions and IDE refactoring prompts.
Gateway & Seat Auth
:4000 · In-Memory RLS
Per-seat virtual key validation and developer budget caps.
Code Guardrail Filter
Secret & License Scan
Inline scanning redacts hardcoded API keys and corporate credentials.
Smart Tier Router
Latency & Task Strategy
40–70% cost reduction steering autocomplete to fast models vs reasoning.
Model Providers
Claude · GPT-4o · Gemini
99.99% multi-provider failover with zero-buffering native token streaming.
The gateway adds about 95 ms at p50 — LLM inference is the dominant latency factor. Streaming runs natively with zero hot-path buffering.
A coding assistant just sets the model field per call — fast for completion, strong for reasoning. These snippets come from the same SDK examples the playground uses; change the model string and the catalog does the rest.
pip install openai| 1 | # Cache: enabled (org default). Pass nrouter_cache: false to skip. |
| 2 | from openai import OpenAI |
| 3 | import os |
| 4 | |
| 5 | client = OpenAI( |
| 6 | api_key=os.environ["NROUTER_API_KEY"], |
| 7 | base_url="https://api.nrouter.ai/v1", |
| 8 | ) |
| 9 | |
| 10 | response = client.chat.completions.create( |
| 11 | model="gpt-5.4-mini", |
| 12 | temperature=1, |
| 13 | max_completion_tokens=1024, |
| 14 | messages=[ |
| 15 | {"role": "user", "content": "Hello! What models do you support?"}, |
| 16 | ], |
| 17 | extra_body={ |
| 18 | # "nrouter_cache": False, # Uncomment to skip cache |
| 19 | }, |
| 20 | ) |
| 21 | |
| 22 | print(response.choices[0].message.content) |
One key reaches every model in the catalog — no per-model provider account to manage.
Yes. The model catalog exposes every model behind one key. Set the model per request (a fast model for inline completion, a stronger reasoning model for refactors or test generation) with no separate provider account or SDK swap. Smart tier routing can also automate this selection based on prompt complexity, saving 40% to 70% in compute costs.
Routing decisions happen in-memory in native Rust and add under a millisecond (<1ms). LLM inference time dominates. Latency-based routing steers each request to the model with the lowest recent p95 latency for your organization, and streaming responses are dispatched with zero hot-path buffering.
Yes. Issue one virtual key per developer or per engineering team and set a hard budget ceiling. The moment spend would exceed the cap, the gateway returns HTTP 402. Spend, rate limits, and the request log all scope to that key, so engineering leaders have total visibility into cost per seat.
Before prompts reach commercial foundation models, inline guardrail filters inspect code payloads for embedded API keys, private certificates, database connection strings, and sensitive credentials. Identified secrets are redacted or blocked in-flight, preventing accidental leakage into external provider logs.
Yes. nRouter is fully OpenAI-compatible and delivers streaming responses with zero buffering, so editor extensions (VS Code, JetBrains, Cursor, Neovim) get tokens as they are generated. Multi-provider failover transparently reroutes to an alternate model provider if the primary provider throttles or fails mid-session.
Model choice, low latency, per-seat budgets
Pick the model per task, route for latency, and cap spend per developer — all unlocked on every plan.
Explore other use cases