Browse documentation

Router Settings

Configure intelligent routing rules, model weights, automated failovers, and caching policies in nRouter to optimize latency, cost, and reliability goals.

Last updated

Router Settings (/router-settings) creates organization-scoped Smart Router aliases. It is writable by Owner and Organization Admin roles; other members can read the configuration.

Routing strategy

Point a router alias at a set of candidate models and choose how each request resolves to one of them:

StrategyBehavior
PriorityTries candidates in the configured order — ordered failover straight down the chain; it is the default, and the strategy every other one falls closed to when its own signal is missing or incomplete
CostRoutes to the cheapest candidate, ranked on the sum of the input and output list price per million tokens
LatencyOrders candidates by the gateway-observed latency EWMA; falls closed to Priority until every candidate has a signal
WeightedUses request-ID-seeded weighted rendezvous selection across candidates with positive weights

Each alias is per-org and resolves at request time — change the set or the strategy in the dashboard, no redeploy. A concrete model you name directly is never re-routed (routing is opt-in).

What the Cost strategy ranks on. It is the sum of the input and output price per million tokens, with no assumption about the shape of your prompt — a candidate with cheap input and expensive output is ranked on both together. The price is read live from the served catalogue at request time, so a catalogue price change takes effect on the next request rather than when you next edit the router. The price saved alongside the router when you created it is only a fallback, used for a candidate the live catalogue no longer prices.

A Smart Router alias is not the same thing as a catalogue alias in the model catalog, even though both go in the model field. A catalogue alias is a friendly name for one model; a Smart Router alias is a name you create here that resolves across a set of candidates by the strategy you picked.

Calling your router

Creating the router is half the job. The other half is one line of client code: put the Smart Router alias in the model field, exactly where you would otherwise name a model.

curl -X POST https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "my-router-alias",
    "messages": [{ "role": "user", "content": "Summarize this incident report." }]
  }'
import { nRouter } from "@nrouter_ai/sdk";

const client = new nRouter(); // reads NROUTER_API_KEY

const completion = await client.chat.completions.create({
  model: "my-router-alias",
  messages: [{ role: "user", content: "Summarize this incident report." }],
});

Nothing else changes: the same key, the same request shape, the same response shape. Until a call names the alias, a Smart Router is configured but never exercised — a router no request points at does nothing at all. Check x-nr-model on the response to see which candidate actually served it.

Fallback chains

The selected candidates form an addressable deployment chain. A later candidate is attempted only when retrying is known not to risk a second model bill: an upstream admission refusal (429, 503, or Anthropic 529) or a connect-phase transport failure. Other statuses and failures after a connection may already have generated billable work, so they fail loudly instead of being replayed.

Reorder by priority (drag or arrow keys) and delete chains you no longer need. Duplicate primaries are rejected.

There is one chain per Smart Router alias, not one chain per failure type. A concrete model named directly follows its direct route and does not inherit a hidden platform fallback.

Fallback models for a single request

The chain above is your standing configuration. A single call can name its own fallback order instead, by putting nrouter_fallbacks in the JSON request body:

{
  "model": "gpt-5.4-mini",
  "messages": [{ "role": "user", "content": "Summarize this incident report." }],
  "nrouter_fallbacks": ["claude-sonnet-4-5", "gemini-2.5-pro"]
}

Live. nrouter_fallbacks is available on api.nrouter.ai as of the 2026-09-18 gateway release.

  • The primary stays model — the list only says what is tried after it.
  • For that one call the list replaces whatever fallback policy your organization has for the alias, rather than extending it. Your saved configuration is untouched and applies again on the next request that names none.
  • A Smart Router alias keeps its own chain. Name one in model and the chain configured above is what runs; a request that also carries nrouter_fallbacks is refused with 400 fallback_not_allowed instead of being served off a chain it did not ask for. The same holds for the allowance models, whose plan is the platform's. Per-request lists on top of a stored chain are a v1 limitation, and the refusal is the point: silently ignoring the list would be worse.
  • At most four entries, trimmed and de-duplicated, none equal to model.
  • Every target must be a model your key can already route to. One it cannot is refused with 400 fallback_not_allowed naming the target — never dropped in silence, because you asked for it by name.
  • Failover conditions are the same narrow set as a configured chain: an upstream 429, 503, Anthropic 529, or a connect-phase failure, and nothing else.
  • The money and throughput budgets do not change: one credit reservation, one rate-limit slot, and at most three provider calls for the whole request, however far it walks.
  • You can see which entry answered. x-nr-routing indexes the chain the request actually walked — the one it named — so direct means model answered and fallback:1 means the first name in your list did. x-nr-attempts reports how many provider calls it took.

Full field reference, including the guardrail and caching fields that travel with it, is in Per-Request Options.

Retries & timeouts

SettingScopeDefault
Maximum provider attemptsper request3 total attempts, including the first
Cumulative retry wait budgetper request20 seconds

On a retry, standard Retry-After plus the millisecond provider variants are honored when they fit inside the remaining wait budget. Request-body retry/timeout overrides are not accepted.

Per-model weights

Assign each candidate a non-negative weight; at least one must be positive. Selection is deterministic for one request ID and distributable across different request IDs. Weights are relative, so 70/30 and 7/3 describe the same ratio.

Response caching

When two requests are byte-identical, nRouter can answer the second one from a short-lived serving cache instead of calling the provider again. You save the round trip. A cache hit is still metered, but at the model's cache-read input rate for the prompt with no output charge — cheaper than the provider call it replaced, never free — so caching saves both latency and money, and the response's x-nr-request-cost shows what the hit actually cost.

Caching is on by default. It applies to identical non-streaming requests on /v1/chat/completions, /v1/completions, /v1/responses and /v1/messages with complete buffered, successfully priced responses, and your organization can switch it off for every request with the caching toggle on the Router Settings page or under Settings → Privacy. Streaming requests are not cached and always go to the provider.

What the cache is, precisely:

  • Tenant-keyed. An entry is keyed to your organization and team, alongside the model, the exact request body, and the guardrail chain that produced it. One organization's completion is never served to another, and a response produced under one guardrail configuration is never replayed under a different one.
  • Short-lived. Entries expire on their own after a few minutes. It is a serving cache, not a record: it is not queryable, not exported, and not part of your request log. Prompt and completion content is never written to the log — see Observability & Logs.

Turning it off for a request

Send nrouter_cache: false in the JSON request body:

{
  "model": "gpt-5.5",
  "messages": [{ "role": "user", "content": "Summarize this contract." }],
  "nrouter_cache": false
}

That request is not answered from the cache and does not fill it — the response never becomes an entry another call could be served from. It is an nRouter control field: the gateway removes it before forwarding, so it never reaches the provider. Most SDKs pass it through an extra_body-style escape hatch; see the snippet for your language under SDKs.

The setting is per request, so you can leave caching on everywhere and opt individual calls out — the ones carrying regulated or user-identifying content, say — without changing anything globally. To keep every call out, switch caching off for the organization in Router Settings instead of relying on each caller to send the field.

What it costs you. Every opted-out request pays full provider latency, every time, including the repeats a cache would have absorbed. It also gives up the cheaper price: a hit is billed at the model's cache-read input rate with no output charge, while an opted-out request pays the full provider price every time.

Reading the response header

Every eligible response says what happened, so you never have to infer it:

x-nr-response-cacheMeaning
hitServed from the cache. No provider call; billed at the model's cache-read input rate with no output charge, never $0.
missWent to the provider. Nothing eligible was cached.
bypassYou sent nrouter_cache: false. Not read from the cache, not written to it.

On a hit, x-nr-response-cache-age gives the age of the entry in seconds, so you can tell a two-second-old replay from a much older one. Both headers sit in the same x-nr-* namespace as x-nr-request-cost and x-nr-request-id.

Which model answered, and after how many tries

Routing is only trustworthy if you can see it happen. Three response headers report it on every served request, so you never have to infer a failover from a latency spike:

HeaderValueMeaning
x-nr-routingdirectThe first entry in the chain answered. Nothing was skipped.
x-nr-routingfallback:<n>The entry at 0-based chain index n answered — fallback:1 is the second model in the chain, fallback:2 the third.
x-nr-attemptsinteger ≥ 1Provider calls this request actually made, counting retries and failovers alike.
x-nr-response-cachehit / miss / bypassWhether the response came from the cache, went to the provider, or skipped the cache because you opted out.

The number after fallback: is the 0-based chain index of the entry that answered, so a chain whose second model answered reports fallback:1, not fallback:2. The first fallback is fallback:1; index 0 is direct.

x-nr-routing and x-nr-attempts are absent on a cache hit and on any refusal. In neither case did a chain entry answer a provider call, so there is nothing to report and nothing is invented. x-nr-model names the physical model that served the request, which is how you audit what an alias resolved to.

A direct call that named a concrete model, made no retry, and went to the provider therefore reads x-nr-routing: direct, x-nr-attempts: 1, x-nr-response-cache: miss.

Saving changes

Creating a Smart Router writes the customer-facing configuration and its gateway deployment rows in one database transaction. Duplicate aliases are rejected rather than silently replacing a live chain. The Rust gateway resolves the database configuration on each request, so no application redeploy is required.

Next steps

FAQ

Who on my team can change router settings?

Router Settings is writable by Owner and Organization Admin. Members and Viewers can review but cannot mutate it. Team roles are separate and do not create organization-wide routing authority.

Do I have to change my application code to use routing?

No. Routing is opt-in: point a router alias at a set of candidate models and call that alias. A concrete model you name directly is never re-routed. You change the model set or strategy in the dashboard and it takes effect at request time — no redeploy and no SDK change.

What's the difference between the Cost, Latency, and Weighted strategies?

Cost orders candidates by their direct catalog input/output prices. Latency uses the gateway-observed latency EWMA and falls closed to priority until every candidate has data. Weighted uses deterministic request-ID-seeded weighted rendezvous selection.

What happens when a model fails mid-request?

Failover requires a configured Smart Router. On that alias, an upstream 429, 503, Anthropic 529, or connect-phase failure can advance to another eligible deployment. Generic 5xx, read timeouts, and other post-connect failures are not replayed because the provider may already have generated and billed for work. A direct model request has no hidden fallback chain. Use the shared x-nr-request-id to correlate all attempts in routing decisions and support logs.

Can one request choose its own fallback models?

Yes, on a direct model. Put nrouter_fallbacks in the JSON request body with up to four model names: for that call the list replaces whatever fallback policy your organization has for the alias, the primary stays whatever is in model, and your saved configuration is untouched for every other call. A Smart Router alias and the allowance models keep their own plan and refuse a per-request list with 400 fallback_not_allowed in v1. A target your key cannot route to is refused with 400 fallback_not_allowed naming it, never dropped silently. See Per-Request Options.

Does a per-request fallback cost me extra credits or rate limit?

No. One request takes one credit reservation and one rate-limit slot however far it walks its chain, and it is capped at three provider calls in total. You are billed for the model that actually answered, reported on x-nr-request-cost.

How do I tell which model in my chain answered?

Read x-nr-routing: direct means the first entry answered, and fallback:<n> means the entry at 0-based chain index n answered — so the first fallback is fallback:1, the second model in the chain. x-nr-attempts reports how many provider calls the request made. Both headers are absent on a cache hit and on a refusal, because no chain entry answered a provider call. x-nr-model names the physical model that served it.

How many times will a request retry, and can I change it?

The gateway permits 3 total attempts, including the first, and caps cumulative retry waiting at 20 seconds. These money-safety limits are not request overrides. A provider Retry-After value is honored only when it fits inside the remaining budget.

How do the per-model weights work?

Weights apply to a Smart Router alias. They are relative non-negative values, at least one must be positive, and selection is deterministically seeded by the request ID.

Does picking the Cost strategy actually lower my bill?

The Cost strategy sends each request to the cheapest eligible model by list price, so you're billed for whichever model actually runs, reported on the x-nr-request-cost response header. Router Settings only decides which model handles a request — your spend ceilings live separately under Budget Controls and are enforced independently.

Is response caching enabled, and how do I turn it off?

Yes — caching is on by default for identical non-streaming requests; streaming requests are not cached. Your organization can switch it off for every request in Router Settings or under Settings → Privacy, and a single call can skip it by putting "nrouter_cache": false in the JSON request body. That field is an nRouter control field and never reaches the provider. An opted-out call is neither answered from the cache nor added to it, and the response comes back with x-nr-response-cache: bypass so you can confirm it.

How do I tell whether a response came from the cache?

Read the x-nr-response-cache response header: hit means it was served from the cache with no provider call (metered at the model's cache-read input rate, no output charge), miss means it went to the provider, and bypass means the request was opted out, by nrouter_cache: false or by your organization's caching toggle. On a hit, x-nr-response-cache-age reports the entry's age in seconds. No header at all means the request was not eligible for the cache.

What is actually held in the cache, and for how long?

A cache entry is a complete, successfully priced response body, keyed to your organization and team along with the model, the exact request body, and the guardrail chain that produced it — so it is only ever replayed for an identical request from the same tenant under the same guardrails. Entries are short-lived and expire on their own after a few minutes. The cache is not queryable, not exported, and separate from your request log, which never carries prompt or completion content.

What do I give up by turning caching off?

Latency and money. An opted-out request always pays the provider round trip and the provider charge, including on repeats an entry would have absorbed for free. Turn it off where you want the guarantee that a response was freshly generated, or where you would rather a body not sit in a serving cache at all, and leave it on everywhere else.

Do my changes take effect the moment I edit them?

The create action is atomic: configuration and deployment candidates either both commit or neither does. The gateway reads the saved Smart Router at request time, with no app or gateway redeploy.

Was this page helpful?