Embeddings and chat, one key
The same endpoint serves the embeddings call that builds your index and the chat call that synthesizes the answer. Zero provider config.
A RAG pipeline calls two model families: embeddings to index and query, a chat model to synthesize the answer. Route both through one nRouter endpoint, cache the repeats, and track cost per stage.
One pipeline, two model families
Via context caching & smart tier routing
Cross-cloud fallback for embeddings & chat
p50 added proxy latency in native Rust
models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic
Two model families, repetitive traffic, real cost pressure, and providers that occasionally fail. nRouter handles all four behind one key.
The same endpoint serves the embeddings call that builds your index and the chat call that synthesizes the answer. Zero provider config.
FAQs re-asked, identical context windows. Exact-match repeats are served from cache and skip the provider call entirely. Override with nrouter_cache: false.
Tag index vs. query traffic in request metadata and the dashboard attributes real per-call cost to each stage of the pipeline.
A degraded embedding or chat provider triggers the next fallback link transparently. Your index build finishes and the query path keeps answering.
Index once, then query: embed the question, retrieve context from your own vector store, and synthesize with a chat model. nRouter sits on the two LLM hops (embeddings and synthesis) and logs the cost of each.
RAG pipeline flow
Query & Context Ingress
inbound RAG turn
User query plus retrieved chunks from PGVector, Pinecone, or Qdrant.
Gateway & Pipeline Auth
In-Memory RLS
Embed & chat virtual key validation, stage budget isolation, and rate pacing.
RAG Guardrail Filter
Context & Prompt Defense
Scans retrieved documents and user queries to block indirect prompt injections.
Smart Router & Cache
Semantic Vector Cache
50–70% cost reduction; repeat context hits return in <15ms without LLM charges.
Synthesis Providers
Claude · GPT-4o · Gemini
99.99% multi-provider synthesis failover with streaming token delivery.
Nearest-neighbour search stays in your database. nRouter handles the two LLM hops (embeddings and synthesis) with caching, failover, and per-stage cost & usage tracking.
A RAG pipeline is just two endpoint calls against one key. These snippets come straight from the SDK examples the playground and dashboard use. Set NROUTER_API_KEY and the chat call runs as-is; the embeddings call uses the same client and base URL.
pip install openai| 1 | # Cache: enabled (org default). Pass nrouter_cache: false to skip. |
| 2 | from openai import OpenAI |
| 3 | import os |
| 4 | |
| 5 | client = OpenAI( |
| 6 | api_key=os.environ["NROUTER_API_KEY"], |
| 7 | base_url="https://api.nrouter.ai/v1", |
| 8 | ) |
| 9 | |
| 10 | response = client.chat.completions.create( |
| 11 | model="gpt-5.4-mini", |
| 12 | temperature=1, |
| 13 | max_completion_tokens=1024, |
| 14 | messages=[ |
| 15 | {"role": "user", "content": "Hello! What models do you support?"}, |
| 16 | ], |
| 17 | extra_body={ |
| 18 | # "nrouter_cache": False, # Uncomment to skip cache |
| 19 | }, |
| 20 | ) |
| 21 | |
| 22 | print(response.choices[0].message.content) |
The same client object also calls client.embeddings.create() — one key covers the whole pipeline.
Yes. The same OpenAI-compatible gateway serves embeddings (/v1/embeddings) and chat synthesis (/v1/chat/completions). Your corpus indexing pipelines and query-time retrieval both authenticate with one virtual key, eliminating the need to configure separate embedding vendors.
RAG workloads frequently process duplicate questions and overlapping document context windows. nRouter’s response cache serves exact and semantic matches in under 15 milliseconds at $0 provider cost. Caching combined with intelligent model tiering cuts overall RAG infrastructure spend by 50% to 70%.
Retrieved third-party documents often contain adversarial instructions, untrusted links, or hidden jailbreaks. nRouter’s inline guardrails screen retrieved context before passing it to the synthesis model, detecting and neutralizing indirect prompt injections without corrupting legitimate facts.
Yes. nRouter reports exact provider token usage and dollar spend for every call. By tagging indexing requests and query synthesis requests with stage metadata, you can analyze cost-per-search, embedding spend, and generation costs directly in the dashboard.
nRouter’s circuit breakers instantly route around degraded endpoints. If your primary embedding or chat model throttles or suffers an outage, requests fail over to backup providers in milliseconds. Your knowledge retrieval pipelines and user-facing answers stay online with 99.99% reliability.
One key for the whole pipeline
Embeddings, chat, caching, and per-stage cost & usage tracking. All behind one nRouter key. Every feature is unlocked on every plan.
Building autonomous workflows on top of retrieval? See the AI agents use case.
Explore other use cases