Use Case · RAG

Retrieval-augmented generation, one key for embed and chat.

A RAG pipeline calls two model families: embeddings to index and query, a chat model to synthesize the answer. Route both through one nRouter endpoint, cache the repeats, and track cost per stage.

rag-pipeline · request trace

One pipeline, two model families

Index embedgemini-embedding
Query embedgemini-embedding
Synthesisgemini-2.5-flash
Cachehit · 0 ms
Stages taggedindex · query
Provider confignone
embed + chatcachedcost-tracked
RAG cost savings
50–70%

Via context caching & smart tier routing

Synthesis uptime
99.99%

Cross-cloud fallback for embeddings & chat

Gateway overhead
95 ms

p50 added proxy latency in native Rust

Model catalog
169+

models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic

Why nRouter for RAG

The four things a RAG pipeline needs

Two model families, repetitive traffic, real cost pressure, and providers that occasionally fail. nRouter handles all four behind one key.

Embeddings and chat, one key

The same endpoint serves the embeddings call that builds your index and the chat call that synthesizes the answer. Zero provider config.

Caching for repetitive traffic

FAQs re-asked, identical context windows. Exact-match repeats are served from cache and skip the provider call entirely. Override with nrouter_cache: false.

Cost per pipeline stage

Tag index vs. query traffic in request metadata and the dashboard attributes real per-call cost to each stage of the pipeline.

Failover keeps retrieval answering

A degraded embedding or chat provider triggers the next fallback link transparently. Your index build finishes and the query path keeps answering.

How it works

A RAG request, end to end

Index once, then query: embed the question, retrieve context from your own vector store, and synthesize with a chat model. nRouter sits on the two LLM hops (embeddings and synthesis) and logs the cost of each.

RAG pipeline flow

  1. Query & Context Ingress

    inbound RAG turn

    User query plus retrieved chunks from PGVector, Pinecone, or Qdrant.

  2. Gateway & Pipeline Auth

    In-Memory RLS

    Embed & chat virtual key validation, stage budget isolation, and rate pacing.

  3. RAG Guardrail Filter

    Context & Prompt Defense

    Scans retrieved documents and user queries to block indirect prompt injections.

  4. Smart Router & Cache

    Semantic Vector Cache

    50–70% cost reduction; repeat context hits return in <15ms without LLM charges.

  5. Synthesis Providers

    Claude · GPT-4o · Gemini

    99.99% multi-provider synthesis failover with streaming token delivery.

Nearest-neighbour search stays in your database. nRouter handles the two LLM hops (embeddings and synthesis) with caching, failover, and per-stage cost & usage tracking.

The code

Same client for embeddings and chat

A RAG pipeline is just two endpoint calls against one key. These snippets come straight from the SDK examples the playground and dashboard use. Set NROUTER_API_KEY and the chat call runs as-is; the embeddings call uses the same client and base URL.

Installpip install openai
1# Cache: enabled (org default). Pass nrouter_cache: false to skip.
2from openai import OpenAI
3import os
4
5client = OpenAI(
6 api_key=os.environ["NROUTER_API_KEY"],
7 base_url="https://api.nrouter.ai/v1",
8)
9
10response = client.chat.completions.create(
11 model="gpt-5.4-mini",
12 temperature=1,
13 max_completion_tokens=1024,
14 messages=[
15 {"role": "user", "content": "Hello! What models do you support?"},
16 ],
17 extra_body={
18 # "nrouter_cache": False, # Uncomment to skip cache
19 },
20)
21
22print(response.choices[0].message.content)

The same client object also calls client.embeddings.create() — one key covers the whole pipeline.

FAQ

Common RAG questions

Does nRouter support embedding models alongside chat completions?

Yes. The same OpenAI-compatible gateway serves embeddings (/v1/embeddings) and chat synthesis (/v1/chat/completions). Your corpus indexing pipelines and query-time retrieval both authenticate with one virtual key, eliminating the need to configure separate embedding vendors.

How does semantic caching lower RAG operational costs?

RAG workloads frequently process duplicate questions and overlapping document context windows. nRouter’s response cache serves exact and semantic matches in under 15 milliseconds at $0 provider cost. Caching combined with intelligent model tiering cuts overall RAG infrastructure spend by 50% to 70%.

How do guardrails prevent indirect prompt injection in RAG contexts?

Retrieved third-party documents often contain adversarial instructions, untrusted links, or hidden jailbreaks. nRouter’s inline guardrails screen retrieved context before passing it to the synthesis model, detecting and neutralizing indirect prompt injections without corrupting legitimate facts.

Can I track and attribute cost for each stage of a RAG pipeline?

Yes. nRouter reports exact provider token usage and dollar spend for every call. By tagging indexing requests and query synthesis requests with stage metadata, you can analyze cost-per-search, embedding spend, and generation costs directly in the dashboard.

What happens if an embedding or chat provider goes down during production?

nRouter’s circuit breakers instantly route around degraded endpoints. If your primary embedding or chat model throttles or suffers an outage, requests fail over to backup providers in milliseconds. Your knowledge retrieval pipelines and user-facing answers stay online with 99.99% reliability.

One key for the whole pipeline

Ship a RAG pipeline without juggling providers

Embeddings, chat, caching, and per-stage cost & usage tracking. All behind one nRouter key. Every feature is unlocked on every plan.

Building autonomous workflows on top of retrieval? See the AI agents use case.