LlamaIndex
Connect LlamaIndex to nRouter for retrieval-augmented generation (RAG) and agent workflows. Configure LLM and embedding endpoints with enterprise guardrails.
Last updated
LlamaIndex is a data framework designed to connect external data sources to large language models for complex retrieval-augmented generation (RAG) and autonomous data agents. By integrating LlamaIndex with nRouter (https://api.nrouter.ai/v1), data pipelines inherit unified model access, real-time cost attribution, semantic query caching, and pre-execution guardrails across both generation and embedding endpoints.
Because LlamaIndex's OpenAI and OpenAIEmbedding classes adhere to standard REST contracts, you simply configure the gateway endpoint and virtual key. Guardrail policies configured in your nRouter dashboard automatically scan incoming queries for prompt injection and sensitive data before retrieval or synthesis runs.
Prerequisites & Installation
LlamaIndex requires Python 3.10 or higher. For reliable dependency resolution, install packages inside an active virtual environment.
Install the modular LlamaIndex OpenAI packages along with the core library:
pip install llama-index-core llama-index-llms-openai llama-index-embeddings-openai nrouter-sdkSetup & Configuration
Store your nRouter virtual key in your environment:
export NROUTER_API_KEY="sk-nrouter-your-virtual-key"Global Settings Configuration
In LlamaIndex, Settings provides global defaults for all query engines, indexes, and agents:
import os
from llama_index.core import Settings
from llama_index.llms.openai import OpenAI
from llama_index.embeddings.openai import OpenAIEmbedding
NROUTER_KEY = os.environ["NROUTER_API_KEY"]
NROUTER_BASE = "https://api.nrouter.ai/v1"
# Configure global generation model
Settings.llm = OpenAI(
model="claude-sonnet-4-5-20250929",
api_key=NROUTER_KEY,
api_base=NROUTER_BASE,
timeout=60.0,
max_retries=2,
)
# Configure global embedding model
Settings.embed_model = OpenAIEmbedding(
model_name="text-embedding-3-small",
api_key=NROUTER_KEY,
api_base=NROUTER_BASE,
timeout=30.0,
)Configuration Parameters
Configure client settings to match your latency and reliability requirements:
| Parameter | Type | Default | Description |
|---|---|---|---|
api_base | str | https://api.nrouter.ai/v1 | Unified gateway base URL. Must include /v1. |
api_key | str | None | nRouter virtual API key (sk-nrouter-...). |
model | str | Required | Model identifier, custom alias, or fallback list. |
timeout | float | 60.0 | Maximum HTTP connection and generation timeout in seconds. |
max_retries | int | 2 | Client-side retry limit. nRouter handles upstream retries automatically. |
default_headers | dict | {} | HTTP headers sent with every request, including x-nr-routing. |
additional_kwargs | dict | {} | Gateway overrides such as nrouter_cache or template settings. |
import os
from llama_index.llms.openai import OpenAI
# Production client configuration with latency optimization
llm = OpenAI(
model="gpt-5.4-mini",
api_key=os.environ["NROUTER_API_KEY"],
api_base="https://api.nrouter.ai/v1",
timeout=45.0,
max_retries=3,
default_headers={
"x-nr-routing": "latency",
},
additional_kwargs={
"nrouter_cache": True,
},
)Implementation Patterns
1. Basic Chat & Generation
Interact with the configured model directly:
from llama_index.core.llms import ChatMessage
messages = [
ChatMessage(role="system", content="You are a senior data architect."),
ChatMessage(role="user", content="Explain vector embedding quantization."),
]
response = Settings.llm.chat(messages)
print(response.message.content)2. Document Indexing & RAG Query Engine
Construct an in-memory vector index and run guarded queries:
from llama_index.core import VectorStoreIndex, Document
documents = [
Document(text="nRouter routes inference across OpenAI, Anthropic, Bedrock, and open-source models."),
Document(text="Virtual keys enforce budget limits and rate limits per team or project."),
Document(text="Server-side guardrails block prompt injections before requests hit downstream models."),
]
# Build index using Settings.embed_model
index = VectorStoreIndex.from_documents(documents)
# Create query engine using Settings.llm
query_engine = index.as_query_engine()
response = query_engine.query("How does nRouter protect against prompt injection?")
print(response)3. Streaming Query Engine
Deliver low time-to-first-token by streaming response tokens:
query_engine = index.as_query_engine(streaming=True)
streaming_response = query_engine.query("What are the key benefits of virtual keys?")
for text in streaming_response.response_gen:
print(text, end="", flush=True)
print()4. Per-Request Gateway Overrides
Attach prompt template identifiers or bypass semantic cache using additional_kwargs:
from llama_index.llms.openai import OpenAI
custom_llm = OpenAI(
model="gpt-5.5",
api_key=os.environ["NROUTER_API_KEY"],
api_base="https://api.nrouter.ai/v1",
additional_kwargs={
"nrouter_prompt_template_id": "tmpl_rag_synthesis_v2",
"nrouter_prompt_variables": {"audience": "executives"},
"nrouter_cache": False,
},
)Production Best Practices
Deterministic Routing & Fallbacks
Prevent service downtime by providing fallback model lists and setting deterministic routing headers:
Settings.llm = OpenAI(
# Primary model with automatic fallback
model="gpt-5.4-mini,claude-haiku-4-5-20251001",
api_key=os.environ["NROUTER_API_KEY"],
api_base="https://api.nrouter.ai/v1",
default_headers={
"x-nr-routing": "latency",
},
)x-nr-routing: latency: Routes queries to the lowest-latency operational provider deployment.x-nr-routing: cost: Directs traffic to the lowest-cost deployment matching the request.- Model Fallback List: If the primary provider experiences capacity constraints or elevated 5xx errors, nRouter seamlessly transfers the query to the fallback model.
Telemetry & FinOps Tracking
Every call through nRouter produces metadata headers:
x-nr-request-id: Trace ID linking application logs with nRouter spend ledgers.x-nr-model: Actual model deployment executing the request.x-nr-cost-status: Indication of pricing certainty (exactorunpriced).x-nr-request-cost: USD spend incurred for this call.x-nr-input-tokens/x-nr-output-tokens: Exact provider token usage.
When using raw_response or inspecting the underlying HTTP response, capture these headers for cost accounting.
Troubleshooting & Error Handling
nRouter reports errors using standard HTTP status codes.
Common Error Codes
| Status | Code | Cause | Recommended Action |
|---|---|---|---|
400 | guardrail_blocked | Query violated content policies or prompt injection filters | Sanitize user input; review guardrail settings in nRouter dashboard. |
401 | authentication_error | Virtual API key is missing, invalid, or revoked | Verify NROUTER_API_KEY in environment variables. |
402 | insufficient_credits | Organization balance exhausted or virtual key limit reached | Add credits in the dashboard or increase key spending limits. |
429 | rate_limit_exceeded | RPM or TPM throughput ceilings reached | Implement backoff logic; check retry-after response headers. |
500 / 503 | service_unavailable | Downstream provider error or transient network issue | Use fallback model lists (model1,model2) for high availability. |
Error Catching Example
import openai
from llama_index.core import Settings
try:
response = Settings.llm.complete("Summarize sensitive corporate data.")
print(response.text)
except openai.BadRequestError as e:
if "guardrail" in str(e).lower():
print("Blocked by nRouter safety guardrail policy.")
else:
print(f"Bad request: {e}")
except openai.AuthenticationError:
print("Authentication error: Check your NROUTER_API_KEY.")
except openai.RateLimitError:
print("Rate limit reached: Backing off before retry.")
except openai.APIError as e:
print(f"Gateway error ({e.code}): {e.message}")Next Steps
- LangChain Integration — Multi-step chains and agent systems
- Python SDK Guide — Official nRouter Python SDK documentation
- Chat Completions API — HTTP endpoint specifications
LangChain
Use nRouter with LangChain in Python and JavaScript. Integrate smart routing, automatic retries, guardrails, and cost tracking into your LangChain pipelines.
Vercel AI SDK
Integrate nRouter with Vercel AI SDK in Next.js and React apps. Stream completions, utilize guardrails, and track costs while routing to any major AI model.