Browse documentation
FrameworksLlamaIndex

LlamaIndex

Connect LlamaIndex to nRouter for retrieval-augmented generation (RAG) and agent workflows. Configure LLM and embedding endpoints with enterprise guardrails.

Last updated

LlamaIndex is a data framework designed to connect external data sources to large language models for complex retrieval-augmented generation (RAG) and autonomous data agents. By integrating LlamaIndex with nRouter (https://api.nrouter.ai/v1), data pipelines inherit unified model access, real-time cost attribution, semantic query caching, and pre-execution guardrails across both generation and embedding endpoints.

Because LlamaIndex's OpenAI and OpenAIEmbedding classes adhere to standard REST contracts, you simply configure the gateway endpoint and virtual key. Guardrail policies configured in your nRouter dashboard automatically scan incoming queries for prompt injection and sensitive data before retrieval or synthesis runs.

Prerequisites & Installation

LlamaIndex requires Python 3.10 or higher. For reliable dependency resolution, install packages inside an active virtual environment.

Install the modular LlamaIndex OpenAI packages along with the core library:

pip install llama-index-core llama-index-llms-openai llama-index-embeddings-openai nrouter-sdk

Setup & Configuration

Store your nRouter virtual key in your environment:

export NROUTER_API_KEY="sk-nrouter-your-virtual-key"

Global Settings Configuration

In LlamaIndex, Settings provides global defaults for all query engines, indexes, and agents:

import os
from llama_index.core import Settings
from llama_index.llms.openai import OpenAI
from llama_index.embeddings.openai import OpenAIEmbedding

NROUTER_KEY = os.environ["NROUTER_API_KEY"]
NROUTER_BASE = "https://api.nrouter.ai/v1"

# Configure global generation model
Settings.llm = OpenAI(
    model="claude-sonnet-4-5-20250929",
    api_key=NROUTER_KEY,
    api_base=NROUTER_BASE,
    timeout=60.0,
    max_retries=2,
)

# Configure global embedding model
Settings.embed_model = OpenAIEmbedding(
    model_name="text-embedding-3-small",
    api_key=NROUTER_KEY,
    api_base=NROUTER_BASE,
    timeout=30.0,
)

Configuration Parameters

Configure client settings to match your latency and reliability requirements:

ParameterTypeDefaultDescription
api_basestrhttps://api.nrouter.ai/v1Unified gateway base URL. Must include /v1.
api_keystrNonenRouter virtual API key (sk-nrouter-...).
modelstrRequiredModel identifier, custom alias, or fallback list.
timeoutfloat60.0Maximum HTTP connection and generation timeout in seconds.
max_retriesint2Client-side retry limit. nRouter handles upstream retries automatically.
default_headersdict{}HTTP headers sent with every request, including x-nr-routing.
additional_kwargsdict{}Gateway overrides such as nrouter_cache or template settings.
import os
from llama_index.llms.openai import OpenAI

# Production client configuration with latency optimization
llm = OpenAI(
    model="gpt-5.4-mini",
    api_key=os.environ["NROUTER_API_KEY"],
    api_base="https://api.nrouter.ai/v1",
    timeout=45.0,
    max_retries=3,
    default_headers={
        "x-nr-routing": "latency",
    },
    additional_kwargs={
        "nrouter_cache": True,
    },
)

Implementation Patterns

1. Basic Chat & Generation

Interact with the configured model directly:

from llama_index.core.llms import ChatMessage

messages = [
    ChatMessage(role="system", content="You are a senior data architect."),
    ChatMessage(role="user", content="Explain vector embedding quantization."),
]

response = Settings.llm.chat(messages)
print(response.message.content)

2. Document Indexing & RAG Query Engine

Construct an in-memory vector index and run guarded queries:

from llama_index.core import VectorStoreIndex, Document

documents = [
    Document(text="nRouter routes inference across OpenAI, Anthropic, Bedrock, and open-source models."),
    Document(text="Virtual keys enforce budget limits and rate limits per team or project."),
    Document(text="Server-side guardrails block prompt injections before requests hit downstream models."),
]

# Build index using Settings.embed_model
index = VectorStoreIndex.from_documents(documents)

# Create query engine using Settings.llm
query_engine = index.as_query_engine()

response = query_engine.query("How does nRouter protect against prompt injection?")
print(response)

3. Streaming Query Engine

Deliver low time-to-first-token by streaming response tokens:

query_engine = index.as_query_engine(streaming=True)
streaming_response = query_engine.query("What are the key benefits of virtual keys?")

for text in streaming_response.response_gen:
    print(text, end="", flush=True)
print()

4. Per-Request Gateway Overrides

Attach prompt template identifiers or bypass semantic cache using additional_kwargs:

from llama_index.llms.openai import OpenAI

custom_llm = OpenAI(
    model="gpt-5.5",
    api_key=os.environ["NROUTER_API_KEY"],
    api_base="https://api.nrouter.ai/v1",
    additional_kwargs={
        "nrouter_prompt_template_id": "tmpl_rag_synthesis_v2",
        "nrouter_prompt_variables": {"audience": "executives"},
        "nrouter_cache": False,
    },
)

Production Best Practices

Deterministic Routing & Fallbacks

Prevent service downtime by providing fallback model lists and setting deterministic routing headers:

Settings.llm = OpenAI(
    # Primary model with automatic fallback
    model="gpt-5.4-mini,claude-haiku-4-5-20251001",
    api_key=os.environ["NROUTER_API_KEY"],
    api_base="https://api.nrouter.ai/v1",
    default_headers={
        "x-nr-routing": "latency",
    },
)
  • x-nr-routing: latency: Routes queries to the lowest-latency operational provider deployment.
  • x-nr-routing: cost: Directs traffic to the lowest-cost deployment matching the request.
  • Model Fallback List: If the primary provider experiences capacity constraints or elevated 5xx errors, nRouter seamlessly transfers the query to the fallback model.

Telemetry & FinOps Tracking

Every call through nRouter produces metadata headers:

  • x-nr-request-id: Trace ID linking application logs with nRouter spend ledgers.
  • x-nr-model: Actual model deployment executing the request.
  • x-nr-cost-status: Indication of pricing certainty (exact or unpriced).
  • x-nr-request-cost: USD spend incurred for this call.
  • x-nr-input-tokens / x-nr-output-tokens: Exact provider token usage.

When using raw_response or inspecting the underlying HTTP response, capture these headers for cost accounting.

Troubleshooting & Error Handling

nRouter reports errors using standard HTTP status codes.

Common Error Codes

StatusCodeCauseRecommended Action
400guardrail_blockedQuery violated content policies or prompt injection filtersSanitize user input; review guardrail settings in nRouter dashboard.
401authentication_errorVirtual API key is missing, invalid, or revokedVerify NROUTER_API_KEY in environment variables.
402insufficient_creditsOrganization balance exhausted or virtual key limit reachedAdd credits in the dashboard or increase key spending limits.
429rate_limit_exceededRPM or TPM throughput ceilings reachedImplement backoff logic; check retry-after response headers.
500 / 503service_unavailableDownstream provider error or transient network issueUse fallback model lists (model1,model2) for high availability.

Error Catching Example

import openai
from llama_index.core import Settings

try:
    response = Settings.llm.complete("Summarize sensitive corporate data.")
    print(response.text)
except openai.BadRequestError as e:
    if "guardrail" in str(e).lower():
        print("Blocked by nRouter safety guardrail policy.")
    else:
        print(f"Bad request: {e}")
except openai.AuthenticationError:
    print("Authentication error: Check your NROUTER_API_KEY.")
except openai.RateLimitError:
    print("Rate limit reached: Backing off before retry.")
except openai.APIError as e:
    print(f"Gateway error ({e.code}): {e.message}")

Next Steps

Was this page helpful?