LangChain
Use nRouter with LangChain in Python and JavaScript. Integrate smart routing, automatic retries, guardrails, and cost tracking into your LangChain pipelines.
Last updated
LangChain provides a comprehensive framework for developing applications powered by language models, including chains, agents, and retrieval-augmented generation (RAG) systems. Because LangChain standardizes on OpenAI-compatible components, you can point LangChain directly to https://api.nrouter.ai/v1 without modifying prompt structures or execution graphs.
Connecting LangChain through nRouter enhances your pipelines with enterprise guardrails (blocking prompt injections and sensitive data before it reaches the model), automatic cross-provider failover, semantic response caching, and granular cost attribution per step.
Prerequisites & Installation
LangChain supports both Python and TypeScript/JavaScript environments.
Python Environment
Ensure Python 3.10+ is installed. Install langchain-openai along with the core library and vector store utilities:
pip install langchain-openai langchain-core langchain-community nrouter-sdkTypeScript / JavaScript Environment
For Node.js and Next.js applications, install @langchain/openai and @langchain/core:
npm install @langchain/openai @langchain/coreSetup & Configuration
Configure your environment with your nRouter virtual API key:
export NROUTER_API_KEY="sk-nrouter-your-virtual-key"Basic Initialization (Python)
Initialize ChatOpenAI with the nRouter gateway endpoint:
import os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="claude-sonnet-4-5-20250929",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
temperature=0.2,
timeout=60.0,
max_retries=2,
)
response = llm.invoke("Explain quantum computing in one sentence.")
print(response.content)Configuration Parameters
When deploying LangChain pipelines in production, configure the following client arguments:
| Parameter | Type | Default | Description |
|---|---|---|---|
base_url | str | https://api.nrouter.ai/v1 | Unified gateway base URI. Must include /v1. |
api_key | str | None | nRouter virtual key (sk-nrouter-...). |
model | str | Required | Model identifier, custom alias, or comma-separated fallback chain. |
timeout | float | 60.0 | Maximum wait time in seconds for HTTP responses. |
max_retries | int | 2 | Client-side retry limit. nRouter handles upstream retries automatically. |
default_headers | dict | {} | HTTP headers attached to every call, including x-nr-routing. |
model_kwargs | dict | {} | Additional payload fields, including nrouter_* parameters. |
import os
from langchain_openai import ChatOpenAI
# Advanced configuration with deterministic latency routing
llm = ChatOpenAI(
model="gpt-5.4-mini",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
timeout=45.0,
max_retries=3,
default_headers={
"x-nr-routing": "latency",
},
model_kwargs={
"nrouter_cache": True,
},
)Implementation Patterns
1. LangChain Expression Language (LCEL)
Build composable, streaming-ready chains using standard LCEL syntax:
import os
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
llm = ChatOpenAI(
model="gpt-5.4-mini",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
)
prompt = ChatPromptTemplate.from_messages([
("system", "You are an expert distributed systems engineer. Answer concisely."),
("user", "Explain why {concept} is critical for high-availability systems."),
])
chain = prompt | llm | StrOutputParser()
result = chain.invoke({"concept": "idempotency keys"})
print(result)2. Streaming Responses
Stream tokens chunk by chunk with minimal time-to-first-token:
for chunk in chain.stream({"concept": "circuit breakers"}):
print(chunk, end="", flush=True)
print()3. Retrieval-Augmented Generation (RAG)
Use OpenAIEmbeddings against nRouter alongside ChatOpenAI. Because nRouter inspects traffic before generation, user prompts attempting prompt injection in RAG contexts are rejected at the gateway:
import os
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_community.vectorstores import FAISS
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
# 1. Embeddings via nRouter
embeddings = OpenAIEmbeddings(
model="text-embedding-3-small",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
)
# 2. In-memory vector store
docs = [
"nRouter provides server-side guardrails and prompt-injection filtering.",
"nRouter charges exactly provider list prices with zero per-token markup.",
"Deterministic routing allows developers to prioritize cost or latency.",
]
vectorstore = FAISS.from_texts(docs, embeddings)
retriever = vectorstore.as_retriever(search_kwargs={"k": 2})
# 3. Generation model
llm = ChatOpenAI(
model="gpt-5.4-mini",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
)
template = """Answer the question based only on the following context:
{context}
Question: {question}
"""
prompt = ChatPromptTemplate.from_template(template)
rag_chain = (
{"context": retriever, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
print(rag_chain.invoke("Does nRouter add markups to model prices?"))4. Per-Request Gateway Overrides
Pass template IDs or cache preferences via model_kwargs:
llm_with_template = ChatOpenAI(
model="gpt-5.5",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
model_kwargs={
"nrouter_prompt_template_id": "tmpl_rag_synthesizer_v1",
"nrouter_prompt_variables": {"tone": "technical", "verbosity": "concise"},
"nrouter_cache": False,
},
)Production Best Practices
Deterministic Routing & Fallbacks
Configure failover chains directly in the model name, and enforce routing rules via headers:
llm = ChatOpenAI(
# Primary: gpt-5.4-mini; Fallback: claude-haiku-4-5-20251001
model="gpt-5.4-mini,claude-haiku-4-5-20251001",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
default_headers={
"x-nr-routing": "latency",
},
)x-nr-routing: latency: Routes to the lowest latency operational provider region.x-nr-routing: cost: Directs requests to the lowest-cost available provider.x-nr-routing: throughput: Routes requests to providers with highest concurrent token bandwidth.
Observability & FinOps Tracking
Each response through nRouter emits metadata headers tracking spend:
x-nr-request-id: Distributed trace identifier across client, gateway, and logs.x-nr-model: The specific model instance that served the completion.x-nr-cost-status: Status indicator (exactorunpriced).x-nr-request-cost: Exact USD cost incurred for the request.
Extract headers in Python via raw_response or response_metadata:
response = llm.invoke("Hello, nRouter!")
print("Content:", response.content)
# Inspect LangChain response metadata
meta = response.response_metadata
if "token_usage" in meta:
print(f"Tokens: {meta['token_usage']}")Troubleshooting & Error Handling
nRouter signals errors via standard HTTP status codes.
Common Error Codes
| Status | Code | Cause | Recommended Action |
|---|---|---|---|
400 | guardrail_blocked | Content failed safety, PII, or prompt injection guardrails | Inspect input content; verify rules in the nRouter dashboard. |
401 | authentication_error | Missing, incorrect, or revoked virtual API key | Verify NROUTER_API_KEY matches an active key in your organization. |
402 | insufficient_credits | Zero organization balance or key spending ceiling hit | Refill account credits or raise key budget ceilings. |
429 | rate_limit_exceeded | Organization RPM/TPM quota exceeded | Apply exponential backoff; check retry headers. |
500 / 503 | service_unavailable | Downstream provider outage or network disruption | Specify multiple models in a comma-separated fallback list. |
Robust Error Catching in Python
import openai
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="gpt-5.4-mini",
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
)
try:
response = llm.invoke("Execute standard pipeline task.")
except openai.BadRequestError as e:
if "guardrail" in str(e).lower():
print("Prompt was blocked by nRouter server-side guardrails.")
else:
print(f"Bad request: {e}")
except openai.AuthenticationError:
print("Authentication failed: Check NROUTER_API_KEY.")
except openai.RateLimitError:
print("Rate limit reached: Backing off before retry.")
except openai.APIError as e:
print(f"nRouter API error ({e.code}): {e.message}")Next Steps
- LlamaIndex Guide — Document indexing and advanced RAG
- Vercel AI SDK — Next.js and React full-stack integration
- Python SDK Guide — Official nRouter Python SDK
Frameworks
Connect nRouter to popular AI frameworks including LangChain, LlamaIndex, Vercel AI SDK, CrewAI, AutoGen, and Instructor via OpenAI-compatible endpoints.
LlamaIndex
Connect LlamaIndex to nRouter for retrieval-augmented generation (RAG) and agent workflows. Configure LLM and embedding endpoints with enterprise guardrails.