Browse documentation

Python (OpenAI client)

Use nRouter as a drop-in proxy with the official OpenAI Python SDK by setting the base URL, enabling team budgets, guardrails, and multi-model routing.

Last updated

If your Python codebase already uses the official openai SDK, you do not need to install additional proprietary libraries or refactor application code. By updating your client's base_url to https://api.nrouter.ai/v1 and configuring your virtual API key (NROUTER_API_KEY), nRouter operates as a transparent, high-performance edge proxy.

Routing through nRouter immediately unlocks enterprise-grade gateway controls: centralized team budgets via virtual keys, server-side prompt injection defenses, content moderation guardrails, multi-provider model routing (accessing OpenAI, Anthropic, Gemini, and open-source weights through a single interface), and exact list-price cost tracking without token markups.

Prerequisites & Installation

The OpenAI Python client requires Python 3.10 or higher.

Install the official OpenAI package via pip:

pip install openai httpx

Setup & Configuration

Store your nRouter virtual key in your environment:

export NROUTER_API_KEY="sk-nrouter-your-virtual-key"

Basic Initialization

Point the client to the nRouter gateway:

import os
from openai import OpenAI

# Drop-in configuration pointing to nRouter gateway
client = OpenAI(
    api_key=os.environ["NROUTER_API_KEY"],
    base_url="https://api.nrouter.ai/v1",
)

Configuration Parameters

Configure connection timeouts, retries, and HTTP transport settings:

ParameterTypeDefaultDescription
base_urlstrhttps://api.nrouter.ai/v1nRouter unified gateway endpoint. Must include /v1.
api_keystrEnvironmentnRouter virtual key (sk-nrouter-...).
timeoutfloat60.0Maximum request duration in seconds.
max_retriesint2Client-side retry attempts. nRouter handles upstream provider retries automatically.
default_headersdict{}Custom HTTP headers sent on every request (e.g. x-nr-routing).
http_clienthttpx.ClientNoneCustom httpx.Client configured with connection limits and keep-alives.
import os
import httpx
from openai import OpenAI

# Production client configuration with connection pooling
http_client = httpx.Client(
    timeout=httpx.Timeout(connect=5.0, read=45.0, write=10.0, pool=60.0),
    limits=httpx.Limits(max_keepalive_connections=50, max_connections=200),
)

client = OpenAI(
    api_key=os.environ["NROUTER_API_KEY"],
    base_url="https://api.nrouter.ai/v1",
    http_client=http_client,
    max_retries=3,
    default_headers={
        "x-nr-routing": "latency",
    },
)

Implementation Patterns

1. Synchronous Chat Completion

Execute standard chat completions across any supported model:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["NROUTER_API_KEY"],
    base_url="https://api.nrouter.ai/v1",
)

response = client.chat.completions.create(
    model="gpt-5.4-mini",
    messages=[
        {"role": "system", "content": "You are a concise technical writer."},
        {"role": "user", "content": "Explain zero-markup model pricing."},
    ],
)

print(response.choices[0].message.content)

2. Provider Switching

Because nRouter normalizes model endpoints, you can switch providers without changing client libraries or endpoints:

# OpenAI model
res_openai = client.chat.completions.create(
    model="gpt-5.4-mini",
    messages=[{"role": "user", "content": "Hello OpenAI!"}],
)

# Google model via the same client and gateway
res_gemini = client.chat.completions.create(
    model="gemini-2.5-flash",
    messages=[{"role": "user", "content": "Hello Gemini!"}],
)

Anthropic Claude Models: To call Anthropic Claude models on nRouter, use the native Anthropic Messages wire format (/v1/messages) via the official nrouter-sdk or standard Anthropic client library pointed at https://api.nrouter.ai.

3. Server-Sent Events (SSE) Streaming

Stream tokens in real-time with low latency:

stream = client.chat.completions.create(
    model="claude-haiku-4-5-20251001",
    messages=[{"role": "user", "content": "Write a short poem about distributed queues."}],
    stream=True,
)

for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
print()

4. Asynchronous Completions (AsyncOpenAI)

Use the asynchronous client for non-blocking I/O in FastAPI or asyncio applications:

import asyncio
import os
from openai import AsyncOpenAI

async def main():
    async_client = AsyncOpenAI(
        api_key=os.environ["NROUTER_API_KEY"],
        base_url="https://api.nrouter.ai/v1",
    )

    response = await async_client.chat.completions.create(
        model="gpt-5.4-mini",
        messages=[{"role": "user", "content": "Explain async I/O in Python."}],
    )
    print(response.choices[0].message.content)

asyncio.run(main())

5. Per-Request Gateway Overrides

Attach prompt template identifiers or bypass caching via extra_body:

response = client.chat.completions.create(
    model="gpt-5.5",
    messages=[{"role": "user", "content": "Summarize customer incident report."}],
    extra_body={
        "nrouter_prompt_template_id": "tmpl_incident_summary_v1",
        "nrouter_prompt_variables": {"environment": "production"},
        "nrouter_cache": False,
    },
)

Production Best Practices

Deterministic Routing & Fallbacks

Ensure high reliability by specifying fallback chains and routing priorities:

response = client.chat.completions.create(
    # Primary model with automatic fallback
    model="gpt-5.4-mini,claude-haiku-4-5-20251001",
    messages=[{"role": "user", "content": "Analyze security logs."}],
    extra_headers={
        "x-nr-routing": "latency",
    },
)
  • x-nr-routing: latency: Directs traffic to the lowest-latency active provider deployment.
  • x-nr-routing: cost: Prioritizes the most cost-effective deployment matching the model specification.
  • Model Fallback Chain: If the primary provider encounters rate limits or upstream 5xx errors, nRouter fails over instantly to the fallback model.

Telemetry & FinOps Tracking

Extract nRouter gateway headers using the client's with_raw_response helper:

raw_response = client.chat.completions.with_raw_response.create(
    model="gpt-5.4-mini",
    messages=[{"role": "user", "content": "Hello!"}],
)

# Parsed completion object
completion = raw_response.parse()
print(completion.choices[0].message.content)

# Access raw gateway headers
headers = raw_response.headers
print(f"Request ID: {headers.get('x-nr-request-id')}")
print(f"Model: {headers.get('x-nr-model')}")
print(f"Cost Status: {headers.get('x-nr-cost-status')}")
if headers.get("x-nr-cost-status") == "exact":
    print(f"Request Cost: ${headers.get('x-nr-request-cost')}")

Troubleshooting & Error Handling

nRouter signals errors via standard HTTP status codes.

Common Error Codes

StatusCodeCauseRecommended Action
400guardrail_blockedInput rejected by server-side content or injection guardrailsVerify prompt safety; review guardrail settings in nRouter dashboard.
401authentication_errorMissing, incorrect, or expired virtual API keyCheck NROUTER_API_KEY environment variable.
402insufficient_creditsZero organization balance or key spending limit reachedAdd funds in dashboard or update key ceiling.
429rate_limit_exceededClient exceeded RPM/TPM quotaImplement exponential backoff; check retry headers.
500 / 503service_unavailableDownstream provider error or network disruptionUse fallback model lists (model1,model2).

Error Catching Pattern

import openai

try:
    response = client.chat.completions.create(
        model="gpt-5.4-mini",
        messages=[{"role": "user", "content": "Execute task."}],
    )
except openai.BadRequestError as e:
    if "guardrail" in str(e).lower():
        print("Blocked by nRouter safety guardrail policy.")
    else:
        print(f"Invalid request parameters: {e}")
except openai.AuthenticationError:
    print("Authentication failed: Check your NROUTER_API_KEY.")
except openai.RateLimitError:
    print("Rate limit reached. Apply exponential backoff.")
except openai.APIError as e:
    print(f"Gateway error ({e.code}): {e.message}")

Next Steps

Was this page helpful?