Vision-capable models, by access
The catalog marks which models are vision-capable, and a virtual key can be restricted to just those. A key without access to a text-only model cannot reach it. Combine vision with long-context for big PDFs.
Document workloads run vision-capable models over scans and PDFs in large batches. nRouter keeps every request on a vision model, paces the batch against provider quotas, and tracks cost per job.
A document batch in flight
Via smart vision tier routing
Auto-failover through provider blips
p50 added proxy latency
models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic
Document processing has three constants: it needs vision, it runs in bulk, and it has a budget. nRouter handles all three behind one key.
The catalog marks which models are vision-capable, and a virtual key can be restricted to just those. A key without access to a text-only model cannot reach it. Combine vision with long-context for big PDFs.
Each page is a standard chat.completions call, so a batch is a loop. Per-key RPM/TPM limits pace it against provider quotas; caching skips identical pages.
Tag the batch with a job id and the dashboard attributes total spend to it. Cost-per-document becomes a measured number, capped by per-team and per-key budgets.
A provider degradation halfway through hours of pages shouldn’t mean restarting. The fallback chain retries the next link transparently and the run continues.
A batch is a loop of standard requests. Each document names a vision-capable model, the loop is rate-paced and failover-protected, and total spend rolls up by job id.
Document batch flow
Document Ingress
batch job loop
Invoices, receipts, legal scans, medical records, multi-page PDFs.
Gateway & Rate Pacing
:4000 · In-Memory RLS
RPM/TPM pacing against quotas and job-id budget allocation.
PII & Document Guardrail
Pre-Inference Sanitizer
Inline masking of tax IDs, SSNs, bank accounts, and credentials.
Smart Vision Router
Complexity & Cost Optimizer
Standard forms route to fast vision; complex tables to frontier reasoning (45–65% ROI).
Vision Model Providers
Claude · GPT-4o · Gemini
99.99% multi-provider failover keeping multi-hour batch runs alive.
Per-key model access is the safety rail — a key scoped to the vision-capable set cannot reach a text-only model, and a job-id tag rolls the batch's spend up on its own line.
A document batch is a loop of standard chat.completions calls — image content in, structured fields out. These snippets are generated from the SDK examples the playground uses; add a job-id to metadata to attribute the batch.
pip install openai| 1 | # Cache: enabled (org default). Pass nrouter_cache: false to skip. |
| 2 | from openai import OpenAI |
| 3 | import os |
| 4 | |
| 5 | client = OpenAI( |
| 6 | api_key=os.environ["NROUTER_API_KEY"], |
| 7 | base_url="https://api.nrouter.ai/v1", |
| 8 | ) |
| 9 | |
| 10 | response = client.chat.completions.create( |
| 11 | model="gpt-5.4-mini", |
| 12 | temperature=1, |
| 13 | max_completion_tokens=1024, |
| 14 | messages=[ |
| 15 | {"role": "user", "content": "Hello! What models do you support?"}, |
| 16 | ], |
| 17 | extra_body={ |
| 18 | # "nrouter_cache": False, # Uncomment to skip cache |
| 19 | }, |
| 20 | ) |
| 21 | |
| 22 | print(response.choices[0].message.content) |
Tag each call with a job id in request metadata — spend analytics rolls the whole batch up to that job.
Yes. The model catalog marks vision-capable models, and per-key model access keeps batch workloads restricted to tested vision models. Requests to text-only models are blocked before dispatch, avoiding wasted inference attempts.
Yes. Each document is processed through standard OpenAI-compatible endpoints. The Rust gateway’s per-key RPM and TPM token-bucket algorithms pace the batch loop smoothly against upstream provider limits, eliminating 429 throttling errors.
High-volume standard receipts and straightforward invoices route to ultra-fast, low-cost multimodal models (e.g. Gemini 2.0 Flash or Claude 3.5 Haiku). Only dense legal contracts and handwritten tables route to heavy reasoning models, cutting total document extraction spend by 45% to 65%.
Before document text or image OCR representations reach external providers, inline guardrails scan and mask tax IDs, social security numbers, banking credentials, and credit card numbers. Zero-content logging policies guarantee extracted text is never retained in disk logs.
nRouter’s automated fallback chain transparently retries failed requests across an alternative provider (e.g., failing over from OpenAI to Anthropic or Google Cloud). The batch continues without requiring an expensive manual restart, and failover events are tagged in the job audit log.
Vision, at volume, with a cost number
Access-scoped vision models, batch-friendly rate limits, failover, and per-job cost & usage tracking — all unlocked on every plan.
Explore other use cases