Use Case · Document Processing

Extract and classify documents at volume.

Document workloads run vision-capable models over scans and PDFs in large batches. nRouter keeps every request on a vision model, paces the batch against provider quotas, and tracks cost per job.

batch · job-id invoices-0418

A document batch in flight

Documents4,200 pages
Capabilityvision
Modelgemini-2.5-pro
Cache hits312 repeats
Fallback used2 pages
Job costtracked
vision-onlybatch-pacedcost-tracked
Batch savings
45–65%

Via smart vision tier routing

Batch reliability
99.99%

Auto-failover through provider blips

Gateway overhead
95 ms

p50 added proxy latency

Document catalog
169+

models on Alibaba US, OpenAI, Azure Foundry, Google Vertex AI & Anthropic

Why nRouter for documents

The right model, at volume, with a cost number

Document processing has three constants: it needs vision, it runs in bulk, and it has a budget. nRouter handles all three behind one key.

Vision-capable models, by access

The catalog marks which models are vision-capable, and a virtual key can be restricted to just those. A key without access to a text-only model cannot reach it. Combine vision with long-context for big PDFs.

Built for batch workloads

Each page is a standard chat.completions call, so a batch is a loop. Per-key RPM/TPM limits pace it against provider quotas; caching skips identical pages.

Cost & usage tracking per job

Tag the batch with a job id and the dashboard attributes total spend to it. Cost-per-document becomes a measured number, capped by per-team and per-key budgets.

Failover for long batches

A provider degradation halfway through hours of pages shouldn’t mean restarting. The fallback chain retries the next link transparently and the run continues.

How it works

A document batch, end to end

A batch is a loop of standard requests. Each document names a vision-capable model, the loop is rate-paced and failover-protected, and total spend rolls up by job id.

Document batch flow

  1. Document Ingress

    batch job loop

    Invoices, receipts, legal scans, medical records, multi-page PDFs.

  2. Gateway & Rate Pacing

    :4000 · In-Memory RLS

    RPM/TPM pacing against quotas and job-id budget allocation.

  3. PII & Document Guardrail

    Pre-Inference Sanitizer

    Inline masking of tax IDs, SSNs, bank accounts, and credentials.

  4. Smart Vision Router

    Complexity & Cost Optimizer

    Standard forms route to fast vision; complex tables to frontier reasoning (45–65% ROI).

  5. Vision Model Providers

    Claude · GPT-4o · Gemini

    99.99% multi-provider failover keeping multi-hour batch runs alive.

Per-key model access is the safety rail — a key scoped to the vision-capable set cannot reach a text-only model, and a job-id tag rolls the batch's spend up on its own line.

The code

One request per document, in a loop

A document batch is a loop of standard chat.completions calls — image content in, structured fields out. These snippets are generated from the SDK examples the playground uses; add a job-id to metadata to attribute the batch.

Installpip install openai
1# Cache: enabled (org default). Pass nrouter_cache: false to skip.
2from openai import OpenAI
3import os
4
5client = OpenAI(
6 api_key=os.environ["NROUTER_API_KEY"],
7 base_url="https://api.nrouter.ai/v1",
8)
9
10response = client.chat.completions.create(
11 model="gpt-5.4-mini",
12 temperature=1,
13 max_completion_tokens=1024,
14 messages=[
15 {"role": "user", "content": "Hello! What models do you support?"},
16 ],
17 extra_body={
18 # "nrouter_cache": False, # Uncomment to skip cache
19 },
20)
21
22print(response.choices[0].message.content)

Tag each call with a job id in request metadata — spend analytics rolls the whole batch up to that job.

FAQ

Common document-processing questions

Does nRouter support vision-capable models for document work?

Yes. The model catalog marks vision-capable models, and per-key model access keeps batch workloads restricted to tested vision models. Requests to text-only models are blocked before dispatch, avoiding wasted inference attempts.

Can I run large batches of documents without hitting provider rate limits?

Yes. Each document is processed through standard OpenAI-compatible endpoints. The Rust gateway’s per-key RPM and TPM token-bucket algorithms pace the batch loop smoothly against upstream provider limits, eliminating 429 throttling errors.

How does tier routing reduce document processing costs?

High-volume standard receipts and straightforward invoices route to ultra-fast, low-cost multimodal models (e.g. Gemini 2.0 Flash or Claude 3.5 Haiku). Only dense legal contracts and handwritten tables route to heavy reasoning models, cutting total document extraction spend by 45% to 65%.

How do inline guardrails protect confidential financial and personal records?

Before document text or image OCR representations reach external providers, inline guardrails scan and mask tax IDs, social security numbers, banking credentials, and credit card numbers. Zero-content logging policies guarantee extracted text is never retained in disk logs.

What happens if a provider degrades in the middle of an overnight batch run?

nRouter’s automated fallback chain transparently retries failed requests across an alternative provider (e.g., failing over from OpenAI to Anthropic or Google Cloud). The batch continues without requiring an expensive manual restart, and failover events are tagged in the job audit log.

Vision, at volume, with a cost number

Process documents without minding provider quotas

Access-scoped vision models, batch-friendly rate limits, failover, and per-job cost & usage tracking — all unlocked on every plan.

Explore other use cases

Production AI workflows powered by nRouter