Cutting LLM costs: what published discounts actually save, and where the model is an assumption
A rebuilt savings model for a $40k/month LLM bill. Batch and prompt caching carry published rates; provider capacity reservations do not, so that lever is a labelled assumption with a stated range. Every line is separated into sourced and assumed.

The Direct Answer: On a $40,000/month LLM bill, verifiable provider levers—batch processing (50% off), prompt caching (up to 90% discount on cache hits), and zero-markup gateway execution—consistently reduce annual spend from $480,000 to ~$368,000 (a 23% immediate cut). Pairing this with intent-based tier routing and prompt compression pushes realistic net savings to 35%–46% without proprietary model lock-in.
Figure 1: Five Realistic Cost Levers — Separating published vendor discounts (batching, prompt caching) from architectural optimization (intent routing, prompt compression).
If you're a VP of Engineering or Head of Platform looking at a $10k–$50k/month LLM bill, the same conversation is happening in every staff-engineering Slack: the model spend is real, the per-customer attribution is opaque, and every gateway that promises to fix it wants you on a $50k+/year enterprise contract before they hand over guardrails, evals, or per-team budgets.
What follows is the arithmetic, rebuilt so that every line says where its number came from. Some lines quote a rate a provider prints on a page you can open right now. Others quote a rate we chose, because no provider prints one. Mixing those two kinds of number in a single savings table is how cost posts end up unfalsifiable, and this one used to do it.
If you'd rather skip the narrative: jump to what the providers actually publish, the worked example, the plan note, or the recipe summary.
The setup: where the savings actually come from
For the complete LLM cost optimization framework, combine routing with caching, prompt reduction, retry controls, attribution, and hard budgets.
Most "save on LLM costs" posts pick one lever — caching, smaller model, batching — and pretend it's the whole story. We're stacking several, in the order a buyer can adopt them without re-architecting. The important split is not the order, though. It's this:
Levers with a published rate.
- Batch inference. Latency-tolerant work — nightly enrichment, backfills, eval runs, bulk classification — bills at half price on every major surface, and each vendor says so in writing.
- Prompt caching. A repeated prefix (system prompt, tool schema, retrieved context) bills at a published multiplier of the base input rate once it is cached, and at a published premium the first time it is written.
- The platform fee. A gateway either takes a cut of what you spend with providers or it doesn't. Ours is a flat 4% of the credits on every plan — added on top of the credits you buy, which works out to 4% of spend, with no minimum fee.
Levers with no published rate.
- Provider capacity reservations. Azure PTU, Google Vertex provisioned throughput, AWS Bedrock Provisioned Throughput. Real products, real discounts — and, as the next section shows, not one of the three publishes a headline percentage.
- Attribution-driven prompt cleanup. Deleting the prompts you can finally see are dead weight. There is no vendor page for this at all; the number is whatever your codebase happens to contain.
Levers 1–3 are things you can price today from public documents. Levers 4 and 5 are modelling inputs. Both belong in the model. Only one kind belongs in a sentence that begins "provider documentation says."
What the three providers actually publish
We re-read all three capacity pages on 2026-08-23 looking for the number this post used to quote. It is not there.
| Provider | What the page says about commitments | Percentage published |
|---|---|---|
| Azure | Azure Reservations are "a financial discount applied to the PTU billing meter"; in exchange for a 1-month or 1-year commitment you get "a discounted effective $/PTU/hr rate" | None |
| Google Cloud | Generative-AI throughput reservations and committed-use discounts, priced per model on the pricing page | None |
| AWS | Bedrock Provisioned Throughput, sold in model units with commitment terms | None |
Sources: Azure's provisioned throughput concepts page and Azure OpenAI pricing, Google's Vertex AI generative-AI pricing, and AWS Bedrock pricing.
The only percentages any of those pages state near capacity language run the other way — Azure's own pricing page prices the Batch API at "a 50% discount on Global Standard Pricing," and its priority-processing tier is a premium, not a discount. So a reservation number is not a thing you can cite. It is a thing you can model, and label.
By contrast, here is what the batch and caching pages state outright:
| Lever | Published rate | Where |
|---|---|---|
| Batch inference | 50% of synchronous price | OpenAI Batch API, Anthropic Message Batches, Azure OpenAI pricing, AWS Bedrock pricing |
| Cache read | 0.1× base input price (Anthropic); Bedrock cache reads "75% less than on-demand input token price" | Anthropic prompt caching, AWS Bedrock pricing |
| Cache write | 1.25× base input price at the 5-minute TTL (Anthropic) | Anthropic prompt caching |
| Platform fee | A flat 4% of the credits on every plan, no minimum fee | /pricing |
Four vendors, one rate, printed in four places. That is what a citable lever looks like, and it is why the batch line below carries more weight in the model than the reservation line does.
We take the same position our routing-strategies guide takes: we quote no headline reservation percentage, because none of the three publishes one. Read the term you would actually buy.
The wedge: features are not the upsell
A nRouter customer can run all of these levers without buying an "enterprise tier", because on nRouter the plan never varies the feature set — guardrails, evals and per-team budgets are on the $0-subscription plan, and the platform fee is the same 4% on every plan. That is what makes the attribution lever included rather than an upgrade. Quick competitive context, drawn straight from public pricing pages on 2026-05-16:
| Capability | OpenRouter | Portkey | Helicone | nRouter |
|---|---|---|---|---|
| Per-team budgets | Not offered | Enterprise tier | Not offered | ✅ Included, every plan |
| Eval pipelines | Not offered | Enterprise tier | Pro / Enterprise | ✅ Included, every plan |
| Guardrails (PII / jailbreak / regex) | Not offered | Pro / Enterprise | Pro / Enterprise | ✅ Included, every plan |
| Platform fee on self-serve credit purchases | 5.5%, $0.80 minimum | Annual contract | Annual contract | 4% on every plan, no minimum fee |
OpenRouter, Portkey, and Helicone are trademarks of their respective owners. nRouter is not affiliated with or endorsed by any of these vendors. Every comparison row is sourced from each vendor's public pricing or documentation page on 2026-05-16; if any has changed, email us and we'll update.
If you've already read our OpenRouter alternative comparison, the matrix above is a subset — included here so this post stands alone as a buyer's reference.
Worked example: Customer A, $40k/month OpenAI direct
To make this concrete, we'll use a hypothetical customer profile that mirrors our ICP 2 spend band ($5k–$50k/month).
Customer A — anonymized opaque ID
acct_a1b2(no real customer names in
This article focuses on cut LLM costs in production, including the practical trade-offs around LLM gateway.
committed content)
- Mid-market SaaS, 50–500 engineers, $5M–$50M ARR
- $40,000/month spend on OpenAI direct (mix of GPT-class chat + embeddings)
- Three teams: AI Features, Data Platform, Customer Support
- Compliance asking for per-team cost attribution and per-customer redaction logs by Q3
Year-1 cost at the status quo (OpenAI direct): $40,000/month × 12 = $480,000/year of LLM spend, $0 of routing or governance tooling, and roughly 0.5 FTE (~$90k/yr loaded) maintaining a home-grown attribution dashboard on top of OpenAI's billing CSVs. That FTE line is excluded from every percentage below and called out separately, so you can validate it against your own loaded cost.
Step 1 — the traffic shape (ASSUMPTION)
Batch and caching rates are published, but how much of your bill they touch is a property of your workload, not of a vendor page. So this split is ours, and it is the first thing to replace with your own numbers:
| Slice of the bill | Share (assumed) | Annual spend |
|---|---|---|
| Latency-tolerant: enrichment, backfills, eval runs, bulk classification | 25% | $120,000 |
| Interactive with a repeated prefix: system prompt + retrieved context | 45% | $216,000 |
| Interactive, no reusable prefix | 30% | $144,000 |
Step 2 — the sourced levers
| Line item | Rate | Basis | Amount |
|---|---|---|---|
| Annualized LLM spend at provider list | — | Customer A profile | $480,000 |
| Batch inference on the latency-tolerant slice | 50% off $120,000 | Published — four vendor pages | -$60,000 |
| Prompt caching on the repeated-prefix slice | 44% off $151,200 | Published multipliers, assumed hit rate | -$66,528 |
| nRouter platform fee on the $353,472 of credits that remain | 4%, every plan | Published — /pricing | +$14,139 |
| Net year-1 LLM cost on published rates alone | $367,611 | ||
| Cash savings vs. status quo | -$112,389 (23.4%) |
The caching line is the only one that needs unpacking, because 44% is not a
number any vendor prints — it is derived from two numbers they do print.
Anthropic bills a cache read at 0.1× the base input price and a 5-minute
cache write at 1.25×. Assume 70% of the repeated-prefix slice is input
tokens ($151,200) and a 60% cache hit rate, and the blended multiplier is
0.60 × 0.10 + 0.40 × 1.25 = 0.56, i.e. 44% off. The hit rate is our
assumption; Anthropic's own docs put observed hit rates for batched requests
anywhere from 30% to 98%, so 60% sits well inside the band rather than at the
flattering end of it.
Note what the platform-fee line does here: moving from provider-direct onto a gateway adds $14,139/year — 4% of the credits you buy. It is not a saving. The fee lever only pays when your alternative is a marked-up gateway — against a 5.5% fee on the same $353,472 of credits ($19,441), it is worth $5,302/year at this spend. The old version of this post quietly counted that delta as savings against a provider-direct baseline it did not apply to.
Step 3 — the assumed lever: capacity reservations
This line is a modelling input, not a citation. No provider publishes a headline reservation discount, so the number below is ours.
- What it applies to: the interactive spend that remains after caching — $149,472 + $144,000 = $293,472. You would not reserve capacity for batch traffic that already bills at half price.
- How much of it sits on reserved capacity: we assume half, $146,736.
- Assumed discount range: 0–30%, central case 15%.
- Stated basis: a reservation is billed per hour of capacity regardless of how many tokens you push through it, so the effective saving is a utilization outcome, not a rate. At full utilization it is whatever spread the reserved hourly rate happens to carry; at low utilization it trends to zero, and a stranded reservation makes it negative. We model the midpoint and show the floor.
| Reservation case | Saving on $146,736 | Running provider spend (before the fee) |
|---|---|---|
| 0% — reservation stranded or not taken | $0 | $353,472 |
| 15% — central | -$22,010 | $331,462 |
| 30% — top of our assumed range | -$44,021 | $309,451 |
Step 4 — the second assumed lever: prompt cleanup
Also ours, also not a vendor number. Once per-team attribution shows which feature owns which slice of the bill, some of it gets deleted. We model 10% of what remains, with a 0–20% range, and we are explicit that this is a property of your prompts and not an observed industry rate.
- 10% of $331,462 = -$33,146, leaving $298,316 of provider spend
- 4% platform fee on those credits = +$11,933
- Central-case year-1 total: $310,249
- Central-case savings vs. status quo: -$169,751 (35.4%)
The three numbers to carry away
| Case | Year-1 LLM cost | Saving | % |
|---|---|---|---|
| Published rates only (batch + caching + fee; both assumptions at zero) | $367,611 | $112,389 | 23.4% |
| Central (reservations 15%, cleanup 10%) | $310,249 | $169,751 | 35.4% |
| Top of both assumed ranges (30% / 20%) | $257,463 | $222,537 | 46.4% |
What this number is not. It is not a guarantee, and the two assumed levers are not hedged citations — they are our inputs, printed as such so you can overwrite them. Bring your last 90 days of LLM invoices to a 30-min call (/community) and we'll rerun every row against your actual traffic shape.
Plans: the fee is the same on every tier
Quick reference table:
| Plan | Subscription | Platform fee | What the subscription adds |
|---|---|---|---|
| Pay as you go | $0 ($5 minimum credit purchase) | 4% of the credits | — |
| Starter | $20/mo | 4% of the credits | A $60/mo nrouter/auto allowance and higher rate limits |
| Pro | $50/mo | 4% of the credits | A $100/mo nrouter/auto allowance and higher rate limits |
| Max | $200/mo | 4% of the credits | A $400/mo nrouter/auto allowance and higher rate limits |
| Enterprise | Custom, contact sales | Custom terms | — |
The fee is charged on top with no minimum: $100 of credits is a $104.00 charge,
and all $100 lands as credit. There is no spend level at which a subscription
lowers it, so there is no breakeven to compute on the fee — a subscription buys
a monthly nrouter/auto allowance and higher limits, not a lower fee. That is
why the worked example above charges the same 4% whichever plan Customer A picks.
How the provider-reservation pass-through works
The published levers are things you can do on day one. Reservations are what nRouter does with aggregated customer volume, post-$10k ARR. It is the same model every cloud hyperscaler uses for compute, applied to LLM tokens — and it is worth understanding precisely because we can't quote you a rate for it.
- Azure OpenAI PTU sells dedicated capacity by the hour, and Azure Reservations discount that hourly meter in exchange for a 1-month or 1-year commitment. The discounted effective $/PTU/hr rate is what you actually buy; Azure does not publish it as a percentage.
- Google Vertex provisioned throughput and Vertex committed-use discounts are the same shape, priced per model on Google's own page.
- AWS Bedrock Provisioned Throughput bills model units at a discounted hourly rate against the on-demand per-token rate.
The catch with all three is that a reservation is a single-tenant commit — it either gets used or it gets stranded. If your traffic dips below the reservation, you've burned the discount, which is exactly why the effective saving is a utilization outcome rather than a number on a rate card. nRouter pools traffic across customers, which means the gateway can sustain reservation utilization that no single customer could. That spread funds the platform rather than a promo discount, which keeps customer pricing — the same 4% fee on every plan — simple and stable.
This is the entire reason features are not the upsell: every customer needs the features (guardrails, evals, budgets) to understand their spend, and they need published discounts today and reservations tomorrow to reduce it. Locking either side behind an enterprise paywall would gate the people who most need the savings out of the savings.
The switch is one base URL and one API key
A common objection is that the savings are real but the migration cost dwarfs
them. That isn't the case for nRouter. The OpenAI client and Anthropic
client both accept a custom base_url and an alternate api_key. The diff for
an existing OpenAI integration is two lines:
# Before — direct to OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
# After — through nRouter (everything else stays the same)
client = OpenAI(
api_key=os.environ["NROUTER_API_KEY"],
base_url="https://api.nrouter.ai/v1",
)Same API surface, same SDK, same response shape. The difference is that with nRouter you also gain the multi-provider model registry, guardrails, A/B routing, evals, and per-team budgets in the same SDK call — and the batch and caching levers above become a routing decision rather than a rewrite.
In production deploys we've seen, the migration is typically scoped at one engineer for half a day, plus 24–72 hours of dual-run for confidence.
What this is not
A small number of buyers should not switch:
- If your total LLM spend is small, the absolute-dollar savings (fee: a flat 4% of the credits on every plan) are real but small. Worth doing for the free guardrails + budgets, not for the cost lever.
- If your traffic is entirely interactive with no reusable prefix, the two published levers in the model above touch none of your bill, the 23.4% floor collapses, and the 4% fee makes the move a net cost. Check your traffic shape before you check our arithmetic.
- If you are contractually committed to a single provider for compliance reasons (some BAA / FedRAMP profiles), the multi-provider routing benefit doesn't apply — but the published batch and caching levers above still do.
- If you've already negotiated direct-deal discounts with a provider, the reservation lever is muted. We can't tell you whether your deal beats a reservation, because neither figure is public.
We'd rather you skip the switch than have you switch on a math error. That includes ours: this post carried an "up to 70% annual / 30% monthly" reservation claim attributed to provider documentation, and on re-reading all three pages in August 2026 we could not find it. It has been replaced with the labelled assumption above rather than re-sourced, because there is nothing to re-source it to.
The recipe, summarized in five lines
- Move billing from N providers to nRouter so every dollar carries attribution metadata — and price the move honestly: it costs 4% of the credits you buy against provider-direct and saves against any gateway that charges more.
- Pick the plan on allowance and limits, not on the fee — every plan pays the
same 4%; Starter, Pro and Max add a monthly
nrouter/autoallowance and higher rate limits. - Move every latency-tolerant workload to batch. This is the single best-sourced lever on the list: 50%, printed by four vendors.
- Cache your repeated prefixes and do the multiplier arithmetic for your hit rate, rather than assuming the cache-read rate applies to every input token.
- Treat capacity reservations and prompt cleanup as modelling inputs with ranges, re-run monthly against the attribution view, and never quote either as a provider-documented percentage.
If your bill is bigger than the worked example above, the dollar savings scale
with it. If your bill is smaller, the percentages hold and the absolute dollars
shrink — start with Pay as you go and revisit when you need a monthly nrouter/auto allowance or higher rate limits.
Try it on your own numbers
Load the $5 minimum in credits, with the platform fee on top. That is enough to route 5–10 production prompts on Pay as you go, look at the per-team budget UI with your real traffic shape, and decide whether the math holds for you before you sign anything.
→ Get started at app.nrouter.ai/signup — Pay as you go from $5, with the platform fee on top. No subscription. Mid-market SaaS or larger? Bring your last 90 days of LLM invoices to a 30-min walk-through (book through /community) and we'll redo the worked example above against your actual spend.
See also
The LLM cost optimization guide connects these levers into a repeatable production process.
- OpenRouter alternative: every enterprise LLM-gateway feature, on every plan — the head-to-head that shows which of these savings come from the fee and which from routing.
- LLM gateway buyer's guide 2026 — the buyer-stage taxonomy for deciding which gateway shape fits before you optimise cost at all.
- LLM routing strategies 2026: benchmark-anchored vs ML-classifier vs operator-controlled — the sibling post that takes the same position on unpublished reservation rates.
- From Credits to Pro: when a flat subscription beats per-call billing — how pay-as-you-go credits and a subscription plan compare.
- Cost-vs-Quality LLM Routing: Which Tasks Can Go Cheap — the routing half of the saving, as a procedure you can run this afternoon.
- 5% of requests, 60% of the bill: reading cost against usage — how to find the quietly expensive model before you try to cut anything.
- Cost honesty: unpriced is never $0 on your LLM bill — why the baseline number this post optimises against is trustworthy in the first place.
- /pricing — the canonical Pay as you go / Starter / Pro / Max / Enterprise table these figures come from.
Sources
Capacity and commitment pages, re-read 2026-08-23. None of the three states a headline reservation discount percentage, which is why this post models that lever instead of citing it:
- Azure Foundry provisioned throughput (the hourly PTU meter and the 1-month / 1-year Azure Reservations that discount it): learn.microsoft.com
- Azure OpenAI Service pricing (PTU meters, reservations, and the Batch API's 50% discount on Global Standard pricing): azure.microsoft.com
- Google Vertex AI generative-AI pricing (throughput reservations + committed-use discounts): cloud.google.com
- AWS Bedrock pricing (Provisioned Throughput commitment terms, 50% batch inference, cache reads at 75% less than on-demand input): aws.amazon.com
Rate pages the sourced levers above are taken from, verified 2026-08-23:
- OpenAI Batch API — "50% cost discount compared to synchronous APIs": platform.openai.com
- Anthropic Message Batches — "All usage is charged at 50% of the standard API prices": docs.anthropic.com
- Anthropic prompt caching — cache reads at 0.1× and 5-minute cache writes at 1.25× the base input price, plus the 30–98% observed hit-rate band: docs.anthropic.com
Competitor pricing pages, verified 2026-05-16:
- OpenRouter pricing: openrouter.ai/pricing
- Portkey pricing: portkey.ai/pricing
- Helicone pricing: helicone.ai/pricing
See also
If you are comparing production routing, pricing, and governance, see our nRouter vs OpenRouter comparison.
Frequently asked questions
How does cut LLM costs address setup: where the savings actually come from?
Most "save on LLM costs" posts pick one lever — caching, smaller model, batching — and pretend it's the whole story. We're stacking several, in the order a buyer can adopt them without re-architecting.
What should teams know about the three providers actually publish?
We re-read all three capacity pages on 2026-08-23 looking for the number this post used to quote. It is not there.
How does cut LLM costs address wedge: features are not the upsell?
A nRouter customer can run all of these levers without buying an "enterprise tier", because on nRouter the plan never varies the feature set — guardrails, evals and per-team budgets are on the $0-subscription plan, and the platform fee is the same 4% on every plan. That is what makes the attribution lever included rather than an upgrade.
The team building nRouter.
More from Comparison
All posts →
Langfuse alternative: observability AND routing AND governance included on every plan
Head-to-head: nRouter vs Langfuse. A self-hostable observability + prompt + evals specialist vs a hosted LLM gateway that bundles observability, routing, and governance — every feature included on every plan. Platform fee of 4% of your credits on every plan. Models available in your live catalog behind one API key.

TrueFoundry AI Gateway alternative: buying one module of a platform
TrueFoundry's gateway is one module of a platform that also sells model deployment, GPU serving, and agent, MCP and skills registries — metered by requests and by seat. nRouter sells the gateway alone, priced as a share of model spend. A procedure for deciding which purchase you are making.

NotDiamond alternative: router versus managed LLM gateway
Not Diamond returns a model recommendation and charges $0.05 per million tokens routed — you still hold every provider key and make the call yourself. What that leaves you to build, and how deterministic A/B tests compare to a trained router when you have to reproduce a decision.