← All posts
Comparison

Cutting LLM costs: what published discounts actually save, and where the model is an assumption

A rebuilt savings model for a $40k/month LLM bill. Batch and prompt caching carry published rates; provider capacity reservations do not, so that lever is a labelled assumption with a stated range. Every line is separated into sourced and assumed.

nRouter team
nRouter team
· 14 min readUpdated · Verified
Cutting LLM costs: what published discounts actually save, and where the model is an assumption

The Direct Answer: On a $40,000/month LLM bill, verifiable provider levers—batch processing (50% off), prompt caching (up to 90% discount on cache hits), and zero-markup gateway execution—consistently reduce annual spend from $480,000 to ~$368,000 (a 23% immediate cut). Pairing this with intent-based tier routing and prompt compression pushes realistic net savings to 35%–46% without proprietary model lock-in.

Cutting LLM Costs: 5 Realistic Levers from Published Discounts to Intent Routing

Figure 1: Five Realistic Cost Levers — Separating published vendor discounts (batching, prompt caching) from architectural optimization (intent routing, prompt compression).

If you're a VP of Engineering or Head of Platform looking at a $10k–$50k/month LLM bill, the same conversation is happening in every staff-engineering Slack: the model spend is real, the per-customer attribution is opaque, and every gateway that promises to fix it wants you on a $50k+/year enterprise contract before they hand over guardrails, evals, or per-team budgets.

What follows is the arithmetic, rebuilt so that every line says where its number came from. Some lines quote a rate a provider prints on a page you can open right now. Others quote a rate we chose, because no provider prints one. Mixing those two kinds of number in a single savings table is how cost posts end up unfalsifiable, and this one used to do it.

If you'd rather skip the narrative: jump to what the providers actually publish, the worked example, the plan note, or the recipe summary.


The setup: where the savings actually come from

For the complete LLM cost optimization framework, combine routing with caching, prompt reduction, retry controls, attribution, and hard budgets.

Most "save on LLM costs" posts pick one lever — caching, smaller model, batching — and pretend it's the whole story. We're stacking several, in the order a buyer can adopt them without re-architecting. The important split is not the order, though. It's this:

Levers with a published rate.

  1. Batch inference. Latency-tolerant work — nightly enrichment, backfills, eval runs, bulk classification — bills at half price on every major surface, and each vendor says so in writing.
  2. Prompt caching. A repeated prefix (system prompt, tool schema, retrieved context) bills at a published multiplier of the base input rate once it is cached, and at a published premium the first time it is written.
  3. The platform fee. A gateway either takes a cut of what you spend with providers or it doesn't. Ours is a flat 4% of the credits on every plan — added on top of the credits you buy, which works out to 4% of spend, with no minimum fee.

Levers with no published rate.

  1. Provider capacity reservations. Azure PTU, Google Vertex provisioned throughput, AWS Bedrock Provisioned Throughput. Real products, real discounts — and, as the next section shows, not one of the three publishes a headline percentage.
  2. Attribution-driven prompt cleanup. Deleting the prompts you can finally see are dead weight. There is no vendor page for this at all; the number is whatever your codebase happens to contain.

Levers 1–3 are things you can price today from public documents. Levers 4 and 5 are modelling inputs. Both belong in the model. Only one kind belongs in a sentence that begins "provider documentation says."


What the three providers actually publish

We re-read all three capacity pages on 2026-08-23 looking for the number this post used to quote. It is not there.

ProviderWhat the page says about commitmentsPercentage published
AzureAzure Reservations are "a financial discount applied to the PTU billing meter"; in exchange for a 1-month or 1-year commitment you get "a discounted effective $/PTU/hr rate"None
Google CloudGenerative-AI throughput reservations and committed-use discounts, priced per model on the pricing pageNone
AWSBedrock Provisioned Throughput, sold in model units with commitment termsNone

Sources: Azure's provisioned throughput concepts page and Azure OpenAI pricing, Google's Vertex AI generative-AI pricing, and AWS Bedrock pricing.

The only percentages any of those pages state near capacity language run the other way — Azure's own pricing page prices the Batch API at "a 50% discount on Global Standard Pricing," and its priority-processing tier is a premium, not a discount. So a reservation number is not a thing you can cite. It is a thing you can model, and label.

By contrast, here is what the batch and caching pages state outright:

LeverPublished rateWhere
Batch inference50% of synchronous priceOpenAI Batch API, Anthropic Message Batches, Azure OpenAI pricing, AWS Bedrock pricing
Cache read0.1× base input price (Anthropic); Bedrock cache reads "75% less than on-demand input token price"Anthropic prompt caching, AWS Bedrock pricing
Cache write1.25× base input price at the 5-minute TTL (Anthropic)Anthropic prompt caching
Platform feeA flat 4% of the credits on every plan, no minimum fee/pricing

Four vendors, one rate, printed in four places. That is what a citable lever looks like, and it is why the batch line below carries more weight in the model than the reservation line does.

We take the same position our routing-strategies guide takes: we quote no headline reservation percentage, because none of the three publishes one. Read the term you would actually buy.


The wedge: features are not the upsell

A nRouter customer can run all of these levers without buying an "enterprise tier", because on nRouter the plan never varies the feature set — guardrails, evals and per-team budgets are on the $0-subscription plan, and the platform fee is the same 4% on every plan. That is what makes the attribution lever included rather than an upgrade. Quick competitive context, drawn straight from public pricing pages on 2026-05-16:

CapabilityOpenRouterPortkeyHeliconenRouter
Per-team budgetsNot offeredEnterprise tierNot offered✅ Included, every plan
Eval pipelinesNot offeredEnterprise tierPro / Enterprise✅ Included, every plan
Guardrails (PII / jailbreak / regex)Not offeredPro / EnterprisePro / Enterprise✅ Included, every plan
Platform fee on self-serve credit purchases5.5%, $0.80 minimumAnnual contractAnnual contract4% on every plan, no minimum fee

OpenRouter, Portkey, and Helicone are trademarks of their respective owners. nRouter is not affiliated with or endorsed by any of these vendors. Every comparison row is sourced from each vendor's public pricing or documentation page on 2026-05-16; if any has changed, email us and we'll update.

If you've already read our OpenRouter alternative comparison, the matrix above is a subset — included here so this post stands alone as a buyer's reference.


Worked example: Customer A, $40k/month OpenAI direct

To make this concrete, we'll use a hypothetical customer profile that mirrors our ICP 2 spend band ($5k–$50k/month).

Customer A — anonymized opaque ID acct_a1b2 (no real customer names in

This article focuses on cut LLM costs in production, including the practical trade-offs around LLM gateway.

committed content)

  • Mid-market SaaS, 50–500 engineers, $5M–$50M ARR
  • $40,000/month spend on OpenAI direct (mix of GPT-class chat + embeddings)
  • Three teams: AI Features, Data Platform, Customer Support
  • Compliance asking for per-team cost attribution and per-customer redaction logs by Q3

Year-1 cost at the status quo (OpenAI direct): $40,000/month × 12 = $480,000/year of LLM spend, $0 of routing or governance tooling, and roughly 0.5 FTE (~$90k/yr loaded) maintaining a home-grown attribution dashboard on top of OpenAI's billing CSVs. That FTE line is excluded from every percentage below and called out separately, so you can validate it against your own loaded cost.

Step 1 — the traffic shape (ASSUMPTION)

Batch and caching rates are published, but how much of your bill they touch is a property of your workload, not of a vendor page. So this split is ours, and it is the first thing to replace with your own numbers:

Slice of the billShare (assumed)Annual spend
Latency-tolerant: enrichment, backfills, eval runs, bulk classification25%$120,000
Interactive with a repeated prefix: system prompt + retrieved context45%$216,000
Interactive, no reusable prefix30%$144,000

Step 2 — the sourced levers

Line itemRateBasisAmount
Annualized LLM spend at provider list—Customer A profile$480,000
Batch inference on the latency-tolerant slice50% off $120,000Published — four vendor pages-$60,000
Prompt caching on the repeated-prefix slice44% off $151,200Published multipliers, assumed hit rate-$66,528
nRouter platform fee on the $353,472 of credits that remain4%, every planPublished — /pricing+$14,139
Net year-1 LLM cost on published rates alone$367,611
Cash savings vs. status quo-$112,389 (23.4%)

The caching line is the only one that needs unpacking, because 44% is not a number any vendor prints — it is derived from two numbers they do print. Anthropic bills a cache read at 0.1× the base input price and a 5-minute cache write at 1.25×. Assume 70% of the repeated-prefix slice is input tokens ($151,200) and a 60% cache hit rate, and the blended multiplier is 0.60 × 0.10 + 0.40 × 1.25 = 0.56, i.e. 44% off. The hit rate is our assumption; Anthropic's own docs put observed hit rates for batched requests anywhere from 30% to 98%, so 60% sits well inside the band rather than at the flattering end of it.

Note what the platform-fee line does here: moving from provider-direct onto a gateway adds $14,139/year — 4% of the credits you buy. It is not a saving. The fee lever only pays when your alternative is a marked-up gateway — against a 5.5% fee on the same $353,472 of credits ($19,441), it is worth $5,302/year at this spend. The old version of this post quietly counted that delta as savings against a provider-direct baseline it did not apply to.

Step 3 — the assumed lever: capacity reservations

This line is a modelling input, not a citation. No provider publishes a headline reservation discount, so the number below is ours.

  • What it applies to: the interactive spend that remains after caching — $149,472 + $144,000 = $293,472. You would not reserve capacity for batch traffic that already bills at half price.
  • How much of it sits on reserved capacity: we assume half, $146,736.
  • Assumed discount range: 0–30%, central case 15%.
  • Stated basis: a reservation is billed per hour of capacity regardless of how many tokens you push through it, so the effective saving is a utilization outcome, not a rate. At full utilization it is whatever spread the reserved hourly rate happens to carry; at low utilization it trends to zero, and a stranded reservation makes it negative. We model the midpoint and show the floor.
Reservation caseSaving on $146,736Running provider spend (before the fee)
0% — reservation stranded or not taken$0$353,472
15% — central-$22,010$331,462
30% — top of our assumed range-$44,021$309,451

Step 4 — the second assumed lever: prompt cleanup

Also ours, also not a vendor number. Once per-team attribution shows which feature owns which slice of the bill, some of it gets deleted. We model 10% of what remains, with a 0–20% range, and we are explicit that this is a property of your prompts and not an observed industry rate.

  • 10% of $331,462 = -$33,146, leaving $298,316 of provider spend
  • 4% platform fee on those credits = +$11,933
  • Central-case year-1 total: $310,249
  • Central-case savings vs. status quo: -$169,751 (35.4%)

The three numbers to carry away

CaseYear-1 LLM costSaving%
Published rates only (batch + caching + fee; both assumptions at zero)$367,611$112,38923.4%
Central (reservations 15%, cleanup 10%)$310,249$169,75135.4%
Top of both assumed ranges (30% / 20%)$257,463$222,53746.4%

What this number is not. It is not a guarantee, and the two assumed levers are not hedged citations — they are our inputs, printed as such so you can overwrite them. Bring your last 90 days of LLM invoices to a 30-min call (/community) and we'll rerun every row against your actual traffic shape.


Plans: the fee is the same on every tier

Quick reference table:

PlanSubscriptionPlatform feeWhat the subscription adds
Pay as you go$0 ($5 minimum credit purchase)4% of the credits—
Starter$20/mo4% of the creditsA $60/mo nrouter/auto allowance and higher rate limits
Pro$50/mo4% of the creditsA $100/mo nrouter/auto allowance and higher rate limits
Max$200/mo4% of the creditsA $400/mo nrouter/auto allowance and higher rate limits
EnterpriseCustom, contact salesCustom terms—

The fee is charged on top with no minimum: $100 of credits is a $104.00 charge, and all $100 lands as credit. There is no spend level at which a subscription lowers it, so there is no breakeven to compute on the fee — a subscription buys a monthly nrouter/auto allowance and higher limits, not a lower fee. That is why the worked example above charges the same 4% whichever plan Customer A picks.


How the provider-reservation pass-through works

The published levers are things you can do on day one. Reservations are what nRouter does with aggregated customer volume, post-$10k ARR. It is the same model every cloud hyperscaler uses for compute, applied to LLM tokens — and it is worth understanding precisely because we can't quote you a rate for it.

  • Azure OpenAI PTU sells dedicated capacity by the hour, and Azure Reservations discount that hourly meter in exchange for a 1-month or 1-year commitment. The discounted effective $/PTU/hr rate is what you actually buy; Azure does not publish it as a percentage.
  • Google Vertex provisioned throughput and Vertex committed-use discounts are the same shape, priced per model on Google's own page.
  • AWS Bedrock Provisioned Throughput bills model units at a discounted hourly rate against the on-demand per-token rate.

The catch with all three is that a reservation is a single-tenant commit — it either gets used or it gets stranded. If your traffic dips below the reservation, you've burned the discount, which is exactly why the effective saving is a utilization outcome rather than a number on a rate card. nRouter pools traffic across customers, which means the gateway can sustain reservation utilization that no single customer could. That spread funds the platform rather than a promo discount, which keeps customer pricing — the same 4% fee on every plan — simple and stable.

This is the entire reason features are not the upsell: every customer needs the features (guardrails, evals, budgets) to understand their spend, and they need published discounts today and reservations tomorrow to reduce it. Locking either side behind an enterprise paywall would gate the people who most need the savings out of the savings.


The switch is one base URL and one API key

A common objection is that the savings are real but the migration cost dwarfs them. That isn't the case for nRouter. The OpenAI client and Anthropic client both accept a custom base_url and an alternate api_key. The diff for an existing OpenAI integration is two lines:

# Before — direct to OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

# After — through nRouter (everything else stays the same)
client = OpenAI(
    api_key=os.environ["NROUTER_API_KEY"],
    base_url="https://api.nrouter.ai/v1",
)

Same API surface, same SDK, same response shape. The difference is that with nRouter you also gain the multi-provider model registry, guardrails, A/B routing, evals, and per-team budgets in the same SDK call — and the batch and caching levers above become a routing decision rather than a rewrite.

In production deploys we've seen, the migration is typically scoped at one engineer for half a day, plus 24–72 hours of dual-run for confidence.


What this is not

A small number of buyers should not switch:

  • If your total LLM spend is small, the absolute-dollar savings (fee: a flat 4% of the credits on every plan) are real but small. Worth doing for the free guardrails + budgets, not for the cost lever.
  • If your traffic is entirely interactive with no reusable prefix, the two published levers in the model above touch none of your bill, the 23.4% floor collapses, and the 4% fee makes the move a net cost. Check your traffic shape before you check our arithmetic.
  • If you are contractually committed to a single provider for compliance reasons (some BAA / FedRAMP profiles), the multi-provider routing benefit doesn't apply — but the published batch and caching levers above still do.
  • If you've already negotiated direct-deal discounts with a provider, the reservation lever is muted. We can't tell you whether your deal beats a reservation, because neither figure is public.

We'd rather you skip the switch than have you switch on a math error. That includes ours: this post carried an "up to 70% annual / 30% monthly" reservation claim attributed to provider documentation, and on re-reading all three pages in August 2026 we could not find it. It has been replaced with the labelled assumption above rather than re-sourced, because there is nothing to re-source it to.


The recipe, summarized in five lines

  1. Move billing from N providers to nRouter so every dollar carries attribution metadata — and price the move honestly: it costs 4% of the credits you buy against provider-direct and saves against any gateway that charges more.
  2. Pick the plan on allowance and limits, not on the fee — every plan pays the same 4%; Starter, Pro and Max add a monthly nrouter/auto allowance and higher rate limits.
  3. Move every latency-tolerant workload to batch. This is the single best-sourced lever on the list: 50%, printed by four vendors.
  4. Cache your repeated prefixes and do the multiplier arithmetic for your hit rate, rather than assuming the cache-read rate applies to every input token.
  5. Treat capacity reservations and prompt cleanup as modelling inputs with ranges, re-run monthly against the attribution view, and never quote either as a provider-documented percentage.

If your bill is bigger than the worked example above, the dollar savings scale with it. If your bill is smaller, the percentages hold and the absolute dollars shrink — start with Pay as you go and revisit when you need a monthly nrouter/auto allowance or higher rate limits.


Try it on your own numbers

Load the $5 minimum in credits, with the platform fee on top. That is enough to route 5–10 production prompts on Pay as you go, look at the per-team budget UI with your real traffic shape, and decide whether the math holds for you before you sign anything.

→ Get started at app.nrouter.ai/signup — Pay as you go from $5, with the platform fee on top. No subscription. Mid-market SaaS or larger? Bring your last 90 days of LLM invoices to a 30-min walk-through (book through /community) and we'll redo the worked example above against your actual spend.


See also

The LLM cost optimization guide connects these levers into a repeatable production process.

Sources

Capacity and commitment pages, re-read 2026-08-23. None of the three states a headline reservation discount percentage, which is why this post models that lever instead of citing it:

  • Azure Foundry provisioned throughput (the hourly PTU meter and the 1-month / 1-year Azure Reservations that discount it): learn.microsoft.com
  • Azure OpenAI Service pricing (PTU meters, reservations, and the Batch API's 50% discount on Global Standard pricing): azure.microsoft.com
  • Google Vertex AI generative-AI pricing (throughput reservations + committed-use discounts): cloud.google.com
  • AWS Bedrock pricing (Provisioned Throughput commitment terms, 50% batch inference, cache reads at 75% less than on-demand input): aws.amazon.com

Rate pages the sourced levers above are taken from, verified 2026-08-23:

  • OpenAI Batch API — "50% cost discount compared to synchronous APIs": platform.openai.com
  • Anthropic Message Batches — "All usage is charged at 50% of the standard API prices": docs.anthropic.com
  • Anthropic prompt caching — cache reads at 0.1× and 5-minute cache writes at 1.25× the base input price, plus the 30–98% observed hit-rate band: docs.anthropic.com

Competitor pricing pages, verified 2026-05-16:

See also

If you are comparing production routing, pricing, and governance, see our nRouter vs OpenRouter comparison.

Frequently asked questions

How does cut LLM costs address setup: where the savings actually come from?

Most "save on LLM costs" posts pick one lever — caching, smaller model, batching — and pretend it's the whole story. We're stacking several, in the order a buyer can adopt them without re-architecting.

What should teams know about the three providers actually publish?

We re-read all three capacity pages on 2026-08-23 looking for the number this post used to quote. It is not there.

How does cut LLM costs address wedge: features are not the upsell?

A nRouter customer can run all of these levers without buying an "enterprise tier", because on nRouter the plan never varies the feature set — guardrails, evals and per-team budgets are on the $0-subscription plan, and the platform fee is the same 4% on every plan. That is what makes the attribution lever included rather than an upgrade.

Share
nRouter team
Written by nRouter team

The team building nRouter.