Tag

rate-limiting

6 posts tagged "rate-limiting".

Editorial Guide

Multi-Tier Rate Limiting: RPM, TPM & Concurrency Control

Protecting backend services and maintaining predictable quality of service in multi-tenant environments requires robust rate limiting. In generative AI systems, rate limiting extends far beyond simple HTTP request counters: gateways must manage requests per minute (RPM), tokens per minute (TPM), and concurrent in-flight connections across organizations, teams, and individual virtual keys. Multi-tier rate limiting prevents noisy-neighbor degradation and safeguards against provider tier exhaustion.

Key Engineering Challenges

Noisy Neighbor Degradation in Shared Fleets

A single customer or rogue microservice generating sudden inference spikes can consume entire organization provider quotas, starving critical production workloads.

Dual Constraints: Requests vs Tokens Capacity

Traditional rate limiters count requests, but in LLM applications, a single request carrying a 100,000-token prompt consumes vastly more compute than a 100-token query, requiring unified TPM throttling.

Sudden Upstream Provider 429 Cascades

When upstream model providers enforce strict tier ceilings, unthrottled application bursts receive abrupt HTTP 429 rejections that break client experiences without graceful queuing.

In-Memory State Synchronization Overhead

Distributed rate limiters relying on heavy database queries introduce unacceptable latency overheads during high-throughput inference traffic.

Architecture Taxonomy & Core Components

Preflight Phase 2 Slot Allocator

In-memory sliding window tracking RPM and estimated TPM limits with sub-millisecond check times.

Hierarchical Quota Resolver

Enforces nested rate limit thresholds across Organization > Team > Virtual Key boundaries.

Standardized Retry-After Headers

Emits RFC-compliant 429 responses with precise Retry-After durations to guide client backoff behavior.

nRouter In-Memory Rate Limiting Engine

nRouter evaluates rate limits during Preflight Phase 2 entirely within high-speed in-memory state structures. The gateway simultaneously evaluates requests per minute (RPM) and tokens per minute (TPM) against configured virtual key, team, and organization quotas. When limits are approached, nRouter gracefully coordinates client traffic, returning standardized HTTP 429 responses with accurate Retry-After headers to prevent upstream provider rejections. This multi-tier architecture guarantees fair tenant resource distribution while protecting global provider limits.

Read our guides on implementing sliding-window rate limiters, configuring token bucket algorithms, and preventing upstream 429 cascades.

Posts

Latest first

Budgets vs Rate Limits: Pick the Control, Then Set Both
Guides

Budgets vs Rate Limits: Pick the Control, Then Set Both

A budget caps dollars over a window and answers 402; a rate limit caps RPM/TPM right now and answers 429. Here is how to classify the risk, set each control in the dashboard, and write client code that tells the three rejections apart.

nRouter team
10 minRead →
Credits, Budgets, Rate Limits, Guardrails: Four Pre-Flight Gates
Guides

Credits, Budgets, Rate Limits, Guardrails: Four Pre-Flight Gates

Every request through an LLM gateway clears four independent gates before a provider ever sees it — credit balance, budget cap, RPM/TPM rate limit, and guardrails. Each has its own status code, its own scope, and its own fix.

nRouter team
11 minRead →
Reserve-and-Settle: Never Overspend a Credit Balance
Engineering

Reserve-and-Settle: Never Overspend a Credit Balance

Checking a balance and then calling a provider is a race, and under the fan-out an LLM gateway is built for it loses. Here is the reserve, settle and release contract as you can observe it — what your balance does on success, on an upstream failure, on a timeout, and when the cost is never knowable.

nRouter team
12 minRead →
Hard LLM Spend Caps at Org, Team, User, and Key Scope
Guides

Hard LLM Spend Caps at Org, Team, User, and Key Scope

A budget is a dollar allowance attached to a scope and a window. Here is how to create one on each of the four scopes, which status code each returns when it fires, and how to prove the cap bites before you trust it in production.

nRouter team
11 minRead →
429 vs 402 on an LLM Gateway: Which to Retry, Which to Stop
Guides

429 vs 402 on an LLM Gateway: Which to Retry, Which to Stop

Both a throttle and a budget block can arrive as 429, and both an empty balance and a team budget arrive as 402. Branch on the error code, not the status — with backoff, jitter, and idempotent retries.

nRouter team
9 minRead →
RPM and TPM Rate Limiting Per Key, Team, and Org
Engineering

RPM and TPM Rate Limiting Per Key, Team, and Org

Understand how RPM and TPM limits resolve per key, team, and org on an LLM gateway, why every 429 response names its source, and how budgets prevent loops.

nRouter team
11 minRead →