Count Tokens
Calculate token count for an Anthropic Messages payload without an inference call. Essential for preflight budget verification and context window packing.
Last updated
The /v1/messages/count_tokens endpoint calculates the exact input token count that a specific Anthropic Messages payload would consume when sent to Claude models. Because no generative tokens are produced and no provider inference is triggered, the operation completes with sub-10ms edge turnaround times and incurs $0 in inference cost.
POST https://api.nrouter.ai/v1/messages/count_tokensRuns token counting locally or via lightweight tokenizers with zero provider token billing.
Instant edge tokenization enables pre-call validation in real-time agent loops.
Verify document context fits model ceilings before attempting full inference.
Accurately counts JSON tool schema overhead and system instructions.
Architectural Role & Lifecycle
In complex agent workflows and retrieval-augmented generation (RAG) pipelines, context packing is critical. Sending prompts that exceed a model's context ceiling results in failed requests and degraded user experiences:
- Phase 1: In-Memory Key Auth & Model Verification: Validates the virtual key hash (
sk-nrouter-...) and checks tenant entitlements for the target Claude model. - Phase 2: Rate Limit Accounting: Applies lightweight rate-limiting to protect gateway tokenizers against denial-of-service spikes.
- Payload Tokenization: Computes the exact tokenizer representation across all input structures, including
systemblocks, multi-turnmessages, multimodal image data, andtoolsfunction schemas. - Immediate Edge Return: Returns the integer
input_tokenscount immediately without invoking provider inference or reserving financial credits.
Request Parameters
The request body must be a JSON object:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | — | Target Claude model identifier (e.g. claude-sonnet-4-5-20250929, claude-opus-4-20250514). |
messages | array | Yes | — | Array of conversational messages containing role and content. |
system | string or array | No | — | System context string or content block array. |
tools | array | No | — | Tool definitions whose JSON schemas will be parsed into token counts. |
thinking | object | No | — | Extended thinking configuration object. |
Response Payload
The response returns a lightweight JSON object containing the exact token count:
{
"input_tokens": 142
}Headers Reference
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
anthropic-version | string | No | Compatibility header: 2023-06-01. |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique UUID correlation identifier for tracing. |
x-nr-latency-ms | integer | Gateway edge turnaround time in milliseconds. |
x-nr-model | string | The target model whose tokenizer was evaluated. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const tokenCount = await client.messages.countTokens({
model: "claude-sonnet-4-5-20250929",
system: "You are an enterprise legal contract analysis agent.",
messages: [
{ role: "user", content: "Analyze this 50-page Master Services Agreement..." }
],
});
console.log(`Estimated Input Tokens: ${tokenCount.input_tokens}`);
if (tokenCount.input_tokens > 180000) {
console.warn("Approaching context ceiling; apply chunking.");
}Error Handling & Status Codes
{
"error": {
"type": "gateway_error",
"message": "messages: field required",
"code": "invalid_request"
}
}| HTTP Status | Error Code | Root Cause | Remediation |
|---|---|---|---|
| 400 Bad Request | invalid_request | Missing required model or messages field, or malformed message structure. | Validate payload against Anthropic Messages structure. |
| 401 Unauthorized | invalid_api_key | Virtual key missing, expired, or invalid. | Check sk-nrouter-... key in organization dashboard. |
| 404 Not Found | model_not_found | Target model not recognized in catalog. | Verify model identifier against GET /v1/models. |
| 429 Too Many Requests | rate_limit_exceeded | Account RPM ceiling exceeded for token counting operations. | Apply exponential backoff. |
| 500 / 503 Gateway Error | service_unavailable | Internal tokenizer service unavailable. | Retry request; no provider inference costs are incurred. |
Token Preflight & Cost Estimation Architecture
Accurately calculating input token count before dispatching expensive LLM inference calls is fundamental to predictable FinOps and latency management.
Exact Byte-Pair Encoding (BPE) Emulation
The count_tokens endpoint runs the provider's exact native tokenizer within the Rust gateway runtime without sending any egress traffic to the model provider:
- Zero Provider Cost: Token counting queries are completely free from upstream model token charges.
- Context Ceiling Enforcement: Verifies that combined system prompts, user queries, document embeddings, and tool declarations do not exceed the model's maximum context window (e.g. 200,000 tokens for Claude 3.5 Sonnet).
- Pre-Call Budget Protection: Combine token counts with
pricing.promptto calculate the exact upfront cost of an inference call before making a commitment. - Dynamic Sliding Window Pruning: Use the token count response to truncate chat conversation histories dynamically when approaching context ceilings.
Token Accounting for Multimodal & Tool Payloads
Modern LLM workflows involve complex payloads beyond simple text strings:
- Tool and Function Declarations: JSON schema parameters for tools consume tokens. The
count_tokensendpoint accurately serializes and counts all function schemas according to provider-specific token formatting rules. - Multimodal Image Tokens: When messages include base64 or hosted image URLs, the token counter calculates exact visual token consumption based on image dimensions and tile segmentation algorithms.
- Prompt Caching Discounts: For models supporting prompt caching (e.g. Anthropic Claude 3.5 Sonnet, Claude 3 Opus),
count_tokensprovides visibility into the cache-eligible prefix length, allowing teams to optimize cache breakpoints and cut recurring input costs by up to 90%. - Zero-Latency In-Memory Execution: Because tokenization runs on compiled Rust SIMD routines locally inside the gateway instance,
count_tokensrequests resolve with sub-5ms latency.
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Header Parameters
Request prompt compression: on to compress eligible prompts, off to skip
Value in
- "on"
- "off"
Target MCP server ID when routing MCP tool calls or prompts through the gateway
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/messages/count_tokens" \ -H "Content-Type: application/json" \ -d '{ "model": "string", "messages": [ {} ] }'{ "input_tokens": 9}