Browse documentation
API ReferenceCount Tokens POST

Count Tokens

Calculate token count for an Anthropic Messages payload without an inference call. Essential for preflight budget verification and context window packing.

Last updated

The /v1/messages/count_tokens endpoint calculates the exact input token count that a specific Anthropic Messages payload would consume when sent to Claude models. Because no generative tokens are produced and no provider inference is triggered, the operation completes with sub-10ms edge turnaround times and incurs $0 in inference cost.

POST https://api.nrouter.ai/v1/messages/count_tokens
Cost Invariant
$0 Inference Spend

Runs token counting locally or via lightweight tokenizers with zero provider token billing.

Turnaround Time
Sub-10ms Latency

Instant edge tokenization enables pre-call validation in real-time agent loops.

Context Safeguard
Prevent 400 Errors

Verify document context fits model ceilings before attempting full inference.

Tool & System
Schema Accounting

Accurately counts JSON tool schema overhead and system instructions.


Architectural Role & Lifecycle

In complex agent workflows and retrieval-augmented generation (RAG) pipelines, context packing is critical. Sending prompts that exceed a model's context ceiling results in failed requests and degraded user experiences:

  1. Phase 1: In-Memory Key Auth & Model Verification: Validates the virtual key hash (sk-nrouter-...) and checks tenant entitlements for the target Claude model.
  2. Phase 2: Rate Limit Accounting: Applies lightweight rate-limiting to protect gateway tokenizers against denial-of-service spikes.
  3. Payload Tokenization: Computes the exact tokenizer representation across all input structures, including system blocks, multi-turn messages, multimodal image data, and tools function schemas.
  4. Immediate Edge Return: Returns the integer input_tokens count immediately without invoking provider inference or reserving financial credits.

Request Parameters

The request body must be a JSON object:

ParameterTypeRequiredDefaultDescription
modelstringYes—Target Claude model identifier (e.g. claude-sonnet-4-5-20250929, claude-opus-4-20250514).
messagesarrayYes—Array of conversational messages containing role and content.
systemstring or arrayNo—System context string or content block array.
toolsarrayNo—Tool definitions whose JSON schemas will be parsed into token counts.
thinkingobjectNo—Extended thinking configuration object.

Response Payload

The response returns a lightweight JSON object containing the exact token count:

{
  "input_tokens": 142
}

Headers Reference

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be application/json.
anthropic-versionstringNoCompatibility header: 2023-06-01.

Outbound Response Headers

HeaderTypeDescription
x-nr-request-idstringUnique UUID correlation identifier for tracing.
x-nr-latency-msintegerGateway edge turnaround time in milliseconds.
x-nr-modelstringThe target model whose tokenizer was evaluated.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const tokenCount = await client.messages.countTokens({
  model: "claude-sonnet-4-5-20250929",
  system: "You are an enterprise legal contract analysis agent.",
  messages: [
    { role: "user", content: "Analyze this 50-page Master Services Agreement..." }
  ],
});

console.log(`Estimated Input Tokens: ${tokenCount.input_tokens}`);
if (tokenCount.input_tokens > 180000) {
  console.warn("Approaching context ceiling; apply chunking.");
}

Error Handling & Status Codes

{
  "error": {
    "type": "gateway_error",
    "message": "messages: field required",
    "code": "invalid_request"
  }
}
HTTP StatusError CodeRoot CauseRemediation
400 Bad Requestinvalid_requestMissing required model or messages field, or malformed message structure.Validate payload against Anthropic Messages structure.
401 Unauthorizedinvalid_api_keyVirtual key missing, expired, or invalid.Check sk-nrouter-... key in organization dashboard.
404 Not Foundmodel_not_foundTarget model not recognized in catalog.Verify model identifier against GET /v1/models.
429 Too Many Requestsrate_limit_exceededAccount RPM ceiling exceeded for token counting operations.Apply exponential backoff.
500 / 503 Gateway Errorservice_unavailableInternal tokenizer service unavailable.Retry request; no provider inference costs are incurred.

Token Preflight & Cost Estimation Architecture

Accurately calculating input token count before dispatching expensive LLM inference calls is fundamental to predictable FinOps and latency management.

Exact Byte-Pair Encoding (BPE) Emulation

The count_tokens endpoint runs the provider's exact native tokenizer within the Rust gateway runtime without sending any egress traffic to the model provider:

  1. Zero Provider Cost: Token counting queries are completely free from upstream model token charges.
  2. Context Ceiling Enforcement: Verifies that combined system prompts, user queries, document embeddings, and tool declarations do not exceed the model's maximum context window (e.g. 200,000 tokens for Claude 3.5 Sonnet).
  3. Pre-Call Budget Protection: Combine token counts with pricing.prompt to calculate the exact upfront cost of an inference call before making a commitment.
  4. Dynamic Sliding Window Pruning: Use the token count response to truncate chat conversation histories dynamically when approaching context ceilings.

Token Accounting for Multimodal & Tool Payloads

Modern LLM workflows involve complex payloads beyond simple text strings:

  • Tool and Function Declarations: JSON schema parameters for tools consume tokens. The count_tokens endpoint accurately serializes and counts all function schemas according to provider-specific token formatting rules.
  • Multimodal Image Tokens: When messages include base64 or hosted image URLs, the token counter calculates exact visual token consumption based on image dimensions and tile segmentation algorithms.
  • Prompt Caching Discounts: For models supporting prompt caching (e.g. Anthropic Claude 3.5 Sonnet, Claude 3 Opus), count_tokens provides visibility into the cache-eligible prefix length, allowing teams to optimize cache breakpoints and cut recurring input costs by up to 90%.
  • Zero-Latency In-Memory Execution: Because tokenization runs on compiled Rust SIMD routines locally inside the gateway instance, count_tokens requests resolve with sub-5ms latency.
POST
/v1/messages/count_tokens

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Header Parameters

x-nr-compress?string

Request prompt compression: on to compress eligible prompts, off to skip

Value in

  • "on"
  • "off"
x-nr-tags?string

Custom spend and attribution tags (comma-separated key=value pairs)

x-nr-mcp-server?string

Target MCP server ID when routing MCP tool calls or prompts through the gateway

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/messages/count_tokens" \  -H "Content-Type: application/json" \  -d '{    "model": "string",    "messages": [      {}    ]  }'
{  "input_tokens": 9}
Was this page helpful?