Embeddings
Generate high-dimensional vector embeddings for text with nRouter. Compatible with OpenAI embeddings endpoints for semantic search, clustering, and RAG.
Last updated
The /v1/embeddings endpoint generates dense numerical vector representations of text strings, enabling semantic search, retrieval-augmented generation (RAG), vector similarity search, and automated text clustering. Operating with full OpenAI API wire compatibility, nRouter lets developers route embedding requests across providers like OpenAI, Cohere, Voyage, and Google Vertex AI with unified key governance, automated cloud failover, and zero-markup pricing.
POST https://api.nrouter.ai/v1/embeddingsEmbed single sentences or arrays of up to 2,048 text documents in a single request.
Customize output embedding dimensions to balance vector DB storage and retrieval accuracy.
All embedding models are billed at exact provider token rates with no per-token surcharge.
Every response reports exact token counts, turnaround latency, and cost attribution headers.
Architectural Role & Lifecycle
Embedding generation is typically the highest-throughput endpoint in modern AI pipelines, feeding real-time RAG ingestion and vector indexing workloads. nRouter optimizes this path for minimum overhead:
- Phase 1: In-Memory Virtual Key Validation: Authenticates the
sk-nrouter-...bearer token against memory structures in microseconds and validates model ACLs. - Phase 2: High-Throughput TPM & RPM Checks: Validates tenant sliding-window rate limits, accommodating massive indexing batches without blocking.
- Phase 3: Input Screening & Preflight Safety: Evaluates input strings for malicious injection payloads. If a safety threshold is crossed, the call is rejected with HTTP 400 (
x-nr-guardrails: blocked). - Phase 4: Token-Level Credit Reservation: Calculates the exact input token count across the batch and reserves the corresponding credit balance. Blocked requests cost $0.
- Upstream Execution & Response Stamping: Dispatches the batch to the upstream provider cloud. Upon successful return, exact spend is settled and edge telemetry headers are stamped on the response.
Request Parameters
The request body must be a JSON object:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Yes | — | ID of the model to use (e.g. text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002). |
input | string or string[] | Yes | — | Input text to embed, encoded as a string or array of strings. Maximum input length depends on the model (e.g. 8,192 tokens for text-embedding-3-small). |
dimensions | integer | No | — | The number of dimensions the resulting output embeddings should have. Only supported in text-embedding-3-* and select modern embedding models. |
encoding_format | string | No | float | The format to return the embeddings in. Can be float or base64. Base64 significantly reduces JSON transfer size for large batches. |
user | string | No | — | A unique identifier representing your end-user, helping track usage and detect platform abuse. |
Response Payloads
Standard Embedding Response
{
"object": "list",
"data": [
{
"object": "embedding",
"index": 0,
"embedding": [
-0.0069292834,
-0.005336422,
-0.024047505,
0.008456123,
0.012398451
]
}
],
"model": "text-embedding-3-small",
"usage": {
"prompt_tokens": 8,
"total_tokens": 8
}
}Headers Reference
Inbound Request Headers
| Header | Type | Required | Description |
|---|---|---|---|
Authorization | string | Yes | Bearer authentication format: Bearer sk-nrouter-.... |
Content-Type | string | Yes | Must be application/json. |
x-nr-tags | string | No | Cost attribution metadata (e.g. index=kb_v2,env=prod). |
Outbound Response Headers
| Header | Type | Description |
|---|---|---|
x-nr-request-id | string | Unique correlation UUID assigned to this embedding request. |
x-nr-latency-ms | integer | Gateway edge turnaround time in milliseconds. |
x-nr-request-cost | float | Exact USD cost computed from provider token rates. |
x-nr-cost-status | string | exact when priced or unpriced if rate metadata is pending. |
x-nr-model | string | Physical model that served the embedding computation. |
x-nr-routing | string | Routing chain outcome: direct or fallback:<n>. |
x-nr-attempts | integer | Number of provider attempts made. |
x-nr-guardrails | string | Guardrail evaluation result: none, monitor, pass, or blocked. |
x-nr-input-tokens | integer | Total input tokens billed across all items in the batch. |
x-nr-total-tokens | integer | Total tokens consumed. |
SDK Code Examples
import { nRouter } from "@nrouter_ai/sdk";
const client = new nRouter({
apiKey: process.env.NROUTER_API_KEY,
});
const embedding = await client.embeddings.create({
model: "text-embedding-3-small",
input: ["Semantic search query", "Vector database document chunk"],
dimensions: 512,
});
console.log(`Generated ${embedding.data.length} vectors.`);
console.log(`Vector dimensions: ${embedding.data[0].embedding.length}`);
console.log(`Total tokens: ${embedding.usage.total_tokens}`);Error Codes & Troubleshooting
{
"error": {
"type": "invalid_request",
"message": "Input text exceeds maximum token context length of 8192.",
"code": "input_too_large"
}
}| HTTP Status | Error Code | Cause | Recommended Action |
|---|---|---|---|
| 400 Bad Request | invalid_request | Input exceeds maximum token ceiling, batch size exceeds limit, or invalid dimensions. | Truncate document text or chunk into smaller paragraphs before submitting. |
| 400 Bad Request | guardrail_blocked | Input string violated safety policies or triggered injection detection. | Inspect input content; blocked requests incur $0 cost. |
| 401 Unauthorized | invalid_api_key | Virtual key missing, expired, or invalid. | Verify sk-nrouter-... key in organization dashboard. |
| 402 Payment Required | insufficient_credits | Organization prepaid balance is insufficient for token hold. | Add credit balance via the dashboard billing interface. |
| 404 Not Found | model_not_found | Requested embedding model is unrecognized or unauthorized. | Query GET /v1/models to confirm active provider catalog entitlements. |
| 429 Too Many Requests | rate_limit_exceeded | RPM or TPM limit exceeded for tenant or key. | Inspect x-nr-limit-source header and apply exponential backoff. |
| 500 / 503 Provider Error | service_unavailable | Upstream embedding provider experiencing degradation. | nRouter automatically initiates fallback routing if alternative models are configured. |
Authorization
NRouterApiKey Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….
In: header
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
application/json
curl -X POST "https://example.com/v1/embeddings" \ -H "Content-Type: application/json" \ -d '{ "model": "text-embedding-3-small", "input": "The quick brown fox jumps over the lazy dog." }'{ "object": "list", "model": "text-embedding-3-small", "data": [ { "object": "embedding", "index": 0, "embedding": [ 0.0023, -0.009, 0.015 ] } ], "usage": { "prompt_tokens": 5, "total_tokens": 5 }}Count Tokens POST
Calculate token count for an Anthropic Messages payload without an inference call. Essential for preflight budget verification and context window packing.
Models
List all available LLMs and multimodal AI models supported by nRouter with real-time list pricing, context window limits, and tenant capability details.