Browse documentation

Embeddings

Generate high-dimensional vector embeddings for text with nRouter. Compatible with OpenAI embeddings endpoints for semantic search, clustering, and RAG.

Last updated

The /v1/embeddings endpoint generates dense numerical vector representations of text strings, enabling semantic search, retrieval-augmented generation (RAG), vector similarity search, and automated text clustering. Operating with full OpenAI API wire compatibility, nRouter lets developers route embedding requests across providers like OpenAI, Cohere, Voyage, and Google Vertex AI with unified key governance, automated cloud failover, and zero-markup pricing.

POST https://api.nrouter.ai/v1/embeddings
Input Support
Batched Texts

Embed single sentences or arrays of up to 2,048 text documents in a single request.

Dimensionality
Flexible Vectors

Customize output embedding dimensions to balance vector DB storage and retrieval accuracy.

Cost Invariant
Zero Markup

All embedding models are billed at exact provider token rates with no per-token surcharge.

Telemetry
Token Auditing

Every response reports exact token counts, turnaround latency, and cost attribution headers.


Architectural Role & Lifecycle

Embedding generation is typically the highest-throughput endpoint in modern AI pipelines, feeding real-time RAG ingestion and vector indexing workloads. nRouter optimizes this path for minimum overhead:

  1. Phase 1: In-Memory Virtual Key Validation: Authenticates the sk-nrouter-... bearer token against memory structures in microseconds and validates model ACLs.
  2. Phase 2: High-Throughput TPM & RPM Checks: Validates tenant sliding-window rate limits, accommodating massive indexing batches without blocking.
  3. Phase 3: Input Screening & Preflight Safety: Evaluates input strings for malicious injection payloads. If a safety threshold is crossed, the call is rejected with HTTP 400 (x-nr-guardrails: blocked).
  4. Phase 4: Token-Level Credit Reservation: Calculates the exact input token count across the batch and reserves the corresponding credit balance. Blocked requests cost $0.
  5. Upstream Execution & Response Stamping: Dispatches the batch to the upstream provider cloud. Upon successful return, exact spend is settled and edge telemetry headers are stamped on the response.

Request Parameters

The request body must be a JSON object:

ParameterTypeRequiredDefaultDescription
modelstringYes—ID of the model to use (e.g. text-embedding-3-small, text-embedding-3-large, text-embedding-ada-002).
inputstring or string[]Yes—Input text to embed, encoded as a string or array of strings. Maximum input length depends on the model (e.g. 8,192 tokens for text-embedding-3-small).
dimensionsintegerNo—The number of dimensions the resulting output embeddings should have. Only supported in text-embedding-3-* and select modern embedding models.
encoding_formatstringNofloatThe format to return the embeddings in. Can be float or base64. Base64 significantly reduces JSON transfer size for large batches.
userstringNo—A unique identifier representing your end-user, helping track usage and detect platform abuse.

Response Payloads

Standard Embedding Response

{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "index": 0,
      "embedding": [
        -0.0069292834,
        -0.005336422,
        -0.024047505,
        0.008456123,
        0.012398451
      ]
    }
  ],
  "model": "text-embedding-3-small",
  "usage": {
    "prompt_tokens": 8,
    "total_tokens": 8
  }
}

Headers Reference

Inbound Request Headers

HeaderTypeRequiredDescription
AuthorizationstringYesBearer authentication format: Bearer sk-nrouter-....
Content-TypestringYesMust be application/json.
x-nr-tagsstringNoCost attribution metadata (e.g. index=kb_v2,env=prod).

Outbound Response Headers

HeaderTypeDescription
x-nr-request-idstringUnique correlation UUID assigned to this embedding request.
x-nr-latency-msintegerGateway edge turnaround time in milliseconds.
x-nr-request-costfloatExact USD cost computed from provider token rates.
x-nr-cost-statusstringexact when priced or unpriced if rate metadata is pending.
x-nr-modelstringPhysical model that served the embedding computation.
x-nr-routingstringRouting chain outcome: direct or fallback:<n>.
x-nr-attemptsintegerNumber of provider attempts made.
x-nr-guardrailsstringGuardrail evaluation result: none, monitor, pass, or blocked.
x-nr-input-tokensintegerTotal input tokens billed across all items in the batch.
x-nr-total-tokensintegerTotal tokens consumed.

SDK Code Examples

import { nRouter } from "@nrouter_ai/sdk";

const client = new nRouter({
  apiKey: process.env.NROUTER_API_KEY,
});

const embedding = await client.embeddings.create({
  model: "text-embedding-3-small",
  input: ["Semantic search query", "Vector database document chunk"],
  dimensions: 512,
});

console.log(`Generated ${embedding.data.length} vectors.`);
console.log(`Vector dimensions: ${embedding.data[0].embedding.length}`);
console.log(`Total tokens: ${embedding.usage.total_tokens}`);

Error Codes & Troubleshooting

{
  "error": {
    "type": "invalid_request",
    "message": "Input text exceeds maximum token context length of 8192.",
    "code": "input_too_large"
  }
}
HTTP StatusError CodeCauseRecommended Action
400 Bad Requestinvalid_requestInput exceeds maximum token ceiling, batch size exceeds limit, or invalid dimensions.Truncate document text or chunk into smaller paragraphs before submitting.
400 Bad Requestguardrail_blockedInput string violated safety policies or triggered injection detection.Inspect input content; blocked requests incur $0 cost.
401 Unauthorizedinvalid_api_keyVirtual key missing, expired, or invalid.Verify sk-nrouter-... key in organization dashboard.
402 Payment Requiredinsufficient_creditsOrganization prepaid balance is insufficient for token hold.Add credit balance via the dashboard billing interface.
404 Not Foundmodel_not_foundRequested embedding model is unrecognized or unauthorized.Query GET /v1/models to confirm active provider catalog entitlements.
429 Too Many Requestsrate_limit_exceededRPM or TPM limit exceeded for tenant or key.Inspect x-nr-limit-source header and apply exponential backoff.
500 / 503 Provider Errorservice_unavailableUpstream embedding provider experiencing degradation.nRouter automatically initiates fallback routing if alternative models are configured.
POST
/v1/embeddings

Authorization

NRouterApiKey
AuthorizationBearer <token>

Your nRouter virtual key (sk-nrouter-…). Sent as Authorization: Bearer sk-nrouter-….

In: header

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/v1/embeddings" \  -H "Content-Type: application/json" \  -d '{    "model": "text-embedding-3-small",    "input": "The quick brown fox jumps over the lazy dog."  }'
{  "object": "list",  "model": "text-embedding-3-small",  "data": [    {      "object": "embedding",      "index": 0,      "embedding": [        0.0023,        -0.009,        0.015      ]    }  ],  "usage": {    "prompt_tokens": 5,    "total_tokens": 5  }}
Was this page helpful?