Browse documentation

Java SDK

Use the official nRouter Java SDK with Maven Central dependencies, automated environment configuration, live cost tracking, and full OpenAI API parity.

Last updated

The official nRouter Java SDK (ai.nrouter:nrouter-sdk) provides enterprise-grade, thread-safe bindings for the nRouter unified AI gateway (https://api.nrouter.ai/v1). Available on Maven Central, it connects directly to the edge gateway, auto-resolves your virtual API key (NROUTER_API_KEY), and provides drop-in compatibility with the official OpenAI Java client while enforcing organization guardrails, budget limits, and real-time cost tracking.

By proxying enterprise Java applications through nRouter, development teams gain centralized governance over LLM usage: automated prompt injection filtering, unified provider billing without per-token markups, and cross-provider failover without modifying client service logic.

Prerequisites & Installation

The SDK requires Java 17 or higher (Java 21 LTS supported).

Maven Installation

Add the dependency to your pom.xml:

<dependency>
  <groupId>ai.nrouter</groupId>
  <artifactId>nrouter-sdk</artifactId>
  <version>2.2.1</version>
</dependency>

Gradle Installation

Add to your build.gradle:

implementation 'ai.nrouter:nrouter-sdk:2.2.1'

Or for Gradle Kotlin DSL (build.gradle.kts):

implementation("ai.nrouter:nrouter-sdk:2.2.1")

Setup & Configuration

Store your nRouter virtual API key in your environment:

export NROUTER_API_KEY="sk-nrouter-your-virtual-key"

Initializing the Client

Initialize the client using the static factory method or the builder:

import ai.nrouter.sdk.NRouter;
import com.openai.client.OpenAIClient;

// Automatically reads NROUTER_API_KEY from environment and targets https://api.nrouter.ai/v1
OpenAIClient client = NRouter.create();

Configuration Parameters

Configure explicit keys, custom endpoints, and HTTP timeouts via the builder:

ParameterTypeDefaultDescription
apiKeyStringEnvironmentnRouter virtual key (sk-nrouter-...).
baseUrlStringhttps://api.nrouter.ai/v1Unified gateway base URL. Must include /v1.
timeoutDurationDuration.ofSeconds(60)HTTP connection and socket read timeout.
maxRetriesint2Client-side retry count for network failures.
headerString, StringEmptyCustom HTTP headers sent on every request (e.g. x-nr-routing).
import java.time.Duration;
import ai.nrouter.sdk.NRouter;
import com.openai.client.OpenAIClient;

OpenAIClient client = NRouter.builder()
    .apiKey("sk-nrouter-your-virtual-key")
    .baseUrl("https://api.nrouter.ai/v1")
    .timeout(Duration.ofSeconds(45))
    .maxRetries(3)
    .header("x-nr-routing", "latency")
    .build();

Implementation Patterns

1. Synchronous Chat Completion

Create a structured chat completion request:

import ai.nrouter.sdk.NRouter;
import com.openai.client.OpenAIClient;
import com.openai.models.chat.completions.ChatCompletion;
import com.openai.models.chat.completions.ChatCompletionCreateParams;
import com.openai.models.ChatCompletionMessageParam;
import com.openai.models.ChatCompletionUserMessageParam;

public class NRouterExample {
    public static void main(String[] args) {
        OpenAIClient client = NRouter.create();

        ChatCompletion response = client.chat().completions().create(
            ChatCompletionCreateParams.builder()
                .model("gpt-5.4-mini")
                .addMessage(ChatCompletionMessageParam.ofUser(
                    ChatCompletionUserMessageParam.builder()
                        .content("Explain garbage collection in modern JVMs.")
                        .build()
                ))
                .build()
        );

        String content = response.choices().get(0).message().content().orElse("");
        System.out.println("Response:\n" + content);
    }
}

2. Asynchronous Completions

For non-blocking reactive microservices (Spring Boot WebFlux, Quarkus, Micronaut):

import java.util.concurrent.CompletableFuture;
import ai.nrouter.sdk.NRouter;
import com.openai.client.OpenAIClient;
import com.openai.models.chat.completions.ChatCompletion;
import com.openai.models.chat.completions.ChatCompletionCreateParams;
import com.openai.models.ChatCompletionMessageParam;
import com.openai.models.ChatCompletionUserMessageParam;

public class AsyncExample {
    public static void main(String[] args) {
        OpenAIClient client = NRouter.create();

        ChatCompletionCreateParams params = ChatCompletionCreateParams.builder()
            .model("gpt-5.4-mini")
            .addMessage(ChatCompletionMessageParam.ofUser(
                ChatCompletionUserMessageParam.builder()
                    .content("Summarize Spring Boot 3 virtual threads.")
                    .build()
            ))
            .build();

        CompletableFuture<ChatCompletion> future = client.async().chat().completions().create(params);

        future.thenAccept(res -> {
            System.out.println(res.choices().get(0).message().content().orElse(""));
        }).join();
    }
}

3. Server-Sent Events (SSE) Streaming

Stream token chunks in real-time:

import ai.nrouter.sdk.NRouter;
import com.openai.client.OpenAIClient;
import com.openai.models.chat.completions.ChatCompletionCreateParams;
import com.openai.models.ChatCompletionMessageParam;
import com.openai.models.ChatCompletionUserMessageParam;

public class StreamingExample {
    public static void main(String[] args) {
        OpenAIClient client = NRouter.create();

        ChatCompletionCreateParams params = ChatCompletionCreateParams.builder()
            .model("claude-haiku-4-5-20251001")
            .addMessage(ChatCompletionMessageParam.ofUser(
                ChatCompletionUserMessageParam.builder()
                    .content("Write a haiku about enterprise architecture.")
                    .build()
            ))
            .build();

        client.chat().completions().createStreaming(params).stream().forEach(chunk -> {
            chunk.choices().get(0).delta().content().ifPresent(System.out::print);
        });
        System.out.println();
    }
}

4. Per-Request Gateway Overrides

To pass prompt template IDs or bypass cache, attach parameters via raw HTTP or extra parameters:

{
  "model": "gpt-5.5",
  "messages": [{"role": "user", "content": "Summarize Q1 earnings report."}],
  "nrouter_prompt_template_id": "tmpl_corporate_brief_v1",
  "nrouter_prompt_variables": {"market": "EMEA"},
  "nrouter_cache": false
}

Production Best Practices

Deterministic Routing & Fallbacks

Ensure enterprise reliability by configuring fallback models and routing preferences:

ChatCompletionCreateParams params = ChatCompletionCreateParams.builder()
    // Specify primary and secondary fallback models
    .model("gpt-5.4-mini,claude-haiku-4-5-20251001")
    .addMessage(ChatCompletionMessageParam.ofUser(
        ChatCompletionUserMessageParam.builder()
            .content("Verify customer identity record.")
            .build()
    ))
    .build();
  • x-nr-routing: latency: Directs requests to the fastest cloud region and provider.
  • x-nr-routing: cost: Ensures execution on the lowest-cost available provider deployment.
  • Automatic Fallback: If the primary provider reports 5xx errors or capacity exhaustion, nRouter fails over seamlessly to the secondary model.

Connection Pooling & Thread Safety

  1. Singleton Client: The OpenAIClient is thread-safe and should be registered as a singleton Spring bean or Guice provider.
  2. HTTP Client Tuning: The SDK utilizes java.net.http.HttpClient with persistent HTTP/2 connection pooling.

Response Telemetry & FinOps Tracking

Every call through nRouter produces metadata headers:

  • x-nr-request-id: Trace UUID linking backend request spans to nRouter spend records.
  • x-nr-model: The specific model instance that served the completion.
  • x-nr-cost-status: Status indicator (exact or unpriced).
  • x-nr-request-cost: Exact USD cost charged for the call.
  • x-nr-input-tokens / x-nr-output-tokens: Exact provider token usage figures.

Troubleshooting & Error Handling

Errors map to standard HTTP status codes.

Common Error Codes

StatusCodeCauseRecommended Action
400guardrail_blockedInput rejected by server-side content or injection guardrailsVerify prompt safety; review guardrail settings in nRouter dashboard.
401authentication_errorMissing, incorrect, or expired virtual API keyCheck NROUTER_API_KEY environment variable.
402insufficient_creditsZero organization balance or key spending limit reachedAdd funds in dashboard or update key ceiling.
429rate_limit_exceededClient exceeded RPM/TPM quotaImplement exponential backoff; check retry headers.
500 / 503service_unavailableDownstream provider error or network disruptionUse fallback model lists (model1,model2).

Error Catching Pattern

import com.openai.errors.OpenAIException;
import com.openai.errors.BadRequestException;
import com.openai.errors.AuthenticationException;
import com.openai.errors.RateLimitException;

try {
    ChatCompletion response = client.chat().completions().create(params);
} catch (BadRequestException e) {
    if (e.getMessage().contains("guardrail")) {
        System.err.println("Rejected by nRouter safety guardrail policy.");
    } else {
        System.err.println("Invalid request parameters: " + e.getMessage());
    }
} catch (AuthenticationException e) {
    System.err.println("Authentication failed: Verify NROUTER_API_KEY.");
} catch (RateLimitException e) {
    System.err.println("Rate limit reached. Apply exponential backoff.");
} catch (OpenAIException e) {
    System.err.println("Gateway error: " + e.getMessage());
}

Next Steps

Was this page helpful?