Browse documentation

Ruby SDK

Connect Ruby applications to nRouter using the ruby-openai gem with custom base URI settings, virtual keys, server-side guardrails, and budget ceilings.

Last updated

Ruby and Ruby on Rails applications can interact directly with the nRouter unified AI gateway (https://api.nrouter.ai/v1) using the standard community ruby-openai gem. By configuring the base URI and providing your nRouter virtual key, your Rails controllers, background ActiveJob workers, and CLI scripts gain access to all premier model providers through a single integration.

Routing through nRouter applies organization-level safeguards automatically: server-side guardrails block prompt injection attempts before tokens reach downstream models, virtual keys enforce team-level monthly budget ceilings, and automatic failovers prevent downtime during upstream cloud outages.

Prerequisites & Installation

The package requires Ruby 3.0 or higher.

Add ruby-openai to your application's Gemfile:

gem "ruby-openai"

Then execute bundler:

bundle install

Or install directly via rubygems:

gem install ruby-openai

Setup & Configuration

Store your nRouter virtual API key in your environment or Rails credentials:

export NROUTER_API_KEY="sk-nrouter-your-virtual-key"

In Rails applications using encrypted credentials (credentials.yml.enc):

nrouter:
  api_key: sk-nrouter-your-virtual-key

Initializing the Client

Configure the client to point to the nRouter gateway:

require "openai"

client = OpenAI::Client.new(
  access_token: ENV.fetch("NROUTER_API_KEY"),
  uri_base: "https://api.nrouter.ai/v1",
)

In a Rails initializer (config/initializers/nrouter.rb):

OpenAI.configure do |config|
  config.access_token = ENV.fetch("NROUTER_API_KEY") { Rails.application.credentials.dig(:nrouter, :api_key) }
  config.uri_base = "https://api.nrouter.ai/v1"
  config.request_timeout = 45 # seconds
end

Configuration Parameters

Configure connection timeouts, retries, and custom Faraday middleware:

ParameterTypeDefaultDescription
access_tokenStringEnvironmentnRouter virtual key (sk-nrouter-...).
uri_baseStringhttps://api.nrouter.ai/v1Unified gateway endpoint. Must include /v1.
request_timeoutInteger120Request read timeout in seconds.
extra_headersHash{}Custom HTTP headers sent on every request (e.g. x-nr-routing).
require "openai"

client = OpenAI::Client.new(
  access_token: ENV.fetch("NROUTER_API_KEY"),
  uri_base: "https://api.nrouter.ai/v1",
  request_timeout: 30,
  extra_headers: {
    "x-nr-routing" => "latency"
  }
)

Implementation Patterns

1. Synchronous Chat Completion

Send structured user messages and parse the completion:

require "openai"

client = OpenAI::Client.new(
  access_token: ENV.fetch("NROUTER_API_KEY"),
  uri_base: "https://api.nrouter.ai/v1",
)

response = client.chat(
  parameters: {
    model: "gpt-5.4-mini",
    messages: [
      { role: "system", content: "You are a senior Ruby on Rails architect." },
      { role: "user", content: "Explain how to structure background jobs with Sidekiq." },
    ],
  }
)

puts response.dig("choices", 0, "message", "content")

2. Server-Sent Events (SSE) Streaming

Stream tokens in real-time using a Ruby block:

require "openai"

client = OpenAI::Client.new(
  access_token: ENV.fetch("NROUTER_API_KEY"),
  uri_base: "https://api.nrouter.ai/v1",
)

client.chat(
  parameters: {
    model: "claude-haiku-4-5-20251001",
    messages: [{ role: "user", content: "Write a short poem about clean code." }],
    stream: proc do |chunk, _bytesize|
      delta = chunk.dig("choices", 0, "delta", "content")
      print delta if delta
    end
  }
)
puts

3. Per-Request Gateway Overrides

ruby-openai forwards unknown parameters to nRouter, allowing prompt templates and cache settings:

# Execute a dashboard prompt template with variables
response = client.chat(
  parameters: {
    model: "gpt-5.5",
    messages: [{ role: "user", content: "Summarize Q1 churn metrics." }],
    nrouter_prompt_template_id: "tmpl_churn_analysis_v1",
    nrouter_prompt_variables: { segment: "enterprise" },
    nrouter_cache: true,
  }
)

# Explicitly bypass semantic cache
response = client.chat(
  parameters: {
    model: "gpt-5.5",
    messages: [{ role: "user", content: "What is the current system time?" }],
    nrouter_cache: false,
  }
)

Production Best Practices

Deterministic Routing & Fallbacks

Ensure continuous uptime during provider maintenance by designating fallback models:

response = client.chat(
  parameters: {
    # Primary model with automatic fallback
    model: "gpt-5.4-mini,claude-haiku-4-5-20251001",
    messages: [{ role: "user", content: "Verify customer account credentials." }],
  }
)
  • x-nr-routing: latency: Directs requests to the lowest-latency active provider deployment.
  • x-nr-routing: cost: Prioritizes the most cost-effective deployment matching the model specification.
  • Model Fallback Chain: If the primary provider reports 5xx errors or capacity exhaustion, nRouter fails over instantly to the fallback model.

Response Telemetry & FinOps Tracking

Every successful response through nRouter carries transparent telemetry headers:

  • x-nr-request-id — Unique ID for the call and join key for the spend ledger.
  • x-nr-model — The specific provider model that fulfilled the request.
  • x-nr-cost-status — exact when fully priced, unpriced if provider is unlisted.
  • x-nr-request-cost — Exact USD spend incurred for this call.
  • x-nr-input-tokens, x-nr-output-tokens, x-nr-total-tokens — Token accounting figures.

In Rails background jobs, log these headers to attribute AI costs to specific tenants or accounts.

Troubleshooting & Error Handling

Errors returned by nRouter map directly to HTTP status codes.

Common Error Codes

StatusCodeCauseRecommended Action
400guardrail_blockedInput rejected by server-side content or injection guardrailsVerify prompt safety; review guardrail settings in nRouter dashboard.
401authentication_errorMissing, incorrect, or expired virtual API keyCheck NROUTER_API_KEY environment variable.
402insufficient_creditsZero organization balance or key spending limit reachedAdd funds in dashboard or update key ceiling.
429rate_limit_exceededClient exceeded RPM/TPM quotaImplement exponential backoff; check retry headers.
500 / 503service_unavailableDownstream provider error or network disruptionUse fallback model lists (model1,model2).

Error Catching Pattern

require "openai"

client = OpenAI::Client.new(
  access_token: ENV.fetch("NROUTER_API_KEY"),
  uri_base: "https://api.nrouter.ai/v1",
)

begin
  response = client.chat(
    parameters: {
      model: "gpt-5.4-mini",
      messages: [{ role: "user", content: "Run background task" }],
    }
  )
rescue Faraday::ClientError => e
  status = e.response[:status]
  body = e.response[:body]

  case status
  when 400
    puts "Bad request / guardrail violation: #{body}"
  when 401
    puts "Authentication failed: Verify NROUTER_API_KEY."
  when 402
    puts "Payment required: Virtual key budget or account credits exhausted."
  when 429
    puts "Rate limit exceeded. Apply backoff retry."
  else
    puts "Gateway error [#{status}]: #{body}"
  end
rescue Faraday::ServerError => e
  puts "Downstream provider error: #{e.message}"
end

Next Steps

Was this page helpful?