Browse documentation

Meta Llama Models

Access Meta Llama 3.3 70B, 3.1 405B, and Llama 4 Scout via nRouter's unified gateway with real-time streaming, guardrails, and zero-markup list pricing.

Last updated

Meta Llama models provide industry-standard open weights performance across reasoning, coding, multilingual tasks, and multimodal vision. Through nRouter (https://api.nrouter.ai/v1), your applications can call Meta's flagship instruction-tuned models with one virtual API key, enterprise budget controls, pre-flight safety guardrails, and flat list pricing with zero per-token markup.


Available Meta Models

nRouter serves three primary Meta Llama models, each tailored for specific architectural requirements:

Model IDArchitecture & ParametersModalityContext WindowMax OutputFlat List Price (Input / Output per 1M)Primary Workload
meta/llama-3.3-70b-instructDense 70B parametersText in / Text out128k tokens4,096 tokens$0.72 / $0.72High-throughput reasoning, code generation, tool use
meta/llama-3.1-405b-instructDense 405B parametersText in / Text out128k tokens4,096 tokens$2.00 / $2.00Frontier-class logic, synthetic data generation, model distillation
meta/llama-4-scout-17b-16e-instructMixture of Experts (17B active / 16 experts)Text + Vision in / Text out128k tokens8,192 tokens$0.30 / $0.60Multimodal document analysis, low-latency agentic loops

Model Highlights

meta/llama-3.3-70b-instruct

The default workhorse for enterprise applications. It rivals previous-generation 405B models on benchmark evaluation suites while operating at a fraction of the compute cost and latency. Supports native function calling, structured JSON output, and complex multi-turn dialogs.

meta/llama-3.1-405b-instruct

Meta's flagship open-weights model. Delivering frontier-class performance comparable to top proprietary models, Llama 3.1 405B is optimized for difficult math and algorithmic reasoning, cross-domain knowledge synthesis, and generating high-quality training and evaluation datasets.

meta/llama-4-scout-17b-16e-instruct

Built on Meta's next-generation Llama 4 Mixture-of-Experts (MoE) architecture. With 17B active parameters out of 16 experts, it delivers ultra-fast time-to-first-token (TTFT) and supports native multimodal image inputs alongside text prompts.


Flat List Price Guarantee

Under nRouter's pricing policy (Rule #28), all models are served at flat provider list rates with zero per-token markup.

  • Exact List Settlement: Your organization pays exactly the listed token price ($0.72/$0.72 for 70B, $2.00/$2.00 for 405B, $0.30/$0.60 for Llama 4 Scout). nRouter makes zero profit on tokens.
  • Phase 4 Credit Reservation: When a request is dispatched, nRouter reserves an estimated credit envelope against your organization balance. Upon response completion, the exact token count is calculated and the remaining hold is immediately returned to your balance.
  • Cost Header Audit: Every response includes the exact billed amount in the x-nr-request-cost header (e.g. x-nr-request-cost: $0.000412).

Endpoint Format & Request Structure

Meta Llama models are served via the standard OpenAI-compatible /v1/chat/completions endpoint:

POST https://api.nrouter.ai/v1/chat/completions

Standard Text Completion

curl https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.3-70b-instruct",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise engineering assistant."
      },
      {
        "role": "user",
        "content": "Explain the difference between dense and sparse MoE architectures."
      }
    ],
    "temperature": 0.2,
    "max_tokens": 1024
  }'

Multimodal Vision (meta/llama-4-scout-17b-16e-instruct)

Llama 4 Scout accepts image inputs via standard URL or Base64 data encodings:

curl https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-4-scout-17b-16e-instruct",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "text", "text": "Analyze the architecture diagram in this image:" },
          {
            "type": "image_url",
            "image_url": {
              "url": "https://nrouter.ai/diagrams/gateway-flow.png"
            }
          }
        ]
      }
    ],
    "max_tokens": 1000
  }'

Real-Time Streaming

Set "stream": true to receive real-time Server-Sent Events (SSE). nRouter forwards tokens with sub-millisecond gateway overhead while maintaining observability and credit tracking.

Streaming with cURL

curl -N https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.3-70b-instruct",
    "messages": [
      { "role": "user", "content": "Write a fast Rust binary search implementation." }
    ],
    "stream": true
  }'

The gateway streams chunks formatted as data: {...}:

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1709251200,"model":"meta/llama-3.3-70b-instruct","choices":[{"index":0,"delta":{"content":"pub fn"},"finish_reason":null}]}

The final chunk includes complete token usage and nRouter spend headers:

x-nr-request-id: req_98a7bc12e4f0
x-nr-model: meta/llama-3.3-70b-instruct
x-nr-request-cost: $0.000144
x-nr-latency-ms: 412

SDK Integration Examples

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get("NROUTER_API_KEY"),
    base_url="https://api.nrouter.ai/v1",
)

response = client.chat.completions.create(
    model="meta/llama-3.3-70b-instruct",
    messages=[
        {"role": "system", "content": "You are an expert SQL engineer."},
        {"role": "user", "content": "Optimize a query joining 100M rows with partitioned dates."},
    ],
    temperature=0.3,
    stream=True,
)

for chunk in response:
    delta = chunk.choices[0].delta.content or ""
    print(delta, end="", flush=True)

Enterprise Guardrails & Caching with Llama

You can apply nRouter's request-path controls directly to Meta Llama requests using body options:

{
  "model": "meta/llama-3.3-70b-instruct",
  "messages": [{ "role": "user", "content": "Audit customer billing records..." }],
  "nrouter_guardrails": ["pii-strict", "prompt-injection-shield"],
  "nrouter_cache": true
}
  • Prompt Injection Defense: Scans inputs with sub-millisecond overhead. If an injection attempt is detected, nRouter blocks the request before it reaches the provider, guaranteeing $0 held and $0 spent.
  • PII Redaction: Automatically redacts credit cards, emails, and sensitive identifiers before tokenization.
  • Response Caching: Identical queries return from cache at cache-read rates with zero model generation delay.

Next Steps

Was this page helpful?