Meta Llama Models
Access Meta Llama 3.3 70B, 3.1 405B, and Llama 4 Scout via nRouter's unified gateway with real-time streaming, guardrails, and zero-markup list pricing.
Last updated
Meta Llama models provide industry-standard open weights performance across reasoning, coding, multilingual tasks, and multimodal vision. Through nRouter (https://api.nrouter.ai/v1), your applications can call Meta's flagship instruction-tuned models with one virtual API key, enterprise budget controls, pre-flight safety guardrails, and flat list pricing with zero per-token markup.
Available Meta Models
nRouter serves three primary Meta Llama models, each tailored for specific architectural requirements:
| Model ID | Architecture & Parameters | Modality | Context Window | Max Output | Flat List Price (Input / Output per 1M) | Primary Workload |
|---|---|---|---|---|---|---|
meta/llama-3.3-70b-instruct | Dense 70B parameters | Text in / Text out | 128k tokens | 4,096 tokens | $0.72 / $0.72 | High-throughput reasoning, code generation, tool use |
meta/llama-3.1-405b-instruct | Dense 405B parameters | Text in / Text out | 128k tokens | 4,096 tokens | $2.00 / $2.00 | Frontier-class logic, synthetic data generation, model distillation |
meta/llama-4-scout-17b-16e-instruct | Mixture of Experts (17B active / 16 experts) | Text + Vision in / Text out | 128k tokens | 8,192 tokens | $0.30 / $0.60 | Multimodal document analysis, low-latency agentic loops |
Model Highlights
meta/llama-3.3-70b-instruct
The default workhorse for enterprise applications. It rivals previous-generation 405B models on benchmark evaluation suites while operating at a fraction of the compute cost and latency. Supports native function calling, structured JSON output, and complex multi-turn dialogs.
meta/llama-3.1-405b-instruct
Meta's flagship open-weights model. Delivering frontier-class performance comparable to top proprietary models, Llama 3.1 405B is optimized for difficult math and algorithmic reasoning, cross-domain knowledge synthesis, and generating high-quality training and evaluation datasets.
meta/llama-4-scout-17b-16e-instruct
Built on Meta's next-generation Llama 4 Mixture-of-Experts (MoE) architecture. With 17B active parameters out of 16 experts, it delivers ultra-fast time-to-first-token (TTFT) and supports native multimodal image inputs alongside text prompts.
Flat List Price Guarantee
Under nRouter's pricing policy (Rule #28), all models are served at flat provider list rates with zero per-token markup.
- Exact List Settlement: Your organization pays exactly the listed token price ($0.72/$0.72 for 70B, $2.00/$2.00 for 405B, $0.30/$0.60 for Llama 4 Scout). nRouter makes zero profit on tokens.
- Phase 4 Credit Reservation: When a request is dispatched, nRouter reserves an estimated credit envelope against your organization balance. Upon response completion, the exact token count is calculated and the remaining hold is immediately returned to your balance.
- Cost Header Audit: Every response includes the exact billed amount in the
x-nr-request-costheader (e.g.x-nr-request-cost: $0.000412).
Endpoint Format & Request Structure
Meta Llama models are served via the standard OpenAI-compatible /v1/chat/completions endpoint:
POST https://api.nrouter.ai/v1/chat/completionsStandard Text Completion
curl https://api.nrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $NROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "meta/llama-3.3-70b-instruct",
"messages": [
{
"role": "system",
"content": "You are a concise engineering assistant."
},
{
"role": "user",
"content": "Explain the difference between dense and sparse MoE architectures."
}
],
"temperature": 0.2,
"max_tokens": 1024
}'Multimodal Vision (meta/llama-4-scout-17b-16e-instruct)
Llama 4 Scout accepts image inputs via standard URL or Base64 data encodings:
curl https://api.nrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $NROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "meta/llama-4-scout-17b-16e-instruct",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "Analyze the architecture diagram in this image:" },
{
"type": "image_url",
"image_url": {
"url": "https://nrouter.ai/diagrams/gateway-flow.png"
}
}
]
}
],
"max_tokens": 1000
}'Real-Time Streaming
Set "stream": true to receive real-time Server-Sent Events (SSE). nRouter forwards tokens with sub-millisecond gateway overhead while maintaining observability and credit tracking.
Streaming with cURL
curl -N https://api.nrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $NROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "meta/llama-3.3-70b-instruct",
"messages": [
{ "role": "user", "content": "Write a fast Rust binary search implementation." }
],
"stream": true
}'The gateway streams chunks formatted as data: {...}:
data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1709251200,"model":"meta/llama-3.3-70b-instruct","choices":[{"index":0,"delta":{"content":"pub fn"},"finish_reason":null}]}The final chunk includes complete token usage and nRouter spend headers:
x-nr-request-id: req_98a7bc12e4f0
x-nr-model: meta/llama-3.3-70b-instruct
x-nr-request-cost: $0.000144
x-nr-latency-ms: 412SDK Integration Examples
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get("NROUTER_API_KEY"),
base_url="https://api.nrouter.ai/v1",
)
response = client.chat.completions.create(
model="meta/llama-3.3-70b-instruct",
messages=[
{"role": "system", "content": "You are an expert SQL engineer."},
{"role": "user", "content": "Optimize a query joining 100M rows with partitioned dates."},
],
temperature=0.3,
stream=True,
)
for chunk in response:
delta = chunk.choices[0].delta.content or ""
print(delta, end="", flush=True)Enterprise Guardrails & Caching with Llama
You can apply nRouter's request-path controls directly to Meta Llama requests using body options:
{
"model": "meta/llama-3.3-70b-instruct",
"messages": [{ "role": "user", "content": "Audit customer billing records..." }],
"nrouter_guardrails": ["pii-strict", "prompt-injection-shield"],
"nrouter_cache": true
}- Prompt Injection Defense: Scans inputs with sub-millisecond overhead. If an injection attempt is detected, nRouter blocks the request before it reaches the provider, guaranteeing $0 held and $0 spent.
- PII Redaction: Automatically redacts credit cards, emails, and sensitive identifiers before tokenization.
- Response Caching: Identical queries return from cache at cache-read rates with zero model generation delay.
Next Steps
- Model Catalog — Browse all supported models across providers
- TypeSafe Jev Guide — Pair Llama with fast System One decision routing
- Guardrails Guide — Enforce zero-trust safety on incoming prompts
- API Reference — Complete schema documentation
Model Catalog
Browse, filter, and compare all LLM models and multimodal providers available to your organization, with real-time pricing, capabilities, and live latency.
TypeSafe Jev (System One Decision Model)
TypeSafe Jev is a System One decision model delivering sub-100ms latency, zero-token generation overhead, typed scoring, and guaranteed $0 output token billing.