Llama 4 Scout 17B (16E MoE)
Meta chat model served via Azure Foundry.
Meta
Azure Foundry
Pricing
Rates are read live
Pricing for Llama-4-Scout-17B-16E-Instruct is fetched from the live catalogue on load, so it is never served from a cache that could outlive a repricing.
Specifications
—
—
chat
Context window, compared
Bars are scaled to 1.05M tokens. A model with no published context window reads Unknown, never zero.
nRouter SDK Integration
curl https://api.nrouter.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${NROUTER_API_KEY}" \
-d '{
"model": "azure_ai/Llama-4-Scout-17B-16E-Instruct",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant powered by Llama-4-Scout-17B-16E-Instruct via nRouter."
},
{
"role": "user",
"content": "Explain distributed consensus algorithms."
}
]
}'Availability
Not enough data yet
We have not collected enough health probes for this model to publish an uptime or latency figure. A percentage from a handful of samples is fabricated precision, so none is shown until the sample count supports one.
Interactive Playground
Send a real test request to Llama-4-Scout-17B-16E-Instruct using your virtual key.
Test it
curl https://api.nrouter.ai/v1/chat/completions \
-H "Authorization: Bearer $NROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"azure_ai/Llama-4-Scout-17B-16E-Instruct","messages":[{"role":"user","content":"Hello!"}]}'Managed Gateway Routing
Unified single-key access with automatic multi-region fallback, real-time credit enforcement, and sub-millisecond routing latency.