Models / Qwen3 30b A3b Instruct 2507ID: dashscope/qwen3-30b-a3b-instruct-2507
by Qwen·via Alibaba US

Qwen3 30b A3b Instruct 2507

Qwen’s high-throughput, low-latency foundation model. Delivers high inference efficiency, rapid token generation, and cost-effective performance for production scale APIs and edge workloads.

Maker

Qwen

Served On

Alibaba US

Architecture

High-Throughput Compact Transformer

Latency Profile

Ultra-Fast (< 200ms TTFT)

Optimized For:High-volume batch processingLow-cost real-time streamingEdge deployment compatibilityContext: 128.0K

Industry Benchmarks & Evaluations

Standard Verified Metrics
MMLU (5-shot)81.2%

General knowledge and multi-task language understanding.

GSM8K (Grade School Math)89.6%

Multi-step arithmetic and mathematical reasoning.

HumanEval (Python Code)78.4%

Zero-shot functional code generation efficiency.

Inference Throughput> 140 tps

High-speed streaming token generation for real-time applications.

Needle In A Haystack (NIAH)99.4%

Context fidelity and accurate extraction across 128K tokens.

Pricing

Rates are read live

Pricing for qwen3-30b-a3b-instruct-2507 is fetched from the live catalogue on load, so it is never served from a cache that could outlive a repricing.

Specifications

Context window

128.0K

tokens
Max output

16K

tokens
Modality

chat

Context window, compared

qwen3-30b-a3b-instruct-2507qwen3-30b-a3b-instruct-2507 — 128K tokens128Kqwen-mt-turboqwen-mt-turbo — 128K tokens128Kqwen-plus-latestqwen-plus-latest — 128K tokens128Kqwen-plus-2025-01-25qwen-plus-2025-01-25 — 128K tokens128Kdeepseek-v4-pro-0813deepseek-v4-pro-0813 — 128K tokens128Kqwen3-vl-plus-2025-09-23qwen3-vl-plus-2025-09-23 — 128K tokens128Kclaude-sonnet-4-5-20250929claude-sonnet-4-5-20250929 — 1M tokens1M

Bars are scaled to 1M tokens. A model with no published context window reads Unknown, never zero.

Vision
Function calling
System prompt

Usage on nRouter

Usage is temporarily unavailable

Our aggregate read failed, so we cannot report traffic for this model right now.

nRouter SDK Integration

v2.2.1
curl https://api.nrouter.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ${NROUTER_API_KEY}" \
  -d '{
    "model": "dashscope/qwen3-30b-a3b-instruct-2507",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant powered by qwen3-30b-a3b-instruct-2507 via nRouter."
      },
      {
        "role": "user",
        "content": "Explain distributed consensus algorithms."
      }
    ]
  }'

Availability

Availability is temporarily unavailable

Our own telemetry read failed, so we cannot report uptime or latency for this model right now. This says nothing about the model itself — it is not an outage, and we would rather show you nothing than a figure we did not measure.

Interactive Playground

Send a real test request to qwen3-30b-a3b-instruct-2507 using your virtual key.

Test it

Prompt
cURL
curl https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"dashscope/qwen3-30b-a3b-instruct-2507","messages":[{"role":"user","content":"Hello!"}]}'

Managed Gateway Routing

Unified single-key access with automatic multi-region fallback, real-time credit enforcement, and sub-millisecond routing latency.

More from Qwen