Models / DeepSeek V4 Flash 0731ID: dashscope/deepseek-v4-flash-0731
by DeepSeek·via Alibaba US

DeepSeek V4 Flash 0731

DeepSeek’s high-throughput, low-latency foundation model. Delivers high inference efficiency, rapid token generation, and cost-effective performance for production scale APIs and edge workloads.

Maker

DeepSeek

Served On

Alibaba US

Architecture

High-Throughput Compact Transformer

Latency Profile

Ultra-Fast (< 200ms TTFT)

Optimized For:High-volume batch processingLow-cost real-time streamingEdge deployment compatibilityContext: 128.0K

Industry Benchmarks & Evaluations

Standard Verified Metrics
SWE-bench Verified65.2%

Autonomous software engineering & real-world GitHub bug resolution.

MMLU-Pro (Reasoning)83.1%

Multi-disciplinary advanced reasoning across STEM and humanities.

GPQA Diamond67.5%

PhD-level biology, physics, chemistry, and domain-expert logic.

AIME 2024 (Math)76.4%

American Invitational Mathematics Examination competition logic.

Needle In A Haystack (NIAH)99.9%

Long-context document retrieval accuracy across 128K tokens.

Pricing

Rates are read live

Pricing for deepseek-v4-flash-0731 is fetched from the live catalogue on load, so it is never served from a cache that could outlive a repricing.

Specifications

Context window

128.0K

tokens
Max output

16K

tokens
Modality

chat

Context window, compared

deepseek-v4-flash-0731deepseek-v4-flash-0731 — 128K tokens128Kqwen-mt-turboqwen-mt-turbo — 128K tokens128Kqwen-plus-latestqwen-plus-latest — 128K tokens128Kqwen-plus-2025-01-25qwen-plus-2025-01-25 — 128K tokens128Kdeepseek-v4-pro-0813deepseek-v4-pro-0813 — 128K tokens128Kqwen3-vl-plus-2025-09-23qwen3-vl-plus-2025-09-23 — 128K tokens128Kclaude-sonnet-4-5-20250929claude-sonnet-4-5-20250929 — 1M tokens1M

Bars are scaled to 1M tokens. A model with no published context window reads Unknown, never zero.

Vision
Function calling
System prompt

Usage on nRouter

Usage is temporarily unavailable

Our aggregate read failed, so we cannot report traffic for this model right now.

nRouter SDK Integration

v2.2.1
curl https://api.nrouter.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ${NROUTER_API_KEY}" \
  -d '{
    "model": "dashscope/deepseek-v4-flash-0731",
    "messages": [
      {
        "role": "system",
        "content": "You are a helpful assistant powered by deepseek-v4-flash-0731 via nRouter."
      },
      {
        "role": "user",
        "content": "Explain distributed consensus algorithms."
      }
    ]
  }'

Availability

Availability is temporarily unavailable

Our own telemetry read failed, so we cannot report uptime or latency for this model right now. This says nothing about the model itself — it is not an outage, and we would rather show you nothing than a figure we did not measure.

Interactive Playground

Send a real test request to deepseek-v4-flash-0731 using your virtual key.

Test it

Prompt
cURL
curl https://api.nrouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $NROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"dashscope/deepseek-v4-flash-0731","messages":[{"role":"user","content":"Hello!"}]}'

Managed Gateway Routing

Unified single-key access with automatic multi-region fallback, real-time credit enforcement, and sub-millisecond routing latency.

More from DeepSeek