DeepSeek

DeepSeek v4 Flash 0731

DeepSeekBalanced
ThinkingTool UseStructured Output

About this model

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Performance Tier

Balanced

DeepSeek v4 Flash 0731 is a balanced model from DeepSeek : strong performance at a reasonable price.

Strong cost-performance ratio. Reliable for most professional use cases without premium pricing.

Pricing

This model is included in Elosia plans
Eco

Minimal cost. Ideal for very high volume or simple tasks.

Typeper 1M tokens
Input (prompt)$0.080
Output (completion)$0.180
Cache read$0.016

Capabilities

Context Length1.0M
Max Output Tokens384K
TokenizerDeepSeek
Inputtext
Outputtext
Release DateJuly 31, 2026

Benchmarks

General Intelligence
MMLU
Not reported
GPQA Diamond
90.8%
Mathematics
MATH-500
Not reported
Programming
HumanEval
Not reported
Reasoning
Humanity's Last Exam
38.6%
Agentic
Terminal-Bench 2.1
82.7%

Where does DeepSeek v4 Flash 0731 stand?

Compare its performance index against every other model.

View the performance leaderboard

Recommended Use Cases

CodingAnalysisResearchMathematics

Strengths

  • Mixture of experts with 13B active parameters per token, drawn from 256 routed experts plus one shared, six engaged at a time
  • Agentic and coding execution evaluated across four separate suites, covering terminal workflows, repository patches, security tasks and tool use
  • Speculative decoding built into the checkpoint itself: the DSpark module ships with the weights rather than being bolted on at serving time
  • MIT open weights that are genuinely self-hostable, with official vLLM and SGLang recipes, tool calling with parallel_tool_calls, a 1M-token context and 384K of output

Limitations

  • Weak factual recall: on questions whose answer it does not know, it answers anyway far more often than it abstains. The strength is on agentic and code, not on knowledge density
  • Official figures are not reproducible, the DeepSeek Harness not being published, so the announced results cannot be replayed in house. Verbose reasoning as well at maximum effort, which weighs on latency and on the token budget

Frequently asked questions

Resources

This model may use your data for training

Similar Models