DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Performance Tier
Balanced
DeepSeek v4 Flash 0731 is a balanced model from DeepSeek : strong performance at a reasonable price.
Strong cost-performance ratio. Reliable for most professional use cases without premium pricing.
Pricing
This model is included in Elosia plans
Eco
Minimal cost. Ideal for very high volume or simple tasks.
Mixture of experts with 13B active parameters per token, drawn from 256 routed experts plus one shared, six engaged at a time
Agentic and coding execution evaluated across four separate suites, covering terminal workflows, repository patches, security tasks and tool use
Speculative decoding built into the checkpoint itself: the DSpark module ships with the weights rather than being bolted on at serving time
MIT open weights that are genuinely self-hostable, with official vLLM and SGLang recipes, tool calling with parallel_tool_calls, a 1M-token context and 384K of output
Limitations
Weak factual recall: on questions whose answer it does not know, it answers anyway far more often than it abstains. The strength is on agentic and code, not on knowledge density
Official figures are not reproducible, the DeepSeek Harness not being published, so the announced results cannot be replayed in house. Verbose reasoning as well at maximum effort, which weighs on latency and on the token budget