DeepSeek V4 Flash 0731 โ 90.8% GPQA, Agent Benchmarks Transformed
A re-post-trained revision of V4 Flash is now the primary backend behind DeepSeek-V4-Flash on OpenAdapter. Same model name, same quota cost โ but dramatically better at agent workflows.
What It Is
DeepSeek V4 Flash 0731 is a re-post-trained revision of V4 Flash. It is a sparse mixture-of-experts model: 284B total parameters, 13B active per token, tuned for coding, reasoning, and agent workflows. Only a fraction of the network fires per token, so it answers at small-model speed while reasoning like a much larger one.
| Attribute | Value |
|---|---|
| Architecture | Sparse MoE ยท 284B total / 13B active |
| Context window | 1M tokens upstream (we serve 256K) |
| Best at | Coding, tool use, long-context agents |
| Quota cost | Unchanged |
Independent Evaluation (Artificial Analysis)
Measured by Artificial Analysis, who run models themselves rather than republishing vendor claims:
| Benchmark | Score |
|---|---|
| Reasoning โ GPQA Diamond (graduate-level science) | 90.8% |
| Reasoning โ HLE (Humanity's Last Exam) | 36.8% |
| Reasoning โ AA-LCR (long-context reasoning) | 65.7% |
| Reasoning โ GDPval-AA (economically valuable tasks) | 52.9% |
| Reasoning โ CritPt (research-level physics) | 16.6% |
| Coding โ SciCode (scientific Python) | 49.9% |
| Knowledge โ AA-Omniscience (accuracy) | 37.2% |
| Knowledge โ AA-Omniscience (non-hallucination rate) | 15.6% |
The low CritPt and non-hallucination figures are equally worth knowing: it is strong at structured reasoning, weaker at research-frontier physics, and like most models it will still assert things confidently when it does not know. Do not use it as an oracle.
What the Retrain Changed
The weights are the same size. The post-training is not. DeepSeek's own agent-benchmark figures for this revision:
| Benchmark | Previous build | 0731 |
|---|---|---|
| DeepSWE | 7.3 | 54.4 |
| Terminal Bench 2.1 | 61.8 | 82.7 |
| Cybergym | 38.7 | 76.7 |
| SWE-bench Verified | 79.0 | 79.0 |
DeepSeek claims 0731 now beats V4 Pro Preview on every agent benchmark they published โ Terminal Bench 82.7 against Pro's 72.1.
How Fast It Is Here
Measured through our gateway, median of three runs each:
| Request size | Median response |
|---|---|
| Short prompt | 1.1s |
| ~6K tokens of context | 1.5s |
| ~34K tokens of context | 4.7s |
Thirty-four thousand tokens of code answered in under five seconds is what makes it usable as an agent backend rather than just a chat model.
Capability Check
We run every new model through the same suite before promoting it. This one passed all six:
| Test | Result |
|---|---|
| Instruction following | Pass |
| Arithmetic | Pass |
| Code generation | Pass |
| Logic (trick question) | Pass |
| Long-context retrieval | Pass |
| Tool / function calling | Pass |
How to Use It
Nothing to change. Point at DeepSeek-V4-Flash and you get it:
curl https://api.openadapter.in/v1/chat/completions \
-H "Authorization: Bearer $OPENADAPTER_API_KEY" \
-d '{"model":"DeepSeek-V4-Flash","messages":[{"role":"user","content":"Refactor this module"}]}'
The other backends stay in place as automatic failover, so if this one is ever busy your request still completes โ you just will not notice.
Referral code BDPBCR3R applied automatically