๐Ÿ—“๏ธ 2026-08-01 AI News โฑ 5 min read

DeepSeek V4 Flash 0731 โ€” 90.8% GPQA, Agent Benchmarks Transformed

A re-post-trained revision of V4 Flash is now the primary backend behind DeepSeek-V4-Flash on OpenAdapter. Same model name, same quota cost โ€” but dramatically better at agent workflows.

DeepSeek V4 Flash Agent Benchmark OpenAdapter #AICoding

What It Is

DeepSeek V4 Flash 0731 is a re-post-trained revision of V4 Flash. It is a sparse mixture-of-experts model: 284B total parameters, 13B active per token, tuned for coding, reasoning, and agent workflows. Only a fraction of the network fires per token, so it answers at small-model speed while reasoning like a much larger one.

AttributeValue
ArchitectureSparse MoE ยท 284B total / 13B active
Context window1M tokens upstream (we serve 256K)
Best atCoding, tool use, long-context agents
Quota costUnchanged
90.8% on GPQA Diamond โ€” graduate-level science reasoning from a model with only 13B active parameters. That is the headline number.

Independent Evaluation (Artificial Analysis)

Measured by Artificial Analysis, who run models themselves rather than republishing vendor claims:

BenchmarkScore
Reasoning โ€” GPQA Diamond (graduate-level science)90.8%
Reasoning โ€” HLE (Humanity's Last Exam)36.8%
Reasoning โ€” AA-LCR (long-context reasoning)65.7%
Reasoning โ€” GDPval-AA (economically valuable tasks)52.9%
Reasoning โ€” CritPt (research-level physics)16.6%
Coding โ€” SciCode (scientific Python)49.9%
Knowledge โ€” AA-Omniscience (accuracy)37.2%
Knowledge โ€” AA-Omniscience (non-hallucination rate)15.6%

The low CritPt and non-hallucination figures are equally worth knowing: it is strong at structured reasoning, weaker at research-frontier physics, and like most models it will still assert things confidently when it does not know. Do not use it as an oracle.

What the Retrain Changed

The weights are the same size. The post-training is not. DeepSeek's own agent-benchmark figures for this revision:

BenchmarkPrevious build0731
DeepSWE7.354.4
Terminal Bench 2.161.882.7
Cybergym38.776.7
SWE-bench Verified79.079.0
Note the shape: single-shot coding barely moved (SWE-bench flat at 79.0) while multi-step agent ability transformed. This model got better at operating, not at writing a function. Driven in a loop with tools you will feel it; on one-off questions you may not.

DeepSeek claims 0731 now beats V4 Pro Preview on every agent benchmark they published โ€” Terminal Bench 82.7 against Pro's 72.1.

Treat that second table with care: those are DeepSeek testing DeepSeek, on a harness that is not public yet, and two of the nine are internal sets the company built. The Artificial Analysis figures above and our own measurements below are the parts with independent backing.

How Fast It Is Here

Measured through our gateway, median of three runs each:

Request sizeMedian response
Short prompt1.1s
~6K tokens of context1.5s
~34K tokens of context4.7s

Thirty-four thousand tokens of code answered in under five seconds is what makes it usable as an agent backend rather than just a chat model.

Capability Check

We run every new model through the same suite before promoting it. This one passed all six:

TestResult
Instruction followingPass
ArithmeticPass
Code generationPass
Logic (trick question)Pass
Long-context retrievalPass
Tool / function callingPass

How to Use It

Nothing to change. Point at DeepSeek-V4-Flash and you get it:

curl https://api.openadapter.in/v1/chat/completions \
  -H "Authorization: Bearer $OPENADAPTER_API_KEY" \
  -d '{"model":"DeepSeek-V4-Flash","messages":[{"role":"user","content":"Refactor this module"}]}'

The other backends stay in place as automatic failover, so if this one is ever busy your request still completes โ€” you just will not notice.

Try OpenAdapter โ€” 20% OFF with my link

Referral code BDPBCR3R applied automatically


#DeepSeek #V4Flash #AICoding #AgentBenchmark #OpenAdapter