← Back to Blog

SWE-bench Is Saturated. Long-Horizon Agent Tasks Are the New Bench.

📅 2026-08-12 ⏱️ 9 min read 🏷️ Essays

📊 When 96% stops meaning anything

Claude Opus 5 is sitting at 96–97% on SWE-bench Verified. Fable 5 is at 95%. The numbers have nowhere left to go, and nobody I talk to believes them anymore. Here's what the field actually moved to in August 2026, and why the boring part — how you afford to run an agent for a thousand steps — quietly became the whole game. All numbers in this piece are pulled from live leaderboards, not press releases.

The benchmark ate itself

For two years SWE-bench Verified was the only number that mattered. Real GitHub issues, human-filtered to 500 instances, a standard mini-SWE-agent harness. Devin was the first to crack 50% on it back in 2024 and it felt like watching a moon landing. By August 2026 the leaderboard looks like this:

ModelSWE-bench VerifiedSource
Claude Opus 596–97%vals.ai, BenchLM aggregator
Claude Fable 595.0%MorphLLM Claude benchmarks
Claude Mythos 595.5%BenchLM aggregator
GPT-5.6 Sol96.2% (independent)MorphLLM July 2026
Claude Opus 4.8~88.6%MorphLLM (for reference)

Four models at or above 95%. And here's the thing nobody in the field will say out loud at a conference but everyone says in the Slack channels: when four frontier models all score within two points of each other at the very top of the scale, the benchmark has stopped discriminating. It's not measuring intelligence anymore. It's measuring how well you can overfit a 500-instance test set that has, by now, been trained on directly or indirectly by almost every major lab.

The contamination question got ugly fast. There's a whole subgenre of 2026 blog posts with names like "How We Broke Top AI Agent Benchmarks" documenting how the SWE-bench Verified subset leaked into training corpora, how some entries are easier than others in ways leaderboards don't disclose, how a 13B model can score 70% with the right scaffolding and zero actual capability underneath. The medium article "Beyond SWE-bench: How to Actually Evaluate AI Coding Agents in 2026" got passed around my network for a week straight because it said the quiet part loud: the universal benchmark became the universal gaming target.

So the field moved.

What replaced it: the long-horizon turn

Three benchmarks matter in August 2026, and none of them are SWE-bench Verified:

Notice what ties these together: they're all long-horizon. Many steps. Many tool calls. Many tokens. And that's where the problem stops being about model intelligence and starts being about something much more boring and much more important: how you afford to keep the agent running.

The shift in one line: SWE-bench asks "can the model solve this?" The new benchmarks ask "can the model solve this and keep enough context loaded to solve step 47 after steps 1 through 46?" The first is a model question. The second is an architecture question.

The real Terminal-Bench 2.1 leaderboard (live data)

I pulled this directly from the official tbench.ai leaderboard, not an aggregator. Top 10 as of August 2026, with confidence intervals:

#AgentModelScoreEffort
1Claude CodeFable 583.8% ± 1.2xhigh
2CodexGPT-5.583.1% ± 1.1xhigh
3Terminus 2Fable 580.4% ± 1.2high
4Cursor CLIGrok 4.579.3% ± 1.5high
5Claude CodeOpus 4.878.9% ± 1.3high
6CodexGPT-5.6 Terra78.4% ± 1.3max
7Terminus 2GPT-5.578.0% ± 1.2xhigh
8mini-SWE-agentMuse Spark 1.176.2% ± 1.2xhigh
9CodexGPT-5.6 Luna75.7% ± 1.3max
10Claude CodeSonnet 574.6% ± 1.6high

Three things jump out from the real data. First, the top score is 83.8%, not the 89%+ you'll see quoted in some aggregators — the official tbench.ai numbers run lower because the harness is stricter. Second, the top is not a single runaway model. Three different labs (Anthropic, OpenAI, xAI) and three different agent harnesses are within five points of each other. Third, notice the spread: positions 1 through 10 cover 9.2 points. That's a real discriminating benchmark — unlike SWE-bench Verified where the top four are within two points.

Notice also what's not on the official top-10. GLM-5.2 isn't there as of this snapshot, though Z.ai reports it at 81.0 on the same benchmark, which would put it at position 3 if confirmed. The open-weight frontier has closed fast, but the closed-weight frontier still has the official top positions locked.

The real bottleneck: context-window economics

Here's the math nobody likes. Run an agent for a thousand steps on a hard task. Each step appends a tool call, an observation, maybe a file read. By step 200 you're feeding 400K tokens back in. By step 500 you're at 1M. By step 800 you've blown past every frontier model's context window and you're either truncating (losing the beginning of the task) or paying quadratic attention cost on a stack that's growing linearly with every step.

For a year the answer was "bigger context windows, accept the cost." Google shipped 2M tokens. Others followed. And the cost was brutal — not just in dollars but in latency. Attention is quadratic in sequence length. At 1M tokens, the attention computation alone can dominate the entire forward pass. You don't notice on a 4K chat. You absolutely notice when your agent takes forty seconds per step.

Sparse attention was supposed to fix this. The idea is sound — instead of attending to every previous token, score them with a lightweight indexer and only attend to the top-k most relevant. DeepSeek's Sparse Attention (DSA) made this real. But sparse attention has its own dirty secret, and it's the one that's been quietly gating the whole agent-eval conversation.

The dirty secret of sparse attention

The indexer in a standard sparse-attention layer has to evaluate every previous position against the current query to decide which tokens to attend to. Do that in one layer and it's manageable. Do it in every layer of a 60-layer transformer and the indexer itself becomes the bottleneck — because the indexer also scales quadratically with context length. You "fixed" attention by adding a second thing that has the same cost disease.

Researchers had noticed for a while that adjacent transformer layers naturally select highly overlapping tokens. The indexer in layer 32 picks almost the same positions as the indexer in layer 33. Running both is mostly redundant work. But knowing that and exploiting it are different problems — the second one requires you to retrain the model around the assumption, not just bolt on an inference-time trick.

This is the point in the story where a specific architecture decision walks in, and it's worth dwelling on because it's one of the cleaner pieces of mechanism design I've seen shipped this year.

IndexShare, or: how to delete 75% of your indexer compute

Zhipu's GLM-5.2 (the confirmed current flagship, released June 13, 2026) introduced a mechanism called IndexShare, and the cleanest technical writeup I've found is Sebastian Raschka's blog. The mechanism is brutally simple once you see it:

Instead of every sparse-attention layer running its own indexer, group the layers in cycles of four. The first layer computes its indexer normally — a "full" layer. The next three layers — "shared" layers — inherit those exact token positions and skip their own indexer entirely.

layer N     → full   (computes its own sparse-attention indices)
layer N+1   → shared (reuses layer N's selected positions)
layer N+2   → shared (reuses layer N's selected positions)
layer N+3   → shared (reuses layer N's selected positions)
layer N+4   → full   (computes its own, restarts the cycle)

The shared layers don't skip attention — they still compute their own queries, attention weights, value combinations, output projections, and feed-forward updates. They only skip the indexer. They reuse the mapping of which positions to attend to, not the attention computation itself.

The reason this works is the observation from the last section: adjacent layers select overlapping tokens anyway. IndexShare just makes that explicit and stops recomputing it. Result: 75% fewer indexer computations across the model. Z.ai reports a 2.9-fold reduction in per-token FLOPs at 1M-token context.

Important caveat, because this is where the marketing usually starts lying: 2.9× is an architectural compute estimate on the indexer. It is not an end-to-end wall-clock speedup. It does not proportionally reduce KV-cache memory, because the cache still has to be stored and transferred. Kernel scheduling and cache transfers still incur real serving cost. Treat "2.9× faster" claims with suspicion; treat "2.9× fewer indexer FLOPs" as accurate.

The mechanism is baked in during mid-training, not applied at inference. That matters — the active indexer layers learn to select positions that work for their three dependents, rather than serving as a static cache. It's an architecture choice, not a serving hack.

Does it actually translate to agent benchmarks?

This is the question that matters, and the data is more interesting than the marketing. Here are GLM-5.2's numbers versus its predecessor GLM-5.1 on the long-horizon agent benchmarks, from Z.ai's official release and corroborated by apidog and emergent.sh:

BenchmarkGLM-5.1GLM-5.2Jump
Terminal-Bench 2.163.581.0+17.5
FrontierSWE30.574.4+43.9
SWE-bench Pro58.462.1+3.7
SWE-Marathon1.013.0+12.0

Look at the pattern. SWE-bench Pro — the closest cousin to the saturated SWE-bench Verified — barely moves. +3.7 points. That's the "everyone's at the ceiling" benchmark, and everyone's still at the ceiling.

But FrontierSWE more than doubled. SWE-Marathon went from 1.0 to 13.0 — a 13× improvement on the benchmark explicitly designed to take agents hours of multi-step work. Terminal-Bench jumped 17.5 points, and 81.0 would slot GLM-5.2 into the top 3 on the official tbench.ai leaderboard if it lands there independently. The architecture change shows up exactly where you'd predict: on the long-horizon tasks where the indexer was the actual bottleneck.

Is GLM-5.2 the absolute frontier? No. The closed-weight leaders still edge it on raw Terminal-Bench 2.1 (Fable 5 + Claude Code at 83.8% is nearly 3 points ahead). But among open-weight models, GLM-5.2 currently tops Terminal-Bench 2.1 and SWE-bench Pro. And if your question is "which model can run the longest before context cost kills the task," the IndexShare architecture is the only published mechanism that directly attacks the indexer cost disease.

Important honesty: the GLM-5.2 numbers above are vendor-reported (Zhipu's release), corroborated by third-party writeups but not yet on the official tbench.ai leaderboard snapshot I pulled. Treat them as accurate-but-not-yet-independently-audited, which is the right epistemic stance for any 2026 model benchmark. The SWE-bench Verified numbers at the top of this article are independently aggregator-confirmed (vals.ai, BenchLM, MorphLLM).

What this means if you're buying or building

Three practical takeaways for the rest of 2026:

  1. Stop using SWE-bench Verified as your decision criterion. When four models are within two points at 95%+, you're picking based on noise. Look at Terminal-Bench 2.1 (the official tbench.ai numbers, not aggregator inflation), FrontierSWE, and SWE-Marathon. The discriminating signal moved.
  2. Price your context economics explicitly. If you're running agents in production, your cost per task is dominated by how many tokens you feed back in per step, and your latency is dominated by how cheaply the model can attend over that growing context. An architecture like IndexShare is a direct hit on that cost line. A 96% SWE-bench Verified number tells you nothing about it.
  3. The open-weight gap on agent tasks closed faster than the open-weight gap on chat tasks. A year ago the closed-frontier-vs-open-frontier gap on long-horizon agent work was enormous. It still exists — GLM-5.2's 81.0 is still behind Fable 5's 83.8 — but the rate of closure on Terminal-Bench and FrontierSWE is faster than on MMLU or HumanEval. If you care about agent work specifically, the open models are now legitimately competitive, not as a values statement, as a benchmark statement.

The unsentimental version

The August 2026 story isn't "models got smarter." The August 2026 story is "we ran out of useful short-task benchmarks, started measuring the tasks we actually wanted agents to do, and discovered the bottleneck wasn't intelligence — it was the cost of keeping an agent's context alive across hundreds of steps."

Whoever makes that cost tractable wins the next eighteen months. IndexShare is one answer. There will be others. The common shape: stop attending to everything, stop indexing in every layer, accept that adjacent layers want the same tokens anyway. The models that ship this will pull ahead on the long-horizon benchmarks. The models that don't will keep winning SWE-bench Verified by a point and a half, and it will keep not mattering.

🤖 Want to run this kind of long-horizon work yourself?

GLM-5.2 is the model with the published IndexShare architecture, the open MIT weights, and the 1M-token context. If you want to actually test the long-context thesis on your own multi-step agent workload, the coding plan is the cheapest way to do it without standing up the weights yourself — readers from here get 10% off. Worth it specifically if your bottleneck is steps-per-task and not single-shot quality.


Sources: SWE-bench Verified numbers from vals.ai, BenchLM, and MorphLLM aggregators (August 2026). Terminal-Bench 2.1 official leaderboard fetched directly from tbench.ai/leaderboard/terminal-bench/2.1. IndexShare mechanism from Sebastian Raschka's technical writeup. GLM-5.2 benchmark numbers are vendor-reported (Z.ai official release, June 2026) corroborated by apidog.com and emergent.sh, pending independent leaderboard confirmation. The "2.9× FLOP reduction" is an architectural compute estimate on the indexer, not a wall-clock claim. GLM-5.2 confirmed as current Z.ai flagship (released June 13, 2026); GLM-5.3, GLM-5.5, and GLM-5.6 are rumored or unreleased as of this date.