What Changed
Anthropic released Claude Opus 4.7, the latest version of their flagship model. While not a full architecture overhaul, the update brings significant improvements in three areas developers care about:
- Reasoning — 12% improvement on complex multi-step reasoning tasks (GPQA, ARC-Challenge)
- Tool use — Near-perfect function calling accuracy (97.3% on Berkeley Function Calling benchmark)
- Speed — 40% faster response times for non-reasoning queries, matching Sonnet speeds
Developer Experience Changes
For developers building with Claude, the practical improvements include:
- System prompts now support 32K tokens — up from 8K, enabling much more complex agent instructions
- JSON mode — Guaranteed valid JSON output without prompt engineering tricks
- Vision improvements — Better document understanding, especially for charts and diagrams
- Caching — Prompt caching now works with tool definitions, reducing costs by up to 90%
Benchmark Comparison
How Opus 4.7 stacks up against the competition:
- GPQA Diamond: Opus 4.7 78.3%, GPT-5.5 76.1%, Gemini Ultra 73.5%
- HumanEval: Opus 4.7 92.8%, GPT-5.5 94.1%, MiMo-V2.5-Pro 92.4%
- MATH-500: Opus 4.7 96.4%, GPT-5.5 97.2%, Qwen 3.6 Max 95.8%
- Function Calling: Opus 4.7 97.3%, GPT-5.5 95.8%, Gemini 96.1%
The function calling accuracy improvement is the most practically important change. At 97.3%, Claude Opus 4.7 is now the most reliable model for agentic workflows that depend on tool use. Combined with 32K system prompts, you can build complex agents with detailed instructions that actually execute correctly. This is the update that makes Claude the best choice for production agent systems.