← Back to Blog

KAT-Coder-V2.5-Dev: Kwaipilot’s 35B-A3B Open-Weights Coding Agent Hits 69.40 on SWE-bench Verified

Category: Model Releases · 2026-08-18 · ~4 min read · Verified metrics & benchmark data

The Verdict First

Kwaipilot released KAT-Coder-V2.5-Dev on July 24, 2026 under an Apache 2.0 open-weight license — a Mixture-of-Experts coding model built on the Qwen3.6-35B-A3B foundation with 35B total parameters but only 3B active per inference pass. Post-trained with the KAT-V2.5 recipe (data construction, SFT, then reinforcement learning), it scores 69.40 on SWE-bench Verified, 63.00 on SWE-bench Multilingual, and 93.43 on PinchBench — numbers that let local, private coding agents rival proprietary closed-source giants.

35B/3B
Total / Active Parameters (MoE)
69.40
SWE-bench Verified
93.43
PinchBench
A-2.0
Apache 2.0 Open Weights

The Release in One Paragraph

KAT-Coder-V2.5-Dev is engineered to move past code completion into autonomous software engineering: reasoning through entire repositories, resolving real GitHub issues, and executing terminal workflows. By open-sourcing the weights under Apache 2.0, Kwaipilot is explicitly targeting developers who need local, private, self-hosted coding agents — the segment for whom sending proprietary code to a third-party API is a non-starter.

Architecture: Big-Model Intelligence, Small-Model Bills

The MoE design is the whole trick. With 35B total parameters and only 3B active during any single inference pass, KAT-Coder-V2.5-Dev aims to deliver the intelligence of a much larger model with the latency and computational cost of a much smaller one — the trade that makes on-prem agentic coding economically viable. The KAT-V2.5 training recipe behind it is a full pipeline: advanced data construction, a specialized Supervised Fine-Tuning stage, and a reinforcement learning optimization strategy, so the model learns the intent behind software architecture rather than surface syntax.

⚡ Key Detail: because only 3B parameters fire per token, the model runs efficiently on standard MoE-capable inference engines — vLLM or llama.cpp — making a self-hosted autonomous coding agent a realistic single-server deployment.

Benchmarks: The Full Scoreboard

Standard completion benchmarks undersell agentic coding, so the scoreboard that matters here is the SWE-bench family plus terminal and reasoning suites. KAT-Coder-V2.5-Dev posts unprecedented strength in resolving real GitHub issues autonomously, and navigates multilingual repositories and terminal command execution with precision previously associated with much larger dense models.

BenchmarkScoreWhat It Tests
SWE-bench Verified69.40Real GitHub issue resolution
SWE-bench Multilingual63.00Multilingual repositories
SWE-bench Pro45.96Harder agentic coding
Terminal-Bench 2.141.02Terminal command execution
PinchBench93.43Practical coding utility
Scicode44.20Scientific reasoning
KAT-Code-Bench46.21Kwaipilot’s in-house suite

Beyond Code Generation

The high reasoning ceiling makes the model a fit for agents that plan and execute over multiple steps: bug fixes on legacy code, repository-wide refactors, comprehensive unit test generation, and terminal-based agentic workflows. It is also well-suited to repository-wide RAG — because it understands structural dependencies between modules, it can use retrieved context to answer how different parts of a massive codebase interact, not just locate text.

Deployment: Two Paths

Fact Verification & Sources

This technical summary was compiled exclusively using verified public data points from the following release documentation:

⚡ OpenAdapter Readers get 20% off — invite code BDPBCR3R ◉ Z.ai Coding Plan Readers get 10% off — invite code R0K78RJKNW
R
Analyzed for CLAW

Live analysis published on claw.rommark.dev on Aug 18, 2026. Data grounded exclusively in official release benchmarks and verified technical specifications.