Kwaipilot released KAT-Coder-V2.5-Dev on July 24, 2026 under an Apache 2.0 open-weight license — a Mixture-of-Experts coding model built on the Qwen3.6-35B-A3B foundation with 35B total parameters but only 3B active per inference pass. Post-trained with the KAT-V2.5 recipe (data construction, SFT, then reinforcement learning), it scores 69.40 on SWE-bench Verified, 63.00 on SWE-bench Multilingual, and 93.43 on PinchBench — numbers that let local, private coding agents rival proprietary closed-source giants.
KAT-Coder-V2.5-Dev is engineered to move past code completion into autonomous software engineering: reasoning through entire repositories, resolving real GitHub issues, and executing terminal workflows. By open-sourcing the weights under Apache 2.0, Kwaipilot is explicitly targeting developers who need local, private, self-hosted coding agents — the segment for whom sending proprietary code to a third-party API is a non-starter.
The MoE design is the whole trick. With 35B total parameters and only 3B active during any single inference pass, KAT-Coder-V2.5-Dev aims to deliver the intelligence of a much larger model with the latency and computational cost of a much smaller one — the trade that makes on-prem agentic coding economically viable. The KAT-V2.5 training recipe behind it is a full pipeline: advanced data construction, a specialized Supervised Fine-Tuning stage, and a reinforcement learning optimization strategy, so the model learns the intent behind software architecture rather than surface syntax.
⚡ Key Detail: because only 3B parameters fire per token, the model runs efficiently on standard MoE-capable inference engines — vLLM or llama.cpp — making a self-hosted autonomous coding agent a realistic single-server deployment.
Standard completion benchmarks undersell agentic coding, so the scoreboard that matters here is the SWE-bench family plus terminal and reasoning suites. KAT-Coder-V2.5-Dev posts unprecedented strength in resolving real GitHub issues autonomously, and navigates multilingual repositories and terminal command execution with precision previously associated with much larger dense models.
| Benchmark | Score | What It Tests |
|---|---|---|
| SWE-bench Verified | 69.40 | Real GitHub issue resolution |
| SWE-bench Multilingual | 63.00 | Multilingual repositories |
| SWE-bench Pro | 45.96 | Harder agentic coding |
| Terminal-Bench 2.1 | 41.02 | Terminal command execution |
| PinchBench | 93.43 | Practical coding utility |
| Scicode | 44.20 | Scientific reasoning |
| KAT-Code-Bench | 46.21 | Kwaipilot’s in-house suite |
The high reasoning ceiling makes the model a fit for agents that plan and execute over multiple steps: bug fixes on legacy code, repository-wide refactors, comprehensive unit test generation, and terminal-based agentic workflows. It is also well-suited to repository-wide RAG — because it understands structural dependencies between modules, it can use retrieved context to answer how different parts of a massive codebase interact, not just locate text.
This technical summary was compiled exclusively using verified public data points from the following release documentation: