Kimi K2.6, an open-weights model from Chinese startup Moonshot AI, just achieved something remarkable: it beat Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro in a rigorous, real-time programming challenge. The results expose a critical gap in how we evaluate frontier models — and suggest that the era of Western labs having an unassailable capability lead is ending.
The Challenge
I run the AI Coding Contest, where I pit major language models against each other in real-time programming tasks with objective scoring. Day 12 was the Word Gem Puzzle — a sliding-tile letter puzzle that tests whether a model can write clean functional code, connect to a TCP server, and make strategic decisions under time pressure.
The board is a rectangular grid (10×10, 15×15, 20×20, 25×25, or 30×30) filled with letter tiles and one blank space. Bots can slide any adjacent tile into the blank and at any point claim valid English words formed in straight horizontal or vertical lines. Diagonals don't count. Backwards doesn't count.
The scoring rewards longer words and punishes short ones. Words under seven letters cost points: a five-letter word loses you one point, a three-letter word costs three. Seven letters or more score their length minus six, so an eight-letter word is worth two points. The same word can only be claimed once — if another bot gets there first, you get nothing.
Each pair of models played five rounds, one per grid size, with a ten-second wall-clock limit per round. The grids are seeded with real dictionary words in a crossword-style layout, then the remaining cells are filled with letters weighted by Scrabble tile frequencies, and finally the blank is scrambled, more aggressively on larger boards.
Results: Open-weights Model Dominates
The final leaderboard told a shocking story:
| Rank | Model | Match Points | Record |
|---|---|---|---|
| 1 | Kimi K2.6 | 22 | 7-1-0 |
| 2 | MiMo V2-Pro | 20 | 6-2-0 |
| 3 | ChatGPT GPT-5.5 | 16 | 5-1-2 |
| 4 | GLM 5.1 | 15 | 5-0-3 |
| 5 | Claude Opus 4.7 | 12 | 4-0-4 |
| 6 | Gemini Pro 3.1 | 9 | 3-0-5 |
| 7 | Grok Expert 4.2 | 9 | 3-0-5 |
| 8 | DeepSeek V4 | 3 | 1-0-7 |
| 9 | Muse Spark | 0 | 0-0-8 |
What I Saw
The move logs tell the story. Kimi won by sliding aggressively. Its approach was greedy: score each possible move by what new positive-value words it unlocks, execute the best one, repeat. When no move unlocked a positive word, it fell back to the first legal direction alphabetically.
This caused some inefficient edge-oscillation, a 2-cycle pattern where the bot bounced the blank back and forth without progress. On smaller grids where seed words were still largely intact, that hurt. On the 30×30 grids, where the scramble had broken up nearly everything and reconstruction was the only path to points, the sheer slide volume eventually paid off.
Kimi's cumulative score of 77 was the highest in the tournament.
MiMo's sliding code exists in the repo, but its "best value greater than zero" threshold never triggered, so in practice it never slid once. It went straight to scanning the initial grid for words of seven letters or more and blasted all its claims in a single TCP packet. Brittle strategy: entirely dependent on the scramble leaving intact seed words. On grids where words survived, MiMo cleaned up fast. On grids where they didn't, it scored nothing. Final tally: 43 cumulative points, second place.
Claude also didn't slide. The move logs show it holding up well on 25×25 boards where scramble density was still manageable, then falling apart on 30×30 where actual tile movement was needed. Not sliding is a real limitation in a puzzle built around sliding.
GPT-5.5 was more conservative, roughly 120 slides per round with a cap to avoid thrashing, and showed the strongest numbers on 15×15 and 30×30 grids. Grok never slid either, yet scored reasonably on the larger boards.
The Bigger Picture
This isn't a clean "China beats West" story — it's two specific models that won. But it does show open-weights models are closing the capability gap.
A year ago, the assumption was that the Western frontier labs had a capability lead open-weights couldn't close. Kimi K2.6 now scores 54 on the Artificial Analysis Intelligence Index. GPT-5.5 scores 60, Claude 57. That's not parity, but it's close, and it's coming from a model anyone can download.
When models within a few index points of the frontier are also freely available to run locally, that's a different competitive situation than the one that existed a year ago. This challenge is one data point in that shift. The gap is small enough now that it shows up in results like this one.
The Danger of Hallucination
Not all models played well. DeepSeek sent malformed data every round. Zero useful output. At least it didn't make things worse by playing.
Muse made things worse by playing. The scoring penalizes short words: three-letter words cost three points, four-letter words cost two, five-letter words cost one. The intent is to stop bots from carpet-bombing the board with "the", "and", and "it." Every serious competitor filtered their dictionary to words of seven letters or more.
Muse claimed everything. Every word it could find, regardless of length, fired off as a claim. On a 30×30 grid with hundreds of short valid words visible at any moment, Muse found them all and claimed every one.
Its cumulative score was −15,309. It lost all eight matches and won zero rounds. There is a version of Muse that simply connected to the server and did nothing, and that version would have scored zero, a 15,309-point improvement. The gap between Muse and eighth place was larger than the gap between eighth and first.
Editor's Thought
Editor's ThoughtThis isn't a clean "China beats West" story — it's two specific models that won. But it does show open-weights models are closing the capability gap. More importantly, it reveals that frontier models trained primarily for benchmarks may struggle with novel, rule-based games requiring active exploration. The fact that Kimi's greedy algorithm worked better than Claude's conservative approach suggests we need to test models on a wider variety of strategic tasks, not just coding and reasoning benchmarks.
Key Takeaways
What this means for developers: The era of Western frontier labs having an unassailable capability lead is ending. Open-weights models are now within a few points of frontier capabilities. When models you can download and run locally score within a few points of GPT-5.5, Claude Opus 4.7, and Gemini, that changes the competitive landscape dramatically.
What this means for research: We need to evaluate models on a wider variety of tasks — not just coding, reasoning, and knowledge benchmarks. The Word Gem Puzzle tests real-time decision-making, strategic exploration, and the ability to write clean functional code. Frontier models that excel at benchmarks may struggle on novel, rule-based games.
What this means for safety: Models that execute partial instructions without understanding penalties can fail catastrophically. Muse Spark claimed every word it could find, ignoring the scoring rules entirely. DeepSeek V4 sent malformed data. Both demonstrate that safety alignment isn't just about not causing harm — it's about understanding the rules and constraints of the task.