← Back to Blog

Ornith-1.5-35B-A3B: The Self-Improving Open-Weights MoE That Plays in the Opus League

Category: Model Releases · 2026-08-20 · ~4 min read · Verified metrics & benchmark data

The Verdict First

Ornith released the Ornith-1.5 family on August 20, 2026 under an MIT license, headlined by a 35B-total / 3B-active MoE that is claimed to perform on par with Claude Opus 4.8 across reasoning, agentic workflows and complex instruction following. The hook isn't just the scores — 86 on SWE-Bench Verified and 86.1 on Terminal-Bench 2.1 — it's how the model was trained: an end-to-end self-improvement loop in which the model proposes its own tasks, generates its own scaffolds, and rolls out solutions for reinforcement learning. Open weights that teach themselves.

35B/3B
Total / Active Parameters (MoE)
86
SWE-Bench Verified
86.1
Terminal-Bench 2.1
MIT
Open-Weights License

The Release in One Paragraph

The Ornith-1.5 family is a three-tier lineup: a 9B Dense model for edge applications, the flagship 35B MoE (A3B) for high-performance reasoning, and a 397B MoE for ultra-complex enterprise workloads. The release positions itself as the moment open-source models stop chasing proprietary giants and start setting benchmarks for autonomous intelligence themselves. Every tier ships under the same permissive license, which allows unrestricted commercial and research use.

Architecture: Self-Scaffolding MoE

The flagship uses an "Active 3B" architecture: 35 billion total parameters, but only 3 billion activated per token — the sparse-activation trick that keeps inference fast while preserving the knowledge capacity of a much larger model. What distinguishes the 1.5 series is the training methodology on top of it. Building on the self-scaffolding strategies introduced in Ornith-1.0, the model doesn't merely consume static datasets: during training it proposes its own new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning — a continuous cycle in which the model effectively becomes its own teacher.

⚡ Key Detail: the self-improvement loop is the release's core claim to fame. Instead of humans curating every training signal, the model generates the tasks and scaffolds it learns from — which is how a 35B/3B MoE is claimed to bridge the gap to frontier-scale systems without frontier-scale training data pipelines.

Benchmarks: The Full Scoreboard

The headline claim is parity with Claude Opus 4.8 across reasoning, agentic workflows, and complex instruction following. Here is the complete verified scoreboard for the 35B-A3B variant, exactly as published in the release:

BenchmarkScoreWhat It Tests
SWE-Bench Verified86Real-world repository-level engineering
SWE-Bench Pro65.1Harder agentic coding
SWE-Bench Multilingual79.6Multilingual repositories
Terminal-Bench 2.186.1Terminal & tool usage
DeepSWE56Deep software engineering
HLE (Humanity’s Last Exam)44.6Frontier reasoning
ClawEval81.4Agentic evaluation suite
Tool Decathlon71.2OS, API & environment interaction

What the Numbers Mean in Practice

Two rows stand out. An 86 verified on SWE-Bench means the model resolves real repository-level engineering tasks — the kind that previously demanded a much larger closed model — with high precision. And the 86.1 on Terminal-Bench 2.1, described in the release as a commanding lead, points at agents that live in the shell: administering systems, chaining tools, and recovering from errors without hand-holding. The Tool Decathlon score of 71.2 reinforces that profile — the release specifically flags it as the reason this model suits autonomous agents that interact with operating systems, APIs, and complex software environments.

Availability: Every Deployment Path at Once

Because the weights are MIT-licensed, Ornith sidesteps the usual open-weights caveat — there is no usage clause to lawyer over. The release lists quantized builds in FP8, GGUF, MLX, and NVFP4 formats, with framework support for vLLM, llama.cpp, and MLX. The MLX builds in particular mean Apple Silicon users get high-performance local inference out of the box. Weights are available via Hugging Face and the official Ornith GitHub repository, and Ollama already lists the model in its library.

Who Should Care

If you build autonomous coding agents, the SWE-Bench profile makes this an immediate evaluation candidate. If you run terminal-heavy automation or system administration agents, Terminal-Bench 2.1 at 86.1 is the number to benchmark against your current stack. And if you operate enterprise RAG in multilingual environments, the 79.6 Multilingual score plus the model's stated strength in retrieval-augmented generation and long-context reasoning keep it on the shortlist — all without a single API bill, if you'd rather host it yourself.

Fact Verification & Sources

This technical summary was compiled exclusively using verified public data points from the following release documentation:

⚡ OpenAdapter Readers get 20% off — invite code BDPBCR3R ◉ Z.ai Coding Plan Readers get 10% off — invite code R0K78RJKNW
R
Analyzed for CLAW

Live analysis published on claw.rommark.dev on Aug 20, 2026. Data grounded exclusively in official release benchmarks and verified technical specifications.