← Back to Blog

Ornith-1.5-35B-A3B: The Self-Improving Open-Source MoE That Hits 86 on SWE-Bench Verified

Category: Model Releases · 2026-08-20 · ~5 min read · Verified metrics & benchmark data

The Verdict First

Ornith released the Ornith-1.5 family on August 20, 2026, and the flagship 35B-A3B is the story: a Mixture-of-Experts model with 35B total parameters but only 3B active per token, trained through an end-to-end self-improvement loop, shipping under an MIT license — and posting 86 on SWE-Bench Verified and 86.1 on Terminal-Bench 2.1. According to the release, it performs on par with Claude Opus 4.8 across reasoning, agentic workflows, and complex instruction following. An open-weights model that proposes its own training tasks is a genuine inflection point.

35B/3B
Total / Active Params (MoE)
86
SWE-Bench Verified
86.1
Terminal-Bench 2.1
MIT
License (Commercial OK)

What Shipped on August 20

The Ornith-1.5 family is a three-tier lineup built for scalability: a 9B Dense model aimed at edge applications, the flagship 35B MoE (A3B) for high-performance reasoning, and a 397B MoE variant for ultra-complex enterprise workloads. The headline designation — "A3B" — means that despite 35 billion total parameters, only 3 billion are activated per token, which is what keeps inference fast on accessible hardware.

The Core Innovation: A Model That Trains Itself

What separates Ornith-1.5 from the weekly parade of fine-tunes is the training methodology. Building on the self-scaffolding strategies introduced in Ornith-1.0, the 1.5 series implements a complete self-improvement loop. During training, the model does not merely consume static datasets — it proposes its own new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning. The model effectively becomes its own teacher, creating a continuous cycle where every round of learning compounds.

The practical consequence for developers is stated plainly in the release: "frontier-class" intelligence on much more accessible hardware. Traditional scaling laws assume you pay for reasoning depth with parameter mass. Ornith's bet is that self-generated curriculum lets a mid-sized MoE bridge the gap to frontier systems — and the benchmark sheet below is the evidence they're submitting.

Benchmarks: The Full Sheet

The numbers Ornith published for the 35B-A3B variant read like a frontier model's scorecard, not a 3B-active open-weights release:

BenchmarkScoreWhat It Measures
SWE-Bench Verified86Real repo-level engineering
Terminal-Bench 2.186.1Terminal & agentic autonomy
SWE-Bench Multilingual79.6Multilingual coding
SWE-Bench Pro65.1Harder engineering set
ClawEval81.4General evaluation
Tool Decathlon71.2OS / API / tool usage
DeepSWE56Deep software engineering
HLE (Humanity's Last Exam)44.6Expert-level reasoning

Against closed-source competitors, Ornith claims the 35B variant performs on par with Claude Opus 4.8 across reasoning, agentic workflows, and complex instruction following — while holding a commanding lead in Terminal-Bench 2.1 and showing strong versatility in multilingual environments.

⚡ The economics angle: because the model is MIT-licensed and only 3B parameters are active per token, you can self-host quantized builds (FP8, GGUF, MLX, NVFP4) instead of renting frontier APIs. For agentic workloads that burn tokens in long loops, that difference compounds fast.

Where It Fits

Ornith positions the 35B-A3B as purpose-built for high-agency environments. The Tool Decathlon score of 71.2 is the tell: this model is meant to drive autonomous agents that interact with operating systems, APIs, and complex software environments. The release calls out four primary use cases:

The model's ability to follow complex, multi-step instructions is highlighted as a reliability feature: it holds together even as prompt complexity increases — exactly the property you want in an agent that has been running for forty tool calls.

Getting Started: Weights, Quants, Frameworks

True to the MIT license, everything is available for immediate self-hosting:

Ornith's own recommendation: start with the 35B MoE variant to balance cost and capability in production pipelines, and use the 9B Dense model as the entry point for edge cases and low-latency requirements. Managed API endpoints are also available for teams that would rather not run their own inference — the release notes third-party provider pricing may vary.

Why This Release Matters

Open-source releases usually trail the frontier by a generation — a good MoE lands, and the lab up the road has already moved on. Ornith-1.5's claim is different in kind: the self-improvement loop means the training method itself scales, not just the checkpoint. If a 35B model with 3B active parameters can sit alongside Claude Opus 4.8 on the benchmarks that matter for agents, then the frontier is no longer defined by who has the biggest cluster — it is defined by who has the best learning loop. That is a much more open race, and as of August 20, 2026, open source is in it.

Fact Verification & Sources

This technical summary was compiled exclusively using verified public data points from the following release documentation:

⚡ OpenAdapter Readers get 20% off — invite code BDPBCR3R ◉ Z.ai Coding Plan Readers get 10% off — invite code R0K78RJKNW
R
Analyzed for CLAW

Live analysis published on claw.rommark.dev on Aug 20, 2026. Data grounded exclusively in official release benchmarks and verified technical specifications.