Ornith released the Ornith-1.5 family on August 20, 2026, and the flagship 35B-A3B is the story: a Mixture-of-Experts model with 35B total parameters but only 3B active per token, trained through an end-to-end self-improvement loop, shipping under an MIT license — and posting 86 on SWE-Bench Verified and 86.1 on Terminal-Bench 2.1. According to the release, it performs on par with Claude Opus 4.8 across reasoning, agentic workflows, and complex instruction following. An open-weights model that proposes its own training tasks is a genuine inflection point.
The Ornith-1.5 family is a three-tier lineup built for scalability: a 9B Dense model aimed at edge applications, the flagship 35B MoE (A3B) for high-performance reasoning, and a 397B MoE variant for ultra-complex enterprise workloads. The headline designation — "A3B" — means that despite 35 billion total parameters, only 3 billion are activated per token, which is what keeps inference fast on accessible hardware.
What separates Ornith-1.5 from the weekly parade of fine-tunes is the training methodology. Building on the self-scaffolding strategies introduced in Ornith-1.0, the 1.5 series implements a complete self-improvement loop. During training, the model does not merely consume static datasets — it proposes its own new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning. The model effectively becomes its own teacher, creating a continuous cycle where every round of learning compounds.
The practical consequence for developers is stated plainly in the release: "frontier-class" intelligence on much more accessible hardware. Traditional scaling laws assume you pay for reasoning depth with parameter mass. Ornith's bet is that self-generated curriculum lets a mid-sized MoE bridge the gap to frontier systems — and the benchmark sheet below is the evidence they're submitting.
The numbers Ornith published for the 35B-A3B variant read like a frontier model's scorecard, not a 3B-active open-weights release:
| Benchmark | Score | What It Measures |
|---|---|---|
| SWE-Bench Verified | 86 | Real repo-level engineering |
| Terminal-Bench 2.1 | 86.1 | Terminal & agentic autonomy |
| SWE-Bench Multilingual | 79.6 | Multilingual coding |
| SWE-Bench Pro | 65.1 | Harder engineering set |
| ClawEval | 81.4 | General evaluation |
| Tool Decathlon | 71.2 | OS / API / tool usage |
| DeepSWE | 56 | Deep software engineering |
| HLE (Humanity's Last Exam) | 44.6 | Expert-level reasoning |
Against closed-source competitors, Ornith claims the 35B variant performs on par with Claude Opus 4.8 across reasoning, agentic workflows, and complex instruction following — while holding a commanding lead in Terminal-Bench 2.1 and showing strong versatility in multilingual environments.
⚡ The economics angle: because the model is MIT-licensed and only 3B parameters are active per token, you can self-host quantized builds (FP8, GGUF, MLX, NVFP4) instead of renting frontier APIs. For agentic workloads that burn tokens in long loops, that difference compounds fast.
Ornith positions the 35B-A3B as purpose-built for high-agency environments. The Tool Decathlon score of 71.2 is the tell: this model is meant to drive autonomous agents that interact with operating systems, APIs, and complex software environments. The release calls out four primary use cases:
The model's ability to follow complex, multi-step instructions is highlighted as a reliability feature: it holds together even as prompt complexity increases — exactly the property you want in an agent that has been running for forty tool calls.
True to the MIT license, everything is available for immediate self-hosting:
Ornith's own recommendation: start with the 35B MoE variant to balance cost and capability in production pipelines, and use the 9B Dense model as the entry point for edge cases and low-latency requirements. Managed API endpoints are also available for teams that would rather not run their own inference — the release notes third-party provider pricing may vary.
Open-source releases usually trail the frontier by a generation — a good MoE lands, and the lab up the road has already moved on. Ornith-1.5's claim is different in kind: the self-improvement loop means the training method itself scales, not just the checkpoint. If a 35B model with 3B active parameters can sit alongside Claude Opus 4.8 on the benchmarks that matter for agents, then the frontier is no longer defined by who has the biggest cluster — it is defined by who has the best learning loop. That is a much more open race, and as of August 20, 2026, open source is in it.
This technical summary was compiled exclusively using verified public data points from the following release documentation: