Ornith released the Ornith-1.5 family on August 20, 2026 under an MIT license, headlined by a 35B-total / 3B-active MoE that is claimed to perform on par with Claude Opus 4.8 across reasoning, agentic workflows and complex instruction following. The hook isn't just the scores — 86 on SWE-Bench Verified and 86.1 on Terminal-Bench 2.1 — it's how the model was trained: an end-to-end self-improvement loop in which the model proposes its own tasks, generates its own scaffolds, and rolls out solutions for reinforcement learning. Open weights that teach themselves.
The Ornith-1.5 family is a three-tier lineup: a 9B Dense model for edge applications, the flagship 35B MoE (A3B) for high-performance reasoning, and a 397B MoE for ultra-complex enterprise workloads. The release positions itself as the moment open-source models stop chasing proprietary giants and start setting benchmarks for autonomous intelligence themselves. Every tier ships under the same permissive license, which allows unrestricted commercial and research use.
The flagship uses an "Active 3B" architecture: 35 billion total parameters, but only 3 billion activated per token — the sparse-activation trick that keeps inference fast while preserving the knowledge capacity of a much larger model. What distinguishes the 1.5 series is the training methodology on top of it. Building on the self-scaffolding strategies introduced in Ornith-1.0, the model doesn't merely consume static datasets: during training it proposes its own new tasks, generates task-specific scaffolds, and produces solution rollouts for reinforcement learning — a continuous cycle in which the model effectively becomes its own teacher.
⚡ Key Detail: the self-improvement loop is the release's core claim to fame. Instead of humans curating every training signal, the model generates the tasks and scaffolds it learns from — which is how a 35B/3B MoE is claimed to bridge the gap to frontier-scale systems without frontier-scale training data pipelines.
The headline claim is parity with Claude Opus 4.8 across reasoning, agentic workflows, and complex instruction following. Here is the complete verified scoreboard for the 35B-A3B variant, exactly as published in the release:
| Benchmark | Score | What It Tests |
|---|---|---|
| SWE-Bench Verified | 86 | Real-world repository-level engineering |
| SWE-Bench Pro | 65.1 | Harder agentic coding |
| SWE-Bench Multilingual | 79.6 | Multilingual repositories |
| Terminal-Bench 2.1 | 86.1 | Terminal & tool usage |
| DeepSWE | 56 | Deep software engineering |
| HLE (Humanity’s Last Exam) | 44.6 | Frontier reasoning |
| ClawEval | 81.4 | Agentic evaluation suite |
| Tool Decathlon | 71.2 | OS, API & environment interaction |
Two rows stand out. An 86 verified on SWE-Bench means the model resolves real repository-level engineering tasks — the kind that previously demanded a much larger closed model — with high precision. And the 86.1 on Terminal-Bench 2.1, described in the release as a commanding lead, points at agents that live in the shell: administering systems, chaining tools, and recovering from errors without hand-holding. The Tool Decathlon score of 71.2 reinforces that profile — the release specifically flags it as the reason this model suits autonomous agents that interact with operating systems, APIs, and complex software environments.
Because the weights are MIT-licensed, Ornith sidesteps the usual open-weights caveat — there is no usage clause to lawyer over. The release lists quantized builds in FP8, GGUF, MLX, and NVFP4 formats, with framework support for vLLM, llama.cpp, and MLX. The MLX builds in particular mean Apple Silicon users get high-performance local inference out of the box. Weights are available via Hugging Face and the official Ornith GitHub repository, and Ollama already lists the model in its library.
If you build autonomous coding agents, the SWE-Bench profile makes this an immediate evaluation candidate. If you run terminal-heavy automation or system administration agents, Terminal-Bench 2.1 at 86.1 is the number to benchmark against your current stack. And if you operate enterprise RAG in multilingual environments, the 79.6 Multilingual score plus the model's stated strength in retrieval-augmented generation and long-context reasoning keep it on the shortlist — all without a single API bill, if you'd rather host it yourself.
This technical summary was compiled exclusively using verified public data points from the following release documentation: