Meta Superintelligence Labs has released Muse Glimmer (August 10, 2026) under the fully open Apache 2.0 license — a dense 30B parameter multimodal model (approximately 29.6B) built for autonomous agents, multi-step planning, and local deployment rather than general-purpose chat. With a dedicated ViT-G/14 vision encoder (~1.8B params), a 131,072+ token context window, and a 4-bit quantized footprint of under 20GB, it is a 30B-class agent model that runs on 24GB consumer GPUs and Apple Silicon — with adjustable reasoning levels from Low to X-High.
While the Mixture-of-Experts trend dominated 2024 and 2025, Meta went the other way: Muse Glimmer is a dense 30B parameter model (approximately 29.6B). The stated rationale is high parameter utilization and consistent performance across complex reasoning tasks. The model is natively multimodal, integrating a dedicated ViT-G/14 vision encoder with roughly 1.8B parameters, letting it process visual data with the same fluidity as text. Reasoning effort is adjustable across four levels — Low, Medium, High, and X-High.
In head-to-head evaluations, Muse Glimmer shows significant performance gains over both Gemma4-31B and Qwen3.6-27B in the 30B parameter class — particularly in tasks requiring long-horizon planning and error recovery. In SWE-bench evaluations, its ability to self-correct during coding tasks sets a new benchmark for open-weight models, per the release documentation.
| Model | Class | Context | Glimmer vs. Field |
|---|---|---|---|
| Muse Glimmer | 30B Dense + ViT-G/14 | 131,072+ | Apache 2.0 · Agents-First |
| Gemma4-31B | ~31B | — | Glimmer stronger in reasoning |
| Qwen3.6-27B | ~27B | — | Glimmer wins multimodal integration |
💡 Key Finding: A 30B model would typically require over 55GB of VRAM; with 4-bit quantization Muse Glimmer drops to under 20GB — in reach of 24GB/32GB consumer RTX GPUs, Mac Studio, and Mac Pro.
The headline engineering feat is consumer-hardware optimization. A 30B model would typically demand over 55GB of VRAM; Meta's engineers used advanced quantization and speculative decoding to collapse that requirement. With 4-bit quantization the footprint falls under 20GB, suited to 24GB-or-better consumer GPUs (RTX series) and Apple Silicon Macs. DFlash speculative decoding pushes high-speed local inference on top of that.
Muse Glimmer was trained end-to-end around the "agent loop": tool calling, multi-step planning, and — the part most chat models fail at — error recovery. When a tool call fails or environment state changes unexpectedly, Glimmer analyzes the error and re-plans its trajectory without human intervention. That makes it a candidate backbone for autonomous software engineers, research assistants, and complex workflow automation.
Glimmer is one stop on a larger open-weights roadmap. Meta has already announced that weights for the more advanced Muse Spark 1.2 foundation model will be released in the near future — a signal that the company intends to keep treating open weights as a primary driver of AI innovation, not a side channel. Weights are available now via Meta's official model repositories and supported local inference engines.
This technical summary was compiled exclusively using verified public data points from the following release documentation: