Meta bets on local AI with Muse Glimmer 30B for PCs and Macs
New Delhi: Meta has introduced Muse Glimmer, a 30-billion-parameter open-weight AI model aimed at running agent-style workloads directly on personal computers. The model has been released under the Apache 2.0 licence and is designed for tasks such as coding, tool use, document analysis, local assistants and multi-step workflows.

The big pitch here is local AI. Meta says Muse Glimmer can run within a 24GB or 32GB memory envelope after quantisation, which brings a 30B model within reach of high-end consumer PCs and Macs. It supports text and images, more than 100 languages, long-context tasks and different reasoning levels.
1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon.
we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
— Alexandr Wang (@alexandr_wang) August 10, 2026
How Muse Glimmer performs against Gemma4 and Qwen3.6
Meta compared Muse Glimmer 30B against Gemma4-31B and Qwen3.6-27B across agentic, coding, multimodal and reasoning benchmarks.
Muse Glimmer leads several agent-focused tests in the shared benchmark table. It scores 75.5 on MCP Atlas, ahead of Gemma4 at 54.2 and Qwen3.6 at 62.5. On DeepSearch QA, Muse Glimmer reaches 74.6, compared with 61.7 and 71.1 respectively.
It also posts 43.3 on GAIA2, beating Gemma4 at 36.4 and Qwen3.6 at 40.0.
Coding results are more mixed.
Muse Glimmer scores 51.2 on SWE-Bench Pro, narrowly ahead of Qwen3.6 at 50.2. On SWE-Bench Verified, Qwen3.6 leads with 77.2, followed by Muse Glimmer at 76.0. Qwen3.6 also leads TerminalBench 2.1 with 60.7 against Muse Glimmer’s 51.7.
For SciCode, Muse Glimmer takes a small lead at 43.6, compared with 43.4 for Gemma4 and 39.8 for Qwen3.6.
Reasoning scores show a mixed race
Muse Glimmer performs strongly on several general reasoning tests.
It scores 94.7 on AIME 2026, slightly ahead of Qwen3.6 at 94.1 and Gemma4 at 89.2. On IFBench, Muse leads with 77.0.
Gemma4 performs better on GPQA Diamond with 85.7, compared with Muse Glimmer’s 83.5. It also leads Humanity’s Last Exam with 23.6, against 22.0 for Muse Glimmer and 23.1 for Qwen3.6.
Qwen3.6 stays ahead on several multimodal tests, including ScreenSpot Pro at 76.1 and OmniDocBench v1.5 at 77.8.
How Meta made a 30B model fit on consumer hardware
A full-precision 30B model would need more than 55GB of memory. Meta says it reduced the model to under 20GB using roughly 4-bit quantisation.
Muse Glimmer was trained using outputs from the larger Muse Spark model, followed by longer-context training, supervised fine-tuning, reinforcement learning and further distillation.
Meta is also using a smaller DFlash drafter model for speculative decoding. According to the company, this improves decoding speed by 3.1 times on an RTX 5090, 1.8 times on an M5 Max and 1.5 times on an M4 Max.
Muse Glimmer weights are available through Hugging Face, with support for llama.cpp, MLX and ExecuTorch expected to follow. The wider significance is fairly clear. Meta is betting that capable AI agents do not always need to live in the cloud, and a growing share of them could soon run directly on a user’s own machine.
