Meta drops Muse Glimmer — 30B open agentic weights for one GPU

Local agents just got a real open-weights option that fits on a desk. Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter agentic model under Apache 2.0, built to run on one consumer GPU or a Mac without calling the cloud.
What shipped today
Per Meta’s research blog, Muse Glimmer is distilled from Muse Spark and tuned for always-on local agent workflows: function calling, local coding, multi-step tool use, and LLM-as-a-judge work. Weights are on Hugging Face now. Optimized integrations for llama.cpp, MLX, and ExecuTorch are promised in the coming days.
The hardware pitch is the point. At full precision, Meta says a 30B would need over 55 GB. Quantized to roughly 4-bit, the language model drops under ~20 GB, leaving room for KV cache, the perception encoder, and a speculative-decoding drafter inside a 24 GB or 32 GB envelope. Speed comes from DFlash, a block-level drafter that proposes tokens the main model verifies in parallel. Input is text and image. Context is 131k+. Output is text.
Where the benches split
MarkTechPost’s breakdown of Meta’s comparisons against Gemma4-31B and Qwen3.6-27B (thinking mode) shows a clean split:
- Leads: MCP Atlas, DeepSearch QA, SWE-Bench Pro
- Trails: OSWorld-Verified, TerminalBench
In plain terms: stronger on agentic orchestration, search-style QA, and software-engineering scaffolds. Weaker on computer-use and terminal work. That is not a small asterisk if your agent is supposed to drive a desktop or a shell.
Why one-GPU agents matter
Cloud agents are a tax and a trust problem. Per-token bills scale with every tool loop. Personal context (calendar, files, screenshots) wants to stay on the machine. A 30B Apache 2.0 model that fits a single consumer card changes who can ship local agents: solo builders, startups avoiding API bills, and teams that need air-gapped or offline runs.
It also resets the open-weight race in the mid-30B class. Meta is not asking you to rent its stack. It is handing you weights and telling you to run them next to the rest of your toolchain.
Strong agent benches, real desktop gaps
This is Meta doing what it still does better than anyone at the frontier: dumping capable weights into the open and letting the ecosystem finish the product. The agentic benches Meta likes look real. The computer-use and terminal gaps look equally real. If you need a local coding or tool-calling agent today, Glimmer is suddenly on the shortlist. If you need a model that reliably drives OSWorld-style desktop work, do not let the launch blog paper over the loss column.
Apache 2.0 plus Hugging Face day one is the distribution win. The runtime win lands when llama.cpp, MLX, and ExecuTorch packs actually ship and hold the claimed latency on M-series Macs and mid-range NVIDIA cards.
When llama.cpp and MLX actually land
Track the promised llama.cpp / MLX / ExecuTorch drops and independent runs on MCP Atlas, SWE-Bench Pro, OSWorld-Verified, and TerminalBench. Vendor benches opened the door. Community numbers decide whether Glimmer is the default one-GPU agent or just another 30B with a good press day.



