Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Meta drops Muse Glimmer โ€” 30B open agentic weights for one GPU

Listen to this article

Local agents just got a real open-weights option that fits on a desk. Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter agentic model under Apache 2.0, built to run on one consumer GPU or a Mac without calling the cloud.

What shipped today

Per Meta’s research blog, Muse Glimmer is distilled from Muse Spark and tuned for always-on local agent workflows: function calling, local coding, multi-step tool use, and LLM-as-a-judge work. Weights are on Hugging Face now. Optimized integrations for llama.cpp, MLX, and ExecuTorch are promised in the coming days.

The hardware pitch is the point. At full precision, Meta says a 30B would need over 55 GB. Quantized to roughly 4-bit, the language model drops under ~20 GB, leaving room for KV cache, the perception encoder, and a speculative-decoding drafter inside a 24 GB or 32 GB envelope. Speed comes from DFlash, a block-level drafter that proposes tokens the main model verifies in parallel. Input is text and image. Context is 131k+. Output is text.

Where the benches split

MarkTechPost’s breakdown of Meta’s comparisons against Gemma4-31B and Qwen3.6-27B (thinking mode) shows a clean split:

  • Leads: MCP Atlas, DeepSearch QA, SWE-Bench Pro
  • Trails: OSWorld-Verified, TerminalBench

In plain terms: stronger on agentic orchestration, search-style QA, and software-engineering scaffolds. Weaker on computer-use and terminal work. That is not a small asterisk if your agent is supposed to drive a desktop or a shell.

Why one-GPU agents matter

Cloud agents are a tax and a trust problem. Per-token bills scale with every tool loop. Personal context (calendar, files, screenshots) wants to stay on the machine. A 30B Apache 2.0 model that fits a single consumer card changes who can ship local agents: solo builders, startups avoiding API bills, and teams that need air-gapped or offline runs.

It also resets the open-weight race in the mid-30B class. Meta is not asking you to rent its stack. It is handing you weights and telling you to run them next to the rest of your toolchain.

Strong agent benches, real desktop gaps

This is Meta doing what it still does better than anyone at the frontier: dumping capable weights into the open and letting the ecosystem finish the product. The agentic benches Meta likes look real. The computer-use and terminal gaps look equally real. If you need a local coding or tool-calling agent today, Glimmer is suddenly on the shortlist. If you need a model that reliably drives OSWorld-style desktop work, do not let the launch blog paper over the loss column.

Apache 2.0 plus Hugging Face day one is the distribution win. The runtime win lands when llama.cpp, MLX, and ExecuTorch packs actually ship and hold the claimed latency on M-series Macs and mid-range NVIDIA cards.

When llama.cpp and MLX actually land

Track the promised llama.cpp / MLX / ExecuTorch drops and independent runs on MCP Atlas, SWE-Bench Pro, OSWorld-Verified, and TerminalBench. Vendor benches opened the door. Community numbers decide whether Glimmer is the default one-GPU agent or just another 30B with a good press day.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile