↓ Skip to main content
  1. Articles/

Meta Ships Muse Glimmer: An Open-Weight 30B Agent Model That Runs on Your Desk

·1095 words·6 mins·
Florent Clairambault
Author
Florent Clairambault
CTO & software engineer — writing daily about spec-driven development and agentic coding

Meta Ships Muse Glimmer: An Open-Weight 30B Agent Model That Runs on Your Desk
Meta shipped Muse Glimmer on August 10 — a 30-billion-parameter agentic model released under Apache 2.0, distilled down from Meta’s closed frontier model, Muse Spark, via logit distillation. The pitch is narrow and specific: not a chat model, not a benchmark flex, but an agent model sized to run entirely on the machine sitting in front of you, with no API key and no network call required.

That framing is worth pausing on, because it’s a reversal for Meta. Muse Spark, the model Glimmer is distilled from, is closed-weight and API-only. Muse Code, Meta’s terminal coding agent, charges for tokens unless you hand over your prompts for training. Muse Glimmer is the opposite move: full weights, a permissive license, and a design brief built around running without Meta’s servers in the loop at all. Whether that’s principle or just a different monetization lane — device makers and edge deployments are the obvious beneficiaries — is a fair question. Either way, the model itself is real and worth understanding on its own terms.

What Muse Glimmer actually is
#

The architecture is a 28B text decoder paired with a 2B vision encoder (a ViT-style “Perception Encoder” that handles both images and video, sampled at 2 fps up to 96 frames), for a combined 30B parameters. It uses hybrid attention — sliding-window layers (2,048 tokens) alternating with full-attention layers — plus Gated Grouped-Query Attention at a 16:1 query-to-key/value ratio, which is where most of the memory savings versus a dense model of similar quality come from. Meta and NVIDIA both cite a context window in the 120K+ token range, enough for long-running agent loops without constant summarization.

On capability, Meta’s own claim is that Glimmer beats Gemma4-31B and Qwen3.6-27B — the obvious open-weight comparison class — across agentic, coding, multimodal, safety, and reasoning evaluations. Selected published numbers: 75.5 on MCP Atlas, 74.6 on DeepSearch QA, 51.2 on SWE-Bench Pro, 43.6 on SciCode, 94.7 on AIME 2026. Treat self-reported numbers from any lab with the usual skepticism — Meta’s own Muse Spark 1.1 benchmarks didn’t survive an independent Vals AI rerun back in July — but the SWE-Bench Pro number in particular is a reasonable sanity check since it’s the benchmark most resistant to contamination.

The feature set is agent-shaped rather than chat-shaped: reliable function calling against precise schemas, multi-step reasoning across extended workflows, failure recovery and error diagnosis when a tool call goes wrong, controllable reasoning effort, and multilingual support across 100+ languages.

What kind of computer can actually run it
#

This is the part that matters if you’re deciding whether to bother. At full BF16 precision, 30B parameters need north of 55GB of memory — not a laptop workload. Meta’s answer is 4-bit quantization (what they call K-Quant-Dynamic), which brings the weights under 20GB, leaving headroom in a 24GB or 32GB memory budget for the KV cache, the vision encoder, and speculative-decoding state.

In practice, that puts three tiers of hardware in range:

  • Apple Silicon laptops. Meta validated on MacBook M4 Max and M5 Max, using unified memory to hold the quantized model. Combined with DFlash speculative decoding, Meta reports 1.5x faster decoding on M4 Max and 1.8x on M5 Max versus non-speculative inference.
  • A single high-end consumer GPU. An RTX 5090 (32GB VRAM) runs the model with no sharding or CPU offload, and gets the largest speculative-decoding boost Meta reported — 3.1x. AMD has published a companion guide for Ryzen AI Max “agentic PCs” and Radeon GPUs in the same class.
  • Workstation and edge silicon. NVIDIA lists DGX Spark and DGX Station for on-prem workstation deployment, and Jetson for embedded/robotics use, where the appeal is a single dense model that fits on one chip without MoE-style routing overhead. On Blackwell Ultra, NVIDIA reports over 20,000 tokens/sec per GPU at BF16/NVF4 — workstation-tier throughput, not laptop-tier.

The floor, in short: a 2026 flagship Apple Silicon laptop or a single 24–32GB consumer GPU. Below that, you’re back to quantizing further or renting a GPU — at which point some of the “runs entirely offline, no API key” appeal starts to erode.

Day-0 support landed across the usual local-inference stack — Ollama, LM Studio, llama.cpp, MLX, ExecuTorch, vLLM, SGLang — plus hosted options through Hugging Face Inference Endpoints, Together AI, Fireworks AI, and OpenRouter for anyone who wants the model without the hardware.

Sample use cases
#

Meta and its launch partners are pointing at a specific category: agents that run continuously on a machine you already own, rather than a chatbot you visit.

  • Local coding assistant. Code generation, project scaffolding, and repo-scoped edits without sending source to a third-party API — the same pitch as self-hosting an open coding model, but sized for a single workstation instead of a GPU cluster.
  • Desktop automation agents. Scheduling, messaging, and file organization tasks that need reliable tool calling and can run “always-on” in the background without per-token API cost.
  • Document and screenshot understanding. The multimodal encoder handles screenshots and documents directly, useful for agents that need to read a UI or a PDF as part of a workflow rather than requiring pre-extracted text.
  • LLM-as-a-judge and evaluation harnesses. A capable local model is cheap to run at the volume evaluation loops require, without burning API budget on judge calls.
  • Edge and robotics. The Jetson deployment path targets industrial automation and robotics use cases where offline operation and predictable latency matter more than peak throughput.

Where this fits
#

None of this makes Muse Glimmer a Claude Code competitor — it’s a model, not an agentic harness, and Meta ships no equivalent to worktree isolation, persistent bash sessions, or multi-agent orchestration around it. But it’s a genuinely useful data point: the same lab that closed the door on open frontier weights with Muse Spark is now shipping a real, permissively-licensed, locally-runnable agent model with legitimate coding and tool-use numbers. If your workload is bounded, latency-sensitive, or has to run without network access — an internal tool, an edge device, a privacy-constrained coding assistant — Glimmer is worth a serious look. If your workload is autonomous multi-step engineering work across a real codebase, the model is only half the story, and the harness around it still isn’t Meta’s strength.


Sources: Meta AI Research: Introducing Muse Glimmer, Hugging Face: Meta is back with Muse Glimmer, NVIDIA Technical Blog: Run Local Agentic AI Workflows with Muse Glimmer, AMD Blogs: Run Muse Glimmer 30B on Ryzen AI Max and Radeon GPUs, Techzine: Meta releases Muse Glimmer as an open local agent model, Engadget: Meta’s ‘open source’ Muse Glimmer model can run on a single computer

Related