Muse Glimmer 30B Review 2026: Meta's Apache 2.0 Agent Model

muse-glimmermetaollamaselfhostedagents

TL;DR: Muse Glimmer 30B (released August 10, 2026) is Meta’s first open-weight model since Llama 4, and the license is clean Apache 2.0 — no Llama-style community license, no user caps, no regional carve-outs. The 4-bit GGUF is ~17 GB and runs on a 24 GB GPU. On agentic benchmarks (MCP Atlas, DeepSearch QA, long-context reasoning) it beats Qwen3.6-27B; on computer-use and terminal benchmarks Qwen3.6 still wins. If your workload is tool-calling loops on a single 24 GB card, this is the new default. If it’s general chat or terminal-heavy coding, it isn’t.

Muse Glimmer 30BQwen3.6-35B-A3BQwen3.6-27B
LicenseApache 2.0Apache 2.0Apache 2.0
ArchitectureDense, 2B vision encoder + 28B decoderMoE, ~3B activeDense
4-bit GGUF size~17 GB~19 GB~16 GB
Context131K262K131K
Best atTool-calling agent loops, long-contextCheap fast inference per tokenTerminal/computer use

Is Muse Glimmer actually Apache 2.0?

Yes — and that is the headline, not the benchmarks. Meta’s open releases since 2023 shipped under Llama Community Licenses with the 700M-MAU clause and acceptable-use addenda. Muse Glimmer 30B, the first open-weight release from the restructured Meta Superintelligence Labs, ships under plain Apache 2.0 with weights on Hugging Face (RuntimeWire, August 2026). Commercial use, fine-tuning, redistribution, and building paid products on top are all unrestricted.

Two things worth knowing before you treat it as a drop-in Llama successor:

  • It is not a Llama. It’s distilled from Muse, Meta’s closed frontier family — the same lineage as the closed Muse Spark we covered in Meta Muse Spark: closed-source, and the open-weight alternatives. Glimmer is the open distillate of that stack.
  • The Apache 2.0 grant covers the weights and inference code. Meta’s hosted Muse API has separate terms — irrelevant for self-hosters, but don’t confuse the two when reading the docs.

What is Muse Glimmer 30B, exactly?

Muse Glimmer is a 30B-parameter dense multimodal model built for always-on local agent workflows: function calling, tool use over long horizons, coding inside agent loops, and LLM-as-judge. The 30B splits into a ~2B ViT-style Perception Encoder (image input) bolted onto a ~28B text decoder. Context window is 131K tokens.

The design choice that matters for self-hosters is block-level speculative decoding: Meta ships an optional DFlash drafter that predicts blocks of tokens ahead, aimed at making agent loops — where the model generates short structured outputs over and over — feel responsive on consumer hardware. vLLM and SGLang support the drafter; llama.cpp runs the model fine without it.

Launch-day runtime support was unusually broad: llama.cpp, Ollama, LM Studio, MLX, vLLM, SGLang, and ExecuTorch, plus hosted access via Together AI, Fireworks, and OpenRouter if you want to try it before downloading 17 GB.

How much VRAM does Muse Glimmer 30B need?

The 4-bit GGUF is ~17 GB and the documented envelope is a 24 GB GPU; full precision needs over 55 GB. Per the Unsloth GGUF documentation (September 2026), the quantization tiers break down like this:

Quant tierFile sizeFits onKnowledge risk
8-bit~34 GB48 GB cards, M4/M5 Max 64GB+None
6-bit20–22 GB32 GB (RTX 5090)Negligible
4-bit (~Q4_K_M)~17 GB24 GB (RTX 3090/4090)Low — the sane default
3-bit14–15 GB16 GB cards, tightNoticeable
2-bit12–14 GB16 GB cardsSevere — avoid

Meta also publishes a “K-Quant-Dynamic” tier with a documented 32 GB VRAM envelope — it keeps sensitive layers at higher precision. On a 24 GB card, pull the 17 GB tier, not the dynamic one.

A problem you will actually hit: the 17 GB quant loads fine on a 24 GB card, but an agent session that runs for hours accumulates KV cache on top of the weights, and at 131K context that cache alone can eat the remaining ~7 GB and spill to CPU — at which point generation speed falls off a cliff mid-session rather than at load time. The fix is to cap context below the maximum (--ctx-size 32768 covers most tool loops) or enable KV cache quantization (--cache-type-k q8_0 --cache-type-v q8_0 in llama.cpp). Treat 131K as a ceiling for occasional long documents, not a default allocation.

Below 20 GB of VRAM, this model is the wrong pick — see the “when not to use” section.

How do you run Muse Glimmer with Ollama and llama.cpp?

Ollama was in the launch-day runtime matrix, and the model lives in the Ollama library as muse-glimmer:

$ ollama pull muse-glimmer
pulling manifest
pulling 8f2a1c9e04b1... 100% ▕████████████████▏  17 GB
success

$ ollama run muse-glimmer "List three risks of exposing an Ollama port to the internet. Return JSON."

For llama.cpp with the KV-cache guard rails from the previous section:

./llama-server -m muse-glimmer-30b-Q4_K_M.gguf \
  --ctx-size 32768 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --port 8080

That exposes an OpenAI-compatible endpoint, so anything that speaks to Ollama or OpenAI — Open WebUI, Continue, Cline — can use Glimmer as a backend. For high-throughput serving with the DFlash speculative drafter, vLLM is the path instead; the tradeoffs are the same ones in our Ollama vs vLLM comparison.

Speed, from bandwidth math rather than vendor numbers: a dense ~28B decoder at 4-bit reads roughly 15–16 GB of weights per token, so an RTX 3090 (936 GB/s) has a hard ceiling around 58 tok/s and will realistically land in the 40–55 tok/s range before speculative decoding. That’s comfortable for agent loops. On a 16 GB card forced down to 3-bit with partial CPU offload, expect single digits — not comfortable.

How does Muse Glimmer compare to Qwen3.6 for agent work?

Muse Glimmer wins the agentic columns and loses the computer-use ones. Meta’s own launch tables put it ahead of Gemma4-31B and Qwen3.6-27B on MCP Atlas, GAIA2, SWE-Bench Pro, and AIME 2026 — vendor numbers, treat accordingly. The independent comparison at BenchLM (September 2026) broadly agrees on direction but shows the split:

BenchmarkMuse Glimmer 30BQwen3.6-27B
MCP Atlas (tool use)75.5lower
DeepSearch QA74.671.1
AA-LCR (long-context reasoning)80.073.3
OSWorld-Verified (computer use)~10 pts behindwins
TerminalBench 2.19 pts behindwins

BenchLM’s summary line is the honest verdict: “the better agent, not the better model.” If your loop is MCP tool calls, retrieval, and structured output over long horizons, Glimmer’s wins are exactly in that column. If your agent drives a terminal or a GUI, Qwen3.6-27B is still the stronger pick at the same VRAM budget.

Against Qwen3.6-35B-A3B the tradeoff is architectural: the MoE activates ~3B parameters per token, so it decodes several times faster on the same card and leaves more headroom for KV cache. Glimmer is dense — slower per token, but with no MoE routing variance, which matters for judge-style evaluations where consistency beats speed.

When should you NOT use Muse Glimmer?

  • Under 20 GB of VRAM. The 17 GB quant leaves no agent-loop headroom on a 16 GB card, and the 3-bit/2-bit tiers trade away stored knowledge fast — the degradation is non-linear, as the 55-quant study we covered showed at 27B scale. On 12–16 GB, run Qwen3.6-35B-A3B or a 12B-class dense model instead.
  • Terminal- or GUI-driving agents. Qwen3.6-27B beats it by ~9–10 points on TerminalBench 2.1 and OSWorld-Verified. Same VRAM, better fit.
  • General chat and writing. Nothing is wrong with it, but you’re paying dense-30B decode speed for agentic training you won’t use; the MoE alternatives are faster for free-form text.
  • Anything needing vision output. The Perception Encoder handles image input only. It does not generate images.
  • Day-one production. The weights are under two months old. Quant releases from third parties (Unsloth, bartowski) stabilized within weeks, but if you need a model with a year of known failure modes, this isn’t it yet.

What hardware does it actually take?

Prices verified September 2026; the DRAM/GPU repricing of 2025–26 means street prices, not MSRPs:

Your situationRun it onPriceWhere
Want the 4-bit tier with headroomUsed RTX 3090 24GB$1,150–$1,350Check price
Want the 6-bit tier or dynamic quantRTX 5090 32GB$3,822–$5,000 (MSRP $1,999)Check price
Undecided — test the workload firstRented RTX 3090, from $0.07/hrpay per hourVast.ai

A used 3090 remains the floor for this model done right. If you’re weighing a unified-memory box instead (Strix Halo class), the bandwidth there (256 GB/s) caps a dense 28B decoder near 16 tok/s — workable, slower; see runaihome’s Ryzen AI Max 395 Strix Halo guide for that math. And if the end goal is a local coding-agent backend, the editor-side comparison at aicoderscope — Kilo Code vs OpenCode vs Cline — covers the tools you’d point at it.

Verdict

Muse Glimmer 30B is the best Apache 2.0 agent model you can run on a 24 GB card as of October 2026, and the license alone makes it the most significant Meta release since Llama 3. It is not the best general model at its size — Qwen3.6 keeps the terminal and computer-use crown, and MoE rivals decode faster. Pull the 17 GB quant, cap your context, and judge it on your own tool loop; that’s the workload it was built for and the one where it earns the download.

FAQ

Does Muse Glimmer run on a 16 GB GPU? Only at 3-bit or 2-bit quantization (12–15 GB files), and the knowledge loss at those tiers is severe at this model scale. A 16 GB card is better served by Qwen3.6-35B-A3B, whose MoE design needs less bandwidth per token.

Is Muse Glimmer multimodal? Input-only. The ~2B Perception Encoder accepts images alongside text (screenshots in an agent loop, document pages), but the model outputs text only — it does not generate images.

Can I use Muse Glimmer commercially? Yes. Apache 2.0 places no restrictions on commercial use, redistribution, or fine-tuning, with no user-count thresholds — unlike the Llama Community Licenses on Meta’s earlier open models.

Sources

Was this article helpful?

What self-hosting actually costs

Real cost breakdowns for self-hosted AI: hardware floors, power, maintenance hours, and the honest comparison against paying for it. No spam, unsubscribe anytime.