Skip to content
← All news
4 min read

Meta ships Muse Glimmer, a 30B open model built for offline AI agents

The 4-bit build fits under 20GB and runs on one consumer GPU or a Mac, with no cloud account or metered tokens required.

The quantized build is under 20GB. The unquantized model is closer to 60, three times the pitch.

Meta Superintelligence Labs released Muse Glimmer on August 10, 2026: a 30 billion parameter open-weight model, licensed Apache 2.0, built specifically to run AI agents on your own machine instead of someone else's cloud. The 4-bit quantized build fits under 20GB, small enough for a single consumer GPU or a Mac, and it handles tool calling, multi-step reasoning, coding agents, and LLM-as-judge work out of the box. Mark Zuckerberg tied the release to Meta's push for what he is calling decentralized, open-source superintelligence, positioning a model that runs on your own hardware as the alternative to agent products that live entirely inside someone else's cloud. It is the first agent-focused open release from the lab that Alexander Wang, Scale AI's former CEO, now runs at Meta.

what's actually inside the release

Muse Glimmer is a dense 30B transformer, not a mixture-of-experts model, and Meta trained it to be quantized from the start instead of shrinking a full-size model after the fact. That quantization-aware training is why the 4-bit build holds up well enough to ship as the default download. Meta also built in DFlash, a speculative decoding method the company says speeds up generation by 1.5x to 3.1x depending on the hardware, tested across an RTX 5090 and Apple's M4 Max and M5 Max chips.

The model went up on Hugging Face the same day, with day-one support across the tools agent builders already use:

  • llama.cpp
  • MLX
  • ExecuTorch
  • Ollama
  • vLLM
  • SGLang

the pitch: skip the cloud entirely

Meta's stated goal is what it calls always-on local agent workflows: an agent that keeps running with no cloud dependency, no account to create, and no metered tokens counting against you per call. Forbes read the release the same way, framing Muse Glimmer as a direct shot at cloud-hosted agent products that charge per token and require an account just to start. For a developer running an agent continuously to watch a codebase, monitor logs, or run scheduled jobs, that pricing model matters as much as the model's raw capability.

Muse Glimmer turns running an AI agent from a subscription and a network connection into a laptop and a free download.

The honest caveat showed up fast in the 1,182-point Hacker News thread under the announcement. Several commenters benchmarked Muse Glimmer against Alibaba's Qwen3.6-27B and found it barely edges out that model, and only on tool-calling tasks specifically, not across the board. Others pointed out that the marketed 20GB figure only holds for the 4-bit quantized build. The unquantized weights run close to 60GB, three times the number in the headline, and quantization always costs some accuracy even when a model is trained around it from the start.

Why a build studio cares

This is squarely our third pillar: AI workflows and custom agents, built and shipped for clients rather than written about from a distance. A lot of that work involves an agent running against a client's codebase, internal docs, or logs, and plenty of those clients do not want that data leaving their network or a per-token cloud bill attached to an agent that runs all day. A 30B model that fits on hardware someone already owns, with day-one support in Ollama and llama.cpp, is a real option to put in front of a client for that exact case. It will not replace every job: Qwen3.6-27B is similarly sized and sitting right there, and picking between them comes down to whether tool-calling is the actual bottleneck in a given agent, not a default either way.

Read Meta's announcement for the full technical writeup, then read the 1,182-point Hacker News discussion for the benchmarks Meta did not publish itself. If you are weighing a self-hosted agent against a cloud API for a client build, write to us at hello@gattyworks.com.

Open Weight ModelsAI AgentsLocal AIMetaMuseGlimmerAlexanderWangHuggingFaceOpenSourceAIAIAgentsLocalAIOllamaOpenWeightModelsMachineLearning

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.