Skip links

Muse Glimmer tested: Meta’s open-weight model on real agency tasks

Cover reading Muse Glimmer, tested on agency tasks, on a pale blue background with soft blue spheres
Muse Glimmer, Meta’s open-weight 30B model, tested by AIWIZ against Qwen3.8-27B on routine agency work

Meta released Muse Glimmer on 10 August 2026. It is a 30-billion-parameter model that reads text and images, writes text, and comes with downloadable weights under the Apache License 2.0. Of everything Meta has launched under the Muse name, it is the only one you can run on your own hardware.

We wanted to know whether it could do real agency work, so we tested it. On three routine jobs (sorting search terms, checking ad copy against brand rules and auditing landing-page screenshots) it matched Alibaba’s Qwen3.8-27B answer for answer, replied more than five times faster and cost about 45% less. Independent benchmarks tell a different story about harder work, where Qwen is well ahead. Both results are true, and the gap between them is the most useful thing in this guide.

Checked against Meta’s research blog, developer docs and Hugging Face model card, and Artificial Analysis, on 4 October 2026. Meta updates its docs without notice.

The answer in brief

  • What it is: a 30B dense multimodal model (29.6B parameters on the Hugging Face card, including the vision encoder)
  • Licence: Apache License 2.0 on the released weights
  • Input and output: text and images in, text out
  • Context: 128K tokens by default (the card lists 131,072+)
  • Where to get it: Hugging Face, under meta-models, with no API key needed
  • Our test: level with Qwen3.8-27B on routine agency checks, with median replies of 6.4 seconds against 35.8
  • Independent benchmarks: Qwen3.8-27B scores twice as high on Artificial Analysis’s Intelligence Index (34 against 17) and leads on agentic work

Meta built Glimmer by distilling it from Muse Spark, its flagship hosted model. Spark itself stays closed. Meta has promised Spark weights more than once and not yet delivered them, which we cover below. Our Muse Spark 1.3 guide deals with the hosted side.

Glimmer is not an image model and has nothing to do with Ads Manager. If you came here from the advertising end, our Meta Muse overview maps the whole family and what is live in the UK.

This guide is for two readers: marketers who need to explain Meta’s AI line-up to clients without mixing up the names, and implementers deciding whether to run an open-weight model themselves. If you only manage Meta ads, or want a ready-made office agent such as Claude Cowork, you can stop here.

What did Meta ship with Muse Glimmer?

According to Meta’s launch post and developer docs, Glimmer was trained on Muse Spark’s outputs using logit distillation, then given further training on longer, agent-heavy tasks. Its knowledge stops at 4 January 2026. Meta says it was trained on data from more than 100 languages, but warns that performance may drop outside its strongly supported set, so test any non-English client work before relying on it.

Meta pitches it at local agents: calling tools, writing code, judging other models’ output, and reading screenshots, charts and documents. You can dial its reasoning effort up or down. Quantised, it is meant to run on a Mac or PC with a single consumer GPU.

The weights come in four official Hugging Face repositories: full-precision BF16, quantised GGUF files, a small “drafter” model that speeds up generation, and ExecuTorch exports for on-device use. Launch partners cover most of the usual tools, from Ollama, LM Studio and llama.cpp for desktops to vLLM and SGLang for servers. Together AI, Fireworks AI and OpenRouter offer it as a hosted service.

How did Muse Glimmer do in our agency test?

We built a test around three jobs an agency does every week, for two invented clients: a fitted-kitchen retailer and an accountancy firm. Each client had its own written policies, so the models had to apply rules rather than general knowledge.

  • Search-term triage: 300 search terms, each needing an intent label and an action (keep, exclude or review) under the client’s keyword policy.
  • Brand-rule checks: 40 ad and landing-page drafts, each to be marked compliant or not, with every broken rule named and quoted.
  • Screenshot audits: 20 landing-page screenshots, each checked against an eight-point checklist covering the headline, call-to-action button, phone number, price, trust signals, form, cookie notice and service area.

We compared Glimmer with Qwen3.8-27B, the open-weight model closest to it in size and date. Both ran with reasoning set to medium and the sampling settings their makers recommend, three times over, through OpenRouter pinned to the same provider (DeepInfra), with the provider’s data collection switched off. Every answer was scored automatically against a fixed answer key, labelled twice independently with full agreement.

AIWIZ agency test: Muse Glimmer 30B and Qwen3.8-27B (432 requests, 4 October 2026)

If the table extends beyond the screen, scroll sideways to view all columns.

Our test (4 October 2026) Muse Glimmer 30B Qwen3.8-27B
Search terms right, intent and action (300 terms, 3 runs) 99.0% 99.2%
Brand-rule drafts exactly right (40 drafts, 3 runs) 100% 100%
Screenshot checklist items right (160 items, 3 runs) 100% 100%
Median time per reply 6.4 seconds 35.8 seconds
Slowest 5% of replies took at least 28.9 seconds 113.4 seconds
Cost of the 216 requests $0.44 $0.80

Source: AIWIZ test, 432 requests run on 4 October 2026 through OpenRouter (DeepInfra). Costs are what OpenRouter billed, in US dollars.

On accuracy the two were level. Across three runs Glimmer got 9 of 900 search-term answers wrong and Qwen 7, a gap well inside the margin of error. Both got every brand-rule draft and every screenshot checklist right. Every reply was valid, in the right format and for the right item. Neither model had a failed request.

The difference was speed and cost. Glimmer’s median reply came back in 6.4 seconds against Qwen’s 35.8, even though both used almost the same number of reasoning tokens.

What our test does not show

Our tasks were within both models’ reach, so they hit the ceiling and can’t be ranked on accuracy. The data was synthetic and the clients invented. A test of harder work would separate them, and the independent benchmarks below suggest Qwen would come out ahead. We also dropped a third contender, Gemma 4 31B, after its hosting provider rate-limited a third of its requests. NVIDIA’s Nemotron 3 Nano Omni had no route that met our data policy.

How we ran it: a controlled test on synthetic data, 216 requests per model over three runs on 4 October 2026, scored automatically against a fixed answer key.

How does Glimmer compare on independent benchmarks?

Artificial Analysis runs the same tests on every model, which makes it the fairest outside comparison. On its Intelligence Index v4.3.2, Qwen3.8-27B scores 34 and Glimmer 17.

Artificial Analysis Intelligence Index v4.3.2: Muse Glimmer (high) and Qwen3.8-27B (xhigh), checked 4 October 2026

If the table extends beyond the screen, scroll sideways to view all columns.

Artificial Analysis (Intelligence Index v4.3.2) Muse Glimmer (high) Qwen3.8-27B (xhigh)
Intelligence Index 17 34
GDPval-AA v2.1 (office tasks, Elo) 791 1423
AutomationBench-AA (workflow automation) 7% 48%
Terminal-Bench 4.0 1% 6%
Humanity’s Last Exam 22% 34%
SciCode 45% 47%
AA-LCR v1.1 (long documents) 83% 82%
Output speed 124 tokens a second 46 tokens a second
Output tokens per task 14k 67k

Source: Artificial Analysis model comparison, checked 4 October 2026. Artificial Analysis tested Glimmer at high reasoning and Qwen at xhigh.

Qwen leads clearly on the agentic and office-work tests, which are closest to handing a model a real multi-step job. The two are level on long-document reading and scientific coding. Glimmer’s advantage is speed, and on Artificial Analysis’s numbers it uses about a fifth of the output tokens per task.

If you have seen a Glimmer score of 35 from Artificial Analysis, that comes from its launch article of 10 August 2026, written on an earlier version of the index. Version 4.3, announced on 7 September, upgraded Terminal-Bench to 4.0 and added AutomationBench-AA, so scores from before and after can’t be compared. The same launch article recorded an 82% hallucination rate for Glimmer on its AA-Omniscience knowledge test, against 49% for Qwen3.6 27B. For research or drafting, that matters more than any ranking. Check every fact Glimmer gives you.

The makers’ own tables point the same way. Meta’s compares Glimmer with Gemma4-31B and Qwen3.6-27B. Glimmer’s clearest lead there is tool orchestration, 75.5 on MCP Atlas against Qwen3.6’s 62.5. Alibaba’s card for Qwen3.8-27B, released four days after Glimmer under Apache 2.0, puts the newer Qwen ahead of Glimmer on SWE-bench Pro coding (61.7 against 51.2), OSWorld-Verified computer use (84.3 against 65.9) and Terminal Bench 2.1 (73.0 against 51.7). The two are closest on IFBench instruction following, 79.5 against 77.0. Those are each vendor’s own numbers.

Why our results and the benchmarks differ

The benchmarks measure long, open-ended work: completing a client brief, automating a workflow, operating a terminal. Our test measured short, closed tasks with written rules and one right answer. That is a large share of agency operations, and on it Glimmer was as accurate as a model that scores twice as high on harder work. Choose by the job: Qwen3.8-27B for multi-step agentic work, Glimmer for high-volume checking where speed and cost count.

How is Muse Glimmer different from Muse Spark, Muse Image and Ads Manager?

These four get confused constantly, and the confusion leads to bad advice.

Muse Glimmer, Muse Spark, Muse Image and Ads Manager compared

If the table extends beyond the screen, scroll sideways to view all columns.

Question Muse Glimmer Muse Spark (1.3) Muse Image Ads Manager / Advantage+
What it is Open-weight model for local agents Meta’s flagship hosted model Image generation and editing Meta’s advertising tools
How you get it Download the weights, or use a hosted provider Muse Code and the Meta Model API Meta Model API and Meta AI apps Your Meta ads account
Terms Apache 2.0 on the weights Meta’s API and product terms Meta’s API and product terms Meta’s advertising policies
What you pay for Your hardware and time, or a provider’s fees Usage, on Standard or Contributor tiers Usage Media spend
Open weights? Yes, since 10 August 2026 Promised, not shipped No Not applicable
Affects your ads? No Not directly Meta said in July 2026 it would reach advertisers through Advantage+ creative; check your own account Yes

Meta’s models overview puts it plainly: unlike Spark, Muse Image, Muse Voice Transcribe and SAM, you don’t call Glimmer over the Model API. You download the weights and serve it yourself.

What happened to open-weight Spark?

Because Glimmer learned from Spark, it is easy to assume Spark is open too. It isn’t. Distillation passes on some of the teacher’s ability, not its weights.

Meta has also moved the goalposts. On launch day Mark Zuckerberg wrote on X: “Soon we’ll also release the weights for Muse Spark 1.2”. Meta then released Spark 1.3 on 2 September as a hosted model only. Its launch post now promises “the Muse Spark open weights release” with no version or date, and ends: “Stay tuned.” On 4 October 2026 Meta’s Hugging Face organisation still listed only Glimmer repositories. Until Spark weights appear there, plan as if they won’t.

Is Muse Glimmer open source, and can UK teams use it commercially?

The weights are released under Apache 2.0, an OSI-approved open-source licence, and Meta describes the release as open source. The training data and training code are not public, though. Meta’s card says only that the data includes publicly available material, data from third parties and information from Meta’s own products and services. Clients with strong views on where training data came from should know that before choosing it. When you brief a client, “open weights under Apache 2.0” is the phrase that won’t come back to bite you.

For commercial work, four things matter:

  1. The Apache License 2.0 itself, including its patent clause.
  2. Meta’s Muse Glimmer Usage Policy. The model card says Glimmer is intended for commercial and research use, but the policy lists prohibited uses and says the model is not intended for anyone under 18. Meta publishes it alongside the licence rather than inside it, so ask your legal adviser how to treat it for client work. Several clauses land squarely on marketing. The policy bars spam, “false online engagement, including fake reviews”, and inferring private or sensitive details about people, such as health or demographics, without the right to do so. It also prohibits “representing that the use of Muse Glimmer or outputs are human generated”. If you plan to use Glimmer for copy that goes out under a client’s or a person’s name, get legal advice on what that clause means for you first.
  3. Who carries the risk. Apache 2.0 provides the model “as is”, without warranties, and we found no indemnity from Meta. Liability for what Glimmer produces sits with you and whatever your client contract says.
  4. Your own data obligations. Downloading a model doesn’t change your UK GDPR position. The ICO says that “in the vast majority of cases” using AI with personal data will require a data protection impact assessment, and self-hosting doesn’t remove that.

We found no Meta restriction on downloading or running Glimmer in the UK on the pages we checked. Sector rules and client contracts still apply. If you build a customer-facing assistant that people in the EU use, the EU AI Act’s transparency rules, which apply from 2 August 2026, can reach a UK business too: people must be told they are talking to an AI.

What should UK marketers take from Glimmer?

You don’t need to run a 30B model to get something out of understanding it. The immediate value is being able to answer clients who have heard “Muse” used for several different things since Meta Connect. Glimmer is the one that runs on your own machine. Spark is the hosted model. The consumer Muse app is something else again: Meta says it is “powered by Muse Spark”, and it launched on 8 September in the US, then reached Canada and Mexico. There is still no UK date. Our Meta Muse overview keeps track of where it is available.

Two things Glimmer will not do. It won’t change your Meta ads, because it has no link to Advantage+ or Muse Image. It won’t make ChatGPT or Google’s AI Overviews mention your brand either. That is a separate job, which our GEO guide covers.

Where it fits in agency work

Our test covered three of these, and Glimmer handled them as well as Qwen3.8-27B in a fraction of the time. The others are still ideas to test.

  • Search-term and keyword triage. Write the client’s keyword policy down and have Glimmer label each term and suggest keep, exclude or review. In our test it got 99% right. A person still approves the negative keyword list before it goes into the account.
  • Brand and compliance checks. Give Glimmer the client’s rules and have it flag and quote every breach in a draft. It got all 40 of our drafts right, but they were written for the test, and real copy is messier. Treat its flags as a sorting aid, not a sign-off.
  • Screenshot and layout checks. Glimmer reads screenshots well enough to run a landing-page checklist. Write the checklist precisely, because a model reads it literally: decide edge cases such as prices shown with “+ VAT” before you judge the results.
  • Research on material that can’t leave the building. Load a fixed set of approved documents into a self-hosted Glimmer and ask for a structured brief with quotations. Given the hallucination figures, someone senior checks every claim. This earns its place only when a client forbids sending material to an outside AI service.
  • Drafting to house style. Keep your style guide on disk and ask for variants. Judge them as in our Grok 4.7 and Claude Opus 5.5 guides: same brief, count what you accept, count the rework.

If a client mainly wants an assistant that organises files and writes reports, a hosted product will get them there faster. Glimmer makes sense when the requirement is local weights, Apache terms, or a lot of quick, rule-based checking.

What should implementers check before using it for client work?

  • Where the data goes. Self-hosting keeps prompts and files on your own machines, and there is no Meta data processing agreement to look for, because Meta never sees them. A hosted provider is a different arrangement. Get its processing location, retention period, subprocessors and training policy in writing before any personal data goes in.
  • Prompt injection. Meta’s own card reports a 28.4% attack success rate on the AgentDojo prompt-injection test, against 25.6% for Gemma4-31B and 40.3% for Qwen3.6-27B (lower is better). It recommends guardrails and human confirmation before irreversible actions. An agent with shell, browser or file access needs a sandbox and an approval step for anything that sends, spends, publishes or changes production systems. With no vendor watching, logging and review are entirely your job. For context, Meta’s card says Glimmer falls below the “Frontier AI” threshold in its own safety framework because it is less capable than Muse Spark.
  • Hardware. Meta offers two 4-bit builds. The smaller needs about 20 GB with image input and the drafter, for 24 GB cards, and Meta estimates it loses about 1% quality. The larger needs about 23 GB, for 32 GB cards, and loses about 0.2%. On Meta’s own tests the drafter lifts output from 74.9 to 233.4 tokens a second on an RTX 5090, but only from 26.6 to 50.2 on an M5 Max Mac. Those are Meta’s measurements, so test on your own hardware. Our test used a hosted provider, so its speed figures don’t tell you how fast Glimmer runs on your machine.
  • Software versions. Use the official meta-models repositories, not mirrors: at least one unofficial quantisation shipped with broken tool calling until it was rebuilt. Meta’s docs require transformers 5.15 or later, and llama.cpp build b10353 or newer, so update Ollama, LM Studio or llama.cpp before you start. Run your quality and security tests on the exact file and runtime you will deploy, not the full-precision model.
  • Settings. The model card sets reasoning strength (low, medium, high or xhigh) in the system prompt and lists sampling defaults of temperature 1.0, top_p 0.95 and top_k 64. We used those defaults with medium reasoning. If your framework overrides them, check the effect on quality.
  • Fine-tuning. Meta’s docs cover fine-tuning and reinforcement learning, and Unsloth says Glimmer can be fine-tuned on a 24 GB card. Any fine-tune invalidates your earlier test results, so run them again.
  • What it can’t do. There is no audio input or output. Video is handled as individual frames, so don’t promise a voice agent.

How should you trial Muse Glimmer?

Start with one question you can answer yes or no, such as “Can Glimmer sort our weekly search-term report to our policy, without the data leaving our network?” Without that, a trial tends to produce demos rather than decisions.

Then pick one route:

Self-host or use a hosted provider?

If the table extends beyond the screen, scroll sideways to view all columns.

Route Choose it when Watch for
Self-host Data must stay on your machines, or you want full control of the set-up GPU or Mac memory, maintenance, access control
Hosted provider You want to test quickly through an API Your data goes to the provider; its pricing and rate limits apply

If you self-host, start from Meta’s get-the-model page and choose the GGUF build that matches your GPU memory. Run the same few tasks at two quantisation levels and compare the results before settling.

A two-week trial is enough to learn something:

  • Week one. Pick tasks you already know the right answers to, and write the rules down as precisely as you would for a new member of staff. Run them on Glimmer and on the model you use today, three times each, and score the answers against your key. Check a sample of the wrong answers by hand: some will be your instructions, not the model.
  • Week two. Give Glimmer one tool, such as web fetch or a read-only internal API, inside a sandbox. Re-run the same tasks, and log tool and injection failures separately from plain wrong answers.

Throughout, record whether each task was finished, how long each reply took, how many tool calls failed, how long the rework took, and whether the sources checked out. Keep Glimmer only if it matches your current model on accepted work at a total cost, hardware and people included, that you can justify.

If you already run Muse Spark through Muse Code or the Model API, treat Glimmer as a parallel experiment, not a replacement. Meta’s models page lists video and PDF input for Spark and a 1,048,576-token context window, against Glimmer’s 128K default. Our Muse Spark 1.3 guide for UK builders covers its pricing, the Contributor tier and Muse Code.

Is Muse Glimmer worth trying?

Yes, for the right job. In our test it handled routine, rule-based agency checks as accurately as Qwen3.8-27B, more than five times faster and for a little over half the cost. If your team spends hours on search-term reports, brand checks or landing-page audits, Glimmer is worth a trial, especially where the data must stay on your own machines.

It is not the model for open-ended work. The independent benchmarks put Qwen3.8-27B well ahead on agentic and office tasks, and the launch hallucination figures are too high for unsupervised research or content. For everyday drafting, hosted models such as Claude Opus 5.5 or Grok 4.7 remain the fairer comparison. Meta deserves credit for choosing Apache 2.0 over a licence of its own, which makes Glimmer easy to adopt for the narrower job it does well.

If you want help deciding between local and hosted models, or running a fair trial for a client, that is part of our digital marketing and AI adoption work. Get in touch if a second opinion would help.

Frequently asked questions

Is Muse Glimmer open source?
The weights are. Meta released them on Hugging Face under the Apache License 2.0, which is an OSI-approved licence, and Meta itself calls the release open source. The training data and training code are not public, so “open weights under Apache 2.0” is the more accurate description.
How did Muse Glimmer do in AIWIZ’s test?
On three routine agency tasks (sorting search terms, checking ad copy against brand rules and auditing landing-page screenshots) it matched Qwen3.8-27B: 99.0% against 99.2% on search terms, and every brand-rule draft and screenshot checklist right for both. Its median reply took 6.4 seconds against Qwen’s 35.8, and it cost $0.44 against $0.80 for the same 216 requests. The test used synthetic data for two invented clients, so it shows what the models can do on this kind of work, not on your accounts.
Is Muse Glimmer better than Qwen3.8-27B?
Not overall. On Artificial Analysis’s independent Intelligence Index (v4.3.2, checked 4 October 2026) Qwen3.8-27B scores 34 and Glimmer 17, and Qwen leads clearly on agentic and office-work tests. Glimmer is much faster and uses far fewer tokens. In our own test of routine agency tasks the two were level on accuracy. Choose Qwen for hard multi-step work and test Glimmer where speed and cost matter more.
Has Meta released Muse Spark as open weights?
Not yet. On 10 August 2026 Mark Zuckerberg said Meta would “soon” release the Muse Spark 1.2 weights. As of 4 October 2026 none have shipped: Meta’s Hugging Face organisation lists only Glimmer repositories, and Meta’s Spark 1.3 post lists the release as future roadmap, with no version or date.
Can UK teams download and use Muse Glimmer commercially?
Yes. Meta’s docs say the weights download from Hugging Face with no API key and no gated access, and the model card says Glimmer is intended for commercial and research use. We found no UK restriction on the Meta pages we checked. Read Meta’s Usage Policy alongside the Apache licence, and treat client personal data exactly as you would with any other system.
What hardware does Muse Glimmer need?
Meta’s smallest official build needs about 20 GB with image input and the speed-up drafter, so it targets a 24 GB graphics card or a Mac with enough unified memory. Meta estimates it loses about 1% quality against full precision. Most office laptops have nothing like that.
Does Muse Glimmer generate images or change Meta ads?
No. Glimmer reads text and images and writes text. Muse Image is Meta’s picture model, and Glimmer has no connection to Ads Manager or Advantage+.
Is the Muse app the same as Muse Glimmer?
No. Meta says the consumer Muse agent is powered by Muse Spark, a hosted model. Meta launched it in the US on 8 September 2026, has since listed it in Canada and Mexico, and has given no UK date.
Is Muse Glimmer the same as Llama?
No. Glimmer belongs to Meta’s Muse line and ships under Apache 2.0. Llama used Meta’s own community licence, which for Llama 4’s multimodal models excluded companies based in the EU. Glimmer’s licence and Usage Policy name no region. Don’t assume the terms of one apply to the other.
Should we self-host Glimmer or use a hosted provider?
Self-host if the point is keeping data on your own machines. Meta’s model page lists hosted routes through Together AI, Fireworks AI and OpenRouter, which are quicker to start with but send your prompts to that provider under its own terms, pricing and data handling. We ran our test through OpenRouter, pinned to one provider with data collection switched off.
Explore
Drag