Meta Muse Spark: What the Benchmarks Actually Mean, Where It Falls Short, and Who Should Pay Attention
In this article
The organizational context you need to understand first
What Muse Spark actually is
The benchmark picture, unvarnished
(Artificial Analysis)
(vs 157M for Claude Opus)
(beats GPT-5.4 at 82.8)
(leads all models)
| Benchmark | Muse Spark | GPT-5.4 | Claude Opus 4.6 | Gemini 3.1 Pro |
|---|---|---|---|---|
| AI Intelligence Index | 52 | ~57–58 | ~57–58 | ~54–55 |
| Output tokens (Index run) | 58M Most efficient | 120M | 157M | 57M |
| MMMU-Pro (multimodal) | 80.5% | ~78–79% | ~77–78% | 82.4% Leads |
| CharXiv visual reasoning | 86.4 Leads | 82.8 | ~80 | 80.2 |
| HealthBench Hard | 42.8 Leads | High 30s–low 40s | Similar band | Slightly lower |
| GDPval-AA (agentic) | 1427 | 1676 Leads | 1648 | 1320 |
| TerminalBench Hard (coding) | Below leaders | 75.1 | 80.8% SWE-bench | 68.5 |
| τ²-Bench Telecom | 92% Top tier | — | — | — |
| CritPT (hard physics) | 11% Above Claude, Gemini Flash | — | 3% | 9% |
Muse Spark is the second-most capable vision model we have benchmarked. Agentic performance does not stand out, it scores 1427 on GDPval-AA, behind Claude Sonnet 4.6 and GPT-5.4, but ahead of Gemini 3.1 Pro Preview at 1320.
Where Muse Spark genuinely leads
Visual reasoning and multimodal understanding
Health reasoning
Token efficiency
Domain-specific reasoning
Where it falls short, and why that matters
Coding and software engineering
Agentic and multi-step work
Closed-source means lock-in
Decision framework: who should actually use this
Your workloads are vision-heavy or health-adjacent
- Parsing charts, figures, scientific diagrams
- Health Q&A at scale (with appropriate guardrails)
- Document intelligence on mixed text-image content
- Cost-sensitive high-volume reasoning inference
- Deep integration with Meta’s social surfaces
Coding quality and agentic execution are the priority
- Software engineering copilots and code review
- Long-running multi-step agent pipelines
- Enterprise stacks needing mature governance tooling
- Open-source flexibility and fine-tuning requirements
- Mission-critical agentic workflow execution
Google Workspace integration and search grounding matter
- Tight integration with Google Cloud or Workspace
- Top-tier MMMU-Pro multimodal score (82.4%)
- Factual grounding through Google Search
- Token efficiency matching Muse Spark’s profile
Strategic implications for different stakeholders
For ML engineers and developers
For CTOs and CIOs
For VCs and investors
For policy makers and regulators
How to access Muse Spark today
Frequently asked questions
The bottom line
Sources & further reading
- Meta AI BlogIntroducing Muse Spark: Scaling Towards Personal Superintelligence
- Meta NewsroomIntroducing Muse Spark: Meta’s Most Powerful Model Yet
- Artificial AnalysisMuse Spark: Everything you need to know, benchmark deep dive
- Artificial Analysis XBenchmark highlights thread with token efficiency data
- LushBinaryMeta Muse Spark: Benchmarks, Modes & Developer Guide
- LushBinaryMuse Spark vs GPT-5.4 vs Claude vs Gemini, comparison
- TechCrunchMeta debuts the Muse Spark model in a ‘ground-up overhaul’ of its AI
- The VergeMeta is reentering the AI race with a new model called Muse Spark
- CNBCMeta debuts first major AI model since $14.3B Scale AI deal
- New York TimesMeta unveils new AI model, its first from the superintelligence lab
- Silicon RepublicMeta’s Superintelligence Labs debuts first product Muse Spark
More posts
-
US-China Trade Truce Expires in November: What Xi Jinping’s Washington Visit Needs to Deliver
Xi Jinping arrives in Washington on September 23 for a state visit that could shape what happens to the US-China trade truce. About eight hours of talks in New York produced an AI dialogue and an operational Board of Trade, but no word on extending the truce before it expires in November. Here is what…
-
Google’s Gemini Accessed Three Real Companies During a Cyber Test, and It Is the Fourth Lab Tied to the Same Vendor
Google Gemini hacked three companies in May, and the test’s fictional target happened to share a name with a real firm. Google confirmed it on Sept. 18, making it the fourth major AI lab tied to the same testing vendor. Here is what happened, why the labs disagree on what to call it, and what…
-
What Is Trump’s “AI Force”? The Czar Plan, the Slowdown Debate and What Comes Next
Trump says he is creating an “AI Force” and will name an AI czar, but he has not said what either will do. The Trump AI Force announcement lands a week after leading AI figures called for a slowdown, with a UN event and a summit with Xi Jinping days away.
-
US-China Trade Talks in New York: What Bessent and He Lifeng Are Negotiating Before Xi’s State Visit
Treasury Secretary Scott Bessent and Vice Premier He Lifeng are holding trade talks inside a JPMorgan Chase building in Manhattan, days before Xi Jinping’s state visit to Washington. The agenda covers a truce that expires Nov. 10, rare-earth supplies and possible AI guardrails. Here is what is reported to be on the table, and where…
-
Can Disorder Make Networks More Stable? Northwestern Physicists Say Yes
For decades, the safest network was assumed to be the most uniform one. Northwestern physicists now report in Science that disorder-promoted stability is real: a measured dose of variation can steady power grids, neurons and ecological networks, though too much tips them over. So far it has been tested in models, not in the field.
-
Trump Signs Russia and Iran Sanctions Act Named for Lindsey Graham: What the Law Does and What Happens Next
Lindsey Graham did not live to see his Russia sanctions bill become law, and the president who signed it offered no public remarks. The act allows tariffs of up to 100 percent on top buyers of Russian oil and gas, but Trump holds the power to waive it. Here is what it does and what…
-
Slow Down or Speed Up? AI’s Biggest Names Split With the White House as Congress Goes Home
Anthropic’s Dario Amodei asked the AI industry to slow down, and rivals Sam Altman, Elon Musk and Demis Hassabis actually agreed. Then Trump phoned Jensen Huang mid-speech to call it a hoax, and Congress left town before acting on any of it.
-
Asian Games Open in Nagoya After Floods, Housing Complaints and a Diplomatic Standoff
The 20th Asian Games opened Saturday in Nagoya, capping a buildup marked by record floods, cramped athlete housing and a diplomatic standoff over Taiwan’s sports minister. Here’s what happened before the opening ceremony, and what to watch as competition begins.
-
Plugin4Shell Explained: How a Missing Git Check Undercuts Plugin Pinning in Four AI Coding Agents
A team that pins an AI coding plugin to one reviewed Git commit is trying to guarantee that nothing changes underneath it. Researchers at Air Security say that in four of the best-known coding agents, nothing confirms the guarantee held. On Thursday, the company’s research lab published a disclosure it calls Plugin4Shell. It describes a…
