Ground Truth.
AI, checked against the source.

News · 2026-09-02

Meta's Muse Spark 1.3 caught GPT-5.6 on one scoreboard and still trails Claude

Meta released Muse Spark 1.3 on September 2, 2026, and the independent scoring puts it level with OpenAI's flagship on one index while still four points behind Anthropic's. Artificial Analysis scores the publicly available Muse Spark 1.3 xhigh tier at 61 on its Intelligence Index and the limited-preview max variant at 62, against 66 for Claude Fable 5.1. That is a genuine comeback for a lab that spent 2025 being written off, and it is not the frontier-topping result the launch-day commentary described.

Key facts

Meta's own framing is about endurance rather than raw score. The research post describes a model built to hold a single long-running thread of work, pull a coherent picture out of messy and conflicting sources, ask clarifying questions instead of guessing, confirm before taking consequential actions, handle being interrupted mid-task, and know more accurately what it cannot do. That last item is quietly the most useful. A model that stops and asks is worth more in an agent loop than a model that scores two points higher and confidently does the wrong thing for forty minutes.

The product page fills in the consumer side. Meta describes Muse Spark as natively multimodal, reading images, charts, and text together rather than through a bolted-on vision encoder, and it ships a setting called Contemplating mode in which, in Meta's words, "multiple agents reason in parallel before answering, reaching deeper and more reliable results on complex problems." The same model drives image generation, website and mini-game creation, recommendations across Instagram, Facebook, and Threads, and the real-time visual understanding in Meta's AI glasses. It is a research result and a consumer feature pipeline at once, which is a different bet from the one OpenAI and Anthropic are making.

Artificial Analysis locates the gains precisely, and the location is the story. The biggest lifts are on agentic and scientific work, with the largest movements on evaluations that measure economically valuable task completion, multi-turn banking workflows, and terminal-based engineering tasks. Coding and science improved by smaller margins, and the model became slightly more cautious about producing confident wrong answers. That is a profile of a model tuned for work, not for exam scores, which is consistent with what Meta says it built.

Two things circulating about this release do not survive checking. The first is that Muse Spark 1.3 beat Claude. Artificial Analysis's direct comparison page shows Fable 5.1 at 66 and Muse Spark's best variant at 62. The second is a widely repeated $0.10 and $0.20 per-million "contributor tier" price, which would make this the cheapest frontier model by an order of magnitude. That figure does not appear on any first-party pricing page that could be retrieved. The verified public price is $1.25 and $4.25, which Artificial Analysis notes is unchanged from version 1.2.

The open-weights question is where Meta's reputation is actually on the line. The company that made open weights into a strategy has now shipped three hosted proprietary Muse Spark releases in a row. The 1.3 post says the roadmap includes "the Muse Spark open weights release." The August 1.2 post said the same thing, describing that release as coming ahead of the open-weights one. Two posts, one commitment, no date. It is a real commitment and it is worth tracking, but it is not a shipped artifact, and there is no download size to quote because there is nothing to download.

The community reaction reflects that gap. The Hacker News thread contains genuine enthusiasm for the model's speed and for a discounted access tier that hobbyists found compelling, including a tester reporting better results on a generation task than 1.2 produced. It also contains a substantial group of commenters who said plainly that they would rather pay a competitor more than route their work through Meta, citing surveillance and privacy concerns. That is the counterweight, and no benchmark score addresses it. Meta has built a model people respect and a brand a meaningful slice of developers will not touch, and 1.3 does not change the second half of that sentence.


Primary source, verified: read the paper →

Key questions

Did Muse Spark 1.3 beat Claude Fable 5.1?

No. Artificial Analysis scores Claude Fable 5.1 at 66 on its Intelligence Index versus 62 for the limited-preview Muse Spark 1.3 max variant, so Meta's model trails.

Are Muse Spark's weights available to download?

Not yet. Meta's own post says the roadmap includes a Muse Spark open weights release but gives no date, and the shipping product is a hosted proprietary model.

What is Contemplating mode?

It is Meta's name for a setting in which multiple agents reason in parallel before the model answers, which Meta says produces deeper results on complex science and research questions.
Cite this

APA

Ground Truth. (2026, September 2). Meta's Muse Spark 1.3 caught GPT-5.6 on one scoreboard and still trails Claude. Ground Truth. https://groundtruth.day/news/meta-muse-spark-1-3-ties-one-index-and-trails-another.html

BibTeX

@misc{groundtruth:meta-muse-spark-1-3-ties-one-index-and-trails-another,
  title  = {Meta's Muse Spark 1.3 caught GPT-5.6 on one scoreboard and still trails Claude},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {sep},
  url    = {https://groundtruth.day/news/meta-muse-spark-1-3-ties-one-index-and-trails-another.html}
}

Topics: models · meta · muse-spark · agents · multimodal · benchmarks

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.