Meta Muse Spark 1.3 Max Is Now Available: Benchmarks, Safety and Early Signals
Meta’s Muse Spark 1.3 launch page now says max reasoning is available. Here is what changed, what independent benchmarks show, and what remains unverified.
What changed
Meta’s September 2 launch page for Muse Spark 1.3 has been updated: at verification time on September 5 UTC, the official page states that Muse Spark 1.3 with max reasoning is now available in Muse Code and the Meta Model API. The earlier launch language said max reasoning would arrive only after additional safety testing, so this is a meaningful availability change rather than a new model family.
“Max” should be treated as a higher reasoning-effort configuration of Muse Spark 1.3, not as a separately named base model. Meta’s page also says the model is aimed at longer-horizon agentic and coding work, including multi-step workflows, tool use and instruction following.
Source: Meta AI Research — Introducing Muse Spark 1.3.
Meta’s efficiency and safety claims are vendor claims
Meta says its engineers found Muse Spark 1.3 more efficient than Muse Spark 1.2 in coding workflows, using about 20% fewer tool calls and about 25% fewer tokens in their comparisons. Those figures are useful, but they are internal measurements rather than an independent cross-provider benchmark.
Meta also says the release improves adversarial robustness, prompt-injection resistance and calibration around irreversible actions. The company’s launch page does not publish enough independent evidence to turn those safety statements into a universal ranking against other frontier models. They should therefore be read as vendor-reported improvements.
Independent benchmark evidence: max gains one point on the Artificial Analysis index
Artificial Analysis reports 62 for Muse Spark 1.3 max and 61 for Muse Spark 1.3 xhigh on its Intelligence Index. A one-point aggregate difference is not proof that max will be better for every workload. The evaluator says the gains are concentrated mainly in agentic work and scientific capabilities, and individual benchmarks can move differently.
The xhigh configuration has substantially more independent operating data. Artificial Analysis currently lists xhigh at $1.25 per million input tokens and $4.25 per million output tokens, with an 88% cache discount and about $0.55 per Intelligence Index task. It also lists a 1-million-token context window and multimodal text, image and video input. The evaluator’s speed and latency measurements are live metrics and can change as infrastructure changes.
For max, the same evaluator publishes the 62 Intelligence Index score but does not yet show comparable public speed or cost-per-task measurements. It would be wrong to assume that max has the same latency or effective cost as xhigh until measured.
Sources: Artificial Analysis release comparison and Muse Spark 1.3 analysis.
Benchmark methodology matters more than the headline table
Meta’s evaluation methodology is unusually useful because it documents task counts and harness choices. It says Muse Spark 1.3, Claude Opus 5 and GPT-5.6 Sol are generally evaluated at max reasoning effort, while Muse Spark 1.2 uses xhigh. It also warns that third-party models are run with best-effort common settings that may not reproduce each vendor’s optimized environment.
The report describes several distinct benchmark families:
- GDPVal-AA v2: 220 professional tasks spanning 44 occupations and nine U.S. industries, run in Artificial Analysis’s Stirrup harness with shell and web access.
- OSWorld 2.0: 108 long-horizon desktop workflows. Meta notes that most models use the 08.08 release while Muse Spark 1.2 used 06.24, so that older comparison is not perfectly version-matched.
- DeepSearchQA: 900 browsing questions using the same search backend and browser harness across compared models.
- AutomationBench: 600 business workflows with deterministic end-state checks rather than an LLM judge.
- DeepSWE v1.1: 113 software-engineering tasks across 91 repositories and five languages. Meta says Muse Spark 1.3 max was run with mini-swe-agent, while other model results were taken from the official Datacurve leaderboard.
- MRCR v2: 100 long-context examples in each of the 256K–512K and 512K–1M bands.
These are not one interchangeable score. A model can lead in a browser or coding harness without being the best choice for professional-document work, long-context retrieval or cost-sensitive production agents.
Source: Meta Muse Spark 1.3 evaluation methodology.
What the first public integration report can—and cannot—tell us
One public GitHub issue filed on September 4 reports intermittent “Invalid upload request” failures when using the muse-spark-1.3-contributor-free route through OpenCode, while another free model worked in the same setup. The reporter supplied timestamps and reproduction steps.
That is useful operational feedback, but it is one integration report, not evidence that the Meta Model API or Muse Spark 1.3 itself is broadly unreliable. The failure may sit in the free-tier routing or provider integration. It should be treated as an anecdotal compatibility signal until maintainers or the provider identify the root cause.
Source: OpenCode issue #47237.
X feedback remains unscored
A public X post associated with the max rollout was discoverable by URL, but direct retrieval returned an access error during this review. Because the primary post could not be independently read in the verification environment, this article does not quote it or use it as evidence for sentiment. No broad X consensus is claimed.
That matters because launch-day social feeds are heavily selection-biased: people who received access early, hit an error, or saw unusually strong results are more likely to post than typical users. Measured repeated-task results are stronger evidence than anecdotes.
Practical tradeoffs: when max may be worth it
Max is most interesting when a task is expensive to fail: long coding changes, research agents, multi-application workflows, or complex instructions where extra reasoning can reduce retries. But the correct production metric is cost per successful task, not just tokens per request or one aggregate benchmark.
Teams evaluating Muse Spark 1.3 should compare max and xhigh on the same representative task set, with the same tool permissions and stop conditions. Track success rate, wall-clock latency, tool-call count, total tokens, cache hit rate, retries and human correction time. Until independent max latency and cost measurements are published, xhigh has the clearer operating profile.
Confidence and what to watch next
Confidence is high that Muse Spark 1.3 max is now available in Muse Code and Meta Model API because Meta’s own launch page says so. Confidence is medium on cross-model performance conclusions because benchmark harnesses, reasoning settings and result provenance differ. Confidence is low on community-wide sentiment because accessible public feedback is still sparse and self-selected.
The next useful evidence is a stable independent measurement of max pricing, latency and cost-per-success, plus same-harness coding and computer-use evaluations against other September frontier releases. Any future comparison should keep vendor claims, third-party measurements and anecdotal user reports visibly separate.
This article is built from the source material below. Open the originals for full context and the latest updates.