Analysis
Analysis

Muse Spark 1.3 Reality Check: 75.4 DeepSWE Uses Max, xhigh Wins Terminal-Bench 89.2, and AA Shifted 61→45

Published Sep 8, 2026 Sources checked Sep 8, 2026

Meta's own Muse Spark 1.3 scorecard shows 75.4 DeepSWE on max reasoning but 89.2 Terminal-Bench on xhigh. SWE-bench Verified and Pro remain unreported, while Artificial Analysis's composite moved from 61/62 at launch to 45/48 under v4.3 because the benchmark changed.

What Meta actually released

Meta released Muse Spark 1.3 on September 2, 2026 for long-horizon agentic and coding work. Meta's current launch page says max reasoning is now available in Muse Code and the Meta Model API. The same page says Meta engineers measured roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on their coding comparisons.

Those efficiency numbers are useful, but they are vendor measurements, not a public same-task reproduction. Meta's launch page also says open weights are on the roadmap rather than available today.

Primary sources:

The launch scorecard mixes two reasoning settings, so the setting matters

Meta's scorecard contains separate Muse Spark 1.3 (max) and Muse Spark 1.3 (xhigh) columns. That matters because the two settings do not move in lockstep.

For DeepSWE v1.1, Meta reports 75.4% for max and leaves the xhigh cell blank. Meta says DeepSWE v1.1 contains 113 tasks across 91 repositories and five languages and that it ran Muse Spark 1.3 max with a mini-swe agent. The primary metric is task pass rate: both the functional and regression tests must pass.

For Terminal-Bench 2.1, Meta reports 88.8% for max but 89.2% for xhigh. In other words, the higher reasoning setting does not automatically produce the higher observed score. Meta describes Terminal-Bench 2.1 as 89 terminal-environment tasks run with each model's native coding harness inside Meta's internal agent-evaluation framework and isolated cloud sandboxes, with the official executable verifier grading final container state.

The same max-versus-xhigh split appears elsewhere in the visible scorecard:

  • GDPVal-AA v2: 1754 max vs 1709 xhigh
  • JobBench: 64.9 vs 61.2
  • OSWorld 2.0 partial score: 66.9 vs 59.0
  • DeepSearchQA: 90.3 vs 89.4
  • AutomationBench: 49.6 vs 48.3
  • SWE-Atlas Codebase QnA: 59.4 vs 54.0
  • MRCR 512K–1M: 98.1 vs 93.1

These numbers support the narrower claim that max often helps on Meta's selected evaluations. They do not support a blanket rule that max is always better, because Terminal-Bench 2.1 goes the other way.

There is also a documentation bookkeeping issue worth noticing: the generic "Overall Methodology" text describes a default reasoning configuration, while the scorecard and per-benchmark sections explicitly contain both Muse Spark 1.3 max and xhigh results. For fair comparisons, the benchmark-level row and reasoning setting should take precedence over a generic model-family label.

SWE-bench Verified and SWE-bench Pro: no exact 1.3 score should be invented

Meta's current Muse Spark 1.3 evaluation report does not publish an exact SWE-bench Verified result for Muse Spark 1.3. It also does not publish an exact SWE-bench Pro result.

That means the 75.4 DeepSWE v1.1 score must not be relabeled as SWE-bench Verified, and it cannot be compared numerically with a competitor's SWE-bench Pro percentage as though they were the same test.

This is a recurring benchmark-hygiene problem in third-party model tables: DeepSWE, SWE-bench Verified and SWE-bench Pro all test software engineering, but they differ in task construction, repositories, harnesses, grading and evaluation history. The defensible status for Muse Spark 1.3 today is:

  • DeepSWE v1.1: 75.4%, Meta-run, max reasoning, 113 tasks, mini-swe agent
  • SWE-bench Verified: no exact 1.3 result accepted here
  • SWE-bench Pro: no exact 1.3 result accepted here

If a later independent run appears, it should be pinned to the exact model endpoint, reasoning setting, suite revision, task count, harness, retry policy and date before being compared.

Artificial Analysis changed the measuring stick, so 61→45 is not a model collapse

Artificial Analysis's September 2 launch article placed Muse Spark 1.3 at 61 for xhigh and 62 for max on the Intelligence Index available at launch. It also reported $0.55 per Index task for xhigh, with $1.25/M input, $4.25/M output and $0.15/M cached input.

Artificial Analysis has since moved to Intelligence Index v4.3, which uses a different ten-evaluation composition including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.

On the current v4.3 model pages, the same named model now reads:

  • Muse Spark 1.3 xhigh: 45
  • Muse Spark 1.3 max: 48

That apparent 61→45 or 62→48 drop should not be described as the model suddenly losing capability. The benchmark version and composition changed.

Current Artificial Analysis operational measurements are also useful:

  • xhigh: about 188.8 output tokens/s, 35.27 s TTFT, $1.37 per v4.3 Index task, and 140M output tokens across the Index
  • max: about 229.5 output tokens/s, $1.60 per v4.3 Index task, and 170M output tokens across the Index
  • both current pages list $1.25/M input, $4.25/M output, an 88% cache discount and a 1M-token context window

These are point-in-time API measurements from Artificial Analysis. They are not service-level guarantees, and the v4.3 composite is not a SWE-bench score.

Sources:

Access changed after launch coverage

Early September 2 coverage described max as limited or pending additional safety testing. Meta's current first-party launch page now says Muse Spark 1.3 with max reasoning is available on Muse Code and the Meta Model API.

That makes dated availability important. A launch-day article can be correct for September 2 and stale a few days later.

Meta's developer landing page currently requires login for detailed account access, so this article does not invent region-specific quotas, subscription allowances or a public max-token output limit that could not be verified from an accessible first-party page.

Public reaction: useful anecdotes, not a benchmark

Mark Zuckerberg's September 2, 2026, 19:26 UTC X post called the release "frontier performance almost too cheap to meter" and described it as Meta's biggest coding-and-agentic jump so far. That post is primary evidence for Meta's launch positioning, not independent validation.

Community feedback is mixed. A September 6 r/opencode thread includes a heavy Claude Code user saying the free Muse Spark 1.3 plus OpenCode harness felt surprisingly good, while another commenter argued the benchmarks looked inflated and that longer programming use exposed weaknesses. A separate September 3 r/opencode thread discusses uncertainty around Muse Code subscription usage and routing.

Those are self-selected practitioner anecdotes from small discussions. They are useful for identifying questions worth testing—long-run reliability, rate limits, harness quality and real cost—but they are not a controlled evaluation and do not establish consensus.

Sources:

Practical verdict

Muse Spark 1.3 has credible evidence of a substantial Meta-side improvement, especially on long-horizon coding and long-context retrieval. But the exact reasoning setting and harness matter more than the headline.

The strongest fair reading today is:

Muse Spark 1.3 max has a Meta-run 75.4% DeepSWE v1.1 result, while xhigh actually posts the higher Terminal-Bench 2.1 score in Meta's own table at 89.2%. Neither SWE-bench Verified nor SWE-bench Pro has an exact 1.3 result accepted here. Artificial Analysis's headline score changed from 61/62 at launch to 45/48 under v4.3 because the benchmark changed, not because a directly comparable rerun proved a collapse.

For teams evaluating the model, the next useful evidence is not another mixed leaderboard. It is a same-endpoint, same-harness comparison of xhigh versus max on a version-pinned coding suite, plus an independent SWE-bench Verified or Pro run with the full task count, retries, trajectories, cost and latency published.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books