Analysis
Analysis

Meta Muse Spark 1.3 Reality Check: 53 on AA v4.2, 75.4% DeepSWE and a Mixed-Provenance Scorecard

Published Sep 7, 2026 Sources checked Sep 7, 2026

Muse Spark 1.3 posts 75.4% on Meta’s DeepSWE run and strong long-context scores, but its launch AA score of 62 became 53 after v4.2 changed the benchmark. We separate mixed harnesses, current independent measurements and missing SWE-bench results.

What Meta actually released

Meta published Muse Spark 1.3 on September 2, 2026 as an updated version of its Spark family focused on agentic work and coding. The official launch page says the model is designed to sustain longer tasks, ask clarifying questions when requirements are ambiguous, invoke the user when it becomes stuck, confirm consequential actions, and preserve user instructions over long trajectories. Meta’s current page says Muse Spark 1.3 with max reasoning is available through Muse Code and the Meta Model API.

Meta also says its engineers observed about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 while completing the same tasks. That is useful directional evidence, but it is a vendor-internal comparison. The public launch post does not provide a complete sample size, confidence interval, task-by-task distribution, or a separately reproducible harness for those two efficiency percentages. They should therefore be described as Meta’s measured internal result, not as an independent universal efficiency claim.

The Artificial Analysis score did not simply fall from 62 to 53

At launch, Artificial Analysis reported 62 for Muse Spark 1.3 at max reasoning and 61 for xhigh on the then-current Artificial Analysis Intelligence Index. Its current model pages now show 53 for max and 52 for xhigh. Taken without context, that looks like a nine-point regression. It is not evidence that the same deployed snapshot suddenly became much less capable.

Artificial Analysis released Intelligence Index v4.2 on September 4, changing the benchmark itself. The revision added AA-Briefcase and GDP.pdf, removed saturated GPQA Diamond, changed grading, and increased the privately held-out share to 40%. The current v4.2 methodology weights Agents at 30%, Coding at 20%, Scientific Reasoning at 20%, and General at 30%.

The responsible interpretation is therefore:

  • 62/61 are launch-era scores under the preceding Index version;
  • 53/52 are current v4.2 scores;
  • the difference is primarily a change of evaluation version, task mix and grading, so it should not be plotted as if it were a longitudinal capability loss of an unchanged model.

Artificial Analysis says repeated evaluations of selected models suggest the overall v4.2 Index confidence interval is usually within roughly one point, although individual component evaluations can be noisier. That uncertainty estimate still does not make scores from two different Index versions directly interchangeable.

Current independent serving measurements: inexpensive, fast, but route-dependent

Artificial Analysis’ current dedicated max-reasoning page reports approximately 190.1 output tokens per second, 18.69 seconds to first answer token, and about $0.96 per Intelligence Index task for Muse Spark 1.3 max. Its xhigh page reports about 165.4 output tokens per second, 46.13 seconds to first answer token, and approximately $0.84 per Index task.

The same current pages list a 1 million-token context window and standard pricing of $1.25 per million input tokens and $4.25 per million output tokens, with an 88% cached-input discount.

These figures are useful operational measurements, but they are not immutable properties of the model weights. Time to first token and throughput depend on provider capacity, route, batching, request shape, reasoning effort and measurement time. Cost per Index task additionally depends on how many tokens and agent turns that benchmark induces. A developer should therefore treat them as a current serving snapshot rather than a latency guarantee.

Meta’s scorecard is not one uniform head-to-head benchmark suite

Meta’s own four-page Muse Spark 1.3 evaluation report contains a broad table spanning professional tasks, GUI interaction, browsing, business automation, long context, software engineering and terminal work. The most important methodology note appears before the scores: the result provenance varies by row and model. Depending on the benchmark, a number may come from Meta’s internal evaluation, an official leaderboard, or a provider’s self-reported result. Meta also says third-party model runs are best-effort and may not use the provider’s most optimized setup.

That means small score gaps should not automatically be read as controlled model-vs-model wins. A 1- or 2-point difference can combine model capability with harness, version, reasoning effort, provider optimization and provenance differences.

The report is still valuable because it documents task counts and evaluation details for many rows. It simply needs to be read as a mixed-provenance scorecard, not a single laboratory experiment in which every model was rerun under an identical stack.

DeepSWE v1.1: 75.4%, with a harness caveat

One of the strongest coding numbers in the report is 75.4% on DeepSWE v1.1 for Muse Spark 1.3 max. Meta describes DeepSWE v1.1 as 113 software-engineering tasks across 91 repositories in TypeScript, Go, Python, JavaScript and Rust, with handwritten functional and regression tests.

Meta says it evaluated Muse Spark 1.3 max using mini-swe-agent, while comparator numbers were sourced from the official DataCurve leaderboard. The table reports 73.0 for GPT-5.6 Sol max and 74.0 for Claude Opus 5 max.

That makes 75.4 a legitimate Meta-reported DeepSWE result, but the comparison is not equivalent to rerunning all three systems from scratch in one pinned environment with the same agent revision, tool budget, retry policy and provider serving path. The 1.4-point lead over Opus 5 is therefore interesting, not definitive evidence of a universally better coding agent.

Terminal-Bench 2.1 shows that more reasoning is not automatically better

Meta reports 88.8 on Terminal-Bench 2.1 for Muse Spark 1.3 max and 89.2 for xhigh. The benchmark contains 89 tasks verified by executable checks. The table also lists 88.8 for GPT-5.6 Sol max and 86.7 for Opus 5 max.

Two cautions matter. First, Meta’s methodology again says result provenance may vary; the GPT-5.6 Sol value is identified as coming from OpenAI’s model card. Second, the xhigh Spark score is slightly higher than the max score. That is a useful reminder that a larger reasoning budget does not monotonically improve every agent benchmark. Different reasoning settings can change tool behavior, time allocation and failure modes.

Terminal-Bench should also remain separate from repository-repair benchmarks such as SWE-bench. A strong terminal score is evidence about a particular terminal-task distribution, not a substitute SWE-bench result.

OSWorld 2.0: a strong gain, but the predecessor row uses an older environment release

For OSWorld 2.0, Meta’s table reports two scoring variants. Muse Spark 1.3 max scores 66.9 partial / 32.0 binary, while xhigh scores 59.0 / 26.7. Muse Spark 1.2 xhigh is shown at 47.6 / 17.9, producing a large apparent generation-over-generation improvement.

The methodology adds an important version detail: most OSWorld results use the 08.08 release, but Muse Spark 1.2 uses 06.24. The benchmark covers 108 computer-use workflows and Meta says models are run in a common internal framework.

Because the predecessor used a different OSWorld release, the 1.2-to-1.3 jump is not a perfect same-environment A/B. The new model may genuinely be much better, but the exact magnitude should not be attributed solely to the model revision without a matched rerun.

Long context: 98%+ on MRCR is impressive but narrowly scoped

Meta reports 98.5 on MRCR v2 at 256K–512K and 98.1 at 512K–1M for Muse Spark 1.3 max. Those are striking long-context numbers, especially beside 1.2 xhigh at 66.3 and 55.5 in the two bands.

The methodology explains exactly what is being measured. MRCR v2 uses 100 examples per context band, places eight needles into long contexts, and asks the model to retrieve/corefer the requested items. A rule-based sequence matcher scores the answers.

That is a useful stress test of retrieval fidelity and reference resolution over very long contexts. It is not a general 1-million-token reasoning benchmark. It does not prove that the model can maintain equal accuracy on a million-token codebase refactor, multi-hour agent plan, legal synthesis or arbitrary long-horizon workflow. The task distribution and scorer are much narrower.

The broader agent table is competitive, not uniformly dominant

Other Meta-reported rows reinforce the mixed picture:

  • GDPVal-AA v2: Muse Spark 1.3 max 1754 Elo; xhigh 1709; GPT-5.6 Sol max 1710; Opus 5 max 1824. The benchmark has 220 tasks across 44 occupations and 9 industries and uses blind pairwise judging against a 1000-point human baseline.
  • JobBench: 64.9 for 1.3 max versus 65.7 for Opus 5 max. JobBench contains 65 professional tasks spanning 35 occupations.
  • DeepSearchQA: 90.3 for 1.3 max, 93.1 for GPT-5.6 Sol max and 90.4 for Opus 5 max. Meta describes 900 browsing questions.
  • AutomationBench: 49.6 for 1.3 max, 46.7 for GPT-5.6 Sol and 50.3 for Opus 5, over 600 business workflows with deterministic end-state assertions.
  • SWE-Atlas CodeBase QnA: 59.4 for 1.3 max, 53.5 for GPT-5.6 Sol and 52.7 for Opus 5, on 124 tasks across 11 repositories.

These results support the claim that Muse Spark 1.3 is a serious frontier agent model. They do not support saying it “wins every benchmark,” because it does not, and because several rows are not strictly identical-provenance reruns.

SWE-bench Verified and SWE-bench Pro: no exact Spark 1.3 score found here

Meta’s report includes DeepSWE v1.1, Terminal-Bench 2.1 and SWE-Atlas CodeBase QnA. Those are all useful software-engineering evaluations, but none is SWE-bench Verified or SWE-bench Pro.

In this verification pass I did not find an exact, primary, pinned Muse Spark 1.3 result for either SWE-bench Verified or SWE-bench Pro in the reviewed launch material. Therefore this article does not:

  • rename DeepSWE as SWE-bench;
  • transfer a score from Muse Spark 1.2 or another Meta model;
  • infer a SWE-bench result from Terminal-Bench;
  • rank Muse Spark 1.3 on SWE-bench Pro using an unofficial or unpinned number.

If a future result appears, it should be recorded with the exact model snapshot, benchmark revision, agent harness, tool policy, retry budget and date. Verified and Pro should remain separate because their task sets and contamination properties differ.

Public feedback is mixed and highly dependent on the client route

Launch-week Reddit feedback illustrates why a benchmark should not be replaced by anecdote. In an r/opencodeCLI thread titled “Muse Spark 1.3 is nowhere near its Artificial Analysis scores for me,” the author described early tests in which the model repeatedly reread and rewrote the same files, consumed context rapidly, did well on frontend work and performed poorly for their Rust tasks. The author explicitly framed this as first impressions.

A separate r/opencode discussion from a heavy Claude Code user described the free Spark 1.3 experience as “surprisingly great” before the user later encountered very large rate limits. Another thread reported empty responses and upstream 429 errors through an OpenCode Go route.

Those are real user reports, but they do not isolate the model from the surrounding agent client, prompt, provider gateway, rate-limit tier, tool implementation or task selection. In particular, a third-party 429 is evidence about availability on that route at that moment, not proof that Muse Spark 1.3 itself is intrinsically unstable.

Artificial Analysis also posted its launch results on X on September 2. That is a useful dated measurement snapshot, not community sentiment. I did not find enough directly attributable, reproducible user X tests in this bounded pass to claim an X consensus, so no such consensus is invented here.

Practical tradeoffs

For developers, Muse Spark 1.3’s appeal is clear: current independent pricing is low relative to many frontier models, throughput is high on the measured route, the context window is large, and Meta’s coding/agent scorecard is competitive. The model also appears to have made substantial gains over its predecessor on several long-horizon and long-context tasks.

The tradeoff is that the headline benchmark story is unusually sensitive to version and harness details. The Artificial Analysis Index changed two days after launch; Meta’s own scorecard mixes result provenance; one major predecessor comparison uses an older OSWorld release; and the strongest coding result uses a Meta-run mini-swe-agent configuration compared with public leaderboard values.

That does not invalidate the scores. It changes what they can support. The safest conclusion is that Muse Spark 1.3 is a strong, cost-efficient frontier agent/coding model with promising long-context retrieval, while the exact size of its advantage over GPT-5.6 Sol, Claude Opus 5 or other frontier systems remains benchmark- and harness-dependent.

Bottom line

Muse Spark 1.3 deserves attention, but not because one leaderboard number settles the question. Its current Artificial Analysis v4.2 score is 53 at max reasoning, not the launch-era 62, because the benchmark was revised. Meta reports 75.4% DeepSWE v1.1, 88.8 Terminal-Bench 2.1, and 98%+ MRCR long-context retrieval under documented but differing methodologies.

The most important next evidence would be an independent replay of the DeepSWE result with the same model snapshot and harness, an exact SWE-bench Pro or Verified run with pinned methodology, and cost-per-success measurements under a stable public serving route. Until then, Muse Spark 1.3 looks highly competitive—but the most defensible comparison is a conditional one, not a universal rank.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books