Analysis
Analysis

Muse Spark 1.3 Max Reality Check: Public Release, Agent Benchmarks, Token Costs and SWE-bench Gaps

Published Sep 6, 2026 Sources checked Sep 6, 2026

Meta’s Muse Spark 1.3 max is now publicly available, but reasoning-effort tradeoffs, benchmark-version drift and missing exact SWE-bench Verified/Pro results complicate simple leaderboard claims.

Meta released Muse Spark 1.3 on September 2, 2026 with a focus on agentic workflows and long-horizon coding. The most important change since launch is availability: Meta’s current launch page now says Muse Spark 1.3 with max reasoning is available in Muse Code and the Meta Model API. That differs from launch-day independent coverage, which described max as a limited partner preview while xhigh was the broadly available variant. The defensible reading is that access changed after launch, not that the earlier report should be retroactively treated as false.

This matters because benchmark tables already contained both max and xhigh configurations. Developers can now evaluate the higher-effort configuration directly instead of treating it as a preview-only score.

The official benchmark table is broad, but the harness matters

Meta’s four-page evaluation report covers professional work, computer use, web research, automation, coding, long-context retrieval and instruction following. The methodology is not one universal benchmark harness. It mixes Meta runs, official leaderboard results and some provider-reported comparison numbers, and Meta explicitly describes third-party-model runs as best-effort where provider-optimized settings may differ.

For Muse Spark 1.3 max, the report lists:

  • GDPval-AA v2: 1,754 Elo. GDPval-AA v2 contains 220 professional tasks across 44 occupations and nine U.S. industries, evaluated in Artificial Analysis’s Stirrup agent harness with web and shell access.
  • JobBench: 64.9. JobBench contains 65 professional tasks across 35 white-collar occupations.
  • OSWorld 2.0: 66.9 partial / 32.0 strict. OSWorld contains 108 stateful desktop workflows. Meta notes that Muse Spark 1.3 and current comparator runs use the 08.08 release, while Muse Spark 1.2 used 06.24, so the predecessor comparison is not perfectly version-matched.
  • DeepSearchQA: 90.3. The benchmark contains 900 browsing questions and uses the same browser backend and harness across models in Meta’s evaluation.
  • AutomationBench v3: 49.6 pass@1. It contains 600 simulated business workflows with deterministic end-state checks.
  • MRCR v2: 98.5 at 256K–512K and 98.1 at 512K–1M. Meta uses 100 examples in each band with the 8-needle variant.
  • DeepSWE v1.1: 75.4. This is 113 software-engineering tasks across 91 repositories and five languages. Meta says it ran Muse Spark 1.3 max with mini-swe-agent and sourced comparison-model numbers from the official Datacurve leaderboard.
  • SWE-Atlas Codebase QnA: 59.4. The public QnA split has 124 tasks across 11 production repositories.
  • Terminal-Bench 2.1: 88.8 pass@1. The benchmark contains 89 terminal tasks; Meta says each model uses its named/native coding harness in an isolated cloud sandbox.

Those values should not be compressed into a single “coding score.” They test different capabilities, use different task sets and in some cases use different model-specific harnesses.

Max is not automatically better on every task

The official table is useful because it exposes both reasoning configurations. Max beats xhigh on most reported agentic and long-context rows—for example GDPval-AA v2 is 1,754 versus 1,709; OSWorld partial is 66.9 versus 59.0; and the 512K–1M MRCR result is 98.1 versus 93.1.

But max does not dominate every row. Terminal-Bench 2.1 is 88.8 for max and 89.2 for xhigh in Meta’s table. That small reversal is a reminder that “more reasoning effort” is not a guarantee of higher task success, especially once an agent harness, tool policy, stopping behavior and execution environment enter the loop.

A practical evaluation should therefore compare max and xhigh under the same repository, task set, tool permissions, retry budget and success criteria—not assume that the highest reasoning setting is automatically optimal.

SWE-bench Verified and SWE-bench Pro are separate—and neither has an exact 1.3 result here

SWE-bench Verified: I did not verify an exact Muse Spark 1.3 result from Meta or the benchmark operator in this pass.

SWE-bench Pro: I likewise did not verify an exact Muse Spark 1.3 result from Meta or a benchmark-operator source in this pass.

That means DeepSWE 75.4, SWE-Atlas 59.4 and Terminal-Bench 88.8 must not be substituted for either SWE-bench variant. It also means older Muse Spark results should not be inherited by Muse Spark 1.3. Several secondary pages currently circulate unsupported or conflicting SWE-bench numbers; without an exact model revision, split, scaffold, internet/tool policy, retries and run count, those values are not strong enough for a fair comparison.

Meta’s efficiency claim is narrower than “25% cheaper”

Meta says that in comparisons by its engineers, Muse Spark 1.3 used about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 while improving coding workflow usability. That is a vendor-reported internal comparison and Meta does not publish enough task-level detail on the launch page to turn it into a universal cost reduction.

Artificial Analysis found a more complicated pattern in its launch evaluation. For the xhigh configuration, it reported $1.25 per million input tokens, $4.25 per million output tokens and $0.15 per million cached-input tokens, with about $0.55 per Intelligence Index task under the then-current Index. It also reported that xhigh used roughly 57% more input tokens per agentic task than Muse Spark 1.2 in its own evaluation, even though output tokens rose only modestly. The two findings are not necessarily contradictory: Meta’s figure is about its internal coding comparison, while Artificial Analysis is measuring a different composite workload.

The practical metric is therefore cost per successful workflow, not token tariff or vendor token-reduction percentage alone.

Current independent measurements changed after the Index changed

Artificial Analysis’s September 2 launch article scored Muse Spark 1.3 at 62 for max and 61 for xhigh on the Intelligence Index it was using that day. On September 4, Artificial Analysis moved to Intelligence Index v4.2, adding AA-Briefcase and GDP.pdf, removing GPQA Diamond, increasing the weight of private/held-out evaluations and changing grading infrastructure.

Its current pages now report 53 for max and 52 for xhigh. That 62→53 or 61→52 movement should not be described as a model regression. The evaluation yardstick changed.

The current max page reports approximately 190.1 output tokens/s, 18.69 seconds time to first token, 1M context, and $1.25/M input + $4.25/M output on Meta’s API. The xhigh page reports approximately 135.2 output tokens/s and 42.55 seconds TTFT at the same listed token prices. These are Artificial Analysis measurements for a particular provider and test setup, not guaranteed latency for every region, prompt length or workload.

One useful result is that max is currently faster in Artificial Analysis’s measured output throughput and TTFT while also using more output tokens across the Index. That makes end-to-end cost and latency task-dependent rather than reducible to a single “max is slower” rule.

Public feedback is mixed and selection-biased

Public OpenCode discussions from September 2–5 are mixed. Some users describe Muse Spark 1.3 as fast with good tool use or a capable orchestrator, while others call it mediocre compared with alternatives or report that another model worked better for their tasks. Separate OpenCode threads report provider-specific free-tier limits and region restrictions.

Those are useful operational anecdotes, especially for discovering access and integration problems, but they are not a statistically representative satisfaction survey. OpenCode’s availability policy can also differ from direct Meta Model API availability, so a region restriction in a third-party route should not be presented as proof that Meta’s own API is unavailable in that country.

I also checked for attributable X discussion. I did not find a reproducible independent technical X thread strong enough in this bounded pass to support a performance or reliability claim. Vendor posts announcing the launch or wider max availability are evidence of the vendor’s announcement, not independent user consensus.

Open weights are promised, not released

Meta’s launch page says a Muse Spark open-weights release is on the roadmap. The current Artificial Analysis model page still classifies Muse Spark 1.3 max as proprietary, and I found no public weights/license release for 1.3 in this pass. “Open weights coming” should therefore remain a future commitment until a downloadable model and license actually appear.

What to test before switching

Muse Spark 1.3 max is now materially more relevant because the configuration has moved from launch-time preview status to current first-party availability. The strongest official evidence is in agentic professional work, computer use, long-context retrieval and coding-agent benchmarks, but the evaluation details show why direct leaderboard ranking can mislead.

For a production coding or agent workload, compare max and xhigh on a frozen task set and record: accepted-task success rate, wall-clock time, input/output/reasoning tokens, tool-call count, retries, cache-hit rate, total cost, and failure class. Keep SWE-bench Verified and SWE-bench Pro separate if exact 1.3 runs are later published. Re-run comparisons after benchmark-version changes rather than interpreting a composite-index shift as a model capability change.

The evidence supports a measured conclusion: Muse Spark 1.3 max is now publicly usable through Meta’s first-party products and posts strong agentic benchmark results, but max does not win every task, the most prominent independent composite benchmark changed days after launch, and exact SWE-bench Verified/Pro evidence for 1.3 remains unresolved.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books