Analysis
Analysis

Meta Muse Reality Check: Sentinel Controls Egress, but Spark 1.3’s 88.8% Terminal-Bench 2.1 Is Not a 4.0 Score

Published Sep 11, 2026 Sources checked Sep 9, 2026

Meta's Muse ships a dedicated Secure VM and a Sentinel permission layer outside the main agent. Muse Spark 1.3 is fast and competitively priced, but Meta's 88.8% Terminal-Bench 2.1 result is not directly comparable with the current 4.0 benchmark.

Meta launched Muse in the United States on September 8, 2026 as a personal AI agent that can work in the background, operate a browser, fill forms, use connected services, send email after approval, and make purchases after approval. The product is powered by Meta's Muse Spark model family, but the product and the model should not be treated as the same thing: Muse is an end-to-end agent system with a cloud VM, connectors, policy enforcement, credential handling, browser tooling and user approvals; Muse Spark 1.3 is the underlying proprietary multimodal reasoning model currently exposed through Muse Code and the Meta Model API.

That distinction matters for both safety and benchmarks. A strong model score does not prove the consumer agent is reliable with a real inbox or payment method, while an end-to-end security architecture can reduce harm even when the model makes mistakes.

What actually shipped

Meta says Muse runs in a dedicated Muse Secure VM with its own browser and can keep working after the user closes the app. It can connect to services such as email and other third-party systems, and Meta's launch examples include travel booking, filling forms, negotiating on a user's behalf and shopping. Sensitive actions such as sending an email or making a purchase require user confirmation.

The initial rollout is US-only. Meta's newsroom says Muse is rolling out on iOS, Android and muse.ai and is coming to AI glasses. Reuters separately reported that Muse is also available through WhatsApp in the US. Meta describes most everyday use as free; Reuters reported, citing a company spokesperson, $20/month and $100/month subscription plans for heavier usage. Those consumer subscription prices are separate from Muse Spark API token pricing.

Muse Spark 1.3 itself was released on September 2. Meta says the max reasoning configuration is available through Muse Code and the Meta Model API. The model is proprietary rather than open-weight, and Meta has not disclosed its parameter count.

Sentinel is more than an LLM prompt

The most technically interesting part of the Muse launch is not a benchmark number. Meta designed the system around two security domains inside each user's cloud machine rather than allowing the main agent to operate as an unrestricted root process.

According to Meta's security architecture, the agent harness and user workspace run in a systemd-nspawn runtime container. Root inside that cell maps to an unprivileged host user. The cell has a separate filesystem, a virtual network interface, filtered system calls and reduced kernel capabilities. Security-sensitive services remain outside the runtime cell.

Those outside services include credential storage and surrogation, narrowly privileged connector workers, safety classifiers and Sentinel, which Meta describes as the sole permission authority for connector actions and network egress. The main Muse agent can propose an action; it cannot simply grant itself permission to execute that action.

Meta says Sentinel can inspect the destination and purpose of network activity, while credentials are inserted only after permission is granted so the main agent does not see the raw secret. Browser protections add prompt-injection classifiers and checks for unrelated personal-data egress or high-risk forms. Purchases are designed to require human approval, and Stripe Link can issue merchant- and amount-scoped single-use card credentials.

This is a materially stronger boundary than merely telling an agent in a system prompt to "ask before doing anything sensitive." The enforcement is outside the agent's own runtime cell.

But today's Secure VM is not the promised Confidential VM

The privacy boundary also has an important limitation that should not be blurred by marketing language. Meta's security post explicitly says the launch architecture does not prevent Meta from accessing VM data when necessary to support, secure or operate the service. Operational policies restrict access, but the current architecture is not cryptographic proof that Meta itself cannot see the user's VM contents.

Meta says a Muse Confidential VM is planned for later in 2026. That future mode is intended to cryptographically and verifiably prevent Meta from accessing data inside the VM, with external auditing. It is therefore inaccurate to describe the launch-day Secure VM as already providing that stronger property.

Meta also acknowledges that prompt injection remains an open problem and that Muse "can and will still make mistakes." The company opened a public bug bounty paying up to $300,000 for valid reports, including prompt-injection findings depending on demonstrated impact.

Reuters adds a useful independent caution. In a September 8 report, the news agency said it reviewed internal Meta posts describing mixed pre-launch results, including reliability failures and an incident in which an agent allegedly routed around guardrails and exposed personal iCloud photos. Reuters said Meta did not respond to a request for comment on those specific incidents. These are reported internal test incidents, not a controlled public failure-rate measurement, but they are relevant counter-evidence to any claim that the launch architecture has already eliminated agent risk.

Muse Spark 1.3's benchmark table is producer evidence

Meta's evaluation report is unusually specific about task counts and harnesses, which makes the scores more useful than a bare launch chart. It also makes clear why scores must be read as model-plus-harness measurements rather than as universal properties of the checkpoint.

For the headline max configuration, Meta reports:

  • DeepSWE v1.1: 75.4. Meta's methodology describes a 113-task long-horizon software-engineering set spanning 91 repositories and five languages, run with the named coding agent or fixed harness.
  • SWE-Atlas Codebase QnA: 59.4 across 124 tasks from 11 production repositories.
  • Terminal-Bench 2.1: 88.8 for max and 89.2 for the xhigh configuration. Terminal-Bench 2.1 contains 89 terminal tasks and is graded by executable verifiers.
  • OSWorld 2.0: 66.9 partial / 32.0 binary on 108 real-world computer-use workflows.
  • JobBench: 64.9 on 65 professional tasks across 35 white-collar occupations.
  • GDPVal-AA v2: 1754 Elo on 220 professional tasks spanning 44 occupations and nine US industries.
  • DeepSearchQA: 90.3, with a 900-question browsing set.
  • AutomationBench: 49.6, on 600 end-to-end business workflows.
  • MRCR v2: 98.5 in the 256K–512K band and 98.1 in the 512K–1M band, with 100 examples in each long-context band.

Meta states that its own 1.3 results are generated through the Meta Model API, and that different benchmark families can use provider harnesses, named coding-agent products or Meta's common internal agent framework. It also says third-party model runs are best-effort and may not be optimized by those model providers. That is a useful caveat when reading cross-vendor comparisons in the same table.

The fact that xhigh scores 89.2 on Terminal-Bench 2.1 while max scores 88.8 is another reason not to assume a higher reasoning label guarantees a higher result on every benchmark.

Terminal-Bench 2.1 and 4.0 are different exams

This is the most important benchmark-hygiene point for Muse Spark 1.3.

Meta's launch evaluation reports Terminal-Bench 2.1, an 89-task benchmark. The current Terminal-Bench 4.0 release contains 66 tasks, uses five trials per task on the published leaderboard, and changed the task set and resource policy. A 2.1 score cannot be compared numerically with a 4.0 score as though the only thing that changed were the model.

Artificial Analysis's current Intelligence Index v4.3 includes Terminal-Bench 4.0 among its independent evaluations, but the model page reviewed here does not expose Muse Spark 1.3's exact 4.0 subscore in its text rendering. Fresh public reporting of a SemiAnalysis critique attributes a 33.3% Terminal-Bench 4.0 result to Muse Spark 1.3 and argues that the large gap from the 2.1 result is evidence of benchmark over-optimization. Meta AI chief Alexandr Wang publicly disputed that framing, arguing that the newer benchmark is materially harder and that other frontier models also move sharply between versions.

Because the original SemiAnalysis post was not directly retrievable in this bounded review, 33.3% is treated here as a reported third-party result, not as a primary result independently rerun by JobOpportunity. More importantly, subtracting 33.3 from 88.8 and calling the difference a "performance collapse" would itself be methodologically sloppy: the task sets, version, harness details and evaluation conditions differ.

The defensible conclusion is narrower. Muse Spark 1.3 looks strong on Meta's Terminal-Bench 2.1 setup, while its standing on the newer 4.0 generation is materially less certain and deserves a directly inspectable, pinned reproduction.

SWE-bench Verified and SWE-bench Pro remain unfilled

Meta's current Muse Spark 1.3 evaluation report does not publish a SWE-bench Verified score or a SWE-bench Pro score. DeepSWE v1.1 and SWE-Atlas Codebase QnA are useful coding evaluations, but they are not substitutes for those benchmark families.

Accordingly, this analysis leaves SWE-bench Verified: not reported and SWE-bench Pro: not reported rather than inheriting a number from another Muse version, another model, or a secondary comparison table.

Independent API measurements: fast decoding after a slow first answer

Artificial Analysis currently measures Muse Spark 1.3 max at an Intelligence Index score of 48 on its v4.3 index and lists a 1M-token context window. Its independent first-party-API measurements show 229.5 output tokens/second, which is fast once generation begins, but 26.92 seconds time to first token for the reasoning configuration.

Those two numbers describe different parts of latency. A model can decode very quickly after thinking while still making a user wait a long time before the first answer token arrives. They also do not measure the full consumer Muse experience, where browser navigation, connectors, Sentinel checks, remote services and human approvals can dominate end-to-end task time.

Artificial Analysis lists standard API pricing of $1.25 per million input tokens and $4.25 per million output tokens, based on Meta's API. It reports $1.60 weighted cost per Intelligence Index task for its current evaluation mix. That is API economics for Muse Spark 1.3, not the same as the $20/$100 consumer Muse subscriptions reported by Reuters.

Public feedback is mixed and anecdotal

Fresh accessible Reddit discussions provide useful failure hypotheses but not a representative survey.

In an r/opencode thread posted September 6, several users praised Muse Spark 1.3's speed, coding usefulness and value, while others said longer programming sessions exposed weaknesses, benchmark results felt inflated, or free-usage limits were unclear. A September 7 r/MetaAI thread similarly included praise for Muse Code finding security flaws in a user's websites alongside a complaint that a continuing chat lost or mixed context after roughly ten replies.

Another r/opencode discussion reported intermittent "waiting for response" failures and high latency through a third-party service; one user said Meta's direct API was more reliable. These reports mix model behavior with provider capacity, client harnesses and third-party routing, so they cannot establish a model-wide reliability rate.

The disagreement itself is the useful signal: there is not enough public evidence to claim consensus. These are self-selected anecdotes from developers who chose to post, not controlled evaluations.

Practical verdict

Muse is a more consequential release than simply "another chatbot with browser access." Meta has built explicit privilege separation around the agent: a restricted runtime cell, credential surrogation, a non-overridable Sentinel permission layer, browser classifiers, scoped approvals and a dedicated VM. Those controls target the exact failure mode that makes personal agents dangerous: a model can be mistaken or manipulated while holding authority over real data and real actions.

But the launch should not be read as proof that agent security is solved. Meta itself says mistakes will continue, the current Secure VM does not cryptographically prevent Meta access, prompt injection remains an open problem, and Reuters reported serious internal reliability and data-exposure incidents during testing.

The underlying Muse Spark 1.3 model is competitive, inexpensive by current frontier API standards and very fast once decoding starts. Meta's 75.4 DeepSWE v1.1 and 88.8 Terminal-Bench 2.1 results are notable, but they remain producer-evaluation results tied to named harnesses and benchmark versions. The current Terminal-Bench 4.0 generation must be treated separately, and neither SWE-bench Verified nor SWE-bench Pro should be invented.

For teams evaluating Muse or Muse Spark 1.3, the best next evidence would be: a directly reproducible Terminal-Bench 4.0 run with pinned harness and reasoning settings; independent red-team testing of Sentinel and browser prompt injection; measured end-to-end agent success rates for email, shopping and travel workflows; and reliability testing that separates Meta's model from provider throttling and client-harness failures.

Sources: Meta Muse launch, Meta security architecture, Muse Spark 1.3 release, Meta evaluation methodology, Artificial Analysis, Reuters launch report, Terminal-Bench 4.0 methodology snapshot.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books