Claude Mythos 5.1: Restricted Access, Benchmarks and Safeguard Tradeoffs
Claude Mythos 5.1 shares an underlying model with Fable 5.1 but uses different safeguards and restricted access. Here is what its benchmarks and science claims do—and do not—prove.
Mythos 5.1 is not a separate set of model weights
Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026. The important identity detail is easy to miss: Anthropic says Fable 5.1 and Mythos 5.1 are the same underlying model, deployed with different safeguard policies. Fable is the broadly available product. Mythos is the restricted variant intended for vetted cybersecurity and life-sciences work where some of Fable's cyber and biology restrictions would otherwise interfere with legitimate research.
That distinction matters when reading benchmark tables. A higher Mythos score does not automatically mean Anthropic trained a smarter model with different weights. It can instead reflect fewer safeguard interventions on tasks that the underlying model could already solve.
Primary sources: Anthropic launch announcement and Claude Mythos product page.
Access is limited, and the programs are not identical
Anthropic currently says Mythos 5.1 is available to a small set of vetted organizations. Its Life Sciences Verification Program is an invite-only beta designed to give qualifying researchers reduced biology safeguards, while other safeguards remain in place. Anthropic says initial participants have been enrolled in partnership with the US government.
The Cyber Verification Program currently provides reduced cyber safeguards for certain other Claude model classes and is expected to include Mythos-class models in the near future. Anthropic also says Claude Security, its codebase vulnerability-scanning product, now runs on Mythos 5.1. At the time of verification, Mythos 5.1 access was limited to a set of US organizations, with broader domestic and international expansion described as future work.
This means a normal Claude subscriber or API customer should not assume that selecting Fable 5.1 grants Mythos behavior. Access is a separate vetted program.
Pricing is published, but several operational numbers are not
Anthropic says Mythos 5.1 pricing starts at $10 per million input tokens and $50 per million output tokens. It also states that Mythos use carries a default 30-day data-retention policy for safety monitoring.
The current Mythos page does not publish a Mythos-specific context-window specification, representative output speed, time-to-first-token figure, or cost-per-success benchmark. Those values are therefore left unknown here rather than inherited from Fable 5.1.
That caution is important because Artificial Analysis currently measures the generally available Fable 5.1 service, not Mythos. Its Fable test uses Anthropic's default server-side fallback configuration. The independent page currently reports a 1-million-token context window and measured Fable service performance, but those service-level measurements should not be silently transferred to a restricted Mythos deployment.
Independent reference: Artificial Analysis — Claude Fable 5.1.
Terminal-Bench 4.0: 60.9% for Mythos, 55.8% for Fable in Anthropic's table
Anthropic's launch table reports 60.9% for Mythos 5.1 and 55.8% for Fable 5.1 on Terminal-Bench 4.0. Anthropic explicitly says the two are the same underlying model and attributes the gap to tasks where Fable's safeguards intervene. It also says improved safeguard precision should make the difference smaller.
This is useful evidence about the effect of a deployment policy, but it is still a vendor-run comparison. It should not be read as a clean 5.1-point gain in base-model intelligence. The result also should not be mixed with older Terminal-Bench versions or with another coding benchmark under a different agent harness, task set, trial count, tool configuration, or timeout.
The same Anthropic launch page says Fable 5.1 was evaluated with production safeguards enabled. On some benchmark interventions, requests were either scored as zero or routed to other Claude models. That makes safeguard and routing configuration part of the measured system.
SWE-bench Verified and SWE-bench Pro remain separate—and neither supplies a Mythos result here
SWE-bench Verified and SWE-bench Pro are different benchmark families and should not be collapsed into a single coding score. SWE-bench Verified is the human-validated 500-task subset of the original SWE-bench. SWE-bench Pro is a separate Scale benchmark with different repositories, task construction, splits, and evaluation design.
In the primary and independent sources checked for this article, I did not find a directly attributable, standardized Mythos 5.1 result on either SWE-bench Verified or SWE-bench Pro. Therefore this analysis does not substitute the Terminal-Bench 4.0 number, a Fable 5.1 result, or an unverified third-party row as a Mythos SWE-bench score.
References: SWE-bench Verified and SWE-bench Pro public leaderboard.
The scientific-research claims are notable, but they are not an independent leaderboard
Anthropic reports that Mythos 5.1 designed protein binders using open-source protein-design and folding tools and sent designs to two external organizations for experimental validation. Anthropic says that, on three targets, binding affinities were ten times higher than the best designs from referenced Adaptyv Bio competitions, and that viable-binder hit rate approached 50% across 12 targets.
Anthropic also reports a computational-biology exercise in which Mythos 5.1 optimized seven open-source deep-learning models by up to 2.5 times on GPU workloads while preserving identical outputs, with estimated GPU-cost reductions of 30–60% for example genome-scale analyses.
These results are more concrete than a marketing adjective because they describe artifacts and external lab validation, but they remain Anthropic-organized experiments rather than a standardized independent Mythos leaderboard. Reproducible code, task definitions, full cost accounting, failure cases, and third-party replication would make them much easier to compare with other systems.
Safety evidence is part of the product tradeoff
Anthropic describes Mythos 5.1 as its strongest released model for cybersecurity while still placing it below the next risk tier in its own frameworks. It also reports that Mythos 5.1 was its most robust model to date on an external prompt-injection benchmark and that automated behavioral auditing found lower rates of reward-hacking behavior than Mythos 5.
Those statements should not be read as a claim that the system cannot fail. Anthropic's own release material says Mythos still requires restricted access for sensitive biology and cyber capabilities. The value proposition is deliberately a tradeoff: more permissive safeguards for vetted defensive or research work, with additional monitoring and program controls.
For teams deciding whether Mythos would be useful, the relevant questions are therefore not only model quality. They include eligibility for trusted access, data-retention requirements, allowed use cases, safeguard behavior, auditability, latency, and cost per successfully completed research task.
Public feedback is too sparse to call a consensus
Because access is restricted, public first-hand Mythos 5.1 usage reports are much thinner than Fable 5.1 reports. An r/Anthropic launch discussion confirms broad interest around the announcement, but discussion threads are self-selected and many comments are reactions to launch material rather than measured Mythos use.
A Claude account post on X was surfaced indirectly by public launch coverage, but I could not retrieve a directly verifiable X source suitable for quoting in this review. No X sentiment score or supposed community consensus is therefore reported.
That absence of evidence matters. Restricted-access models are especially vulnerable to reputation being built from vendor demos, second-hand descriptions, and users who may actually be testing the generally available Fable variant.
Practical conclusion
Claude Mythos 5.1 is best understood as a restricted deployment of the Fable 5.1 underlying model with more permissive safeguards for approved cyberdefense and life-sciences work—not as a separate consumer frontier model.
The strongest quantified launch distinction is Anthropic's 60.9% versus 55.8% Terminal-Bench 4.0 comparison, but Anthropic itself says safeguard intervention explains the gap. Independent evaluators currently provide useful Fable 5.1 measurements, not a comparable public Mythos service row. There is also no verified Mythos 5.1 SWE-bench Verified or SWE-bench Pro score in the evidence checked here.
Confidence is high on model identity, launch date, access structure, starting token price, default retention policy, and Anthropic's stated benchmark numbers because those come from current primary sources. Confidence is medium on the scientific capability claims because the work includes external validation but remains vendor-organized. Confidence is low on community-wide sentiment and real-world latency or cost-per-success because public access and reproducible field data remain sparse.
This article is built from the source material below. Open the originals for full context and the latest updates.