Claude Fable 5.1 Reality Check: 1M Context, Cache Pricing, Agent Benchmarks and SWE-bench Gaps
Claude Fable 5.1 brings a 1M context window, 128K output, cheaper cache reads and strong agent benchmarks. We separate vendor claims from independent measurements, explain benchmark-version drift, and keep SWE-bench Verified and Pro distinct.
Anthropic released Claude Fable 5.1 on September 1, 2026 as its generally available model for demanding reasoning and long-horizon agentic work. The exact Claude API model ID is claude-fable-5-1. Anthropic also released Claude Mythos 5.1, which it describes as the same underlying model with a different safeguard configuration and restricted trusted access for vetted cybersecurity and life-sciences work.
That distinction matters. Fable 5.1 and Mythos 5.1 should not be treated as two independently trained foundation models, and benchmark differences between them can reflect safeguard intervention rather than different base capabilities. Anthropic's launch materials explicitly say the production safeguards were enabled during its evaluations and that some security-sensitive tasks could be redirected to earlier Claude models.
Availability, context and pricing
Anthropic's current model documentation lists Fable 5.1 as active, released September 1, 2026, with a 1 million-token context window, 128K maximum output, text-and-image input, text output, adaptive thinking that is always on, and high as the default effort setting. The model is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. AWS separately confirmed general availability on September 1.
The headline API rates remain $10 per million input tokens and $50 per million output tokens. The important price change is prompt-cache reuse: Anthropic reduced cache reads to $0.25 per million tokens, 75% below the previous Fable 5 cache-read rate. Five-minute cache writes are $12.50 per million tokens and one-hour cache writes are $20 per million tokens. Anthropic says Batch API input and output receive a 50% discount.
Anthropic estimates that the lower cache-read price reduces the total cost of a typical Fable workload by about 25%, and can reduce highly agentic, context-heavy workload cost by up to roughly 45%. Those are workload estimates, not a promise that every request is 25–45% cheaper. Base uncached input and output rates did not fall.
This is an important practical distinction because long-running agents repeatedly reuse large histories, tool results and project context. A model can keep the same headline token prices while becoming materially cheaper for a cache-heavy workload.
Anthropic's agent and reasoning benchmarks
Anthropic reports a large gain on Terminal-Bench-Science 0.1: Fable 5.1 scores 52.6%, compared with 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol in Anthropic's table. The company also discloses useful uncertainty information here: its setup used the Claude Code harness and the public benchmark uses three trials per task; Anthropic says the standard error is roughly ±3.5 to ±4.5 points per model.
On Terminal-Bench 4.0, Anthropic reports 55.8% for Fable 5.1 and 60.9% for Mythos 5.1. Because the two are the same underlying model, Anthropic says the gap reflects tasks where the Fable safeguard layer intervened. That is a reason not to read 60.9 versus 55.8 as evidence that Mythos is a separately stronger base model.
Other launch results include 1,853 Elo on GDPval-AA v2, 77.9% OSWorld 2.0 partial / 41.7% strict, 60.9% Humanity's Last Exam without tools / 65.0% with tools, 31.4% AutomationBench, and 73.4% CursorBench 3.2.0.
The OSWorld result has a versioning caveat that should travel with the number. Anthropic says these scores use the benchmark authors' August 2026 task release, and that earlier published OSWorld 2.0 results are not directly comparable because the task files changed. Benchmark name alone is therefore not enough; task release, harness and safeguard behavior are part of the result identity.
Independent evaluation: strong results, but cost and latency tradeoffs remain
Artificial Analysis received pre-release access and initially reported Fable 5.1 Max at 66 on the then-current Artificial Analysis Intelligence Index. It also found that the five effort levels spanned roughly 11× in output-token usage, from 13.1 million tokens across the low-effort evaluation set to 143.7 million at max effort. In that launch evaluation, Fable 5.1 Max cost about $3.76 per Intelligence Index task, around 20% more than Fable 5 Max, even after the cache-read reduction, because it generated substantially more output tokens.
That is a useful reminder that cheaper cache reads do not guarantee a cheaper completed task. For an agentic workload, total cost depends on uncached input, cache writes, cache hits, reasoning/output volume, retries and whether the task succeeds.
Artificial Analysis changed its Intelligence Index to v4.2 on September 4, 2026. The new version added AA-Briefcase and GDP.pdf, removed saturated GPQA Diamond, doubled the held-out/private-data weighting to 40%, changed grading infrastructure and re-anchored some ratings. Its current Fable 5.1 Max page reports an Index score of 57, not the launch-week 66.
That 66-to-57 change should not be called a nine-point Fable regression. The evaluation yardstick changed. A defensible time-series comparison must hold the benchmark version fixed or re-evaluate both snapshots under the same harness.
Artificial Analysis' current first-party API measurements report about 69 output tokens per second for Fable 5.1 Max and a very high 273.6-second time to first answer token in its reasoning-heavy test configuration. The page defines that latency metric as including the model's thinking time before the first answer token. Those figures are useful for that measured setup, but they are not universal service-level guarantees: effort level, prompt, region, load, streaming behavior and provider route can change observed latency substantially.
SWE-bench Verified and SWE-bench Pro must stay separate
SWE-bench Verified is a human-validated subset of 500 SWE-bench instances. The benchmark maintainers say the language-model comparison track uses a standardized mini-SWE-agent bash environment, and they warn that results from mini-SWE-agent 1.x and 2.x are not necessarily comparable because the action mechanism changed.
In this bounded verification pass, I did not find a primary Anthropic or benchmark-operator publication of an exact Claude Fable 5.1 SWE-bench Verified score. That field should therefore remain unknown rather than being filled with a number from a different SWE-bench variant.
SWE-bench Pro is a different benchmark. A current third-party BenchLM leaderboard lists 81.2% for Claude Fable 5.1 and explicitly warns that its rows can come from different splits, scaffolds, tool budgets, retry policies and run counts. I could not trace that 81.2 row to a primary Anthropic or benchmark-operator run with enough frozen harness details during this pass. It is therefore best treated as a provisional third-party published row, not silently upgraded into a primary-source-verified Anthropic benchmark claim.
The practical rule is simple: never substitute a Pro score for Verified, never substitute a CursorBench or Terminal-Bench result for either, and do not rank small score differences unless model snapshot, dataset split, agent harness, tool/internet policy, retry count and inference settings match.
Vendor claims versus reproducible evidence
Anthropic calls Fable 5.1 its strongest model for coding and knowledge work and publishes several early-customer anecdotes. Those testimonials can identify promising workloads, but they are not controlled benchmark results. Anthropic's own benchmark page is more useful where it discloses harness details, uncertainty or task-release caveats.
The strongest evidence stack for choosing Fable 5.1 today is therefore layered:
- Use Anthropic's documentation for exact model identity, context, output limits, pricing and availability.
- Use Anthropic's benchmark table for named launch results, while retaining safeguard, harness and task-version caveats.
- Use independent evaluations such as Artificial Analysis for separately measured cost, token use, throughput and composite-evaluation results.
- Re-run your own workload at multiple effort settings before deciding that the Max setting is worth its token and latency cost.
Public feedback is mixed and self-selected
Fresh Reddit discussion illustrates why anecdotal feedback needs caution. One September 2026 user reported that Fable 5.1 appeared to consume a weekly Claude allowance much faster than Fable 5. In the broader usage-limits discussion, other users reported different consumption patterns and tried to attribute changes to cache reads, output volume or session behavior. These posts do not expose a controlled denominator, fixed workload or billing trace that can establish a model-wide rate.
They are useful as hypotheses to test, not as evidence of consensus. Usage-plan limits are also not the same thing as API per-token pricing.
I did not find a sufficiently detailed, attributable first-hand X post in this bounded search that added reproducible technical evidence beyond Anthropic's own launch announcement and the accessible benchmark/developer discussions. No X consensus is claimed.
Practical tradeoffs
Fable 5.1 is most defensible for workloads where its stronger long-horizon behavior can offset its premium base price: difficult codebase work, research, complex document production and agents that reuse large cached context. Anthropic itself recommends starting with the cheaper Opus 5 for most workloads and moving to Fable 5.1 when higher-effort Opus evaluations still fall short.
Its 1M context and 128K output ceiling are substantial, but maximum context is not the same as economical context. Cache-hit rate, reasoning effort and retry behavior can dominate total cost. Likewise, an impressive Terminal-Bench or CursorBench score does not guarantee better performance on a company's own repository.
The current evidence supports a nuanced conclusion: Fable 5.1 is a genuine frontier release with strong agentic and research results and dramatically cheaper cache reads, but it remains expensive at uncached base rates, high-effort latency can be substantial, benchmark versions are already moving, and exact SWE-bench Verified evidence should not be invented. The best comparison is a frozen workload with the same harness, effort setting, tool policy and success criterion, measured on cost and latency per successful task, not a single leaderboard number.
This article is built from the source material below. Open the originals for full context and the latest updates.