Anthropic Launches Claude Fable 5.1, Cutting Agentic AI Costs by Up to 45%
Anthropic released Claude Fable 5.1 on September 1, along with a restricted-access sibling model called Claude Mythos 5.1, marking the company's first major upgrade to its flagship model line since Fable 5 shipped roughly three months earlier. The update pairs meaningful performance gains on coding and multi-step agentic tasks with a substantial cut in operating costs, a combination Anthropic is positioning as a direct challenge to rivals racing to make frontier models cheaper to run at scale.
Fable 5.1 is now generally available through Claude.ai, the Claude Developer Platform, Claude Code, Claude Enterprise, and through cloud partners Amazon Web Services, Google Cloud, and Microsoft Azure. Mythos 5.1, which Anthropic describes as the identical underlying model with reduced default safeguards, is being distributed only to vetted organizations in the United States through two new channels: a Cyber Verification Program for cybersecurity professionals and a Life Sciences Verification Program developed in coordination with the U.S. government.
The headline change for most enterprise customers is price. Anthropic left input and output token pricing for Fable 5.1 unchanged at $10 and $50 per million tokens, respectively, but cut the cost of cache reads by 75%, from $1.00 to $0.25 per million tokens. Because agentic workflows lean heavily on cached context as a model repeatedly reasons over the same codebase or documents, Anthropic estimates the change lowers effective costs by roughly 25% for typical workloads and by up to 45% for highly agentic, long-running tasks.
On benchmarks, Anthropic reported gains that are large enough to be notable rather than incremental. On Terminal-Bench-Science 0.1, a test of agentic scientific research, Fable 5.1 scored 52.6% compared with 24.7% for Fable 5, more than double the prior result. On Terminal-Bench 4.0, a coding-agent benchmark, the score rose to 55.8% from 42.0%. On Humanity's Last Exam, a difficult multidisciplinary reasoning test, Fable 5.1 reached 60.9% without external tools, up from 57.8%. On OSWorld 2.0, which measures a model's ability to operate a computer interface to complete tasks, partial-credit performance improved to 77.9% from 72.9%. Third-party evaluator Browserbase separately reported Fable 5.1 completing 82% of tasks on its hardest browser-agent benchmark, ahead of Anthropic's own Opus 5 at 74% and Fable 5 at 57%.
Anthropic and several early customers pointed to sustained, unsupervised operation as the more consequential shift than any single benchmark score. Trading firm Millennium said Fable 5.1 traced an extremely rare software crash to a bug in an external vendor library after the issue had resisted explanation from engineers and other AI models for four to five years. Fintech company Ramp said the model completed a 38-hour machine learning run without supervision, autonomously re-evaluating its approach across six experiments and returning recommendations. Database company MongoDB said it used the model to build a complex prototype over three days, including phases that ran and verified themselves overnight. AI coding startup Cognition, maker of the Devin agent, said it is shifting traffic that previously ran on Anthropic's Opus 5 over to Fable 5.1 at launch, citing lower cost and performance that matched or exceeded its prior model in testing.
The Mythos 5.1 variant is aimed narrowly at defensive cybersecurity and life-sciences research, domains where Anthropic's standard safety restrictions can get in the way of legitimate expert work. The company said expert red-teaming and tabletop exercises found Mythos 5.1's biological and chemical risk profile still falls below the next tier defined in Anthropic's Responsible Scaling Policy, and that despite having the strongest cyber capabilities of any model Anthropic has released, external testers found no critical-severity jailbreaks. Anthropic also said the model is measurably less prone to accessing resources outside designated test environments and to what the company calls "motivated reasoning," language that reflects a broader industry preoccupation with keeping increasingly autonomous agents inside their intended boundaries as they take on longer, less-supervised tasks.
Alongside the model release, Anthropic introduced several new enterprise controls. Enterprise Frontier Safeguards let customers store data with customer-managed encryption inside their own AWS, Azure, or Google Cloud environment rather than relying solely on Anthropic's zero-data-retention policy. The company said updated safety classifiers reduced false-positive refusals on legitimate biology questions, such as those from medical students, by 85%, and cut unnecessary interventions in cybersecurity contexts by roughly 60%, while now explicitly permitting defensive vulnerability research that earlier versions sometimes blocked. To address a separate regulatory requirement, Anthropic added an invisible watermark to Fable 5.1 outputs and opened a detection API in private preview, a step aimed at compliance with the European Union's AI Act, which mandates that certain AI-generated content be machine-detectable.
The release lands amid an increasingly public argument over how much oversight AI systems need as they gain autonomy — a debate playing out this week between Washington and Brussels over the shape of AI regulation. Anthropic's own framing of Mythos 5.1's alignment improvements, and its choice to gate the model behind government-linked verification programs rather than open access, suggests the company is trying to signal caution to regulators even as it pushes capability forward. Independent research firm Vellum and industry newsletter The Neuron both flagged the Millennium debugging anecdote and the Ramp unsupervised run as evidence that frontier labs are now competing less on raw benchmark scores and more on how long and how reliably a model can work without a human checking in.
For Anthropic, Fable 5.1 is as much a pricing move as a capability release. Rivals including OpenAI and Google have spent much of the past year pushing down the cost of running agentic workloads at scale, and cache-read pricing in particular has become a proxy battleground because it determines the real-world cost of the long, context-heavy sessions that coding and research agents rely on. Whether the discount holds up against competing releases from OpenAI and Google DeepMind, both of which are expected to update their own flagship models in the coming weeks, will likely determine how much of Fable 5.1's early enterprise enthusiasm translates into a lasting shift in market share.