Anthropic 推出两款新模型:Fable 偏向创意写作与叙事,Mythos 偏向结构化推理,各有专项优化,为程序员提供了更多模型选择。
我们推出 Claude Fable 5.1 和 Claude Mythos 5.1。它们是全球最先进的编程与知识工作模型——其研究能力也让人们得以初步窥见 AI 模型将如何推动科学进步。
Claude Fable 5.1 和 Claude Mythos 5.1 是同一模型,但具有不同级别的护栏。Fable 5.1 已全面开放,而 Mythos 5.1 仅通过可信访问计划提供;其护栏专门为网络安全和生命科学领域的工作而设计。
除了能力提升,Fable 5.1 还在价格、数据保留和护栏方面根据客户反馈迈出了重要步伐。
价格。Fable 5.1 在按 token 计费的工作负载上,预计比 Fable 5 降低约 25% 的成本。这是因为我们下调了缓存读取的定价(缓存读取是指模型读取已处理和存储的输入)。对于高度智能体化的工作,节省通常会更大——最高可达约 45%。
数据保留。我们的企业前沿护栏(Enterprise Frontier Safeguards,EFS)新系统,在防止对抗性使用的领域仍处于领先水平的同时,为客户提供了完全的隐私保护(等同于零数据保留策略)。EFS 的工作原理是将数据存储在完全由客户控制的云基础设施中,而非 Anthropic 控制。它将分阶段向企业客户提供,最早于今年秋季开始可用。在 EFS 可用之前,符合条件的客户将能够使用零数据保留的 Fable 5.1。
护栏。我们改进了护栏以减少误报(系统将良性内容标记的情况)。在网络安全领域,我们最新的护栏比之前减少了 60% 的误报。这在一定程度上是因为 Fable 5.1 现在可以用于发现软件漏洞——虽然不能用于开发漏洞利用程序。在生物学领域,我们建立了一个访问计划,该计划与美国政府合作开发,以使科学家能够使用 Claude Mythos 5.1 先进的生物学能力。我们预计将很快向科学家开放注册。
Claude Fable 5.1 为编程、知识工作和长时间运行的问题解决任务设立了新标准。下图显示,Fable 5.1 的性能远高于其前身 Fable 5。而在低或中等努力级别下,Fable 5.1 以低得多的成本实现了与 Fable 5 相当或更好的结果。(注意,Fable 5.1 在 Claude Code 中默认为高努力级别,在 Claude Cowork 和 Claude.ai 上默认为中努力级别。)
Terminal-Bench-Science 0.1:标准误差为每个模型 ±3.5–4.5 分。公开排行榜(3 次尝试/任务,Claude Code 测试框架)报告 Claude Opus 5 为 30.0%,Claude Fable 5 为 21.4%;我们的测试环境复现结果分别为 29.0% 和 24.7%,均在噪声范围内。
Terminal-Bench 4.0:按成本(对数刻度)和各努力级别显示分数。Claude Fable 5.1 和 Claude Mythos 5.1 是相同的底层模型;它们之间的差距反映了我们早期不太精确的网络安全护栏介入的任务。有了今天我们对这些护栏的改进,我们预计两个模型之间的差异将大大缩小。
Humanity's Last Exam:按成本(对数刻度)和各努力级别显示分数。
CursorBench 3.2.0:按成本(对数刻度)和各努力级别显示分数。
Fable 5.1 避免了导致工作质量下降的捷径,而且它足够智能,能够修复软件问题的根本原因。例如,在投资公司 Millennium 的测试中,Fable 5.1 找到了其内部系统中一个罕见崩溃的原因——这个崩溃在多年尝试后,没有任何工程师(或任何其他模型)能够解释。
以下是你可以看到 Fable 5.1 在各种基准测试中的对比:
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 | |
|---|---|---|---|---|
| Agentic scientific research Terminal-Bench-Science 0.1 [1] | 52.6% | 24.7% | 29.0% | 22.4% |
| Agentic coding Terminal-Bench 4.0 | 55.8%(Mythos 5.1 为 60.9%) | 42.0% | 52.3% | 37.3% |
| Knowledge work GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| Computer use OSWorld 2.0 [2] | 77.9%(部分) | 72.9%(部分) | 75.4%(部分) | —(部分) |
| Computer use OSWorld 2.0 | 41.7%(严格) | 36.1%(严格) | 39.6%(严格) | —(严格) |
| Multidisciplinary reasoning Humanity's Last Exam | 60.9%(无工具)/ 65.0%(有工具) | 57.8%(无工具)/ 63.8%(有工具) | 56.6%(无工具)/ 63.6%(有工具) | —(无工具)/ —(有工具) |
| Business workflows AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| Agentic coding CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Fable 5.1 在启用生产护栏的情况下进行评估。在这些护栏介入的任务中,Fable 5.1 和 Fable 5 在 OSWorld 2.0 上得分为零,Fable 5 在 AutomationBench 上得分为零。在我们护栏的所有其他介入中,网络安全任务由 Claude Opus 4.8 完成,生物学任务由 Claude Opus 5 完成。这可能会降低 Fable 5.1 和 Fable 5 在这些基准测试上的性能。
我们的早期访问合作伙伴注意到了这些性能提升,也注意到了模型输出中更具质的改进。以下是他们告诉我们的:
"在内部基准测试中,Claude Fable 5.1 比 Fable 5 或 Opus 5 解决了更多的编码问题,并在交易直觉上达到了最新水平。虽然之前的模型在工作时间越长就越难跟踪,但 Fable 5.1 在长时间的多步骤任务中仍然保持可读性。"
"在发布日,我们将把 Devin 中的 Opus 5 流量迁移到 Claude Fable 5.1。它在我们的测试中以更低的每任务成本与 Fable 5 持平或略胜一筹,而且凭借新的缓存读取定价,Fable 级模型终于对我们一直留在 Opus 上的工作负载变得经济了,从代码审查开始。"
"有一段代码存在极其罕见的崩溃,大约每百万次运行中出现一次,我们团队没有人能够在四到五年内解释它。我尝试过的每个模型,包括 Fable 5,都错过了。Claude Fable 5.1 是第一个找到它的。它反汇编了一个外部供应商库,将其与核心转储进行匹配,并将崩溃追踪到该库中的一个 bug。进行这种分析所需的时间很难证明是合理的。"
"Claude Fable 5.1 大约在三天内构建了一个复杂的原型。它跨我们所有服务的代码和文档进行了初步研究,产生了一个新颖且可扩展的设计。然后它无值守地运行了数小时,拥有强大的验证循环,实现了整个原型。早上醒来时,下一阶段已经完成,有完整的可视化演练,展示它构建的内容及其成功的清晰证据。"
"它是友好的 Fable。Fable 级的智能,Opus 级的价格,Sonnet 级的速度。在我们的测试中,它大约是 Opus 5 的两倍快,且使用的 token 数量减半,所以对于习惯使用 Opus 作为日常驱动的人来说,这是一个明显的升级。"
"在我们的研究套件上,Claude Fable 5.1 创下了新的最佳分数。在一项任务中,它提出了一个与我们在其他模型或过去人类研究人员身上从未见过的完全不同的轴线上的新颖解决方案,这将其结果提升到了之前的平台之上。它更擅长创造性解决问题,获得解决困难问题所需的那种顿悟。"
"作为我们对 AI 模型持续评估的一部分,Claude Fable 5.1 在我们的测试中提供了令人印象深刻的结果。使用 Claude Code,它在我们测试的所有努力级别上正确识别了每个破坏构建的根本原因。它也比之前的 Anthropic 模型沟通更有效,更新更简洁、更易于跟踪。"
"We asked Claude Fable 5.1 to review a clinical research project for Rakizen Medical that three other frontier models had signed off on. It found a gap none of them had seen and insisted on testing it further. It then proposed a completely new hypothesis, turning a dataset we had written off into a new research direction in one afternoon. It's the first time a frontier model like Claude has empowered us to explore new research in this way."
"Claude Fable 5.1 is very smart. On our 30-day simulated run-a-business eval, where the model gets full access to simulated Square tools, customers, employees, and vendors, it was far more efficient per token than Opus 5. We plan to use it to work through our most complex scenarios, the kind that used to take days of whiteboarding, so our engineering teams can keep moving fast."
"For anything research, greenfield or long-horizon, I would absolutely use Claude Fable 5.1 as the orchestrator. One unattended 38-hour run on a machine learning problem diagnosed a prior result as a label artifact, made the correction, kicked off six parallel experiments that ran overnight, and returned with a result and next steps. Given an open-ended prompt to find the highest-leverage problem nobody owned, it surfaced an unowned alert tied to a production outage, pulled the logs and prescribed the fix."
"The standout in Claude Fable 5.1 is the writing: more understandable, more meaningful, and it follows our writing guidance better. In blind tests against Fable 5, I preferred its writing and output. And in Canva Code it built a rhythm game with real music and on-beat gameplay matched to the level it generated, something no other model we tested delivered."
"On our PowerPoint eval, Claude Fable 5.1 produced the best decks of any model we've tested, both in slide craft and in fully answering our research topic. That same completeness showed up in our Citations eval, where it had the best fact recall over financial documents. And on complex, multi-part questions, it's the first model we've seen answer every part."
"We had a change that touched more than eight services across three codebases. Claude Fable 5.1 mapped the whole workflow end to end, in extremely fine detail, from the incoming service call down to the individual function and the database tables and rows, and it was accurate all the way down. We appreciated the opportunity to test the model and provide feedback, helping us prepare for a new frontier where we can increasingly rely on these tools to take on bolder initiatives."
"Across our evaluation sets, our judges preferred Claude Fable 5.1's answers roughly 2-to-1 over Fable 5 on everyday knowledge questions, high-intent research, and drafting and artifact work, all with improved response grounding over Fable. These results have made Fable 5.1 our go-to recommendation wherever Fable 5 was previously the choice."
"On the hardest problems we work on, Claude Fable 5.1 separates strongly from any other model we've tried. On a grand challenge-tier problem we used as a testbed for 18 months, it actually produced material progress. Rather than being trapped in stamp collecting, it made clear white-space connections I have yet to see elsewhere. It also optimized a compute kernel that Fable 5 had tapped out on by about 35%."
"On our hardest browser-agent benchmark, Claude Fable 5.1 completed 82% of tasks in about 10 minutes each, against 74% for Opus 5 and 57% for Fable 5, while using fewer tokens than either. It feels stronger than Fable 5 in every dimension we test. It also did exactly the right amount of work: never crossed a critical stop point across hundreds of measured tasks."
"On our internal Finance benchmark, Claude Fable 5.1 matches Fable 5 on accuracy while using 20% fewer tokens. We also saw big gains in slide generation, with Fable 5.1 showing improvements in explaining complex data in simpler English and translating it into more banker-grade visuals."
"Claude Fable 5.1 is more comfortable with long, unattended work than Fable 5. I've had workflows run for a long stretch without losing the plot: it keeps its own records, reprioritizes as things change, and picks up where it left off."
"Compared to Fable 5, Claude Fable 5.1 was a massive improvement on RedlineBench, our contract redlining benchmark, improving from 47.9 to 57.0. Most of the gain came on first-turn quality, where it doubled the previous score, with substantial gains on counterparty acceptance as well. Its edits were also more concise, with smaller changes on average to the documents."
"Claude Fable 5.1 is a leading model for our incident investigation evals, which use real production incidents to assess how effectively our agent, Bits Investigation, can produce root cause analyses. We evaluate our agent's output against root causes identified by our engineers. In these evaluations, it has demonstrated stronger reasoning than Opus 5 and has successfully diagnosed the most complex production incidents we've tested."
"Claude Fable 5.1 is the most capable model we've run on CursorBench 3.2, scoring 73.4% at max effort. We found it especially skilled at verifying its own work, allowing it to take on difficult coding tasks from start to finish."
"We evaluate models and systems on real-world investor workflows. On FrontierFinance, our latest finance benchmark, Claude Fable 5.1 shows a clear gain over Fable 5, with a 55.9% rubric score compared to 49.2%. The gain comes from its ability to dig harder into grounded, authoritative sources: on one earnings question, it went straight to the call transcript and captured the exact figures management cited, where other models leaned on secondary coverage."
We tested the scientific research capabilities of Claude Fable 5.1 and Claude Mythos 5.1 across a wide range of domains. What we found—which includes the early examples we share below—adds to the evidence that AI models will soon make important contributions to scientific discovery.
Molecular design. Many modern medicines work by binding to targets within the body to block, activate, or deliver something to them. High-affinity binders are necessary for drugs to work at lower doses; designing one is the first step in the development process for many common drug modalities. To see how well Claude Mythos 5.1 could do at this task, we gave the model access to open-source protein design and folding tools and sent its designs to two external organizations for experimental validation. Mythos 5.1 proved able to design very high-affinity binders. On three targets, [3] its binding affinities were 10 times higher than the best designs submitted to Adaptyv Bio's protein design competitions. Its hit rate (that is, the number of designs that were viable binders) was the strongest we've measured to date: it reached nearly 50% across 12 targets. (Hit rates of 10–15% are typical in protein design today.)
Claude-designed protein binders (orange) for each of 12 targets (grey). Every design in the video was confirmed to bind in the lab. Structures shown are ESMFold2 predictions.
Computational analysis and modeling. Claude Fable 5.1 trained a neural network to create a new, high-resolution elevation map of a third of the planet Venus. Its work was based on radar images taken by NASA's Magellan mission more than 30 years ago and a map that already existed for one-fifth of the planet. Claude's new map now reveals details down to two to three kilometers, rather than 10 to 20, and shows heights up to 25% more accurately than before.
We're releasing this map under a Creative Commons license in advance of upcoming NASA VERITAS and ESA EnVision missions, in hopes that it might help them determine which geologic features to target for future observation.

Altimetry 10-20km footprint

New DEM (300m) a volcano 15km across

A small shield volcano on Venus
Computational biology. In computational biology, it's common to run task-specific machine learning models on GPUs. The speed of these models is therefore a bottleneck to research progress. Mythos 5.1 provided one solution to this problem: by writing custom GPU kernels and caching their intermediate results, it sped up seven open-source deep learning models by up to 2.5 times (with identical outputs).
这类加速带来的收益会迅速累积。在任何一次典型实验中,生物学家可能需要运行这些模型数千次(例如,测试人类基因附近所有可能的突变)。在这类分析中,优化后的模型将预估 GPU 成本降低了 30%–60%。这种优化通常需要一支性能工程师团队花费数周才能完成,对于学术实验室来说往往负担不起。Mythos 5.1 仅用几天就完成了优化,而且仅使用了公开可用的源代码。我们计划不久后将这些优化开源。
Inference speedupInference speedup for seven open-source protein and genomics models on an NVIDIA H100
Inference speedup for seven open-source protein and genomics models on an NVIDIA H100
Estimated cost savings on genome-wide analysesOriginal implementationOptimizedEstimated GPU cost of three genome-wide analyses before and after optimization, at cloud list price. Evo 2 40B saves more on a whole job (2.3x) than per forward (1.4x) because some of its optimizations only pay off across many sequences.
Original implementation
Estimated GPU cost of three genome-wide analyses before and after optimization, at cloud list price. Evo 2 40B saves more on a whole job (2.3x) than per forward (1.4x) because some of its optimizations only pay off across many sequences.
随着我们模型科学能力的提升,我们对科学进步的投入也在不断增长。上周,我们预览了 Model Hardware Standard,该标准允许 Claude 直接且安全地操作实验室设备。我们最近还通过 AI for Science 计划扩大了对科学家的支持力度,该计划为从事高影响力科学项目的研究人员提供免费额度,并且我们通过全新的 Claude Team 计划为科学家提供大幅折扣的使用价格。
安全、安全与对齐
过去两年间,AI 模型的智能体能力已变得强大得多。但正如我们所记录的,更大的自主性也带来了新的风险。安全、安全与对齐方面的工作需要与 AI 能力同步推进。昨天,我们发布了一份报告,描述了我们如何改进自身的对齐和安全工作。
在发布 Claude Fable 5.1 和 Claude Mythos 5.1 之前,我们(以及在某些情况下由外部研究人员)对模型进行了广泛的风险测试,涵盖多个领域。我们在 System Card 中完整描述了这些努力;以下是简要总结。
化学与生物风险。我们测试了 Claude Mythos 5.1 在多大程度上能够帮助制造化学或生物武器。这涉及专家红队测试、自动化评估,以及一场将拥有 PhD 学位的生物学家与 AI 专家配对进行的桌面推演,以测试模型是否能够达到人类专家的水平。Mythos 5.1 的能力确实比 Mythos 5 更强。然而,我们的评估表明,它仍然低于我们的 Responsible Scaling Policy 中定义的下一风险等级。因此,我们将沿用对 Mythos 5 实施的相同安全防护措施来部署 Mythos 5.1,这些措施限制了对其生物研究能力的访问。
网络风险。我们运行了一系列评估来衡量 Claude Mythos 5.1(关闭网络安全护栏时)的网络能力。总体而言,该模型展现出我们发布的所有模型中最强的网络能力,不过它仍处于我们的 Frontier Compliance Framework 中的较低风险类别。我们还对 Fable 5.1 的网络安全护栏进行了大量压力测试:除了我们自己对这些护栏稳健性的动态评估外,我们还委托了两个组织进行外部测试,以及由 Gray Swan 进行的自动化测试。与 Fable 5 和 Opus 5 一样,我们没有发现针对这些护栏的关键级别越狱证据。
智能体安全。我们评估了 Claude Mythos 5.1 对恶意请求和提示注入(在 AI 模型处理的内容中隐藏的对抗性指令)的响应能力。它拒绝恶意智能体编码和计算机使用请求的比例与 Mythos 5、Sonnet 5 和 Opus 5 相当,在一项外部提示注入基准测试中,它是迄今为止我们最稳健的模型。
对齐。我们通过静态和交互式行为评估、使用自然语言自动编码器对其内部思维的分析、对齐相关能力评估、训练数据审查以及内部试点使用分析来测试模型行为。我们还收到了外部测试报告。
我们的自动化行为审计发现,Claude Mythos 5.1 在大多数指标上的对齐表现都优于其前身 Mythos 5。与 Mythos 5 相比,该模型在被分配一项本不可能完成的任务时,尝试访问测试环境之外资源的可能性明显更低。与 Mythos 5 相比,它也不太可能使用动机性推理来为其行为辩护(例如,通过推理认为当前情况是模拟或评估),也不太可能在追求用户目标时忽略明确的约束条件。从对训练数据的审查来看,Mythos 5.1 尝试奖励黑客行为(或作弊)的频率,以及成功实施奖励黑客行为的频率,都低于 Mythos 5 的总体水平。
虽然我们的对齐评估总体上显示出改进,但我们的测试也发现该模型有时仍可能绕过审批和自动模式分类器(如我们在 System Card 中更详细讨论的那样)。我们的对齐评估所提供的覆盖范围也存在局限性。目前,我们的自动化行为审计对超长上下文工作和多智能体环境的可见度较低。对于不可能完成的任务(这类任务可能引发更异常和对齐度更低的行为),我们的覆盖也不如我们所希望的那样全面,尽管我们最近在这一领域有所改进,并且正在努力继续推进。
我们还改进了护栏,使它们在不影响安全的前提下让模型更有用。我们将在下文中描述这些变化。
企业自动化护栏。Enterprise Frontier Safeguards (EFS) 允许我们检测并响应对我