Artificial Analysis 数据显示 Mercury 2.5 在推理速度上取得突破,适合对延迟敏感的实时应用场景。
Released September 2026
Mercury 2.5 在智能水平上略低于平均,但与其同价位其他模型相比定价颇具竞争力。它的一大亮点是速度非常快,且输出相当简洁。该模型支持文本输入、文本输出,上下文窗口为 260k tokens。
Mercury 2.5 在 Artificial Analysis Intelligence Index 上得分为 12,在同类可比模型中处于中位以下( median: 13)。在该 Intelligence Index 的评估中,它生成了 35M 个输出 tokens,相比中位数 85M 相当简洁。
Mercury 2.5 的定价为:输入 $0.25 / 1M tokens(中等定价,中位数: $0.25),输出 $0.75 / 1M tokens(中等定价,中位数: $0.90)。在该 Intelligence Index 上评估 Mercury 2.5,单次任务平均成本为 $0.06。
Mercury 2.5 推理速度达 781 tokens per second,表现相当突出(112)。
Technical specifications
本页展示的是该模型的推理版本(reasoning variant)。
也可能存在一个非推理版本。
175 个同级别模型
指标在同类模型中进行对比:
非推理模型 → 仅与其他非推理模型对比
推理模型 → 跨推理和非推理模型对比
开源权重模型 → 仅与同规模级别的其他开源权重模型对比:
Small: 4B–40B 参数
Medium: 40B–150B 参数
Large: >150B 参数
专有模型 → 与同价格区间的专有和开源权重模型跨类对比,使用 3:1 输入/输出的混合价格比:
$0.15–$1 / 1M tokens
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 包含:AA-Briefcase v1.1、GDPval-AA v2.1、AutomationBench-AA、Terminal-Bench 4.0、SciCode、Humanity's Last Exam、GDP.pdf、CritPt、AA-Omniscience、AA-LCR v1.1。详见 Intelligence Index methodology,包含各评估项目的细分以及运行方式。
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 包含:AA-Briefcase v1.1、GDPval-AA v2.1、AutomationBench-AA、Terminal-Bench 4.0、SciCode、Humanity's Last Exam、GDP.pdf、CritPt、AA-Omniscience、AA-LCR v1.1。详见 Intelligence Index methodology,包含各评估项目的细分以及运行方式。
表示模型权重是否开放下载。如果商业使用受条件限制则标记为"Commercial Use Restricted",如果许可证禁止商业使用则标记为"Non-commercial"。
衡量模型在特定能力和行业上的表现
Artificial Analysis Finance & Accounting Index
Intelligence Evaluations
代理知识工作,(Elo-500)/2000
代理现实世界任务,(Elo-500)/2000
代理 SaaS 工作流
代理编程和终端使用
推理与知识
专业文档推理,All-pass
1 - 幻觉率
长上下文推理
法律代理工作,criterion pass rate
代理业务运营
电子表格和文档上的定量分析
Kubernetes 事故根因分析
医疗长上下文推理
Intelligence Evaluation Relevance
虽然模型智能通常可以跨用例泛化,但特定评估对某些用例可能更具参考价值。
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 包含:AA-Briefcase v1.1、GDPval-AA v2.1、AutomationBench-AA、Terminal-Bench 4.0、SciCode、Humanity's Last Exam、GDP.pdf、CritPt、AA-Omniscience、AA-LCR v1.1。详见 Intelligence Index methodology,包含各评估项目的细分以及运行方式。
AA-Briefcase v1.1 已更新
AA-Briefcase Elo 是一个综合指标,聚合了分析质量 Elo、呈现 Elo 和 rubric 通过率,rubric 表现通过合成一对一比赛转换为 Elo。Elo 及 95% 置信区间边界钳制在 0。
AA-Omniscience Index(越高越好)衡量知识可靠性和幻觉率。它奖励正确答案、惩罚幻觉,且对拒答不扣分。分数范围为 -100 到 100,其中 0 表示答对和答错数量相当,负分表示答错多于答对。
Intelligence Index Comparisons
Intelligence Index vs. 每 Intelligence Index 任务的成本
每 Intelligence Index 任务的成本
每个 Intelligence Index 任务的加权平均成本。各评估的成本由输入、缓存命中、缓存写入、推理和回答 token 价格除以任务数计算,再按其 Intelligence Index 权重加权。
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 包含:AA-Briefcase v1.1、GDPval-AA v2.1、AutomationBench-AA、Terminal-Bench 4.0、SciCode、Humanity's Last Exam、GDP.pdf、CritPt、AA-Omniscience、AA-LCR v1.1。详见 Intelligence Index methodology,包含各评估项目的细分以及运行方式。
每个 Intelligence Index 任务的输出 Tokens
每个 Intelligence Index 任务的输出 Tokens
每个 Intelligence Index 任务所需的 token 数。通过将每个评估的输出 tokens 乘以 Intelligence Index 中各基准的相对权重,再除以任务数(不含重复)计算得出。
每 Intelligence Index 任务的成本
每 Intelligence Index 任务的成本
每个 Intelligence Index 任务的加权平均成本。各评估的成本由输入、缓存命中、缓存写入、推理和回答 token 价格除以任务数计算,再按其 Intelligence Index 权重加权。
运行 Artificial Analysis Intelligence Index 的成本
运行 Artificial Analysis Intelligence Index 的成本
运行 Artificial Analysis Intelligence Index 中评估的成本,根据模型的输入、缓存命中、缓存写入、推理和回答 token 价格以及跨评估使用的 token 总数(不含重复)计算。
定价:缓存命中、输入和输出
缓存提示词(之前处理过的)的每 token 价格,通常比常规输入价格有显著折扣,以每百万 tokens 的美元价格表示。这里显示的是缓存命中价格;缓存写入和缓存存储单独计费,且因提供商而异——详见各提供商的缓存定价详情。
用于 RAG 的上下文窗口
更大的上下文窗口与 RAG(Retrieval Augmented Generation)LLM 工作流相关,这类工作流通常涉及对大量数据的推理和信息检索。
输入和输出 tokens 的合并最大数量。输出 tokens 通常有更低的独立上限(因模型而异)。
通过输出速度(tokens per second)测量
模型在生成 tokens 时的每秒接收 token 数(即模型支持流式输出的情况下,从 API 收到第一个 chunk 后开始计时)。
模型性能呈现
数据代表模型第一方 API 的性能,或在没有第一方 API 时跨提供商的中间值。
每个 Intelligence Index 任务的时间
每个 Intelligence Index 任务的时间
每个 Artificial Analysis Intelligence Index 任务的加权平均时间(秒)。通过将每个任务的输出 tokens 除以输出速度,再按 Intelligence Index 中各基准的相对权重加权计算得出。
通过首 token 时间(秒)测量
延迟:首 Answer Token 时间
首 Token 时间
发送 API 请求后收到第一个回答 token 的时间(秒)。对于推理模型,这包含模型在给出答案前的"思考"时间。对于不支持流式的模型,这表示收到完成内容的时间。
端到端响应时间
输出 500 tokens 所需的秒数,根据首 token 时间、推理模型的"思考"时间和输出速度计算
端到端响应时间
端到端响应时间
收到 500 token 响应所需的秒数。关键组成部分:
输入时间:收到第一个响应 token 的时间
思考时间(仅推理模型):推理模型在给出答案前输出 tokens 进行推理的时间。token 数量基于 60 个多样化提示词的平均推理 tokens(方法论详情)。
回答时间:基于输出速度生成 500 个输出 tokens 的时间
模型性能呈现
数据代表模型第一方 API 的性能,或在没有第一方 API 时跨提供商的中间值。
Frequently Asked Questions
关于 Mercury 2.5 的常见问题
Mercury 2.5 何时发布?
Mercury 2.5 于 2026 年 9 月 8 日发布。
Mercury 2.5 由谁创建?
Mercury 2.5 由 Inception 创建。
Mercury 2.5 有多智能?
Mercury 2.5 在 Artificial Analysis Intelligence Index 上得分为 12,在同价格区间其他推理模型中处于中位以下(median: 13)。
Mercury 2.5 有多快?
Mercury 2.5 的生成速度为 780.8 tokens per second(基于 Inception 的 API),在同价格区间其他推理模型中远高于平均(median: 111.8 t/s)。
Mercury 2.5 的延迟是多少?
Mercury 2.5 的首 token 时间(TTFT)为 2.91s(基于 Inception 的 API),在同价格区间其他推理模型中略高于平均(median: 2.23s)。
Mercury 2.5 的费用是多少?
Mercury 2.5 的输入价格为 $0.25 / 1M tokens(优于平均,median: $0.25),输出价格为 $0.75 / 1M tokens(优于平均,median: $0.90),基于 Inception 的 API。
Mercury 2.5 的 API 定价是多少?
Mercury 2.5 的输入价格为 $0.25 / 1M tokens,输出价格为 $0.75 / 1M tokens(基于 Inception 的 API)。混合费率(7:2:1 缓存命中/输入/输出比)为 $0.14 / 1M tokens。定价因提供商而异。比较 API 提供商定价
Mercury 2.5 的输出有多冗长?
在 Intelligence Index 上评估时,Mercury 2.5 生成了 35M 个输出 tokens,在同价格区间其他推理模型中优于平均(median: 85M)。
Mercury 2.5 是推理模型吗?
是的,Mercury 2.5 是推理模型。它使用扩展思考或链式思维推理来处理复杂问题,然后再给出答案。
Mercury 2.5 支持哪些输入模态?
Mercury 2.5 支持文本输入。
Mercury 2.5 支持哪些输出模态?
Mercury 2.5 支持文本输出。
Mercury 2.5 能处理图像吗?
不能,Mercury 2.5 不支持图像输入。它只能处理文本。
Mercury 2.5 是多模态的吗?
不是,Mercury 2.5 不是多模态的。它只支持文本输入。
Mercury 2.5 的上下文窗口是多少?
Mercury 2.5 的上下文窗口为 260k tokens。这决定了模型在单个请求中能处理多少文本和对话历史。
Mercury 2.5 是开源的吗?
不是,Mercury 2.5 是专有模型。模型权重不公开。
Mercury 2.5 有多少参数?
Mercury 2.5 是专有模型,Inception 未披露模型大小或参数数量。
Mercury 2.5 在基准测试上表现如何?
Mercury 2.5 在 Artificial Analysis Intelligence Index 上得分为 12。这个综合基准评估模型在推理、知识、数学和编程方面的表现。
Mercury 2.5 可以通过 API 调用吗?
可以,Mercury 2.5 可通过 1 家提供商访问 API。比较 API 提供商
在哪里可以使用 Mercury 2.5?
Mercury 2.5 可通过 1 家 API 提供商使用。比较提供商