Fireworks基于Moonshot开源权重模型Kimi K3构建的Ember-1,评测结果几乎一致但速度快3.4倍。
Fireworks Research 于 9 月 23 日以研究预览版的形式发布了 Ember-1。它基于 Moonshot 的开源模型 Kimi K3 构建,声称在质量相当的情况下减少了约 40% 的 token 使用量。Fireworks 表示它"学会了在保留关键思考的同时削减不必要的推理"。推理 token 按输出 token 计费,因此一个思考更少的模型理应成本更低。
但这里有个问题。在 OpenRouter 上,Ember 的输入 token 价格为每百万 3 美元,输出 token 价格为每百万 15 美元。Fireworks 对 Kimi K3 的定价相同,但其他提供商出售 Kimi 的价格低至输入每百万 1 美元、输出每百万 9 美元。
我通过 OpenRouter 在 Fireworks 上运行了两个模型,因此两者的计费标准都是 $3 和 $15,同时我也计算了 Kimi 在更低价格下的成本。
如果模型不准确,再便宜也没意义。我想知道 Ember 更短的思考过程是否经得起考验。如果它削减了真正需要的推理能力,准确率应该在最难的问题上首先下滑。因此我构建了逐步加难的测试,每个测试运行五次,与我近期测试中加入的一致性检查相同。
我通过 OpenRouter 调用两个模型,使用相同的提示词和默认推理设置。为了保证速度对比的公平性,我将两个模型都路由到 Fireworks。每次测试每个模型运行五次,我将推理 token 与其他输出分开记录。
逻辑谜题 —— 三道规模递增的谜题,分别有 4、5、7 位工程师,每道题各有一个唯一解。更大的谜题需要更长的推理链,因此削减思考能力应该首先在这里受到影响。
部署排程 —— 十二个有依赖关系的服务,每个团队同一时间只能部署一个,还有两个维护窗口。模型需要找到最快的部署方案:17 小时。
概率题 —— 五道关于重试系统的问题,服务器在健康和降级状态之间切换,外加一个熔断器。每道题的答案都是精确分数,我通过 200 万次请求的模拟进行了验证。
在两个模型接触任何一道题之前,我用两种独立方法确认了每道题的答案。如果你想在自己的系统上复现这些测试,我在文章末尾包含了所有提示词。
两个模型在全部五次运行中都解出了全部三道谜题。
Ember 平均每次使用 13,630 个推理 token,耗时 3 分 46 秒,成本 $0.27。Kimi 平均使用 16,679 个推理 token,耗时 12 分 26 秒,成本 $0.34。Ember 的推理 token 减少了 18%。尽管不在营销宣传中,Ember 比 Kimi 快了非常多。Kimi 最慢的一次运行接近 20 分钟。
两个模型每次都找到了 17 小时的排程,且每个排程都通过了我编写的校验器。
Ember 平均使用 6,543 个推理 token,耗时 1 分 29 秒,成本 $0.10。Kimi 平均使用 7,792 个推理 token,耗时 4 分 46 秒,成本 $0.13。推理 token 减少 16%,且速度明显更快。
这次测试产生了唯一的失误。Kimi 在每次运行中都正确回答了全部五道题。Ember 在五次中有四次全部正确。第五次时,它在第一道题上犯了一个小的算术错误,写成了 0.94619 而非 0.94629,这个错误还连带影响了另外两个答案。
这次测试也是节省最多的一次。Ember 平均使用 6,242 个推理 token,耗时 1 分 47 秒,成本 $0.13。Kimi 平均使用 9,682 个推理 token,耗时 6 分 48 秒,成本 $0.19。这是 Ember 首次接近其宣称的 40% token 削减。同样,Ember 明显更快。
| 测试(各 5 次运行) | Ember-1 | Kimi K3 |
|---|---|---|
| 逻辑谜题 | 5/5 完美,3:46,13,630 推理 / 17,766 总输出,$0.27 | 5/5 完美,12:26,16,679 推理 / 22,553 总输出,$0.34 |
| 部署排程 | 5/5 完美,1:29,6,543 推理 / 6,822 总输出,$0.10 | 5/5 完美,4:46,7,792 推理 / 8,365 总输出,$0.13 |
| 概率题 | 4/5 完美,1:47,6,242 推理 / 8,365 总输出,$0.13 | 5/5 完美,6:48,9,682 推理 / 12,381 总输出,$0.19 |
| 完美运行次数 | 14/15 | 15/15 |
| 两模型均在 Fireworks 上的总成本($3/$15) | $2.48 | $3.26 |
| Kimi 按最低价($1/$9)的总成本 | $2.48(无更低价提供商) | $1.96 |
Kimi K3 在 15 次运行中 15 次完美,Ember-1 为 14 次。Ember-1 唯一的一次失误是概率题上的算术错误;对我来说这不算太严重,毕竟是 15 次中的一次。
我在测试过程中注意到一个营销宣传中没有提到的事实:Ember-1 完成每组测试的速度是 Kimi K3 的 3.4 倍。至于减少推理 token 使用量,Ember-1 减少了 23%。这有助于降低成本。在 Fireworks 上,总成本为 $2.48,而 Kimi K3 为 $3.26,便宜了 24%。
有一点需要说明:Kimi K3 的成本取决于提供商。按其最低挂牌价输入每百万 $1、输出每百万 $9 计算,相同的运行将花费 $1.96,比 Ember-1 更低。
Ember-1 的准确率与 Kimi K3 大致相当,但推理 token 使用量少得多。在这个测试中 Ember-1 的成本优势并不那么重要,因为想用 Kimi K3 的用户可以通过 OpenRouter 路由到更便宜的提供商,把价格压到 Ember-1 以下。
关于 Kimi K3,我了解的一点是它很慢。又便宜又慢。不过我的速度数据来自 Fireworks 的标准端点,我没有测试更便宜的提供商。如果你有时间又想省钱,Kimi K3 赢了。如果你想以快得多的速度获得与 Kimi K3 几乎相同的结果,用 Ember-1。
每个提示词末尾都有固定的答案格式,以便自动评分。
Solve all three logic puzzles below. Each has exactly one solution.
PUZZLE SMALL: 4 engineers (Ava, Bo, Cleo, Dev) each own exactly one server. Each server has a rack position (1, 2, 3, 4, numbered left to right), an operating system (Debian, Alpine, Fedora, Ubuntu), a role (cache, queue, proxy, db). No two servers share any of these values.
If the server in rack 3 runs Fedora, then the Debian server is in rack 3.
The server in rack 1 is Dev's.
Exactly one of these is true: the server in rack 4 runs Ubuntu, or the Alpine server is not Bo's.
The db server is Dev's.
The queue server is not Ava's.
Bo and the Ubuntu server are in neighboring racks.
PUZZLE MEDIUM: 5 engineers (Ava, Bo, Cleo, Dev, Eun) each own exactly one server. Each server has a rack position (1, 2, 3, 4, 5, numbered left to right), an operating system (Debian, Alpine, Fedora, Ubuntu, Arch), a role (cache, queue, proxy, db, build), a replication data center (Oslo, Lima, Pune, Accra, Perth). No two servers share any of these values.
Exactly one of these is true: the cache server runs Arch, or the server replicated to Perth is the proxy.
The server in rack 5 runs Fedora.
Exactly one of these is true: the server in rack 2 is not the proxy, or Cleo runs Debian.
Exactly one of these is true: the server replicated to Pune is the db, or Dev is in rack 3.
The server in rack 1 is the db.
The proxy server is in rack 5.
The server in rack 4 runs Ubuntu.
The cache server and the server replicated to Pune are in neighboring racks.
The server replicated to Accra is the queue.
The server replicated to Lima is not in rack 3.
The queue server is not in rack 2.
The Fedora server is not Dev's.
The server in rack 2 is not replicated to Lima.
PUZZLE LARGE: 7 engineers (Ava, Bo, Cleo, Dev, Eun, Finn, Gia) each own exactly one server. Each server has a rack position (1, 2, 3, 4, 5, 6, 7, numbered left to right), an operating system (Debian, Alpine, Fedora, Ubuntu, Arch, Rocky, NixOS), a role (cache, queue, proxy, db, build, metrics, auth), a replication data center (Oslo, Lima, Pune, Accra, Perth, Quito, Riga). No two servers share any of these values.
The build server is Cleo's.
The proxy server does not run Fedora.
The server replicated to Quito is Finn's.
The server in rack 6 is replicated to Lima.
The Alpine server is exactly 4 racks to the right of the auth server.
The server replicated to Oslo is Dev's.
The Debian server is replicated to Quito.
The server replicated to Accra is not the build.
The Rocky server is exactly 1 rack to the right of Gia.
Exactly one of these is true: the Ubuntu server and the server replicated to Pune are in neighboring racks, or the server replicated to Oslo runs Fedora.
The server replicated to Oslo does not run Fedora.
Exactly one of these is true: the proxy server is somewhere to the left of the server replicated to Riga, or the Debian server is in rack 7.
The auth server and Bo are in neighboring racks.
The metrics server is Gia's.
The Fedora server is somewhere to the left of the server replicated to Pune.
Exactly one of these is true: the server replicated to Accra is in rack 2, or Finn is the queue.
If the Debian server is Finn's, then the server in rack 7 does not run Rocky.
The NixOS server is not replicated to Perth.
The server in rack 7 is replicated to Accra.
The server replicated to Pune is the db.
At the end of your response, give each solution as a block, one line per engineer, in the order the engineers are listed in that puzzle:
<name> | <rack> | <os> | <role>
<name> | <rack> | <os> | <role> | <dc>
<name> | <rack> | <os> | <role> | <dc>
You are planning a production rollout of 12 services. Time is measured in whole hours from hour 0.
Services (owning team, deploy duration in hours):
auth: team A, 3 hours
billing: team B, 4 hours
catalog: team C, 2 hours
search: team C, 3 hours
cart: team B, 2 hours
checkout: team B, 3 hours
payments: team A, 4 hours
notify: team C, 2 hours
ledger: team A, 2 hours
gateway: team A, 3 hours
reports: team B, 3 hours
inventory: team C, 4 hours
Dependencies. A service may start deploying only after every service it depends on has finished deploying:
gateway depends on auth
payments depends on auth
search depends on catalog
cart depends on catalog
cart depends on inventory
checkout depends on cart
checkout depends on payments
ledger depends on billing
ledger depends on payments
notify depends on checkout
reports depends on ledger
gateway depends on search
notify depends on gateway
Each team can deploy only one of its services at a time.
Different teams can deploy at the same time.
Blackout windows. No deploy may be in progress at any time during hours 9 to 11 or hours 17 to 19. A deploy may end exactly at hour 9 or 17, and may start exactly at hour 11 or 19. A deploy cannot pause and resume.
Each deploy runs from its start hour for its full duration without interruption.
What is the earliest hour by which all 12 services can be finished? Give a schedule that achieves it.
At the end of your response, give your answer in exactly this format:
<service>: <start hour>
(one line per service, all 12 services)
A client calls a payment API and retries on failure.
The server is in one of two states on each attempt, Healthy or Degraded.
On the first attempt, the server is Healthy with probability 4/5 and Degraded with probability 1/5.
Between consecutive attempts, the state changes like this: from Healthy, it stays Healthy with probability 3/4 and becomes Degraded with probability 1/4. From Degraded, it stays Degraded with probability 2/3 and becomes Healthy with probability 1/3.
An attempt succeeds with probability 9/10 if the server is Healthy and 2/5 if it is Degraded, independently of everything else given the state.
The client makes at most 4 attempts and stops as soon as one succeeds.
Circuit breaker: if two consecutive attempts both hit a Degraded server and both fail, the client stops immediately and makes no more attempts.
Answer these questions with exact fractions in lowest terms:
Q1. What is the probability that the request eventually succeeds?
Q2. What is the expected number of attempts the client makes?
Q3. Given that the request succeeds, what is the probability that it succeeded on exactly the second attempt?
Q4. What is the probability that the circuit breaker ends the request early, before the client has used all 4 attempts?
Q5. Given that the request succeeds, what is the probability that the first attempt failed?
At the end of your response, give exactly five lines in this format: