前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
返回 AI 情报前线
All News · 全部资讯8614
  • TradingAgents v0.3.1:多Agent金融交易框架更新
  • 编程Agent每次会话都在重复发现你的代码库
  • AI编程助手在给自己的作业打分
  • H3-Metal:Apple Silicon 原生多模态 AI 推理引擎
  • Qwen 3.8-27b本周发布
  • Cline五个月实测:免费VS Code AI代理的真实成本
  • Muse开源模型发布:单卡RTX 3090可运行
  • Claude 刷新黎曼猜想下界记录,新模型身份未公开
  • 本地跑通电影实时配音工具LiveDub,Whisper+LM Studio+TTS全链路
  • OpenAPI上下文批处理:LLM Token消耗降低81.7%
  • 设计一个不会 hallucinate 的AI Agent监督者
  • 本地AI Agent时代到来:Meta Muse Glimmer可在消费级硬件运行
  • 将OpenAPI规范转换为MCP工具的实践指南
  • Paperclip:开源AI Agent团队编排平台
  • Anthropic 为 Claude 生成文本添加隐形水印
  • Celery+Redis异步任务队列设计:压缩任务实战
  • AI Agent工作区架构设计指南
  • 浏览器Agent安全加固实战:SSRF防护与内存治理
  • MiniMax-H3在ComfyUI的W4A8量化与VAE加速实战
  • Claude Code自动模式首发:10个可调用的x402付费API实战
  • DeepSeek流量超Google,token成本降13.6%
  • Zapier详解:如何把Claude接入各种应用
  • SpaceX 以 600 亿美元收购 AI 编程独角兽 Cursor,品牌将逐步淘汰
  • Claude全量嵌入「隐形水印」,所有输出文字均可溯源
  • MCP 服务器 GitHub Token 权限安全:只读工具也有写风险
  • Spring Boot 生产级 AI Agent 实战:金丝雀发布、模型回退与成本上限
  • 写代码前先写契约:用TLA+形式化验证需求
  • Qarinah:编程Agent上下文压缩方案
  • K3模型MoE稀疏性:896专家仅激活1.8%
  • AI 自动生成 Git patch 修复 SAST 告警,告别逐条手工排查
  • Anthropic 将 Claude Code 自动模式设为 Pro 用户默认
  • 一个HTML文件搞定24个AI编程prompt库:零后端改动的设计思路
  • Coding Agent 不听 AGENTS.md 的 7 个原因
  • Google Cloud发布agents-cli:自改进Agent循环的7条工程法则
  • Claude Sonnet 5 取消涨价维持首发价,API成本预期更稳定
  • Text-to-SQL演示只要一天,产品化那90%才是真正的工程挑战
  • OpenAI发布GPT-5.6-Cyber网络安全专用模型,漏洞研究能力大幅提升
  • Meta开源本地AI智能体模型Muse Glimmer
  • MCP协议:AI应用的USB-C接口
  • 两款并行编程Agent工具对比:Cezar vs Sculptor
  • OpenAI 推出专用于网络攻防的 GPT-5.6-Cyber,仅限 Red 级合作客户
  • AI SDK新增Grok Build支持:统一接口切换AI编程引擎
  • AI 沙箱安全:计算隔离不等于网络隔离
  • OpenAI 推出网络安全专项模型并扩大防御计划
  • Meta 发布 30B 参数开源模型 Muse Glimmer,Apache 2.0 许可
  • AI 编程工具的可靠性危机:光鲜表象下的技术债
  • 不做规格说明书直接上AI:一次跨租户安全漏洞的教训
  • 一次错误的 bug 修复:幂等性 key 的设计陷阱
  • LLM 分布式训练实战:CUDA/ROCm 多卡并行优化详解
  • 我如何重写错误提示,直到陌生人都能看懂
  • 14MB 二进制跑 Agent 任务:45M 参数 Needle 2 深度解析
  • 已加载 51 / 8614
8.0
热点
AI SCORE
行业动态2026-08-11 12:00

DeepSeek流量超Google,token成本降13.6%

Vercel Blog#DeepSeek#Moonshot#成本
Editor brief · 编辑速览

Vercel AI Gateway数据显示Moonshot K3高速增长,DeepSeek开源份额在7月翻倍至8.6%。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

AI Gateway Production Index — August 2026

Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from May, June, and July.

August 2026 Summary

The August index reports on AI Gateway data collected through July 2026.

The average price paid per token fell 13.6% in July, after rising almost 20% in May and holding steady in June. Token consumption grew so quickly that even with 37% growth in spend, cost per token saw a double-digit drop.

The average price paid per token fell 13.6% in July, after rising almost 20% in May and holding steady in June. Token consumption grew so quickly that even with 37% growth in spend, cost per token saw a double-digit drop.

DeepSeek became the second-largest lab by token volume, now running more than twice Google's volume.

DeepSeek became the second-largest lab by token volume, now running more than twice Google's volume.

Anthropic collected 65% of gateway spending on 30% of token volume, at 4.4 times the average price of every other lab's tokens.

Anthropic collected 65% of gateway spending on 30% of token volume, at 4.4 times the average price of every other lab's tokens.

Both media leaderboards changed hands. Google's Nano Banana took the lead in image volume from OpenAI's GPT Image, and ByteDance's Seedance led video in both volume and dollars.

Both media leaderboards changed hands. Google's Nano Banana took the lead in image volume from OpenAI's GPT Image, and ByteDance's Seedance led video in both volume and dollars.

Kimi K3 launched into agent work

Moonshot released Kimi K3 on July 16. Like Z.ai's GLM 5.2 released in June, it is built for long-horizon agent work, so its usage was heavy from the start at about twelve times the tokens per request of its predecessor, K2.5.

It scaled quickly. K3's daily volume tripled between launch week and the final week of July, and by month end its requests were as heavy as Claude Opus 4.8's. On the last full day of July it ranked eighth on the gateway by token volume.

The demand was new rather than diverted. K3 processed nearly two-thirds of all Kimi tokens within two weeks and 82% by the final week, while the rest of the family's volume fell only slightly.

Open weight's share of gateway spend more than doubled in July to 8.6%. More than 90% of that growth is Moonshot and Z.ai. Moonshot's share of total gateway spend quadrupled, to 2.3%. Cheap open-weight models have been taking volume for months without taking revenue. Kimi K3 and GLM 5.2 are the first to capture significant volume at more than eleven times DeepSeek's rate per token.

DeepSeek is now second by volume

In June, this report said DeepSeek had entered the fight for token volume. Last month, we said an open-weight lab would soon be second by volume. In July, DeepSeek surpassed Google to take that place.

Google ran nearly 40% of the gateway's token volume in April, and DeepSeek less than 1%. By July, DeepSeek ran a quarter, more than twice Google's 11%. DeepSeek V4 Flash, its cheapest model, ran more tokens alone than all of Google.

The reversal is concentrated in consumer-facing work. Google's share of personal-assistant tokens fell by more than half in a month while DeepSeek's more than tripled, and most of the token share Google gave up went to DeepSeek. DeepSeek's V4 Flash ran more tokens than any other model on the gateway in July, nearly a fifth of the total and 70% more than the next model.

Open-weight models previously owned the cheap end of the market. Their share of gateway token volume nearly tripled between April and June, from 11% to 29%, while share of spend stayed under four cents of every gateway dollar.

In July, volume continued its growth trend, increasing to 36%. Spend broke its trend, more than doubling to nearly nine cents of every dollar, the highest in the index's history.

The four largest frontier labs' combined share of token spend, which had not fallen below 93% in seven months, fell to 89%. Google accounted for most of the decline. Almost none of the spend it lost went to DeepSeek, whose share barely moved even as its volume grew. Z.ai and Moonshot's latest models account for the entire open-weight spend increase.

Buyers sort models by what the task requires, then choose inside that tier according to price. DeepSeek and Opus were never competing. GLM 5.2 and Kimi K3, both released in the last two months, are the first open-weight models running a meaningful share of the workloads historically owned by closed-weight labs.

Average price per token fell 13.6%

Companies bought more inference in July and ran more of it on cheap models. Volume grew 59%, spend grew 37%, and the average price paid per token fell 13.6%.

Holding June's mix of models constant, the average price would have held essentially flat instead of declining. OpenAI is one example: the average cost of an OpenAI token fell to 58% of its June level, because 85% of the volume it added went to GPT-5-Nano, the cheapest model in the GPT-5 family.

The entire decline in average price came from what companies chose to route. Among the thousands of teams that ran more than 10 million tokens in both months, three in four changed at least a tenth of their model mix. Three in five changed at least a quarter.

The median team's cost per token fell 2.9%, but only one team in six stayed within five percent of where it started.

A quarter pared cost by more than 30%, typically teams that entered the month paying well above the gateway average per token. Another quarter paid at least 20% more, typically teams that had been paying below the average. Concurrently, the expensive end shifted down and the cheap end migrated up.

81% of July's tokens ran on models that were not on the gateway six months ago. Open-weight models crossed a third of all volume, at about a seventh of frontier rates. Google fell from 24.0% of gateway token volume to 10.7%. Anthropic slipped two points to 29.8%, and OpenAI's share rose to 12.8%.

Gateway spend is up 74% since May and volume has doubled, with July's token growth running at nearly double June's pace. Spend on inference continues to increase even as each dollar buys more tokens than it did in the previous month.

Anthropic's premium widened as prices fell

Anthropic has held more than 60% of gateway spend in every month we have measured, through a period in which open weight tripled its share of volume. In July, it collected 65.1% of all spend on 30% of total volume.

The average price per Anthropic token ran 4.4 times the average across every other lab, up from 3.4 in June. Part of that rise was Claude Fable 5, which returned on July 1 after a three-week export-control suspension and quickly grew to 13.2% of all gateway spend, second only to Opus 4.8. In coding agents, the gateway's largest use case by tokens, Anthropic collected more than 80% of spend.

Anthropic has no model at the bottom of the market. Haiku 4.5, its least expensive, runs just over two-thirds of the gateway average price per token. OpenAI's GPT-5-Nano runs a sixth; DeepSeek's V4 Flash a sixteenth.

Customers cutting costs on OpenAI or Google can stay in the catalog. Cutting costs on Anthropic means leaving Anthropic, and in July it still took a majority of all spending. Anthropic's premium is the gap between what the cheap tier can do and what its customers need done.

Switching models on the AI Gateway is a one-line change, so nothing locks a customer in. When Fable 5 came back, daily volume returned to the pre-ban level almost exactly, but nine in ten of the teams running it in July had not used it before. The work returned; the customers were different.

The rest of the market's tokens averaged less than a quarter of Anthropic's price, and Anthropic still took two of every three dollars spent on the gateway. Price competition is happening, but it's happening in the part of the market with the least money in it.

Google took images, ByteDance took video

In June, OpenAI's GPT Image generated most of the gateway's images. In July, Google's Nano Banana led, at 45% of images to GPT Image's 42%, with spend split almost evenly between them. Nearly all of Nano Banana's gain came from one model, Gemini 3.1 Flash Lite Image. Google took the image lead in the same month it fell to fourth by token volume.

ByteDance's Seedance led video on both counts, in videos generated and dollars spent. xAI's Grok Imagine, June's volume leader, fell to second. Chinese labs collected about seven of every ten dollars spent on video.

Also in July's data

Within coding, the gateway's largest workload by tokens, DeepSeek ran nearly a third of the volume and Anthropic collected more than four of every five dollars.

Within coding, the gateway's largest workload by tokens, DeepSeek ran nearly a third of the volume and Anthropic collected more than four of every five dollars.

Back-office agents remain the most expensive work per token, with a share of spending roughly two and a half times their share of volume.

Back-office agents remain the most expensive work per token, with a share of spending roughly two and a half times their share of volume.

About this report

This analysis is based on anonymized, aggregate routing data from the Vercel AI Gateway through July 2026.

A few notes on measurement:

Volume counts all tokens routed through AI Gateway.

Volume counts all tokens routed through AI Gateway.

Spend values every request at the lab's published list price.

Spend values every request at the lab's published list price.

Price per token is calculated as total spend divided by total volume.

Price per token is calculated as total spend divided by total volume.

Open-weight volume and spend shares count the four open-weight labs serving their own models at scale (DeepSeek, MiniMax, Moonshot, and Z.ai).

Open-weight volume and spend shares count the four open-weight labs serving their own models at scale (DeepSeek, MiniMax, Moonshot, and Z.ai).

Image and video figures count media generated, not requests or tokens.

Image and video figures count media generated, not requests or tokens.

All figures use the most recent data available; prior months may be revised as methodology is updated.

All figures use the most recent data available; prior months may be revised as methodology is updated.

Jerilyn Zheng, Eric Dodds

Original source

本文由 AI 翻译整理自 Vercel Blog,原文版权归原作者所有。

阅读英文原文
上一篇
Claude Code自动模式首发:10个可调用的x402付费API实战
下一篇
Zapier详解:如何把Claude接入各种应用