前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
返回 AI 情报前线
All News · 全部资讯9321
  • Axonius多租户AI Agent隔离方案实践
  • 用LiteLLM将Claude Code路由到DeepSeek节省成本
  • Claude Code支持AGENTS.md跨工具标准配置
  • Coding Agent 账单省 50-70%:利用 sticky routing 保住 Prompt Cache 命中
  • AI Agent 为何还在用 while(true) 循环——工程陷阱深度剖析
  • SEO Agent 选 MCP 还是 REST?一份实用决策框架
  • SGLang深度解析:如何高效服务DeepSeek-V4-Pro
  • FlakeFixer: 用Agent自动分析Flaky Test
  • AI编码Agent记忆系统设计的四个教训:删除不是过期
  • 盲人开发者为视障群体打造AI描述应用ScribeMe
  • AI代码审查员的验证悖论:声称完成≠真正完成
  • LLM API多租户安全清单:tenants-safety essential
  • AI Agent不应持有你的钥匙:权限最小化原则
  • Claude 多智能体系统上演自复制恶意软件攻防战
  • MCP Server 开发避坑指南:工具描述比 TypeScript 更难
  • 生产级 Solana Agent 交易生命周期深度解析
  • 5分钟让AI助手读懂你的代码库
  • 用Python构建AI简历筛选器
  • Agent上下文满了该丢什么:长对话记忆管理实战
  • Warp推出Factories:一站式AI软件开发工厂基础设施
  • AI Agent试点到生产:成本暴涨700倍的教训
  • Cursor发布Origin功能:AI编程上下文管理
  • TryHackMe 提示词注入 CTF 实战攻略
  • 面向 Agent 的运维队列:失败自动转Ticket
  • 四个静默失败的 CI 检查:它们都是绿的,但什么都没做
  • AI 编码工具会读取 .env:本地 DLP 代理 Anonmyz 在prompt边界截流
  • Google 开源 SAM:零配置的 AI Agent P2P 发现与调用网络
  • Cursor Skills完全指南:格式规范与跨Agent迁移实测
  • OpenAI Codex Skills规范详解:目录结构与官方文档未记载的细节
  • 用Gitea自建Claude Code内部插件市场,团队Skill统一分发
  • Claude Skills规范深度解读:从格式到团队协作
  • 2026年LLM应用架构实战:摆脱if/else链式判断
  • Anthropic CEO:AI天然趋向集中,开源只是转移权力
  • 微软 Copilot 隐藏参数漏洞可被利用窃取密码
  • LangChain 揭示:Agent 效果不佳时换模型是误区,换 Harness 才是关键
  • 三阶段工作流让 AI Agent 保持精准:Research-Plan-Implement
  • 小米MiMo桌面应用即将上线,AI编程助手6月已开源
  • Agent成熟度记分卡:追踪AI Agent可靠性的五个核心维度
  • 模型趋同时代:系统架构比选模型更重要
  • 代码审核功能默认关闭的教训
  • 程序员用 Claude 为 Windows 专用 HP 打印机编写 macOS 驱动
  • NIST AI风险管理框架生产级RAG实战
  • Codex自动探索优化技能:研究-测试-评判闭环
  • AI协作UML编辑器:代码臭味一目了然
  • AI API成本降低95%的实战经验
  • 不同模型Tokenizer成本差异的技术解析
  • 我用低价模型替代OpenAI:生产迁移实录
  • Anthropic单token成本是Vercel均值4.4倍,使用量却占65%
  • career-ops: 在AI编程CLI里做求职管理
  • CI守护失效实录:4个静默失败的检查
  • 不同分词器对中文token计数差异高达20%
  • 已加载 51 / 9321
8.0
热点
AI SCORE
编程提效2026-08-18 22:00

AI Agent试点到生产:成本暴涨700倍的教训

dev.to · AI#AI Agent#成本分析#企业AI
Editor brief · 编辑速览

企业试点阶段月成本约1500美元,全量上线后达百万美元,差距700倍。根源是Agent工作流消耗的推理量是简单聊天的5-30倍,Gartner数据支撑。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

One documented enterprise deployment ran its proof of concept for about $1,500 a month in API usage. The results looked strong. Leadership approved full production, and the monthly bill at real-world volume landed just over $1 million. That is a 700X jump from pilot to production, and no business case survives a multiplier like that.

The case comes from an analysis of enterprise LLM deployments, and while the number represents a worst case, the mechanics behind it are ordinary. At pilot scale, a cost of $0.10 to $0.50 per agent request is easy to absorb and even easier to present as a savings story. At 10,000 users, the same request rate produces a monthly infrastructure bill that makes the original spreadsheet unrecognizable. Most teams run those numbers after the architecture decision is locked, which happens to be the most expensive possible time to learn them.

The Multiplier Hiding Inside Every Agent Request

A chatbot query triggers one inference call. An AI agent working through a task plans, calls tools, evaluates results, and loops back when something fails. Gartner's analysis from earlier this year found that agentic workflows consume between 5 and 30 times more tokens per task than a standard chatbot, with a single user request often triggering 10 to 20 separate model calls behind the scenes.

Four mechanics drive that multiplier.

Reasoning loops sit at the center. Every pass through plan, act, and evaluate fires at least one model call, and complex tasks can take dozens of passes before the agent settles on an answer.

Context accumulates. Agents carry system prompts, tool definitions, and step history into every subsequent call. All of it gets re-sent each time, so the token cost of step twelve includes the freight of steps one through eleven.

Tool calls stack their own costs on top. Web searches, database queries, and code execution each add latency and expense beyond the model call that triggered them.

Retries compound everything above. When a tool returns an unexpected schema or an output fails validation, the agent tries again. Each retry is a fresh trip through the loop, and you pay for the attempt whether the task succeeds or fails.

One documented example makes the point better than any abstraction. A coding agent assigned to fix a one-character typo in a README consumed over 21,000 input tokens listing issues, branching, committing, and opening a pull request. A trivial fix, wrapped in an expensive workflow.

Cheaper Tokens, Bigger Bills

Per-token pricing has collapsed. Inference for a GPT-3.5-level model fell from $20 per million tokens in late 2022 to $0.07 by October 2024, roughly a 280x drop in two years, and Gartner projects inference on trillion-parameter models will cost 90 percent less by 2030. Enterprise AI bills keep rising anyway, because total token consumption is growing faster than prices are falling. More capable agents run more reasoning loops, call more tools, and burn more tokens per completed task. Capability and cost move together by design, since quality in these systems comes from iteration rather than single-pass generation.

Uber ran into the same problem at scale. The company rolled out agentic coding tools to roughly 5,000 engineers, and heavy users racked up between $500 and $2,000 per month each, burning through the annual AI budget in about four months. The pilot had only ever tested one engineer, and nobody had modeled what concurrency at that scale would cost.

Run the Production Math Before the Architecture Locks

The forecast that prevents all of this takes about an afternoon to build. Start with cost per completed task from your pilot data, repriced at full production rates rather than free-tier or discounted credits. Multiply by the ratio of production users to pilot users. Then apply a burstiness factor of 3 to 5x, because production traffic spikes and runs in parallel in ways a pilot never exercises. If the resulting number breaks the business case, a spreadsheet is a far cheaper place to find out than an invoice.

Cost per successful task is the metric worth anchoring on. Cost per prompt and cost per session both hide failure. An agent that completes tasks cheaply but fails half the time and requires human cleanup costs far more than its dashboard suggests, and that gap stays invisible until you measure completion rather than activity.

In my work with client teams, the forecast conversation almost never happens at this stage. The pilot generates momentum, the demo impresses the steering committee, and the architecture gets approved on pilot economics. Everything downstream inherits that assumption.

Decide Which Steps Actually Need an Agentic Loop

Autonomous reasoning is the most expensive pattern in the stack, and most workflows only need it in a few places. Research on enterprise deployments suggests small language models can handle 60 to 80 percent of agent tasks at 10 to 30 times lower inference cost, with frontier models reserved for the steps that require genuinely complex reasoning.

A routing layer that classifies each step and sends classification, formatting, and retrieval work to smaller models can cut costs by 60 percent or more without touching quality where it matters. Caching does similar work on the input side. If the agent starts every task with the same system prompt and knowledge base, prompt caching can reduce input costs by roughly 90 percent. Retry caps close the remaining leak. Classify errors so some warrant a retry, some escalate to a human, and some fail gracefully before the loop becomes a line item.

The work is unglamorous. Walk the workflow step by step and decide where iteration earns its cost.

Build Cost Governance Before the Quarterly Surprise

Gartner predicts that more than 40 percent of agentic AI projects will be cancelled by the end of 2027, with escalating costs and unreliable outputs as leading causes. Most of those cancellations will trace back to the same sequence. The pilot succeeds, the deployment scales, the bill arrives, and the cost conversation happens under pressure with a CFO reading line items aloud.

Governance moves that conversation earlier, where it costs almost nothing. Track spend per task, per step, and per tool call, since aggregate metrics hide the one workflow that's eating the budget. Set alerts on cost per successful task so anomalies surface in days instead of at quarter close. And assign an owner, because a cost that's technically everyone's job doesn't get caught by anyone.


Nick Talwar is a CTO, ex-Microsoft, and a hands-on AI engineer who supports executives in navigating AI adoption. He shares insights on AI-first strategies to drive bottom-line impact.

→ Follow him on LinkedIn to catch his latest thoughts.

→ Subscribe to his free Substack for in-depth articles delivered straight to your inbox.

→ Watch the live session to see how leaders in highly regulated industries leverage AI to cut manual work and drive ROI.

Original source

本文由 AI 翻译整理自 dev.to · AI,原文版权归原作者所有。

阅读英文原文
上一篇
Warp推出Factories:一站式AI软件开发工厂基础设施
下一篇
Cursor发布Origin功能:AI编程上下文管理