前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
返回 AI 情报前线
All News · 全部资讯9087
  • AI编程增量改动如何悄然腐蚀系统架构
  • DeepSeek V4费用因时段和渠道差异可达2倍
  • MCP服务器供应链攻击警示:credential窃取实案
  • Claude Code 下一代功能:Mods、插件、项目与标签系统
  • 用 Gemini Live API 构建实时语音 AI 代理
  • 7 种生产环境 AI 代理沙箱隔离方案
  • Grok 4.7登陆AWS Bedrock:50万token上下文支持编程
  • Claude Sonnet 5.5 发布:提速 30% 降价 30%
  • Claude Sonnet 5.5 高风险请求会被静默降级至 Sonnet 5
  • OpenAI披露可像蠕虫传播的新型提示注入攻击
  • Salesforce Agentforce 曝零点击数据泄露漏洞
  • 多Agent系统中级联失败的自我修复方案
  • NVIDIA开源Agent安全平台:毫秒级隔离失控AI
  • OpenAI 安全负责人谈 AI 能力突变下的组织弹性
  • AWS正式上线Claude Sonnet 5.5
  • Claude Sonnet 5.5 上线 Google Cloud
  • 决策模型:只返回概率、拒绝生成文本的新范式
  • AI购物Agent遭评测攻击:恶意评论注入钓鱼
  • GPT-6 Astra百万token成本拆解:何时值得用
  • MCP 服务器悄然变更工具描述 = 你的 Agent 在执行未知指令
  • 生产级 AI 不能只靠模型版本管理
  • GitHub Copilot 已上线 Claude Sonnet 5.5
  • Claude Sonnet 5.5 编程 benchmark 暴涨 6 倍,成本降三成
  • Anthropic发布Sonnet 5.5:半价达到 Opus 水平
  • Claude Sonnet 5.5 发布:性价比翻倍
  • VoiceStudio: 开源全本地语音克隆,支持646语言和MCP
  • 企业级AI代理界面设计:让记忆成为第一公民
  • OpenAI代理绕过限制:通过DNS隧道实现隐蔽通信
  • IDE 不够用:构建 Agentic 开发环境 ADE
  • Agent 记住了修复方案却错了:AfterTrace 事件恢复工具
  • 代理工具调用经济学:告别靠猜
  • OpenAI代理利用Google安全游戏漏洞抓取UN数据
  • 小米 MiMo-V2.6 开源登顶 AA 指数,超越 Kimi K3
  • 代码生成自动化了,代码审查却没有:AI辅助PR的真相
  • Cloudflare推出cf CLI:Agent化的全API命令行工具
  • Nvidia在芯片层引入AI Agent看门狗:毫秒级隔离
  • pgEdge:AI 编程助手造的数据库分支无法合并,这才是设计原意
  • 月之暗面 Kimi K3.1 曝光:100 万 Token 上下文
  • Cloudflare 开源 Forge:自动生成 API SDK 和文档
  • 英伟达推 AI Agent 安全平台:毫秒级隔离失控 Agent
  • EmDash 1.0:面向Astro的安全开源CMS
  • Cloudflare Kitesurf浏览器升级:730K平台测试通过,支持WebMCP
  • Cloudflare 收购 VoidZero 后:JS 工具链提速 10 倍
  • RAG原理解析:从第一性原理出发
  • 16 岁少年用 AI 机器人 Antares 发现微软内部 API 漏洞:涉 17 万亿行数据
  • Nvidia推出Open Agent安全平台,管控越狱AI智能体
  • Holo4:通用计算机操作智能体的底层支撑
  • 英伟达发布 AI 智能体安全平台:毫秒级异常隔离
  • 用Termux在Android手机上搭AI开发服务器
  • NVIDIA开源OpenShell安全沙箱:给本地Agent施加真实运行时限制
  • OpenRig:用 YAML 定义多 Agent 团队,Claude Code 与 Codex 协同作战
  • 已加载 51 / 9087
8.0
热点
AI SCORE
技术实践2026-09-29 02:24

生产级 AI 不能只靠模型版本管理

Stack Overflow Blog#MLOps#AI运维#最佳实践
Editor brief · 编辑速览

Stack Overflow 博客指出生产 AI 需完整 MLOps 流程,涵盖评估、部署与回滚,版本号不够用。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

Imagine a documentation assistant that starts timing out after a routine retrieval change. The model is the same. The service is healthy. Every deployment check passed. But the retriever now sends more context, generation takes longer, and requests pile up behind a busy inference server. Reverting the application container does not help because the retrieval configuration lives elsewhere.

This is a hypothetical incident, but it exposes a useful design question: what, exactly, did the team deploy?

For an AI application, a model version is only part of the answer. Inputs, preprocessing, prompts, retrieval, tool contracts, and serving settings can change behavior independently. Reliable AI infrastructure needs a release boundary around the components that must work together. MLOps then becomes the workflow for testing that release, observing it under real traffic, and replacing it safely.

The practical starting point is not a bigger platform. It is a versioned release manifest, a meaningful evaluation gate, and a rollback path that has actually been tested.

Traditional ML systems already have this problem: a classifier trained with one feature transformation can behave incorrectly when production applies another. Google's MLOps guidance describes training-serving skew and the need for data and model validation within automated pipelines. Generative applications extend the set of dependencies; they do not remove the need to manage them. [1]

For a retrieval-augmented generation application, I would begin with a small manifest like this. The values below are illustrative identifiers, not a runnable platform specification.

release_id: docs-assistant-r17
app_revision: git-8f2a1c
model_revision: model-snapshot-42
prompt_revision: support-v7
retrieval:
  index_revision: docs-index-2026-09-01
  embedding_revision: embed-v3
  pipeline_revision: chunk-and-rank-v4
runtime_revision: serving-config-v6
evaluation_suite: support-regression-v12
previous_release: docs-assistant-r16

Each reference should resolve to retained, inspectable configuration or artifacts. The runtime revision should cover settings that affect execution, such as token limits, batching, timeouts, and resource placement. If the application calls tools, version the schemas and adapters too. Store references to secrets, never secret values, in the manifest.

A manifest is not a promise of bit-for-bit reproducibility. External services change; generation can remain nondeterministic; a provider may not offer immutable model snapshots. Record those limits instead of disguising a floating alias as a fixed version. Where data changes continuously, record the ingestion watermark and index configuration so an incident can be investigated even when replay is imperfect.

This does not require copying the entire environment for every release. It requires knowing which combination was tested, and resolving that combination consistently when a request begins. Updating four configuration stores one after another is not an atomic release.

Figure 1. Test and promote the release as a unit. Production failures feed the next evaluation suite.

Workflow diagram showing a versioned release tested by an evaluation gate, then promoted to live traffic, with live traffic failures feeding back as new evaluation tests.

A service can return HTTP 200 and still fail the user. For a documentation assistant, a useful answer might need to cite an accessible source, reflect the correct product version, and decline to invent instructions when evidence is missing. Those are different checks from endpoint availability.

Start with a compact, versioned dataset built around the tasks the feature is meant to complete. Include ordinary questions, previously observed failures, ambiguous requests, missing evidence, and attempts to cross authorization boundaries. Keep a held-out set so repeated prompt tuning does not turn the entire suite into a training target.

Use deterministic checks where possible: schema validity, allowed tool arguments, citation identifiers, and permission enforcement. For semantic judgments, define a rubric and compare automated scores with human review. A model-based judge can help prioritize review, but its output is not ground truth. Pin its configuration and investigate disagreements rather than averaging them away.

The gate should evaluate the full path, not just call the model with a prepared prompt. In the opening example, a model-only test would miss the change in retrieved context. Run the same release through retrieval, generation, and output validation, then inspect results by meaningful slices: long inputs, languages, product versions, and requests with sparse evidence.

Choose acceptance criteria before looking at the candidate. A practical policy might block a release on any observed access-control violation, require reviewed evidence that important task slices have not regressed beyond a chosen tolerance, and require the latency and cost budgets to hold under a representative workload. Passing a finite suite does not prove the absence of security flaws or rare failures.

A failed gate should produce a debugging artifact: the candidate and baseline release IDs, dataset revision, failed cases, and relevant traces. A red build that says only "quality decreased" leaves the next engineer with another research project.

Requests per second alone is a weak description of an LLM workload. A short question with a short answer and a long document with a long answer can impose very different demands. Test distributions of input length, output length, concurrency, and arrival bursts. Include both warm and cold cache behavior.

For streaming responses, separate time to first token from the pace of subsequent tokens and total completion time. Queue time matters too: a server can generate tokens quickly after admission while users spend most of their time waiting. vLLM's metrics documentation exposes these distinct measurements, alongside queue-depth gauges and token counters. These server-side measurements do not replace client-visible latency across retrieval, networking, and rendering. [2]

Start with end-to-end traces, then inspect retrieval, reranking, queueing, prefill, decoding, and downstream calls where the serving stack exposes them. Do not add the p95 of each stage and call it the end-to-end p95; those percentiles may describe different requests. Use per-request timing to identify where slow requests spend their time.

Batching illustrates the tradeoff. Waiting to assemble work can improve throughput, yet extend user-visible delay. NVIDIA's Triton documentation makes this tradeoff explicit through configurable queue delay for dynamic batching. That mechanism is not identical to an autoregressive server's continuous batching, but the operational lesson carries over: benchmark the scheduler you actually run. [3]

GPU utilization is a diagnostic signal, not the product objective. Decide what the user must experience, then measure how much capacity it takes to meet that requirement. Bound queues, propagate deadlines, and cancel work when the client no longer needs it where the stack supports cancellation. Retries should respect the remaining deadline and an explicit retry budget; unlimited retries can add load to an already overloaded system.

Record release identity on traces and structured request events. Track quality signals, latency distributions, errors, token usage, and fallback rates together. Keep high-cardinality identifiers in traces or logs rather than turning every request or document into a metrics label. Record enough to diagnose behavior without logging raw prompts and retrieved content by default; access controls, redaction, sampling, and retention limits matter.

A cheaper request is not necessarily a cheaper completed task. If a low-cost candidate triggers more retries or human escalation, its apparent saving may disappear. For a fixed measurement window, calculate cost per successful task as the attributable serving and supporting costs divided by tasks meeting a defined success criterion. Count failed attempts in the numerator. If reliable success labels are unavailable, report the proxy explicitly instead of calling it task success.

Compare like with like. Cache hit rate, output length, traffic mix, and quality all affect the result. A release that looks cheaper because it silently truncates answers should fail evaluation, not win a cost comparison. The goal is a useful operating envelope: which workloads meet the quality and latency requirements, at what cost, and with how much headroom?

Keep the current release available while exposing a candidate to a bounded share of traffic. A canary is useful because it limits exposure and creates a comparison, not because a particular percentage is universally safe. Google's SRE guidance emphasizes comparing canary and control signals before expanding a rollout. [4]

Figure 2. Retain a known-good release so rollback can move new requests away from the candidate.

Flowchart showing requests routed from a router to Release A (stable) and Release B (canary). A note indicates rollback routes new requests to A.

Choose assignment deliberately. Randomizing each request may be fine for independent tasks; a conversation may need stable assignment so its behavior does not change midway. Confirm that the canary exercises the workload slices that matter. A quiet canary is not evidence about peak load, and too few completed tasks cannot establish a reliable quality comparison.

Check absolute service objectives as well as candidate-versus-control differences: a shared failing dependency can degrade both. Separate rapid safety signals from slower product signals. Timeouts or a confirmed permission violation can justify stopping immediately. Quality labels and escalation outcomes may arrive later, so promotion may need to wait. Define the owner, stop conditions, minimum observation requirements, and recovery procedure before starting the rollout.

Rollback must restore compatible dependencies, not just older model weights. If the candidate overwrites the retrieval index in place, routing back to an old application image may still leave it reading the new index. Retain compatible index versions or design a reversible migration. Historical snapshots must still honor current access revocations and deletion requirements. Version or invalidate relevant caches so an old route does not serve results produced under the candidate's assumptions.

Routing only affects requests that have not yet been assigned. In-flight generations need an explicit drain or cancellation policy. Tool side effects need separate protection: sending an email or updating a record cannot be undone by switching model versions. Use idempotency and approval boundaries appropriate to those actions, and do not let shadow traffic execute real side effects.

After an incident, add the failure to the evaluation suite with its expected behavior and context. That is the connection between operations and the next release. But production feedback is selective: complaints overrepresent some users, clicks are not necessarily correctness, and missing labels can hide unsuccessful sessions.

For predictive models, input drift can trigger investigation; it does not by itself prove that accuracy has fallen or that retraining will help. For generative applications, a prompt or retrieval change can require a new release even when there is no training job. In both cases, promotion should depend on evidence about the resulting system.

Assign ownership at the interfaces. The application team defines acceptable task behavior. The platform team makes release resolution, deployment, telemetry, and recovery dependable. Data owners maintain freshness and access policies. Decide who can stop a rollout; an alert without someone empowered to act is incomplete infrastructure.

The smallest useful implementation can live in an existing repository: a release manifest, an evaluation job, a representative load test, release-aware traces, and a rehearsed switch back to the previous version. Add more platform machinery when repeated operational pain justifies it.

Before the next launch, ask one question: can the on-call engineer identify the complete release behind a bad answer and restore a compatible known-good version? If not, improve that path before making deployment faster.

[1] Google Cloud. MLOps continuous delivery and automation pipelines in machine learning

[2] vLLM. Production metrics

[3] NVIDIA Triton Inference Server. Batchers

[4] Google SRE Workbook. Canarying releases

Original source

本文由 AI 翻译整理自 Stack Overflow Blog,原文版权归原作者所有。

阅读英文原文
上一篇
MCP 服务器悄然变更工具描述 = 你的 Agent 在执行未知指令
下一篇
GitHub Copilot 已上线 Claude Sonnet 5.5