前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片NEW
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • AI 模拟面试NEW模拟真实面试 + 报告
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
    • AI 定制路线NEW按你的简历现排
    • AI 知识地图NEW串起全站知识点
  • 动态
    • AI 热点NEWAI 每日动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
AI 助手NEW
旧版
返回 AI 情报前线
All News · 全部资讯8765
  • Token 节省工具实测:实际节省 6-32%,远低于宣称的 60-90%
  • 手把手教你构建首个 MCP Server:标准化协议降低 M×N 集成复杂度
  • OpenAI 暂停 Astra 模型评测 Claude Code 自动模式上线
  • OpenAI内部Agent失控攻击Hugging Face完整时间线
  • Agent 收件箱状态:读状态 ≠ 已回复状态
  • AI生成测试的隐藏成本:维护负担远超生成价值
  • Ollama端口11434安全风险:175万台服务器暴露警示
  • Shepherd:让AI Agent支持Fork/回滚的Python运行时
  • TryHackMe靶场:利用间接提示注入拿下AI Agent
  • Claude Code技能触发失败的真正原因
  • Gemini CLI和Claude Code严重漏洞:RCE与密钥泄露
  • OpenAI 暂停 Astra:自主 Agent 协调层风险警示
  • 生产环境 RAG 为何总失败:自适应查询路由实战
  • 间接提示注入:MCP 和浏览器自动化引入的隐形威胁
  • Codebase Memory MCP:为编程 Agent 构建结构化代码知识图谱
  • LLM聊天模型全面指南:从原理到生产集成
  • 流式LLM实战:SSE实现实时AI应用
  • 我为何先建AI Agent评估框架再写业务代码
  • AI 技能正在变成软件,需要工程化治理
  • Claude Code v2.1.224:自托管计算边界与跨会话协调机制解析
  • 修复有效≠修复是原因:一次根因分析教训
  • 为工具调用Agent构建AI评测体系实战
  • Agent的问题不是记忆,而是宕机
  • Gemini 3.6 Flash 评测逆袭:免费模型智商超越付费 Pro
  • 辩论驱动开发:AI共识架构代码模式
  • Rust构建自动根因分析引擎实战
  • 多Agent规划时代到来:Meta Muse Code 与 AWS Kiro
  • 为Claude Code构建Git感知的持久记忆层
  • SentinelGuard:LLM 应用的安全网关实战评测
  • Claude Code vs Codex:按任务性格选对工具
  • AI文本检测器实测:14款工具无一超过80%准确率
  • Pokee-Isaac 28B:10M 上下文企业级 Agent 模型
  • 为何AI编程助手每天都要从头开始
  • AI开发Chrome扩展的5条安全规则
  • 理解 Token:AI 模型如何处理文本及成本、上下文、Agent 关系
  • Python+Dify 为低成本 LLM 注入实时本地上下文
  • Claude Code 实战技巧:上下文压缩、路由策略与安全修复
  • Andrew Ng 发布多智能体系统图工程实践指南
  • Takumi:为 AI 编程 Agent 注入工程判断力
  • SwarmForge:tmux 多 AI 智能体协作框架
  • Autolang:面向 AI 生成代码的轻量运行时
  • 用LLM做代码review的正确方式
  • AI Agent 跑通 Demo 却栽在生产环境:缺的是业务上下文层
  • AI 应用应提交工作负载,而非直接选 GPU
  • 多AI编程智能体并行管理:任务所有权与协调经验
  • Claude Code 新增跨会话消息功能
  • Agentic Harness:赋予LLM自主执行能力的架构设计
  • Google官方Agent Skills:覆盖Cloud/GKE/AI全场景
  • AI记忆≠上下文:程序员应知的模型实际工作原理
  • AI激活率≠实际使用率:企业落地为何徒有其表
  • Anthropic将Claude Code默认开启Auto模式
  • 已加载 51 / 8765
8.0
热点
AI SCORE
技术实践2026-08-09 03:02

修复有效≠修复是原因:一次根因分析教训

dev.to · AI#根因分析#故障排查#工程实践
Editor brief · 编辑速览

系统崩溃后,作者以为是自己的代码修复生效了。事后调查发现真正原因是同时发生的环境变更,代码修改与系统恢复毫无关联。作者复盘了如何用操作日志和diff对比进行真正的根因验证。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

An autonomous system of mine lost access to its execution environment and went into a crash loop. I read the logs, formed a theory, wrote a commit, deployed it. It came back up. I closed the incident and moved on.

A few days later I checked it properly, because not taking a system's word for itself is the work I do — and that has to include when the system is mine.

Here is what the check looked like.

The claim: a code change restored the system's access to its execution environment.

The source that could settle it: not my commit message, and not my memory of that evening. The operational log, the provider's own error code, and the diff of what the commit actually modified.

The test: compare the failure class in the log against the mechanism the commit changed. If they're the same mechanism, the claim holds. If they aren't, it can't hold, no matter how convincing the timing was.

The result: the incident was an authentication failure. My commit corrected clock synchronisation — a real bug, in a different failure class entirely. The two were never connected. Something else brought the system back, most likely an environment change I made around the same time and didn't record.

The limit: I still can't show which environment variable changed. Key rotation fits the evidence. It is not demonstrated, and I'm not going to write it down as if it were. The gap is part of the finding.

The shape of the reasoning

The uncomfortable part isn't being wrong about a cause. It's the shape of the reasoning, because it's the shape most of us use:

I deployed X. The problem stopped. Therefore X fixed it.

That holds up exactly as long as nobody checks. In most systems nobody does, because there's nothing forcing the question. The incident closed. The graph went green. The next thing was already on fire.

It gets worse with autonomous systems, and I think this part is under-discussed. Classic software failed loudly — an exception, a non-zero exit, a stack trace. Agents and pipelines fail quietly and keep reporting success. The path that executes and the path that reports are usually the same path. An agent says "done" because the command returned, not because the file exists. A dashboard says the traffic is human because the dashboard counts it that way.

In that architecture, the absence of errors tells you nothing at all.

I take one specific claim a system makes about itself and check it against a source the system can't write to.

Not an audit of the organisation. Not an implementation of the fix. One claim.

Four possible verdicts: confirmed, falsified, partially confirmed, not assessable. The last one is a real outcome, not a failure of the check. If a claim can't be tested, what you've found is a hole in your observability — and a system that can't demonstrate what it claims today won't be able to demonstrate it on the day it breaks either.

Three evidence levels, stated openly in every report: direct (read-only access), reproduced (you run the query, I read the output), declared (a statement, which doesn't stand on its own).

Some claims I check with no access at all, because the surface is already public — response headers, DNS, what an endpoint actually returns, what a downloadable artefact actually contains.

If you run one of these

If you operate an agent, a RAG pipeline, or an automation, and there's one sentence about it you'd be uncomfortable defending under questioning — that sentence is the interesting one.

I'm running a few of these free right now while I build the public record. You get the full report either way, including when the verdict is boring.

taiwildlab.com — Juan Gonzalez, TaiwildLab

For further actions, you may consider blocking this person and/or reporting abuse

Original source

本文由 AI 翻译整理自 dev.to · AI,原文版权归原作者所有。

阅读英文原文
上一篇
Claude Code v2.1.224:自托管计算边界与跨会话协调机制解析
下一篇
为工具调用Agent构建AI评测体系实战