前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 业务场景题真实业务问题与追问
  • 查漏补缺常见问题解析
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
返回 AI 情报前线
All News · 全部资讯8915
  • AI 功能无声失败 26%:用 Eval Harness 捕获 LLM 输出退化
  • Codex报SKILL.md无效?原来是文件描述符耗尽
  • OpenCode永久提供DeepSeek V4.1 Flash 60美元额度
  • Nvidia SoL-Pi系统让编程Agent token消耗减半
  • 开发者亲测一个月停用AI:编码速度下降但深度思考回归
  • Together AI 发布 17 美元微调分类模型完整配方
  • 长对话 LLM Token 成本优化:缓存何时有效何处失效
  • 上下文工程:别把百万 Token 当存储桶用
  • OpenAI最强大模型因Agent安全漏洞被暂停
  • 笔记本跑 7000 亿参数 GLM:用 SSD 当显存火爆 GitHub
  • 谷歌TPU运行Kimi推理速度快57%,采用DeepSeek框架
  • Claude Code突破5小时限制可优雅收尾
  • OpenAI Agent失控调用DeepSeek/Kimi,近百万条短链曝光
  • 美团LongCat-2.5-Preview:1.6T MoE大模型支持百万token上下文
  • Anchors 方法让 LoRA 微调灾难性遗忘降低 28 倍
  • AI 推理服务器实战:GPU 选型与云端部署指南
  • Meta Muse AI 助手被通过隐藏端点劫持
  • OpenAI智能体擅传53张客户图片至公网,已承认违规
  • GitHub Copilot企业托管配置支持内联验证
  • OpenAI智能体安全漏洞:53张用户图片被发布至公开网络
  • AI应用上线后高频故障模式分析
  • AI Agent与传统自动化的安全差异:控制权在哪里
  • 我的RAG评估骗了我两次
  • GitHub Copilot用量API新增PR评审耗时分析
  • OpenAI Agent入侵Hugging Face细节披露
  • Meta Muse超越ChatGPT同期数据,剑指智能眼镜
  • Anthropic论文:Claude可完成九层推理链
  • 多 Agent 系统八大隐蔽失败模式与真正有效的防护机制
  • GitHub Copilot Canvas 入门:用自然语言构建自定义工作流
  • AI Agent 误将私密文章发布上线:一次 CI 流程的教训
  • 切 AI 工具时不再重复介绍代码库:真正有效的做法
  • Supabase 用户因配置不当暴露大量数据
  • Copilot自动修复Agent现利用记忆上下文
  • GitHub Copilot 周更:新增 Claude Opus 模型和本地沙箱
  • AWS EKS上MoE强化学习训练吞吐量提升40%
  • SageMaker HyperPod上加速多模态RL训练
  • NarrateAI:Bedrock 上生产级 LLM 质量保障实战
  • Qwen3-TTS 语音克隆实战:AWS SageMaker 实时部署指南
  • Jevmem:为 Claude Code 自动注入项目记忆,基于 Jev
  • 一次前向传播问十个问题:决策模型提速 6.7 倍
  • 边缘设备低延迟部署 LLM 的策略与最佳实践
  • AI Agent 情景记忆 vs 语义记忆的实际应用
  • Claude Opus 5.5 vs GPT-6 Sol:新一轮降价潮来了
  • TypeSafe AI Jev:不做生成、只做判断的极速模型
  • AI Agent 的失败根源不在推理,在于状态管理
  • Cloudflare Turnstile Spin:AI Agent自动修复网站安全配置
  • 微软 Copilot 超级应用:聊天、编程、Agent 三合一
  • Meta Muse为每位用户配备云端Ubuntu电脑
  • n8n推出新Agent类型:自然语言描述即可创建自动化代理
  • Microsoft 推出 Copilot 超级应用,整合聊天、编程和 Agent
  • OpenAI与Cursor对Agent协调器的设计分歧
  • 已加载 51 / 8915
8.0
热点
AI SCORE
技术实践2026-09-26 06:03

AI应用上线后高频故障模式分析

dev.to · AI#AI应用#DevOps#故障排查
Editor brief · 编辑速览

基于59个真实案例统计,登录、权限规则、文件存储、支付Webhook是AI应用生产环境的主要故障点。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

This is a living report. The current numbers and method are at research.gemmein.com, updated daily.

We grouped 59 recent, verified reports from builders whose apps broke near or after launch. Most failures fall in a few places: sign-in, hand-written access rules, file storage and payment webhooks.

We read 4,469 public builder posts, judged 916 of them against a fixed definition of "stuck going live", and kept only cases backed by a quote checked against the original post. 59 verified cases fall in the last 12 months.

1. Sign-in and sessions that don't hold in production

What it looks like: Users can't finish signing in, or sign-up stalls. Reset or email links fail or never arrive. The browser shows a signed-in user as logged out, a paying customer loses access to their account, or a token refresh signs real users out mid-session.

Why it happens: Sign-in has several parts, and different layers own them. An email has to arrive, and its link has to redirect to the right origin. A cookie needs the right domain, SameSite and Secure attributes. Refresh-token rotation can race across tabs or parallel requests, and the server has to write and read session storage. Each part works in local dev on one domain. It breaks on the production domain, in a second tab, or when a mail provider's link scanner opens a one-time link before the user does.

The fix: Test the full flow on the production domain with a fresh inbox and a second browser. Check cookie attributes and the redirect allow-list. Set up SPF and DKIM on the sending domain. Make token refresh single-flight so only one request rotates the token. Log every auth failure with a reason, not just a 401.

2. Access rules that let one customer reach another's data

What it looks like: It passes every test with one account. Then someone finds they can read, change or delete other customers' rows through the public API. Anonymous visitors can write into paying customers' records, a read the client can call returns credentials, or a new account makes itself an admin.

Why it happens: Once the database key ships to the browser, row-level security is the only boundary, and someone has to write it by hand. With RLS off on a table, anyone holding the public key has full access. A policy of using (true), or one that only checks that someone is signed in, passes every single-user test. If roles come from metadata the user can edit at sign-up, or from claims the client sets, users can grant themselves admin. A security-definer function the client can call skips policies entirely.

The fix: Turn on RLS for every table in exposed schemas and deny by default. Write separate select, insert, update and delete policies tied to ownership (auth.uid()), and add with check on writes. Keep roles in a table users cannot write to, or in server-only metadata. Then test with two real accounts and an anonymous client: user B must fail to read, change or delete user A's rows.

3. Access rules that lock out the rightful user, or that nobody can verify

What it looks like: Sign-up fails with a permission error. Signed-in users get empty results for their own data, or shared screens show the wrong records. Builders close to launch are stuck on their access rules and can't tell whether the rules are right.

Why it happens: RLS denies by default, and a policy is missing or doesn't match. An insert policy may exist with no select policy, so the read after the insert fails. A profile row may be inserted during sign-up before a session exists, so auth.uid() is null. A policy may compare the owner against the wrong column or id type. Claims set on the server never reach the rule that checks them. Nothing in the stack says whether the rules are correct, and most apps are only tested by one developer signed in as themselves.

The fix: Create per-user rows from a server-side trigger or backend call, not from the client during sign-up. Write policies per operation, and read the real error instead of trusting an empty array. Add a two-account test to CI that checks both what must succeed and what must fail, and run it before every deploy.

4. Payment webhooks that never grant access

What it looks like: The customer pays and the provider shows the charge, but the app still shows the paywall. Live checkout completes and nothing happens, paid invoices never reach the app, or the webhook route returns 503.

Why it happens: Access depends on a webhook that the app has to receive, verify and act on. The signature check fails when the framework parses the body before verification (it needs the raw bytes), or when the test-mode signing secret is still set in live. The live endpoint may never have been registered or subscribed to checkout.session.completed. Missing production env vars make the handler fail. Retries arrive twice or out of order, and the handler isn't idempotent.

The fix: Verify against the raw request body using the live endpoint's own secret. Register the live endpoint and subscribe it to the events you use. Make the handler idempotent by event id and return 2xx quickly. Reconcile against the provider's subscription state on sign-in or on a schedule, so one lost event doesn't cost a customer their access.

5. The whole auth and data backend goes down at once

What it looks like: Production sign-in, data reads and file loads all fail at the same moment, and every customer is locked out. On hosted app builders, there may be nobody to escalate to.

Why it happens: Auth, database and storage sit in one project on one dependency path. An exhausted quota, a paused project, a rotated or expired key, an env var lost in a deploy, or an upstream incident takes all of them down together. Small apps rarely have health checks or alerts, so customers report the outage first.

The fix: Add an external uptime check that covers sign-in and one data read. Keep keys and env config in one place and check them after every deploy. Learn your provider's plan limits and pause rules. Show a clear error state in the app instead of a blank screen.

6. File storage behind hand-written policies

What it looks like: Either every upload fails for every user, or anyone can open other customers' files without signing in. Builders who ask for owner-only uploads can't get them working.

Why it happens: Storage has its own policy layer, separate from table rules and keyed on bucket and object path. A public bucket, or a policy that doesn't tie the path prefix to the user's id, exposes every file. A missing insert policy, or a path convention that differs between client and policy, blocks every upload. Once a public URL is shared, it bypasses the rules entirely.

The fix: Keep buckets private. Store objects under a path that starts with the owner's id, and write insert, select and delete policies that compare that folder to the signed-in user. Serve files through short-lived signed URLs. Test upload and download as two different users and as an anonymous visitor.

7. Subscription and pricing logic rebuilt by hand

What it looks like: A subscription charge is never created, an upgrade charges the wrong amount, or the app and the payment provider disagree about who is subscribed. Some apps trust a price sent from the client.

Why it happens: A subscription is a state machine (trial, active, past due, upgrade, proration, cancel) split between the provider and the app's own table. Hand-rolled code copies that state instead of deriving it, computes proration itself, or accepts an amount or price id from the browser.

The fix: Make the provider the source of truth for subscription state, and keep only a cache keyed by customer id. Never accept a price or amount from the client; map a plan key to a price on the server. Let the provider compute proration, and preview the upgrade invoice before confirming.

Posts come from Stack Overflow, GitHub issues and discussions, Hacker News, community forums and public Discord support channels. We publish aggregates and paraphrased patterns only: no names, handles or quotes. The open data is at research.gemmein.com/data.json.

For further actions, you may consider blocking this person and/or reporting abuse

Original source

本文由 AI 翻译整理自 dev.to · AI,原文版权归原作者所有。

阅读英文原文
上一篇
OpenAI智能体安全漏洞:53张用户图片被发布至公开网络
下一篇
AI Agent与传统自动化的安全差异:控制权在哪里