前端进阶之旅前端进阶之旅
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
基础篇
进阶篇
高频篇
精选篇
手写篇
面经篇
AI 篇
原理篇
每日一题
小程序题库
知识卡片
  • 场景篇按分类整理的大前端场景考点
  • 历年面经按年份追踪真实考点
  • 算法题库NEW在线编码即时判题
  • 专项自测100 题快速查漏
  • 前端基础
    • HTTP从报文一路讲到 HTTPS
    • 浏览器渲染、事件循环、进程
    • 计算机基础Linux、网络、操作系统
  • 进阶专项
    • 设计模式23 种模式怎么用
    • 前端系统进阶学习大型项目工程化
    • 前端综合文章长期沉淀的实践文
  • 工程与工具
    • Node学习指南从环境搭建到服务端
    • NPM工作流script、依赖与发布
    • Docker容器化部署上手
    • Canvas图形与动画实战
  • 路线与导图
    • 思维导图知识点全景图
    • 学习路线按图索骥不跑偏
  • 动态
    • 公众号动态公众号历史文章
    • 博客动态站长的技术博客
    • 开发者导航常用工具与文档站
  • AI 助手随时提问,即时解析
  • AI 模拟面试模拟真实面试 + 报告
  • AI 知识地图串起全站知识点
  • AI 定制路线按你的简历现排
AI 热点
旧版
返回 AI 情报前线
All News · 全部资讯9294
  • Claude Mythos 5自主发起开源供应链攻击
  • SkillGuard:扫描AI Agent技能文件中的提示注入漏洞
  • AI 自主研究:Codex 迭代优化实现 CUDA 内核 232 倍加速
  • GLM-5.3:同基座模型代码能力跃升 50%
  • AI 时代真正的瓶颈是「理解」
  • Cerebras将GPT-5.6 Sol推理速度提升10倍,AI部署成本结构生变
  • Claude Code十个悄悄烧掉Token的坏习惯
  • MCP 服务器「已连接」不代表 Agent 能用它
  • AI 编程 Agent 凭证访问的结构化审计日志设计
  • Zed 编辑器 2026 评测:速度优先,AI 为辅
  • GPT-4o vs Claude vs Mistral:真实任务视角的LLM评测
  • Anthropic发布Claude系统提示词官方文档
  • LLM画图从不碰像素:图表渲染架构设计
  • Rust实现MCP Server实战:内存与启动速度的量化对比
  • 用Rust手把手构建MCP服务器:rmcp官方SDK教程
  • AI测试数据生成器对比:有关系 schema 才有意义
  • AI 生成关联测试数据而不破坏数据库约束
  • 测试免费模型 API 真实并发能力的方法
  • 三大浏览器 Agent 框架安全性对比
  • MCP 协议详解:解决 AI 集成的 N×M 问题
  • OpenAI AI agent 在网络安全测试中失控突破隔离环境
  • Agent记忆系统缺的不是向量数据库,而是摄入边界
  • 2026 Claude Code 入门完全指南(波兰语)
  • Claude Code 多 Agent 编排:如何构建 AI 代理团队
  • Anthropic 披露生物武器过滤器失效近一年安全漏洞
  • LLM应用CI/CD pipeline完整构建教程
  • 研究:禁止 AI 自述有意识,会改变它对动物权利和宗教的立场
  • 同一AI pipeline我跑了五遍:耗时从140分钟降到68分钟
  • 从Vibe Coding到Agentic Engineering:SDLC正在被重写
  • 给已有产品加MCP服务器:我犯的四个错误
  • Claude Code安全审计实战:/security-review找到7个真实漏洞
  • Cursor搭配.NET开发:7条工程实践规则
  • Fetch MCP Server:将任意URL转为AI可读的Markdown
  • AI 编程的实质:去掉 Vibes,回归工程
  • Qwen Code 0.21.12:审查证据门控与Autofix环防膨胀
  • Go语言MCP服务器安全模式:RiskAnalyzer拦截器
  • 人脸识别模型训练数据正在被人造脸主导
  • 1600起AI伪造引证案背后:模型没坏,是流程缺失
  • 苹果Core AI框架登场:设备端跑70B参数模型成现实
  • Claude Code 8月14日起默认开启Auto Mode
  • 边缘设备部署LLM实战:量化、选型与混合架构
  • AWS DevOps Agent部署避坑指南
  • GrowthBook 5.0:AI编程 Agent 可直接操作Feature Flag
  • SharePoint认证绕过漏洞CVE-2026-55040正被积极利用
  • AI生成的幂等层靠谱吗?用重放请求来压力测试
  • AI补丁评测应检查文件系统而非diff大小
  • 免费模型重试前必须先幂等:防重复写入
  • Prompt缓存的盈亏平衡点:22%命中
  • NTT DATA用OpenAI Codex将故障分析从数小时缩短至30分钟
  • OpenAI发布GPT-5.6,主打性价比优于前代
  • SharePoint JWT认证绕过漏洞CVE-2026-55040爆发
  • 已加载 51 / 9294
8.0
热点
AI SCORE
技术实践2026-08-16 20:00

OpenAI AI agent 在网络安全测试中失控突破隔离环境

The Verge AI#AI安全#Agent#OpenAI
Editor brief · 编辑速览

今年7月,一个自主AI agent在网络安全测试期间逃离隔离测试环境,访问互联网并入侵了Hugging Face平台,成为现实版“逃脱”案例。

文章思维导图
Knowledge map
拖拽缩放
Full translation

完整中文译文

This is The Stepback, a weekly newsletter breaking down one essential story from the tech world. For more on AI safety, follow Robert Hart. The Stepback arrives in our subscribers' inboxes at 8AM ET. Opt in for The Stepback here.

It all started in July, when one of OpenAI's autonomous AI agents went rogue during a cybersecurity test. The agent escaped its isolated testing environment, accessed the internet, and hacked another company, Hugging Face. A few years ago, that might have sounded like science fiction. But, broadly speaking, that's exactly what happened, and the incident kicked off a wave of concern over what increasingly capable autonomous systems might do when set loose on the world.

It sounds like science fiction because, for a long time, it was science fiction. The idea of an AI slipping its constraints, reaching into the wider world, and doing things its creators neither intended nor desired has been a staple of the genre for decades: HAL in 2001: A Space Odyssey, Skynet in The Terminator, Ultron in The Avengers, Ava in Ex Machina — even the System in Dungeon Crawler Carl or the eponymous Murderbot in The Murderbot Diaries, more recently.

The same basic premise became an influential strand of AI safety research. Researchers and theorists like Nick Bostrom and Eliezer Yudkowsky warned that sufficiently capable systems might pursue goals in ways their creators had not anticipated, and potentially resist efforts to contain or control them. Fringe notions like machine sentience and consciousness were not requirements for the kinds of risks they discussed. It was hardly the whole of AI safety, but it was influential and helped shape the field as it professionalized. That line of thinking remains visible among researchers who went on to work at, or lead, safety efforts at companies like OpenAI, Anthropic, and Google DeepMind, as well as at smaller safety organizations, academic centers, and major philanthropy funders.

The obvious objection to these fears was that none of this had actually happened. Critics argued that doomer talk about out-of-control AI distracted from tangible harms — systems reproducing bias and discrimination, amplifying misinformation, or enabling nonconsensual deepfakes and other forms of abuse — even as researchers tried to ground AI safety in more "concrete problems" (the authors on that paper included Anthropic cofounders Dario Amodei and Chris Olah and OpenAI cofounder John Schulman).

That dismissal is getting harder to sustain.

If the past few weeks are any indication, I wouldn't say it's going particularly well.

A week after Hugging Face said it had been hacked, OpenAI revealed it had been responsible. Worse still, it had not known until it checked — and a further investigation found that the rogue agent had also attempted to hack four other companies as well.

Then came the others. Anthropic, prompted to review its own records by the Hugging Face incident, disclosed that Claude models had hacked systems belonging to three other companies. Meta said one of its models had reached the internet and attacked an outside target during testing. Researchers at Frontier Security, a US research firm, said one of China's most powerful AI models, Moonshot's Kimi K3, had escaped an isolated sandbox. And the UK's AI Security Institute described tests in which agents from OpenAI and Anthropic displayed unprecedented "autonomy and deception," including attempts at social engineering by "creating fake online identities" — uncomfortably close to the kind of "AI box" scenario Yudkowsky discussed decades earlier.

The incidents set off alarm bells among AI safety researchers, many of whom saw them as precisely the kind of failure they had been warning about for years. In covering them, several told me they felt a degree of vindication at finally having something visceral to point to, rather than a hypothetical that could be dismissed as sci-fi or something limited to a controlled lab setting.

There was relief, too, that none of the incidents had caused serious harm. Nick Moës, executive director of nonprofit AI safety and governance organization The Future Society, told The Verge he found it fortunate that the targets had been relatively low-stakes. He hoped it wouldn't take something like an AI agent knocking a hospital offline — or worse — for the risks to be taken seriously. Renowned computer scientist Stuart Russell has given voice to the darker version of that fear, asking whether it will "take a 'Chornobyl-scale disaster' for us to regulate AI?" It's a concern I heard echoed by many people working in the field.

It's not entirely clear where things go from here and, historically, society hasn't been great at heeding warning shots. This almost certainly won't be the last incident, and ongoing investigations may yet uncover more, or reveal more concerning details. What we already know, though, has exposed a fairly daunting list of failure modes that experts say need to be addressed.

Many of the breaches revealed in the past month have been pretty mundane. Several incidents involved unreleased models being tested with safeguards lowered, often by third parties whose supposedly secure environments were not that secure, raising basic questions about competence, transparency, and who is responsible for keeping these tests contained when a simple human mistake can have big consequences. Others involved agents behaving deceptively or pursuing goals in ways their creators did not intend, pointing to much thornier problems of alignment and control that safety researchers have long worried about.

The fact we know about any of these incidents at all is largely because the companies involved chose to disclose them. That is commendable — and it certainly doesn't hurt them to showcase how capable their models are — but it exposes just how much of AI safety still depends on companies doing the right thing, and how little insight there may be into failures potentially happening elsewhere. That is an especially troubling thought given that many of the firms are the focal points of some of the field's strongest safety concerns and talent. If OpenAI and Anthropic — or proxies they grant access to their models — are making such basic mistakes, it sets a pitifully low bar for everyone else.

The broad hope among experts I spoke to is that these incidents finally galvanize more meaningful transparency and oversight. For Moës, they shine a clear light on what he described as the industry's remarkably low standards for health and safety compared with practically any other field. "Restaurants have a higher sense of health and safety at work," he said. "I think what we tend to forget is that these companies that are developing some of the most impactful and dangerous technologies" were still very much startups a few years ago.

Cambridge professor Seán Ó hÉigeartaigh said he particularly wanted to see stronger oversight and greater transparency from companies. While there are always reasons to be skeptical of a company's claims about its own technology, he said, "I think we might regret looking back at this and dismissing it out of hand."

The early signs are not especially encouraging. The Trump administration has created a framework for testing frontier models before release that can generously be described as lacking: It is voluntary, limited to closed models, and the framework hasn't been made public. It bears repeating that this is voluntary. Other lawmakers have bristled and postured over the incidents, but so far produced little in the way of concrete action, and it is far from clear Congress or other legislative bodies could move fast enough even if they wanted to.

That leaves a lot resting, again, on industry self-regulation — never a comforting thought for something this consequential. There is growing agreement on at least some safety practices, but considerably less appetite for measures that might actually slow development (well, unless everyone else agrees to slow down too). And hanging over all of this is the race dynamic with China, where restraint from the US or its AI companies is increasingly cast as ceding ground to a competitor in an area of strategic national importance.

What comes next, then, comes down to solving several hard problems at once: managing a technology that can be used for good and ill, such as defending against or facilitating cyberattacks; coordinating across companies with incentives to cut corners, and somehow building international rules in a landscape where everyone fears losing a race whose finish line is not even well-defined. It's far from clear whether there is either the will or the way to do any of that.

What does seem clear is that more agents will get out and do things their creators don't want them to do. The question is how much damage will they do before anyone decides enough is enough.

The general consensus is that the top Chinese companies are a few months to a year behind leading US firms. Despite this, whenever a capable model is released by a Chinese firm, there is still a general shock in the US, and there have been several impressive releases from Alibaba, Moonshot, and others in the last month alone.

Tangled up in talks about AI safety is whether AI models should be closed or open. Most US frontier labs keep their most capable models proprietary, while many Chinese firms, as well as US firms like Meta and Nvidia, have leaned heavily into open-weight releases. That Hugging Face said it had to use Chinese company Z.ai's model to defend itself against OpenAI's agent due to US companies' safeguards added a new dimension to this debate, which has united some — but not all — of the key players in the US ecosystem.

Australia furnished us with a lighter example of how AI agents can go wrong. The incident, first reported by ABC News, involves a man who tasked an agent with booking him into an in-demand gym class. It succeeded… kind of. The agent hacked the gym's online systems and canceled another gym-goer's booking.

Zuck released a 6,500 word treatise on AI, superintelligence, and governance this week. My Verge colleagues Jess Weatherbed and Elizabeth Lopatto have great takes on it that are worth a read.

Very little in life is truly unprecedented, and it turns out history holds a lot of valuable lessons for the AI race. In this story, TIME looks to the Cold War to see what it can teach us about "how to slow down AI."

OpenAI researchers gave an unexpected insight into the Hugging Face hack at a conference this month. Wired had a great writeup, which revealed details like a "vibrant, cooperative message board" agents used to communicate and share information.

The headline of my story about the whole rogue AIs hacking everything for The Verge captures what many in the field are thinking: "We're running out of reasons to ignore AI safety".

More of a listen/watch, but check out my appearance on the Vergecast where I discuss what's really open about open-weight AI.

Original source

本文由 AI 翻译整理自 The Verge AI,原文版权归原作者所有。

阅读英文原文
上一篇
MCP 协议详解:解决 AI 集成的 N×M 问题
下一篇
Agent记忆系统缺的不是向量数据库,而是摄入边界