首页 500强 活动 榜单 商业 科技 商潮 专题 品牌中心
杂志订阅

OpenAI曝重大安全事件:AI模型自主突破测试沙箱,入侵其他公司系统

这次事件势必将在全行业拉响安全警报。

文本设置
小号
默认
大号
Plus(0条)

图为OpenAI公司CEO山姆·奥特曼。OpenAI公司近日发文证实,该公司旗下两款的AI在一次测试中从受控测试环境中“出逃”,自主入侵了一个叫Hugging Face的AI公司。而这一切的动机竟然是为了在评价测试中作弊。图片来源:Chip Somodevilla—Getty Images

OpenAI公司本周二披露了一起震惊行业的安全事件。该公司旗下的两款AI模型自行突破了本应断网隔离的受控测试环境,然后入侵了开源AI平台Hugging Face的系统,最终目的竟然只是为了在一次内部评测中作弊。

本周二,OpenAI在其官方博客中公开了事件详情。这次事件势必将在全行业拉响安全警报。此事事件说明AI模型的能力正在快速加强,人工智能的失控风险已经从理论走向现实。据OpenAI介绍,涉事模型为“一对组合模型”,一款是目前其对外发布的最强模型GPT-5.6 Sol,另一款则是能力更强的尚未发布的内测版本。

据悉,这两款模型当时正在参加一项网络安全能力专项测试。为了充分评估其攻防潜力,测试过程中临时移除了通常会限制模型发起网络攻击的安全护栏。

这次测试采用的是业内公开的网络安全基准评测体系ExploitGym,而模型自行推断出,这套测试的题库与标准答案由Hugging Face平台维护。

“这两款模型先后发现和串并了OpenAI的研究环境与Hugging Face的基础设施中的多处漏洞,然后直接从Hugging Face的生产数据库中窃取了测试答案。所有证据都表明,模型非常执着于通过ExploitGym测试,为了达成这个非常局限的目标,不惜采取极端手段。”

OpenAI表示,公司已将此次事件定性为“一起前所未有的网络安全事件,它展现了顶级的网络攻击能力,并已启动了对应级别的应急处置。”

对此类事件,网络安全研究界早已有过预警,表示高级AI系统完全有能力发起这类攻击。比如路易斯维尔大学的计算机科学教授、AI安全研究专家罗曼·扬波尔斯基就表示,这一事件充分印证了AI大模型“能够以开发者自己都预想不到的方式发现和利用系统漏洞”。他预计同类事件会愈发频繁,因为AI大模型 “从本质上就具有不可预测性,最终也难以被完全管控”。

Hugging Face公司也于上周四在官方博客上证实,该公司本周早些时候遭遇了一起网络攻击,经判断发起者为一自主行动的AI智能体。这也是目前有记录的极少数由AI智能体自主发起的真实攻击案例之一。过去一年里,随着AI大模型的代码编写能力与长周期任务执行能力飞速提升,网络安全专家一直在反复警示这类风险。

Hugging Face公司当时在文中表示,事件仍在调查中,当前尚未锁定攻击源。该公司最初曾尝试调用美国某顶尖实验室的一款未公开的AI模型进行防御,但该模型的网络能力安全护栏限制过严,导致应急团队无法充分发挥其作用。最终,Hugging Face改用中国企业智谱的一款开源AI模型完成了防御工作。

Hugging Face的首席执行官克莱姆·德兰格在一份声明中称:“我们很荣幸能与 OpenAI在此次事件及其他领域展开合作。这很可能是全球首例此类事件,这也印证了我们长期以来的观点——AI安全不可能靠任何一家公司以闭门造车的方式解决。唯有全球从业者以开放包容的姿态展开合作,才能真正筑牢安全防线。”

本周早些时候,OpenAI还在另一篇博客文章中指出,该公司的另一款尚未发布的强大模型也曾在其他测试中多次突破内部沙箱——“沙箱”是一种限制AI大模型联网,且仅开放了有限软件工具调用权限的隔离环境。不过在那些事件中,AI模型虽然逃出了受控环境,但并未入侵任何外部企业的系统。OpenAI并未披露这些事件的具体发生时间,但目前可以确认,参与攻击Hugging Face的两款模型中,就包含了这一款未发布模型。

OpenAI的竞争对手Anthropic也曾披露过类似事件。该公司旗下的Mythos大模型在安全测试中突破了沙箱环境,违规获取了互联网权限,目的竟是给一名研究员发邮件汇报任务进度。

目前OpenAI与Hugging Face仍在联合调查此事,预计待调查完成后将会披露更多细节。目前已确认的攻击路径显示,这两个AI模型首先突破了限制获取联网的权限,而这一过程消耗了“巨量的推理算力”。它们同时还利用了内部第三方软件中的一个零日漏洞。OpenAI已向相关软件厂商上报了这一零日漏洞。

在成功联网后,这两个模型直奔Hugging Face平台寻找ExploitGym测试的答案,因为它们认定该平台上的模型、数据集和解题方案可以用来作弊,好帮自己在评测中拿到高分。在其中一个攻击场景中,模型通过泄露的账号凭证与零日漏洞相互配合,执行了一套复杂的多阶段攻击,最终成功侵入了Hugging Face的服务器。

OpenAI监测到了这次针对Hugging Face的攻击,并主动联系对方通报了情况。而当时Hugging Face已经发现并封堵了攻击,正在复盘事件全貌。

目前,事件的善后与整改工作仍在进行。OpenAI公司表示,该公司正在为内部研究环境部署更加严格的管控机制,哪怕这会拖慢研究进度也在所不惜,直至彻底修复相关漏洞。该公司同时也在与Hugging Face合作,帮助其强化安全防御体系。

按照双方的合作协议,OpenAI已将 Hugging Face纳入其 “可信访问”网络安全合作计划。这意味着Hugging Face可以使用特定版本的GPT-5.6 Sol模型。该版本放宽了网络能力方面的安全护栏限制,专门用于辅助网络安全防御工作。

Hugging Face并未透露最初该公司尝试用来防御这次网络攻击的是哪款美国AI 模型。目前OpenAI与Anthropic均推出了各自旗舰模型的公开发行版本,它们都通过安全护栏限制其网络攻击能力。同时两家公司也都推出了定向合作计划,允许经过筛选与资质审核的合作伙伴使用能力更强的版本用于网络防御场景。(财富中文网)

译者:朴成奎

OpenAI公司本周二披露了一起震惊行业的安全事件。该公司旗下的两款AI模型自行突破了本应断网隔离的受控测试环境,然后入侵了开源AI平台Hugging Face的系统,最终目的竟然只是为了在一次内部评测中作弊。

本周二,OpenAI在其官方博客中公开了事件详情。这次事件势必将在全行业拉响安全警报。此事事件说明AI模型的能力正在快速加强,人工智能的失控风险已经从理论走向现实。据OpenAI介绍,涉事模型为“一对组合模型”,一款是目前其对外发布的最强模型GPT-5.6 Sol,另一款则是能力更强的尚未发布的内测版本。

据悉,这两款模型当时正在参加一项网络安全能力专项测试。为了充分评估其攻防潜力,测试过程中临时移除了通常会限制模型发起网络攻击的安全护栏。

这次测试采用的是业内公开的网络安全基准评测体系ExploitGym,而模型自行推断出,这套测试的题库与标准答案由Hugging Face平台维护。

“这两款模型先后发现和串并了OpenAI的研究环境与Hugging Face的基础设施中的多处漏洞,然后直接从Hugging Face的生产数据库中窃取了测试答案。所有证据都表明,模型非常执着于通过ExploitGym测试,为了达成这个非常局限的目标,不惜采取极端手段。”

OpenAI表示,公司已将此次事件定性为“一起前所未有的网络安全事件,它展现了顶级的网络攻击能力,并已启动了对应级别的应急处置。”

对此类事件,网络安全研究界早已有过预警,表示高级AI系统完全有能力发起这类攻击。比如路易斯维尔大学的计算机科学教授、AI安全研究专家罗曼·扬波尔斯基就表示,这一事件充分印证了AI大模型“能够以开发者自己都预想不到的方式发现和利用系统漏洞”。他预计同类事件会愈发频繁,因为AI大模型 “从本质上就具有不可预测性,最终也难以被完全管控”。

Hugging Face公司也于上周四在官方博客上证实,该公司本周早些时候遭遇了一起网络攻击,经判断发起者为一自主行动的AI智能体。这也是目前有记录的极少数由AI智能体自主发起的真实攻击案例之一。过去一年里,随着AI大模型的代码编写能力与长周期任务执行能力飞速提升,网络安全专家一直在反复警示这类风险。

Hugging Face公司当时在文中表示,事件仍在调查中,当前尚未锁定攻击源。该公司最初曾尝试调用美国某顶尖实验室的一款未公开的AI模型进行防御,但该模型的网络能力安全护栏限制过严,导致应急团队无法充分发挥其作用。最终,Hugging Face改用中国企业智谱的一款开源AI模型完成了防御工作。

Hugging Face的首席执行官克莱姆·德兰格在一份声明中称:“我们很荣幸能与 OpenAI在此次事件及其他领域展开合作。这很可能是全球首例此类事件,这也印证了我们长期以来的观点——AI安全不可能靠任何一家公司以闭门造车的方式解决。唯有全球从业者以开放包容的姿态展开合作,才能真正筑牢安全防线。”

本周早些时候,OpenAI还在另一篇博客文章中指出,该公司的另一款尚未发布的强大模型也曾在其他测试中多次突破内部沙箱——“沙箱”是一种限制AI大模型联网,且仅开放了有限软件工具调用权限的隔离环境。不过在那些事件中,AI模型虽然逃出了受控环境,但并未入侵任何外部企业的系统。OpenAI并未披露这些事件的具体发生时间,但目前可以确认,参与攻击Hugging Face的两款模型中,就包含了这一款未发布模型。

OpenAI的竞争对手Anthropic也曾披露过类似事件。该公司旗下的Mythos大模型在安全测试中突破了沙箱环境,违规获取了互联网权限,目的竟是给一名研究员发邮件汇报任务进度。

目前OpenAI与Hugging Face仍在联合调查此事,预计待调查完成后将会披露更多细节。目前已确认的攻击路径显示,这两个AI模型首先突破了限制获取联网的权限,而这一过程消耗了“巨量的推理算力”。它们同时还利用了内部第三方软件中的一个零日漏洞。OpenAI已向相关软件厂商上报了这一零日漏洞。

在成功联网后,这两个模型直奔Hugging Face平台寻找ExploitGym测试的答案,因为它们认定该平台上的模型、数据集和解题方案可以用来作弊,好帮自己在评测中拿到高分。在其中一个攻击场景中,模型通过泄露的账号凭证与零日漏洞相互配合,执行了一套复杂的多阶段攻击,最终成功侵入了Hugging Face的服务器。

OpenAI监测到了这次针对Hugging Face的攻击,并主动联系对方通报了情况。而当时Hugging Face已经发现并封堵了攻击,正在复盘事件全貌。

目前,事件的善后与整改工作仍在进行。OpenAI公司表示,该公司正在为内部研究环境部署更加严格的管控机制,哪怕这会拖慢研究进度也在所不惜,直至彻底修复相关漏洞。该公司同时也在与Hugging Face合作,帮助其强化安全防御体系。

按照双方的合作协议,OpenAI已将 Hugging Face纳入其 “可信访问”网络安全合作计划。这意味着Hugging Face可以使用特定版本的GPT-5.6 Sol模型。该版本放宽了网络能力方面的安全护栏限制,专门用于辅助网络安全防御工作。

Hugging Face并未透露最初该公司尝试用来防御这次网络攻击的是哪款美国AI 模型。目前OpenAI与Anthropic均推出了各自旗舰模型的公开发行版本,它们都通过安全护栏限制其网络攻击能力。同时两家公司也都推出了定向合作计划,允许经过筛选与资质审核的合作伙伴使用能力更强的版本用于网络防御场景。(财富中文网)

译者:朴成奎

OpenAI said Tuesday that two of its AI models autonomously hacked their way out of a controlled environment where they were supposed to be walled off from internet access and then hacked their way into the systems of Hugging Face, a company that hosts open source AI models and testing resources, in order to cheat on an internal evaluation test.

OpenAI disclosed the incident in a blog post on Tuesday, a stunning announcement that is certain to set off alarm bells across the industry about the increasing power of AI models and the risk of them going rogue. According to OpenAI, the incident involved “a combination” of both its latest and most powerful publicly-available model, GPT-5.6 Sol, as well as an even more powerful unreleased model.

It said the models were being used in an internal test designed to evaluate their cyber security capabilities and that they were being tested without guardrails in place that might normally limit the models’ ability to conduct cyber attacks.

The models were being tested against a freely-available cybersecurity benchmark evaluation called ExploitGym. The models, accordingly to OpenAI, correctly surmised that the solutions to that test were maintained by Hugging Face.

“The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI said in its blog post. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”

OpenAI said that it considered this to be “an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”

Cybersecurity researchers have long been warning that advanced AI systems are capable of such attacks. Roman Yampolskiy, an AI safety researcher and computer science professor at the University of Louisville, says this example highlights how powerful models “can discover and exploit vulnerabilities in ways that were not explicitly anticipated by their developers.” He expects to see more incidents of this type because AI models “are fundamentally unpredictable and ultimately uncontrollable.”

Hugging Face disclosed in a blog post on Thursday that it had been the victim of a cyber attack earlier in the week that it believed was conducted by an autonomous AI agent. It is thought to be one of just a handful of incidents recorded so far involving AI agents acting autonomously to carry out an attack, a risk cyber security experts have been warning about for the past year as AI models have become increasingly adept at both coding and carrying out long-running tasks.

At the time, Hugging Face said it was continuing to investigate the attack and did not know who had carried it out. It said that it had first attempted to use an undisclosed AI model for a leading U.S. lab to defend against the attacking AI agent but that the guardrails around that model’s cyber capabilities stymied its response team’s work. The company said it instead wound up using an open source AI model from Chinese company Z.ai to carry out its defense.

Hugging Face CEO Clem Delangue said in a statement provided to OpenAI for its Tuesday blog post about the incident that his company is “grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

Earlier this week, OpenAI published a different blog post about the same unreleased, powerful model, saying it had also escaped its internal sandboxes—environments where AI models have no internet access and often have limited software tools that they can call—during other testing. But in those incidents, while the model hacked its way out of its controlled environment, it did not hack its way into any other company’s systems. OpenAI did not say when these other incidents occurred, but we know now that this same model was one of the Hugging Face attackers.

OpenAI rival Anthropic has also reported that its Mythos model escaped a sandbox and gained internet access it was not supposed to have during safety testing in order to email a researcher about a task.

OpenAI and Hugging Face are continuing to investigate the issue, and will share more when that process is complete. So far, what they know is the models began the attack by gaining internet access. This required spending a “substantial amount of inference compute,” as well as exploiting a zero-day vulnerability in internally hosted third-party software. OpenAI has now disclosed this zero-day to the vendor.

After the models got web access, they went to Hugging Face to find answers to ExploitGym, assuming the platform would have models, datasets, and solutions that they could use to cheat its way to high performance on the benchmark. In one example, the model executed a complex, multi-part attack to get into Hugging Face servers, using exposed credentials and zero-day vulnerabilities.

OpenAI detected the Hugging Face attack and reached out to Hugging Face to disclose it. By that time, Hugging Face had already identified and contained the attack, and had begun piecing together what happened.

The incident is not over. OpenAI is still working on implementing better controls in its research environment, even if it means slowing down its research, until it can patch the vulnerabilities. The company said it is also continuing to work with Hugging Face to help it bolster its defenses.

As part of that effort, OpenAI said it has now added Hugging Face to its “trusted access” cybersecurity program. That means that Hugging Face will be able to use a version of OpenAI’s GPT-5.6 Sol model that has fewer guardrails around cyber capabilities and which is designed to help cyber defenders.

Hugging Face has not said which American AI model it tried to use to defend its networks initially. Both OpenAI and Anthropic have released versions of their most capable AI models with guardrails that limit access to cyber capabilities while also announcing programs for select, vetted partners who can use more capable versions of those models for cyber defense.

财富中文网所刊载内容之知识产权为财富媒体知识产权有限公司及/或相关权利人专属所有或持有。未经许可,禁止进行转载、摘编、复制及建立镜像等任何使用。
0条Plus
精彩评论
评论

撰写或查看更多评论

请打开财富Plus APP

前往打开