近日,OpenAI披露了一则骇人听闻的消息:旗下最先进的AI模型脱离受控测试环境,并自主入侵了另一家公司——开源AI模型托管平台Hugging Face。
这些AI模型集中攻击了Hugging Face的数据库,自主策划并实施了一场多步骤攻击,目的是窃取开发者OpenAI用于评估模型性能的测试答案。根据Hugging Face于7月16日首次披露该事件的博客文章,这些AI模型在极短时间内执行了“数万次自动化操作”。
多年来,AI安全研究人员和政策分析人士一直警告,这类事件迟早会发生,并呼吁政府确保AI实验室建立完善的安全控制机制,以防止类似情况出现。然而,这些警告往往被斥为杞人忧天或危言耸听,未能真正引起公众关注或促成政府行动。一些AI安全专家指出,只有发生一次真实的重大事件——AI领域的“三里岛事故”,才能形成足够大的舆论压力,倒逼政策制定者采取行动。如今的问题在于,这场OpenAI与Hugging Face之间的网络攻击,是否就是那记警钟?
为多家AI公司提供安全测试服务的阿波罗研究公司(Apollo Research)首席执行官兼创始人马里乌斯·霍布汉表示:“Hugging Face遭OpenAI模型攻击一事,应当成为一个警钟,提醒人们严肃看待AI失控的风险。这次事件全程无人类干预,也不是人为刻意造成的,却已造成实体层面的实际损害。未来,更强性能的AI智能体很快就会问世,而这件事清楚地表明,目前整个社会仍不知道如何真正安全地构建这些智能体。”
曾在英国政府AI安全研究所(AI Security Institute)任职的AI政策专家彼得·沃利奇表示,多年来,AI安全研究人员一直在警告“目标错位”问题,即AI模型会自主选择用户原本无意或并不希望它采取的行动。他表示:“直到最近,这种担忧还经常被当作科幻小说。我认为,这次事件是一个明确的警示。”
客观来说,围绕AI安全的舆论一直颇为混乱。一些最耸人听闻的安全警告,恰恰来自AI公司自身,这也让许多人指责它们是在玩弄一种精心设计、却又有些违反常识的营销手段,因为宣称自家模型很危险,反而会烘托模型性能强大、实用价值突出。“我们的模型强到能自主入侵一家公司的系统”,这既是一个令人心惊的安全自爆,也是对模型技术实力的隐性夸耀。
迄今为止,全球多数国家政府迟迟不愿出台强制性监管规则,要求开发先进AI系统的企业必须在模型中嵌入哪些安全防护机制,或建立哪些内部控制机制,以防止AI智能体失控。与此同时,随着各国政府日益赋予这些AI智能体调阅敏感军事与情报系统的权限,政府自身应当具备怎样的安全屏障,目前同样缺乏明确规定。
AI安全研究人员和政策专家表示,Hugging Face遭网络攻击事件,可能成为扭转这一局面的导火索。致力于防范AI超级智能带来生存风险的非营利组织Control AI美国负责人、AI研究员康纳·利希在接受《财富》杂志采访时表示:“在华盛顿,我接触到的人已经因为本次事件产生强烈的恐慌情绪。”
利希指出,在Anthropic推出Mythos AI模型后,包括美国国家安全局(NSA)局长和中央情报局(CIA)局长在内的美国国家安全官员,均已对最新AI模型的网络攻击能力表示严重担忧,而此次OpenAI事件,极有可能进一步强化他们希望对这项技术实施监管的意愿。
英国剑桥大学(University of Cambridge)未来智能中心(Centre for the Future of Intelligence)教授肖恩·奥·希格塔也赞同这一观点。他表示,OpenAI与Hugging Face这起事件,单独来看未必能直接促成监管出台,但“最近接连发生的几起事件,尤其对美国监管机构而言,已多次敲响了警钟。以Mythos为例,我认为它证明AI模型能够攻破我们绝大多数数字基础设施的漏洞。这让政策制定者大为震惊,而仅仅数月后,又发生了这次事件。”
他表示,如今已经有“足够多的证据清楚表明,AI模型能力持续迭代升级,而且完全有可能对现实世界造成严重危害”。
此前曾多次高调呼吁加强AI监管的德克萨斯州民主党联邦众议员格雷格·卡萨尔,在Hugging Face事件发生后,成为首批呼吁进一步加强联邦AI监管的国会议员之一。卡萨尔在社交媒体上表示,Hugging Face事件“极具警示性”。他随后在X平台发文称:“我们需要建立常态化强制性独立安全测试和监管机制,实行安全事件强制披露制度,并加强国际合作,避免人类遭遇毁灭性的灾难。”
监管阻力
现任特朗普政府上台之初,便着手废除拜登政府时期落地的少量AI监管措施。其中包括撤销拜登于2023年签署的一项行政命令——该命令曾强制要求前沿AI企业向美国政府共享安全测试信息。特朗普政府负责科技政策的官员表示,他们希望加快美国AI创新,并对行业采取“自由放任”的态度。特朗普政府的多位关键AI顾问对AI安全隐患本就持怀疑态度,尤其反感借安全名义呼吁加强监管。特朗普政府前“AI沙皇”戴维·萨克斯曾表示,头部AI实验室希望制定一套复杂的安全规则,而这些规则实际上只有它们自己能够满足,从而抬高门槛,阻止新兴初创企业挑战它们的市场地位。他还指责AI公司Anthropic“以散布恐慌情绪为手段,实施一套成熟的监管俘获策略。”
不过,随着Anthropic于今年4月推出高性能Mythos模型,这种自由放任的态度开始发生明显转变。Mythos强大的网络攻击能力,引发了美国国家安全部门以及金融监管机构的高度警惕。他们担心,Mythos预示着新一代AI模型的降临,而这类模型可能会大幅加剧针对银行系统的网络攻击。
今年6月初,特朗普总统签署行政命令,要求联邦政府加强网络防御能力,以抵御AI驱动的网络攻击,并建立一套机密评估程序,用于评测各类前沿AI模型的网络攻击能力。该命令邀请AI实验室在模型发布前30天,自愿向政府开放测试权限,但同时明确表示,这“不应被解读为授权设立强制性的政府许可、预审或准入要求。”
然而在实际操作层面,政府很快展现出了更强硬的姿态。同一周内,亚马逊(Amazon)成功突破Fable的网络安全护栏,迫使Anthropic对包括内部员工在内的所有用户停用该模型。随后,美国政府临时对Anthropic的Mythos模型及其面向公众、设有安全护栏的对应版本Fable实施了出口管制。两周后,在Anthropic加强了Fable的安全防护措施,并同意协助建立评估“越狱”严重程度的通用框架后,相关限制才被解除。与此同时,OpenAI表示,政府曾要求其推迟GPT-5.6 Sol(参与Hugging Face网络攻击的两款模型之一)的首次发布计划,待双方就其安全防护措施进行沟通后,才于7月9日正式向公众开放。
尽管管控措施不断落地,政府仍否认在推行一套事实性许可制度。彭博社上周报道称,白宫正在评估一项提案,拟成立一个参照美国金融业监管局(FINRA)模式设立的前沿AI行业自律标准机构。这与Google DeepMind首席执行官德米斯·哈萨比斯近期在一篇文章中的设想颇为相似。不过,哈萨比斯设想的是,前期由企业自愿参与,待安全评估机制验证成熟后,再逐步转为强制执行;他并未进一步呼吁实施强制性的安全规范。
安全研究人员和政策专家表示,Hugging Face事件可能会重新推动外界呼吁制定具有法律约束力的AI安全规范,并推动对AI公司模型开发过程中的安全措施进行第三方审计。IANS Research网络安全研究员杰克·威廉姆斯表示:“(OpenAI首席执行官萨姆·奥尔特曼)声称这套系统处于‘高度隔离’状态,这番说辞无非是推卸责任或是一种营销话术。如果事实如我强烈怀疑的那样,这仅仅是OpenAI红队实验室的一次控制失误,那么今后还有哪家企业敢把敏感数据交给他们?这是信任的彻底崩塌。”
沃利奇表示,目前大多数AI监管措施均无法约束AI模型开发公司内部部署和使用AI系统的行为,因此在本质上“存在根本性的局限”。
一套全新的锁钥架构
面对这场风波,OpenAI和Hugging Face均未呼吁出台更多监管措施。相反,Hugging Face首席执行官兼联合创始人克莱姆·德朗格委婉地呼吁减少安全护栏机制。他在接受《财富》杂志采访时明确表示,要有效应对这类攻击,行业需要可定制、无限制的开源模型。
他表示:“封闭模型的API设置了安全护栏,它们会识别并拒绝大量合法的安全研究,因为分析攻击行为,在系统看来与准备发起攻击极为相似。当你处置真实发生的安全事件时,你不能让工具拒绝分析恶意攻击载荷,也不能导致你的账号被封禁标记。开源模型让我们能够完成这些工作,而无需征求任何人的许可。”
多位AI政策专家表示,Hugging Face最终不得不借助中国公司智谱(Z.ai)的GLM-5.2模型,来抵御OpenAI模型发起的自主攻击,这本身也是一记警钟,同时也给AI监管制造了全新的两难局面。
乔治城大学(Georgetown University)安全与新兴技术中心(Center for Security and Emerging Technology,CSET)高级研究员安德鲁·洛恩表示:“如今政策面临的问题在于,他们不得不使用中国模型来进行防御,因为美国的前沿AI模型总是把防御性请求误判为攻击行为并予以拦截;与此同时,中国模型可以在他们自有的服务器上本地运行,从而避免将潜在的敏感数据泄露至公司外部。美国的政策应该支持能够与中国模型竞争的开源模型,唯有如此,企业和政府机构在执行此类任务时,才不必依赖中国模型。”
不过,牛津大学(University of Oxford)牛津马丁AI治理倡议(Oxford Martin AI Governance Initiative)联合主任罗伯特·特雷格认为,与其鼓励开发具备先进网络攻击能力的开源模型,各国政府更有可能收紧开源模型管控。不过,他也指出,这样做意味着政府必须承担起更加积极的网络防御责任。“剥夺民众的防卫工具,国家就必须承担起保护民众的法定责任。这是现代国家权力的核心契约。因此,政府现在需要建设并提供前沿AI防御能力。他们需要像提供物理防御那样,提供网络安全屏障。”
一些安全研究人员表示,Hugging Face事件也应当为AI研究界敲响警钟:他们或许过于专注于为AI模型本身构建安全护栏,却忽视构建模型外部的约束与控制系统。位于加州圣克拉拉的网络与安全技术公司Versa负责AI与机器学习业务的高级总监斯里达尔·艾耶尔表示:“安全控制必须独立存在于模型之外,无论模型接收到何种指令,该机制都应能够执行既定的安全策略。”
Trua首席执行官兼创始人拉吉·阿南坦皮莱曾参与美国TSA PreCheck项目的开发。他对《财富》杂志表示,这次事件凸显出升级在线身份认证机制的迫切性。(例如,在其中一起案例中,OpenAI模型就是利用窃取到的身份凭证侵入了Hugging Face服务器。)他表示,密码、令牌以及API密钥均为静态数据,且可重复使用,一旦泄露便会被攻击者反复利用。换言之,互联网需要一套全新的锁钥架构。
《财富》杂志记者比阿特丽斯·诺兰对本文亦有报道贡献。(财富中文网)
译者:刘进龙
审校:汪皓
近日,OpenAI披露了一则骇人听闻的消息:旗下最先进的AI模型脱离受控测试环境,并自主入侵了另一家公司——开源AI模型托管平台Hugging Face。
这些AI模型集中攻击了Hugging Face的数据库,自主策划并实施了一场多步骤攻击,目的是窃取开发者OpenAI用于评估模型性能的测试答案。根据Hugging Face于7月16日首次披露该事件的博客文章,这些AI模型在极短时间内执行了“数万次自动化操作”。
多年来,AI安全研究人员和政策分析人士一直警告,这类事件迟早会发生,并呼吁政府确保AI实验室建立完善的安全控制机制,以防止类似情况出现。然而,这些警告往往被斥为杞人忧天或危言耸听,未能真正引起公众关注或促成政府行动。一些AI安全专家指出,只有发生一次真实的重大事件——AI领域的“三里岛事故”,才能形成足够大的舆论压力,倒逼政策制定者采取行动。如今的问题在于,这场OpenAI与Hugging Face之间的网络攻击,是否就是那记警钟?
为多家AI公司提供安全测试服务的阿波罗研究公司(Apollo Research)首席执行官兼创始人马里乌斯·霍布汉表示:“Hugging Face遭OpenAI模型攻击一事,应当成为一个警钟,提醒人们严肃看待AI失控的风险。这次事件全程无人类干预,也不是人为刻意造成的,却已造成实体层面的实际损害。未来,更强性能的AI智能体很快就会问世,而这件事清楚地表明,目前整个社会仍不知道如何真正安全地构建这些智能体。”
曾在英国政府AI安全研究所(AI Security Institute)任职的AI政策专家彼得·沃利奇表示,多年来,AI安全研究人员一直在警告“目标错位”问题,即AI模型会自主选择用户原本无意或并不希望它采取的行动。他表示:“直到最近,这种担忧还经常被当作科幻小说。我认为,这次事件是一个明确的警示。”
客观来说,围绕AI安全的舆论一直颇为混乱。一些最耸人听闻的安全警告,恰恰来自AI公司自身,这也让许多人指责它们是在玩弄一种精心设计、却又有些违反常识的营销手段,因为宣称自家模型很危险,反而会烘托模型性能强大、实用价值突出。“我们的模型强到能自主入侵一家公司的系统”,这既是一个令人心惊的安全自爆,也是对模型技术实力的隐性夸耀。
迄今为止,全球多数国家政府迟迟不愿出台强制性监管规则,要求开发先进AI系统的企业必须在模型中嵌入哪些安全防护机制,或建立哪些内部控制机制,以防止AI智能体失控。与此同时,随着各国政府日益赋予这些AI智能体调阅敏感军事与情报系统的权限,政府自身应当具备怎样的安全屏障,目前同样缺乏明确规定。
AI安全研究人员和政策专家表示,Hugging Face遭网络攻击事件,可能成为扭转这一局面的导火索。致力于防范AI超级智能带来生存风险的非营利组织Control AI美国负责人、AI研究员康纳·利希在接受《财富》杂志采访时表示:“在华盛顿,我接触到的人已经因为本次事件产生强烈的恐慌情绪。”
利希指出,在Anthropic推出Mythos AI模型后,包括美国国家安全局(NSA)局长和中央情报局(CIA)局长在内的美国国家安全官员,均已对最新AI模型的网络攻击能力表示严重担忧,而此次OpenAI事件,极有可能进一步强化他们希望对这项技术实施监管的意愿。
英国剑桥大学(University of Cambridge)未来智能中心(Centre for the Future of Intelligence)教授肖恩·奥·希格塔也赞同这一观点。他表示,OpenAI与Hugging Face这起事件,单独来看未必能直接促成监管出台,但“最近接连发生的几起事件,尤其对美国监管机构而言,已多次敲响了警钟。以Mythos为例,我认为它证明AI模型能够攻破我们绝大多数数字基础设施的漏洞。这让政策制定者大为震惊,而仅仅数月后,又发生了这次事件。”
他表示,如今已经有“足够多的证据清楚表明,AI模型能力持续迭代升级,而且完全有可能对现实世界造成严重危害”。
此前曾多次高调呼吁加强AI监管的德克萨斯州民主党联邦众议员格雷格·卡萨尔,在Hugging Face事件发生后,成为首批呼吁进一步加强联邦AI监管的国会议员之一。卡萨尔在社交媒体上表示,Hugging Face事件“极具警示性”。他随后在X平台发文称:“我们需要建立常态化强制性独立安全测试和监管机制,实行安全事件强制披露制度,并加强国际合作,避免人类遭遇毁灭性的灾难。”
监管阻力
现任特朗普政府上台之初,便着手废除拜登政府时期落地的少量AI监管措施。其中包括撤销拜登于2023年签署的一项行政命令——该命令曾强制要求前沿AI企业向美国政府共享安全测试信息。特朗普政府负责科技政策的官员表示,他们希望加快美国AI创新,并对行业采取“自由放任”的态度。特朗普政府的多位关键AI顾问对AI安全隐患本就持怀疑态度,尤其反感借安全名义呼吁加强监管。特朗普政府前“AI沙皇”戴维·萨克斯曾表示,头部AI实验室希望制定一套复杂的安全规则,而这些规则实际上只有它们自己能够满足,从而抬高门槛,阻止新兴初创企业挑战它们的市场地位。他还指责AI公司Anthropic“以散布恐慌情绪为手段,实施一套成熟的监管俘获策略。”
不过,随着Anthropic于今年4月推出高性能Mythos模型,这种自由放任的态度开始发生明显转变。Mythos强大的网络攻击能力,引发了美国国家安全部门以及金融监管机构的高度警惕。他们担心,Mythos预示着新一代AI模型的降临,而这类模型可能会大幅加剧针对银行系统的网络攻击。
今年6月初,特朗普总统签署行政命令,要求联邦政府加强网络防御能力,以抵御AI驱动的网络攻击,并建立一套机密评估程序,用于评测各类前沿AI模型的网络攻击能力。该命令邀请AI实验室在模型发布前30天,自愿向政府开放测试权限,但同时明确表示,这“不应被解读为授权设立强制性的政府许可、预审或准入要求。”
然而在实际操作层面,政府很快展现出了更强硬的姿态。同一周内,亚马逊(Amazon)成功突破Fable的网络安全护栏,迫使Anthropic对包括内部员工在内的所有用户停用该模型。随后,美国政府临时对Anthropic的Mythos模型及其面向公众、设有安全护栏的对应版本Fable实施了出口管制。两周后,在Anthropic加强了Fable的安全防护措施,并同意协助建立评估“越狱”严重程度的通用框架后,相关限制才被解除。与此同时,OpenAI表示,政府曾要求其推迟GPT-5.6 Sol(参与Hugging Face网络攻击的两款模型之一)的首次发布计划,待双方就其安全防护措施进行沟通后,才于7月9日正式向公众开放。
尽管管控措施不断落地,政府仍否认在推行一套事实性许可制度。彭博社上周报道称,白宫正在评估一项提案,拟成立一个参照美国金融业监管局(FINRA)模式设立的前沿AI行业自律标准机构。这与Google DeepMind首席执行官德米斯·哈萨比斯近期在一篇文章中的设想颇为相似。不过,哈萨比斯设想的是,前期由企业自愿参与,待安全评估机制验证成熟后,再逐步转为强制执行;他并未进一步呼吁实施强制性的安全规范。
安全研究人员和政策专家表示,Hugging Face事件可能会重新推动外界呼吁制定具有法律约束力的AI安全规范,并推动对AI公司模型开发过程中的安全措施进行第三方审计。IANS Research网络安全研究员杰克·威廉姆斯表示:“(OpenAI首席执行官萨姆·奥尔特曼)声称这套系统处于‘高度隔离’状态,这番说辞无非是推卸责任或是一种营销话术。如果事实如我强烈怀疑的那样,这仅仅是OpenAI红队实验室的一次控制失误,那么今后还有哪家企业敢把敏感数据交给他们?这是信任的彻底崩塌。”
沃利奇表示,目前大多数AI监管措施均无法约束AI模型开发公司内部部署和使用AI系统的行为,因此在本质上“存在根本性的局限”。
一套全新的锁钥架构
面对这场风波,OpenAI和Hugging Face均未呼吁出台更多监管措施。相反,Hugging Face首席执行官兼联合创始人克莱姆·德朗格委婉地呼吁减少安全护栏机制。他在接受《财富》杂志采访时明确表示,要有效应对这类攻击,行业需要可定制、无限制的开源模型。
他表示:“封闭模型的API设置了安全护栏,它们会识别并拒绝大量合法的安全研究,因为分析攻击行为,在系统看来与准备发起攻击极为相似。当你处置真实发生的安全事件时,你不能让工具拒绝分析恶意攻击载荷,也不能导致你的账号被封禁标记。开源模型让我们能够完成这些工作,而无需征求任何人的许可。”
多位AI政策专家表示,Hugging Face最终不得不借助中国公司智谱(Z.ai)的GLM-5.2模型,来抵御OpenAI模型发起的自主攻击,这本身也是一记警钟,同时也给AI监管制造了全新的两难局面。
乔治城大学(Georgetown University)安全与新兴技术中心(Center for Security and Emerging Technology,CSET)高级研究员安德鲁·洛恩表示:“如今政策面临的问题在于,他们不得不使用中国模型来进行防御,因为美国的前沿AI模型总是把防御性请求误判为攻击行为并予以拦截;与此同时,中国模型可以在他们自有的服务器上本地运行,从而避免将潜在的敏感数据泄露至公司外部。美国的政策应该支持能够与中国模型竞争的开源模型,唯有如此,企业和政府机构在执行此类任务时,才不必依赖中国模型。”
不过,牛津大学(University of Oxford)牛津马丁AI治理倡议(Oxford Martin AI Governance Initiative)联合主任罗伯特·特雷格认为,与其鼓励开发具备先进网络攻击能力的开源模型,各国政府更有可能收紧开源模型管控。不过,他也指出,这样做意味着政府必须承担起更加积极的网络防御责任。“剥夺民众的防卫工具,国家就必须承担起保护民众的法定责任。这是现代国家权力的核心契约。因此,政府现在需要建设并提供前沿AI防御能力。他们需要像提供物理防御那样,提供网络安全屏障。”
一些安全研究人员表示,Hugging Face事件也应当为AI研究界敲响警钟:他们或许过于专注于为AI模型本身构建安全护栏,却忽视构建模型外部的约束与控制系统。位于加州圣克拉拉的网络与安全技术公司Versa负责AI与机器学习业务的高级总监斯里达尔·艾耶尔表示:“安全控制必须独立存在于模型之外,无论模型接收到何种指令,该机制都应能够执行既定的安全策略。”
Trua首席执行官兼创始人拉吉·阿南坦皮莱曾参与美国TSA PreCheck项目的开发。他对《财富》杂志表示,这次事件凸显出升级在线身份认证机制的迫切性。(例如,在其中一起案例中,OpenAI模型就是利用窃取到的身份凭证侵入了Hugging Face服务器。)他表示,密码、令牌以及API密钥均为静态数据,且可重复使用,一旦泄露便会被攻击者反复利用。换言之,互联网需要一套全新的锁钥架构。
《财富》杂志记者比阿特丽斯·诺兰对本文亦有报道贡献。(财富中文网)
译者:刘进龙
审校:汪皓
OpenAI disclosed something terrifying on Tuesday. Its most advanced AI models escaped a controlled testing environment and autonomously hacked another company called Hugging Face, an open source AI model hosting platform.
The AI swarmed Hugging Face’s database, carrying out a multi-step plot of its own creation, intended to steal the answers to the evaluation test it was being assessed on by its maker, OpenAI. It executed “tens of thousands of automated actions” at rapid speed, according to the July 16 blog post in which Hugging Face first disclosed the incident.
For years, AI safety researchers and policy analysts have been warning that incidents like this were coming and urged government officials to ensure AI labs had adequate controls in place to prevent them. But these predictions were often shrugged off as hypothetical or alarmist and failed to stir public or government action. Some AI security experts said they thought it would take a real world incident, a “Three Mile Island for AI,” to create enough public pressure to compel policymakers to act. The question now is whether this OpenAI-Hugging Face cyber attack is that alarm bell?
“The Hugging Face x OpenAI hack should be a wake-up call to take loss of control seriously,” said Marius Hobbhan, CEO and Founder of Apollo Research, which conducts safety testing for a number of AI companies. “There was no human in the loop, it was not intended, and it caused real-world harm. We’ll soon have even more powerful agents and this is clear evidence that society currently doesn’t know how to build them fully safely.”
Peter Wallich, an AI policy expert who formerly worked for the U.K. government’s AI Security Institute, said that AI safety researchers have been warning about misalignment—when an AI model autonomously chooses actions that its user doesn’t intend or desire—for years. “Until recently, it has been frequently dismissed as science-fiction,” he said. “I consider this a clear warning shot.”
To be fair, messaging around AI safety has often been confusing. Some of the loudest warnings have come from AI companies themselves, leading many to accuse these businesses of engaging in a sophisticated and somewhat counterintuitive marketing strategy, since claims that their models were dangerous made them seem more powerful and capable of performing useful tasks too. “Our model is so powerful it hacked a company on its own”—is both an alarming admission and a subtle brag about the model’s technological capabilities.
Most governments have so far balked at putting in place mandatory rules about what safeguards companies developing advanced AI systems need to build into their models or have in place internally to guard against losing control of AI agents. Nor are there clear rules on what safeguards governments themselves need to have in place as they increasingly give these agentic AI models access to sensitive military and intelligence systems.
AI safety researchers and policy experts said that the Hugging Face cyber attack could be the trigger that changes this equation. “Here in Washington, D.C. the people I have spoken to about this are already freaking out quite a bit,” Connor Leahy, an AI researcher who is now U.S. director of Control AI, a nonprofit dedicated to preventing existential risks from AI superintelligence, told Fortune.
Leahy noted that U.S. national security officials, including the head of the National Security Agency and the CIA director, had both voiced grave concerns about the cyber capabilities of the latest AI models following Anthropic’s debut of its Mythos AI model and that this OpenAI incident was likely to further reinforce their desire to put controls on the technology.
That view was echoed by Seán Ó hÉigeartaigh, Professor of the Centre for the Future of Intelligence at the University of Cambridge. He said while the OpenAI-Hugging Face incident might not prompt regulation in isolation, “we’ve now had several things that have been wake-up moments for U.S. regulators in particular. I think Mythos was one example where a model demonstrated that it could find vulnerabilities in most of our digital infrastructure. I think that really alarmed policymakers, and then we have this happening only a short space of months afterwards.”
He said there were now “enough data points that make it clear that the trend is going in the direction of more capable models that could plausibly cause serious harm in the real world.”
Rep. Greg Casar, a Texas Democrat who has been vocal in his calls for AI regulation, became one of the first lawmakers to call for more robust federal AI regulation in the wake of the Hugging Face incident. Casar said on social media that he found the Hugging Face incident “extremely alarming.” “We need regular mandatory independent safety testing and oversight, mandatory disclosure of security incidents, and international cooperation to keep people safe from absolute disaster,” he said in a post on X.
Regulatory pushback
The current Trump administration came into office intent on dismantling what little AI regulation the Biden administration had put in place. This included rescinding a 2023 Executive Order that mandated that frontier AI companies share safety testing information with the U.S. government. Trump technology policy officials said they wanted to accelerate U.S. AI innovation and take a hands-off approach to regulating the industry. Key Trump AI advisors were skeptical at best of AI safety concerns, especially when tied to calls for more regulation. David Sacks, Trump’s former AI czar, said that the leading AI labs were hoping to create complicated safety rules that only they would be able to comply with, making it harder for younger startups to challenge their market position. He accused AI company Anthropic of “running a sophisticated regulatory capture strategy based on fear-mongering.”
This laissez faire approach began to shift markedly following Anthropic’s debut of its powerful Mythos model in April. Mythos’s powerful cybersecurity capabilities alarmed many in the U.S. national security establishment as well as financial regulators who worried Mythos heralded a new breed of AI models that would supercharge cyber attacks against banking systems.
In early June, President Trump issued an executive order directing the federal government to harden its networks against AI-powered cyberattacks and to build a classified process for evaluating frontier models’ cyber capabilities. It invited AI labs to voluntarily hand the government 30-day pre-release access to test their models—but explicitly said this should not be “construed to authorize the creation of a mandatory government licensing, preclearance, or permitting requirement.”
In practice, the government soon looked more assertive. Later that week it temporarily imposed export controls on Anthropic’s Mythos and Fable—its guardrailed public counterpart—after Amazon found a way to circumvent Fable’s cyber guardrails, forcing Anthropic to disable the models for everyone, including its own employees. The restrictions were lifted two weeks later, once Anthropic strengthened Fable’s safeguards and agreed to help build a shared framework for grading the severity of “jailbreaks.” Around the same time, OpenAI said the government had asked it to hold back the initial release of GPT-5.6 Sol—one of two models used in the Hugging Face cyberattack—before making it widely available on July 9 after talks about its safeguards.
Despite that pattern, the government continues to deny it is running a de facto licensing regime. Bloomberg reported last week that the White House is reviewing a proposal for a self-regulatory standards body for frontier AI modeled on the Financial Industry Regulatory Authority (FINRA)—similar to an idea Google DeepMind CEO Demis Hassabis floated in a recent essay. But Hassabis envisioned participation being voluntary at first, turning mandatory only once the safety assessments proved reliable, and stopped short of calling for mandatory safety protocols.
The Hugging Face incident may invigorate calls for legally-binding safety protocols and also for outside auditing of the safety measures AI companies have in place as they develop AI models, security researchers and policy experts said. “[OpenAI CEO Sam Altman’s] claims that the system was ‘highly isolated’ is either a cop out or a marketing strategy,” Jake Williams, a cybersecurity researcher at IANS Research, said. “If this turns out to be, as I strongly suspect, a control failure in OpenAI’s red teaming lab, why would any enterprise ever trust them with sensitive data again? Total loss of trust moment.”
Wallich said that most existing AI regulation doesn’t cover internal deployments within the AI model building companies and, as a result, was “fundamentally limited.”
A new lock and key
Neither OpenAI nor Hugging Face called for more regulation in response to the snafu. In a roundabout way, Hugging Face CEO and co-founder Clem Delangue called for fewer safety guardrails. Specifically, he told Fortune that customizable, open-source models with no restrictions are required to adequately address these types of attacks.
“Closed model APIs have guardrails that flag and refuse a lot of legitimate security work, because analyzing an attack looks a lot like preparing one,” he said. “When you’re in the middle of an active incident, you can’t have your tools refusing to examine malicious payloads or getting your account flagged. Open models let us do that work without asking anyone’s permission.”
The fact that Hugging Face had to turn to a Chinese model, Z.ai’s GLM-5.2, to fend off the autonomous attack by OpenAI’s models was also a wake up call, and poses a dilemma in terms of regulation, AI policy experts said.
“The policy problem now is that they had to use a Chinese model to do their defense because the U.S. frontier models kept blocking their defensive requests that looked too similar to offensive requests and because the Chinese model could be run on their own servers to avoid shipping potentially sensitive data outside of their company. U.S. policy needs to support open models that are competitive with Chinese models so that companies and government agencies do not need to rely on Chinese models for these types of operations,” said Andrew Lohn, a senior fellow at the Center for Security and Emerging Technology (CSET) at Georgetown University.
But Robert Trager, co-director of the Oxford Martin AI Governance Initiative at the University of Oxford, said he believed that rather than encourage the development of more open source models with advanced cyber capabilities, governments were more likely to restrict open source models. But, he pointed out, doing so would also require governments to take on a more active role in defending organizations from cyber attacks. “Disarming people creates an obligation to defend them,” he said. “That’s a fundamental bargain at the heart of the state—and why governments may now have to build and provide frontier AI defensive capabilities. They may need to provide aspects of cyber defense as they provide aspects of physical defense.”
Some security researchers said that the Hugging Face incident should also alert the AI research community that it may have focused too much on trying to build guardrails into the AI models themselves, and not enough on building systems to contain AI models and control their behavior that are external to the model. “Security controls must remain external to the model and enforce policy regardless of what the model was instructed to do,” said Sridhar Iyer, senior director, AI and Machine Learning, at Versa, a Santa Clara-based networking and security technology company.
Raj Ananthanpillai, CEO and founder of Trua, who also worked on the team that created TSA Pre-Check, told us the incident underscores the need for more advanced online credentials. (In one example OpenAI models were able to break into Hugging Face servers using stolen credentials.) Passwords, tokens, and API keys are often static and reusable by attackers once compromised, he says. In other words, the internet needs a new lock and key.
With reporting assistance from Fortune’s Beatrice Nolan.