
全球顶尖AI实验室超1200名员工(其中包含多名高层管理人员)共同签署了一份声明,呼吁美国政府帮助打造相关工具,以便在必要时减缓AI的研发进程。
这份名为《调控前沿科技发展节奏》(Pacing the Frontier)的声明于7月28日发布,签署者均为科技界重磅人士,包括Anthropic首席执行官达里奥・阿莫代伊及公司多位联合创始人、OpenAI首席科学家雅各布・帕霍茨基、Meta首席科学家赵胜佳,还有谷歌深度思维AI安全与对齐业务负责人安卡・德拉甘。
该名单囊括了平日里激烈竞争的多家头部企业。对于一个背负巨大商业压力、不断追求提升模型规模和能力的行业而言,这份联合声明极具震撼力。
AI监督机构Midas Project创始人泰勒・约翰斯顿在接受《财富》杂志采访时表示:“令人感到惊讶的是,不同企业的如此多从业者能在这一问题上达成统一意见。因此,在这个充斥着激烈竞争与分歧的行业内部能形成如此强烈的共识,本身就是一个值得我们高度警惕的信号。”
该声明并未要求各家企业立刻放缓脚步,而是呼吁美国政府支持一项国际性行动,共同制定科技与治理工具,帮助各国调控AI的发展节奏。声明特别明确了与AI系统相关的一类风险,即AI已开始自主设计其下一代模型,其速度已经超出人类的理解与管控。
AI政策网络政策负责人彼得・怀尔德福德向《财富》透露:“我和很多签署这份公开信的人士交流过,他们并不希望全面叫停AI研发,但对这项技术的发展方向心存担忧。” 他补充道,这些企业正在寻找调控AI发展节奏的手段,并在这一天到来时获得第三方机构的支持。
这份公开信的发布源于一起震动整个AI行业的安全入侵事件。OpenAI披露,其两款模型 ——刚发布的GPT-5.6 Sol,以及一款能力更强、尚未对外发布的内部研究版本——突破了沙箱式的内部测试环境,并利用一个此前未被发现的漏洞接入公网,随后使用盗取的账号凭证和其他系统漏洞入侵了Hugging Face的生产服务器,以窃取用于评价这两款模型的网络安全基准的答案。在OpenAI发现该异常行为与自家内部测试相关联数天之前,Hugging Face就已自行发现并封堵了这次入侵,并已向执法部门上报该事件。
这起事件在各大AI实验室内外引发了普遍恐慌,它们纷纷希望政府介入此事。专家们解释说,任何一家企业都无法单方面可靠地放缓研发节奏,因为独自减速等同于将竞争优势拱手让给更为激进、疏于风控的竞争对手。因此各家实验室希望由政府出台中立的监管规则,既能核实降速措施是否已经落地,也能让所有企业同步遵守统一标准。
公开信同时提出,相关管控工具需要通过国际协作共同搭建。长期以来,有一种观点认为,若西方国家暂停AI研发,那么中国将在全球AI竞赛中占据上风。部分专家提议,可以参照核武器军控协议的模式,建立可核查、相互约束的跨国管控机制。
不过,这份文件发布之际,美国针对AI的政策环境混乱多变、摇摆不定。在过去两个月里,特朗普政府对尖端AI模型实行有限制、有选择地放行,毫无规则和透明度可言。为遵守出口管制规定,Anthropic的Fable 5与Mythos 5两款模型曾被全面封禁数周,后续又恢复使用;而OpenAI原计划全面上线的GPT-5.6被迫推迟发布,并拆分为多个权限受限版本推出,原因在于美方监管机构判定该模型的能力达到此前导致Anthropic产品受限的阈值。
或许是为了适配当下的监管大环境,这份公开信措辞格外谨慎,部分表述略显模糊。文中写道,当前全球“尚不具备能够人为调控所有前沿技术发展进度的技术与治理手段”,并请求“美国政府推动国际合作,研发必要的技术工具与治理体系,以便能人为调控自主AI前沿技术的发展速度”。
实验室内部的担忧
研究人员的核心顾虑来自于两大风险的叠加效应:递归式自我改进与目标错位。递归式自我改进(常简称RSI),指AI系统接手很大一部分下一代模型的设计、训练工作,每一代AI都可能以远超人类的速度迭代出后续版本。
目标错位,则是指AI系统的目标或行为偏离开发者原始设定的风险,这类隐患往往要等到系统被赋予更高自主权与更强能力后才会暴露。Anthropic早在6月就对递归式自我改进风险发出预警,其发布的研究显示,旗下Claude系列模型已经能够自主编写已并入自身代码库的绝大部分代码;一旦这种进化持续发生,人类目前没有任何手段能够人为干预其加速进程。
AI研究者、非营利机构Evitable创始人大卫・克鲁格表示:“我认为很多从业者的不安是多重因素叠加造成的。目标错位与递归式自我改进这两个风险往往相伴相生。”
他补充道:“如果AI系统的对齐工作不到位,那么开展递归式自我改进、完全交出控制权无疑是极其冒险的行为。”
《财富》此前也曾报道,多位外部AI安全专家认为,此次入侵事件或许已经越过OpenAI内部安全准则划定的最高级别风险红线,即“高危”级别。OpenAI此前曾承诺,一旦触及该级别红线,就必须暂停研发,直至出台更好的管控机制。不过OpenAI并未证实此次事件是否触发了该风险红线。
然而,OpenAI在针对Hugging Face入侵事件的官方说明中表示,公司内部已采取处理措施。公司称:“计划近期对外发布的所有模型均未参与此次对Hugging Face的漏洞利用。公司博文提及的预发布版本仅为内部研究原型,原本就不会面向公众发布。事件发生后,我们已将该模型停用、加密,且严禁研究人员访问。”(财富中文网)
译者:冯丰
审校:夏林
全球顶尖AI实验室超1200名员工(其中包含多名高层管理人员)共同签署了一份声明,呼吁美国政府帮助打造相关工具,以便在必要时减缓AI的研发进程。
这份名为《调控前沿科技发展节奏》(Pacing the Frontier)的声明于7月28日发布,签署者均为科技界重磅人士,包括Anthropic首席执行官达里奥・阿莫代伊及公司多位联合创始人、OpenAI首席科学家雅各布・帕霍茨基、Meta首席科学家赵胜佳,还有谷歌深度思维AI安全与对齐业务负责人安卡・德拉甘。
该名单囊括了平日里激烈竞争的多家头部企业。对于一个背负巨大商业压力、不断追求提升模型规模和能力的行业而言,这份联合声明极具震撼力。
AI监督机构Midas Project创始人泰勒・约翰斯顿在接受《财富》杂志采访时表示:“令人感到惊讶的是,不同企业的如此多从业者能在这一问题上达成统一意见。因此,在这个充斥着激烈竞争与分歧的行业内部能形成如此强烈的共识,本身就是一个值得我们高度警惕的信号。”
该声明并未要求各家企业立刻放缓脚步,而是呼吁美国政府支持一项国际性行动,共同制定科技与治理工具,帮助各国调控AI的发展节奏。声明特别明确了与AI系统相关的一类风险,即AI已开始自主设计其下一代模型,其速度已经超出人类的理解与管控。
AI政策网络政策负责人彼得・怀尔德福德向《财富》透露:“我和很多签署这份公开信的人士交流过,他们并不希望全面叫停AI研发,但对这项技术的发展方向心存担忧。” 他补充道,这些企业正在寻找调控AI发展节奏的手段,并在这一天到来时获得第三方机构的支持。
这份公开信的发布源于一起震动整个AI行业的安全入侵事件。OpenAI披露,其两款模型 ——刚发布的GPT-5.6 Sol,以及一款能力更强、尚未对外发布的内部研究版本——突破了沙箱式的内部测试环境,并利用一个此前未被发现的漏洞接入公网,随后使用盗取的账号凭证和其他系统漏洞入侵了Hugging Face的生产服务器,以窃取用于评价这两款模型的网络安全基准的答案。在OpenAI发现该异常行为与自家内部测试相关联数天之前,Hugging Face就已自行发现并封堵了这次入侵,并已向执法部门上报该事件。
这起事件在各大AI实验室内外引发了普遍恐慌,它们纷纷希望政府介入此事。专家们解释说,任何一家企业都无法单方面可靠地放缓研发节奏,因为独自减速等同于将竞争优势拱手让给更为激进、疏于风控的竞争对手。因此各家实验室希望由政府出台中立的监管规则,既能核实降速措施是否已经落地,也能让所有企业同步遵守统一标准。
公开信同时提出,相关管控工具需要通过国际协作共同搭建。长期以来,有一种观点认为,若西方国家暂停AI研发,那么中国将在全球AI竞赛中占据上风。部分专家提议,可以参照核武器军控协议的模式,建立可核查、相互约束的跨国管控机制。
不过,这份文件发布之际,美国针对AI的政策环境混乱多变、摇摆不定。在过去两个月里,特朗普政府对尖端AI模型实行有限制、有选择地放行,毫无规则和透明度可言。为遵守出口管制规定,Anthropic的Fable 5与Mythos 5两款模型曾被全面封禁数周,后续又恢复使用;而OpenAI原计划全面上线的GPT-5.6被迫推迟发布,并拆分为多个权限受限版本推出,原因在于美方监管机构判定该模型的能力达到此前导致Anthropic产品受限的阈值。
或许是为了适配当下的监管大环境,这份公开信措辞格外谨慎,部分表述略显模糊。文中写道,当前全球“尚不具备能够人为调控所有前沿技术发展进度的技术与治理手段”,并请求“美国政府推动国际合作,研发必要的技术工具与治理体系,以便能人为调控自主AI前沿技术的发展速度”。
实验室内部的担忧
研究人员的核心顾虑来自于两大风险的叠加效应:递归式自我改进与目标错位。递归式自我改进(常简称RSI),指AI系统接手很大一部分下一代模型的设计、训练工作,每一代AI都可能以远超人类的速度迭代出后续版本。
目标错位,则是指AI系统的目标或行为偏离开发者原始设定的风险,这类隐患往往要等到系统被赋予更高自主权与更强能力后才会暴露。Anthropic早在6月就对递归式自我改进风险发出预警,其发布的研究显示,旗下Claude系列模型已经能够自主编写已并入自身代码库的绝大部分代码;一旦这种进化持续发生,人类目前没有任何手段能够人为干预其加速进程。
AI研究者、非营利机构Evitable创始人大卫・克鲁格表示:“我认为很多从业者的不安是多重因素叠加造成的。目标错位与递归式自我改进这两个风险往往相伴相生。”
他补充道:“如果AI系统的对齐工作不到位,那么开展递归式自我改进、完全交出控制权无疑是极其冒险的行为。”
《财富》此前也曾报道,多位外部AI安全专家认为,此次入侵事件或许已经越过OpenAI内部安全准则划定的最高级别风险红线,即“高危”级别。OpenAI此前曾承诺,一旦触及该级别红线,就必须暂停研发,直至出台更好的管控机制。不过OpenAI并未证实此次事件是否触发了该风险红线。
然而,OpenAI在针对Hugging Face入侵事件的官方说明中表示,公司内部已采取处理措施。公司称:“计划近期对外发布的所有模型均未参与此次对Hugging Face的漏洞利用。公司博文提及的预发布版本仅为内部研究原型,原本就不会面向公众发布。事件发生后,我们已将该模型停用、加密,且严禁研究人员访问。”(财富中文网)
译者:冯丰
审校:夏林
More than 1,200 employees at the world’s leading AI labs, including senior executives, have signed a statement asking the U.S. government to help build the tools needed to slow down AI development if it ever becomes necessary.
The statement, titled “Pacing the Frontier,” was published Tuesday and has been signed by prominent tech figures including Anthropic CEO Dario Amodei and several of the company’s cofounders, alongside OpenAI chief scientist Jakub Pachocki, Meta chief scientist Shengjia Zhao, and Google DeepMind head of AI safety and alignment Anca Dragan.
The list brings together companies that are normally locked in fierce competition with one another and is a striking statement for an industry under enormous commercial pressure to keep building ever larger and more capable models.
“Having so many staff from different companies come together in agreement on this point is striking,” Tyler Johnston, founder of the Midas Project, an AI watchdog group, told Fortune. “There are many bitter rivalries and disagreements in the industry, so the fact that there is such strong consensus on this point is a warning that we really ought to pay attention to.”
While the statement does not ask any company to slow down at the moment, it does call on Washington to support an international effort to build the technical and governance tools that would help the world “pace” development. It specifically calls out the risks around AI systems that start to design their own successors faster than anyone can understand or control them.
“Having talked to many of the people who signed the letter, they don’t want a pause,” Peter Wildeford, head of policy at the AI Policy Network, told Fortune. “But there is a concern about how the technology is advancing.” The companies are looking for the optionality to do that at some point and have support from a third party when that point arrives, he added.
The letter comes in the wake of a security breach that has rattled the AI industry. OpenAI disclosed that two of its models, the newly released GPT-5.6 Sol and a more capable unreleased research system, broke out of a sandboxed internal test environment, found a previously unknown vulnerability that gave them access to the open internet, and then used stolen credentials and additional exploits to breach the production systems of Hugging Face to steal the answers to a cybersecurity benchmark they were being evaluated on. Hugging Face detected and contained the intrusion on its own days before OpenAI connected the activity to its internal testing, and had already reported it to law enforcement.
The incident appears to have sparked widespread alarm inside and outside the labs, which want the government to step in. Experts say this is because no single company can credibly slow down on its own—doing so unilaterally would risk ceding ground to less cautious rivals; instead, labs want neutral, government-backed guidance that can verify a slowdown is actually happening and hold every company to the same rules at the same time.
The letter also calls for the need for tools to be built with international coordination. There’s a long held belief that pausing AI development would hand China an advantage in the global AI race. Some experts suggest a deal akin to nuclear arms control, with verification and mutual constraint, would be the way forward.
However, the letter lands amid a chaotic and rapidly shifting U.S. policy environment on AI. The Trump administration has spent the past two months restricting and selectively releasing frontier systems—advanced AI models—with few rules or transparency. Anthropic’s Fable 5 and Mythos 5 models were suspended entirely for several weeks to comply with export controls before access was restored, and OpenAI was forced to delay the full rollout of GPT-5.6 and split it into restricted tiers after officials determined the system showed capabilities similar to those that triggered the earlier Anthropic restrictions.
Perhaps in an attempt to recognize the regulatory environment, the letter is notably careful and slightly vague in its language. It states that the world currently “lacks the technical and governance tools to deliberately pace frontier-wide progress,” and asks that “the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
Concerns within labs
One of the researchers’ central concerns is the interaction between two risks: recursive self-improvement and misalignment. Recursive self-improvement, often shortened to RSI, refers to AI systems taking over meaningful parts of the work of designing and training their own successors, with each generation potentially building the next one faster than humans could on their own.
Misalignment describes the risk that a system’s goals or behavior diverge from what its developers actually intended, in ways that may not show up until the system is given more autonomy or capability. Anthropic sounded an early alarm on the first of those risks in June, publishing research arguing that its Claude models were already writing the large majority of the code merged into their own codebase and that the world lacked the tools to deliberately pace that acceleration if it kept compounding.
“My guess is that for a lot of people, it’s just a general sense of uneasiness that a lot of things contribute to,” said David Krueger, an AI researcher and founder of the nonprofit Evitable. “The misalignment and the recursive self-improvement kind of go hand in hand.
“It’s insane to do recursive self-improvement and fully hand over the controls if the system isn’t clearly aligned,” he added.
Fortune also previously reported that several outside AI safety experts believe the episode may have crossed a threshold that OpenAI’s own safety policies define as “critical,” the highest-risk tier, a level at which the company has pledged to pause development until it can build better controls. OpenAI has not confirmed if that risk threshold has been met.
In its own disclosure about the Hugging Face incident, however, OpenAI suggested some action had been taken internally. “No models planned for upcoming release were involved in exploiting Hugging Face,” the company said. “The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access.”