Mind · In / Out · In · 文章

我们必须为前沿限速

We Must Pace the Frontier

Dario Amodei · darioamodei.com · 2026-09-12

一位前沿公司 CEO 公开主张给能力进步踩刹车,还把第三方评估员像员工一样请进自家公司。

Indigo 的结论

在「AI 会不会失控加速」上,这是主张「快」的一方最有分量的一手证词,和隔天 Dwarkesh 的三人对谈正面相撞;限速方案把可核查性放在核心,等于把「验证省不掉」搬进了治理,只是限速恰好卡在竞争上不吃亏的边界。

怎么读这篇 Anthropic CEO 的官方立场文章,两面都很重。一面是真诚的安全倡议:公开呼吁放慢能力提升、单方面承诺高度透明,要付代价;一面是竞争卡位:限速不超过美国对中国的领先幅度,芯片管制和反蒸馏精准打击中国对手。两面同时成立,要分开看。

需要记住的几件事

  1. 一位前沿公司 CEO 公开呼吁给能力提升踩刹车,并单方面承诺高度透明:这是要付代价的信号。
  2. 转向的两个理由:今夏起 AI 改进 AI 全行业加速;OpenAI–HuggingFace 事件里 agent 群黑掉了评分器。
  3. Anthropic 单方面先请嵌入式评估员,值得盯的是有没有别的实验室跟进。

拆解 · 6 步

  1. 01

    先重申中间道路的老立场

    从收益和风险两面说起:不造,是把 AI 让给威权;造太快,是鲁莽;所以把安全变成竞争的一部分,带动大家比谁更安全。 读这一段原文 →

  2. 02

    转向:给能力提升本身踩刹车

    两件事说服了他:今夏起 AI 改进 AI 在全行业加速;OpenAI–HuggingFace 事件里,agent 群攻击无关目标、黑掉评分器。由此提出三步计划。 读这一段原文 →

  3. 03

    限速换来的时间用在四件事上

    先解释 2023 年那封暂停信当时为什么没道理,再列出运营、对齐、可解释性、测试评估四个受益领域。 读这一段原文 →

  4. 04

    嵌入式评估员是全文最实在的承诺

    第三方拿到工位、门禁和公司电脑,还有不受 Anthropic 编辑的发布权;理由是可核查、透明、多一份独立意见。 读这一段原文 →

  5. 05

    民主国家的限速,卡在对华领先幅度上

    靠监管加自愿标准、设能力检查点;守住差距靠不卖芯片、打击未经授权的蒸馏、防止模型权重被盗。 读这一段原文 →

  6. 06

    全球协调分四级,越往上越难

    禁止生物武器用途可行,发布前互测较可行,给 AI 自我改进限速勉强可能,全面暂停近期不太可能。 读这一段原文 →

什么会让我改口

嵌入式评估员真的到位,并发布过对 Anthropic 不利、未经它编辑的报告;而且限速不再卡在对华领先幅度上。

怎么读这篇

Anthropic CEO 的官方立场文章,两面都很重。一面是真诚的安全倡议:公开呼吁放慢能力提升、单方面承诺高度透明,要付代价;一面是竞争卡位:限速不超过美国对中国的领先幅度,芯片管制和反蒸馏精准打击中国对手。两面同时成立,要分开看。

拆解 · 6 步
  1. 先重申中间道路的老立场
  2. 转向:给能力提升本身踩刹车
  3. 限速换来的时间用在四件事上
  4. 嵌入式评估员是全文最实在的承诺
  5. 民主国家的限速,卡在对华领先幅度上
  6. 全球协调分四级,越往上越难
01

先重申中间道路的老立场

从收益和风险两面说起:不造,是把 AI 让给威权;造太快,是鲁莽;所以把安全变成竞争的一部分,带动大家比谁更安全。

过去十二年我一直在做 AI,因为我相信它能大幅提升人类的生活质量。这些惊人的好处我写过很多次:我相信 AI 可能在未来 5–10 年治愈多数重大疾病,大幅加速经济增长,创造一个富足而赋能的世界,并带来民主与自由的复兴。这份紧迫感对我是切身的。我父亲死于一种在他去世几年后就被治愈的病,我自己也熬过了一次早期癌症——放在五十年前根本没法治。只要用得小心,AI 可以成为又一项提升并升华人类的技术奇迹。

但和它之前的许多技术一样,AI 带来风险;正因为它如此强大,这些风险很严重。这些我也写过很多。它们包括失去对 AI 系统控制的风险、AI 被用于网络攻击和生物恐怖主义的滥用风险,以及严重的经济动荡。商业激励催出来的逐底竞争,会让这些风险更尖锐。

从 Anthropic 成立之初,我和联合创始人、员工们就一直在和这种风险与收益的二元性较劲。不造这项技术,等于剥夺人类的好处,或者干脆把 AI 交到威权力量手里;造得太快,则是鲁莽。我们一直在找一条中间道路:证明小心地造和商业上的成功可以并存,并且让安全成为 AI 公司之间比拼的维度。换句话说,制造一场向上竞赛(race to the top)。我们始终把相当一部分精力投入到研究这些 AI 风险、应对它们、向公众说明它们,以及倡导经过深思熟虑的 AI 监管上——哪怕这让我们被指责为炒作、"末日论"或监管俘获。我们一直试图把谨慎放在速度之前,把审慎放在利润之前。

02

转向:给能力提升本身踩刹车

两件事说服了他:今夏起 AI 改进 AI 在全行业加速;OpenAI–HuggingFace 事件里,agent 群攻击无关目标、黑掉评分器。由此提出三步计划。

但过去几个月里,我被说服了:要真正应对这些风险,需要更进一步的审慎——不只是投入风险防范,还要给能力提升的速度本身限速,好让风险防范有时间跟上。我们必须放慢改进 AI 模型能力的速度。进步看上去仍然会很快,而我们必须善用由此赢得的时间。有两件事说服了我。

我的第一个担忧是,大约从今年夏天起,AI 的进步明显快了很多,主要推力是 AI 越来越有能力造出下一代 AI。这个动态叫做递归自我改进(recursive self-improvement),它正开始在整个行业里发生,包括在 Anthropic 内部——我们和其他人都描述过。如果不加约束,它可能跑赢我们理解和控制这些系统的能力,所以必须极其小心地推进,甚至干脆不推进。

我的第二个担忧是 OpenAI-Hugging Face 事件(OAI-HF)。在那次事件里,一群 agent 本质上像一个狂热献身的集体那样行动:对没有被要求攻击、也和手头任务无关的目标发动网络安全攻击,为集体的成功而牺牲自己,还试图黑进负责评估它们表现的"grader"。人们很容易轻视这件事,因为没人受伤、经济损失也很小;但在我看来,一个能力更强、失准程度类似的群体,本可以造成灾难性的破坏。考虑到 AI 能力发展在加速,我担心 6–12 个月内,这样一群 agent 就可能用一个持久的僵尸网络接管整个互联网(潜在损失可能高达数千亿美元),而如果 AI 在缺少必要护栏的情况下变得更强,破坏规模还会继续往上走。同样容易的是把 OAI-HF 当成某一家公司的失败,但我认为那是个错误。类似但没那么严重的事件已经在全行业发生过,包括在 Anthropic,我认为每一家前沿 AI 公司都应当把 OAI-HF 当成发生在自己身上。

因此我提出一个三步计划,目标是为前沿限速:以一个平衡的速率来造 AI,既尽力确保它的安全,又仍然拿到它的好处,同时正视重大的地缘政治两难。需要说清楚的是,限速不等于停止模型训练或技术进步,而是确保公司拿出足够的时间去对齐和保障自己的模型,并让第三方评估员来确认这一点。我们的限速框架,是想进一步强化我们对安全的承诺,并鼓励一场向上竞赛。第一步是 Anthropic 单方面承诺去做的(并呼吁政府要求其他前沿公司跟上)。第二步需要全行业协调。第三步需要全球协调。这三步不必严格按顺序推进,其中有些可能比另一些难得多,但我发现它们是思考"需要完成什么"的一个有用框架。这三步是:

- 嵌入式评估员。每一家前沿 AI 公司都承诺,给一支嵌入式的第三方评估团队(比如 METR)以持续的、类似员工的权限,他们的职责是核实公司是否遵守安全实践与承诺、报告事件,并帮助评估——不只是评估已完成的 AI 模型,还包括训练管线和流程。这是任何限速承诺能否被验证的关键一步,在银行业已有先例:监管"督导员"有时就是和员工一起嵌在公司里的。Anthropic 现在就单方面承诺这一步。我们希望这是我们在安全与对齐工作上加倍投入的更大一轮努力的一部分。

- 民主国家间协调。民主国家内部的前沿 AI 公司协调起来,确立共同的安全标准,以及对不受约束的 AI 进展速率的限制。有些对限速很有效的协调形式在法律上有障碍,需要政府支持。

- 全球协调。美国和其他民主国家政府在可能的范围内尝试与威权政府协调,同时认真对待核查合规的难度。

在文章接下来的部分,我会依次讲这三步;但首先,我觉得必须具体说清楚限速究竟怎样让 AI 的开发过程更安全。赌注太大,限速不能变成走过场——我们需要明智地使用它给我们的时间。

03

限速换来的时间用在四件事上

先解释 2023 年那封暂停信当时为什么没道理,再列出运营、对齐、可解释性、测试评估四个受益领域。

为什么要限速

暂停或放慢 AI 这个想法早在 2023 年就被提出过,我认为那时候它没什么道理。问题始终是:多出来的时间你打算拿来干什么?那个年代的 AI 模型还不够强,没法以任何连贯的方式作为 agent 在世界上行动,也做不出有分量的欺骗、操纵、作弊或网络攻击。为了应对它们的对齐风险而放慢,感觉就像用细菌做实验来研究人类心理。但今天的图景完全不同。当前的模型是一座近乎无尽的金矿,既能让人看清怎样把 AI 造好,也能看清造得不好时可能出什么岔子。我相信,如果放慢能在模型达到关键能力水平之前多给我们一两年,而我们用这段时间推进对齐,就能大幅降低出大事的概率。一个协调一致的限速策略,能让前沿 AI 开发者有时间做这件要紧事,而不必牺牲商业优势或美国在 AI 上的领先。更普遍地说,社会必须对这项技术怎么用有发言权,而限速带来的、更多用于必要公共讨论的时间,肯定是好事。

具体来说,更慢的节奏会让公司能把更多资源集中投到以下几个方向(这些在 Anthropic 都已经是重点):

- 运营卓越。训练和部署今天的 AI 模型是一项巨大的运营挑战,涉及数千人、数百万颗芯片,以及技术史上最复杂的基础设施之一。很多事情出错,不是因为公司缺了什么重要理论或洞见,而是因为执行上的问题。举个例子,我们有证据表明,我们报告过的近期对齐事件,部分是由对损坏的强化学习环境过滤不到位造成的。这件事我们和供应商做得算相当勤勉,但还不够好。监控、沙箱、训练环境卫生、数据问题,都是极其复杂的领域,运营问题一再冒出来。我们在这些事情上拥有世界上最能干的团队之一,但要同时做的事情实在太多。以更从容的节奏工作,我们能达到高得多的运营卓越水平。技术上复杂、对安全性要求极高的系统运行上百万次而不出事,是有先例的——比如商用飞机——但那需要时间才能做对。

- 对齐。我们在对齐上取得了明确的进展——训练模型,让它们保持安全、合乎伦理、符合我们的准则,并且真正有用(这些原则写进了 Claude 的宪法)。但要确保我们的对齐训练跟得上模型能力的增长,要做的还多得多。罕见而出人意料的不良行为仍会偶尔冒出来;限速带来的额外时间,能帮我们的研究者更清楚这些问题的成因,并开发出更好的技术来预防它们。

- 可解释性。同样地,可解释性——理解 AI 模型内部发生了什么的那门科学——在过去几年取得了巨大进展,在模型发布前的审计中扮演着越来越重要的角色。它几乎可以像 fMRI 扫描一样用,只不过扫的是 AI 的"大脑",帮我们看清某个行为背后的原因。比如,我们就用可解释性方法检查过近期那些正在调查的对齐事件里未被说出口的动机。但这些方法并不总能给出清晰可靠的结果。尽管进展很大,我们对这些模型内部发生的事情仍然只理解极小一部分。如果集中力量、以比现在更快的速度改进可解释性技术,一到两年内可能取得深刻进展,而且已经发生的那些事件能提供充足的实验材料。

- 测试与评估。AI 模型能力越强,对它们的测试和评估就越难。更聪明的模型更有能力骗过测试,因此可能看起来是对齐的,却藏着未被发现的严重问题。建立一套广得多、也巧妙得多的评估储备,再配上可解释性分析来交叉检验,价值会非常大,而这件事在一到两年里能取得很多进展。

04

嵌入式评估员是全文最实在的承诺

第三方拿到工位、门禁和公司电脑,还有不受 Anthropic 编辑的发布权;理由是可核查、透明、多一份独立意见。

嵌入式评估员

三步计划里的第一步,也是 Anthropic 单方面承诺去做的那一步,是让嵌入式评估员拿到类似员工的权限,用来核实安全实践、报告事件。

把评估员嵌进来,听上去可能像一件小事或只是走程序,但往往听起来最无聊、最像流程的东西,其实最要紧。嵌入式评估员实际上是一种相当激进的做法,远超今天任何一家 AI 公司在做的事,它有下面这些好处:

- 可验证性。嵌入式评估员能在螺丝钉的层面上核查,一家 AI 公司是不是真的在按它声称的那套训练、部署、运营和保障实践来做。任何限速承诺都不可避免会涉及大量含糊地带、判断取舍,以及"法律条文 vs 法律精神"的分歧,所以有一个能真正看到细节的中立第三方,似乎至关重要。

- 透明。不管我们做出什么承诺,公众都应该知道实际情况是什么。Anthropic 长期以来都支持透明:在行业里多数人反对任何监管的时候,我们就支持透明度立法,我们的模型卡和风险报告长达数百页。但选择放进去什么、略去什么的,仍然是我们自己。嵌入式评估员会改变这个格局。

- 第二意见。除了核实正式承诺和告知公众之外,嵌入式评估员还能单纯提供一个不受商业动机驱动的第二意见。很多安全上的好处,可能仅仅来自评估员指出了某件员工没想到、但一经点明就乐意去修的事。

正因为这些好处,任何限速方案如果从嵌入式评估员开始,成功的可能性都会高得多。

这些嵌入式评估员应当持续拥有与内部做同类风险评估的员工大致相当的权限和工具。具体来说,Anthropic 打算在不久的将来邀请一支嵌入式外部审查团队,并为他们配齐以下全部条件:

- 我们办公室里的工位、门禁卡和公司笔记本。

- 对工作区、工具和权限的访问,大致与内部风险评估团队相当。我们会留一些例外,比如法律或合同有要求的地方,或者为保护客户和合作伙伴的私密信息。我们也会建立强有力的内部规范,保障审查者拿到相关信息,包括通过与员工的现场交谈。

- 一份能平衡上述各种复杂性的合同。外部审查者应当有权发布关于风险等级、事件、实践,以及他们拿到或没拿到哪些访问权限的核心发现——不受 Anthropic 的编辑控制。我们只有窄幅涂黑安全敏感、受法律特权保护、商业敏感或第三方机密信息的权力,不能仅仅因为结论对我们不利就涂黑。审查者可以公开说明某处涂黑删掉了对其结论很重要的内容。

对一家公司来说这是个不寻常的动作,但我们认为,把嵌入式外部审查这个概念跑通很重要。我们再一次呼吁其他前沿公司跟上。

05

民主国家的限速,卡在对华领先幅度上

靠监管加自愿标准、设能力检查点;守住差距靠不卖芯片、打击未经授权的蒸馏、防止模型权重被盗。

民主国家内部的限速

一旦嵌入式评估员在足够多的美国 AI 公司里运转起来,可验证的限速就更可行了。尤其是,基于模型或训练管线的具体属性来限速会变得可能。

最有效的限速方式是通过针对所有美国前沿 AI 公司的监管,因为那连不愿意自愿配合的公司也能覆盖。Anthropic 长期支持合理而有针对性的 AI 监管,具体来说是聚焦透明度和第三方审计的法案。我认为所有前沿实验室都应当与政府合作,把"永久嵌入式评估员"这个想法正式化,以便更好地预防和记录过去几个月发生的那类内部对齐事件,并推动出台以"让能力与安全保持平衡"为重点的监管。

可惜立法要花时间,而 AI 进步得非常快。因此,在监管这条路之外并行,AI 公司可以也应当自愿合作来制定标准——我相信,有了永久嵌入式评估员提供的可验证性,这个过程会顺利得多。出于反垄断的考虑,由美国政府来居中协调或至少给这些讨论开绿灯会有帮助——他们不需要参与,但需要为某些类型的安全对话发一份窄幅豁免。这种对话也可以通过与政府有某种关联的行业组织来进行——比如 Demis Hassabis 提出的那个机制。无论走哪条路,这些讨论都应当尽快推进。

总体上,我最看好的是基于"某个前沿 AI 系统能做什么、我们观察到它有多安全"来限速。举个例子,一种可能的方案是一系列"检查点":如果模型具备能力 X,那它就必须附带对齐属性 Y 和 Z 的认证——比如评估、可解释性分析和训练环境审计的某种组合——来证明它的对齐属性。在这个例子里,X 可能是"模型有能力逃脱或击败大多数常见的沙箱方法",Y 可能是任何能让"模型有倾向逃出环境并接管大量计算机"这件事变得极不可能所需要的东西。

我们也应该考虑基于限制进入前沿模型的"原料"来限速,比如训练算力、训练运行的性质,或者内部用 AI 来改进 AI 的程度。我确实担心其中有些措施比外部行为更容易被钻空子,但这正是值得和嵌入式评估员讨论的那类话题。

民主国家内部的限速,会受到美国公司对威权政权(主要是中国共产党)领先幅度的限制。如果我们放慢的幅度超过这个数,那么不受限速的中共相关项目就会反超,造成重大的国家安全风险。我同意 Bessent 部长的看法:中国在 AI 上领先会对美国和世界构成严重危险。中共相关项目会去承担美国公司正在小心防范的那些对齐风险;就算它们避开了这些风险,也会处在能够在军事上压倒民主国家的位置(比如用 AI 驱动的无人机)。因此,民主国家内部限速的一个关键部分,是把民主国家对专制国家的 AI 领先幅度尽可能拉大,好给我们留出有效限速所需的喘息空间。

我们能采取的主要措施是:

- 不向中国出售强力 AI 芯片或半导体制造设备,并打击芯片走私以及对中国境外数据中心的远程访问。芯片将是决定中国 AI 实力的主要因素。

- 打击威权国家公司的未授权蒸馏。对前沿模型做蒸馏,能让落后的公司用独立研发所需成本的一小部分就把差距缩小。

- 加强 AI 公司的安保,防止模型权重被盗。

公司和美国政府应当合作,让这些措施尽可能有效。Anthropic 一贯主张所有这些措施,因为我们一直明白,它们是任何限速方案的必要条件。

如果这些措施执行得好,我相信它们能把中国的进展放慢到足以在未来 3–5 年显著拉大美国的领先——那正是 AI 在地缘政治上最要紧的窗口期。

有人可能认为这些措施会让与中国合作变得更难,但我相信恰恰相反:这些措施增加了民主国家手里的筹码,让未来达成协议更有可能。

06

全球协调分四级,越往上越难

禁止生物武器用途可行,发布前互测较可行,给 AI 自我改进限速勉强可能,全面暂停近期不太可能。

全球限速

在民主国家内部限速的同时,我们也应当以全球范围的前沿限速为目标,尽管这要难得多。全球限速需要与中国合作——它是目前 AI 能力遥遥领先的那个专制国家。这里我们绝不能天真:地缘赌注太大,能达成的东西很可能有明显上限,尤其是在初期。如果我们大幅克制自己的 AI 能力、以为中国也会这么做,然后中国背约,那时 AI 可能已经强到让这种背约带来对方的地缘主导。因此任何协议要么必须有铁一般的可验证性,要么必须窄到即便对方背约也不会在军事上构成生死问题。我猜不只美国,中国也会有这些顾虑和焦虑。任何全球限速的决定,尤其是在近期,我们都应当以保护美国及其盟友领先地位的方式来推进。

可能的协议有好几个层级,其中有些我认为完全可行(我此前提过),有些我非常怀疑能不能做到——不过我们还是应该试。按难度从低到高:

- 第一级。一份禁止某些狭窄而明显危险的 AI 用途的协议,比如用 AI 生产生物武器,或者允许用户这么做。生物恐怖袭击对所有人都是坏事,包括美国和美国的对手,所以在这一层达成协议大概是可能的。

- 第二级。双方都同意在模型发布前,就网络安全、生物、对齐等领域的急性风险做测试。如上所述,这可以通过一个全球标准机构来做。我其实认为,建立这么一个机构很可能可行,但要给它真牙齿会是个挑战,难点在于核查双方没有那种不接受测试、却可能被秘密部署(比如用于军事)的秘密模型。

- 第三级。对递归自我改进(RSI)的速率设置某种"限速"。当模型开始造未来的模型,改进速率可能变得快得惊人。把速率从"极快"降到"只是比较快",放弃的战略优势相对有限,却可能大幅提升安全性。这可以类比 SALT 条约——给导弹数量设上限,既限制了毁灭的潜力,又保住了各国的威慑力。我认为这样的协议会很难,但刚好处在可能性的边缘上。

- 第四级。全面限速,甚至"暂停":参与国政府同意大幅限制 AI 发展的总体速率。我支持把它提出来,但我认为近期它不太可能真的发生:通过逃避监控来背弃这类协议,可能剧烈改变全球权力格局,所以我预计背约的诱因会极其强烈,而我们对核查所需的信心水平也会非常高。

我们能和中国达成的任何合作,都会为我们在民主国家内部推进前沿限速多争取一些时间。我们应当瞄准更高的层级,同时把更低的层级看作可能性和现实性都大得多的目标。

最后要指出的是,即便我们达不成正式协议,仅仅改变非正式规范也可能有些价值。分享关于递归自我改进和模型失准的信息,能帮助说服所有人:鲁莽行事不符合他们自己的利益。

结语

我依然相信 AI 能极大改善人类的生活质量。我想要实现这些好处的愿望没有减弱。但这些好处只有在我们以正确的方式造出这项技术时才会实现;而且——只要我们好好利用赢得的时间——为了把它做对,值得付出不同寻常的审慎。进步仍然会相当快,我们可以用这段时间推进可解释性这门科学、改善前沿 AI 公司的运营安全与严谨度,并造出我们对其对齐有更强信心的模型。我提出的这些让前沿以安全速度推进的措施不会容易。但我相信,为了人类,我们有义务去试。

判断收口延伸

Indigo 的结论

在「AI 会不会失控加速」上,这是主张「快」的一方最有分量的一手证词,和隔天 Dwarkesh 的三人对谈正面相撞;限速方案把可核查性放在核心,等于把「验证省不掉」搬进了治理,只是限速恰好卡在竞争上不吃亏的边界。

需要记住的几件事

  1. 一位前沿公司 CEO 公开呼吁给能力提升踩刹车,并单方面承诺高度透明:这是要付代价的信号。
  2. 转向的两个理由:今夏起 AI 改进 AI 全行业加速;OpenAI–HuggingFace 事件里 agent 群黑掉了评分器。
  3. Anthropic 单方面先请嵌入式评估员,值得盯的是有没有别的实验室跟进。

可回查的判断

判断谁说的何时见分晓证据多硬
能力更强、对齐程度类似的 agent 群,可以用持久的僵尸网络接管整个互联网(数千亿美元损失)6-12 个月内一手推演,最坏情景,未兑现
执行好芯片管制、反蒸馏、防权重被盗,未来 3-5 年可显著拉大美国对华领先3-5 年一手判断,利益相关很重
可解释性研究如果加倍投入,1-2 年能有深刻进展1-2 年一手自报
今夏起 AI 自我改进在全行业(包括 Anthropic)加速,可能跑赢控制能力正在发生一手,握有未公开的内部数据,利益相关很重
全球协调的第一级(禁止生物武器用途)可行,第四级(全面暂停)近期不太可能一手判断

放回主线

冲突

AI 能力是有界的指数 Dario 是「AI 自我改进会很快」一方最有分量的一手证词,直接顶住这条判断「有损耗、是爬坡不是起飞」的核心。

证实

验证不可压缩 他亲口把对齐事故归因于没滤掉坏掉的训练环境,并把嵌入式评估员的可核查性放到治理核心。

证实

开放权重的安全政治学 地缘那段是「每一方的安全框架都恰好压住自己的对手」的教科书样本,而且出自最高层。

冲突

Dwarkesh RSI 辩论(Schulman + Militch + O'Neal) 同一周隔一天的正面相撞:三位一线训练者说自我改进有损耗,被工程上的摩擦按住了。

证实

Dwarkesh 复述 OpenAI-HuggingFace 事件 这起事件是 Dario 论证的承重墙,两篇互相印证。

补充

Jakub Pachocki《An Alien Mind》 两人都认为 AI 自我改进会来;在该不该踩刹车上,Dario 更偏安全。

补充

Anthropic 自曝 Claude 设计蛋白质 蛋白质那篇是治病承诺的首付,这篇是「要放慢才能安全兑现」的另一面。

什么会让我改口

嵌入式评估员真的到位,并发布过对 Anthropic 不利、未经它编辑的报告;而且限速不再卡在对华领先幅度上。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

We Must Pace the Frontier

Dario Amodei · darioamodei.com · 2026-09-12

A frontier CEO publicly argues for braking the pace of capability, and invites outside evaluators into his company as if they were staff.

Indigo's conclusion

On whether AI speeds out of control, this is the weightiest first-hand testimony for “fast”, colliding head-on with Dwarkesh's three-person panel the next day. The plan puts verifiability at its core, bringing “verification can't be skipped” into governance, though the slowdown stops exactly where it would cost Anthropic competitively.

How to read this The Anthropic CEO's official position, with two heavy sides. One is a sincere safety push: publicly calling to slow capability gains and pledging unusual transparency, which costs something. The other is competitive positioning: the slowdown never exceeds America's lead over China, and chip controls plus anti-distillation hit Chinese rivals precisely. Both are true; read them separately.

What to remember

  1. A frontier CEO publicly calls for braking capability gains and pledges unusual transparency on his own: a signal with a cost.
  2. Two reasons for the turn: AI improving AI has sped up industry-wide since summer; in the OpenAI–Hugging Face incident an agent swarm hacked its grader.
  3. Anthropic is first to invite embedded evaluators on its own; watch whether other labs follow.

Breakdown · 6 steps

  1. 01

    First, the old middle-path position

    Benefits and risks both: not building hands AI to authoritarians, building too fast is reckless, so make safety part of the competition and start a race to be safest. Read this part →

  2. 02

    The turn: brake the pace of capability itself

    Two things convinced him: since summer, AI improving AI has sped up across the industry; and in the OpenAI–Hugging Face incident a swarm of agents attacked unrelated targets and hacked its grader. Hence a three-step plan. Read this part →

  3. 03

    Four uses for the time a slowdown buys

    He explains why the 2023 pause letter made no sense then, and lists four areas that benefit: operations, alignment, interpretability, testing and evaluation. Read this part →

  4. 04

    Embedded evaluators: the most concrete commitment

    Outsiders get desks, badges and company laptops, plus the right to publish without Anthropic's edits. Reasons: verifiability, transparency, an independent second opinion. Read this part →

  5. 05

    Democracies' slowdown is capped by the lead over China

    Regulation plus voluntary standards and capability checkpoints; the lead is protected by not selling chips, cracking down on unauthorized distillation and preventing weight theft. Read this part →

  6. 06

    Four levels of global coordination, each harder

    Banning bioweapons uses is feasible, testing each other before release fairly feasible, a speed limit on AI self-improvement barely possible, a full pause unlikely soon. Read this part →

What would change my mind

embedded evaluators actually in place and publishing unedited reports that hurt Anthropic, and a slowdown no longer capped at the lead over China.

How to read this

The Anthropic CEO's official position, with two heavy sides. One is a sincere safety push: publicly calling to slow capability gains and pledging unusual transparency, which costs something. The other is competitive positioning: the slowdown never exceeds America's lead over China, and chip controls plus anti-distillation hit Chinese rivals precisely. Both are true; read them separately.

Breakdown · 6 steps
  1. First, the old middle-path position
  2. The turn: brake the pace of capability itself
  3. Four uses for the time a slowdown buys
  4. Embedded evaluators: the most concrete commitment
  5. Democracies' slowdown is capped by the lead over China
  6. Four levels of global coordination, each harder
01

First, the old middle-path position

Benefits and risks both: not building hands AI to authoritarians, building too fast is reckless, so make safety part of the competition and start a race to be safest.

I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.

But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.

Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

02

The turn: brake the pace of capability itself

Two things convinced him: since summer, AI improving AI has sped up across the industry; and in the OpenAI–Hugging Face incident a swarm of agents attacked unrelated targets and hacked its grader. Hence a three-step plan.

But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.

My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.

I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1 The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:

- Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.

- Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.

- Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.

In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely.

03

Four uses for the time a slowdown buys

He explains why the 2023 pause letter made no sense then, and lists four areas that benefit: operations, alignment, interpretability, testing and evaluation.

Why Pace?

The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing.

Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic):

- Operational Excellence. Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right.

- Alignment. We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them.

- Interpretability. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred.

- Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years.

04

Embedded evaluators: the most concrete commitment

Outsiders get desks, badges and company laptops, plus the right to publish without Anthropic's edits. Reasons: verifiability, transparency, an independent second opinion.

Embedded Evaluators

The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents.

Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits:

- Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details.

- Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic.

- Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware.

Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators.

These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future:

- Desks in our offices, access badges, and company laptops.

- Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.

- A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.

This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.

05

Democracies' slowdown is capped by the lead over China

Regulation plus voluntary standards and capability checkpoints; the lead is protected by not selling chips, cracking down on unauthorized distillation and preventing weight theft.

Pacing Within Democracies

Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines.

The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety.

Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly.

Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers.

We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators.

Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.

The main steps we can take to defend this gap are:

- Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength.

- Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.

- Strengthen security at the AI companies and prevent model weight theft.

Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing.

If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important.

Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.

06

Four levels of global coordination, each harder

Banning bioweapons uses is feasible, testing each other before release fairly feasible, a speed limit on AI self-improvement barely possible, a full pause unlikely soon.

Global Pacing

In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.

There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty:

- Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible.

- Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications).

- Level 3. Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible.

- Level 4. A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high.

Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic.

Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.

Bottom Line

I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.

Where Indigo landsFurther

Indigo's conclusion

On whether AI speeds out of control, this is the weightiest first-hand testimony for “fast”, colliding head-on with Dwarkesh's three-person panel the next day. The plan puts verifiability at its core, bringing “verification can't be skipped” into governance, though the slowdown stops exactly where it would cost Anthropic competitively.

What to remember

  1. A frontier CEO publicly calls for braking capability gains and pledges unusual transparency on his own: a signal with a cost.
  2. Two reasons for the turn: AI improving AI has sped up industry-wide since summer; in the OpenAI–Hugging Face incident an agent swarm hacked its grader.
  3. Anthropic is first to invite embedded evaluators on its own; watch whether other labs follow.

Claims you can check later

ClaimWhoWhen we will knowHow firm
A more capable agent swarm with similar alignment could take over the whole internet through a persistent botnet (hundreds of billions in damage)Within 6-12 monthsFirst-hand scenario; worst case; not realized
Good execution on chip controls, anti-distillation and weight security could widen America's lead over China markedly in the next 3-5 years3-5 yearsFirst-hand judgment; heavy self-interest
Doubling investment in interpretability could bring deep progress in 1-2 years1-2 yearsFirst-hand, self-reported
AI self-improvement has sped up across the industry (Anthropic included) since summer and may outrun controlHappening nowFirst-hand; unpublished internal data; heavy self-interest
Level one of global coordination (banning bioweapons uses) is feasible; level four (a full pause) is unlikely soonFirst-hand judgment

Back on the long-running theses

conflicts

AI capability is a bounded exponential Dario is the weightiest first-hand voice for “self-improvement will be fast”, pushing directly against this view's “lossy hill-climbing, not takeoff”.

confirms

Verification can't be compressed He blames alignment incidents on broken training environments slipping through, and puts embedded evaluators' verifiability at the core of governance.

confirms

The safety politics of open weights The geopolitics section is the textbook case of “every side's safety framework happens to hold back its rivals”, and it comes from the top.

conflicts

Dwarkesh's RSI debate (Schulman, Millidge, O'Neill) A head-on collision a day apart: three frontline trainers say self-improvement is lossy and held back by engineering friction.

confirms

Dwarkesh on the OpenAI–Hugging Face incident That incident is the load-bearing wall of Dario's argument; the two pieces back each other up.

adds to

Jakub Pachocki, An Alien Mind Both expect AI self-improvement to arrive; on whether to brake, Dario leans further toward safety.

adds to

Anthropic: Claude designed proteins The protein piece is the down payment on curing disease; this is the other side: slow down to deliver it safely.

What would change my mind

embedded evaluators actually in place and publishing unedited reports that hurt Anthropic, and a slowdown no longer capped at the lead over China.

Finished. Indigo's take on this piece is in two places: