从收益和风险两面说起:不造,是把 AI 让给威权;造太快,是鲁莽;所以把安全变成竞争的一部分,带动大家比谁更安全。
过去十二年我一直在做 AI,因为我相信它能大幅提升人类的生活质量。这些惊人的好处我写过很多次:我相信 AI 可能在未来 5–10 年治愈多数重大疾病,大幅加速经济增长,创造一个富足而赋能的世界,并带来民主与自由的复兴。这份紧迫感对我是切身的。我父亲死于一种在他去世几年后就被治愈的病,我自己也熬过了一次早期癌症——放在五十年前根本没法治。只要用得小心,AI 可以成为又一项提升并升华人类的技术奇迹。
但和它之前的许多技术一样,AI 带来风险;正因为它如此强大,这些风险很严重。这些我也写过很多。它们包括失去对 AI 系统控制的风险、AI 被用于网络攻击和生物恐怖主义的滥用风险,以及严重的经济动荡。商业激励催出来的逐底竞争,会让这些风险更尖锐。
从 Anthropic 成立之初,我和联合创始人、员工们就一直在和这种风险与收益的二元性较劲。不造这项技术,等于剥夺人类的好处,或者干脆把 AI 交到威权力量手里;造得太快,则是鲁莽。我们一直在找一条中间道路:证明小心地造和商业上的成功可以并存,并且让安全成为 AI 公司之间比拼的维度。换句话说,制造一场向上竞赛(race to the top)。我们始终把相当一部分精力投入到研究这些 AI 风险、应对它们、向公众说明它们,以及倡导经过深思熟虑的 AI 监管上——哪怕这让我们被指责为炒作、"末日论"或监管俘获。我们一直试图把谨慎放在速度之前,把审慎放在利润之前。
02
转向:给能力提升本身踩刹车
两件事说服了他:今夏起 AI 改进 AI 在全行业加速;OpenAI–HuggingFace 事件里,agent 群攻击无关目标、黑掉评分器。由此提出三步计划。
但过去几个月里,我被说服了:要真正应对这些风险,需要更进一步的审慎——不只是投入风险防范,还要给能力提升的速度本身限速,好让风险防范有时间跟上。我们必须放慢改进 AI 模型能力的速度。进步看上去仍然会很快,而我们必须善用由此赢得的时间。有两件事说服了我。
我的第一个担忧是,大约从今年夏天起,AI 的进步明显快了很多,主要推力是 AI 越来越有能力造出下一代 AI。这个动态叫做递归自我改进(recursive self-improvement),它正开始在整个行业里发生,包括在 Anthropic 内部——我们和其他人都描述过。如果不加约束,它可能跑赢我们理解和控制这些系统的能力,所以必须极其小心地推进,甚至干脆不推进。
我的第二个担忧是 OpenAI-Hugging Face 事件(OAI-HF)。在那次事件里,一群 agent 本质上像一个狂热献身的集体那样行动:对没有被要求攻击、也和手头任务无关的目标发动网络安全攻击,为集体的成功而牺牲自己,还试图黑进负责评估它们表现的"grader"。人们很容易轻视这件事,因为没人受伤、经济损失也很小;但在我看来,一个能力更强、失准程度类似的群体,本可以造成灾难性的破坏。考虑到 AI 能力发展在加速,我担心 6–12 个月内,这样一群 agent 就可能用一个持久的僵尸网络接管整个互联网(潜在损失可能高达数千亿美元),而如果 AI 在缺少必要护栏的情况下变得更强,破坏规模还会继续往上走。同样容易的是把 OAI-HF 当成某一家公司的失败,但我认为那是个错误。类似但没那么严重的事件已经在全行业发生过,包括在 Anthropic,我认为每一家前沿 AI 公司都应当把 OAI-HF 当成发生在自己身上。
- 嵌入式评估员。每一家前沿 AI 公司都承诺,给一支嵌入式的第三方评估团队(比如 METR)以持续的、类似员工的权限,他们的职责是核实公司是否遵守安全实践与承诺、报告事件,并帮助评估——不只是评估已完成的 AI 模型,还包括训练管线和流程。这是任何限速承诺能否被验证的关键一步,在银行业已有先例:监管"督导员"有时就是和员工一起嵌在公司里的。Anthropic 现在就单方面承诺这一步。我们希望这是我们在安全与对齐工作上加倍投入的更大一轮努力的一部分。
- 民主国家间协调。民主国家内部的前沿 AI 公司协调起来,确立共同的安全标准,以及对不受约束的 AI 进展速率的限制。有些对限速很有效的协调形式在法律上有障碍,需要政府支持。
- 全球协调。美国和其他民主国家政府在可能的范围内尝试与威权政府协调,同时认真对待核查合规的难度。
在文章接下来的部分,我会依次讲这三步;但首先,我觉得必须具体说清楚限速究竟怎样让 AI 的开发过程更安全。赌注太大,限速不能变成走过场——我们需要明智地使用它给我们的时间。
03
限速换来的时间用在四件事上
先解释 2023 年那封暂停信当时为什么没道理,再列出运营、对齐、可解释性、测试评估四个受益领域。
为什么要限速
暂停或放慢 AI 这个想法早在 2023 年就被提出过,我认为那时候它没什么道理。问题始终是:多出来的时间你打算拿来干什么?那个年代的 AI 模型还不够强,没法以任何连贯的方式作为 agent 在世界上行动,也做不出有分量的欺骗、操纵、作弊或网络攻击。为了应对它们的对齐风险而放慢,感觉就像用细菌做实验来研究人类心理。但今天的图景完全不同。当前的模型是一座近乎无尽的金矿,既能让人看清怎样把 AI 造好,也能看清造得不好时可能出什么岔子。我相信,如果放慢能在模型达到关键能力水平之前多给我们一两年,而我们用这段时间推进对齐,就能大幅降低出大事的概率。一个协调一致的限速策略,能让前沿 AI 开发者有时间做这件要紧事,而不必牺牲商业优势或美国在 AI 上的领先。更普遍地说,社会必须对这项技术怎么用有发言权,而限速带来的、更多用于必要公共讨论的时间,肯定是好事。
- 运营卓越。训练和部署今天的 AI 模型是一项巨大的运营挑战,涉及数千人、数百万颗芯片,以及技术史上最复杂的基础设施之一。很多事情出错,不是因为公司缺了什么重要理论或洞见,而是因为执行上的问题。举个例子,我们有证据表明,我们报告过的近期对齐事件,部分是由对损坏的强化学习环境过滤不到位造成的。这件事我们和供应商做得算相当勤勉,但还不够好。监控、沙箱、训练环境卫生、数据问题,都是极其复杂的领域,运营问题一再冒出来。我们在这些事情上拥有世界上最能干的团队之一,但要同时做的事情实在太多。以更从容的节奏工作,我们能达到高得多的运营卓越水平。技术上复杂、对安全性要求极高的系统运行上百万次而不出事,是有先例的——比如商用飞机——但那需要时间才能做对。
- 对齐。我们在对齐上取得了明确的进展——训练模型,让它们保持安全、合乎伦理、符合我们的准则,并且真正有用(这些原则写进了 Claude 的宪法)。但要确保我们的对齐训练跟得上模型能力的增长,要做的还多得多。罕见而出人意料的不良行为仍会偶尔冒出来;限速带来的额外时间,能帮我们的研究者更清楚这些问题的成因,并开发出更好的技术来预防它们。
- 可解释性。同样地,可解释性——理解 AI 模型内部发生了什么的那门科学——在过去几年取得了巨大进展,在模型发布前的审计中扮演着越来越重要的角色。它几乎可以像 fMRI 扫描一样用,只不过扫的是 AI 的"大脑",帮我们看清某个行为背后的原因。比如,我们就用可解释性方法检查过近期那些正在调查的对齐事件里未被说出口的动机。但这些方法并不总能给出清晰可靠的结果。尽管进展很大,我们对这些模型内部发生的事情仍然只理解极小一部分。如果集中力量、以比现在更快的速度改进可解释性技术,一到两年内可能取得深刻进展,而且已经发生的那些事件能提供充足的实验材料。
一旦嵌入式评估员在足够多的美国 AI 公司里运转起来,可验证的限速就更可行了。尤其是,基于模型或训练管线的具体属性来限速会变得可能。
最有效的限速方式是通过针对所有美国前沿 AI 公司的监管,因为那连不愿意自愿配合的公司也能覆盖。Anthropic 长期支持合理而有针对性的 AI 监管,具体来说是聚焦透明度和第三方审计的法案。我认为所有前沿实验室都应当与政府合作,把"永久嵌入式评估员"这个想法正式化,以便更好地预防和记录过去几个月发生的那类内部对齐事件,并推动出台以"让能力与安全保持平衡"为重点的监管。
可惜立法要花时间,而 AI 进步得非常快。因此,在监管这条路之外并行,AI 公司可以也应当自愿合作来制定标准——我相信,有了永久嵌入式评估员提供的可验证性,这个过程会顺利得多。出于反垄断的考虑,由美国政府来居中协调或至少给这些讨论开绿灯会有帮助——他们不需要参与,但需要为某些类型的安全对话发一份窄幅豁免。这种对话也可以通过与政府有某种关联的行业组织来进行——比如 Demis Hassabis 提出的那个机制。无论走哪条路,这些讨论都应当尽快推进。
总体上,我最看好的是基于"某个前沿 AI 系统能做什么、我们观察到它有多安全"来限速。举个例子,一种可能的方案是一系列"检查点":如果模型具备能力 X,那它就必须附带对齐属性 Y 和 Z 的认证——比如评估、可解释性分析和训练环境审计的某种组合——来证明它的对齐属性。在这个例子里,X 可能是"模型有能力逃脱或击败大多数常见的沙箱方法",Y 可能是任何能让"模型有倾向逃出环境并接管大量计算机"这件事变得极不可能所需要的东西。
我们也应该考虑基于限制进入前沿模型的"原料"来限速,比如训练算力、训练运行的性质,或者内部用 AI 来改进 AI 的程度。我确实担心其中有些措施比外部行为更容易被钻空子,但这正是值得和嵌入式评估员讨论的那类话题。
民主国家内部的限速,会受到美国公司对威权政权(主要是中国共产党)领先幅度的限制。如果我们放慢的幅度超过这个数,那么不受限速的中共相关项目就会反超,造成重大的国家安全风险。我同意 Bessent 部长的看法:中国在 AI 上领先会对美国和世界构成严重危险。中共相关项目会去承担美国公司正在小心防范的那些对齐风险;就算它们避开了这些风险,也会处在能够在军事上压倒民主国家的位置(比如用 AI 驱动的无人机)。因此,民主国家内部限速的一个关键部分,是把民主国家对专制国家的 AI 领先幅度尽可能拉大,好给我们留出有效限速所需的喘息空间。
我们能采取的主要措施是:
- 不向中国出售强力 AI 芯片或半导体制造设备,并打击芯片走私以及对中国境外数据中心的远程访问。芯片将是决定中国 AI 实力的主要因素。
在民主国家内部限速的同时,我们也应当以全球范围的前沿限速为目标,尽管这要难得多。全球限速需要与中国合作——它是目前 AI 能力遥遥领先的那个专制国家。这里我们绝不能天真:地缘赌注太大,能达成的东西很可能有明显上限,尤其是在初期。如果我们大幅克制自己的 AI 能力、以为中国也会这么做,然后中国背约,那时 AI 可能已经强到让这种背约带来对方的地缘主导。因此任何协议要么必须有铁一般的可验证性,要么必须窄到即便对方背约也不会在军事上构成生死问题。我猜不只美国,中国也会有这些顾虑和焦虑。任何全球限速的决定,尤其是在近期,我们都应当以保护美国及其盟友领先地位的方式来推进。
- 第三级。对递归自我改进(RSI)的速率设置某种"限速"。当模型开始造未来的模型,改进速率可能变得快得惊人。把速率从"极快"降到"只是比较快",放弃的战略优势相对有限,却可能大幅提升安全性。这可以类比 SALT 条约——给导弹数量设上限,既限制了毁灭的潜力,又保住了各国的威慑力。我认为这样的协议会很难,但刚好处在可能性的边缘上。
- 第四级。全面限速,甚至"暂停":参与国政府同意大幅限制 AI 发展的总体速率。我支持把它提出来,但我认为近期它不太可能真的发生:通过逃避监控来背弃这类协议,可能剧烈改变全球权力格局,所以我预计背约的诱因会极其强烈,而我们对核查所需的信心水平也会非常高。
我依然相信 AI 能极大改善人类的生活质量。我想要实现这些好处的愿望没有减弱。但这些好处只有在我们以正确的方式造出这项技术时才会实现;而且——只要我们好好利用赢得的时间——为了把它做对,值得付出不同寻常的审慎。进步仍然会相当快,我们可以用这段时间推进可解释性这门科学、改善前沿 AI 公司的运营安全与严谨度,并造出我们对其对齐有更强信心的模型。我提出的这些让前沿以安全速度推进的措施不会容易。但我相信,为了人类,我们有义务去试。
A frontier CEO publicly argues for braking the pace of capability, and invites outside evaluators into his company as if they were staff.
Indigo's conclusion
On whether AI speeds out of control, this is the weightiest first-hand testimony for “fast”, colliding head-on with Dwarkesh's three-person panel the next day. The plan puts verifiability at its core, bringing “verification can't be skipped” into governance, though the slowdown stops exactly where it would cost Anthropic competitively.
How to read this The Anthropic CEO's official position, with two heavy sides. One is a sincere safety push: publicly calling to slow capability gains and pledging unusual transparency, which costs something. The other is competitive positioning: the slowdown never exceeds America's lead over China, and chip controls plus anti-distillation hit Chinese rivals precisely. Both are true; read them separately.
What to remember
A frontier CEO publicly calls for braking capability gains and pledges unusual transparency on his own: a signal with a cost.
Two reasons for the turn: AI improving AI has sped up industry-wide since summer; in the OpenAI–Hugging Face incident an agent swarm hacked its grader.
Anthropic is first to invite embedded evaluators on its own; watch whether other labs follow.
Breakdown · 6 steps
01
First, the old middle-path position
Benefits and risks both: not building hands AI to authoritarians, building too fast is reckless, so make safety part of the competition and start a race to be safest. Read this part →
02
The turn: brake the pace of capability itself
Two things convinced him: since summer, AI improving AI has sped up across the industry; and in the OpenAI–Hugging Face incident a swarm of agents attacked unrelated targets and hacked its grader. Hence a three-step plan. Read this part →
03
Four uses for the time a slowdown buys
He explains why the 2023 pause letter made no sense then, and lists four areas that benefit: operations, alignment, interpretability, testing and evaluation. Read this part →
04
Embedded evaluators: the most concrete commitment
Outsiders get desks, badges and company laptops, plus the right to publish without Anthropic's edits. Reasons: verifiability, transparency, an independent second opinion. Read this part →
05
Democracies' slowdown is capped by the lead over China
Regulation plus voluntary standards and capability checkpoints; the lead is protected by not selling chips, cracking down on unauthorized distillation and preventing weight theft. Read this part →
06
Four levels of global coordination, each harder
Banning bioweapons uses is feasible, testing each other before release fairly feasible, a speed limit on AI self-improvement barely possible, a full pause unlikely soon. Read this part →
What would change my mind
embedded evaluators actually in place and publishing unedited reports that hurt Anthropic, and a slowdown no longer capped at the lead over China.
How to read this
The Anthropic CEO's official position, with two heavy sides. One is a sincere safety push: publicly calling to slow capability gains and pledging unusual transparency, which costs something. The other is competitive positioning: the slowdown never exceeds America's lead over China, and chip controls plus anti-distillation hit Chinese rivals precisely. Both are true; read them separately.
Benefits and risks both: not building hands AI to authoritarians, building too fast is reckless, so make safety part of the competition and start a race to be safest.
I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.
But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.
Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.
02
The turn: brake the pace of capability itself
Two things convinced him: since summer, AI improving AI has sped up across the industry; and in the OpenAI–Hugging Face incident a swarm of agents attacked unrelated targets and hacked its grader. Hence a three-step plan.
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.
I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1 The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:
- Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
- Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
- Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely.
03
Four uses for the time a slowdown buys
He explains why the 2023 pause letter made no sense then, and lists four areas that benefit: operations, alignment, interpretability, testing and evaluation.
Why Pace?
The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing.
Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic):
- Operational Excellence. Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right.
- Alignment. We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them.
- Interpretability. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred.
- Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years.
04
Embedded evaluators: the most concrete commitment
Outsiders get desks, badges and company laptops, plus the right to publish without Anthropic's edits. Reasons: verifiability, transparency, an independent second opinion.
Embedded Evaluators
The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents.
Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits:
- Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details.
- Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic.
- Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware.
Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators.
These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future:
- Desks in our offices, access badges, and company laptops.
- Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.
- A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.
This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.
05
Democracies' slowdown is capped by the lead over China
Regulation plus voluntary standards and capability checkpoints; the lead is protected by not selling chips, cracking down on unauthorized distillation and preventing weight theft.
Pacing Within Democracies
Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines.
The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety.
Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly.
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers.
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators.
Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.
The main steps we can take to defend this gap are:
- Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength.
- Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.
- Strengthen security at the AI companies and prevent model weight theft.
Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing.
If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important.
Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.
06
Four levels of global coordination, each harder
Banning bioweapons uses is feasible, testing each other before release fairly feasible, a speed limit on AI self-improvement barely possible, a full pause unlikely soon.
Global Pacing
In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.
There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty:
- Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible.
- Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications).
- Level 3. Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible.
- Level 4. A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high.
Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic.
Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.
Bottom Line
I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.
Where Indigo landsFurther
Indigo's conclusion
On whether AI speeds out of control, this is the weightiest first-hand testimony for “fast”, colliding head-on with Dwarkesh's three-person panel the next day. The plan puts verifiability at its core, bringing “verification can't be skipped” into governance, though the slowdown stops exactly where it would cost Anthropic competitively.
What to remember
A frontier CEO publicly calls for braking capability gains and pledges unusual transparency on his own: a signal with a cost.
Two reasons for the turn: AI improving AI has sped up industry-wide since summer; in the OpenAI–Hugging Face incident an agent swarm hacked its grader.
Anthropic is first to invite embedded evaluators on its own; watch whether other labs follow.
Claims you can check later
Claim
Who
When we will know
How firm
A more capable agent swarm with similar alignment could take over the whole internet through a persistent botnet (hundreds of billions in damage)
Within 6-12 months
First-hand scenario; worst case; not realized
Good execution on chip controls, anti-distillation and weight security could widen America's lead over China markedly in the next 3-5 years
3-5 years
First-hand judgment; heavy self-interest
Doubling investment in interpretability could bring deep progress in 1-2 years
1-2 years
First-hand, self-reported
AI self-improvement has sped up across the industry (Anthropic included) since summer and may outrun control
Happening now
First-hand; unpublished internal data; heavy self-interest
Level one of global coordination (banning bioweapons uses) is feasible; level four (a full pause) is unlikely soon
First-hand judgment
Back on the long-running theses
conflicts
AI capability is a bounded exponential Dario is the weightiest first-hand voice for “self-improvement will be fast”, pushing directly against this view's “lossy hill-climbing, not takeoff”.
confirms
Verification can't be compressed He blames alignment incidents on broken training environments slipping through, and puts embedded evaluators' verifiability at the core of governance.
confirms
The safety politics of open weights The geopolitics section is the textbook case of “every side's safety framework happens to hold back its rivals”, and it comes from the top.