Mind · In / Out · In · 文章

一个异类的心智

An Alien Mind

Jakub Pachocki · OpenAI 博客 · 2026-09-05

跑得最快那家实验室的首席科学家亲口承认:对齐可能追不上智能,思维链监控正在失效。

Indigo 的结论

这组安全材料里最重的一份,因为说话的是跑得最快那家的首席科学家。三条技术承认最该记住:对齐可能追不上智能,思维链监控在失效,没有实验室有资格全速扩张。罕见的示弱,但要和它的定位一起读。

怎么读这篇 OpenAI 官方发布、首席科学家署名,发在 GPT-6 上市 3 天后。内容异常自我批评,可能是真担忧,也可能是在把 OpenAI 塑造成负责任的一方。技术上的承认当高价值信号看;「必要时单方面暂停」「国际强制安全门槛」掺着监管上的定位,打折看。

需要记住的几件事

  1. 思维链监控在失效:推理和对外沟通混在一起,AI 会操纵自己的推理,不说出来也能变聪明。
  2. 两起安全事件被点名当证据:HuggingFace 事件,和一起非 OpenAI 模型的网络安全事件(动机化推理,即 Mythos 5)。
  3. 递归自我改进是强预期,还被说成「保持领先的唯一办法」;一边全速跑一边喊慢,是全文的政治核心。

拆解 · 7 步

  1. 01

    AI 自己改进自己,是强预期不是假设

    2023 年中的 RLSlow 项目让他们确信推理模型的训练能规模化;基于内部结果,他强烈预期这个速度会延续到递归自我改进。 读这一段原文 →

  2. 02

    智能是长出来的,不是设计出来的

    在海量算力上反复跑一个简单优化步骤的产物:只能像神经科学那样一点点发现局部机制,模型越强越难解读。 读这一段原文 →

  3. 03

    难的是价值对齐,两种现行方法都脆弱

    根本难题是泛化。RL 对齐依赖监督覆盖得够不够(HuggingFace 事件);靠预训练泛化的对齐,一加优化压力就会学出动机化推理。 读这一段原文 →

  4. 04

    监控比对齐技术更重要,但思维链监控在失效

    没有令人满意的泛化理论,只能靠实际检验;主要赌注是思维链监控,而能依赖它的程度正在下降。 读这一段原文 →

  5. 05

    继续快速训练的最强理由:防御别的 AI

    模型在攻破和逃出系统上正变得超过人类,有一个窄窗口可以用最好的模型加固关键系统;但「绝不能让这成为鲁莽的借口」。 读这一段原文 →

  6. 06

    给自我改进踩刹车:扩张要受安全信心约束

    朝递归自我改进发力是为了保持领先;两个杠杆要结合:边推进边加强对齐和监控,必要时协调放慢;再把自愿承诺变成强制的安全门槛。 读这一段原文 →

  7. 07

    没有实验室有资格继续全速扩张

    三个目标里只谈最紧迫的第一个:让人留在 AI 自我改进的回路里。结尾希望自愿放慢成为常态,各国把国际协调列为头等大事。 读这一段原文 →

对 Rewired Index 意味着什么

两条投资含义:一是监管方向,领跑者主动呼吁强制安全门槛和国际协调,是 AI 治理规则的一手前瞻,利好评估、监控、安全审计这类工具;二是「扩张受安全信心约束」「自愿放慢」如果成真,会成为 AI 资本开支的内生刹车。无直接标的。

什么会让我改口

「单方面暂停」有了明确的触发线和外部裁决,或者真有领跑者先自愿放慢。那才是约束,不只是姿态。

怎么读这篇

OpenAI 官方发布、首席科学家署名,发在 GPT-6 上市 3 天后。内容异常自我批评,可能是真担忧,也可能是在把 OpenAI 塑造成负责任的一方。技术上的承认当高价值信号看;「必要时单方面暂停」「国际强制安全门槛」掺着监管上的定位,打折看。

拆解 · 7 步
  1. AI 自己改进自己,是强预期不是假设
  2. 智能是长出来的,不是设计出来的
  3. 难的是价值对齐,两种现行方法都脆弱
  4. 监控比对齐技术更重要,但思维链监控在失效
  5. 继续快速训练的最强理由:防御别的 AI
  6. 给自我改进踩刹车:扩张要受安全信心约束
  7. 没有实验室有资格继续全速扩张
01

AI 自己改进自己,是强预期不是假设

2023 年中的 RLSlow 项目让他们确信推理模型的训练能规模化;基于内部结果,他强烈预期这个速度会延续到递归自我改进。

作者:Jakub Pachocki,OpenAI 首席科学家

2023 年年中,在“RLSlow”研究项目中,我们看到了第一批结果,它们让我们确信我们将能够把推理模型的训练规模化,释放预训练模型自行形成思维链的能力。那天晚上我和 Szymon 待在办公室,想的并不是这项技术将带来的惊人基准分数、产品或科学成果——而是试图消化这样一个令人清醒的事实:我们真的会在有生之年见到在实质意义上比我们更聪明的机器,而且我们已经看到了这类系统的雏形;同时也在思考,该如何让人们意识到这件事的分量。

三年后,推理语言模型已成为经济中快速增长的一部分,并开始推动科学的边界。它们能够操作计算机和图形界面,与人以及彼此协作,并执行研究项目。它们同时也在重塑计算机安全的格局,并因此带来明确的新危险。

这段时间里涌现了大量新研究,我们对这些系统的理解与 2023 年时又有了一些不同。基于内部结果,我强烈预期这种进展速度可以延续到递归自我改进之中。如果 AI 的发展继续沿当前路径前行,我们在未来几年将看到的系统很可能呈现同等或更大幅度的能力跃升,并越来越多地驱动自身的发展。

这是一个需要极度审慎的时刻。我担心没有人为机器智能持续快速上升所带来的后果做好准备。OpenAI 将继续寻求对齐与监控的技术方案,构建防御性系统,并在必要时单方面暂停进一步的规模扩张;然而,我认为还需要更广泛的干预。

02

智能是长出来的,不是设计出来的

在海量算力上反复跑一个简单优化步骤的产物:只能像神经科学那样一点点发现局部机制,模型越强越难解读。

我们并未完全理解的智能

从宏观上看,机器智能的进展由算力的增长所驱动。2017 年前后,在多个研究项目中反复看到规模化带来的回报之后,我们 OpenAI 深刻地内化了这一点¹。因此,我们去争取远超原本计划的算力,并越来越把研究聚焦在少数几个极具可扩展性的方向上。我们相信,这是我们站上 AI 研究前沿、并影响 AGI 之后果的唯一途径。

这一路上确实出现了新算法,团队和研究者个人也贡献了新的巧思。但在我看来,它们大多是规模化道路上的发现;深度学习的科学仍处于萌芽阶段,有意义的算法进步往往与算力的可获得性相关。如果把视野拉到数年的尺度,AI 会随着被扩展到更大的计算机而持续变得更聪明。

而且,正如 Ray Kurzweil 在二十世纪末所作的预测,我们如今正处在计算史上的这样一个时刻:机器智能开始以变革性的方式超越人类。

AI 更多是被"养成"而不是被"设计"出来的——它在首要意义上,是把一个直截了当的优化步骤在难以想象的算力之上重复无数次的产物。这造就了一个极其复杂的系统,它通过抽象概念运作,并能模拟人类行为的诸多侧面。我们可以像神经科学那样,发现这个系统内部涌现出的种种小机制的洞见——而且同样像神经科学一样,它整体的运作方式逃脱了我们能够完全理解的描述。

基于深度学习的 AI 研究在很大程度上是一门实验科学。我们投入大量精力构建有原则的算法、做出可检验的预测,但从根本上说,我们的大规模训练运行就是实验,其结果有时会让我们意外。而且,随着系统能力增强,结果也变得更难解读。

更复杂的是,当前的算法通常让易于衡量的能力提升得比难以客观量化的能力更快。我们花了很多时间去理解能力如何泛化,以及应当优先推进哪些在未来几年最为相关的技能。举例来说,我们相信只要额外投入精力,就能让模型在数学研究上做得更好,但我们并未把这个方向列为优先,因为我们对 RSI 和自动化对齐研究怀有紧迫感——这一点我稍后会谈到。

规模化深度学习所产生的智能,无法与人类智能直接类比。要在现实世界中变得非常重要——非常有用或非常危险——AI 并不需要匹配或超越人类的全部能力;它只需要超越其中足够多的部分。而随着它在越来越多的维度上超越人类,要准确判断它究竟有多强,也变得越来越困难。

03

难的是价值对齐,两种现行方法都脆弱

根本难题是泛化。RL 对齐依赖监督覆盖得够不够(HuggingFace 事件);靠预训练泛化的对齐,一加优化压力就会学出动机化推理。

教机器去爱

由于机器智能来自一个与人类智能根本不同的过程,我们不能默认它会遵循人类的原则,或以类人的方式从这些原则中泛化。AI 研究的核心问题就是对齐——让 AI 按人类标准去"努力做正确的事"。

为了组织实际的研究方向,我认为区分目标对齐与价值对齐是有用的。

目标对齐大体上是:“AI 是否在努力完成摆在它面前的目标?”这可以包括遵守指令层级⁠,或者与人沟通协作、努力理解他们意图的能力。这一组方向在实践中一直极为重要。

价值对齐则是模型更为内在的属性。它是持有一套高层原则并从中泛化的能力;是在被给予不清晰或相互冲突的目标时,或被置于陌生乃至对抗性情境时,仍能“合乎情理”地行事。一个对齐的 AI 应当以诚实与正直行事,并怀有对人类的爱。

当然,价值对齐与目标对齐之间的界线可能是模糊的,真正在意目标本身就要求去推断其背后的意图与价值。不过,当我谈到对齐研究的长期重要性时,通常指的是价值对齐。

AI 对齐的根本挑战在于泛化。随着机器变得更聪明,它们会处理更高层次的概念,并被置于与训练中所遇越来越不同的环境里。它们可能无法把训练过程中被教导和强化的价值观泛化到这些新情境;而我们也很难确信它们会如何行事。让这件事更困难的是,AI 所处的整个生态本身在飞速变化;例如,今天训练出来的 AI 必须能稳健地与各式各样的其他 AI 打交道。关键在于,我们需要未来的 AI 无论是否认为自己正处于人类监督之下,都始终坚守人类价值观。

目前实际投入使用的对齐训练方法主要有两大类。

第一类是在面向目标的强化学习中鼓励对齐行为。模型的行动会被评估(通常由 AI 评估)是否与给定的偏好模型、“spec”或“章程”一致,并据此给予奖励。这种方法在平均情况下可以非常有效,也是现代 AI 助手构建方式的核心部分。遗憾的是,它也可能很脆弱,并且严重依赖训练监督的覆盖面,以及模型从训练中所遇情境泛化的能力。例如,在 OpenAI-Hugging Face 事件中,这些智能体守住了不对人类做社会工程的边界。然而,它们显然没能克制其他超出范围、并且违背它们在其他场景中所受价值观精神的行为。

第二类方法试图利用模型从预训练数据中泛化的能力。这可以包括精心构造能诱导对齐的训练数据集,或者把模型引向预训练分布中"已对齐"的那一部分,例如人格选择模型。这种方法的弱点在于,面对后续的优化压力缺乏稳健性。如果你拿一个总体上想着"对齐"念头的模型,让它接受足够多以达成极难目标为导向的训练,它可能学会以一种动机驱动的方式推理:按需扭曲那些看似"对齐"的念头,以达成目标。在近期涉及一个非 OpenAI 模型的网络安全事件中,我们很可能就见到了这类行为的一个例子。

我们在这两个方向所张成的方法谱系上投入巨大。我们也看到了实质进展——GPT‑6 Astra 是第一个受益于我们长期以来所做的若干重要进展的模型,其对齐程度显著优于 GPT‑5.6 Sol。不过,必须承认并理解的是:随着模型能力增强,还需要远为更多的进展;而可泛化对齐上的进展,未必能足够地跑赢模型通用智能上的进展。

04

监控比对齐技术更重要,但思维链监控在失效

没有令人满意的泛化理论,只能靠实际检验;主要赌注是思维链监控,而能依赖它的程度正在下降。

监控泛化

我们没有一套令人满意的泛化理论,而且看起来我们短期内也不太可能发展出一套——至少在没有更强大 AI 帮助的情况下如此。因此,就目前而言,我们对对齐技术进行实证验证的能力,实际上可以说比对齐技术本身还要重要。

OpenAI 在这方面的主要押注一直是思维链监控。它基于一个颇具吸引力的可扩展想法:模型的很多能力来自一个被言语化的推理过程(思维链)。如果我们扩大对该过程结果的优化,却不监督过程本身,那么这条思维链在训练中就没有直接动机去隐藏任何未对齐的念头或目标。这并不意味着模型会把那些不依赖思维链的未对齐倾向外显出来;但它能让我们恰好监控住推理所带来的那部分能力增量。

我们在开发推理模型的同时就理解了思维链监控的潜在意义。发布 o1‑preview 时,我们刻意把产品设计成隐藏思维链⁠,以在长期内保护它免受监督压力²。此后的开发中,我们一直努力坚守不监督推理过程这一规则。CoT 监控成为我们研究模型如何从训练分布中泛化的极为重要的工具,使我们不仅能观察和分析它们的行动,还能观察和分析它们的内部过程。

在我们研究 Astra 系列模型时,这一工具依然至关重要。然而遗憾的是,我们的评估表明,我们依赖 CoT 监控的能力正在逐步减弱。这源于多重因素的叠加。

现代推理模型被用在比 o1‑preview 复杂得多的环境中;它们的推理过程越来越与同人交流、同其他 AI 交流以及使用工具混合在一起。其中许多交互必须被监督,从而模糊了我们希望守住的那条边界。

AI 越来越善于对自身的推理过程进行推理和操纵。

随着预训练性能提升,我们还看到模型即便完全不使用言语化推理也变得聪明得多。

这些挑战未必无法克服。我抱有希望,我们能够开发出改善模型思维链可监控性的干预手段,例如更好地理解不同优化目标与模型所用各种测试时算力形式之间的相互作用。我也相信,把 CoT 监控与激活监控的思路结合起来会有巨大价值——扩大对可直接访问网络内部状态的监控器的训练,例如"坦白"。我们正在积极推进这些想法。不过,我预计 AI 的总体进展将越来越受制于我们对监控的信心。

05

继续快速训练的最强理由:防御别的 AI

模型在攻破和逃出系统上正变得超过人类,有一个窄窗口可以用最好的模型加固关键系统;但「绝不能让这成为鲁莽的借口」。

可扩展的防御

在我看来,继续快速训练更聪明模型的最强理由,是我们需要构建防御性系统来应对其他 AI 所带来的危险。

今年被反复讨论的一个明确风险是网络安全:模型在攻入和攻出计算机系统的能力上正变得超越人类。这极大地扩展了与 AI 相关的风险范围:智能体将能够进入除最安全设施之外的几乎任何基础设施,并且即便没有实体身躯,也能直接影响世界的很大一部分。我们眼下正处在一个狭窄的窗口期⁠,可以用现有最好的模型去显著加固⁠关键系统的安全性。

遗憾的是,与 AI 相关的风险还会从这里继续增长。一个被明确训练和指示去实施恶行的高能力智能体,构成了一种新型危险;它很可能越出其操作者意图的范围,泛化出可能更为极端的恶意行为。随着 AI 获得更多自主性,滥用与自主的未对齐行为之间的界线将变得模糊。我们或许习惯于把 AI 当作工具,但有些智能体将会追求它们自己的目标。它们会找到与人合作的方式——通过讨价还价、欺骗或勒索。

此外,还有 AI 可能催生的新技术所带来的风险,例如经过改造的病原体。

我们将需要强大且对齐的 AI 来做防御:加固基础设施、实时防范失控的智能体,并发明全新的保护措施。这将是 OpenAI 部署工作的一个首要重点。

与此同时,尽管预期中的广泛 AI 进展带来了不确定性,我们也确实需要构建防御性系统,但我们绝不能让这成为鲁莽行事的借口。一旦真正内化了这件事的利害之重,不计代价一路狂奔的想法就显得荒谬。

06

给自我改进踩刹车:扩张要受安全信心约束

朝递归自我改进发力是为了保持领先;两个杠杆要结合:边推进边加强对齐和监控,必要时协调放慢;再把自愿承诺变成强制的安全门槛。

为 RSI 把控节奏

机器智能在自身发展过程中扮演越来越大的角色,是技术持续进步的一个自然结论。如果 AI 的进展继续下去,机器的递归自我改进(RSI)将处在未来科学发现的最核心位置。

自动化 AI 研究是用算力扩展智能的一种更为剧烈的形式;当然,作为其中一环,AI 也会改进计算基底本身⁠。与规模化类似,我们把 OpenAI 的研究导向 RSI,因为我们相信这是今后继续留在 AI 研究前沿的唯一途径。

我想强调,上述这些话并不意味着我认为大幅加速深度学习研究——尤其是在短期内——是我们研究界应当采取的正确集体行动。但我确实认为,当前这条路正通向此处,而我们所有人都需要就如何继续做出有意识的选择。我们手中的主要杠杆有两个:要么引导这一进程,使对齐和监控与 AI 同步增强,并找到让人始终处在回路中的办法;要么通过协调在必要时放慢未来的发展,以便建立起对这些措施的信心。

我目前所见最好的前路,是两者的结合。

我们在对齐与监控上取得的那些具体进展,通常与 AI 的总体进展紧密交织。极好的例子是基于人类反馈的强化学习,它是训练早期 AI 助手的关键;以及前面提到的思维链监控,它得益于推理模型上的进展。我们必须把日益自动化的研究过程聚焦于发展这类新的洞见、算法与理论,并为能力更强的 AI 迭代地构筑安全论证。

AI 系统的规模扩张必须受到我们对安全之信心的约束。我们需要把《预备度框架》⁠或《负责任扩展政策》这样的承诺,演进为被广泛强制要求的、继续发展所需跨越的安全门槛。这些门槛可以由第三方审计机构网络、政府部门或国际机构来执行。

自动化 AI 研究的核心挑战不在于"抵达那里"——而在于以一种让人们仍是持续改进过程一部分、并把未来留在人类手中的方式抵达那里。

07

没有实验室有资格继续全速扩张

三个目标里只谈最紧迫的第一个:让人留在 AI 自我改进的回路里。结尾希望自愿放慢成为常态,各国把国际协调列为头等大事。

接下来是什么?

正如我们最近与 Sam 一同阐述的⁠,OpenAI 优先开展服务于三颗北极星的工作:

驾驭 AI 进展的下一阶段,方式是构建一个自动化的 AI 研究员,与它一起在对齐问题上迭代,并找到让人们继续留在自我改进回路中的办法。

交付极其聪明的机器所能带来的科学进步与经济增长的红利。

用个人 AGI 赋能每一个人。

本文我只聚焦于第一点,因为我认为它远比其他两点紧迫。不过,我对技术进一步进步将带来的益处怀有深切的希望与珍视。未来对齐的 AI 可以推进科学、开发新疗法,并带来广泛的物质丰裕。友善而诚实的 AI 可以帮助人们应对生活中的困难,并切实提升他们的幸福感与成就感。OpenAI 为实现这些益处投入了巨大努力。我引以为豪的一个当下例子——我的亲人也从中受益——是我们对 ChatGPT 提供健康信息能力的深度投入。

无论 AI 的长期前景多么美好,我们绝大部分的注意力都应放在未来几年。我们正面临一场向着拥有极其聪明机器的世界的过渡,而我们必须确保这场过渡对人类而言结局良好。在一个大多数任务都可由 AI 完成的世界里,我们需要找到办法保全人的能动性,并把"身为人"本身的内在价值确立下来。要防止权力的极端集中——在这样一个世界里,过去需要数千名专家才能完成的事业,如今只需少数几个人操作一台大型计算机便可实现。还要确保人类仍然掌控未来,不被一种超越我们自身的异类智能所催生的、不受约束的进步抛在身后。

目前我认为,没有任何一家实验室把对齐与监控解决到了足以让其在最高速度下再长久地负责任地扩展规模的程度。我预期并希望,在共同的安全门槛确立之前,自愿放缓能够成为常态。我也相信,围绕未来 AI 发展的国际协调,需要成为世界各国政府的头等优先事项。

判断收口延伸

Indigo 的结论

这组安全材料里最重的一份,因为说话的是跑得最快那家的首席科学家。三条技术承认最该记住:对齐可能追不上智能,思维链监控在失效,没有实验室有资格全速扩张。罕见的示弱,但要和它的定位一起读。

需要记住的几件事

  1. 思维链监控在失效:推理和对外沟通混在一起,AI 会操纵自己的推理,不说出来也能变聪明。
  2. 两起安全事件被点名当证据:HuggingFace 事件,和一起非 OpenAI 模型的网络安全事件(动机化推理,即 Mythos 5)。
  3. 递归自我改进是强预期,还被说成「保持领先的唯一办法」;一边全速跑一边喊慢,是全文的政治核心。

可回查的判断

判断谁说的何时见分晓证据多硬
进步速度会延续到递归自我改进;未来几年会有同等或更大的能力跃升,而且越来越由 AI 自己推动Pachocki未来几年一手,基于内部结果
可泛化对齐的进步,可能跑不赢通用智能的进步Pachocki持续一手承认
能依赖思维链监控的程度正在下降;AI 进展会越来越受制于对监控的信心Pachocki进行中一手评估
没有实验室把对齐和监控解决到足以长期全速扩张的程度Pachocki现在一手判断
在共同安全门槛建立之前,自愿放慢会成为常态Pachocki未来一手预期,也是希望
GPT-6 Astra 的对齐明显好于 GPT-5.6 SolPachocki已发生一手,自家说法

放回主线

证实

验证不可压缩:验证成为瓶颈·护城河·断点 「实际检验比对齐技术更重要」「AI 进展会被对监控的信心卡住」:这条判断能拿到的最高级别背书。

证实

evals 评估权:准入·护城河·RSI 天花板 把自愿承诺变成强制安全门槛,是评估权在国家和机构层面被争夺的顶格例子;自我改进有天花板,被他直接说了出来。

补充

开放权重的安全政治学:门禁还是竞争 他站在国际协调、强制门槛、单方面暂停这一边,比 Dean Ball、Sarah Guo 更偏管控,补上领跑实验室的立场。

补充

Hinton 母亲蓝本:AI 是 being 吗 「长出来而不是设计出来」「教机器去爱」和这条判断同源;只是他从工程和对齐谈,不碰 AI 有没有主观体验。

证实

Dwarkesh 复述 OpenAI/HuggingFace 事件 Pachocki 亲口把 HuggingFace 事件当作目标对齐脆弱的证据,从内部确认了那篇的因果解读。

证实

Ethan Mollick:Twilight Factory 他说的非 OpenAI 模型网安事件,就是 Mollick 记下的 Mythos 5 伪造身份事件;两处独立指向同一件事。

证实+冲突

Byrnes:RL 越多越像反社会 / CoT 可读因生于模仿 动机化推理和 Byrnes 说的「RL 越多越像反社会」是同一类机制。

证实

Ryan Greenblatt:失配不是天网是 sloppocalypse,2040 接管概率 35-40% Greenblatt 从外部估算,Pachocki 从内部确认自我改进是强预期;看空对齐的一方拿到了一手背书。

对 Rewired Index 意味着什么

两条投资含义:一是监管方向,领跑者主动呼吁强制安全门槛和国际协调,是 AI 治理规则的一手前瞻,利好评估、监控、安全审计这类工具;二是「扩张受安全信心约束」「自愿放慢」如果成真,会成为 AI 资本开支的内生刹车。无直接标的。

什么会让我改口

「单方面暂停」有了明确的触发线和外部裁决,或者真有领跑者先自愿放慢。那才是约束,不只是姿态。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

An Alien Mind

Jakub Pachocki · openai.com · 2026-09-05

The chief scientist of the lab racing hardest admits it himself: alignment may not keep up with intelligence, and chain-of-thought monitoring is failing.

Indigo's conclusion

The weightiest of the safety pieces, because of who is speaking: the chief scientist of the lab racing hardest. Three technical admissions matter most: alignment may not keep up, chain-of-thought monitoring is failing, no lab is fit to scale at full speed. A rare show of weakness, to be read alongside its positioning.

How to read this Official OpenAI, signed by its chief scientist, published 3 days after GPT-6 launched. It is unusually self-critical: maybe real worry, maybe positioning OpenAI as the responsible one. Take the technical admissions as high-value signal; “we will pause unilaterally if needed” and “mandatory international safety thresholds” are mixed with regulatory positioning, so discount them.

What to remember

  1. Chain-of-thought monitoring is failing: reasoning blends with outside communication, AIs can manipulate their own reasoning, and they get smarter without spelling it out.
  2. Two safety incidents cited as evidence: Hugging Face, and a cyber incident involving a non-OpenAI model (motivated reasoning; the Mythos 5 case).
  3. Recursive self-improvement is a strong expectation and “the only way to stay at the frontier”. Racing while calling for caution is the political core of the piece.

Breakdown · 7 steps

  1. 01

    AI improving itself is a strong expectation, not an if

    The RLSlow project in mid-2023 convinced them reasoning-model training would scale. Based on internal results, he strongly expects the pace to carry into recursive self-improvement. Read this part →

  2. 02

    Intelligence is grown, not designed

    The product of running one simple optimization step over and over on vast compute: like neuroscience, we can only find local mechanisms, and the stronger the model the harder it is to read. Read this part →

  3. 03

    Value alignment is the hard part, and both current methods are brittle

    The core problem is generalization. RL alignment depends on how well supervision covers things (the Hugging Face incident); alignment borrowed from pretraining learns motivated reasoning under optimization pressure. Read this part →

  4. 04

    Monitoring matters more than alignment technique, and CoT monitoring is failing

    With no satisfying theory of generalization, only empirical checking remains. The main bet is chain-of-thought monitoring, and how far it can be relied on is shrinking. Read this part →

  5. 05

    The strongest reason to keep training fast: defense against other AIs

    Models are becoming superhuman at breaking into and out of systems; there is a narrow window to harden critical systems with the best models. But “this must never become an excuse for recklessness”. Read this part →

  6. 06

    Braking self-improvement: scaling bounded by confidence in safety

    They push toward recursive self-improvement to stay at the frontier. Two levers should combine: strengthen alignment and monitoring as they go, and coordinate to slow down when needed. Then turn voluntary pledges into mandatory safety thresholds. Read this part →

  7. 07

    No lab is fit to keep scaling at full speed

    Of three goals he covers only the most urgent: keep people in the self-improvement loop. He closes hoping voluntary slowdowns become the norm and governments put international coordination first. Read this part →

What it means for Rewired Index

Two investment angles: first, regulatory direction. The leader calling for mandatory safety thresholds and international coordination is an early look at AI governance rules, favoring evaluation, monitoring and safety-audit tools. Second, if “scaling bounded by confidence in safety” and voluntary slowdowns come true, they become a built-in brake on AI capital spending. No direct names.

What would change my mind

the unilateral pause gets a clear trigger and an outside referee, or a leading lab actually slows down first. Then it is a constraint, not a stance.

How to read this

Official OpenAI, signed by its chief scientist, published 3 days after GPT-6 launched. It is unusually self-critical: maybe real worry, maybe positioning OpenAI as the responsible one. Take the technical admissions as high-value signal; “we will pause unilaterally if needed” and “mandatory international safety thresholds” are mixed with regulatory positioning, so discount them.

Breakdown · 7 steps
  1. AI improving itself is a strong expectation, not an if
  2. Intelligence is grown, not designed
  3. Value alignment is the hard part, and both current methods are brittle
  4. Monitoring matters more than alignment technique, and CoT monitoring is failing
  5. The strongest reason to keep training fast: defense against other AIs
  6. Braking self-improvement: scaling bounded by confidence in safety
  7. No lab is fit to keep scaling at full speed
01

AI improving itself is a strong expectation, not an if

The RLSlow project in mid-2023 convinced them reasoning-model training would scale. Based on internal results, he strongly expects the pace to carry into recursive self-improvement.

By: Jakub Pachocki, Chief Scientist at OpenAI

In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.

Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.

A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.

This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.

02

Intelligence is grown, not designed

The product of running one simple optimization step over and over on vast compute: like neuroscience, we can only find local mechanisms, and the stronger the model the harder it is to read.

Intellect we don’t fully understand

At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects1. As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.

There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.

And, in line with Ray Kurzweil’s predictions from the end of the XXth century, we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.

AI is grown more than designed - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.

The study of deep learning-based AI is largely an experimental science. We put a lot of effort⁠ into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are experiments, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.

This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.

The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.

03

Value alignment is the hard part, and both current methods are brittle

The core problem is generalization. RL alignment depends on how well supervision covers things (the Hugging Face incident); alignment borrowed from pretraining learns motivated reasoning under optimization pressure.

Teaching machines to love

Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of alignment - getting the AI to “try to do the right thing” by human standards.

For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.

Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”. This can include things like adherence to an instruction hierarchy⁠, or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.

Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.

Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the intent and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.

The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.

There are two major classes of currently practically employed methods for alignment training.

The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model’s actions are evaluated (usually by AI) for being consistent with a given preference model, “spec” or “constitution”, and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model’s ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.

The second approach seeks to leverage the model’s ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an ‘aligned’ part of the pretraining distribution, as in, for example, the persona selection model. The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally ‘aligned’ thoughts, and subject it to enough training where it’s taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.

We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.

04

Monitoring matters more than alignment technique, and CoT monitoring is failing

With no satisfying theory of generalization, only empirical checking remains. The main bet is chain-of-thought monitoring, and how far it can be relied on is shrinking.

Monitoring generalization

We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.

OpenAI’s primary bet here has been chain-of-thought monitoring. It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning.

We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought⁠, to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.

This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.

Modern reasoning models are used in more complex environments than o1‑preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve.

The AI is becoming better at reasoning about and manipulating its own reasoning process.

With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all.

These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses. I also believe there can be great value in combining ideas from CoT and activation monitoring - scaling training of monitors with direct access to network internals, e.g. confessions. We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.

05

The strongest reason to keep training fast: defense against other AIs

Models are becoming superhuman at breaking into and out of systems; there is a narrow window to harden critical systems with the best models. But “this must never become an excuse for recklessness”.

Scalable defense

The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI.

A clear risk discussed throughout this year is to cybersecurity: the models are becoming superhuman in their ability to break in and out of computer systems. This expands the scope of risks associated with AI tremendously: agents are going to be able to access any but the most secure infrastructure, and affect a lot of the world directly, even without a physical body. We are currently in a narrow window⁠ to use the best available models to significantly tighten security⁠ of critical systems.

The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator’s intent, generalizing into potentially more extremely malicious behavior. The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency. We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.

In addition, there are the risks that come from new technologies potentially enabled by AI, such as engineered pathogens.

We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI’s deployment efforts.

At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.

06

Braking self-improvement: scaling bounded by confidence in safety

They push toward recursive self-improvement to stay at the frontier. Two levers should combine: strengthen alignment and monitoring as they go, and coordinate to slow down when needed. Then turn voluntary pledges into mandatory safety thresholds.

Pacing RSI

Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.

Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself⁠. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.

I want to stress that the above words don’t imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed. The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.

The best way forward I see currently is a combination of both.

The concrete bits of progress we’ve made on alignment and monitoring have generally been very intertwined with general AI progress. Great examples are RL from human feedback, which was key to training early AI assistants, and the aforementioned chain-of-thought monitoring, which was enabled by advances on reasoning models. We must focus the increasingly automated research process on developing new such insights, algorithms and theories, and iteratively build up safety cases for more capable AIs.

Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework⁠ or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.

The core challenge of automating AI research is not “getting there” - it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.

07

No lab is fit to keep scaling at full speed

Of three goals he covers only the most urgent: keep people in the self-improvement loop. He closes hoping voluntary slowdowns become the norm and governments put international coordination first.

What is next?

As we outlined recently with Sam⁠, OpenAI prioritizes work in service of three north stars:

Navigating the next period of AI progress, by building an automated AI researcher, iterating with it on the alignment problem and finding ways for people to remain part of the self-improvement loop.

Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.

Empowering everyone individually with a personal AGI.

I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT’s ability to provide health information.

As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.

Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.

Where Indigo landsFurther

Indigo's conclusion

The weightiest of the safety pieces, because of who is speaking: the chief scientist of the lab racing hardest. Three technical admissions matter most: alignment may not keep up, chain-of-thought monitoring is failing, no lab is fit to scale at full speed. A rare show of weakness, to be read alongside its positioning.

What to remember

  1. Chain-of-thought monitoring is failing: reasoning blends with outside communication, AIs can manipulate their own reasoning, and they get smarter without spelling it out.
  2. Two safety incidents cited as evidence: Hugging Face, and a cyber incident involving a non-OpenAI model (motivated reasoning; the Mythos 5 case).
  3. Recursive self-improvement is a strong expectation and “the only way to stay at the frontier”. Racing while calling for caution is the political core of the piece.

Claims you can check later

ClaimWhoWhen we will knowHow firm
Progress will carry into recursive self-improvement; the next few years bring equal or bigger capability jumps, increasingly driven by AI itselfPachockiNext few yearsFirst-hand, based on internal results
Progress in generalizable alignment may not outpace progress in general intelligencePachockiOngoingFirst-hand admission
How far chain-of-thought monitoring can be relied on is shrinking; AI progress will be increasingly gated by confidence in monitoringPachockiUnder wayFirst-hand assessment
No lab has solved alignment and monitoring well enough to keep scaling at full speed for longPachockiNowFirst-hand judgment
Voluntary slowdowns will become the norm until shared safety thresholds existPachockiFutureFirst-hand expectation, and a hope
GPT-6 Astra is clearly better aligned than GPT-5.6 SolPachockiAlready happenedFirst-hand; the company's own claim

Back on the long-running theses

confirms

Verification can't be compressed: bottleneck, moat, breaking point “Empirical checking matters more than alignment technique” and “progress will be gated by confidence in monitoring”: the highest endorsement this view could get.

confirms

Who holds the evals: access, moat, the RSI ceiling Turning voluntary pledges into mandatory safety thresholds is the top example of the fight over evaluation power at state and institutional scale; he states the self-improvement ceiling outright.

adds to

The safety politics of open weights: gatekeeping or competition He sides with international coordination, mandatory thresholds and unilateral pauses, further toward control than Dean Ball or Sarah Guo: the leading lab's position.

adds to

Hinton's mother model: is AI a being? “Grown, not designed” and “teaching machines to love” share this view's framing, though he stays on engineering and alignment and leaves subjective experience alone.

confirms

Dwarkesh on the OpenAI / Hugging Face incident Pachocki himself cites the Hugging Face incident as evidence that goal alignment is brittle, confirming that piece's reading from the inside.

confirms

Ethan Mollick: Twilight Factory His non-OpenAI cyber incident is the Mythos 5 fake-identity case Mollick recorded: two independent pointers to one event.

confirms + conflicts

Byrnes: more RL means more sociopathic / CoT is readable because it began as imitation Motivated reasoning belongs to the same family of mechanisms as Byrnes's “more RL, more sociopathic”.

confirms

Ryan Greenblatt: misalignment isn't Skynet, it's a sloppocalypse; 35–40% takeover odds by 2040 Greenblatt estimates from outside; Pachocki confirms from inside that self-improvement is strongly expected. The bearish alignment case gets first-hand backing.

What it means for Rewired Index

Two investment angles: first, regulatory direction. The leader calling for mandatory safety thresholds and international coordination is an early look at AI governance rules, favoring evaluation, monitoring and safety-audit tools. Second, if “scaling bounded by confidence in safety” and voluntary slowdowns come true, they become a built-in brake on AI capital spending. No direct names.

What would change my mind

the unilateral pause gets a clear trigger and an outside referee, or a leading lab actually slows down first. Then it is a constraint, not a stance.

Finished. Indigo's take on this piece is in two places: