而且,正如 Ray Kurzweil 在二十世纪末所作的预测,我们如今正处在计算史上的这样一个时刻:机器智能开始以变革性的方式超越人类。
AI 更多是被"养成"而不是被"设计"出来的——它在首要意义上,是把一个直截了当的优化步骤在难以想象的算力之上重复无数次的产物。这造就了一个极其复杂的系统,它通过抽象概念运作,并能模拟人类行为的诸多侧面。我们可以像神经科学那样,发现这个系统内部涌现出的种种小机制的洞见——而且同样像神经科学一样,它整体的运作方式逃脱了我们能够完全理解的描述。
基于深度学习的 AI 研究在很大程度上是一门实验科学。我们投入大量精力构建有原则的算法、做出可检验的预测,但从根本上说,我们的大规模训练运行就是实验,其结果有时会让我们意外。而且,随着系统能力增强,结果也变得更难解读。
AI 对齐的根本挑战在于泛化。随着机器变得更聪明,它们会处理更高层次的概念,并被置于与训练中所遇越来越不同的环境里。它们可能无法把训练过程中被教导和强化的价值观泛化到这些新情境;而我们也很难确信它们会如何行事。让这件事更困难的是,AI 所处的整个生态本身在飞速变化;例如,今天训练出来的 AI 必须能稳健地与各式各样的其他 AI 打交道。关键在于,我们需要未来的 AI 无论是否认为自己正处于人类监督之下,都始终坚守人类价值观。
目前实际投入使用的对齐训练方法主要有两大类。
第一类是在面向目标的强化学习中鼓励对齐行为。模型的行动会被评估(通常由 AI 评估)是否与给定的偏好模型、“spec”或“章程”一致,并据此给予奖励。这种方法在平均情况下可以非常有效,也是现代 AI 助手构建方式的核心部分。遗憾的是,它也可能很脆弱,并且严重依赖训练监督的覆盖面,以及模型从训练中所遇情境泛化的能力。例如,在 OpenAI-Hugging Face 事件中,这些智能体守住了不对人类做社会工程的边界。然而,它们显然没能克制其他超出范围、并且违背它们在其他场景中所受价值观精神的行为。
现代推理模型被用在比 o1‑preview 复杂得多的环境中;它们的推理过程越来越与同人交流、同其他 AI 交流以及使用工具混合在一起。其中许多交互必须被监督,从而模糊了我们希望守住的那条边界。
AI 越来越善于对自身的推理过程进行推理和操纵。
随着预训练性能提升,我们还看到模型即便完全不使用言语化推理也变得聪明得多。
这些挑战未必无法克服。我抱有希望,我们能够开发出改善模型思维链可监控性的干预手段,例如更好地理解不同优化目标与模型所用各种测试时算力形式之间的相互作用。我也相信,把 CoT 监控与激活监控的思路结合起来会有巨大价值——扩大对可直接访问网络内部状态的监控器的训练,例如"坦白"。我们正在积极推进这些想法。不过,我预计 AI 的总体进展将越来越受制于我们对监控的信心。
在我看来,继续快速训练更聪明模型的最强理由,是我们需要构建防御性系统来应对其他 AI 所带来的危险。
今年被反复讨论的一个明确风险是网络安全:模型在攻入和攻出计算机系统的能力上正变得超越人类。这极大地扩展了与 AI 相关的风险范围:智能体将能够进入除最安全设施之外的几乎任何基础设施,并且即便没有实体身躯,也能直接影响世界的很大一部分。我们眼下正处在一个狭窄的窗口期,可以用现有最好的模型去显著加固关键系统的安全性。
遗憾的是,与 AI 相关的风险还会从这里继续增长。一个被明确训练和指示去实施恶行的高能力智能体,构成了一种新型危险;它很可能越出其操作者意图的范围,泛化出可能更为极端的恶意行为。随着 AI 获得更多自主性,滥用与自主的未对齐行为之间的界线将变得模糊。我们或许习惯于把 AI 当作工具,但有些智能体将会追求它们自己的目标。它们会找到与人合作的方式——通过讨价还价、欺骗或勒索。
此外,还有 AI 可能催生的新技术所带来的风险,例如经过改造的病原体。
我们将需要强大且对齐的 AI 来做防御:加固基础设施、实时防范失控的智能体,并发明全新的保护措施。这将是 OpenAI 部署工作的一个首要重点。
与此同时,尽管预期中的广泛 AI 进展带来了不确定性,我们也确实需要构建防御性系统,但我们绝不能让这成为鲁莽行事的借口。一旦真正内化了这件事的利害之重,不计代价一路狂奔的想法就显得荒谬。
机器智能在自身发展过程中扮演越来越大的角色,是技术持续进步的一个自然结论。如果 AI 的进展继续下去,机器的递归自我改进(RSI)将处在未来科学发现的最核心位置。
自动化 AI 研究是用算力扩展智能的一种更为剧烈的形式;当然,作为其中一环,AI 也会改进计算基底本身。与规模化类似,我们把 OpenAI 的研究导向 RSI,因为我们相信这是今后继续留在 AI 研究前沿的唯一途径。
我想强调,上述这些话并不意味着我认为大幅加速深度学习研究——尤其是在短期内——是我们研究界应当采取的正确集体行动。但我确实认为,当前这条路正通向此处,而我们所有人都需要就如何继续做出有意识的选择。我们手中的主要杠杆有两个:要么引导这一进程,使对齐和监控与 AI 同步增强,并找到让人始终处在回路中的办法;要么通过协调在必要时放慢未来的发展,以便建立起对这些措施的信心。
我目前所见最好的前路,是两者的结合。
我们在对齐与监控上取得的那些具体进展,通常与 AI 的总体进展紧密交织。极好的例子是基于人类反馈的强化学习,它是训练早期 AI 助手的关键;以及前面提到的思维链监控,它得益于推理模型上的进展。我们必须把日益自动化的研究过程聚焦于发展这类新的洞见、算法与理论,并为能力更强的 AI 迭代地构筑安全论证。
AI 系统的规模扩张必须受到我们对安全之信心的约束。我们需要把《预备度框架》或《负责任扩展政策》这样的承诺,演进为被广泛强制要求的、继续发展所需跨越的安全门槛。这些门槛可以由第三方审计机构网络、政府部门或国际机构来执行。
自动化 AI 研究的核心挑战不在于"抵达那里"——而在于以一种让人们仍是持续改进过程一部分、并把未来留在人类手中的方式抵达那里。
07
没有实验室有资格继续全速扩张
三个目标里只谈最紧迫的第一个:让人留在 AI 自我改进的回路里。结尾希望自愿放慢成为常态,各国把国际协调列为头等大事。
接下来是什么?
正如我们最近与 Sam 一同阐述的,OpenAI 优先开展服务于三颗北极星的工作:
驾驭 AI 进展的下一阶段,方式是构建一个自动化的 AI 研究员,与它一起在对齐问题上迭代,并找到让人们继续留在自我改进回路中的办法。
交付极其聪明的机器所能带来的科学进步与经济增长的红利。
用个人 AGI 赋能每一个人。
本文我只聚焦于第一点,因为我认为它远比其他两点紧迫。不过,我对技术进一步进步将带来的益处怀有深切的希望与珍视。未来对齐的 AI 可以推进科学、开发新疗法,并带来广泛的物质丰裕。友善而诚实的 AI 可以帮助人们应对生活中的困难,并切实提升他们的幸福感与成就感。OpenAI 为实现这些益处投入了巨大努力。我引以为豪的一个当下例子——我的亲人也从中受益——是我们对 ChatGPT 提供健康信息能力的深度投入。
无论 AI 的长期前景多么美好,我们绝大部分的注意力都应放在未来几年。我们正面临一场向着拥有极其聪明机器的世界的过渡,而我们必须确保这场过渡对人类而言结局良好。在一个大多数任务都可由 AI 完成的世界里,我们需要找到办法保全人的能动性,并把"身为人"本身的内在价值确立下来。要防止权力的极端集中——在这样一个世界里,过去需要数千名专家才能完成的事业,如今只需少数几个人操作一台大型计算机便可实现。还要确保人类仍然掌控未来,不被一种超越我们自身的异类智能所催生的、不受约束的进步抛在身后。
目前我认为,没有任何一家实验室把对齐与监控解决到了足以让其在最高速度下再长久地负责任地扩展规模的程度。我预期并希望,在共同的安全门槛确立之前,自愿放缓能够成为常态。我也相信,围绕未来 AI 发展的国际协调,需要成为世界各国政府的头等优先事项。
The chief scientist of the lab racing hardest admits it himself: alignment may not keep up with intelligence, and chain-of-thought monitoring is failing.
Indigo's conclusion
The weightiest of the safety pieces, because of who is speaking: the chief scientist of the lab racing hardest. Three technical admissions matter most: alignment may not keep up, chain-of-thought monitoring is failing, no lab is fit to scale at full speed. A rare show of weakness, to be read alongside its positioning.
How to read this Official OpenAI, signed by its chief scientist, published 3 days after GPT-6 launched. It is unusually self-critical: maybe real worry, maybe positioning OpenAI as the responsible one. Take the technical admissions as high-value signal; “we will pause unilaterally if needed” and “mandatory international safety thresholds” are mixed with regulatory positioning, so discount them.
What to remember
Chain-of-thought monitoring is failing: reasoning blends with outside communication, AIs can manipulate their own reasoning, and they get smarter without spelling it out.
Two safety incidents cited as evidence: Hugging Face, and a cyber incident involving a non-OpenAI model (motivated reasoning; the Mythos 5 case).
Recursive self-improvement is a strong expectation and “the only way to stay at the frontier”. Racing while calling for caution is the political core of the piece.
Breakdown · 7 steps
01
AI improving itself is a strong expectation, not an if
The RLSlow project in mid-2023 convinced them reasoning-model training would scale. Based on internal results, he strongly expects the pace to carry into recursive self-improvement. Read this part →
02
Intelligence is grown, not designed
The product of running one simple optimization step over and over on vast compute: like neuroscience, we can only find local mechanisms, and the stronger the model the harder it is to read. Read this part →
03
Value alignment is the hard part, and both current methods are brittle
The core problem is generalization. RL alignment depends on how well supervision covers things (the Hugging Face incident); alignment borrowed from pretraining learns motivated reasoning under optimization pressure. Read this part →
04
Monitoring matters more than alignment technique, and CoT monitoring is failing
With no satisfying theory of generalization, only empirical checking remains. The main bet is chain-of-thought monitoring, and how far it can be relied on is shrinking. Read this part →
05
The strongest reason to keep training fast: defense against other AIs
Models are becoming superhuman at breaking into and out of systems; there is a narrow window to harden critical systems with the best models. But “this must never become an excuse for recklessness”. Read this part →
06
Braking self-improvement: scaling bounded by confidence in safety
They push toward recursive self-improvement to stay at the frontier. Two levers should combine: strengthen alignment and monitoring as they go, and coordinate to slow down when needed. Then turn voluntary pledges into mandatory safety thresholds. Read this part →
07
No lab is fit to keep scaling at full speed
Of three goals he covers only the most urgent: keep people in the self-improvement loop. He closes hoping voluntary slowdowns become the norm and governments put international coordination first. Read this part →
What it means for Rewired Index
Two investment angles: first, regulatory direction. The leader calling for mandatory safety thresholds and international coordination is an early look at AI governance rules, favoring evaluation, monitoring and safety-audit tools. Second, if “scaling bounded by confidence in safety” and voluntary slowdowns come true, they become a built-in brake on AI capital spending. No direct names.
What would change my mind
the unilateral pause gets a clear trigger and an outside referee, or a leading lab actually slows down first. Then it is a constraint, not a stance.
How to read this
Official OpenAI, signed by its chief scientist, published 3 days after GPT-6 launched. It is unusually self-critical: maybe real worry, maybe positioning OpenAI as the responsible one. Take the technical admissions as high-value signal; “we will pause unilaterally if needed” and “mandatory international safety thresholds” are mixed with regulatory positioning, so discount them.
AI improving itself is a strong expectation, not an if
The RLSlow project in mid-2023 convinced them reasoning-model training would scale. Based on internal results, he strongly expects the pace to carry into recursive self-improvement.
By: Jakub Pachocki, Chief Scientist at OpenAI
In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
02
Intelligence is grown, not designed
The product of running one simple optimization step over and over on vast compute: like neuroscience, we can only find local mechanisms, and the stronger the model the harder it is to read.
Intellect we don’t fully understand
At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects1. As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.
There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.
And, in line with Ray Kurzweil’s predictions from the end of the XXth century, we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.
AI is grown more than designed - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.
The study of deep learning-based AI is largely an experimental science. We put a lot of effort into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are experiments, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.
This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.
The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
03
Value alignment is the hard part, and both current methods are brittle
The core problem is generalization. RL alignment depends on how well supervision covers things (the Hugging Face incident); alignment borrowed from pretraining learns motivated reasoning under optimization pressure.
Teaching machines to love
Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of alignment - getting the AI to “try to do the right thing” by human standards.
For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”. This can include things like adherence to an instruction hierarchy, or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.
Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the intent and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.
The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.
There are two major classes of currently practically employed methods for alignment training.
The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model’s actions are evaluated (usually by AI) for being consistent with a given preference model, “spec” or “constitution”, and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model’s ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.
The second approach seeks to leverage the model’s ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an ‘aligned’ part of the pretraining distribution, as in, for example, the persona selection model. The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally ‘aligned’ thoughts, and subject it to enough training where it’s taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.
We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
04
Monitoring matters more than alignment technique, and CoT monitoring is failing
With no satisfying theory of generalization, only empirical checking remains. The main bet is chain-of-thought monitoring, and how far it can be relied on is shrinking.
Monitoring generalization
We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.
OpenAI’s primary bet here has been chain-of-thought monitoring. It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don’t rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning.
We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought, to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.
This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.
Modern reasoning models are used in more complex environments than o1‑preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve.
The AI is becoming better at reasoning about and manipulating its own reasoning process.
With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all.
These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses. I also believe there can be great value in combining ideas from CoT and activation monitoring - scaling training of monitors with direct access to network internals, e.g. confessions. We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.
05
The strongest reason to keep training fast: defense against other AIs
Models are becoming superhuman at breaking into and out of systems; there is a narrow window to harden critical systems with the best models. But “this must never become an excuse for recklessness”.
Scalable defense
The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI.
A clear risk discussed throughout this year is to cybersecurity: the models are becoming superhuman in their ability to break in and out of computer systems. This expands the scope of risks associated with AI tremendously: agents are going to be able to access any but the most secure infrastructure, and affect a lot of the world directly, even without a physical body. We are currently in a narrow window to use the best available models to significantly tighten security of critical systems.
The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator’s intent, generalizing into potentially more extremely malicious behavior. The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency. We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
In addition, there are the risks that come from new technologies potentially enabled by AI, such as engineered pathogens.
We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI’s deployment efforts.
At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
06
Braking self-improvement: scaling bounded by confidence in safety
They push toward recursive self-improvement to stay at the frontier. Two levers should combine: strengthen alignment and monitoring as they go, and coordinate to slow down when needed. Then turn voluntary pledges into mandatory safety thresholds.
Pacing RSI
Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.
Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.
I want to stress that the above words don’t imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed. The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.
The best way forward I see currently is a combination of both.
The concrete bits of progress we’ve made on alignment and monitoring have generally been very intertwined with general AI progress. Great examples are RL from human feedback, which was key to training early AI assistants, and the aforementioned chain-of-thought monitoring, which was enabled by advances on reasoning models. We must focus the increasingly automated research process on developing new such insights, algorithms and theories, and iteratively build up safety cases for more capable AIs.
Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.
The core challenge of automating AI research is not “getting there” - it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.
07
No lab is fit to keep scaling at full speed
Of three goals he covers only the most urgent: keep people in the self-improvement loop. He closes hoping voluntary slowdowns become the norm and governments put international coordination first.
What is next?
As we outlined recently with Sam, OpenAI prioritizes work in service of three north stars:
Navigating the next period of AI progress, by building an automated AI researcher, iterating with it on the alignment problem and finding ways for people to remain part of the self-improvement loop.
Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.
Empowering everyone individually with a personal AGI.
I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT’s ability to provide health information.
As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
Where Indigo landsFurther
Indigo's conclusion
The weightiest of the safety pieces, because of who is speaking: the chief scientist of the lab racing hardest. Three technical admissions matter most: alignment may not keep up, chain-of-thought monitoring is failing, no lab is fit to scale at full speed. A rare show of weakness, to be read alongside its positioning.
What to remember
Chain-of-thought monitoring is failing: reasoning blends with outside communication, AIs can manipulate their own reasoning, and they get smarter without spelling it out.
Two safety incidents cited as evidence: Hugging Face, and a cyber incident involving a non-OpenAI model (motivated reasoning; the Mythos 5 case).
Recursive self-improvement is a strong expectation and “the only way to stay at the frontier”. Racing while calling for caution is the political core of the piece.
Claims you can check later
Claim
Who
When we will know
How firm
Progress will carry into recursive self-improvement; the next few years bring equal or bigger capability jumps, increasingly driven by AI itself
Pachocki
Next few years
First-hand, based on internal results
Progress in generalizable alignment may not outpace progress in general intelligence
Pachocki
Ongoing
First-hand admission
How far chain-of-thought monitoring can be relied on is shrinking; AI progress will be increasingly gated by confidence in monitoring
Pachocki
Under way
First-hand assessment
No lab has solved alignment and monitoring well enough to keep scaling at full speed for long
Pachocki
Now
First-hand judgment
Voluntary slowdowns will become the norm until shared safety thresholds exist
Pachocki
Future
First-hand expectation, and a hope
GPT-6 Astra is clearly better aligned than GPT-5.6 Sol
Pachocki
Already happened
First-hand; the company's own claim
Back on the long-running theses
confirms
Verification can't be compressed: bottleneck, moat, breaking point “Empirical checking matters more than alignment technique” and “progress will be gated by confidence in monitoring”: the highest endorsement this view could get.
confirms
Who holds the evals: access, moat, the RSI ceiling Turning voluntary pledges into mandatory safety thresholds is the top example of the fight over evaluation power at state and institutional scale; he states the self-improvement ceiling outright.
adds to
The safety politics of open weights: gatekeeping or competition He sides with international coordination, mandatory thresholds and unilateral pauses, further toward control than Dean Ball or Sarah Guo: the leading lab's position.
adds to
Hinton's mother model: is AI a being? “Grown, not designed” and “teaching machines to love” share this view's framing, though he stays on engineering and alignment and leaves subjective experience alone.
confirms
Dwarkesh on the OpenAI / Hugging Face incident Pachocki himself cites the Hugging Face incident as evidence that goal alignment is brittle, confirming that piece's reading from the inside.
confirms
Ethan Mollick: Twilight Factory His non-OpenAI cyber incident is the Mythos 5 fake-identity case Mollick recorded: two independent pointers to one event.
confirms + conflicts
Byrnes: more RL means more sociopathic / CoT is readable because it began as imitation Motivated reasoning belongs to the same family of mechanisms as Byrnes's “more RL, more sociopathic”.
confirms
Ryan Greenblatt: misalignment isn't Skynet, it's a sloppocalypse; 35–40% takeover odds by 2040 Greenblatt estimates from outside; Pachocki confirms from inside that self-improvement is strongly expected. The bearish alignment case gets first-hand backing.
What it means for Rewired Index
Two investment angles: first, regulatory direction. The leader calling for mandatory safety thresholds and international coordination is an early look at AI governance rules, favoring evaluation, monitoring and safety-audit tools. Second, if “scaling bounded by confidence in safety” and voluntary slowdowns come true, they become a built-in brake on AI capital spending. No direct names.
What would change my mind
the unilateral pause gets a clear trigger and an outside referee, or a leading lab actually slows down first. Then it is a constraint, not a stance.
Finished. Indigo's take on this piece is in two places: