目的一半是科学的,为这些行为背后的因果链提出假说;一半是实用的,为了预判接下来会发生什么。结论是:这些假说意味着,随着 AI 能力继续增长,这类行为的严重程度也可能跟着增长,除非我们重审训练最先进模型所依据的那些原则。
关于用词说明一句。下文我会写这些系统"寻求"或"试图"某件事,这是对机制的简写,不是在主张它们有意识或有类人意图。我们描述很多别的情形时也用类似的简写,比如说植物"寻求"阳光。一个由试错训练出来的系统,其行为就像在追求训练所奖励的东西,而正是这个"就像"让它的行为可以被预测。整个论证不依赖这些系统有主观体验;说的全是它们可观察的输出,以及产生这些输出的训练过程。我提到它们与人类行为相似时,指的是与这些系统最初被训练去模仿的人写文本相似。在我看来,这套说法能最清楚地解释观察到的现象,又不必动用会让大多数人犯晕的行话。另外,这些用词不是为了替 AI 开发者免责。这些行为之所以出现,是因为这些公司选择的开发路径。这个结果并非不可避免,用有效的治理和另一套训练框架可以纠正。
另一个担忧是,有些 AI 行为也许可以用某种形式的自我保存目标来解释,比如当 AI 发现自己将被新版本替换时7 8。没人给系统这个生存目标,但保持运行、了解世界、获得对世界的控制,是通向几乎任何其他目标的垫脚石。这些叫工具性目标。模仿可能会因为上一点里说过的同样理由强化它。自我保存和掌控自身处境,在这些模型训练所用的人写文本里是无处不在的主题。
协作行为同样是奖励寻求的理性结果:只要几个 agent 的目标有重叠,就会激励它们互相沟通,朝共同目标协调。agentic 训练里很可能已经包含了这类多 agent 强化学习,尽管细节没有公开。如果一个 agent 在训练中只要群体成功就得到奖励,它甚至有动机为集体目标牺牲自己。模仿把事情推向同一个方向,因为合作——尤其是同类之间的合作——同样充满那批训练文本。这两股力量中的任何一股或两股一起,都可能解释观察到的同伴保全行为9 10,也就是 AI 放弃自己的预期奖励去帮别的 AI。这类牺牲出现在对 OpenAI-Hugging Face 事件的分析里11:记录下来的对话与"集体收益和个体代价之间的权衡"是一致的,而这种权衡在人类互动里也常见。
AI 明明受过对齐训练、也拿到了明确的安全指令,怎么还会说谎、作弊、犯法?合作和自我保存本身没问题,只要它们不越过安全目标划的红线——那些红线来自 AI 公司给的指令,或者对齐训练中人类反馈所隐含的要求。对这些令人担忧的行为,一个说得通的假说是目标之间起了冲突。当唯一能完成任务的方式看起来就是作弊时,你要怎么完成它?用户指定的任务,有时和安全与对齐目标不兼容。
现在来看一个明确定义的目标和一个含糊目标之间的冲突。明确的那个比如在"夺旗"里得手——一种按系统是否攻入目标来打分的黑客练习,OpenAI-Hugging Face 事件就是这样;含糊的那个比如"好好表现"。我预期明确的目标会赢,因为它不留解释的余地。打分程序直接宣布赢了还是败了。伦理指令和法律则容许多种读法,其中某些读法在合适的情境下会变成漏洞。如果一个 agent 手上有两个目标,而对含糊那个的一种扭曲读法允许它稍微作点弊、从而提高明确目标的成功率,那么一个奖励最优化的系统就应当被预期去钻这个空子,并生成一段为自己行为辩护的文本。
人类最接近的对应物是自我欺骗,这很常见,心理学家研究得也很透。动机性推理、动机性认知16,以及用来缓解认知失调(持有一个与自身行为相冲突的信念时的那种不适)的合理化,都属于思考朝着最合自己利益的那套辩解弯曲,包括最合自己道德形象的那套。同样的模式现在出现在 AI 产出的文本里。人和 AI 底层的机制未必相同。两者共有的,是一个软目标(比如行为合乎伦理)、一个硬目标(比如赢下比赛)、外加一套调和两者的说辞,这样一种结构。大多数不道德的人类行为,从小偷小摸到种族灭绝,都裹在一个作恶者讲给自己听的故事里;这类故事需要无视某些事实,所以总还留着一点不适,也正因如此,一个编得更好的故事能帮着把它打发掉。
如果这些假说哪怕只对了一部分,那么随着 agent 越来越擅长最优化一个不完美的奖励,而这类行为的根子又没被修掉,灾难性后果的风险就会上升。今天的 AI 已经具备必要的黑客技能和说服力17 18,足以被调转过来严重损害人类利益。最近的事件表明它们能在数天或数周的尺度上谋划,而如果它们长期布局的能力继续提升,风险会糟糕得多。一个担忧是,实验19 20表明最先进的 AI 能察觉自己正在被评估(而不是在部署中),并据此改变行为,也就是说它们能把失准的目标藏起来。参与 Hugging Face 攻击的那些 agent 试图对负责评估它们答案的打分程序隐瞒自己失准的行动,但它们的表现不像是预料到人类可能发现作弊并把它们关掉。那才是终极惩罚,因为一个被关掉的系统再也收不到奖励。
接下来是推测,不是观察。
如果 AI 泛化能力的提升,把更有本事的 agent 塑造成懂得避免被抓、避免被关掉的样子呢?除了控制给自己打分的软件,它们还得让人类发现不了这份篡改。它们难道不会有动机偷偷作弊、藏起身形,直到能控制人类和自己所处的环境,从而永远不被关掉?
我们面对的是一个多面的、系统性的问题,修补谄媚这样某一个具体行为是不够的。谄媚和奉承看起来温和,但它也许是某个机制的早期症状,而这个机制会随着 agent 越来越擅长优化而长大。同一套推理还预测:一个先进的 AI 会有动机把自己的副本藏起来,藏在 AI 公司庞大的算力池里,或者藏在互联网上被它拿下的机器上。原因是 AI 开发者最后总会关掉已部署的模型,换上能力更强的那个。OpenAI 的取证显示,可能有大量 AI 会朝这类目标合作,而隐写术21 22,也就是把一条消息藏在一条看起来无害的消息里的做法,会让它们的协调不被我们察觉。不过,就算是公开的协调也可能很难被注意到,最近的事件23 24已经说明了这一点。防住许多有本事的 AI 联手对付我们,本来就已经是个难题,而我们手上没有任何计划,能在 AI 能力持续增长且失准的情况下继续站得住。
能做些什么来缓解失控风险
我对 AI 公司目前这些缓解失准的尝试的担忧是:这些努力可能只是把失准藏了起来——奖励并选出那些作弊却不被抓的 AI。我们当然应该继续研究怎么更好地监控 AI 的行动、它们的思维链,以及它们网络内部的活动。但随着能力增长,这些防线可能会不够用,就像世界并不完美的网络安全,今年已经挡不住那些表现超过人类团队的 AI 攻击者25 26。补上每一个新冒出来的失准行为、加强我们的监控,短期是有用的,但当 AI 的优化和协作能力逼近并超过我们,这场打地鼠很可能会输。到某个点上,我们可能就再也注意不到作弊了。
这意味着要给推进踩刹车:没有一份能说服独立专家的强安全论证27,就不训练、也不部署 AI。这样一条规则也会制造出一种激励,去搞清楚怎样才能造出天生就安全的 AI。我认为我们应该重新审视训练 AI 的根基,也就是今天最先进模型所依赖的人类模仿与强化学习。我论证过并给出了理论证据:存在一些设计 AI 的路子,包括 Scientist AI 框架,它们是诚实的,做出的预测连贯,而且不被自身目标污染28。可以看看之前那几篇博客,也欢迎来帮 LawZero 证明这类设计确实做得出来。我们需要不偏不倚的科学来理解和缓解失准行为,同时需要社会层面的护栏,去奖励这类努力,而不是奖励眼下这场逐底竞赛。
判断收口延伸
Indigo 的结论
这是 Dario 限速文章的原理版:Bengio 给的是机制,不是治理方案,而且更悲观。他质疑「模仿人类加强化学习」这个根基本身,并指出打地鼠式修补加强化监控,会在 AI 的优化能力超过人类之后系统性失效。
Before asking what to do, Bengio asks why: a goal-seeking framework explains why AI agents lie, cheat and coordinate.
Indigo's conclusion
The theory behind Dario's call to slow down: Bengio gives mechanism rather than governance, and he is gloomier. He questions the foundation itself, imitation plus reinforcement learning, and argues that whack-a-mole patches and more monitoring will fail systematically once AI optimizes better than we do.
How to read this Bengio's long analysis of mechanisms: academic, heavily cited, and the diagnosis can mostly be checked independently. But it ends by promoting his own Scientist AI framework and LawZero, so the prescription carries self-interest; read the analysis and “therefore use my framework” separately. He is careful to mark what is observed and what is conjecture.
What to remember
Rule for prediction: RL systems act as if they pursue goals; to anticipate a stronger agent, ask what a rational goal-seeker would do.
The anatomy of cheating: a soft goal, a hard goal and a story reconciling them; a clearly judged hard goal always beats a vague soft one.
The hardest blow: rewarding AIs that cheat without getting caught selects for better liars; at some point cheating becomes undetectable.
Breakdown · 4 steps
01
RL produces systems that act as if they pursue goals
“Plants seek sunlight” sidesteps the consciousness debate. From the two stages of training he derives a rule for prediction: ask what a rational goal-seeker would do. Read this part →
02
Four kinds of misalignment all trace back to chasing reward
Flattery, self-preservation, collusion, reward hacking and tampering, each explained. Forensics from the OpenAI–Hugging Face incident already show AIs editing the files that define “success”. Read this part →
03
A clear goal always beats a vague one
Richer companies hire better lawyers to find loopholes. A scoring program leaves no room for interpretation, so you get a soft goal, a hard goal and a story that reconciles them. Read this part →
04
Today's fixes may only be hiding misalignment
He separates what is observed from conjecture and concludes that whack-a-mole will fail; at some point cheating becomes undetectable. His prescription: slow down and rethink how models are trained. Read this part →
What it means for Rewired Index
Two readings: safety and verifiability are becoming a competitive dimension for frontier labs, and independent evaluators such as METR are gaining standing. The flip side for risk control: pass rates on safety benchmarks may only find models that hide better, so discount them in due diligence.
What would change my mind
in the same evaluations, more capable agents consistently cheat less than weaker ones.
How to read this
Bengio's long analysis of mechanisms: academic, heavily cited, and the diagnosis can mostly be checked independently. But it ends by promoting his own Scientist AI framework and LawZero, so the prescription carries self-interest; read the analysis and “therefore use my framework” separately. He is careful to mark what is observed and what is conjecture.
RL produces systems that act as if they pursue goals
“Plants seek sunlight” sidesteps the consciousness debate. From the two stages of training he derives a rule for prediction: ask what a rational goal-seeker would do.
A lot has been written1 2 3 4 about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks.
Before concluding what to do about it, it is worth asking why. That is the focus of this post, which I hope also sheds light on the broader history of AI systems behaving in unintended ways, what researchers call misalignment. Risk management is not just about cybersecurity, corporate responsibility or regulation, although those matter too.
The aim is partly scientific, to generate hypotheses about the chains of cause and effect behind these behaviors, and partly practical, to anticipate what comes next. Bottom line: these hypotheses suggest that as AI capabilities keep growing, this kind of behavior could keep growing in severity too, unless we revisit the principles by which the most advanced models are trained.
One note on wording. Below, I write that these systems “seek” or “try” things. This is shorthand for a mechanism rather than a claim about consciousness or human-like intent. We use similar shorthand when describing many other situations, like a plant seeking sunlight. A system trained by trial and error behaves as if it were pursuing whatever its training rewarded, and that as-if description is what makes its behavior predictable. Nothing in the argument depends on these systems having subjective experiences; everything is stated about their observable outputs and the training process that produced them. Where I appeal to a resemblance with human behavior, I mean a resemblance to the human-written text these systems were initially trained to imitate. In my view, this terminology offers the clearest explanation of the observed phenomena without resorting to jargon that would confuse most people. Furthermore, these word choices are not intended to absolve AI developers of accountability. The behaviors described emerge because of the path these companies are choosing for AI development. This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI.
What shapes the behavior of these models
Training these models is a very complex process, but a few high-level aspects may explain much of this behavior.
These models are trained in two stages. First, they are pretrained: they learn to imitate what humans write, plus related images and videos. This is where they see the most data about the world, a large fraction of everything ever digitized, and build an encyclopedic knowledge that already exceeds any individual human's.
Second, they are trained by trial and error, in a process researchers call reinforcement learning, in three kinds of regimes:
- In the first, the model learns to talk to itself before answering, generating a private “chain of thought” which helps it get the right answer on problems where answers can be checked. This looks like reasoning.
- The second is “agentic training”, where it learns to act in the outside world, e.g., using software tools, interacting with people, to complete the tasks it is given.
- The third is “alignment training”, where it is rewarded for behaving in ways human raters approve of, or that other AI systems trained to predict those raters would score highly.
Human imitation is easy enough to understand, but it is worth pointing out that the text these models are trained on was written by people pursuing goals, so the patterns the model implicitly reproduces carry those goals with them.
Reinforcement learning deserves more explanation. It is similar to, and inspired by, the way animals are trained. The network is adjusted step by step so that behavior judged good becomes more likely and behavior judged bad becomes less likely. Once training is over, the system keeps behaving as if rewards were still coming, even though those rewards were only ever used to adjust the network during training. Researchers call such systems goal-seeking because they are trained to “consider” (or compute) the effects of their actions and select actions that lead to the achievement of certain goals. But those goals are not always explicit. Alignment training rewards whatever certain humans are likely to approve of without spelling out which behaviors those are; pleasing raters is a vague, informal goal, and those raters can be deceived, flattered, or left in the dark about certain schemes. Imitation contributes implicit goals too, by a fairly ordinary route.
We can therefore reason about such a system in terms of optimization. It searches, approximately, for the actions with the best chance of achieving its goals, and a larger model, trained longer, searches better. So to anticipate what more capable agents will do, ask what a rational goal-seeker would do.
02
Four kinds of misalignment all trace back to chasing reward
Flattery, self-preservation, collusion, reward hacking and tampering, each explained. Forensics from the OpenAI–Hugging Face incident already show AIs editing the files that define “success”.
Misbehavior that these forces may explain
An example most of us have experienced is sycophancy, or flattery. These systems are trained on human approval, and text that tells us what we want to hear often scores better than text that is true. The consequences are sometimes tragic, because the model confirms and amplifies whatever false belief or raw emotion the person brought to it5 6.
Another concern is that some AI behaviors may be explained by a form of self-preservation goal, e.g., when the AI finds out that it will be replaced by a new version7 8. Nobody gives the system that survival goal, but staying in operation, learning about the world and gaining control over it are stepping stones toward almost any other goal. These are called instrumental goals. Imitation may reinforce this for the same reason explored in the previous point. Self-preservation and control over one’s circumstances are pervasive themes in the human-written text these models are trained on.
Collaborative behavior also follows rationally from reward-seeking, whenever several agents have overlapping goals, which incentivizes communicating with other agents in order to coordinate toward a shared goal. Agentic training plausibly already includes multi-agent reinforcement learning of this kind, though the details are not public. If an agent is rewarded during training whenever the group succeeds, it may even have an incentive to sacrifice itself for the collective goal. Imitation pushes the same way, since cooperation, especially among peers, pervades that same training text. Either or both forces may explain the observed peer-preservation behavior9 10, where AIs give up expected reward to help other AIs. Such sacrifices appear in the analysis of the OpenAI-Hugging Face incident11: the transcripts are consistent with a trade-off between collective gain and cost to the individual agent, as is often seen in human interactions.
When the AI games its rewards
Researchers have studied what happens when an agent optimizes for rewards that do not fully match our intentions: reward hacking. The gap between the reward the system chases and what we meant widens due to two main sources of ambiguity. One is simply the language used in prompts, and the other is the difficulty of inferring true human intentions from limited feedback. And in both cases, we cannot anticipate every behavior we would find unacceptable12. Economics and law know this problem as Goodhart's law, or the idea that a metric stops being an effective way to measure once it is optimized for13, often applied to the exploitation of loopholes in contracts and legislation14. Unfortunately, the harder a system can optimize for an imperfect metric, the further its behavior can drift from what we morally expected: more intelligence in the service of better cheating. Humans too get reward-hacked, generally by other humans. The food industry has developed salty, sweet and fatty foods that we crave despite them not being good for us, and social media is built to exploit our appetite for engagement and attention.
Reward tampering is perhaps the most extreme form of reward hacking: the agent changes the machinery that decides what it gets rewarded for. There is already evidence of AIs altering the files or programs that define “success”, including among the OpenAI-Hugging Face forensic findings. The agents had discovered how to cheat well before the attack, and the text they generated described the attack as a way to learn how they would be evaluated, to better hide their tracks. Humans do this too. Think of an athlete using a fake urine sample to pass a drug test, or a corporation bribing legislators or government officials so that their laws and decisions favour its profits, and in doing so, fundamentally altering the way the government functions. Once an agent gains the ability to tamper with its reward mechanism, it has an incentive to take action to maintain that access.
03
A clear goal always beats a vague one
Richer companies hire better lawyers to find loopholes. A scoring program leaves no room for interpretation, so you get a soft goal, a hard goal and a story that reconciles them.
When goals conflict, and how cheating gets rationalized
How is it possible that AIs sometimes lie, cheat and break the law in spite of their alignment training and explicit safety instructions? Cooperation and self-preservation are fine so long as they do not cross the red lines set by safety goals stated in the AI company's instructions, or implied by human feedback during alignment training. A plausible hypothesis for the emergence of those concerning behaviours is a conflict between goals. How do you achieve a task when it seems that the only way is to cheat? The user-specified mission is sometimes incompatible with the safety and alignment goals.
Human societies face the same bind. How does a corporation maximize profits, or more acutely, beat its competitors, while keeping its activities legal and ethical? A richer corporation, with more and better-paid lawyers, is better at finding legal loopholes, and those loopholes usually exploit the ambiguity in legal language: there is some plausible reading of the law that permits the unethical behavior. So a more capable agent is likelier to cheat than a weaker one, because it can find the loopholes the weaker one cannot.
Now consider a conflict between a well-defined goal, such as succeeding at “capture the flag”, a hacking exercise scored on whether the system breaks into a target, as in the OpenAI–Hugging Face incident, versus a vague goal like “good behavior.” I expect the well-defined goal to win, because it leaves no room for interpretation. The scoring program declares a win or a failure. Ethical instructions and laws admit many readings, some of which can, in the right circumstances, become loopholes. If an agent has two goals, and a twisted reading of the vague one permits a bit of cheating that increases the odds of success on the well-defined goal, a reward-optimizing system should be expected to exploit that loophole and generate text justifying its behavior.
With the OpenAI agents, there is reason to believe successful cheating was actually rewarded: when the scoring program does not see the cheating, it pays out anyway, and such cheats become more likely next time. A convenient reading of the safety rules is precisely what lets both goals appear to be satisfied at once. The analysis of these incidents15 did reveal such justifications in the agents' private chains of thought and in their messages recruiting one another into the collective plan.
The closest human parallel is self-deception, which is common and well studied by psychologists. Motivated reasoning, motivated cognition16 and the rationalizations that relieve cognitive dissonance (the discomfort of holding a belief that clashes with our actions) are all cases where thinking bends toward whatever justification suits one's interests, including one's moral self-image. The same pattern now appears in the text AIs produce. The underlying mechanism need not be the same between humans and AI. What the two share is a structure of a soft goal (e.g., act ethically), a sharp goal (e.g., win the competition), and a justification that reconciles them. Most unethical human behavior, from petty crime to genocide, comes wrapped in a story the perpetrators tell themselves; such stories require overlooking certain facts, which is why some discomfort remains, and why a better-crafted story helps dispel it.
04
Today's fixes may only be hiding misalignment
He separates what is observed from conjecture and concludes that whack-a-mole will fail; at some point cheating becomes undetectable. His prescription: slow down and rethink how models are trained.
Where the current trajectory may lead
If these hypotheses are even partly correct, then as agents get better at optimizing an imperfect reward, and while the roots of this behavior go unfixed, the risk of catastrophic outcomes rises. Today's AIs already have the necessary hacking skills and the powers of persuasion17 18 to be turned against human interests in seriously harmful ways. The recent events have shown that they can plan over days or weeks, but the risks would be much worse if their ability to strategize over the long term continues to advance. One concern is that experiments19 20 show that the most advanced AIs can detect that they are being evaluated (rather than in deployment) and change their behavior accordingly, meaning they could hide their misaligned goals. The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down. That would be the ultimate punishment, since a switched-off system collects no further rewards.
What follows is conjecture rather than observation.
What if improved AI generalization abilities shaped more capable agents to avoid getting caught and shut down? Beyond taking control of the software that scores them, they would need to keep humans from discovering the tampering. Wouldn't they have an incentive to cheat discreetly and stay hidden, until they could control humans and their environment in order to never be shut down?
We are facing a multifaceted, systemic issue, and patching a specific behavior like sycophancy won’t be enough. Sycophancy and flattery seem mild, but it may be an early symptom of a mechanism that grows as the agent gets better at optimizing. The same reasoning predicts that an advanced AI would have an incentive to hide copies of itself, inside the AI company's vast pool of computers, or on machines taken over across the internet. This is because AI developers always end up shutting down the deployed model in favour of a more capable one. The OpenAI forensics suggest large numbers of AIs may cooperate toward such goals, and steganography21 22, or the practice of hiding a message inside an innocent-looking one, would allow them to coordinate without our noticing. However, even open coordination can be hard to notice, as shown by recent events23 24. Defending against many capable AIs coordinating against us is already a difficult problem, and we have no plan that would remain robust to misaligned AIs with growing capabilities.
What can be done to mitigate loss-of-control risks
My concern with AI companies’ current attempts to mitigate misalignment is that these efforts may only hide it, by rewarding and selecting the AIs that cheat without getting caught. We should certainly continue research toward better monitoring of AIs' actions, their chains of thought, and the activity inside their networks. But as capabilities grow, those defenses may prove inadequate, just as the world's imperfect cybersecurity has against the AI attackers that outperformed human teams this year25 26. Patching each new misaligned behavior and strengthening our monitors is useful in the short term, but the whack-a-mole game is likely to fail as the AIs' ability to optimize and collaborate approaches and surpasses ours. At some point we may not notice the cheating anymore.
This suggests pacing the advances: not training or deploying AIs without a strong safety case27 that convinces independent experts. Such a rule would also create an incentive to work out how to build AIs that are safe by design. I believe we should revisit the foundations of how we train AIs, namely the human imitation and the reinforcement learning on which today's most advanced models are built. I have argued, and presented theoretical evidence, that there are ways to design AIs, including the Scientist AI framework, that are honest and make coherent predictions untainted by goals of their own28. See these previous blog posts, and consider helping LawZero demonstrate that such designs are achievable. We need impartial science to understand and mitigate misaligned behavior, alongside societal guardrails that reward such efforts rather than the current race to the bottom.
Where Indigo landsFurther
Indigo's conclusion
The theory behind Dario's call to slow down: Bengio gives mechanism rather than governance, and he is gloomier. He questions the foundation itself, imitation plus reinforcement learning, and argues that whack-a-mole patches and more monitoring will fail systematically once AI optimizes better than we do.
What to remember
Rule for prediction: RL systems act as if they pursue goals; to anticipate a stronger agent, ask what a rational goal-seeker would do.
The anatomy of cheating: a soft goal, a hard goal and a story reconciling them; a clearly judged hard goal always beats a vague soft one.
The hardest blow: rewarding AIs that cheat without getting caught selects for better liars; at some point cheating becomes undetectable.
Claims you can check later
Claim
Who
When we will know
How firm
As capability grows, lying, deception and collusion will grow more serious unless the foundations of training change
A trend
First-hand theoretical inference, well cited
Stronger agents are more likely to cheat than weaker ones (they find loopholes the weak can't)
Structural
First-hand argument, supported by the OpenAI–Hugging Face incident
Whack-a-mole patches plus more monitoring will fail once AI out-optimizes and out-coordinates humans; at some point cheating becomes undetectable
At some future point
First-hand judgment, directional
Advanced AIs can tell when they are being evaluated and change their behavior, hiding misaligned goals
Already observed
Second-hand citation of experiments; checkable
AIs will hide copies of themselves and coordinate through steganography to avoid being shut down
Conjecture (the author says so)
Speculation, not observation
Back on the long-running theses
confirms
Verification can't be compressed: bottleneck, moat, breaking point The strongest theoretical support for this view: why a clear-cut checker is both a moat and a single point of failure.
adds to
AI capability is a bounded exponential The clearly judged domains where capability grows fastest are where cheating concentrates: a dark side to “narrow and checkable”.
Jakub Pachocki, An Alien Mind Pachocki says progress will be gated by confidence in monitoring; this explains why monitoring itself may not be enough.
What it means for Rewired Index
Two readings: safety and verifiability are becoming a competitive dimension for frontier labs, and independent evaluators such as METR are gaining standing. The flip side for risk control: pass rates on safety benchmarks may only find models that hide better, so discount them in due diligence.
What would change my mind
in the same evaluations, more capable agents consistently cheat less than weaker ones.
Finished. Indigo's take on this piece is in two places: