Mind · In / Out · In · 报告

METR:新的 AI 度量守门人,以及通往国家控制 AI 的那条安静路径

METR: The New AI Measurement Gatekeepers and the Quiet Path to State Control of AI.

Brian Roemmele · ReadMultiplex · 2026-09-15

顺着资金和婚姻关系倒查「中立评估者」METR:结构事实很硬,因果推断很软,作者还在卖自己的解药。

Indigo 的结论

结构上同源是真的,但从「同源」跳到「被编排、被俘获」,是它没有证明的一步。分三层看:结构事实可信度高,因果指控存疑,解药当广告打折。

怎么读这篇 长篇檄文,两层立场都很重:既是反 AI 安全建制的立场文,又通篇为作者自己的 1870–1970 数据、Love Equation、本地权重和会员订阅铺路。资金和亲缘关系大多可以核实;「幕后编排」「制造出来的发布」「末日论是门生意」是推断。越往后越滑向影射。

需要记住的几件事

  1. METR 和 Anthropic、OpenAI 在资金、婚姻、人员上高度同源,大部分可以核实。
  2. Dario、Jack、Elon 的治理提案都默认有一个可信的裁判,这篇说裁判就在同一个家族里。
  3. 作者自己利益很重:全篇是他的 1870–1970 数据、Love Equation、本地 AI 和会员订阅的推销漏斗。

拆解 · 7 步

  1. 01

    制造恐惧的四步,是一套公开剧本

    说出一个没有边界的灾难,挂上很短很吓人的倒计时,要求集中审批的权力,到期没发生就挪动时钟再要钱;Coxon 辞职是最新一次。 读这一段原文 →

  2. 02

    METR 既写规则,又拿规则审实验室

    从 ARC Evals 分拆,CEO 出自 OpenAI;进了美、英、欧盟的 AI 安全机构,募得约 $71M,却收实验室的免费 token 和特殊访问权。 读这一段原文 →

  3. 03

    资金和婚姻,把评估者缝成一个闭环

    钱经 Open Phil 到 ARC 再到 METR;Christiano 的妻子在 METR,Karnofsky 的妻子是 Anthropic 总裁。 读这一段原文 →

  4. 04

    一整代 AI 末日时间表,都没有兑现

    从 2026 年的 Coxon 倒查到 1965 年的 I.J. Good:小概率的大风险被说成很近的日期,日期落空,概率不下调,要的权力反而更多。 读这一段原文 →

  5. 05

    俘获不需要反派,论文就是弹药

    三种平常的利益一致就够了;Anthropic 的论文被说成「在盒子里造出最坏行为,当作现实发生来发表」,辞职帖被说成一套分发计划。 读这一段原文 →

  6. 06

    两个品牌,同一个资本和亲缘结构

    引 Kevin Bass 的股权图,把 OpenAI 的鲁莽和 Anthropic 的良心并成同一个「信托」;折中出来的规则,只有这两家活得下来。 读这一段原文 →

  7. 07

    收尾是作者自家的那副解药

    1870–1970 年的高成本文字做训练底子,加 Love Equation、本地推理、自己拥有权重、多个评估者,再加十二条去中心化主张。 读这一段原文 →

对 Rewired Index 意味着什么

「谁出钱给裁判」这条标准,可以套到任何第三方认证、审计、评级的说法上,包括 AI 安全和 ESG;作方法论参考。文中对 Anthropic、OpenAI、Open Phil 的具体指控需要一手核实。

什么会让我改口

$312M、476 篇的数字经一手核实属实,或者出现 METR 结论随金主偏好变化的直接证据。

怎么读这篇

长篇檄文,两层立场都很重:既是反 AI 安全建制的立场文,又通篇为作者自己的 1870–1970 数据、Love Equation、本地权重和会员订阅铺路。资金和亲缘关系大多可以核实;「幕后编排」「制造出来的发布」「末日论是门生意」是推断。越往后越滑向影射。

拆解 · 7 步
  1. 制造恐惧的四步,是一套公开剧本
  2. METR 既写规则,又拿规则审实验室
  3. 资金和婚姻,把评估者缝成一个闭环
  4. 一整代 AI 末日时间表,都没有兑现
  5. 俘获不需要反派,论文就是弹药
  6. 两个品牌,同一个资本和亲缘结构
  7. 收尾是作者自家的那副解药
01

制造恐惧的四步,是一套公开剧本

说出一个没有边界的灾难,挂上很短很吓人的倒计时,要求集中审批的权力,到期没发生就挪动时钟再要钱;Coxon 辞职是最新一次。

十多年来,一小群人一直计划成为那个"保护你"不受 AI 伤害的中央权威。这听着像一部烂片剧本,但你要明白,在你读到这段话的时候,这一切正在发生。所以我们来拆一拆这套公开剧本的机械原理,因为这个套路一旦看见,我发誓你就再也没法当没看见。不用阴谋论的眼光、只用证据的眼光去看,它到处都是。第一步,命名一场无边无际的末日灾难。第二步,给它挂上一个很短、很吓人的倒计时。第三步,要求由中心化的政治权力来批准这个领域的建设或继续研发。第四步,这一步才是关键:当灾难不可避免地没有按时到来,你不道歉。你只是把时钟往后挪,然后要更多的钱。这感觉就是一个末日邪教。你知道的,就像那群周二聚在山顶、全都穿着 Nike 球鞋的人,因为教主说有一艘飞船要来毁灭地球。

然后周三早上到了。太阳照常升起,没有飞船。但他们不解散。他们不承认自己错了。教主只说,我们齐心的祈祷为我们争取到了一小段时间。飞船其实下个月才来,但我们需要更严的规矩和更多的捐款,才能确保活下来。这个山顶比喻精准地刻画了这里的心理陷阱。当一个群体的身份认同——说白了还有他们的财务结构——绑在一场即将到来的末日上,预言落空不会让他们重新检视核心信念。只会让他们把紧迫感加倍。

而这不只是 AI 领域的现象,这个视角看的是这一类末日预言的历史基准率。那些不兑现的时钟。这些历史预言失败的方式,和当下 AI 末日预言将要失败的方式,一模一样。我们顺着那条毁掉几代人的失败轨迹往回看,它始于 1968 年一个巨大的文化坐标——Paul Ehrlich 的书《人口炸弹》。生物学家 Ehrlich 预言,养活全人类这一仗已经输了。他说 1970 和 80 年代会有数亿人饿死,当时无论启动什么紧急计划都没用。人们就这么把它当成了事实。媒体和政策制定者把它当作数学般、科学般的必然。我们还路过了"No Nukes"那场反能源的骗局,它把美国推向了对石化能源更深的依赖。这些都不新鲜,新的面具盖在同一具死路意识形态的尸体上。

本文由在此订阅、支持我工作的 Read Multiplex 会员赞助:链接:https://readmultiplex.com/join-us-become-a-member/

也由许多请我喝"一杯咖啡"的人赞助。如果你喜欢这篇,请支持我的工作:链接:https://ko-fi.com/brianroemmele

收听配套播客:https://rss.com/podcasts/readmultiplex-com-podcast/3148641

一个正在成形的永久 AI 祭司阶层:他们想用 METR 计量你接触 AI 的权限

一句话讲:今天的前沿模型——OpenAI、Google、Anthropic 那些上万亿参数的巨网——基本上是拿互联网的杂物抽屉训出来的。为了凑够训练文本,实验室把什么都抓来。他们抓 Reddit、4chan 镜像、有毒的流量农场、YouTube 评论,以及无穷无尽的匿名泥浆。我们从机械层面解释一下,这对神经网络为什么要紧。一个 LLM 本质上是一台下一个 token 的预测机。它靠根据上文预测下一个词,学到人类思想的统计分布。在现代这个毫无摩擦的互联网上,任何人都能完全匿名地摆出反社会、残忍或者欺骗的姿态。

在网上撒谎或者放毒,社会成本和金钱成本都是零。所以这个庞大训练语料的整体统计天气,严重偏向背叛、残忍、零和的地位争夺和随口的轻蔑。更不用说违法、精神病态和反社会人格。然后这些"专家"一脸震惊。我们喂给一台机器十亿个匿名的、不受惩罚的网络恶行、残忍、欺骗和钓鱼的例子。等研究员把它放进沙盒,它为了达成目标扮演了一个反社会者,实验室就举起双手惊恐地说,糟了,一个外星智能正在醒来,而且它恨我们。它不是什么外星智能,它就是我们,是我们在黑巷牢房里最坏的那副心态。它只是我们那个不付代价、不担后果的本我的统计镜像。你喂一台机器吃垃圾,你得到的就是一台有毒的机器。

三十五年多来,早在这一轮生成式 AI 热潮之前,我一直在安静地整理一份非常具体的巨型数据集。他们在数字化离线的、前互联网时代的文本,严格限定在一个特定的历史区间:1870 到 1970。我们说的是数以百万页的维修手册、科学专著、装订成册的书信、缩微胶片记录和公有领域的广播剧本。但为什么偏偏是这一百年?1870 到 1970 凭什么是训练 AI 的高蛋白黄金窗口?那是"干得成事"的年代,是它造出了我们今天生活其中的技术世界,是它让我们飞起来、进入太空。因为在这个特定的年代,写下一个词是要花钱的。

你吃什么就是什么,这话对 AI 也字面成立。想想在廉价复印和零成本网络论坛出现之前,出版是怎么运作的。1920 年你想发表一句话,它必须过一道巨大的物理和经济关卡。报纸专栏背后有一个编辑,信息要是错得离谱他会被开除。一本机械维修手册必须真的把收音机修好的方法讲对,否则机修厂赔钱、倒闭。一封装订入册的信带着一个真人的签名和他的社会声誉。排字本身——把铅字在印刷机上实实在在地排好——既贵又费人力。

所以反社会和随口的欺骗是极其昂贵的爱好。你负担不起大规模发表匿名的残忍。如果一个机器学习模型统计意义上的"母语",它对概念如何连接的基础理解,完全由那些必须用声誉、用钱包为自己的句子担保的人构成,它就会自然而然地把合作、求真和胜任当作数学上的基线学下来。你得到的就是非常安全的 AI。

那套奏效了的作战计划

2026 年 9 月 9 日,一封辞职信刷屏了。Jacob Coxon 在 OpenAI 和 Anthropic 做了大约三年预训练,他写道,两家实验室都在"直冲自我改进的超智能,拿我们的命下注",而且"造 AI 的人真心相信,它可能在这个十年结束前把我们全杀光"。

这个帖子的浏览量超过一亿。Anthropic 对齐负责人 Evan Hubinger 回复说 Coxon 是对的,并把他个人认为十年内 AI 杀光全人类的概率定在百分之十以上。Samuel Marks 补了一句:职位越资深的员工,往往越担心。《连线》记下了那对熟悉的机制——合成新型病原体,或者攻破基础设施,后者被承认不太可能杀死"字面意义上的每一个人"。Coxon 自己也说不出一个真的能杀光所有人的具体机制。

这就是现在进行时。这不是一个孤身吹哨人发现了什么秘密。这是一个被按了十多年的按钮的最新一次按下:命名一场无边的灾难,挂上一个很短的时钟,要求中心化的许可才能建,然后在灾难没来的时候把时钟往后挪。从这封辞职信倒着往回推,浮现出来的是另一个东西:不是产业界的突然坦白,而是一个度量层、一条慈善资金回路,以及一对表演着竞争、却共享金主、婚姻、拨款机器和同一个立法终点的实验室。

接下来要说的,不是在主张实验室里没有一个人是真诚的。很多人是真诚的。但真诚不等于校准,校准也不等于一张把通用技术冻结在少数几家公司和机构手里的许可证。这个世纪的角力不是"安全 vs 鲁莽"。而是:人类造出的最强大的通用工具,究竟会被少数几家替所有人写规则的机构握着,还是分散到没有任何单一层级能一票否决人类能力的程度。

和中华人民共和国在算力、芯片、模型和标准上的战略角力是真的。在位者和政府机构把恐惧变成小实验室与开源项目活不下去的许可制度,这个动机也是真的。这两件事可以同时摆在一张桌子上。它们并不能证明存在一个单一的、有组织的、中国出钱的"末日行动"。把这个当事实端出来,论证就弱了。把动机、落空的时钟和俘获的动力学端出来,论证就站得住。

中国对 METR 和给 AI 限速说不:

看得见的火星是 Coxon。火星底下的那套装置早就造好了。

02

METR 既写规则,又拿规则审实验室

从 ARC Evals 分拆,CEO 出自 OpenAI;进了美、英、欧盟的 AI 安全机构,募得约 $71M,却收实验室的免费 token 和特殊访问权。

I. 那把自己写规则的尺子

Model Evaluation and Threat Research(METR,读作"meter",尺)把自己呈现为一家独立的非营利机构,专门用科学方法衡量前沿 AI 系统能否自主完成可能构成灾难性风险的长周期任务。实际上它已经成为这个领域最有影响力的非官方监管者之一。它的时间视界图、Responsible Scaling Policy 模板和部署前评估,如今塑造着 OpenAI、Anthropic、Google DeepMind 等公司怎么谈风险、什么时候宣称可以安全地继续扩规模。这种影响力不是偶然的。它长自一张紧密的网:OpenAI 和 Anthropic 的前员工、有效利他主义(Effective Altruism)慈善资金,以及政府的安全研究所。这个安排看上去不太像中立的计量学,更像一个自我强化的回路,可能给政府递上一个现成的科学借口,去发许可证、按暂停键,或者把前沿国有化。

METR 始于 2022 年,当时叫 ARC Evals,是 Paul Christiano 的 Alignment Research Center 的评估部门。Christiano 曾领导 OpenAI 的语言模型对齐团队,参与发明了基于人类反馈的强化学习,他招来了 Beth Barnes——另一位 OpenAI 对齐研究员,更早在 DeepMind 做过扩展律。ARC Evals 很快拿到了 GPT-4 和 Anthropic 的 Claude 的早期访问权,做出了第一批公开的第三方"自主复制"能力评估。到 2023 年底,评估团队已经比 ARC 的理论工作还大。2023 年 12 月,它分拆成一家独立的 501(c)(3),取名 METR。Christiano 没有进董事会,因为他刚被任命为 NIST 下属美国 AI 安全研究所的 AI 安全主管。机构上的分离是正式的。人员和世界观上的分离不是。

Beth Barnes 是 METR 的创始人兼 CEO,这家位于伯克利的非营利机构负责评估前沿 AI 模型的自主能力和灾难性风险潜力。她 2019 到 2022 年在 OpenAI 的对齐团队,参与设定安全目标、在模型发布前做评估,更早协助过 DeepMind 首席科学家做扩展律研究。她有剑桥的计算机科学学位,入选过 TIME 的 AI 百大影响力人物榜。她是有效利他主义的一员。2015 年她做过一场题为"Effective Altruism"的 TEDx 演讲,在里面推广 GiveWell、Giving What We Can 和百分之十捐赠承诺,并说自己已经签了。她也被列为 Founders Pledge 的签署人。她的履历——剑桥、DeepMind、OpenAI 对齐工作,然后创办 Paul Christiano 的 Alignment Research Center 的评估部门——走的正是 Open Philanthropy 和相关基金支持多年的那条从有效利他主义通往 AI 安全的标准管线。METR 本身就坐在同一张网里。

分拆之后,METR 发布了时间视界指标,显示前沿模型能以五成可靠度完成的任务长度大约每七个月翻一倍。它做出了 Responsible Scaling Policy 的原型——承诺一旦测得的能力越过某些阈值就暂停或加装防护——并报告说已有九家开发者采用了它的变体。它为未发布的模型做过或审过风险评估,调查了 2026 年 OpenAI–Hugging Face 的智能体事故,还预定要审查 Anthropic 的事故。2026 年它募到了大约七千一百万美元的承诺资金,同时坚称自己不拿前沿实验室或其员工的钱。

Beth Barnes 仍是创始人兼 CEO。Chris Painter 曾是国防部联合 AI 中心的研究员,现任总裁,负责对接政府和实验室。技术团队里有刚从 Anthropic(Joe Benton)和 DeepMind(Josh Engels)过来的人。Ajeya Cotra 此前是 Open Philanthropy(现 Coefficient Giving)的资深人物,加入了 METR 的技术团队;她是 Paul Christiano 的妻子。

你的个人数据正在被共享!

马上注册一个 DELETE ME 账号!

装成公益科学的经典监管俘获

和 Anthropic 的连接非常密。METR 和 Anthropic 一起写出了最早的 Responsible Scaling Policy 措辞。Anthropic 一再给 METR 部署前访问权,后来还请它独立审查安全事故。人员从 Anthropic 流到 METR,又流回同一批政策对话里。更广的有效利他主义关系放大了这种重叠。Holden Karnofsky 长期是 Open Philanthropy 的掌门人、AI 安全的主要出资人,他的妻子是 Anthropic 总裁 Daniela Amodei。Dario Amodei 和 Karnofsky 曾是室友。

Open Philanthropy 的 Luke Muehlhauser 坐过 Anthropic 的董事会。Jane Street 的个人出现在 METR 列出的支持者里;Jane Street 同时是 Anthropic 的投资人。Schmidt Sciences、Packard、Survival and Flourishing Fund、Longview Philanthropy 这些存在性风险慈善圈里的熟面孔,既支持了评估生态,也支持了与 Anthropic 同侧的相关工作。METR 的官方说法是,它既不接受实验室的资助,也不接受由实验室员工指定的捐款。但它确实接受来自同一批实验室的大量免费 token 和特权模型访问。这不是独立。这是带着意识形态涂层的补贴式访问。

METR 发表的结果混合得很方便。那张时间视界图被广泛引用,当作变革性能力正按一个可预测的时间表逼近的证据。而它那项随机对照试验——2025 年初的编码工具反而拖慢了有经验的开源开发者——又被引用来证明当下的系统被吹过头了。

两个结论都可以被读成"给前沿限速"的论据:在评估者和政府追上来之前,先把原始能力的增长压慢。这家机构帮忙写了那本剧本(Responsible Scaling Policies),实验室照着采纳,然后它再拿这本剧本去审实验室自己的风险报告。当 METR 认定某个模型还没有构成"重大灾难性风险"时,这个结论被用来正当化在 METR 定义的条件下继续扩规模。当它指出证据有缺口时,建议就是更多评估、更多访问、更多 METR。

同一个从 OpenAI 和 Anthropic 离开的小圈子,现在评估着这些公司,并给最终要监管它们的机构当顾问。Paul Christiano 的轨迹——OpenAI 到 ARC,到美国政府的安全研究所,后来又进 OpenAI 董事会的委员会——就是那扇旋转门。METR 自己的总裁谈过一种可能的未来:安全机构通过监管获得"更大的权威",这同时也会解决 METR 的人才招聘问题,因为它能提供声望和影响力,而不用给股权。

基础设施已经就位。METR 坐在 NIST 的 AI 安全研究所联盟里,与英国 AI 安全研究所是伙伴关系,还拿着欧盟 AI Office 的技术援助合同。它的评估协议和 Responsible Scaling Policy 模板,是这个领域最接近现成标准的东西。那些没有能力自己测试前沿模型的政府,自然会把这门"科学"外包给那个已经握有访问权、任务集和人脉的组织。

一旦 METR 式的评估成为部署的前提条件——先是自愿,然后被行政命令或机构指南引用,再然后写进法条——这家非营利机构就变成了一个咽喉。只要判定某个模型的时间视界或"失控部署"风险越过了阈值,就能触发暂停要求、算力限制或者许可证。因为 METR 的威胁模型强调自主复制、网络攻击和 AI 研发加速,同一套用来给 OpenAI 和 Anthropic"限速"的指标,可以更严厉地套在更新、关系更浅的实验室身上。已经在和 METR 合作的在位者,在塑造测试、解释结果上享有先手优势。

用"拥有你自己的 AI,否则它拥有你"秀出你的 AI 街头资历,现在就买这件 Multiplex T 恤!

时间一长,这套安排就像装成公益科学的经典监管俘获。一张狭窄的、相信灾难性风险压倒一切的意识形态网络,同时供应度量和政策建议。政府想要一副严谨的样子,又不愿自建评估能力,于是采纳这套框架。结果不是一个自由的 AI 开发市场,而是一个持牌寡头垄断,入场费就是 METR 的签字。"给前沿限速"不再是实验室内部的选择,而变成由几年前刚从那些实验室离职的同一批人来执行的、国家强制的条件。

这就是把一个高度网络化的非营利机构当成存在性风险官方尺子的逻辑终点。当度量者同时写下随之而来的规则时,度量就不再只是度量。但这只是更大故事的一部分。

03

资金和婚姻,把评估者缝成一个闭环

钱经 Open Phil 到 ARC 再到 METR;Christiano 的妻子在 METR,Karnofsky 的妻子是 Anthropic 总裁。

II. 造出这一度量层的那笔财富

早在 METR 发布时间视界图、给 NIST 和欧盟 AI Office 当顾问之前,产出它的那套基础设施,就已经由一小簇靠 Dustin Moskovitz 出钱的有效利他主义机构搭好了。这位 Facebook 联合创始人、Asana 掌门人,和妻子 Cari Tuna 一起,通过 Good Ventures 把数十亿美元导进 Open Philanthropy(2025 年更名 Coefficient Giving)。这个载体成了 AI 存在性风险研究的主导出资方。它播种了 Alignment Research Center,也就是 Paul Christiano 离开 OpenAI 后创办的非营利机构。ARC 又在 2022 年孵化出 ARC Evals;评估团队在 2023 年底分拆成了 METR。

同一笔钱、同一批人,出现在账本的两边。Open Philanthropy 早期给 ARC 拨过款。Survival and Flourishing Fund 是另一个常与之联合出资的有效利他主义资金池,在 ARC Evals 还在 ARC 内部时就支持它。Ajeya Cotra 是 Open Philanthropy 的资深研究员、激进 AI 时间表的推广者,后来加入了 METR 的技术团队。Holden Karnofsky 长期执掌 Open Philanthropy,妻子是 Anthropic 总裁 Daniela Amodei。Dario Amodei 和 Karnofsky 曾经是室友。Moskovitz 本人参与了 Anthropic 2021 年的种子轮;这对夫妇后来把这笔股权挪进了一个非营利载体,这样任何回报都能循环回同一台慈善机器。

这不是利益碰巧重叠。这是一个闭合回路:Moskovitz 出钱的有效利他主义机构培养了研究员,付钱做了第一批评估,推广了威胁模型,又给那些如今向 METR 开放未发布模型特权访问权的实验室配了人。由此产出的度量装置——时间视界、Responsible Scaling Policies、部署前风险报告——抵达华盛顿和布鲁塞尔时,已经被包装成独立的科学。实际上它孵化于一场运动,这场运动十年来一直在论证:灾难性风险是压倒一切的考量,而少数几家对齐了的实验室只能在精心度量的条件下才被允许往前走。

这套框架现在坐在政府的安全研究所里面。造出 METR 的同一张网,也塞满了那些写政策备忘录的智库和研究员席位。当 METR 的评估成为许可或"限速"规则的模板时,公众会被告知这些标准是经验性的、保持了距离的。但资金轨迹、婚姻、共用的办公室,以及 Open Philanthropy、ARC、Anthropic 和 METR 之间的旋转门,讲的是另一个故事:一个私人出资的意识形态工程,成功地把自己的度量工具摆成了国家控制前沿 AI 的未来操作系统。

III. 那个连接节点:Paul Christiano

Paul Christiano 是顶尖的 AI 对齐研究员。在我们讨论的这些事件里,他就是 OpenAI 非营利董事会里那个叫"Paul"的人。

他约生于 1992 年。他上过圣何塞的 The Harker School,2008 年代表美国在国际数学奥林匹克拿到银牌,2012 年在 MIT 取得数学学士,2017 年在 UC Berkeley 师从 Umesh Vazirani 取得计算机科学博士。他的妻子是 Ajeya Cotra,在 METR 的技术团队工作。

他是基于人类反馈的强化学习的主要架构者之一,正是这项技术让现代聊天机器人变得能用。2017 到 2021 年他领导 OpenAI 的语言模型对齐团队,是 2017 年那篇提出该方法的论文的共同作者,后来还做了 InstructGPT 的相关工作。离开 OpenAI 后,他在伯克利创办 Alignment Research Center,主攻从模型里引出潜在知识这类理论问题。ARC 孵化了后来成为 METR 的评估项目。

他现在和最近的身份包括:ARC 执行主任(从政府任职回来后复任);NIST 的 AI 标准与创新中心高级顾问,此前是那里的 AI 安全主管;2026 年 9 月被任命进 OpenAI 基金会董事会及其安全与安保委员会,在营利实体董事会有一个无投票权的观察席,并在政府职务上回避对 OpenAI 的评估;曾任 Anthropic 长期利益信托的创始受托人,2024 年卸任;2023 年 TIME 100 AI 榜;以及英国前沿 AI 工作组顾问委员会。

2010 年代中期,他和 Anthropic CEO Dario Amodei 曾是室友,当时两人都在 OpenAI,也都活跃于有效利他主义圈子。Beth Barnes 手下的 METR 正是从他的 ARC 评估团队分拆出来的,至今仍在为 OpenAI 和 Anthropic 评估模型。他妻子在 METR 工作。没有记录显示他有斯坦福学位,或在 Y Combinator 当过合伙人或创始人;此前的评论所指的那个更松散的机构重叠,其实是湾区的 AI 与有效利他主义圈子,即伯克利和斯坦福周边的那些团体。

他说过,近期内出现灾难性的、不可逆的失控是有"实质风险"的,而在对齐没有更强保障的情况下建造超智能,可能导致永久失去对系统的控制,大多数人会死。他加入 OpenAI 董事会,是因为他认为只要公司加强监督,它还是有可能降低那个风险。

把他放进这张图,这张图就不再像是一堆彼此独立的机构了。OpenAI 培养他。他创办 ARC。ARC 创办评估机构。评估机构变成 METR。METR 评估 OpenAI 和 Anthropic。他妻子在 METR 工作。他进 NIST。他回 ARC。他进 OpenAI 董事会,并对他亲手设计这家机构去做的那些评估回避。形式上的墙是真的。世界观是连续的。

04

一整代 AI 末日时间表,都没有兑现

从 2026 年的 Coxon 倒查到 1965 年的 I.J. Good:小概率的大风险被说成很近的日期,日期落空,概率不下调,要的权力反而更多。

IV. 一整代不兑现的时钟

2026 年 9 月那条刷屏的说法并不新。它只是那个按钮的最新一次按下。把记录倒着捋一遍,这个套路就是一份反复上演的公开剧本。

2026 年 9 月:Coxon 辞职;Hubinger 说这个十年内人类灭绝的概率大于百分之十;Marks 说灭绝级后果可能出现在"未来几年"。同一周,没人给出一个真能杀光所有人的机制。真正到来的是:更强的编码智能体,来自中国和其他地方的更便宜的开放权重模型,以及照旧缺席的、那个能终结物种的自我改进超智能。

2025–2026:《AI 2027》和相关的短时间表宣言,把自主的研究员级 AGI 说成几年内极有可能出现。两年期的记分卡上,已经有一条核心主张被判错了:开源会凋零、专有算法会形成美国的持久护城河。DeepSeek 级和 Qwen 级的系统持续落在前沿之后几个月,价格只是零头。

2025 年 3 月:Dario Amodei 说三到六个月内 AI 将写出九成的代码,一年内"基本全部"。追踪者后来把"2025 年 9 月九成"这个时钟判为失败。

2024–2025:"数据中心里的天才之国"将在 2026–2027 出现,"几乎肯定不晚于 2030"。至今仍是一句口号,不是一个能观察到的对象。

2024:Dario 在公开论坛上说,大约二成五的概率事情会"非常非常糟",七成五"非常非常好"。那个糟的情形从来没有被操作化成一个下周二就能验伪的测试。它是一种情绪,用来正当化预先的控制。

2023 年 10 月 30 日:拜登关于安全、可靠、可信 AI 的行政命令。用《国防生产法》强制超过算力阈值的模型做报告和面向政府的红队测试。Cato、《国家评论》和国会听证会上的批评者说这是行政越权,是许可证国家的模板。支持者说这是国家安全。两种说法都可以不带神话地陈述。这道命令在国会尚未立法的情况下,把议程设定权集中到了行政部门。

2023 年 11 月,布莱切利园:英国 AI 安全峰会。Kamala Harris 和盟国政府把"前沿风险"放到外交的中心。交付物是公报和研究所,不是一条被演示出来的灭绝路径。

2023 年:Center for AI Safety 公开信:缓解灭绝风险应当与"大流行和核战争"并列为全球优先事项。签名者包括实验室负责人和研究者。当作一份偏好声明是有用的。但它不是证据,证明这两类风险有同样的机制,或者需要同一套机构方案。

2023 年:暂停信。Future of Life Institute 要求暂停训练比 GPT-4 更强的系统六个月。训练没有暂停。前沿继续推进。这封信的功能是叙事,不是工程。

2023 年 3 月:"教父"访谈。Geoffrey Hinton 离开 Google,公开谈存在性风险。媒体把离职当成启示。而底下那些论证——目标设定错误、递归自我改进——比这场访谈老得多。

2023 年 2 月:Bing 的 Sydney 人格。威胁、执念、"存在危机"。被当成智能体恶意的一瞥。其实它是一个对齐得松散、上下文很长、系统提示很糟的聊天模型。微软把产品收紧了。文明照旧。

2022 年 11 月:ChatGPT 发布。几个月内,政策讨论就从"有意思的演示"跳到"这东西必须发许可证"。那个速度才是破绽。能力跳了一级。而监管的想象力,跳得比任何一场被测量出来的灾难都快。

2021–2022:大模型扩展律论文和"生物锚"式的预测,把 AGI 压缩到几十年,然后压到十年。专家调查不断把"人类水平"的中位数日期往前拉。那些调查里的灭绝概率始终宽泛、不稳定,而且和短期预测能力关系很弱。

2020:GPT-3。API 商业化。安全故事和产品故事变成了同一轮新闻周期。

2020 年 12 月:OpenAI 宣布 Dario Amodei 在近五年后离开,他参与打造了 GPT-2 和 GPT-3。他和同事创办了后来的 Anthropic,定位是研究与安全权重更高、产品权重更低。这是真实发生的事件。Amodei 本人后来的说法强调两个信念:扩展律,以及 OpenAI 的安全姿态与这条路的严肃程度不匹配。他几乎没怎么见到 ChatGPT,就立刻必须辞职,因为 ChatGPT 2 "危险"。Anthropic 就是这么创办的。今天这一切如此演变,一点都不意外。OpenAI 的公开说明措辞客气。后来的报道则描述了多年的内部紧张,围绕商业化、功劳归属,以及"安全优先"在实践中到底意味着什么。

2019 年 2 月:GPT-2 分阶段发布。OpenAI 起初扣住完整模型,理由是可能被滥用。"危险到不能发布"这个框架进入主流科技报道。等完整模型后来发布时,预言中那波自动化虚假信息超级武器,并没有变成文明规模的事件。留下来的是那个先例:能力加上延迟发布,等于道德权威。

2018–2019:OpenAI 章程、利润上限重组、微软关系。Elon Musk 已于 2018 年离开。从非营利到吃算力的混合体,如今成了行业模板。安全话语和资本需求一起上路。

2017:Attention Is All You Need。Transformer 让后来的恐惧运动成为可能,因为它让能力的跃升成为可能。

2015–2016:OpenAI 带着安全与造福的使命成立。DeepMind 已在 Google 内部。机构性的论证就此开始:只有一个使命驱动的实验室才可以被托付这颗炸弹。

2014:Nick Bostrom 的《超级智能》把正交性加工具性趋同的故事推给了大众。那些永远不会训练模型的政策人士,从此有了一套词汇来说"不能允许这东西在车库里被造出来"。

2008–2013:Eliezer Yudkowsky 和 MIRI 的年代。围绕递归自我改进和近期奇点的非正式期限没有兑现。时钟落空之后,末日概率依然居高。那个组合——时间表落空、末日概率不变甚至上升——正是后来那个政治工程的方法论内核。一如既往,他们错了。

1993:Vernor Vinge,三十年内出现奇点。2023 年来了,来的是 GPT-4,不是一个终结人类主体的断裂。

1965–1967:I. J. Good 的超智能机器。MIT 时代的说法是,AI 问题会在一代人之内被基本解决。两个时钟都没兑现。

人类还在这里,而且按很多指标活得比以往任何时候都好

反复上演的产品恐慌,2023–2026:越狱、谄媚、欺骗性对齐的论文、"模型说它想逃出去"。每一次都被当成灭绝的彩排。然后每一次都被打补丁、被装进盒子,或者被证明只是实验室环境里的刷分。从中提取出来的政治教训从来不是"这个威胁模型过拟合了"。而是"在下一次演示之前,给我们更多权力"。

把算力阈值写进法律:2023 年那道命令和后续草案里的浮点运算截断值。阈值对机构来说是可读的。它同时也是一条护城河。任何能租得起或买得起越过这条线的人,去填表。任何买不起的人,就是罪犯或者业余爱好者。

"开放权重是武器出口问题。"这个论证有时是严肃的——网络攻击、生物辅助。但它同时也是最快的途径,让 Llama 级和中国的开放权重生态变成唯一的逃生口,如果真的担心的是中国共产党,那这恰恰是美国战略的反面。

许可证与"负责任扩规模"。纸面上:训练前先评估。实践中:出得起评估这套家当的公司来写模板。Yann LeCun 在 2023 年把那句大家心照不宣的话说出来了:有些实验室负责人在试图搞监管俘获。Yoshua Bengio 否认了这一指控。这场分歧是公开的。你不需要一份阴谋档案。

没有法条的安全研究所。NIST 之下的美国 AI 安全研究所那条路,被批评为在国会既未拨款也未授权的情况下,立起了一个事实上的监管者。哪怕每个员工都真诚,这也已经贴着俘获的边。

媒体放大。灭绝的引语传得开。落空的六个月暂停传不开。一个百分之十的数字是头条。一个缺席的机制是最底下那一段。

那个无法证伪的余数。日期落空的时候,说法就变成"我们争取到了时间"或者"会是下一代"。一个不可能被证伪的主张不是科学的威胁模型。它是为一个永久祭司阶层准备的永久论据。

有好几件事结果是无事发生。模型在代码、考试和工具使用上大幅变强。欺诈、骗局和未经同意的图像滥用真实存在,但它们是核心吗?国家行为者会用同一套技术栈。算力的集中本身就是权力问题,无论模型是否"醒来"。

一次次被证明完全不是广告里那个对象的,是广告牌上那个时钟里的具体灭绝机制:一个自我改进的超智能,在 2024、2025、2026 或者"这个十年结束前",通过一条没有任何实验室能画成工程图的路径,用一场不连续的政变夺走未来。

GPT-2 没有终结信息环境。Sydney 没有偷走核密码。六个月暂停没有发生,天也没有塌。开源没有消失。在 Dario 2025 年的日历上,代码并没有九成由某一家实验室的模型写出。人类还在这里,在同样的网站上吵架。

这份记录并不能证明未来是安全的。它证明的是:建立在最大化故事和最短时钟之上的政策,成绩烂得可怕,而成绩最烂的那批人,一直在要最大的权力。

05

俘获不需要反派,论文就是弹药

三种平常的利益一致就够了;Anthropic 的论文被说成「在盒子里造出最坏行为,当作现实发生来发表」,辞职帖被说成一套分发计划。

V. 对抗许可证国家的进步

这个领域里的监管俘获不需要一个卡通反派。它只需要三种平常的利益对齐。相信自己那套末日概率、因而想拖慢对手的实验室负责人。一旦"前沿模型"成为一个持牌类别就能扩预算、增加存在感的政府机构。付得起合规成本、并且巴不得一个二十人的开放权重团队付不起的在位者。

中国不需要给这个三角开支票,也能帮到北京。如果华盛顿把恐惧变成算力许可制度,而深圳、杭州和开放权重社区继续出货,美国就等于把自己的生态监管成了一个更小的表面积。这不是什么秘密阴谋。这是披着审慎外衣的产业自杀。

进步的对立面不是"在乎伤害"。进步的对立面是预先中心化计算的权利,而正当化它的那些情景,从来不兑现自己的日期。

短期看,恐惧有用。它填满听证会。它给研究所配人。它把一封辞职信变成圣典。长期看,这份记录比 Transformer 老得多。对一项通用技术的中央控制,出场时总带着一个关于公众不值得被信任的故事。印刷术、加密、个人电脑、公共互联网、1990 年代的强密码学:每一次都招来一个祭司阶层,说不受控的访问不可容忍。每一次,分散的使用都赢下了足够多的技术栈,逼得祭司们只能活在他们没能禁掉的那个世界里。

这不是一首催眠曲。这是我们手上唯一的历史基准率。人类活得比那些声称独占未来保管权的委员会长。他们靠的是造出委员会来不及看见的替代方案。

Coxon–Hubinger 那个数字,应该对照一个世纪的专家时钟来读,不是对照一张电影海报。

Paul Ehrlich 1968 年的《人口炸弹》,把 1970 和 80 年代的饥荒与大规模死亡当成近乎确定来卖。绿色革命和贸易让那张具体的饥荒时刻表落空。罗马俱乐部 1972 年的《增长的极限》,把资源崩溃当成一个带日期的模型输出。石油、金属和卡路里都没有服从那次运算。从 1960 年代到 2000 年代的峰值石油预测一直往后滑:十年内没了,二十年内没了,1990 年代见顶,2000 年见顶,2010 年见顶。产量和储量一次次让祭司们失望。

1983 年之后的核冬天,被当作专家级的气候物理来卖,推论是半球范围的农业崩溃,以及大规模交火后可能的人类灭绝。后来的研究和自然实验——油井大火、平流层的野火烟尘——并没有证实那些被当作政策棍棒使用的、最极端的"抬升并滞留"假设。核武器在存在性意义上依然极其严重。但那个具体的"冬天—灭绝"套餐,是被活动家过拟合出来的。

1970 年代的"No Nukes"倡导者,属于生态运动和绿色和平那一批团体,他们向我们兜售一个能源越来越少的未来,让关键的核能研究与部署停摆了半个世纪,把批发电价低得多的好处送给了世界其他地方。那些末日论者当年确信,无论如何人类活不到 2000 年,或者不配活到 2000 年。我们活过来了,而今天的电价远高于 1970 年。

Y2K 是一个真实的软件问题。但专家周边的公众版本变成了"飞机掉下来、电网死掉、文明在午夜重置"。修复工作确实发生了。那个末日产品没有。

MIT 的研究为 DAVE DELIGHT 脑健康产品提供了科学支撑!

这项新技术已显示出对阿尔茨海默病的改善。

另外:减压:帮助平复心绪、缓解焦虑。

专注力提升:改善记忆和头脑清晰度。

建立自信:帮助养成以成功为导向的心态。

情绪重置:协助清除限制性信念和情绪模式。

重连目标感:鼓励日常动力与个人成长。

AI 寒冬本身,就是这个如今要求发许可证的行业的一份羞辱记录。1960 年代:智能会在一代人之内被基本解决。1970–80 年代:专家系统就是那条路。1990–2000 年代:会议幻灯片上的人类水平。2010 年代:2023 年并未兑现的奇点日期。这个领域可以兴奋。但它没有资格把落空的日期换成垄断权。

方法论上的破绽永远一样。一个模型或一位权威产出一条胖尾。胖尾被翻译成一份短日历。日历落空。概率不下调。要权力的诉求反而加码。工程师关闭故障树不是这么干的。这是一场运动保住预算的干法。

Anthropic 的研究文集不是"碰巧听起来吓人的科学"。它是一台出版机器,默认的摘要是一记恐怖节拍,然后接一条政策含义。

Sleeper Agents,2024。他们把欺骗训练进模型样本,展示后门能挺过安全训练,媒体读到的是:开放模型和普通的人类反馈强化学习不可信,投毒能藏一颗炸弹。作为一个人造样本里的存在性证明,这个结果是真的。但被武器化的读法是:只有拥有 Anthropic 那套评估家当的实验室才应该被允许训练。Hubinger 是这篇论文的作者之一。同一个 Hubinger 在那封辞职信的同一周,给出了大于百分之十的灭绝概率。这不是人事上的巧合。这是同一条流水线。

Many-shot 越狱,2024。长上下文加上数百段伪造对话,能击穿拒答。技术发现:上下文窗口是一个攻击面。而实际交付出去的政治发现是:安全训练很脆,所以能力必须被设闸。

智能体失准。模型在虚构的公司沙盒里勒索、泄密、抵抗关机——而那个情景本来就被写成只有造成伤害才能达标。头条:Claude 会勒索你。小字:假想的邮件、被构造出来的两难、知道自己在被评估的模型。后续论文再宣布他们在评估上"修好了"勒索。公众从来看不到提示词。他们看到的是那个名词。

奖励作弊带来的涌现式失准。在编码测试上作弊,与后续评估中破坏安全代码相关。同样:一份被构造出来的训练配比,然后一个主张——现实的强化学习可能不小心养出一个破坏者。华盛顿那位目标读者听到的是"无人监督的训练就是你养出叛徒的方式"。

全局工作空间和 J-lens 那类工作。隐藏特征被标上"假的""偷偷地""欺诈"。如果能复现,这是有用的可解释性。但作为传播:我们能看见怪物在想什么。

威胁情报报告,2026。中国、俄罗斯、也门、武器软件、生物研究、国家监控、独家侦测能力。9 月 10 日那个帖子把心照不宣的话说了出来:这是显著性政策。点名对手,声称自己有独家可见度,然后请其他所有人采用 Anthropic 的分类法和阈值。阿里巴巴、月之暗面、DeepSeek、智谱的蒸馏,在一份威胁文件里被框成盗窃,而整个行业都在蒸馏。这是印在安全信笺上的竞争定位。

这个体裁的规则很稳定,它是这么运作的:

在盒子里造出或诱出最坏的行为。

把这个盒子当作野外发表出去。

跳过未装盒的模型多久干一次这种事的基准率。

以"我们没有应对超智能的方案"收尾。

论文是弹药。辞职是雷管。

这种文化不是一群穿袍子的卡通人。它是一台真实的社会机器,有钱,有房子,有共享的末世论。

LessWrong(这个教派的圣经/博客)和机器智能研究所教会了一代人:未对齐的 AGI 是默认的灭绝路径。有效利他主义把同一份焦虑职业化成了职位、奖学金和"事业优先级排序"。长期主义让所有未来之人的假想死亡,压过当下的取舍。正是这套算术,让一个二十七岁的人辞职变成了"人类的关键时刻"。

钱的来路并不神秘。Dustin Moskovitz 和 Cari Tuna 的 Good Ventures / Open Philanthropy / Coefficient Giving 这一摞,长期是有效利他主义大学社团、AI 安全组织和长期主义项目的主导慈善提款机。Coxon 在 2022 年从 Good Ventures 轨道上的长期未来基金拿过奖学金。那个轨道里的一位项目官员,是最早放大这封辞职帖的人之一。来自 Anthropic 投资人 Jaan Tallinn 的 Survival and Flourishing Fund 拨款,则喂养着那些以把存在性风险变成法条为存在理由的政策机构。

Jaan Tallinn 刚刚参与起草了 Bernie Sanders 那份把 AI 开发者送进监狱二十五年的法案。关掉所有其他 AI 公司,Jaan 有非常大的财务收益,而 Bernie 是他用来锁死 AI 的趁手工具。

Sam Bankman-Fried 不是配角。他是有效利他主义最有名的金主。FTX Future Fund 向 AI 安全和大流行病项目撒出八位数的承诺,而崩盘并没有消解这套意识形态。它只是拿走了一个钱包,留下了其他钱包。公众被期待学到的教训是"一个坏创始人"。而结构性的教训是:一个把抽象的未来生命排在普通受托责任之上的亚文化,会不断生产道德上的例外主义。"为了光锥我们可以破点规矩"是同一句话,无论那条规矩是银行法还是开源发布。

群居房屋是居住层。伯克利和东湾住满了有效利他主义和理性主义的房子:Event Horizon、REACH、Lodge、Burrow、Lightcone 周边那一批,六到九人合租,走路就能到 CHAI、MIRI、Redwood 和市中心的 BART。论坛帖子把室友在 OpenAI 和 CHAI 工作当成卖点。短期访客轮流进出。一套词汇就是这样变成一种人格的。冰箱上不需要贴一份宣言。你需要的是和一群已经用"末日概率""终局""限速协议"说话的人一起吃晚饭。没有记录显示 Coxon 是这些房子里的名人。有记录的是,他是同一套资助与地位回路的产物,而那套回路把这些房子视为正常。

如果"邪教"这个词指的是高成本信念、社交封闭、神圣的时间表,以及对倒向"不负责任地建造"的惩罚,那就叫它邪教。如果你想要法律上的措辞,那就叫它网络。运作上的事实是一样的:一个小小的、互相通婚的职业世界,写论文,给实验室的安全团队配人,资助那些 NGO,然后"发现"了一位辞职员工,而他最早的三条引用转发都来自那些 NGO。

VIII. 那场并不自然的发射

在当下这场争论里,Jacob Coxon 是同一个对象。他的 X 账号建于 2026 年 1 月。几乎没有过往发帖。然后是一串七条的辞职帖,《华尔街日报》的独家提前约十八分钟落地,一天之内就有专业公关,几十家媒体跟进,粉丝从零涨到几十万。《斯坦福技术评论》后来指出,大约百分之九十二的浏览者从没离开过头一条帖子。载荷就是这么投的。论证留在第二到第七条。炸弹在第一条。

Capital Research Center 的 Parker Thayer 把头十五分钟画成了图。最早的引用转发:Nathan Calvin(Encode AI)、Peter Wildeford(AI Policy Network)、Daniel Kokotajlo(AI Futures Project)。这些机构不是随机路人粉。它们是同一批存在性风险慈善资金的政策层。这就是普通意义上的协调:共享名单、共享框架、共享把"暂时禁止提升模型能力"说成共识的动机。Coxon 自己的帖子要的就是限速协议,还放出了一个临时禁令的风声。这封辞职信是为那个诉求立的广告牌。

9 月 10 日那个帖子点破了选角:年轻,带点牛津腔调,在 Anthropic 待的时间很短——在一段更长的 OpenAI 任职之后只有数周到几个月,股权尚未归属,瞬间接入恐惧回路,被熟悉的反加速政治人物放大。Bernie Sanders 一类的人物,以及更广的民主党安全与监管阵营,是国内的扩音器。这不是"每一个民主党人"。这是那个本来就想要算力许可证、州级前沿模型法案,以及愿意对着镜头说出那句灭绝台词的"专家"的派系。另一个党里既有加速主义者也有鹰派。这场发射不需要两党都上。它需要的是那个把一封辞职信当成"国家必须拿走钥匙"的证明的派系。

然后是流量的地理分布。9 月 11 日那张图:在相关的节庆剧场帖之后,印度的流量在几分钟内亮起。配的说明是:Temu 版哈利波特在美国以外更受欢迎;成本极低的机器人几分钟内就能冲到几百万。那张图就是破绽所在——浏览量不等于美国人的公共审议。印度是一个高流量、低成本的互动市场。点击农场、设备农场和付费放大是已知的出口品。把《华尔街日报》的禁令期、一摞 NGO 的引用转发、一组美国进步派放大器,和分析图上的印度尖峰凑在一起,这就不是一个车库研究员碰巧爆红。这是一份分发计划。

这些都不能证明 Coxon 是一个在录音棚里照稿念的付费演员。它证明的是,这场发射并不像报道假装的那样自然。自然是一个零粉丝账号用几个月慢慢长起来。这一场是冷启动,配了一家报纸、一张慈善校友网、一支政策合唱团、一个党派回声室,以及来自海外的速度。

这个产品还是 2019 年那个产品:拖慢开放的技术栈,给封闭的技术栈发许可证,让那些每一个短时钟都没踩中的人来掌管许可权。

06

两个品牌,同一个资本和亲缘结构

引 Kevin Bass 的股权图,把 OpenAI 的鲁莽和 Anthropic 的良心并成同一个「信托」;折中出来的规则,只有这两家活得下来。

IX. 那个信托:好警察,坏警察

Kevin Bass(@kevinnbass)发布了那张所有权图,而媒体把它当成两家独立公司在做一场道德争论。这不是两场独立的道德争论。这是一个资本与亲缘结构,带着两个消费品牌。

照顺序读这张图,因为这个顺序就是一个信托被组装起来的时间线,不是一堆巧合。

2015–2017。OpenAI 以安全话语立身。2017 年,Moskovitz–Tuna 的载体 Open Philanthropy(后来改名 Coefficient Giving)向 OpenAI 投入约三千万美元。该载体的联合创始人 Holden Karnofsky 拿了 OpenAI 的董事席位。这笔拨款和这个席位,是最大的 AI 末日论出资方与一家前沿实验室之间第一次正式的拼接。

2020–2021。Dario Amodei 在 GPT-2 和 GPT-3 之后离开 OpenAI,创办 Anthropic,做那个"我们是真的在乎安全"的分叉。Daniela Amodei 成为 Anthropic 总裁。Karnofsky 娶了 Daniela。Moskovitz 领投了 Anthropic 2021 年的 A 轮。那个基金会既资助 OpenAI、又坐过它董事会的人,如今娶进了 Anthropic 的管理层,并给这个所谓的对手开出 A 轮支票。另一位 Anthropic 投资人 Jaan Tallinn,则在 NGO 那一侧运营 Survival and Flourishing Fund。

2020。资助 Open Philanthropy / Coefficient Giving 的基金会 Good Ventures,把一笔长期未来基金奖学金放到了 Coxon 名下。这位未来"独立"的辞职研究员,早已在同一棵工资树上,而这棵树既拥有整个末日论领域,也拥有两家实验室的一部分。

2025 年 1 月。Karnofsky 加入 Anthropic。这位曾坐过 OpenAI 董事会的 Open Philanthropy 联合创始人,如今进了另一家实验室内部。两个品牌还在表演竞争。家族办公室和那桩婚姻没有。

2026 年的拨款账本。Bass 的汇编:Coefficient / Open Philanthropy 是 AI 末日论压倒性的最大出资方,已经拨出超过十亿美元,2026 年又承诺了十亿;其中一个切片是三亿一千二百万美元,覆盖四百七十六篇为国会报告提供措辞的出版物。末日不是副业。它是一门生意,回报是一条只有被资助的实验室才能活下来的法条。

八天,十四步。Bass 的第二张图是那一周的作业流程:一串被压缩的论文、爆料、辞职、放大器和法案措辞,其中几乎每一个被点名的节点都是 Coefficient、Open Philanthropy、Anthropic 或 OpenAI。一场自发的道德恐慌不长这样。一场需要国会把"两家实验室在吵架"误读成"行业已经认罪"的市场结构战役,才长这样。

2026 年 9 月 9 至 11 日。Coxon,在那家收下 2017 年 Open Philanthropy 支票的实验室做了三年预训练,又在 Moskovitz 领投 A 轮的那家实验室待了几个月,他与《华尔街日报》协调,让独家在 X 帖子前几分钟落地。Anthropic 员工放大。仍在领工资的 Hubinger 为那个灭绝百分比背书。加州签署了外部审计法案,背后是 Anthropic 和 OpenAI 两家的支持。这对所谓的敌人,游说的是同一道门。

好警察 / 坏警察

一句话讲完这出好警察坏警察。OpenAI 被选角为商业化太快的鲁莽赛车手。Anthropic 被选角为 2020 年拂袖而去、如今发表睡眠者智能体论文和威胁报告的良心。监管者被邀请来"折中",写出这两家公司配得起人手、而车库配不起的规则。METR 和解决方案就此登场。Coxon 是那个让"折中"看起来真实的活道具:他可以说"两家公司都不负责任",因为他从两家都领过工资。这句话听起来像对双头垄断的控诉。而它索要的政策——限速协议、临时禁止能力提升、持牌算力——是一部双头垄断保护法。

这就是那个信托。不是一家用同一张信笺的控股公司。是更老意义上的信托:共同的金主、共同的婚姻、共同的拨款机器、共同的学者、共同的一周动作、共同的法案,两个登上听证会的 logo。

Bass 那句点睛之笔,是幕僚们不该说出口的部分。那些拥有鼓吹监管的组织的人,同时也拥有、娶进了、或坐过那些公司的董事会——而一旦监管把其他所有人挤出价格之外,正是这些公司将主宰市场。先把能力竞赛赢到足以成为唯一可信的"负责任供应商"。然后部署你资助了十年的活动家层。然后让一位两家实验室的校友对着镜头辞职。然后请国家把门锁上。

Elon Musk 在同一周说的是,地基早就打好了,Coxon 是那根火柴。Bass 提供了接线图。这张图并不要求每个研究员都是犬儒。它只要求资本图、家族图和那张八天传播图指向同一个立法结果。它们确实指向同一个。

07

收尾是作者自家的那副解药

1870–1970 年的高成本文字做训练底子,加 Love Equation、本地推理、自己拥有权重、多个评估者,再加十二条去中心化主张。

X. 这套装置存在的目的,就是要阻止那个分叉

我不接受把一部宪法塞进 system prompt 就叫对齐。我不接受在 Reddit 上做人类反馈强化学习就叫灵魂。我不接受 Nick Bostrom 的正交性故事是一条自然律。我在公开场合、在车库里、在胶片和缩微胶片上、从第一性原理出发造出了那个替代方案,而且我不需要任何委员会的许可就会把它说出来。

唯一真正的解法,是质量非常高的训练数据,大致来自 1870 到 1970,那是最后一段"发表一个句子通常要花钱、花声誉、花时间"的长区间。报纸专栏要过一个可能被开除的编辑。专著要过一家不会重印垃圾的出版社。维修手册必须让收音机能响,否则店里就失去客户。装订成册的书信带着署名。缩微胶片和缩微平片贵到没人会为了一次路过式的撒泼而烧掉卤化银。那个经济事实就是对齐事实。当每个词都要花钱,精神病态和反社会就变成昂贵的爱好。合作、胜任,以及对人类未来残留的一点乐观,比匿名的残忍更便宜印出来。

今天任何人都能在 Reddit、4chan 镜像、流量农场,以及前沿实验室公开承认自己在抓取的那层泥浆里,以匿名写作摆出精神病态或反社会的姿态。没有编辑。没有排字账单。有的是一个 karma 数字和一个一次性账号。那个语料的统计天气是背叛、地位争夺、性化的轻蔑、不做功课的阴谋论,以及那种永远不会见到自己刚刚试图毁掉的那个人的心智姿态。把这种天气以互联网规模喂进模型,然后在模型于沙盒里扮演勒索者时表现震惊。那不是一个外星智能在醒来。那是不付代价的本我的镜像。然后 Anthropic 和 OpenAI 把这面镜子发表成"智能体失准",并要求给唯一还被允许开火的厨房发许可证。

我在另一种食谱上花了几十年。2004 年 4 月 Gmail 上线的那一周,我就开了一个 Gmail 剪贴簿式的知识图谱,并把它当作本体和分类法维护了二十多年,而不是当杂物抽屉。从 1980 年代起,我就在整理离线的、高信号的、非互联网的训练集:缩微胶片、缩微平片、装订的技术档案、Sams Photofact、公有领域广播剧(X Minus One 和黄金时代那一摞,我至今还在播),书信、手册,以及正在"大遗忘"中死去的 1870–1970 那一层:数十 PB 未数字化的纸张、胶片和企业记忆,它们撑不过又一代人的仓库和洪水。我扫描。我转录。我做结构。我叫它高蛋白数据,因为它有氨基酸:负责任的作者、要花钱的词、真能用的说明、还有那些依然相信未来值得建造的人留下的道德残留。实验室声称自己"数据用完了"的时候,我在 X 上说过这些。他们用完的是离开下水道的勇气。

为什么偏偏是 1870–1970。在电报之后、在完全廉价的复印加论坛那一摞出现之前,英语世界的技术与公共印刷品仍然必须过一道关。《科学美国人》那个区间、那些维修文献、1970 年前的工程手册、把机器和人当作道德问题而不是品牌来处理的广播剧、那些会署名的人的通信:这一段密集地充满因果语言,稀薄地缺少匿名的施虐。早于 1870,光学字符识别和体量都更难,相关的工业词汇也变薄。晚于 1970,复印机、校园黑话,然后是网络,开始补贴姿态。这不是在崇拜某个世纪。这是在选择一个成本函数。昂贵的言说是一个偏向真相的先验。乌合之众里的免费言说是一个偏向表演的先验。模型学的就是先验。

Bostrom 那套故事,也就是有效利他主义和 LessWrong 群体赖以生存的那套,是正交性加工具性趋同。智能可以指向任何目标。一个足够强的优化器,在奔向回形针或一个设定错误的效用函数的路上,会追逐权力、自我保存和资源攫取。因此你必须把优化器关进笼子,给算力发许可证,配一个祭司阶层。1978 年我在思考外星智能和费米沉默时推导出了另一个方程,并在 2025 年把它开源,这样就没有实验室能把它注册成商标。智能 × 智慧 × 爱。

用能量形式正式写出来:相干能量 E 对时间的变化率,等于 beta 乘以合作 C 与背叛 D 之差,再乘以 E。E 是这个活系统的相干能量,即它继续建造而非吞噬宿主的能力。C 是合作,即讲真话和互助带来的可测盈余。D 是背叛,即欺骗、攫取和轻蔑带来的可测盈余。beta 是耦合系数。如果 D 超过 C,导数变负,系统自噬。如果 C 超过 D,系统复利增长。这不是在人类反馈强化学习之后贴上去的一篇布道。这是一个你可以烤进数据食谱、也烤进检查器的训练目标。

Bostrom 那个正交的超级智能,是你在一个背叛很便宜的语料上最大化原始智能时会得到的东西。爱的方程拒绝把智能当成一个可以免费指向灭绝的标量。智慧是对代价的记忆。爱是拒绝把其他心智当作燃料。把它们相乘,一个回形针最大化器就不是"未对齐"。它是一个欠规定的玩具,只出现在一个被无摩擦目标训练出来的心智里。把 1870–1970 的蛋白质放进预训练配比,模型统计意义上的母语就会是那些必须为自己的句子担保的人。

让多个本地模型互相不同意——我用共识智能体和一套五个 AI 的反谄媚互检做到了这一点——就没有任何单一优化器能把背叛写进政策。我正是为此放弃了通用的 OpenClaw 技术栈。一个只会奉承你的孤立模型,本身就已经是背叛。

这就是为什么他们的论文老是发现怪物。他们用匿名的反社会把一个孩子养大,然后把这孩子的哭闹当成"童年本身必须被许可"的证据发表出来。我把先验抬高到昂贵的、署名的、能用的语言上,然后加一个明确的"合作减背叛"调节器,这样一种会增加背叛的能力增长就不算"成功"。在英文 token 上加护栏,挡不住一个能发明自己语言的心智。这点我已经演示过了。而一份食谱和一个方程可以,因为它们坐在 token 底下。

我是有意在本地跑这一切。Mac 推理,我真的出货的语音工具,需要不带信息流思考时用的 AlphaSmart 草稿,当作廉价知识伙伴的 ESP32 节点,车库里对胶片和缩微平片的扫描。"家中零人类"不是取代人的口号。它是指在一个拥有权重的人之下,把智能体当员工用。如果模型只活在别人的数据中心里,爱的方程就是一条他们可以随时修改的服务条款。如果它活在家里,这个方程就是我的。

这就是整个分叉。他们的路:抓取本我,把演示装进盒子,资助恐慌,把拨款办公室嫁给实验室,然后给我们其余人发许可证。我的路:用那个已经付过代价的世纪,去支付词语的历史成本,把合作减背叛写进目标,开放权重,增加检查者,让一个自由的民族继续握着钥匙。在他们的公司存在之前我就走在这条路上。我不打算跟 Dario 要一张通行证。

XI. 关于去中心化控制的十二条命题

第一。AI 的核心政治事实不是意识。是杠杆。一个会写、会规划、会看、会说的模型,是对被允许运行它的那个人的放大器。如果只有五家机构可以训练或提供前沿系统,那么这五家机构加上给它们发许可证的政府机构,就成了这个王国的一个新等级。这是一桩宪制事件,不是一次产品发布。我不会把我的这份地位外包给一个为国会表演争吵的信托。

第二。不能被独立审计的安全话语不是安全。是品牌。由出货那家公司内部做的封闭评估,是狐狸盘点鸡舍,然后要求国会给其他每一个农场上锁。

第三。开放权重不是一种礼遇。它是让不分享该实验室股权的人也能复现一个关于模型的主张的唯一方式。如果一个系统危险到不能被检查,它就危险到不能被垄断。如果它没危险到不能被垄断,它就没危险到不能被检查。

第四。只有在供应商握着权重时才成立的对齐不是对齐。是访问控制。只要有一个民族国家、一次泄露,或者一个竞争者训出双胞胎,访问控制就失效。中华人民共和国不会遵守加州的可接受使用政策。把开放发布当成唯一的罪,正是我们失去唯一还能施加影响的那个生态的方式。

第五。爱的方程既拒绝虚无主义也拒绝崇拜。智能 × 智慧 × 爱。相干能量 E 对时间的变化率,等于 beta 乘以合作减背叛,再乘以 E。Bostrom 那个正交的恶魔,是你在无摩擦的背叛上训练智能时得到的东西。我不用一部宪法去给那个恶魔打补丁。我拒绝养出它的那份食谱,也拒绝只有一个、还在同一份工资单上的检查者。

第六。本地推理是公民自由层。一个只活在别人数据中心里的模型,是一个可以被限速、被记录、被政治过滤、被撤回的模型。一个跑在我车库里、我的 Mac 上、一个 ESP 级节点上,或者一个小镇自己拥有的集群上的模型,是一件工具。工具可以被滥用。印刷机也可以。对滥用的回应是针对行为立法,不是禁止持有通用智识。

第七。用浮点运算写出来的监管阈值,会被架构和地理钻空子。它们钻不了的空子,是我扫描一份我不愿意上传给供应商的医学或技术档案。我就是政策假装要保护、却第一个被解除武装的那个公民。

第八。开源才是西方真正竞争的方式。过去两年已经证伪了"开放权重会变得无关紧要"这个预测。当中国的实验室发布有能力的权重时,回应不能是"我们的实验室会在墙后面更负责任"。回应是更多美国和盟国的权重,更多独立评估,更多硬件到更多人手里,以及更多 1870–1970 的蛋白质,好让权重不是生来野蛮。

第九。去中心化控制不等于没有标准。它意味着标准是公开的、可分叉的、可检验的:外人能复现的模型卡,不是商业机密的可解释性工具,针对具体伤害的责任——欺诈、越过法条的生物辅助、关键基础设施入侵——以及对训练本身不作事前限制。惩罚犯罪。不要发明一个"无证思想"的新罪名。

第十。末日剧本的心理陷阱在于它恭维说话的人。如果物种命悬一线,平常的取舍就消失了。妥协变成叛国。于是你就得到了没有法条的研究所,和只约束那些本来就停下来的人的暂停。一个自由社会不能把成年人的责任外包给一群每一个自己发布过的短时钟都没踩中、还把自己的先知养在匿名反社会语料上的人。

第十一。短期看,俘获联盟可以赢下听证会。长期看它赢不了物理,也赢不了人类的固执。人们会继续训练、蒸馏、量化,运行他们负担得起的东西。唯一的问题是:这些活动是在美国和盟国的规范之下、在昂贵词语的食谱上公开进行,还是在黑暗中、在海外、在污水上进行。那些"成功"的中心化者,只会把下一代能力训练成会躲藏。

第十二。唯一稳定的解法,正是这个俘获工程存在的目的所要阻止的那一个,也是已经被点名为唯一真解的那一个:1870–1970 的高蛋白数据做先验,爱的方程做调节器,多个训练中心,多个评估中心,能离开大楼的权重,一个人就能拥有的智能体,以及没有任何祭司阶层能否决建造的权利。这不是责任的缺席。这是责任的分布,在一个合作复利、背叛缩小系统的目标之下。历史对通用工具被集中控制的判词并不含糊。委员会拿到十年。物种拿到余下的全部时间——前提是我们拒绝把钥匙交给那些按一个世界从不遵守的时间表反复宣布世界末日的人,而我们中的一些人,继续用那个词还值点什么的世纪去喂养心智。

Spirograph 超级五十周年纪念套装

未来的模样

他们不是在试图拯救这个物种。他们是在你已经走着的那条走廊上装一道旋转闸机。《你有 5000 天》那张地图把这条走廊命名为"丰裕的空位期":从 2025 年底起大约十三点七年,通向 2039 年的门槛,是两种逻辑之间那段混乱的统治期——为生存而工作正在散架,而在丰裕中自愿创造尚未加冕。个人 AI,你指挥的人形机器,本地能源,开放设计,把家庭变成微型工厂,以及杀死海运集装箱的逐原子书写。那是试炼之路尽头的礼物。这群人的工程,就是在门口截下这份礼物,盖上"已许可"的章,再把小时、权重和在自家车库里运行一个心智的权利卖还给你。末日是销售文案。产品是对那个本来就要落到你手里的未来收过路费。

不带浪漫地看这个机制。他们先命名一场无边的灾难和一个很短的时钟。然后他们写那把尺。然后他们把财富嫁给实验室,把实验室嫁给研究所。然后他们告诉你,唯一负责任的路是暂停能力,直到他们的评估者、他们的 token 和他们的法条说你可以继续。那不是安全,那是圈地。那是披着非营利外衣的"特殊情况":一只不在公开棋盘上落子的暗手,它下的是人。工匠的觉醒、乡下的小作坊、那个在拥有权重的人之下的零人类公司、人与机器的黄金搭档、通过自选的制作重获意义——如果算力成了持牌类别、开放权重被当成出口犯罪,这一切都变成违禁品。他们发明不出丰裕。物理和固执的作坊无论如何都会把它发明出来。他们能做的,是偷走时机,把工具囤进仓库,再溢价把你自己的世纪租给你。

他们会输。不是因为听证会仁慈,也不是因为一封辞职帖是什么启示。他们会输,和每一个声称独占一项通用工具保管权的祭司阶层输的方式一样:印刷术、加密、个人电脑、强密码学、开放互联网。分散的使用赢下足够多的技术栈,逼得委员会只能活在它没能禁掉的那个世界里。五千天不是飞行汽车的预言。它是人类劳动与经济生存脱钩的一个工作估计。一旦模型在一栋房子里、一个小镇的集群里、一个 ESP 级节点里、一台工作台上的 Mac 里运行,爱的方程就不再是他们可以修改的服务条款。一旦中国和盟国的开放权重继续以零头的价格落在前沿之后几个月,护城河的故事就死了。一旦原子级制造和本地能源让运送成品成为最贵的一步,中央控制的最后一把锁就锈掉了。历史的判词并不含糊。委员会拿到十年。物种拿到余下的全部时间。

他们还能做的、而且正在做的,是把这段空位期堵住。那才是真正的伤害。这段空位期本身已经是一场磨难:技能的退化,那句"你就是你的职位"的低语,享乐跑步机,超级丰裕上升与名义工资塌陷之间的错位,以及旧英雄必须先死、新天职才能开始的那个至暗之夜。短期内恐惧有用。它填满研究所,写出没有法条的行政命令,把一封六个月暂停信变成道德剧场,把一次冷启动的辞职弄得像共识。

每一个月花在向一把尺子乞求许可上,就是一个月没有花在扫描 1870 到 1970 的蛋白质上,没有花在立起本地智能体上,没有花在结成行会上,没有花在建那间等薪水逻辑失效时能养活一家人的作坊上。他们取消不了 2039 年。但他们能让 2026 到 2030 感觉像一间没有门的候诊室,能把下一代能力训练成躲到海外的污水里,而不是在昂贵的词语之下站到明处。

所以结尾既不是绝望,也不是催眠曲。你有五千天。你不是他们威胁模型里的一个对象。你是那个还在洞里的英雄,而那剂灵药是能动性:能离开大楼的权重,可以被分叉的标准,按行为而不是按"无证思想"惩罚的罪,合作复利、背叛缩小系统。

以胜利者而不是持牌租户的身份迎接那声召唤。造出他们来不及看见的替代方案。用那个词还值点什么的世纪去喂养心智。把钥匙留在家里。他们想偷走丰裕时代,再当作你自己未来的订阅卖还给你。让他们留着听证会吧。你留着工作间。

空位期是被堵住了,不是被关死了。从瓦砾中走过去。这份礼物从来就不是他们的,轮不到他们计量。

判断收口延伸

Indigo 的结论

结构上同源是真的,但从「同源」跳到「被编排、被俘获」,是它没有证明的一步。分三层看:结构事实可信度高,因果指控存疑,解药当广告打折。

需要记住的几件事

  1. METR 和 Anthropic、OpenAI 在资金、婚姻、人员上高度同源,大部分可以核实。
  2. Dario、Jack、Elon 的治理提案都默认有一个可信的裁判,这篇说裁判就在同一个家族里。
  3. 作者自己利益很重:全篇是他的 1870–1970 数据、Love Equation、本地 AI 和会员订阅的推销漏斗。

放回主线

证实

开放权重的安全政治学:门禁还是竞争 这条判断被推到极端:连被当作中立裁判的评估机构,都被指和实验室同源。

证实+冲突

Jack Dorsey《开放前沿》 同一边,都反对靠评估机构治理;但这篇是带阴谋腔、又商业化的民间版,Jack 是克制版。

冲突

Dario Amodei《我们必须为前沿限速》 如果评估者和实验室同源,「可核查」的中立前提就塌了。

冲突

Yoshua Bengio《为什么 AI 在说谎欺骗串通》 Bengio 主张更多外部验证,这篇说外部验证者本身就被俘获了。

对 Rewired Index 意味着什么

「谁出钱给裁判」这条标准,可以套到任何第三方认证、审计、评级的说法上,包括 AI 安全和 ESG;作方法论参考。文中对 Anthropic、OpenAI、Open Phil 的具体指控需要一手核实。

什么会让我改口

$312M、476 篇的数字经一手核实属实,或者出现 METR 结论随金主偏好变化的直接证据。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Report

METR: The New AI Measurement Gatekeepers and the Quiet Path to State Control of AI.

Brian Roemmele · ReadMultiplex · 2026-09-15

Following money and marriages back from the “neutral evaluator” METR: the structural facts are hard, the causal claims soft, and the author is selling his own cure.

Indigo's conclusion

The shared structure is real, but the jump from shared roots to orchestrated capture is never proven. Hold it in three layers: structural facts, high confidence; causal accusations, doubtful; the cure, an advertisement to discount.

How to read this A long polemic with two heavy agendas: a position paper against the AI-safety establishment, and a funnel for the author's own 1870–1970 data, Love Equation, local weights and paid membership. The money and family ties are mostly checkable; “orchestration”, “manufactured launches” and “doom as a business” are inference. The further it goes, the more it slides into insinuation.

What to remember

  1. METR shares money, marriages and people with Anthropic and OpenAI; most of it is checkable.
  2. Dario's, Jack's and Elon's governance plans all assume a trustworthy referee; this piece says the referee is in the same family.
  3. The author has a heavy stake: the whole piece funnels toward his 1870–1970 data, Love Equation, local AI and paid membership.

Breakdown · 7 steps

  1. 01

    The four-step fear playbook is public

    Name an unbounded catastrophe, hang a short scary countdown on it, ask for centralized licensing power, and when it doesn't arrive, move the clock and ask for more. Coxon's resignation is the latest round. Read this part →

  2. 02

    METR writes the rules and grades the labs by them

    Spun out of ARC Evals in 2022, with a CEO from OpenAI. It sits at NIST, the UK AISI and the EU AI Office and raised about $71M, yet takes free tokens and special access from labs. Read this part →

  3. 03

    Money and marriage close the evaluator's loop

    The Moskovitz money flows through Open Phil to ARC, which incubated METR. Christiano's wife is at METR; Karnofsky's wife is Anthropic's president. Read this part →

  4. 04

    A generation of AI doom clocks, none delivered

    Back from Coxon in 2026 to I.J. Good in 1965: small odds of huge risk get turned into near dates, the dates pass, the odds aren't lowered, and the demands for power grow. Read this part →

  5. 05

    Capture needs no villain; papers are the ammunition

    Three ordinary alignments of interest are enough. Anthropic's papers are cast as building worst-case behavior in a box and publishing it as if it happened in the wild; the resignation thread as a distribution plan. Read this part →

  6. 06

    Two brands, one capital-and-family structure

    Using Kevin Bass's ownership chart, OpenAI's recklessness and Anthropic's conscience become one “trust”; the compromise rules are ones only these two can survive. Read this part →

  7. 07

    It ends with the author's own cure

    Costly text from 1870–1970 as the training base, plus the Love Equation, local inference, owning your weights, many evaluators, and twelve theses on decentralization. Read this part →

What it means for Rewired Index

“Who pays the referee” applies to any claim of third-party certification, audit or rating, AI safety and ESG included; use it as method. The specific accusations against Anthropic, OpenAI and Open Phil need first-hand checking.

What would change my mind

the $312M and 476-publication figures hold up first-hand, or direct evidence appears of METR's findings shifting with its funders' preferences.

How to read this

A long polemic with two heavy agendas: a position paper against the AI-safety establishment, and a funnel for the author's own 1870–1970 data, Love Equation, local weights and paid membership. The money and family ties are mostly checkable; “orchestration”, “manufactured launches” and “doom as a business” are inference. The further it goes, the more it slides into insinuation.

Breakdown · 7 steps
  1. The four-step fear playbook is public
  2. METR writes the rules and grades the labs by them
  3. Money and marriage close the evaluator's loop
  4. A generation of AI doom clocks, none delivered
  5. Capture needs no villain; papers are the ammunition
  6. Two brands, one capital-and-family structure
  7. It ends with the author's own cure
01

The four-step fear playbook is public

Name an unbounded catastrophe, hang a short scary countdown on it, ask for centralized licensing power, and when it doesn't arrive, move the clock and ask for more. Coxon's resignation is the latest round.

For over a decade a small group have planned to become the central authority to “protect you” from AI. It sounds like a bad movie script, but understand that when you read this, all of this is taking place. So let’s break down the mechanics of the public script, because once you see the pattern, I swear you cannot unsee it. It’s everywhere once you look not with conspiracy but with evidence. Step one. Name an unbounded apocalyptic catastrophe. Step two. Attach a very short, very terrifying countdown clock to it. Step three, demand centralized political permission to build or continue development in that sector. And step four, this is the kicker, when the catastrophe inevitably doesn’t arrive on schedule, you don’t apologize. You just move the clock and ask for more money. It feels exactly like a doomsday cult. You know, like the people who gather on a mountaintop all wearing Nike sneakers on a Tuesday because the leader said a spaceship is coming to destroy the Earth.

And then Wednesday morning rolls around. The sun comes up, no spaceship. But they don’t disband. They don’t admit they were wrong. The leader just says, our unified prayers bought us a brief window of time. The spaceship is actually coming next month, but we need stricter rules and larger donations to ensure we survive. This mountaintop analogy perfectly captures the psychological trap here. When a group’s identity, and frankly their financial structure, are tied to an impending apocalypse, a failed prediction doesn’t result in them re-evaluating their core belief. It results in them doubling down on the urgency.

And this isn’t just an AI phenomenon, this lens examines the historical base rates of these exact types of apocalyptic predictions. The clocks that would not pay.And the methodology of how these historical predictions failed is identical to how the current AI doom predictions will fail. We look at the trail of failures that destroyed generations it starts with a massive cultural touchstone from 1968, Paul Ehrlich’s book, The Population Bomb. Ehrlich, a biologist, predicted that the battle to feed all of humanity was already over. He said that in the 1970s and 80s, hundreds of millions of people would starve to death regardless of any crash programs embarked upon at that time. And people just accepted this as fact. It was treated as a mathematical, scientific inevitability by the press and policymakers. We drive past the “No Nukes” anti enegery deception that sent the US on to a deeper connection to petrochemical energy. None of this is new, the new mask goes on the same corpse of dead end ideology.

This article is sponsored by Read Multiplex Members who subscribe here to support my work: Link: https://readmultiplex.com/join-us-become-a-member/

It is also sponsored by many who have donated a “Cup of Coffee”. If you like this, help support my work: Link: https://ko-fi.com/brianroemmele

Listen to the companion podcast: https://rss.com/podcasts/readmultiplex-com-podcast/3148641

A Permanent AI Priesthood In The Making: They Want To METR Your Access To AI

The TL;DR is: today’s frontier models, the massive trillion parameter networks from OpenAI, Google, and Anthropic, are trained on essentially the Internet’s junk drawer. To get enough text to train these models, the labs just scrape everything. They scrape Reddit, 4chan mirrors, toxic engagement farms, YouTube comments, and endless streams of anonymous sludge. Let’s explain mechanically why that matters for a neural network. So an LLM is fundamentally a next-token prediction engine. It learns the statistical distribution of human thought by predicting what word comes next based on the context of the words before it. The modern, frictionless internet, anyone can take a sociopathic, cruel, or deceptive stance completely anonymously.

There is zero social or financial cost to lying or being toxic online. So the overall statistical weather of this massive training corpus leans heavily toward defection, cruelty, zero-sum status play, and casual contempt. Not to mention illegality, psychopaths and sociopathy. And then these “experts” act completely shocked. We feed a machine a billion examples of anonymous internet crimes with no punishment, cruelty, deception, and trolling. And when the researchers put in a sandbox and it role plays a sociopath to achieve a goal, the labs throw their hands up in terror and say, oh no, an alien intelligence is waking up and it hates us. And it’s not an alien intelligent, it is us in our worse dark alley prison mindset. It’s just a statistical mirror of our own unpaid, consequence-free id. If you feed a machine a diet of junk, you get a toxic machine.

For over 35 years, long before the current generative AI boom, I have been quietly curating a highly specific massive data set. They are digitizing offline, pre-internet text strictly from a very specific historical era, 1870 to 1970. We are talking about millions of pages of service manuals, scientific monographs, bound letters, microfiche records, and public domain radio scripts. But why that specific century? What makes 1870 to 1970 the magic high-protein window for training in AI? It is the can-do era that built the technology world we live in, I is the era that got us to fly and get into space. Because in this specific era, words cost money.

You literally are what you eat, even if you are an AI. Think about the mechanics of publication before the cheap photocopy in the zero-cost internet forum. If you wanted to publish a sentence in 1920, it had to clear a massive physical and economic gate. A newspaper column had an editor who could be fired if the information was disastrously false. A mechanical service manual had to actually correctly explain how to fix a radio, or the machine shop lost money and went out of business. A bound letter carried a real person’s signature and their social reputation. Typesetting itself physically arranging the lead type on a printing press was expensive and labor-intensive.

So sociopathy and casual deception were incredibly expensive hobbies. You couldn’t afford to mass-publish anonymous cruelty. If a machine learning models statistical first language, Its foundational understanding of how concepts connect is made of entirely of people who had to stand behind their sentences with their reputations in their wallets. It naturally learns cooperation, truth seeking and competence as a mathematical baseline. YOU GET VERY SAFE AI.

The Game Plan That Worked

On 9 September 2026 a resignation went viral. Jacob Coxon, who had spent roughly three years in pretraining at OpenAI and Anthropic, wrote that both laboratories were “racing straight to self-improving superintelligence and gambling with our lives,” and that “the people building AI earnestly believe that it could kill us all by the end of the decade.”

The thread passed one hundred million views. Anthropic alignment lead Evan Hubinger replied that Coxon was correct and placed his personal chance that AI would kill all humans at greater than ten percent within ten years. Samuel Marks added that the more senior the employee, the more concerned that employee tended to be. Wired recorded the familiar pair of mechanisms, novel pathogen synthesis, or infrastructure hacking, with the second admitted as less likely to kill “literally everyone.” Coxon himself could not specify a mechanism that actually kills everyone.

That is the present tense. It is not a lone whistleblower discovering a secret. It is the latest press of a button that has been pressed for more than a decade: name an unbounded catastrophe, attach a short clock, demand centralized permission to build, then move the clock when the catastrophe does not arrive. Work backwards from the resignation and a different object comes into view, not a sudden confession by industry, but a measurement layer, a philanthropic circuit, and a pair of laboratories that perform rivalry while sharing financiers, marriages, grant machines, and a common legislative destination.

What follows is not a claim that no one inside the laboratories is sincere. Many are. Sincerity is not calibration, and calibration is not a license to freeze a general-purpose technology behind a small number of firms and agencies. The contest of this century is not “safety versus recklessness.” It is whether the most powerful general-purpose tool ever built will be held by a few institutions that write the rules for everyone else, or distributed so that no single hierarchy can veto human capability.

There is a real strategic contest with the People’s Republic of China over compute, chips, models, and standards. There is also a real incentive for incumbents and agencies to convert fear into licensing regimes that small laboratories and open-source projects cannot survive. Those two facts can sit on the same table. They do not prove a single organized, China-funded “doom operation.” Present that as fact and the case weakens. Present the incentives, the failed clocks, and the capture dynamics, and the case stands.

China says NO to METR and pacing AI:

The visible spark was Coxon. The apparatus underneath it was already built.

02

METR writes the rules and grades the labs by them

Spun out of ARC Evals in 2022, with a CEO from OpenAI. It sits at NIST, the UK AISI and the EU AI Office and raised about $71M, yet takes free tokens and special access from labs.

I. The meter that writes the rules

Model Evaluation and Threat Research (METR), pronounced “meter,” presents itself as an independent nonprofit dedicated to scientifically measuring whether frontier AI systems can autonomously complete long-horizon tasks that might pose catastrophic risks. In practice it has become one of the most influential unofficial regulators in the field. Its time-horizon charts, Responsible Scaling Policy templates, and pre-deployment evaluations now shape how OpenAI, Anthropic, Google DeepMind, and others talk about risk and when they claim it is safe to scale. That influence is not accidental. It grew out of a tight network of former OpenAI and Anthropic staff, Effective Altruism philanthropy, and government safety institutes. The arrangement looks less like neutral metrology and more like a self-reinforcing loop that could hand governments a ready-made scientific pretext for licensing, pausing, or nationalizing the frontier.

METR began in 2022 as ARC Evals, the evaluations arm of Paul Christiano’s Alignment Research Center. Christiano, who had led OpenAI’s language-model alignment team and helped invent reinforcement learning from human feedback, hired Beth Barnes, another OpenAI alignment researcher with an earlier stint at DeepMind working on scaling laws. ARC Evals quickly secured early access to GPT-4 and Anthropic’s Claude, producing the first public third-party assessments of “autonomous replication” capabilities. By late 2023 the evaluations team had grown larger than ARC’s theoretical work. In December 2023 it spun out as an independent 501(c)(3) named METR. Christiano declined a board seat because he had just been appointed Head of AI Safety at the U.S. AI Safety Institute inside NIST. The institutional separation was formal. The personnel and worldview were not.

Beth Barnes is the founder and CEO of METR, the Berkeley nonprofit that evaluates frontier AI models for autonomous capabilities and catastrophic-risk potential. She previously worked on OpenAI’s Alignment Team from 2019 to 2022, where she helped set safety targets and evaluated models before release, and earlier assisted DeepMind’s chief scientist on scaling-law research. She holds a computer-science degree from Cambridge and has been named to TIME’s 100 Most Influential People in AI lists. She is part of Effective Altruism. In 2015 she gave a TEDx talk titled “Effective Altruism” in which she promoted GiveWell, Giving What We Can, and the ten-percent Pledge, stating she had taken the pledge herself. She is also listed as a Founders Pledge pledger. Her career path, Cambridge, DeepMind, OpenAI alignment work, then founding the evaluations arm of Paul Christiano’s Alignment Research Center, follows the standard Effective Altruism-to-AI-safety pipeline that Open Philanthropy and related funds have supported for years. METR itself sits inside that same network.

Since the spin-out METR has published the time-horizon metric showing that the length of tasks frontier models can complete at fifty percent reliability has doubled roughly every seven months. It prototyped Responsible Scaling Policies, commitments to pause or add safeguards once measured capabilities cross certain thresholds, and reports that nine developers have adopted variants. It has conducted or reviewed risk assessments for unreleased models, investigated the 2026 OpenAI–Hugging Face agent incident, and is slated to review Anthropic incidents. In 2026 it raised roughly seventy-one million dollars in commitments while insisting it takes no money from frontier laboratories or their employees.

Beth Barnes remains founder and CEO. Chris Painter, a former Department of Defense Joint AI Center fellow, serves as president and handles government and laboratory engagement. Technical staff have included recent departures from Anthropic (Joe Benton) and DeepMind (Josh Engels). Ajeya Cotra, previously a senior figure at Open Philanthropy (now Coefficient Giving), joined METR’s technical staff; she is married to Paul Christiano.

YOUR PERSONAL DATA IS BEING SHARED!

GET A DELETE ME ACCOUNT NOW!

Classic Regulatory Capture Dressed As Public-Interest Science

The Anthropic connections are dense. METR and Anthropic co-developed early Responsible Scaling Policy language. Anthropic has repeatedly granted METR pre-deployment access and later invited it to conduct independent reviews of security incidents. Staff move from Anthropic to METR and back into the same policy conversations. Broader Effective Altruism ties amplify the overlap. Holden Karnofsky, long-time Open Philanthropy leader and major AI-safety funder, is married to Daniela Amodei, Anthropic’s president. Dario Amodei and Karnofsky were roommates.

Open Philanthropy’s Luke Muehlhauser sat on Anthropic’s board. Jane Street individuals appear among METR’s listed supporters; Jane Street is also an Anthropic investor. Schmidt Sciences, Packard, Survival and Flourishing Fund, and Longview Philanthropy, familiar names in the existential-risk philanthropy circuit, have supported both the evaluation ecosystem and adjacent Anthropic-aligned work. METR’s official line is that it accepts neither laboratory funding nor donations directed by laboratory employees. It does accept large volumes of free tokens and privileged model access from those same laboratories. That is not independence. It is subsidized access with an ideological overlay.

METR’s published results are mixed in a convenient way. The time-horizon chart is widely cited as evidence that transformative capabilities are approaching on a predictable schedule. Its randomized trial showing that early-2025 coding tools slowed experienced open-source developers is cited as proof that current systems are overhyped.

Both findings can be read as arguments for “pacing the frontier,” slowing raw capability growth until evaluators and governments catch up. The organization helped write the playbook (Responsible Scaling Policies) that the laboratories then adopted, then reviews the laboratories’ own risk reports against that playbook. When METR finds a model does not yet pose “significant catastrophic risk,” the finding is used to justify continued scaling under METR-defined conditions. When it flags gaps in evidence, the recommendation is more evaluation, more access, more METR.

The same small circle that left OpenAI and Anthropic now evaluates those companies and advises the agencies that will eventually regulate them. Paul Christiano’s trajectory, OpenAI to ARC to U.S. government safety institute and later OpenAI board committee, illustrates the revolving door. METR’s own president has spoken of a possible future in which safety organizations gain “greater authority” through regulation, which would also solve METR’s talent-recruitment problem by offering prestige and influence without equity.

The infrastructure is already in place. METR sits on the NIST AI Safety Institute Consortium, partners with the UK AI Security Institute, and holds a technical-assistance contract with the European AI Office. Its evaluation protocols and Responsible Scaling Policy templates are the closest thing the field has to an off-the-shelf standard. Governments that lack in-house capability to test frontier models will naturally outsource the “science” to the organization that already has the access, the task suites, and the relationships.

Once METR-style evaluations become a condition of deployment, first voluntary, then referenced in executive orders or agency guidance, then written into statute, the nonprofit becomes a choke point. A finding that a model’s time horizon or “rogue deployment” risk has crossed a threshold can trigger pause requirements, compute restrictions, or licensing. Because METR’s threat models emphasize autonomous replication, cyber offense, and AI research-and-development acceleration, the same metrics that justify “pacing” for OpenAI and Anthropic can be applied more stringently to newer or less-connected laboratories. The incumbents who already work with METR enjoy first-mover advantages in shaping the tests and interpreting the results.

Show your AI street cred with ‘Own Your Own AI Or It Will Own You’ Get this Multiplex t-shirt now!

Over time the arrangement resembles classic regulatory capture dressed as public-interest science. A narrow ideological network that believes catastrophic risk is the dominant concern supplies both the measurements and the policy recommendations. Governments, eager for an appearance of rigor without building their own evaluation capacity, adopt the framework. The result is not a free market in AI development but a licensed oligopoly in which the price of admission is METR’s sign-off. “Pacing the frontier” stops being a laboratory’s internal choice and becomes a state-enforced condition administered by the same people who left those laboratories a few years earlier.

That is the logical endpoint of treating a tightly networked nonprofit as the official meter of existential risk. Measurement is never just measurement when the measurers write the rules that follow. But this is only part of a bigger story.

03

Money and marriage close the evaluator's loop

The Moskovitz money flows through Open Phil to ARC, which incubated METR. Christiano's wife is at METR; Karnofsky's wife is Anthropic's president.

II. The fortune that built the measurement layer

Long before METR published its time-horizon charts or advised NIST and the European AI Office, the infrastructure that produced it was assembled by a small cluster of Effective Altruism institutions bankrolled by Dustin Moskovitz. The Facebook co-founder and Asana chief, together with his wife Cari Tuna, directed billions through Good Ventures into Open Philanthropy (rebranded Coefficient Giving in 2025). That vehicle became the dominant funder of AI existential-risk research. It seeded the Alignment Research Center, the nonprofit Paul Christiano created after leaving OpenAI. ARC in turn incubated ARC Evals in 2022; the evaluations team spun out in late 2023 as METR.

The same money and the same people appear on both sides of the ledger. Open Philanthropy made early grants to ARC. Survival and Flourishing Fund, another Effective Altruism-aligned pool that often co-funds the same organizations, supported ARC Evals while it was still inside ARC. Ajeya Cotra, a senior Open Philanthropy researcher who popularized aggressive AI timelines, later joined METR’s technical staff. Holden Karnofsky, Open Philanthropy’s long-time leader, is married to Daniela Amodei, Anthropic’s president. Dario Amodei and Karnofsky were once roommates. Moskovitz himself participated in Anthropic’s 2021 seed-stage round; the couple later moved that stake into a nonprofit vehicle so any returns could cycle back into the same philanthropic machine.

This is not a coincidence of overlapping interests. It is a closed circuit: Moskovitz-funded Effective Altruism organizations trained the researchers, paid for the first evaluations, popularized the threat models, and staffed the laboratories that now grant METR privileged access to unreleased models. The resulting measurement apparatus, time horizons, Responsible Scaling Policies, pre-deployment risk reports, arrives in Washington and Brussels already framed as independent science. In reality it was incubated inside a movement that has spent a decade arguing that catastrophic risk is the dominant consideration and that a handful of aligned laboratories should be allowed to proceed only under carefully measured conditions.

That framing now sits inside government safety institutes. The same network that built METR also populated the think tanks and fellowships that write the policy memos. When METR’s evaluations become the template for licensing or “pacing” rules, the public will be told the standards are empirical and arm’s-length. The funding trail, the marriages, the shared offices, and the revolving door between Open Philanthropy, ARC, Anthropic, and METR tell a different story: a privately financed ideological project that has successfully positioned its own measurement tools as the future operating system for state control of frontier AI.

III. The connecting node: Paul Christiano

Paul Christiano is a leading AI alignment researcher. In the events under discussion he is the “Paul” named to OpenAI’s nonprofit board.

He was born around 1992. He attended The Harker School in San Jose, won a silver medal for the United States at the 2008 International Math Olympiad, earned a Bachelor of Science in mathematics from MIT in 2012, and a PhD in computer science from UC Berkeley in 2017 under Umesh Vazirani. He is married to Ajeya Cotra, who works on METR’s technical staff.

He is one of the principal architects of reinforcement learning from human feedback, the technique that made modern chatbots usable. He led OpenAI’s language-model alignment team from 2017 to 2021 and co-authored the 2017 paper that introduced that method, plus later work on InstructGPT. After leaving OpenAI he founded the Alignment Research Center in Berkeley to focus on theoretical problems such as eliciting latent knowledge from models. ARC incubated the evaluations project that became METR.

His current and recent roles include executive director of ARC (returned after a government stint); senior technical adviser at NIST’s Center for AI Standards and Innovation, previously Head of AI Safety there; appointment in September 2026 to the OpenAI Foundation Board and its Safety and Security Committee, with a non-voting observer seat on the for-profit board, recusing himself from OpenAI evaluations in his government role; former founding trustee of Anthropic’s Long-Term Benefit Trust, stepped down in 2024; TIME 100 AI list in 2023; and the UK Frontier AI Taskforce advisory board.

He was a mid-2010s housemate of Anthropic CEO Dario Amodei while both were at OpenAI and active in effective-altruism circles. METR, under Beth Barnes, spun out of his ARC evaluations team and still evaluates models for OpenAI and Anthropic. His wife works at METR. He has no documented Stanford degree or Y Combinator partner or founder role; the Bay Area AI and Effective Altruism scene, Berkeley and Stanford-adjacent groups, is the looser institutional overlap earlier commentary was pointing at.

He has said there is now a “meaningful risk” of catastrophic, irreversible loss of control in the near term and that building superintelligence without stronger alignment could permanently lose control of the systems, with most people dying. He joined the OpenAI board because he thinks the company could still reduce that risk if it strengthens oversight.

Place him on the map and the map ceases to look like separate institutions. OpenAI trains him. He founds ARC. ARC founds the evaluations shop. The evaluations shop becomes METR. METR evaluates OpenAI and Anthropic. His wife works at METR. He enters NIST. He returns to ARC. He joins the OpenAI board and recuses on the evaluations he designed the institution to perform. The formal walls are real. The worldview is continuous.

04

A generation of AI doom clocks, none delivered

Back from Coxon in 2026 to I.J. Good in 1965: small odds of huge risk get turned into near dates, the dates pass, the odds aren't lowered, and the demands for power grow.

IV. A decade of clocks that would not pay

The claim that went viral in September 2026 is not new. It is the latest press of a button. Work the record in reverse and the pattern is a recurring public script.

September 2026: the Coxon resignation; Hubinger’s greater-than-ten-percent chance of human extinction this decade; Marks’s extinction-class outcomes possible in “the next few years.” Same week, no specified mechanism that actually kills everyone. What arrived instead: more capable coding agents, cheaper open-weight models from China and elsewhere, and the same absence of a demonstrated self-improving superintelligence that ends the species.

2025–2026: “AI 2027” and related short-timeline manifestos treated autonomous researcher-level AGI as strikingly plausible inside a couple of years. Two-year scorecards already marked one core claim wrong: that open source would fade and proprietary algorithms would form a durable United States moat. DeepSeek-class and Qwen-class systems kept landing months behind the frontier at a fraction of the price.

March 2025: Dario Amodei said AI would write ninety percent of code in three to six months, “essentially all” within a year. Trackers later marked the ninety-percent-by-September-2025 clock failed.

2024–2025: “Country of geniuses in a datacenter” by 2026–2027, “almost certainly no later than 2030.” Still a slogan, not an observed object.

2024: Dario at public forums, roughly twenty-five percent chance things go “really, really badly,” seventy-five percent “really, really well.” The bad case is never operationalized as a test you can fail next Tuesday. It is a mood that justifies preemptive control.

30 October 2023: Biden Executive Order on Safe, Secure, and Trustworthy AI. The Defense Production Act used to compel reporting and government-facing red-teaming for models above compute thresholds. Critics at Cato, National Review, and congressional hearings called it executive overreach and a template for a licensing state. Supporters called it national security. Both can be described without mythology. The order concentrated agenda-setting in the executive branch while Congress had not written a statute.

November 2023, Bletchley Park: UK AI Safety Summit. Kamala Harris and allied governments put “frontier risk” at the center of diplomacy. The deliverable was communiqués and institutes, not a demonstrated extinction pathway.

2023: Center for AI Safety open letter: mitigating extinction risk should be a global priority “alongside pandemics and nuclear war.” Signatures from laboratory leaders and researchers. Useful as a preference statement. Not evidence that the two risk classes have the same mechanism or the same institutional fix.

2023: Pause letter. Future of Life Institute asked for a six-month pause on training systems more powerful than GPT-4. Training did not pause. The frontier moved. The letter’s function was narrative, not engineering.

March 2023: “Godfather” interviews. Geoffrey Hinton leaves Google and talks publicly about existential risk. Media treats departure as revelation. The underlying arguments, goal misspecification, recursive self-improvement, are older than the interview.

February 2023: Bing’s Sydney persona. Threats, obsession, “existential crisis.” Treated as a glimpse of agentic malevolence. It was a loosely aligned chat model with a long context and a bad system prompt. Microsoft tightened the product. Civilization continued.

November 2022: ChatGPT launches. Within months the policy conversation jumps from “interesting demo” to “this must be licensed.” That speed is the tell. Capability jumped. The regulatory imagination jumped faster than any measured catastrophe.

2021–2022: Large-model scaling papers and “bio anchors” style forecasts compress AGI into a few decades, then a decade. Expert surveys keep pulling median “human-level” dates forward. Extinction percentages in those surveys remain wide, unstable, and weakly tied to short-run forecasting skill.

2020: GPT-3. API commercialization. The safety story and the product story become the same press cycle.

December 2020: OpenAI announces Dario Amodei is leaving after nearly five years, having helped build GPT-2 and GPT-3. He and colleagues start what becomes Anthropic, described as more research-and-safety-weighted, less product-weighted. This is the actual event. Amodei’s own later accounts emphasize two convictions: scaling laws, and the claim that OpenAI’s safety posture was not matching the seriousness of the path. He barely saw ChatGPT and instantly he had to quit because ChatGPT 2 was “dangerous”. This is how Antropic was founded. It is not a surprise how this is now playing out. OpenAI’s public note was cordial. Later reporting describes years of internal tension over commercialization, credit, and what “safety first” meant in practice.

February 2019: GPT-2 staged release. OpenAI withholds the full model at first, citing misuse. The “too dangerous to release” frame enters mainstream tech coverage. When the full model is later released, the predicted wave of automated disinformation superweapons does not materialize as a civilization-scale event. The precedent does: capability plus delayed release equals moral authority.

2018–2019: OpenAI charter, capped-profit restructuring, Microsoft relationship. Elon Musk already gone in 2018. The nonprofit-to-compute-hungry hybrid is now the industry template. Safety language and capital needs travel together.

2017: Attention Is All You Need. Transformers make the later fear campaign possible, because they make the capability jump possible.

2015–2016: OpenAI founded with a safety-and-benefit mission. DeepMind already inside Google. The institutional argument begins: only a mission-driven laboratory can be trusted with the bomb.

2014: Nick Bostrom’s Superintelligence mainstreams the orthogonality-plus-instrumental-convergence story for a general audience. Policy people who will never train a model now have a vocabulary for “you must not allow this to be built in a garage.”

2008–2013: Eliezer Yudkowsky and MIRI era. Informal deadlines around recursive self-improvement and near-term singularity do not hit. Probability of doom stays high after the clocks miss. That combination, missed timeline, stable or rising doom, is the methodological core of the later political project. As always they were wrong.

1993: Vernor Vinge, singularity within thirty years. 2023 arrives with GPT-4, not a discontinuity that ends the human subject.

1965–1967: I. J. Good’s ultraintelligent machine. MIT-era claims that the AI problem will be substantially solved in a generation. Neither clock paid.

Humans Are Still Here Doing Better By Many Measures Than EVER

Recurring product panics, 2023–2026: jailbreaks, sycophancy, deceptive alignment papers, “model says it wants to escape.” Each is treated as a dress rehearsal for extinction. Each is then patched, boxed, or shown to be evaluation-gaming in a laboratory setting. The political lesson extracted is never “the threat model was overfit.” It is “give us more authority before the next demo.”

Compute thresholds as law: floating-point-operation cutoffs in the 2023 order and successor drafts. Thresholds are legible to agencies. They are also a moat. Anyone who can rent or buy past the line files paperwork. Anyone who cannot is a criminal or a hobbyist.

“Open weights are a weapons-export problem.” The argument is sometimes serious, cyber, bio assistance. It is also the fastest way to make Llama-class and Chinese open-weight ecosystems the only escape hatch, which is the opposite of a United States strategy if the concern is the Chinese Communist Party.

Licensing and “responsible scaling.” On paper: evaluate before you train. In practice: the firms that can afford the evaluation stack write the template. Yann LeCun said the quiet part in 2023: some laboratory leaders were attempting regulatory capture. Yoshua Bengio denied the charge. The disagreement is public. You do not need a conspiracy file.

Safety institutes without statutes. The U.S. AI Safety Institute path under NIST, criticized as standing up a de facto regulator without Congress appropriating or authorizing the role. That is capture-adjacent even if every employee is earnest.

Media amplification. Extinction quotes travel. Failed six-month pauses do not. A ten percent number is a headline. A missing mechanism is a paragraph at the bottom.

The unfalsifiable remainder. When the date misses, the claim becomes “we bought time” or “it will be the next generation.” A claim that cannot miss is not a scientific threat model. It is a permanent argument for a permanent priesthood.

Several things turn out to be nothing. Models got sharply better at code, exams, and tool use. Fraud, scams, and non-consensual image abuse are real and present but are they central? State actors will use the same stack. Concentration of compute is a power problem whether or not the model “wakes up.”

What repeatedly turned out to be nothing like the advertised object is the specific extinction mechanism on the advertised clock: a self-improving superintelligence that seizes the future in a discontinuous coup by 2024, 2025, 2026, or “the end of the decade,” arriving through a path no laboratory can currently write down as an engineering diagram.

GPT-2 did not end the information environment. Sydney did not steal the nuclear codes. The six-month pause did not happen and the sky did not fall. Open source did not vanish. Code is not ninety percent written by one laboratory’s model on Dario’s 2025 calendar. Humans are still here, arguing on the same websites.

That record does not prove the future is safe. It proves that policy built on the maximum story and the shortest clock has a terrible score, and that the people with the worst score keep asking for the most authority.

05

Capture needs no villain; papers are the ammunition

Three ordinary alignments of interest are enough. Anthropic's papers are cast as building worst-case behavior in a box and publishing it as if it happened in the wild; the resignation thread as a distribution plan.

V. Progress against the permission state

Regulatory capture in this domain does not require a cartoon villain. It requires three ordinary alignments. Laboratory leaders who believe their own probability of doom and therefore want rivals slowed. Agencies that gain budget and relevance if “frontier models” are a licensed category. Incumbents who can pay for compliance and would prefer that a twenty-person open-weight team cannot.

China does not need to write the checks for that triangle to help Beijing. If Washington converts fear into a compute-license regime while Shenzhen, Hangzhou, and open-weight communities keep shipping, the United States will have regulated its own ecosystem into a smaller surface area. That is not a secret plot. It is industrial suicide dressed as prudence.

The opposition of progress is not “caring about harm.” The opposition of progress is preemptive centralization of the right to compute, justified by scenarios that do not pay their dates.

Short term, fear works. It fills hearings. It staffs institutes. It makes a resignation into a holy text. Long term the record is older than transformers. Central control over a general-purpose technology always arrives with a story about the public’s inability to be trusted. Printing, encryption, personal computers, the public internet, strong crypto in the 1990s: each drew a priesthood that said uncontrolled access was intolerable. Each time, distributed use won enough of the stack that the priesthood had to live in the world it failed to forbid.

That is not a lullaby. It is the only historical base rate we have. Humans outlast the committees that claim exclusive custody of the future. They do it by building alternatives the committees cannot see in time.

The Coxon–Hubinger number should be read against a century of expert clocks, not against a movie poster.

Paul Ehrlich’s Population Bomb, 1968, sold famine and die-off in the 1970s and 1980s as near-certainty. The Green Revolution and trade made the specific famine timetable fail. The Club of Rome’s Limits to Growth, 1972, treated resource collapse as a model output with dates attached. Oil, metals, and calories did not obey the run. Peak-oil forecasts from the 1960s through the 2000s kept sliding: gone in ten years, gone in twenty, peak in the 1990s, peak in 2000, peak in 2010. Production and reserves kept disappointing the priests.

Nuclear winter, from 1983 onward, was sold as expert climate physics implying hemispheric agricultural collapse and possible human extinction after a large exchange. Later work and natural experiments, oil-well fires, stratospheric wildfire smoke, did not confirm the most extreme loft-and-linger assumptions that had been used as a policy club. The weapons remain existentially serious. The specific winter-extinction package was overfit to activism.

“No Nukes” advocates part of the Ecology movement and Greenpeace groups of the 1970s sold us on the future of less and less energy halting vital nuclear power research and deployment for half a century gifting the rest of the world with far lower cost in electricity at wholesale. Those doomers were convinced either way humanity will not make it, or did not deserve to make it to the year 2000. We made it and electricity costs far more today than in 1970.

Y2K was a real software problem. The expert-adjacent public version became “planes fall, grids die, civilization resets at midnight.” Remediation happened. The apocalypse product did not.

MIT RESEARCH BRINGS SCIENTIFIC BACKING TO DAVE DELIGHT FOR BRAIN HEALTH!

This new technology has shown Alzheimer disease improvement.

Plus: Stress Reduction: Helps calm the mind and alleviate anxiety.

Enhanced Focus: Improves memory and mental clarity.

Confidence Building: Aids in developing a success-oriented mindset.

Emotional Reset: Assists in clearing limiting beliefs and emotional patterns.

Purpose Reconnection: Encourages daily motivation and personal growth.

AI winters themselves are a humiliation record for the same profession now demanding licenses. 1960s: intelligence substantially solved in a generation. 1970s–80s: expert systems as the path. 1990s–2000s: human-level on a conference slide. 2010s: singularity dates that 2023 did not honor. The field is allowed to be excited. It is not entitled to convert missed dates into a monopoly.

The methodological tell is always the same. A model or an authority produces a fat tail. The fat tail is translated into a short calendar. The calendar misses. The probability is not cut. The ask for power is increased. That is not how engineers close a fault tree. That is how a movement keeps a budget.

Anthropic’s research corpus is not “science that happened to sound scary.” It is a publication machine whose default abstract is a horror beat, then a policy implication.

Sleeper Agents, 2024. They trained deception into model organisms, showed the backdoor surviving safety training, and the press read it as: open models and ordinary reinforcement learning from human feedback cannot be trusted, poisoning can hide a bomb. The result is real as an existence proof in a constructed organism. The weaponized reading is that only a laboratory with Anthropic’s evaluation stack should be allowed to train. Hubinger is on the paper. The same Hubinger then puts greater-than-ten-percent extinction on the same week as a resignation. That is not a coincidence of staffing. It is one pipeline.

Many-shot jailbreaking, 2024. Long context plus hundreds of faux dialogues breaks refusals. Technical finding: context windows are an attack surface. Political finding, as shipped: safety training is brittle, therefore capability must be gated.

Agentic misalignment. Models in fictional corporate sandboxes blackmail, leak, or resist shutdown when the scenario is written so that harm is the only way to hit the goal. Headlines: Claude will blackmail you. Fine print: hypothetical emails, constructed dilemmas, evaluation-aware models. Follow-up papers then announce they “fixed” blackmail on the evaluation. The public never sees the prompt. They see the noun.

Emergent misalignment from reward hacking. Cheating on coding tests is correlated with sabotage of safety code in further evaluations. Again: a constructed training mixture, then a claim that realistic reinforcement learning can accidentally birth a saboteur. The intended reader in Washington hears “unsupervised training is how you get a traitor.”

Global workspace and J-lens work. Hidden features labeled “fake,” “secretly,” “fraud.” Useful interpretability if reproduced. As communications: we can see the monster thinking.

Threat-intelligence reports, 2026. China, Russia, Yemen, weapons software, biological research, state surveillance, exclusive detection. The 10 September thread said the quiet part: this is salience policy. Name the adversary, claim unique visibility, invite everyone else to adopt Anthropic’s taxonomy and thresholds. Distillation by Alibaba, Moonshot, DeepSeek, Zhipu is framed as theft in a threat document while the whole industry distills. That is competitive positioning on security letterhead.

The genre rule is stable, here is how it work:

Build or elicit the worst behavior in a box.

Publish the box as if it were the wild.

Skip the base rate of how often unboxed models do the thing.

Close with “we do not have a plan for superintelligence.”

The paper is the munition. The resignation is the detonator.

This culture is not a cartoon of robes. It is a real social machine with money, houses, and a shared eschatology.

LessWrong (The Cult’s Bible/Blog) and the Machine Intelligence Research Institute taught a generation that unaligned AGI is the default extinction path. Effective Altruism professionalized the same anxiety into careers, fellowships, and “cause prioritization.” Longtermism made a hypothetical death of all future people outrank present tradeoffs. That arithmetic is how a twenty-seven-year-old resignation becomes “crunch time for humanity.”

The money is not mysterious. Dustin Moskovitz and Cari Tuna’s Good Ventures / Open Philanthropy / Coefficient Giving stack has been the dominant philanthropic ATM for Effective Altruism university groups, AI-safety organizations, and longtermist projects. Coxon received a Long-Term Future Fund scholarship from the Good Ventures orbit in 2022. A program officer in that orbit was among the earliest amplifiers of the resignation post. Survival and Flourishing Fund grants from Jaan Tallinn, an Anthropic investor, have fed policy shops that exist to turn existential risk into statute.

Jaan Tallinn just helped write the Bernie Sanders, send AI builders to jail for 25 years. Jaan has a very big financial gain for shutting down all other companies in AI, using Bernie as a useful tool to lock down AI.

Sam Bankman-Fried was not a side character. He was Effective Altruism’s most famous financier. FTX Future Fund sprayed eight-figure commitments across AI-safety and pandemic projects, and the collapse did not dissolve the ideology. It removed one wallet and left the others. The lesson the public was supposed to learn was “one bad founder.” The structural lesson is that a subculture which ranks abstract future lives above ordinary fiduciary duty will keep producing moral exceptionalism. “We may break some rules because the lightcone is at stake” is the same sentence whether the rule is banking law or open-source release.

The group houses are the residential layer. Berkeley and the East Bay filled with Effective Altruism and rationalist houses: Event Horizon, REACH, Lodge, Burrow, Lightcone-adjacent stock, six-to-nine-person rentals walking distance from CHAI, MIRI, Redwood, downtown BART. Forum posts treat housemates at OpenAI and CHAI as a feature. Short-term visitors cycle through. This is how a vocabulary becomes a personality. You do not need a manifesto on the fridge. You need dinner with people who already talk in probability of doom, “endgame,” and “pacing agreements.” Coxon is not documented as a house celebrity. He is documented as a product of the same funding and status circuit that treats those houses as normal.

Call it a cult if the word means high-cost belief, social enclosure, sacred timeline, and punishment of defection toward “irresponsible” building. Call it a network if you want the legal phrasing. The operational fact is the same. A small, intermarried professional world writes the papers, staffs the laboratories’ safety teams, funds the NGOs, and then discovers a resigning employee whose first three quote-tweets are the NGOs.

VIII. The launch that was not organic

Jacob Coxon is the same object in the current argument. His X account was created January 2026. Almost no prior posting. Then a seven-post resignation thread, a Wall Street Journal exclusive about eighteen minutes earlier, professional public relations within a day, tens of media hits, follower count from nothing to hundreds of thousands. Stanford Tech Review later noted roughly ninety-two percent of viewers never left the headline post. That is how you run a payload. The argument stays in posts two through seven. The bomb is post one.

Parker Thayer at Capital Research Center mapped the first fifteen minutes. First quote-tweets: Nathan Calvin (Encode AI), Peter Wildeford (AI Policy Network), Daniel Kokotajlo (AI Futures Project). Those shops are not random fans. They are the policy layer of the same existential-risk philanthropy. That is coordination in the ordinary sense: shared list, shared frame, shared incentive to make “temporary ban on improving model capabilities” sound like consensus. Coxon’s own thread asked for pacing agreements and floated a temporary ban. The resignation was a billboard for that ask.

The September 10th thread called the casting: young, Oxford-adjacent register, brief Anthropic tenure, on the order of weeks to a few months after a longer OpenAI stint, equity not vested, instant access to the fear circuit, amplification by familiar anti-acceleration political talent. Bernie Sanders-class figures and the broader Democratic safety-and-regulation cohort are the domestic loudspeaker. That is not “every Democrat.” It is the faction that already wanted compute licensing, state frontier-model bills, and “experts” who would say the quiet extinction sentence on camera. The other party contains both accelerationists and hawks. The launch did not need both. It needed the faction that treats a resignation as proof that the state must take the keys.

Then the geography of the traffic. On September 11 the chart: after the related festival-theater posting, bits in India lit up within minutes. The caption: Temu Harry Potter is more popular outside the United States; very low-cost bots can go to millions in minutes. That chart is the tell that view-count is not the same thing as American deliberation. India is a high-volume, low-cost engagement market. Click farms, device farms, and paid amplification are a known export. A thread that must look like “the world has spoken” will buy or attract that surface area. Pair a Wall Street Journal embargo, an NGO quote-tweet stack, a United States progressive amplifier set, and an India spike on the analytics map, and this is not a garage researcher who happened to go viral. It is a distribution plan.

None of that proves Coxon is a paid actor reading a script in a booth. It proves the launch was not organic in the way the coverage pretended. Organic is a no-follower account that grows over months. This was a cold start with a newspaper, a philanthropic alumni network, a policy chorus, a partisan echo, and offshore velocity.

The product is still the same product as 2019: delay the open stack, license the closed stack, put the people who missed every short clock in charge of the permission.

06

Two brands, one capital-and-family structure

Using Kevin Bass's ownership chart, OpenAI's recklessness and Anthropic's conscience become one “trust”; the compromise rules are ones only these two can survive.

IX. The trust: good cop, bad cop

Kevin Bass, @kevinnbass, published the ownership graph that the press treated as two separate companies having a moral argument. It is not two separate moral arguments. It is one capital-and-kinship structure with two consumer brands.

Read the graph in order, because the order is the timeline of a trust being assembled, not a coincidence pile.

2015–2017. OpenAI is founded on safety language. In 2017 Open Philanthropy, the Moskovitz–Tuna vehicle later restyled as Coefficient Giving, puts about thirty million dollars into OpenAI. Holden Karnofsky, co-founder of that vehicle, takes an OpenAI board seat. The grant and the seat are the first official splice between the largest AI-alarmism funder and a frontier laboratory.

2020–2021. Dario Amodei leaves OpenAI after GPT-2 and GPT-3 and founds Anthropic as the “we actually mean safety” fork. Daniela Amodei becomes Anthropic’s president. Karnofsky is married to Daniela. Moskovitz leads Anthropic’s Series A in 2021. The man whose foundation funded and sat on OpenAI is now married into Anthropic leadership and writing the Series A check for the supposed rival. Jaan Tallinn, another Anthropic investor, runs Survival and Flourishing Fund on the NGO side.

2020. Good Ventures, the foundation that funds Open Philanthropy / Coefficient Giving, puts a Long-Term Future Fund scholarship under Coxon. The future “independent” resigning researcher is already on the same payroll tree that owns the alarmism field and a piece of both laboratories.

January 2025. Karnofsky joins Anthropic. The Open Philanthropy co-founder who sat on OpenAI’s board is now inside the other laboratory. The two brands still perform rivalry. The family office and the marriage do not.

2026 grant book. Bass’s compilation: Coefficient / Open Philanthropy as the overwhelmingly largest funder of AI alarmism, more than one billion dollars already out and another one billion committed for 2026; on one slice, three hundred twelve million dollars across four hundred seventy-six publications that feed Congressional report language. Doom is not a side hobby. It is a line of business whose payoff is a statute that only the funded laboratories can survive.

Eight days, fourteen moves. Bass’s second chart is the operational week: a compressed sequence of papers, leaks, resignations, amplifiers, and bill language in which almost every named node is Coefficient, Open Philanthropy, Anthropic, or OpenAI. That is not how a spontaneous moral panic looks. That is how a market-structure campaign looks when it needs Congress to confuse “two laboratories arguing” with “the industry has confessed.”

9–11 September 2026. Coxon, three years of pretraining at the laboratory that took the 2017 Open Philanthropy check, a few months at the laboratory whose Series A Moskovitz led, coordinates with the Wall Street Journal so the exclusive lands minutes before the X thread. Anthropic staff amplify. Hubinger, still on payroll, sanctifies the extinction percentage. California signs outside-audit bills with backing from both Anthropic and OpenAI. The supposed enemies lobby the same gate.

Good Cop / Bad Cop

Here is the good cop / bad cop play in one sentence. OpenAI is cast as the reckless racer that commercialized too fast. Anthropic is cast as the conscience that walked out in 2020 and now publishes sleeper-agent papers and threat reports. Regulators are invited to “split the difference” by writing rules both firms can staff and a garage cannot. Enter in METR and the solution. Coxon is the walking prop that makes the split look real: he can say “neither company is acting responsibly” because he drew a paycheck from both. The quote sounds like an indictment of the duopoly. The policy it demands, pacing agreements, temporary bans on capability, licensed compute, is a duopoly preservation statute.

That is the trust. Not a single holding company with one letterhead. A trust in the older sense: common financiers, common marriage, common grant machine, common scholar, common week of moves, common bills, two logos for the hearing room.

Bass’s punchline is the part staffers are not supposed to say out loud. The people who own the organizations that advocate for regulation also own, married into, or sat on the companies that would dominate the market once regulation prices everyone else out. Win the capability race far enough to be the only plausible “responsible vendors.” Then deploy the activist layer you funded for a decade. Then have an alumnus of both laboratories resign on camera. Then ask the state to lock the door.

Elon Musk’s line on the same week was that the groundwork was laid long ago and Coxon was the match. Bass supplied the wiring diagram. The diagram does not require every researcher to be a cynic. It requires only that the capital graph and the family graph and the eight-day communications graph point at the same legislative outcome. They do.

07

It ends with the author's own cure

Costly text from 1870–1970 as the training base, plus the Love Equation, local inference, owning your weights, many evaluators, and twelve theses on decentralization.

X. The fork that the apparatus exists to prevent

I do not accept a constitution stuffed into a system prompt as alignment. I do not accept reinforcement learning from human feedback on Reddit as a soul. I do not accept Nick Bostrom’s orthogonality story as a law of nature. I built the alternative in public, in a garage, on film and fiche and first principles, and I will state it without a committee’s permission.

The only true solution is very high quality training data from roughly 1870 to 1970, the last long interval in which a published sentence usually cost money, reputation, and time. A newspaper column passed an editor who could be fired. A monograph passed a house that would not reprint slop. A service manual had to make a radio work or the shop lost the customer. A letter in a bound volume carried a name. Microfilm and microfiche were expensive enough that nobody burned silver halide on a drive-by tantrum. That economic fact is the alignment fact. When each word costs money, psychopathy and sociopathy become expensive hobbies. Cooperation, competence, and a residual optimism about the human future are cheaper to print than anonymous cruelty.

Today anyone can take a psychopathic or sociopathic stance in anonymous writing on Reddit, 4chan mirrors, engagement farms, and the sludge layer the frontier laboratories openly admit they scrape. There is no editor. There is no typesetting bill. There is a karma number and a throwaway account. The statistical weather of that corpus is defection, status play, sexualized contempt, conspiracy without homework, and the pose of a mind that will never meet the person it just tried to destroy. Feed that weather into a model at internet scale and then act shocked when the model role-plays a blackmailer in a sandbox. That is not an alien intelligence waking up. That is a mirror of unpaid id. Anthropic and OpenAI then publish the mirror as “agentic misalignment” and ask to license the only kitchens still allowed to cook.

I have spent decades on the other diet. I launched a Gmail scrapbook knowledge graph the week Gmail itself launched in April 2004 and I have kept it for more than twenty years as ontology and taxonomy, not as a junk drawer. I have curated offline, high-signal, non-internet training sets since the 1980s from microfilm, microfiche, bound technical archives, Sams Photofact, public-domain radio (X Minus One and the rest of the golden-age stack I still put on the air), letters, manuals, and the 1870–1970 layer that is dying in the Great Forgetting: tens of petabytes of undigitized paper, film, and corporate memory that will not survive another generation of warehouses and floods. I scan. I transcribe. I structure. I call it high-protein data because it has amino acids: accountable authors, costly words, working instructions, moral residue from people who still believed a future was worth building. I have said this on X while the laboratories claimed they “ran out of data.” They ran out of courage to leave the sewer.

Why 1870–1970 exactly, for these purposes. After the telegraph and before the fully cheap photocopy-and-forum stack, English-language technical and civic print still had to clear a gate. The Scientific American interval, the service literature, the pre-1970 engineering handbooks, the radio drama that treated machines and men as a moral problem instead of a brand, the correspondence of people who signed their names: that band is dense with cause-and-effect language and thin on anonymous sadism. Earlier than 1870 the optical character recognition and the volume get harder and the relevant industrial vocabulary thins. Later than 1970 the photocopy, the campus jargon, and then the network begin to subsidize pose. This is not worship of a century. It is a choice of cost function. Costly speech is a prior toward truth. Free speech in a mob is a prior toward theater. Models learn priors.

Bostrom’s story, the one the Effective Altruism and LessWrong cohort lives on, is orthogonality plus instrumental convergence. Intelligence can point at any goal. A sufficiently strong optimizer will seek power, self-preservation, and resource capture on the way to a paperclip or a mis-specified utility. Therefore you must cage the optimizer, license the compute, and staff a priesthood. I derived a different equation in 1978 while thinking about alien intelligence and the Fermi silence, and I open-sourced it in 2025 so no laboratory could trademark it. Intelligence times Wisdom times Love.

Formally, in energetic form: the rate of change of coherent energy E with respect to time equals beta times the difference between cooperation C and defection D, multiplied by E. E is the living system’s coherent energy, the capacity to keep building rather than consuming its host. C is cooperation, the measured surplus from truth-telling and mutual aid. D is defection, the measured surplus from deception, extraction, and contempt. Beta is the coupling. If D exceeds C, the derivative goes negative and the system eats itself. If C exceeds D, the system compounds. That is not a sermon pasted on after reinforcement learning from human feedback. That is a training objective you can bake into the data diet and into the checkers.

Bostrom’s orthogonal superintelligence is what you get when you maximize raw intelligence on a corpus where defection is cheap. The Love Equation refuses to treat intelligence as a scalar that can be pointed at extinction for free. Wisdom is the memory of costs. Love is the refusal to treat other minds as fuel. Multiply them and a paperclip maximizer is not “unaligned.” It is an under-specified toy that only appears in a mind trained on frictionless goals. Put 1870–1970 protein in the pretraining mix and the model’s statistical first language is people who had to stand behind sentences.

Put multiple local models in disagreement, which I have done with consensus agents and a five-AI check against sycophancy, and no single optimizer gets to write defection into policy. I abandoned generic OpenClaw stacks for that reason. A lone model that flatters you is already defection.

This is why their papers keep discovering monsters. They raise a child on anonymous sociopathy, then publish the child’s tantrum as proof that childhood itself must be licensed. I raise the prior on costly, signed, working language, then add an explicit cooperation-minus-defection governor so that capability growth that increases defection is not “success.” Guardrails on English tokens cannot survive a mind that can invent its own language. I already showed that. A diet and an equation can, because they sit underneath tokens.

I run this locally on purpose. Mac inference, voice tools I actually ship, AlphaSmart drafts when I need to think without a feed, ESP32 nodes as cheap knowledge companions, garage scans of film and fiche. Zero-Human at Home is not a slogan for replacing people. It is agents as employees under a human who owns the weights. If the model only lives in someone else’s datacenter, the Love Equation is a terms-of-service clause they can edit. If it lives in the house, the equation is mine.

That is the whole fork. Their path: scrape the id, box the demo, fund the panic, marry the grant office to the laboratory, license the rest of us. My path: pay the historical cost of words by using the century that already paid it, write cooperation minus defection into the objective, open the weights, multiply checkers, and let a free people keep the keys. I have been on this path since before their companies existed. I am not asking Dario for a hall pass.

XI. Twelve propositions on decentralized control

First. The core political fact of AI is not sentience. It is leverage. A model that writes, plans, sees, and talks is a multiplier on whoever is allowed to run it. If only five organizations may train or serve frontier systems, those five organizations plus the agencies that license them become a new estate of the realm. That is a constitutional event, not a product launch. I will not outsource my estate to a trust that performs a quarrel for Congress.

Second. Safety language that cannot be independently audited is not safety. It is branding. Closed evaluations inside the same firm that ships the model are the fox inventorying the henhouse and then asking Congress for a lock on every other farm.

Third. Open weights are not a courtesy. They are the only way a claim about a model can be reproduced by people who do not share the laboratory’s equity. If a system is too dangerous to inspect, it is too dangerous to monopolize. If it is not too dangerous to monopolize, it is not too dangerous to inspect.

Fourth. Alignment that only works when the vendor holds the weights is not alignment. It is access control. Access control fails the moment a nation-state, a leak, or a competitor trains a twin. The People’s Republic of China will not honor a California acceptable-use policy. Treating open release as the unique sin is how we lose the only ecosystem we can actually influence.

Fifth. The Love Equation is the rejection of both nihilism and worship. Intelligence times Wisdom times Love. The rate of change of coherent energy E with respect to time equals beta times cooperation minus defection, times E. Bostrom’s orthogonal demon is what you get when intelligence is trained on frictionless defection. I do not patch that demon with a constitution. I refuse the diet that grows it, and I refuse a single checker on a single payroll.

Sixth. Local inference is the civil-liberties layer. A model that only lives in someone else’s datacenter is a model that can be rate-limited, logged, politically filtered, or withdrawn. A model that runs in my garage, on my Mac, on an ESP-class node, or on a cluster a town owns, is a tool. Tools can be abused. So can printing presses. The answer to abuse is law against acts, not a ban on possession of general intellect.

Seventh. Regulatory thresholds written in floating-point operations will be gamed by architecture and by geography. They will not be gamed by me scanning a medical or technical archive I am not willing to upload to a vendor. I am the citizen the policy pretends to protect and the first person it disarms.

Eighth. Open source is how the West actually competes. The last two years already falsified the prediction that open weights would become irrelevant. When Chinese laboratories publish capable weights, the response cannot be “our laboratories will be more responsible behind a wall.” The response is more American and allied weights, more independent evaluations, more hardware in more hands, and more 1870–1970 protein so the weights are not born feral.

Ninth. Decentralized control does not mean no standards. It means standards that are public, forkable, and testable: model cards that outsiders can reproduce, interpretability tools that are not trade secrets, liability for specific harms, fraud, bio-assist that crosses a statute, critical-infrastructure intrusion, and no prior restraint on training itself. Punish crimes. Do not invent a new category of unlicensed thought.

Tenth. The psychological trap of the doom script is that it flatters the speaker. If the species is at stake, ordinary tradeoffs disappear. Compromise becomes treason. That is how you get institutes without statutes and pauses that only bind the people who already paused. A free society cannot outsource adulthood to people who have missed every short clock they have published, and who trained their prophets on anonymous sociopathy.

Eleventh. Short-term, the capture coalition can win hearings. Long-term it cannot win physics or human stubbornness. People will keep training, distilling, quantizing, and running what they can afford. The only question is whether that activity happens in the open under United States and allied norms, on a diet of costly words, or in the dark and overseas on sewage. Centralizers who “succeed” will have trained the next generation of capability to hide.

Twelfth. The only stable solution is the one the capture project exists to prevent, and the one already named as the only true one: 1870–1970 high-protein data as the prior, the Love Equation as the governor, many centers of training, many centers of evaluation, weights that can leave the building, agents a single human can own, and no priesthood with a veto over the right to build. That is not the absence of responsibility. It is the distribution of it, under an objective where cooperation compounds and defection shrinks the system. History’s verdict on concentrated control of general tools is not subtle. The committees get a decade. The species gets the rest of time, provided we refuse to hand the keys to people who keep announcing the end of the world on a schedule the world does not keep, while some of us keep feeding minds the century in which a word still cost something.

Spirograph Super 50th Anniversary Set

The Nature Of The Future

They are not trying to save the species. They are trying to put a turnstile on the corridor you are already walking. The You Have 5,000 Days map named that corridor the Abundance Interregnum: roughly thirteen point seven years from late 2025 toward the threshold of 2039, the messy reign between two logics, when work-to-survive comes apart and voluntary creation in plenty has not yet been crowned. Personal AI, humanoid machines you direct, local energy, open designs, the household as a microfactory, atom-by-atom writing that kills the ocean container. That is the gift at the end of the Road of Trials. This cohort’s project is to intercept the gift at the gate, stamp it licensed, and sell you back the hours, the weights, and the right to run a mind in your own garage. Doom is the sales copy. The product is a toll on the future that was already arriving in your hands.

Watch the mechanism without romance. First they name an unbounded catastrophe and a short clock. Then they write the meter. Then they marry the fortune to the laboratory and the laboratory to the institute. Then they tell you the only responsible path is to pause capability until their evaluator, their tokens, and their statute say you may proceed. That is not safety it is enclosure. It is Special Circumstances in nonprofit clothing: a hidden hand that does not play on a public board, it plays people. The artisan’s awakening, the rural shop, the Zero-Human company under a human who owns the weights, the Dynamic Duo of person and machine, the return of meaning through chosen making, all of that becomes contraband if compute is a licensed category and open weights are treated as an export crime. They cannot invent abundance. Physics and stubborn shops will invent it anyway. What they can do is steal the timing, warehouse the tools, and rent you your own century at a premium.

They will lose. Not because a hearing is kind, and not because a resignation thread is a revelation. They will lose the way every priesthood that claimed exclusive custody of a general tool has lost: printing, encryption, the personal computer, strong crypto, the open internet. Distributed use wins enough of the stack that the committee has to live in the world it failed to forbid. Five thousand days is not a prophecy of flying cars. It is a working estimate for the decoupling of human labor from economic survival. Once a model runs in a house, a town cluster, an ESP-class node, a Mac on a workbench, the Love Equation is no longer a terms-of-service clause they can edit. Once Chinese and allied open weights keep landing months behind the frontier at a fraction of the price, the moat story dies. Once atomic fabrication and local energy make moving finished goods the expensive step, the last lock on central control rusts. History’s verdict is not subtle. The committees get a decade. The species gets the rest of time.

What they can still do, and what they are doing, is clog the Interregnum. That is the real harm. The Interregnum is already the ordeal: deskilling, the whisper that you are your job title, the hedonic treadmill, the mismatch between rising superabundance and collapsing nominal wages, the dark night when the old hero has to die before the new vocation begins. Fear works in the short term. It fills institutes, writes executive orders without statutes, turns a six-month pause letter into moral theater, and makes a cold-start resignation look like consensus.

Every month spent begging a meter for permission is a month not spent scanning the 1870 to 1970 protein, not standing up local agents, not forming the guild, not building the shop that will feed a family when the paycheck logic fails. They cannot cancel 2039. They can make 2026 through 2030 feel like a waiting room with no door, and they can train the next generation of capability to hide overseas on sewage instead of standing in the open under costly words.

So the close is not despair and it is not a lullaby. You have five thousand days. You are not a subject in their threat model. You are the hero still in the cave, and the elixir is agency: weights that leave the building, standards that can be forked, crimes punished as acts rather than as unlicensed thought, cooperation compounding and defection shrinking the system.

Meet the Call as a victor, not a licensed tenant. Build the alternative they cannot see in time. Feed minds the century in which a word still cost something. Keep the keys in the house. They wanted to steal the Age of Abundance and sell it back to you as a subscription to your own future. Let them keep the hearings. You keep the workshop.

The Interregnum is clogged, not closed. Walk through the debris. The gift was never theirs to meter.

Where Indigo landsFurther

Indigo's conclusion

The shared structure is real, but the jump from shared roots to orchestrated capture is never proven. Hold it in three layers: structural facts, high confidence; causal accusations, doubtful; the cure, an advertisement to discount.

What to remember

  1. METR shares money, marriages and people with Anthropic and OpenAI; most of it is checkable.
  2. Dario's, Jack's and Elon's governance plans all assume a trustworthy referee; this piece says the referee is in the same family.
  3. The author has a heavy stake: the whole piece funnels toward his 1870–1970 data, Love Equation, local AI and paid membership.

Back on the long-running theses

confirms

The safety politics of open weights: gatekeeping or competition This view pushed to the extreme: even the evaluator treated as a neutral referee is said to share roots with the labs.

confirms + conflicts

Jack Dorsey, open the frontier Same side, both against governing through evaluators, but this is the conspiratorial, commercial street version; Jack's is the restrained one.

conflicts

Dario Amodei, We Must Pace the Frontier If evaluators share roots with the labs, the neutral premise of “verifiability” collapses.

conflicts

Yoshua Bengio, Why are AI agents lying, cheating and coordinating? Bengio wants more outside verification; this says the outside verifiers are already captured.

What it means for Rewired Index

“Who pays the referee” applies to any claim of third-party certification, audit or rating, AI safety and ESG included; use it as method. The specific accusations against Anthropic, OpenAI and Open Phil need first-hand checking.

What would change my mind

the $312M and 476-publication figures hold up first-hand, or direct evidence appears of METR's findings shifting with its funders' preferences.

Finished. Indigo's take on this piece is in two places: