Mind · In / Out · In · 视频

AI 研究者辩论:我们离递归自我改进还有多远

AI researchers debate how close we are to recursive self-improvement

John Schulman、Beren Millidge、Charlie O'Neill · YouTube · 2026-09-12

三位一线研究者辩论离递归自我改进还有多远:会来,但被定目标和持续学习这些工程摩擦按住。

第 1 段 / 共 12 段 · 0:01
卡住的不是聪明,是下一个不连续

2036 年若一切如常,原因多半不是模型不够聪明,而是泛化和持续学习卡住;摩尔定律那条直线靠一次次不连续续命,当前范式未必能自己找到下一次。

拆解 · 12 步

  1. 01

    0:01 – 7:08

    卡住的不是聪明,是下一个不连续

    2036 年若一切如常,原因多半不是模型不够聪明,而是泛化和持续学习卡住;摩尔定律那条直线靠一次次不连续续命,当前范式未必能自己找到下一次。 读这一段视频稿 →

  2. 02

    7:08 – 18:51

    目标清楚能快 10 倍,定目标本身难

    目标清楚时,当前范式能加速 10 倍;可提出正确的目标没法光靠想出来。留给人最久的工作,是定义目标、决定我们要什么。 读这一段视频稿 →

  3. 03

    18:51 – 28:10

    蒸馏让谁也甩不开谁

    强化学习学到的都能便宜地蒸馏,关键是提示分布;前沿实验室在训练环境上未必有优势,真实部署也许更重要。 读这一段视频稿 →

  4. 04

    28:10 – 33:52

    实验室里的持续学习

    自动化研究员靠人类反馈加练习环境训练出来;现在是把最近三个月的进展蒸馏回模型,所以有人觉得它是渐近的;环境能超出人做得到的范围。 读这一段视频稿 →

  5. 05

    33:52 – 41:27

    在模拟里学会学习,够不够

    实验室的赌注,是在海量模拟环境里练出能坚持、能分轻重的智能体,按领域一个个补;样本效率低时只能靠模拟,和真人实时互动很难模拟。 读这一段视频稿 →

  6. 06

    41:27 – 52:30

    累积型任务比律师助理好做

    部署数据已经一代代回流进模型;RSI 是累积型任务,发现一个守住一个;律所这种分布一直在变的工作反而难。品味也许能从短片段里学出来。 读这一段视频稿 →

  7. 07

    52:30 – 1:00:38

    持续学习在细处会崩

    企业不愿让供应商从自己的部署中学,压力推向可插拔模块;放大看能用,落到一个模型反复小更新,就出现灾难性遗忘,最后只能从零重训。 读这一段视频稿 →

  8. 08

    1:00:38 – 1:10:32

    信号从哪里来

    环境梯子每级越来越难爬,前沿需要的比特连整个世界都给不够;小规模实验里数据约贡献 12.0 倍效率、架构约 3.7 倍,但架构能打开新区间。 读这一段视频稿 →

  9. 09

    1:10:32 – 1:18:06

    模型还会不会变大

    推理效率优先,未来几年参数可能停在一个平台;Schulman 预计模型仍会变大,但幅度取决于规模定律;漂亮的直线背后藏着大量复杂性。 读这一段视频稿 →

  10. 10

    1:18:06 – 1:25:00

    中训练干了八成,强化学习只调策略

    中训练常把模型带到约 80% 的位置,强化学习用信噪比极高的一个比特调策略;得到的是任务时长上的泛化,不是跨领域的泛化。 读这一段视频稿 →

  11. 11

    1:25:00 – 1:28:48

    第 37 手与单一文化

    强化学习没有彻底扼杀创造力,模型曾一次找出好几个零日漏洞;但输出多样性下降,大家都蒸馏 Claude,开放权重模型越写越像。 读这一段视频稿 →

  12. 12

    1:28:48 – 1:37:01

    时间线:两年到十年

    即插即用的远程员工约 1 到 3 年;AI 研究者提效 10 倍,Schulman 说两年,O'Neill 说 5 到 10 年;超级智能 3 到 10 年。 读这一段视频稿 →

Indigo 的结论

这是「能力是有界的指数」的一线从业者辩论版:RSI 大概率会来,但被定目标、持续学习、样本效率按住,不是快速起飞。速度问题仍没有定论。

怎么读这篇 三位一线训练和研究者的圆桌,主持人说选他们,是因为他们在相对开放的实验室、能公开说话。没人卖产品,反而互相拆台、常常承认不确定,是难得的诚实信号。整体看,RSI 会来,但被一堆工程摩擦拖住。

需要记住的几件事

  1. 瓶颈从算力改判到定目标和持续学习:2036 年若一切如常,原因多半不是模型不够聪明,而是不知道该干什么、在部署中学不动。
  2. 对齐是最后一份工作:留给人最久的是定义目标、决定我们要什么;技术活能自动化,指定正确的目标短期内让不出去。
  3. 持续学习在细处会崩,只能重训;RSI 这种累积型任务,反而比分布一直在变的真实工作好做。
  4. 蒸馏是反集中化的力量:可蒸馏的能力守不住,价值上移到独有的真实部署数据,比如 Composer、Harvey 这样的垂直场景。
  5. 强化学习管用,是因为中训练已完成约 80%,强化学习只用一个高信噪比的比特调策略;它延长任务时长,不跨领域。

什么会让我改口

下一个不连续是在当前范式上做加法,而不是推倒重来,当前范式能自己把点连成线。

怎么读这篇

三位一线训练和研究者的圆桌,主持人说选他们,是因为他们在相对开放的实验室、能公开说话。没人卖产品,反而互相拆台、常常承认不确定,是难得的诚实信号。整体看,RSI 会来,但被一堆工程摩擦拖住。

拆解 · 12 步
  1. 卡住的不是聪明,是下一个不连续
  2. 目标清楚能快 10 倍,定目标本身难
  3. 蒸馏让谁也甩不开谁
  4. 实验室里的持续学习
  5. 在模拟里学会学习,够不够
  6. 累积型任务比律师助理好做
  7. 持续学习在细处会崩
  8. 信号从哪里来
  9. 模型还会不会变大
  10. 中训练干了八成,强化学习只调策略
  11. 第 37 手与单一文化
  12. 时间线:两年到十年

据视频字幕整理,按说话人分段。

01

卡住的不是聪明,是下一个不连续

2036 年若一切如常,原因多半不是模型不够聪明,而是泛化和持续学习卡住;摩尔定律那条直线靠一次次不连续续命,当前范式未必能自己找到下一次。

00:00 · 如果 2036 年一切如常,是哪里没走通

0:01Dwarkesh Patel: 今天我请来三位 AI 研究者朋友,每次和他们聊都能学到很多,而且他们恰好在相对开放的实验室和公司,所以能公开说话。Beren Millidge 是 Zyphra 的 CTO,Zyphra 做开源模型。John Schulman 是 Thinking Machines 的首席科学家,之前是 OpenAI 联合创始人,主导了通向 ChatGPT 的 RLHF 工作。Charlie O'Neill 是 Baseten 的模型训练负责人。第一个问题:假如到了 2036 年,世界上并没有几十亿个超级智能到处跑、把世界彻底改变,那么除了政治冲击、战争或者禁止 AI 这类外部原因,最可能的技术原因是什么?

1:16Beren Millidge: 有一个经典的模式,几乎就是莫拉维克悖论:我们总以为 AI 要是能解难的数学题、能下赢国际象棋,那就了不得了;结果它做到了,影响却比想象的小。如果这种情况一直持续,真正的泛化火花始终没有出现,AI 可能会在人们放进基准测试、放进训练环境的一切上都极其厉害,却始终被某种从模拟到现实的落差卡住其他所有事。我认为这不太可能,因为我们已经看到这种泛化,连强化学习都能带来。但如果元学习极难泛化,再加上持续学习解决不了,那就是我的默认剧本。

1:50John Schulman: 我同意。人类现在相对模型仍有很多优势。每出一个新模型,就在某些方面追上来,但你总会被模型更弱、判断更差、没法充分自我核查的地方卡住。有一个循环一直在重复:新模型出来,大家惊呼这就是 AGI 了,用了一个月左右,又开始觉得它笨。这个循环可能还会继续,很难预测会重复多少次。现在能力之所以没有爆炸式增长,是因为做研究和做工程时瓶颈仍然够多:哪怕模型写的代码比人多得多,也没让人的效率提高 100 倍。

2:55Charlie O'Neill: 在我看来,问题在于当前的配方,也就是 Transformer 加强化学习,离「能装进一块芯片的学习者」这个全局最优有多远。人们想象,一旦有个智能体在 AI 研究上比所有人类强,哪怕只强 0.1%,你能并行跑几十万甚至几百万个,芯片变快它也跟着变快,这就压倒了所有其他瓶颈,于是出现非常快的起飞。但想想摩尔定律:一条漂亮的直线维持了非常久,靠的是一次又一次离散的不连续和创新在续命。LLM 也一样:预训练的规模定律撞上收益递减,然后强化学习出来接上,又给出一条新曲线,所以看上去仍是一条向上的直线。如果还需要再来一次这样的不连续,我不确定用强化学习环境训练 LLM,哪怕是专门瞄准 RSI 的环境,能把它发现出来。发现不了,我们大概就会撞上一条渐近线。

4:25Dwarkesh Patel: 你觉得那次不连续会比 2012 年以来的任何一次都难吗?

4:34Charlie O'Neill: 要是知道答案,我们就能把它做出来了。但应该区分两种情况:一种不连续是给当前范式做加法,是累积性的,强化学习之外还有某个东西要发现,模型也许能自己把点连成线;另一种还是那个问题,我们离全局最优有多远:是不是得回头把梯度下降、把神经网络本身都扔掉?如果差得那么远,我不认为继续放大当前范式,不管跑多少个 LLM,能把它发现出来。

5:05Dwarkesh Patel: 说到底,唯一的指望是深度学习压根到不了「至少在研发上碾压人类、包括提出新范式的能力」那一步。可如果把 2012 年到现在的进步直接延续下去,我知道这背后是巨量算力的扩张,它要是没能至少在研发上碾压人类,尤其是在未来几年,那才奇怪。Ryan Greenblatt 最近上节目时提过一点:AI 越来越强,就能在模拟环境上取得进展,而这些环境奖励的不只是 AI 研发能力,还有整体的科学能力,所有实验室和很多创业公司都在瞄准这个。另一个直觉泵是 80 年代以来国际象棋程序的 Elo 分:上升非常线性,但跨过人类区间时是一个巨大的不连续,从人类专家总是赢 AI,变成人类专家再也赢不了。到目前为止 AI 对经济的最终影响不大,是因为它相对人类的 Elo 还在慢慢爬。

6:40Beren Millidge: 我同意,那会非常意外。唯一不发生的可能,是进步恰好在跨越之前停在渐近线上,而在我看来,我们已经很接近开始跨过人类的 Elo 区间了。要落到你说的「2035 年一切如常」,唯一的另一条路是对 AI 的激烈监管。其实我觉得这比技术原因更有可能。

02

目标清楚能快 10 倍,定目标本身难

目标清楚时,当前范式能加速 10 倍;可提出正确的目标没法光靠想出来。留给人最久的工作,是定义目标、决定我们要什么。

07:03 · 目标清楚的研究,和开放式的科学

7:08Charlie O'Neill: 研究分好几种。有一种是自动研究式的,目标已经被很干净地指定,照着优化就行;大家想象的是,把预训练损失继续往下压、把环境里的奖励继续往上推,就会带来进步。Ryan 说的也许是另一种开放得多的科学,范式转变需要的那种:我们没法指定目标,AI 当然也指定不了。无论做哪种事,指定目标都要非常非常小心。

7:49Dwarkesh Patel: John,你当年在一线。想必一个大突破是意识到「预测下一个 token」就是那个该做的事;2014 年你不会想到 nanoGPT 速通是值得优化的东西。现在进了这个范式,你当然会想到去速通它,让 AI 在这上面变得很强。但也许还有下一个内循环,是 AI 预料不到的。外面有一个收入的外循环,最终应该很有力,但它非常慢。

8:29John Schulman: 其实我记得在 OpenAI 早期,我的直觉是光把对数损失降下来到不了智能,因为重要的那些比特在损失里占比太小,会被噪声淹没,所以得设计更好的目标,把更多权重放在重要的东西上。这方面可以找出各种理由:人类大概并不学着去建模环境里的一切;大多数人没法把看过的场景画成照片那样逼真;所以一定需要更好的目标。结果它就是管用了。

9:33Dwarkesh Patel: 就像你说过的,即使在今天的 AI 研究里,后训练基准测试这个内循环,也不一定能转化成用户喜欢的东西。

9:42John Schulman: 对。整个领域很依赖泛化,而什么时候能泛化、尤其是超出分布的泛化,很难预测。在你在乎的任务上训练会更好,但最重要的进展往往是我们本来没理由指望的那种泛化:从很朴素的预测下一个 token,泛化到需要深度理解输入的任务,或者学到预训练数据里很少见的技能;还有从可验证的任务泛化到不那么可验证的任务,这事先也没有理由指望。

10:45Dwarkesh Patel: 有一个直觉泵,说明为什么哪怕不扩大 AI 劳动之外的投入,也可能很快出现奇点:每做一个七位数美元的实验之前,都先花同样多的算力在 AI 劳动上,等于有一群自动化的你们,花一个世纪去想最优的实验是什么,做小规模消融,建立整整一个世纪的理论;实验做完,再花一个世纪分析结果、决定下一个实验。

11:34John Schulman: 想得足够透,其中一些事大概本来就能预见。很可能有某种巧妙的小规模实验,能让人建立起可以推广到大规模实验的理论。所以我认为,研究能做得多好,我们离天花板还远得很。我可以想象 AI 做大量分析和理论构建,花在这上面的算力和花在实验本身上的差不多。

12:17Charlie O'Neill: 在目标指定清楚的时候,确实有很具体的例子。思考能做的,只是根据形成先验之后拿到的比特去更新后验;光靠想,得不到任何新比特。但当目标清楚、数据就摆在那里时,我预计当前范式里会有很大的加速。比如让 AI 去看 Kaplan 的规模定律,它现在就会发现:他们拿中间检查点来用,却没考虑退火,所以结论是错的。这本来能提前好几年发现,光这一个观察就能省下一两年的进展。muP、学习率随模型规模怎么变、模型宽度也很重要,这些都一样:只要多想一想,很多东西都能倒推出来,把低垂的果子摘掉。如果任务只是「把现有目标最大化」,我预计会有 10 倍的加速。但我看不出这怎么能泛化到一开始就提出正确的目标。光靠想,买不来正确的目标。

13:34Beren Millidge: 这正是当前 AI 要实现任何极快 RSI 的关键问题:AI 能在多大程度上泛化到学会自己定目标?要有一个自我推进的自动循环,AI 得能提出目标、优化、弄明白、再提新目标,而且很长很长时间里都不能跑偏。回到莫拉维克悖论,这里也许又是一个:我们觉得这种自主,自己想好该做什么、再去做、如此循环,非常容易,因为我们一直在做;进化当然得造出能长时间独立生存的生物。但出于某种原因,这对 AI 可能恰恰很难,就像运动控制对它很难、数学却很容易,和我们正好相反。

14:00Dwarkesh Patel: 可任务时长一直在变长,难道不说明情况不是这样吗?

14:25Beren Millidge: 没错,我也同意这方面没有明显证据。我们的智能体现在非常能坚持,这件事做起来也挺容易,这其实是反证。但如果它真的难,这可能就是我们没有立刻起飞的原因之一。

14:35 · 留给人的最后一份工作:决定我们要什么

14:39Dwarkesh Patel: 从 2012 年,或者从你们开始做研究算起,这些年的所有创新里,不管是纯工程的还是纯概念的,哪一类看起来会是 AI 完全自动化 AI 研发之前,人类要做的最后一件事?

15:05Beren Millidge: 大概就是一轮一轮地提出对的问题。就算 AI 什么实验都能做,也还得有人决定做哪些实验。现在 AI 在这方面比写实验代码差得多:一聊研究,它们就提出一堆零零碎碎、步子非常小的东西。

15:24Charlie O'Neill: 或者说,从 DeepMind 那种「通过学会以超人水平玩游戏来解决智能」的路子,跳到一位研究者,Radford,决定「我就试着在范围极广的数据上预测下一个 token」。即使 Radford 发现了这一点,也过了一阵子才有人决定把它放大,因为我们得先想出规模定律,想到这些东西可以非常可靠地预测。

15:55John Schulman: 我会说,人类最后的工作,或者说留存最久的角色,是定义目标、决定我们到底想要什么。比如 AI 助手该怎么表现、什么叫有帮助、做基于人类反馈的强化学习时目标是什么,这是一类;后来写宪法、写模型规范,又是一类。就算 AI 能做所有技术活,这类事我们仍然要做很多,要决定我们真正想要什么。

16:31Dwarkesh Patel: 对齐是最后一份工作。

16:33John Schulman: 对齐算是答案。但对齐本身可以拆成两部分:指定目标,也就是弄清正确的目标应该是什么;然后真正去实现、去优化你定下的目标。前一部分短期内不会消失。想想后训练团队为什么需要那么多人:因为有很多不同领域,得有人去想模型在这里该怎么表现,这很难整个自动化。

03

蒸馏让谁也甩不开谁

强化学习学到的都能便宜地蒸馏,关键是提示分布;前沿实验室在训练环境上未必有优势,真实部署也许更重要。

18:39 · 为什么没有哪家实验室一骑绝尘

18:51Dwarkesh Patel: 为什么模型供应商没有出现巨大的集中化?指向集中化的因素太多了。放到好几年的尺度上,有什么东西会阻止它吗?

18:54John Schulman: 我认为蒸馏是对抗集中化的主要力量。基本上,凡是能通过强化学习学到的东西,都能很容易地蒸馏,因为那只是少量比特,能从少量数据里学到。只要能拿到模型展示某种行为的轨迹,就能轻松蒸出来。另外也有可能出现公司专属的模型,能从部署中学习,一家公司持续改进自己的模型。这样的系统可能由现在的几家寡头提供,也可能由某家现在还小的公司提供。但我觉得这会稍微改变格局。

19:56Beren Millidge: 我还想指出,持续学习其实挡不住蒸馏。就算你的模型每天都在变好,别人也可以每天蒸馏它。两个循环可以同速运转。

20:10Dwarkesh Patel: 有道理。可要复制模型的行为,你自己得知道该用什么样的提示分布去问,才能引出相关行为吧?

20:18John Schulman: 对。光用监督学习去蒸馏,提示分布极其重要。即便你能完全访问模型、拿到思维链,要把它所有有用的能力蒸出来也非常不容易,因为你得用东西去问它,得用真实的提示,而且要一个覆盖面非常广的真实提示分布。最近有消息说,一些中国公司可能在用中转服务:这些服务让中国用户能用上本来被屏蔽的美国前沿模型,主要用来写代码,而这些中转服务在收集并出售一部分数据。这对蒸馏是非常有用的数据集,因为它给了你完美的提示分布。

21:29Beren Millidge: 这正是 AI 能帮大忙的地方。看看前沿的训练流程,或者中国模型在论文里写出来的做法:他们先从某处拿到种子提示,来源是人加上这类数据,然后用自己的模型或其他前沿模型,从种子提示合成出覆盖面极广的数据。收集提示分布、搭建环境,很大一部分都能自动化,模型越好,人需要提供的比特就越少。

21:57Dwarkesh Patel: 但看起来你还是会被「有没有一个真有用户在用的服务」卡住。用户会说:给我做个这样的应用;不行;其实我想加这个功能;算了,退一步,做另一件事。把这一整串过程记下来,才是关键。要是这些你本来就能自己生成,那你就已经有 RSI 了。

22:25Beren Millidge: 不一定。有用户当然很有帮助,但理论上,你可以直接去想用户要什么。归根到底,完全自动的循环本质上就是 RSI:AI 决定数据,决定训练,这就是那个循环。但这取决于你需要多少来自人的信息。到了某个程度,如果你只是想要「长这样的轨迹」,就去提示模型,它能给出相当不错的近似。

22:40Dwarkesh Patel: 那如果你想要「给我做一个非常好的政治家」,它得凭空预判参议院里的一场辩论会怎么进行呢?

22:51Beren Millidge: 讽刺的是,这对蒸馏方反而比对前沿实验室容易。蒸馏方只要说「我要一个好政治家」,然后去问前沿模型;前沿模型已经知道怎么当好政治家,直接生成轨迹就行。而如果你要造第一个做到这件事的模型,就得想办法真的拿到政治家每天在做什么的数据。说「我要这样的东西」,再让 AI 生成十亿个变体,远比一开始把这个东西造出来容易。

23:19Charlie O'Neill: 从「中国实验室手里有中转数据」这一点,其实能推出一个很具体的预测。起因是我说过:Sonnet 5 和 Opus 5 几乎客观上比 GLM-5.3 和 Kimi K3 更差,而前者不仅能蒸馏,甚至能从 Mythos 做 logit 蒸馏,这不奇怪吗?对方的反驳是提示分布真的非常重要,你得看到用户在做什么,才能把这些行为蒸进去。由此推出的预测是:前沿实验室现在在强化学习环境上未必有多少优势,甚至可能完全没有。用户分布对一般行为当然重要,但衡量一项能力的最好标尺,是你在前沿造出来的那些极难的强化学习环境。如果 Anthropic 既有这些环境,又能做 logit 蒸馏,做出来的模型还是更差,那也许……

24:12Dwarkesh Patel: 那就是说,真实世界的部署比环境更重要。这很有意思。可他们当初总得先在 Fable,也就是前沿模型里,把这些能力激励出来,那为什么在一个更小的模型上就激励不出来了,这很奇怪。

24:32Charlie O'Neill: 也许我们正处在一个诡异的恐怖谷里:太想照搬前沿模型,师生之间的差距,不管具体是什么,就是太大了。大家在 Opus 身上也说过这点:Opus 4.6 和 Opus 5 的区别在于,Opus 5 感觉像是有一个 AI 评审在检查它做过的每一件事,所以它用那么多 token。它想把所有事都考虑到,却没有 Fable 那种大模型的感觉,不知道什么时候该停,哪条路值得走下去。

25:02John Schulman: 我给一个稍有不同的假设。能造的环境有两个维度:难度和真实度。造大量高难度的环境相对容易,让模型做复杂得多、或者需要多得多聪明才智的任务;可以把这叫「刷榜分布」,因为很多最出名的基准测试就是非常难、像解谜一样、又容易验证的任务。另一个维度是真实度:你希望模型在真实的编程智能体场景里表现好,和人来回好几轮,同时有好几个目标。第一次打造某种模型行为的实验室,两个方向都得推,而要行为好,就得在真实度上使劲,用评分细则或某种人类反馈去喂奖励函数。可如果天真地去蒸馏,你最后只会在刷榜分布上追平老师;没有足够多真正考验这些棘手真实场景的环境,学生模型就学不到这些能力。也许一种情况是,大模型从刁钻的窄任务泛化到更真实任务的能力更强。如果你有一个非常好的真实提示分布去蒸馏,就能把大模型模仿得很像;如果只有一堆容易验证的任务,你能在所有基准上追平大模型,但在更广的分布上会更差。这也许能解释 Anthropic 的小模型,比如 Sonnet 5 的一些情况,虽然很难知道他们具体怎么做后训练。也可能是他们的后训练流程一直在变,某些模型上刚好弄错了几件事,某个东西调得太高,弄出一些大家很不喜欢的怪癖。后训练很容易以一种基准测试看不出来的方式搞砸。

27:49Beren Millidge: 还有一点很基本:前沿实验室的数据都是从大数据公司买来的,中国公司也能从这些公司买到同样的数据,而且他们确实在买。很多人对此很恼火,但如果数据完全一样、又能蒸馏,跟上其实挺容易的。

04

实验室里的持续学习

自动化研究员靠人类反馈加练习环境训练出来;现在是把最近三个月的进展蒸馏回模型,所以有人觉得它是渐近的;环境能超出人做得到的范围。

28:06 · 第一批自动化 AI 研究员会怎么训练出来

28:10Dwarkesh Patel: 另一个问题是,第一批能自动化 AI 研发的模型到底会怎么训练出来。有一个玩具版本,就是 Ryan 说的:让 GPT-8 去造 GPT-3 大小的模型,专攻内循环类的挑战,比如打通需要持续学习的电子游戏,或者用最少的算力把损失降到某个值。John,你说过实际情况可能不是这样。

28:47John Schulman: 我们大概会把两件事结合起来:从人类反馈中学习,吸收研究者的品味;再造大量练习环境,里面是多步骤的研究项目。实践中人们会把两者结合起来,每一轮修补上一轮看起来最糟的地方。研究者会大量使用 AI,注意到它们有一些固定的弱点,然后要么收集人类反馈、要么搭新环境去修补。

29:30Charlie O'Neill: 有一种思路是看沿着模型谱系往回退多少,再从那里开始自我对弈。极端情况下,你只给它一块 GPU、大概再给些神经网络,说:你自己琢磨怎么训练一个模型来做这些任务。现在的做法是走到谱系的最前沿,说:这是 Anthropic 最近几个月在自己训练流程里发现的 bug,把它们变成环境。你得在前沿上训练、在前沿上变强,所以谱系之前的历史自然全部锁定。你也可以想象退回到 GRPO 之前,搭一些环境,让它去发现训练模型的最佳强化学习方式,然后再往回退、再往回退。但我觉得算力还会紧张到让大家继续守在前沿,基本上就是把上个版本以来发现的 bug 和改进做个对比,变成训练环境。

30:30Dwarkesh Patel: 这对模型代际之间拿到不陈旧的新数据也很好。说到底,这基本就是实验室内部的持续学习:通过环境和 RLHF 这类工作,把最近三个月的 AI 研究进展蒸馏回模型里。

30:41Charlie O'Neill: 而且它确实就是在蒸馏。这也许就是为什么我们有些人觉得它是渐近的。你总是在追最近三个月的进展,这些进展当然有 AI 的贡献,但回路里仍然有人,你只是在一点一点地逼近人类研究者发现的、能做到的东西。

31:04Beren Millidge: 不过我要说一点:只在轨迹上蒸馏,显然永远超不过轨迹本身。但环境可以远远超出人能做到的范围:设计一个没有人能解开的环境很容易,AI 照样可以去试着解。这才是超越人类 AI 研究的路径。尤其在 AI 研究里,定义目标非常容易:比如要求损失降到 1.3,现在没有人做得到。这是一个极其可度量、可验证的任务。

31:38Dwarkesh Patel: 或者 nanochat 速通,但比任何人类速通者都快;或者造一个 1 亿参数、能打通《我的世界》的模型。这也许太容易了,那就打通一个复杂得多的游戏。

31:45Charlie O'Neill: 1 亿参数的模型打通《我的世界》,我们居然说这太容易了,这不疯狂吗?五年前这么说试试。

31:52John Schulman: 不过很多研究并不完全是那样,在一个定义清楚的目标上爬坡。更像是:我们有个直觉,觉得模型应该在某方面更好;也有个算法的想法,似乎往这个方向走了一点;那就设计一个任务,专门用来看这个方法有没有生命迹象。有的话,再一步步做出更真实的任务版本。你并没有直接优化你最终关心的目标,或者实际生产中的目标,而是把目标稍微放松一点:在真实度上先放一放,找到真能用的方法,等方法成熟一点再回到真实度上。还有一类研究更偏向解释现象、发展理论。机器学习里我们很少有那么有预测力的数学理论,但有很多更非正式的理论。

33:27Beren Millidge: 想必模型会在所有这些任务的混合上训练:有的非常容易验证,有的用 LLM 当评审,或者直接问人「这看起来合理吗」。希望是这些都能泛化到难得多、更模糊、更含糊的任务上。某种程度上大概会的。至于能不能泛化到让这个循环自己闭合、完全不需要人,就不清楚了。

05

在模拟里学会学习,够不够

实验室的赌注,是在海量模拟环境里练出能坚持、能分轻重的智能体,按领域一个个补;样本效率低时只能靠模拟,和真人实时互动很难模拟。

33:51 · 实验室的赌注:在模拟环境里练长任务

33:52Dwarkesh Patel: 退一步说,我理解的 AI 研究接下来的计划是这样,你们看对不对:在几百个领域、几百万个各不相同的环境里放大可验证奖励的强化学习训练。最后出来的是一个学会了一些基本技能的智能体:能坚持,能对信息和上下文分轻重,最终能端到端地和其他智能体配合。这样的智能体在上下文里会非常高效地学习。最后出来的东西,基本上就像一个能工作一周或一个月的即插即用远程员工。第一,你们同意这就是实验室在下的赌注吗?第二,这够不够:在数据中心的模拟环境里学会怎么学习,然后部署到真实世界,却并不真正从真实部署中学习,只从模拟环境里学到这些元技能?

35:10Charlie O'Neill: 现在很难分清实验室的精力有多少用在直接做 RSI 上,又有多少用在做通用智能的模型、拿去部署赚收入、好为下一次大训练出钱。对后者来说,是的,这大概就是他们下的注,过去几年环境往哪走的模式非常清楚。Anthropic 的环境谱系就是最清楚的例子:先只做编程,把编程做到非常非常好,因为互联网上能拿来造环境的数据和他们自己的内部材料,在编程上最现成;然后借着编程练出来的任务时长去泛化。接下来是金融,强化学习训练里塞进海量 Excel 数据之类;再往后是 PowerPoint,是整个工作经济的长尾。这招看起来很管用,很多其他实验室,连开源实验室,现在都意识到这是对的赌注。

36:09Dwarkesh Patel: 但这意味着什么?我请 Dario 上节目时问过他:如果你真的预期模型会像人一样在工作中学习,为什么还要把用 PowerPoint 之类的技能事先练进去?难道不该指望它部署后自己学会吗?有几种解释。一种是我们预期模型很快能做到,但还没做到,那何不把这些技能先摊进训练里;另一种是,我们根本没把重心放在让它擅长广泛部署的工作上,我们只想让它擅长 RSI,这只是一种赚收入的办法,好把钱投回一个真正擅长 RSI 研发的模型;等奇点发生,出来的东西自然会擅长当前模型卡住的所有事。John,如果路径是靠这种泛化,模型里为什么有这么多针对具体任务的知识,你怎么看?

37:09John Schulman: 如果模型在上下文里学习的能力足够好,理论上就不需要专门训练它做金融:它能当场把书都读了,弄清楚在相应法域里怎么做每件事。但也可以说,做大量领域训练只是为了让它更高效:就算它聪明到能当场弄明白,你可能仍然想用强化学习把这些直觉练进权重里,让它运行时更高效。实践中,模型供应商确实在一个领域一个领域地推进,在价值最高的领域加强模型。我会说这是模型变好这么多的原因之一:供应商已经覆盖了很多高价值领域和最常见的技能。

38:11Beren Millidge: 还有一点,两件事同时做其实也没那么贵。模型非常大,参数上完全负担得起什么都学。而且很可能有迁移:即使金融本身和 RSI 无关,学会判断什么重要、怎么有品味、怎么做长任务,这种通用的元学习可能是能泛化的。世界上 RSI 的数据也没那么多,很难生成,要花很大力气。所以能把其他数据摊进来,就能从中得到一些迁移。你本来就有海量算力和海量参数空间,何不一起做,更何况还有卖模型这个直接的商业目的?

38:55John Schulman: 我再补充一点:现在这种从模拟到现实的范式,是不是会永远占主导,是个问题。先看真实任务是什么样,再搭一堆能在数据中心里模拟的环境,在上面做强化学习。这显然非常成功,但弱点也很多,因为很多事就是很难模拟,尤其是要和一群人实时互动的事。

39:31Beren Millidge: 我认为在样本效率低的时候,从模拟到现实必然是主导框架,因为现在需要和人进行成千上万次互动,没有人会坐在那里当强化学习训练回路的一环。所以现在只能模拟,才能拿到所需的样本。但如果样本效率大幅提升,从部署中学习就会占大得多的比重。

39:56John Schulman: 不过也有别的办法。可以做异策略(off-policy)学习:把所有轨迹拿过来,不用把一切重新模拟一遍,也可能从中学到东西。

06

累积型任务比律师助理好做

部署数据已经一代代回流进模型;RSI 是累积型任务,发现一个守住一个;律所这种分布一直在变的工作反而难。品味也许能从短片段里学出来。

41:19 · 一半的算力没让模型变好

41:27Dwarkesh Patel: 我想多问几句,因为这很奇怪:50% 的算力花在推理上,却没有直接帮模型变好。数字心智最终应有的一大优势是:一个人一辈子有 50 年的真实经验,而一个模型通过它所有的实例,会在经济中各种有价值的工作里积累几百万年的部署经验。可现在,这些数据在任何实质意义上都没有帮它变好。模型最终应该能从这些数据中学习,这太显然了。一旦能做到,就会出现一种近乎「广泛部署式的智能爆炸」,因为模型在所有部署实例上吸收了海量信息。这种蜂巢思维的疯狂事情,你们预计什么时候开始?

42:14Beren Millidge: 我觉得大体上,在很基础的层面,这已经在发生了,只是发生在下一代模型身上。现在你显然可以把部署数据放进未来模型的预训练或中训练里,尤其是经过某种过滤、评判、标注或合成之后。

42:29Dwarkesh Patel: 你觉得这能解释多少代际之间的进步?

42:32Beren Millidge: 我觉得能解释相当一部分。我不知道实验室是不是这么做,因为他们说自己不拿用户数据训练。但中国公司百分之百在这么做,他们肯定拿到了这个优势。这基本上就是蒸馏:拿来模型,调用模型拿到一部分部署数据,再用它训练下一代模型。他们当然也可以对自己的模型这么做,完全没有理由不做。

42:56Charlie O'Neill: 完全同意。把视角拉得足够远,这肯定在发生。我们都在想象的持续学习圣杯,是一个非常自然的实时循环:单个模型获得一段经验,当场实时更新、从中学习。放大到那个粒度,很多东西都会坏掉。但大实验室在做,闭源模型在做。用开源模型的人也有了早期迹象,而且节奏快得多。Composer 大概是个好例子,Harvey 在法律智能体上也在做同样的事。你有某个模型,从这项任务的数据里、从用户在抱怨什么、从你想办法从具体部署中提取的所有反馈里,造出非常具体的环境。很多这类公司相对大实验室的优势,正是它们能把这些数据用得非常非常好。然后造环境,对 Kimi K3 做一次大的后训练,部署出去,也许还做些在线学习,就像 Composer 那样,基本上长时间地跑 REINFORCE。回路里仍然有人:有人在说「这些是我们在乎的信号,这是我们怎么用手头数据造环境」。节奏也许比你想的慢,但它真的在发生,这个循环最终会越来越快。

44:20Dwarkesh Patel: Composer 这件事有意思,在 Cursor 里,人们对模型建议的下一段补全按不按 Tab,基于这个,Composer 每天都在更好地预测下一个……

44:29Charlie O'Neill: 那是以前的 Tab 模型。他们其实不只对 Tab 模型,对真正的生成模型也做了同样的事。难点在于,做在线强化学习时没有「组」:只有一个用户说了一件事,你只得到一次采样,所以方差很大,得想办法压下来。Cursor 的模糊解法是:「我们有很好的启发式方法,能估计这次回答比平均好多少、或者差多少。」然后做一次大的 REINFORCE 更新。至于怎么判断有没有变差,他们的办法是:如果在 CursorBench 上变好了,就每 5 小时部署一个新模型;没变好,就把那个版本扔掉。

45:09John Schulman: 我觉得最大的问题其实是,对自然数据不知道奖励函数该是什么。如果用某种表面信号,比如用户有没有接受这次修改,就可能被模型钻空子。

45:24 · 累积型任务,以及为什么 RSI 比当律师助理容易

45:25Dwarkesh Patel: 可这对从模拟到现实不是更大的问题吗?任务越长,越难在数据中心里模拟。在我看来,就连编程也已经到了这个地步:没有哪个长达一年的编程任务,最后不需要和客户沟通、和公司或用户打交道。想想我们希望 AI 能做的全部事情,超级智能最终应该能经营一家企业,或者新创一家企业并让它盈利,或者在市场上靠日内交易赚钱,或者打赢一场官司。这些都很难在数据中心里模拟,和真实世界打交道本来就是学习的一部分。也许靠从模拟到现实的迁移就能学会;但也可能必须从这类互动中更新权重才能变好。如果迁移不够强、确实需要更新权重,那模型样本效率低,也许就是更深层的问题。我之所以好奇,是因为默认情况下,我看不出未来 10 年怎么会不出现某种疯狂的递归自我改进。唯一可能让它不发生的原因是:就更新权重的样本效率而言,模型似乎远远落后于人类。比较一个人从出生到成年看到的数据量,和一个模型从冷启动到训练完成看到的数据量,模型很可能落后人类百万倍。所以,第一,模拟环境能不能很好地迁移到我们想让 AI 在真实世界里做的那种极长、极复杂的真实事情?第二,如果不能,模型样本效率低这件事,是不是就会反咬我们一口?

47:25Charlie O'Neill: 我会这样区分两类任务,一类模型会做好,一类模型会一直吃力:看任务是累积型的,还是分布一直在变、你得不断重新学习、把很多事重新掂量一遍。RSI 可能就是累积型的。理论上,一个不到一百万 token 的 Python 文件,就能从零训练出一个具备递归自我改进能力的模型。你做出的每个发现,都是一条守住的线。如果做 RSI 真的不需要再发现一种新的注意力变体之类,那么一旦发现了注意力、发现了混合专家、发现了 GRPO,就把它加进训练流程,它就一直在那里。一个好例子是 OpenAI 训练 5.6 Sol、5.6 Terra,或者它告诉我们的无论哪个:它不用回去重新发现注意力,基本上就是调用一堆脚本,比如 pre-training.sh 和 post-training.sh,就完成了。这就是累积型任务。而真实世界,也是大家这么在意持续学习的原因,并不是累积型任务。想象一家律所里有个智能体当律师助理,这是分布变化极大的场景:你得把公司里所有重要人物之间的关系都装进上下文,而这些关系还一直在变;还有各种不成文的做事方式、去哪里找信息等等。这远不像 RSI 那样是干净的累积型任务。我觉得任务之间会出现这种分化。而如果实验室意识到这一点,并且确实相信 RSI 是累积型的,也就是不需要回头发现某种全新的架构,那也许越来越多的精力和算力会集中到它上面,而不是其他任务。

48:56Dwarkesh Patel: RSI 恰好比当律师助理还容易,这真是太不幸了。

49:08John Schulman: 我会说,今天的模型在很多方面都比人弱。有些弱点也许和某个区间里的样本效率有关。在有些区间,模型的样本效率很高,比如在上下文里学习;但在某个中等长度的区间,它们可能不那么高效,因为人类能以比模型更高效的方式更新某种「权重」。某些区间里样本效率更低,可能是弱点来源之一。但我认为还有一些弱点来源和这个完全不同。比如思维多样性比人低,或者在某些长期判断上很差。我觉得人们说的品味,很大一部分是知道什么样的做法从长远看行得通,而且人们已经意识到它行得通。不是全部,但品味有一部分是这样。尤其是软件工程,很多品味就是:什么样的系统在这个项目的长期里可维护、运转良好?模型有各种各样的弱点在限制 RSI 和其他事情,有的和样本效率有关,有的无关。

50:40Charlie O'Neill: 也许可以做个思想实验。假设你能给模型一个一万亿 token 的上下文窗口,或者够装下你在做 RLHF 之前所有经验的长度,它把这些经验都放在上下文里,而且样本效率和上下文学习能力跟在一百万 token 时一样。你觉得品味就解决了吗?它能做出和你一样的判断吗?还是说,除了更长的上下文窗口、同样的样本效率之外,还缺了某种根本的东西?

51:07John Schulman: 它得被训练成能从那样的上下文里学习:要么被训练成能从上下文中得出正确的更新,要么得能泛化过去。

51:26Beren Millidge: 而且你还需要那么长的数据去训练它用长上下文。就算理论上能有一万亿的上下文,也需要一万亿长度的数据来训练。现在上下文是 1 万,你没法直接塞进 100 万。但理论上,我觉得可以。这其实归结为:品味在多大程度上能从更短的片段里元学习出来。我觉得没有明显的理由说它需要特别长的片段,因为人类不知怎么就在没经历多少长片段的情况下形成了品味。我们活不到一万岁,发展得还挺快。想想读博,从博一到博士毕业或博士后,大概五年,总共也就做 10 到 30 个研究项目,但他们不知怎么就从一串相对短的小事里很快形成了品味。理论上可以这样形成。AI 显然会有多得多的经验来形成品味、元学习品味。问题是这能在多大程度上泛化到真正的长任务上,我觉得这目前完全没有解决,我们不知道。

07

持续学习在细处会崩

企业不愿让供应商从自己的部署中学,压力推向可插拔模块;放大看能用,落到一个模型反复小更新,就出现灾难性遗忘,最后只能从零重训。

52:22 · 持续学习在微观层面会崩

52:30Dwarkesh Patel: 回到这个问题,最终应该会有一个阶段,AI 能从每一个部署实例中学到大量东西。现在可以说,模型确实通过某种模糊的元过程从部署中变好,但我觉得这个反馈回路很弱。你们看得到这种非常快速的蜂巢式学习吗?如果看得到,它具体怎么发生?

52:54John Schulman: 我会说,能不能出现一个从所有部署经验中学习的蜂巢思维,很大程度上是激励问题,而不是技术问题。企业不会愿意让模型供应商从自己的全部部署中学习,因为那可能削弱它们的业务优势。

53:21Charlie O'Neill: 我觉得经济压力会推动的,不一定是对一个大的共享模型更新权重,而是可以插拔替换的模块。最明显的例子是 LoRA,但也可能是别的东西。有很多工作在试图把任意长度的上下文塞进固定大小,就是那些线性注意力的研究。还有 cartridges,本质上是被训练得高度压缩、能装下大量信息的 KV 缓存。如果这些东西插进模型、又不真正改变底层的基础模型,企业也许愿意接受。从数据中实时学习有很多种版本,后面这几种其实帮不到大实验室,因为它们只是些模块。但我觉得经济压力会逼实验室先走这条路,然后才能开始……

54:13Beren Millidge: 什么经济压力呢?我觉得就算你有一堆 cartridges 或 LoRA 之类的,你照样可以把所有轨迹拿来,塞进下一代模型的预训练里。

54:20Charlie O'Neill: 对,这可能是大实验室得到的一种更间接的学习,对他们显然仍然很有价值。但我想象不出一个世界,一开始就是「我们直接用拿到的全部原始数据,训练这一个大模型」。

54:43Beren Millidge: 不会,我觉得肯定会分阶段走,因为那种设想是假定会有一个不连续的时刻,我们突然解决了持续更新权重的问题。实际上,更可能的是:cartridges 之类的东西让你在各个部署里做专门化,然后你生成轨迹,放进模型,三个月后推出一个在这方面更好的模型;再专门化,再整合,最终这个跳跃越来越快。从每三个月发布一个模型,变成每周、每天、每小时,到那时我们基本上就解决了。

55:02Charlie O'Neill: 这点说得好,因为你问的是当前范式离能做到这件事还有多远。我们做过一些这方面的研究,其他人也做了很多。在非常大的规模上,噪声被充分平均掉、批次足够大时,把数据放进中训练、自己造环境这个外循环,确实能在某种持续学习的模式下奏效。但问题是,一旦放大到微观层面:我只有一个模型,想为某家律所一遍遍地更新它,而且是用相对少量的数据非常连续地做,所有方法都会有点崩。如果只用成功的轨迹对模型做监督微调,不管异策略还是同策略,在这种高度迭代的模式下,做几百次这样的微更新之后,你会看到灾难性遗忘:很早之前在基础模型之上学到的信息被忘掉,通用能力也在退化。同策略蒸馏似乎能把这个期限往后推一点,但最终还是败给同样的问题。强化学习擅长把能力练进去,却不太擅长把知识放进去,就是那种非常明确的知识:「哦,这家律所里这个人是这么做这件事的,这是我们发现的一个非常具体的流程。」要用强化学习把知识放进去,就得投入大量算力去造合适的环境。

55:46Dwarkesh Patel: 你觉得这里的根本问题,为什么其他技能会变差、会遗忘,根本上是容量问题,还是技术问题?

56:30Charlie O'Neill: 两者都有一点。我觉得监督微调、甚至同策略蒸馏,破坏性都可能太大。强化学习好就好在它对模型的改动非常非常小,这方面有很多证据。它只是在一个非常小的损失谷底里微调一下,把模型挪到对的位置。但这也限制了强化学习能做什么、到底能把模型改多少。

57:08Beren Millidge: 我认为主要是技术问题,肯定不是容量不够。如果你有某个模型和所有这些数据,拿一个大小完全一样的模型,把这些东西都放进中训练、从零预训练一遍,它会更好。我觉得今天很多事情就是这么做的。有一个很实在的瓶颈,让我们没法永远继续训练同一个模型,而是得把旧模型的数据全拿来、从零训练一个新模型。这正是 Charlie 说的:可塑性下降和灾难性遗忘的某种组合。如果天真地在分布变化的数据上训练,因为你一边训一边加新数据,就会搅乱数据分布,旧的东西就被忘掉了。我们其实没有好办法阻止这种事发生。

57:46Dwarkesh Patel: 所以在极限情况下,瓶颈也许就是得用这些新信息从零重训模型。

57:58Beren Millidge: 对,这当然非常昂贵,从零训练一个模型很贵。如果有了持续学习,也许最终你永远不用训练新模型,就一个模型,它一直在学、一直在扩展。

58:06Charlie O'Neill: 问题就在这里。我觉得我们已经把需要从零做的部分往后推了。现在完全可以拿预训练好的基础模型,在上面持续地做很好的中训练,再加上从训练后期不同检查点出发的强化学习。这更像持续学习了,但肯定不是拿最新的模型、做几次非常小的更新、一轮轮迭代还什么都不丢。

58:42Dwarkesh Patel: 抱歉,我有点糊涂,这不就是训练时实际发生的事吗?后训练时,模型已经经过了大量训练,然后你把一个进一步做过强化学习的分支蒸馏进去。这不就是这么回事吗?

58:44Charlie O'Neill: 但那仍然是在足够大的规模上,我觉得很多噪声被平均掉了,而且你不是只盯着一个分布,就像 Beren 说的,那才是问题所在。

58:59Dwarkesh Patel: 可到了最终阶段,会有几十亿个部署实例,你同时从所有实例中学习,希望这样也能把噪声平均掉。

59:13Beren Millidge: 也许在那个规模上可以。就像 Charlie 说的,持续中训练肯定能做很长时间,也可以回退到某个检查点、给它新的中训练数据。但同时,这没法无限做下去。如果一直持续训练同一个基础模型,它到某个点就会停在渐近线上,学不进新东西了。这就是为什么大家最后还是会训练新的基础模型,不然一直在同一个基础模型上做中训练就好了。

08

信号从哪里来

环境梯子每级越来越难爬,前沿需要的比特连整个世界都给不够;小规模实验里数据约贡献 12.0 倍效率、架构约 3.7 倍,但架构能打开新区间。

1:00:33 · 信号从哪里来

1:00:38Dwarkesh Patel: 我们来聊聊数据。我一直很感兴趣:AI 的进步有多少只是由数据的进步来解释的?这并不意味着它一定难以自动化,那是另一个问题。有没有某种数据分布,用当前的架构在上面训练,就能得到一个在每个领域都完全碾压人类专家的超级智能?

1:01:00Charlie O'Neill: 说的是预训练加后训练数据,也包括环境吗?我觉得这种东西的存在是显然的,问题只在于我们能不能造出通往那里的正确环境。

1:01:08Beren Millidge: 在最平凡的情况下,我们可以直接训练它输出那个能训练出真正超级智能的 Python 文件,把它背进权重里。

1:01:15Charlie O'Neill: 对,大概有一架强化学习环境的梯子,是有可能搭出来的,能让你得到一个至少和人类研究者一样好的 AI 研究者。但爬每一级台阶所需的力气差不多是指数增长的。这两件事之间的取舍,决定了我们多快爬到最后那一级、让它比人强。我觉得这点相当清楚。在搭建强化学习环境上,我们还处于相对早期,利用了很多不对称性来造好环境。其中一种我们之前聊过:有些环境倒着走比顺着走容易。意思是,很容易定义一个复杂的数据生成过程,把它作为隐藏变量不让模型看到;你可以生成任意复杂的环境,而模型得花大量省不掉的 token、做大量省不掉的功,才能弄清那个数据生成过程是什么。还有往环境里注入真实世界信息的不对称:Anthropic 靠几万个人加 LLM 一起找到一个 bug,把它变成一个非常精巧的环境,理论上一个 LLM 在几百万 token 之内就能找到。我们在挑拣所有这些不对称性,指望任务时长上的泛化。但我觉得到某个点就会收益递减:一开始造这些环境、想出这些环境本身就越来越难,因为你不一定总能找到倒着走更容易的过程,得真正坐下来和人一起搭出一个时长足够长的东西,这会是非常复杂的工作。然后还有智能体实际完成这些任务所需的算力和时间瓶颈。我觉得你会开始看到这条曲线变平。

1:03:07John Schulman: 我看到有人拿 Talkie 模型,它只用 1930 年以前的数据训练过,在现代编程智能体的数据上做微调,结果在 SWE-bench 上比 Claude 3 Opus 还好。一个对代码一无所知的模型,用中等量的数据一微调,当编程智能体就比一个大得多的预训练模型表现更好,这挺疯狂的。这说明一旦有了正确的专家行为示例,把它复制到一个相对弱的模型里,其实出奇地容易。

1:03:53Charlie O'Neill: 但有个反例:最近一篇论文训练了一个模型,数学只学到五年级,英语是小学水平,算是个还行的语言模型。他们想用强化学习让它做高中后期和大学的数学,差距实在太大,完全爬不上去。但如果一级级来,7 年级数学、8 年级数学……显然能爬到 12 年级。还是那个问题:梯子上每级之间的距离有多大,造起来有多难?

1:04:18Beren Millidge: 这又回到了强化学习的信号问题。强化学习现在不太会探索:如果模型在 128 次采样里都做不对,就几乎拿不到能推动进步的信号。这就是为什么强化学习需要循序渐进的课程,而预训练不需要,那对预训练根本不是问题。

1:04:35Charlie O'Neill: 再说一次,预训练数据和后训练数据不一样。我想随着发展,人的参与会越来越少,但这改变不了一个事实:你受限于能从真实世界提取多少信号。世界上信号很多,没错:有人在做表格,有人在做法律工作,诸如此类。但在模型现在所处的能力前沿,世界上到底有多少比特真的和提升模型能力相关?有多少新的数学题被解出来,正好超出当前模型的能力范围?有多少新的编程问题被提出或解决,超出了当前模型的能力范围?我觉得这就是收益递减出现的原因:即使整个世界,也没有给你那些能把你推进下一个能力盆地的有用比特。

1:05:30Beren Millidge: 完全同意。关键就是信号从哪里来。预训练的信号本来就在 Common Crawl 里,对预训练关心的任务来说,问题根本不是拿不到信号,而是过滤掉所有噪声,这是一个相当能自动化的过程。但随着模型变好,进入中训练和后训练,信号在我们手头的原始数据里根本不存在,怎么过滤都得不到。Common Crawl 里并没有藏着一个千禧年大奖难题的证明,等你过滤出来。到了这一步,就得从别的地方拿比特:要么直接找人,让他们写出推理过程;要么造环境,由人来决定该造什么环境、这些环境的目标是什么;要么用部署中存在的人类数据做某种训练。比特总得从某个地方来。

1:06:10 · 数据和架构,各贡献多少

1:06:15Dwarkesh Patel: 有个问题是,预训练的进步有多少是数据推动的。我和普林斯顿的学生 Jerry Han 做过一个调查:把 2019 年到现在的所有训练配方,和 2019 年到现在的所有数据集两两配对来训练。拿 GPT-2 在最新的数据集,比如 Ultra-FineWeb 上训练;拿最新的开源训练配方 Delphi,在 Pile 这种老数据集上训练,整个网格都做一遍,看达到某个能力水平需要少多少算力。结果看起来,数据大约解释了 12.0 倍的算力效率提升,架构改进大约解释了 3.7 倍,这是在很小的规模上。如果这在大规模上也成立,也就是预训练算力效率的提升大部分来自更好的数据,那它还能持续多久?还能不断加强过滤、造越来越多的合成数据吗?你们对这种预训练进步还能持续多少有感觉吗?

1:07:23Charlie O'Neill: 我的先验是,低垂的果子已经摘得差不多了。互联网是一整块给了我们,它并不会以同样的速度增长,互联网上有用的东西更不会。我们大概还有一堆 0.1% 的损失下降可以争取,但肯定没有过去那么多了。不过也很有意思,你们发现两者合起来是 33 倍的累积提升。好像是 Epoch 还是谁估计过,自 2019 年以来每年 3 倍,那就意味着大约 3 的 7 次方,超过 2000 倍的提升。那缺的那 100 倍左右是从哪来的?这大概能很好地提示有多少来自后训练。

1:08:08Dwarkesh Patel: 我觉得解释只能是,很多算力效率的提升依赖规模,而我们是在极小的规模上做的。这就引出一个问题:数据带来的效率提升和算法带来的效率提升,哪个更依赖规模?我们没有足够的算力去研究这个问题。

1:08:36Charlie O'Neill: 单纯从理论上说,架构对规模的依赖相当清楚,可以拟合出一条直线。而把预训练、中训练和后训练数据合在一起,我完全不知道该怎么做。

1:08:49Beren Millidge: 说来好笑,我倒觉得规模越大,数据越重要,架构更像是一次性的东西。只说提升了百分之几的效率,有点误导,因为架构的作用是让你进入一个旧架构到不了的、质上全新的区间,而在那个区间里,决定结果的主要是数据。如果我们连 GQA 都没有,整天做全注意力,做一百万的上下文就会贵得离谱,也就永远用不上那些真正一百万长度的数据,得不到这些能力。可如果你只是天真地看「在 2K 上下文下它有多大作用」,这时架构什么都没解锁,数据就会显得比它实际上更重要。我不确定这些提升真的是这样简单相乘的。另外,我们现在很多中训练和后训练数据其实是规模越大越好用,因为很多是那种超长时长的环境数据,真要大模型才用得上。你拿一个 1 亿参数的模型去学 SWE-bench 的轨迹,它什么也学不到。

1:10:10Charlie O'Neill: 现在比较起来也更难了,因为很多架构改动,比如 Kimi 或 DeepSeek 的,不只是为了降低预训练损失,而是考虑到模型在真实世界里怎么被使用。DeepSeek 模型里某种压缩注意力,是为了推理效率,并不一定是在根本的取舍上做了改进。

09

模型还会不会变大

推理效率优先,未来几年参数可能停在一个平台;Schulman 预计模型仍会变大,但幅度取决于规模定律;漂亮的直线背后藏着大量复杂性。

1:10:23 · 模型还会不会继续变大

1:10:32Dwarkesh Patel: 为了理解未来,我好奇的一个问题是:随着我们进入更重强化学习的阶段,参数规模会怎么走。看开源架构,可以看到参数增长多快,前沿开源模型大概每年 2 倍。就算前沿闭源模型的激活参数有 1000 亿或 2000 亿,你们觉得还会每年翻倍吗?还是说,既然进入了强化学习阶段,采样时也想省算力……也可能有门槛效应,容量够了之后,再任意增加参数就没那么重要了。你们觉得 2030 年一个前沿模型会有多少激活参数?

1:11:11Charlie O'Neill: 我觉得未来几年,因为我们太专注于做越来越长的强化学习采样,推理效率非常重要,而模型在这方面的能力还不一定饱和,瓶颈仍然是环境,所以可能会看到一点平台期。我感觉 Mythos 和 GPT 模型比大家说的 10 万亿参数级别小得多;单纯拿它们和开源模型比一比,大概就能倒推出这个结论。未来几年,我不觉得参数量会大幅增长。不过这里要权衡的东西太多:你根据手头有多少预训练数据、再根据要训练的强化学习环境有多难来决定模型大小。理想情况是到达那个最优点,在你最难的环境上能拿到还不错的 pass@1;模型再做大就不划算,因为你只是在为超出需要的推理付钱。所以很多取决于 Mercor 和内部团队能多快把强化学习环境的复杂度做上去。

1:12:21John Schulman: 我预计模型会继续变大,只因为大家在扩大算力,GPU 也在变大。但具体变大多少,会以不那么显然的方式取决于规模定律。一点是,既然我们开始进入高质量预训练数据不够用的阶段,我觉得数据效率会比算力效率更能决定大家用什么架构,这可能影响你想把模型做得多稀疏。我也觉得我们对稀疏性理解得不太好:总参数和激活参数是两种不同的资源。稀疏度确实提高了一些,但不清楚会不会无限提高下去,可能存在某个最佳点。有一种说法是稀疏会让数据效率变差,因为同样的东西可能得在多个专家上各学一遍,不过这有争议。我不认为我们有足够好的规模定律理论,能真正理解稀疏为什么有帮助、帮多少、会不会在某个稀疏度上停住。如果横轴放的是数据而不是算力,就会得到另一组最优模型。

1:14:48Charlie O'Neill: 我也不觉得过去几年模型真的每年翻一倍。人们训练 1 万亿参数的模型至少有好几年了,甚至有个开源的叫 Falcon。Periodic Labs 的 Liam 好像昨天在推特上说,一个早期实验就是训练一个非常非常稀疏的 1 万亿参数模型,John 说那是他们去 OpenAI 之前在 Google 做的 Switch Transformer。它因为太稀疏,知识很强,推理却很糟。感觉我们在 1000 亿到 2 万亿参数这个区间已经待了一阵子,肯定不是那种漂亮的线性增长。

1:15:26Beren Millidge: 我觉得有两点。就像 Charlie 说的,推理效率对强化学习采样极其重要,这会把激活参数压下来很多。总参数则很大程度上取决于硬件:要能部署几万亿参数的模型,需要非常高的内存带宽和显存。现在大家还在大量用 H100 之类的,等大家换到 GB 系列、再到 Vera Rubin,我们就更有能力在更大规模上部署、做大规模的强化学习推理。数据这个问题也有意思,因为单纯说,大模型在实际数据点上的样本效率高得多。就算模型没饱和,做大还是更好,因为大模型泛化更好,同样多的数据能降到更低的损失。现在我觉得数据很多,不是约束,算力才是,所以我们在用推理效率很高的小模型。但如果算力不再是瓶颈,可能又会回到更大的模型:没饱和,但因为大得多而有这种泛化能力。

1:16:26Dwarkesh Patel: 如果只看最基本的 Chinchilla 规模定律,把参数拉满,达到同样损失所需的数据其实只减少一点点。就算参数趋于无穷,所需的数据大概也减少不到 10 倍,这是幂律的性质决定的。

1:16:44Beren Millidge: 但我们现在处在 Chinchilla 定律里数据远远过多的那一侧,按 Chinchilla 来说我们在过度训练模型。所以随着数据用完,我们很容易回到 Chinchilla 的最优点,甚至稍微偏向训练不足的一侧。这其实取决于训练算力和推理算力的比例:如果卡在数据上而不是算力上,就该做大;卡在算力上,就该一直做小。还可以用计算机生成的合成数据,所以这是那种非常难预测的事。

1:17:22John Schulman: 我觉得人们当初花了那么久才弄清规模定律,部分原因是如果这些东西没全部做对,就得不到那么干净的关系。图上那些漂亮的直线,背后藏着大量复杂性:你得确保每个超参数都按正确的方式缩放,或者把优化器参数化得能随规模扩展,模型变大时不用改超参数。

1:17:47Charlie O'Neill: bug 也有自己干净的规模定律。比如 Kaplan 忘了余弦退火,或者我记得还没算嵌入参数,这把小模型上的估计搞乱了,因为嵌入参数在小模型里占的比例不小。

10

中训练干了八成,强化学习只调策略

中训练常把模型带到约 80% 的位置,强化学习用信噪比极高的一个比特调策略;得到的是任务时长上的泛化,不是跨领域的泛化。

1:18:03 · 强化学习为什么比预想的管用

1:18:06Dwarkesh Patel: 聊聊强化学习。一年前,很多人都在说强化学习放大到模型上不会特别成功。John,你写过一篇研究论文,指出模型做强化学习时每个回合只学到一个比特:「我答对了还是答错了?」今年早些时候我写了几篇博客说「情况比这还糟」,因为通过率低、模型很难答对的时候,它从一次强化学习回合里几乎什么都学不到。可我看今天的模型,它们显得相当聪明,这似乎就是放大强化学习的结果。Beren,你几周前写了篇文章试图解释这是怎么回事。强化学习为什么比人们天真以为的更成功?

1:19:01Beren Millidge: 我觉得强化学习的成功要归于好几件事。第一,中训练被有点低估了。我们看到的强化学习的成功,很大一部分其实来自非常非常好的中训练数据:本质上还是在做预训练,只不过是在合成的推理数据上,外加一些为强化学习热身的环境。这一步常常已经把模型带到最终强化学习检查点差不多 80% 的位置。强化学习在此基础上做的,本质上是微调策略。这就是它不需要你天真以为的那么多比特的原因之一:它不用从零学这些行为,只需要从这些回合里拿到几个比特,而这些你是拿得到的。第二,我在博客里指出,这些比特和普通预训练相比,信号极高,这就是为什么非要强化学习,而不是直接在成功的推理轨迹上做监督微调。不只是因为它们恰好是「怎么把答案做对」的比特;监督微调的轨迹里也有答案那个 token,那个比特也在。关键是强化学习的目标忽略了其他所有比特。做监督微调,你得去匹配所学模型产生的每一个推理 token,等于拿到了太多关于那个模型具体怎么推理的比特。而强化学习只拿那一个比特,信号就不会被模型其他所有比特的噪声淹没。这是训练中信噪比极其戏剧性的提升,所以强化学习按步数算效率高得惊人。

1:20:31Charlie O'Neill: 关于强化学习对模型做了什么,和中训练、监督微调之类相比,争论非常多。大家都说 pass@1 会上升,pass@256 会下降,非常罕见的正确推理会被压低,被来自简单推理的梯度信号盖过去。我觉得现在看强化学习,一个简单的角度是:如果你有足够的算力,采样一个足够大的组,让拿到一批正确答案的概率超过某个不可忽略的水平,这些答案就会被加权提升。顺着 Beren 的话说,中训练和更多的预训练,也就是 pass@1,强化学习的起点,是随预训练 token 数的对数增长的。

1:21:48Dwarkesh Patel: 我能问几个非常基础的问题吗?这个回答说得通,也许还有实证研究表明确实是这样。可我一看模型本身……我不知道发生了什么。也许你们能告诉我,过去一年 AI 进步的基础是什么。也许只是把那些本来就会正确思考的策略加权提升了。但从质上看,模型的能力强了太多。也许这两者并不矛盾,但按这种看法,强化学习的作用相对很小,这和模型实际上在质上获得的能力,怎么对得上?

1:22:08Beren Millidge: 我想指出的一点是,这并不一定意味着强化学习的作用小。就算只有几个比特、参数只改了一点点,对函数空间,也就是模型学到的从输入到输出的映射,实际影响仍然可能非常巨大。哪怕一个比特,也能大幅改变函数空间,它可以排除一半的假设空间,这是巨大的。我不认为从一个非常好的起点出发,只用少量比特、少量强化学习,就一定不会带来行为上的巨大变化。至少不一定。

1:22:42Charlie O'Neill: 我觉得归结为两件事。第一,大家本来都希望强化学习能把推理能力泛化到各个不同领域,我认为我们并没有得到这种横向泛化。只训练数学,不一定让你成为最好的程序员,你还是得在代码环境上做强化学习。但我们确实得到了时长上的泛化:模型学会了用更多 token、干得更久,还能在任务上持续推进。你可以在越来越长的环境上训练,再把它们放进一个全新的环境;是的,它们也许没有泛化出在那个环境里表现好所需的推理模式,但至少泛化出了在任务上坚持更久的能力,而这和成功相关。有篇论文叫 EdgeBench,显示模型能持续工作的时长每三个月翻一倍,这是泛化的明确证据。最后一种看法是,预训练里有「量子」这个概念:你看到一条非常平滑的预训练损失曲线,可仔细看模型内部,它在学大量非常离散的小任务,有很多涌现点,出现相变:原来没有归纳头,现在有了。这样的东西有几万、几百万,大概几亿个,全部平均起来,就得到一条非常平滑的损失曲线。某种程度上,强化学习也在发生类似的事。就像 Beren 说的,有一个非常慢的外循环:我们训练一个模型,对它做强化学习,然后在下一代模型的训练里,把一堆这样的合成推理轨迹放进中训练数据。我们在各种不同任务上一个个碰到这些「量子」,单看一个任务,可能像一次相变:在某个具体的金融或 Excel 任务上,通过率突然从 0.5% 跳到 90%。但把所有这些平均起来,再加上时长上的泛化,你就会说:「哇,模型在质上变好了。」

1:24:32Beren Millidge: 我觉得还有一点,强化学习确实会泛化一些:数学和代码之间、谜题和数学之间,都有一些迁移。另外,大家瞄准的环境数量多了太多。以前,比如两年前,你想让它做一件日常生活里做的事,实验室根本不在乎,不会为此训练模型。现在范围广得多,有大量环境专门瞄准各种具体的事。

11

第 37 手与单一文化

强化学习没有彻底扼杀创造力,模型曾一次找出好几个零日漏洞;但输出多样性下降,大家都蒸馏 Claude,开放权重模型越写越像。

1:24:54 · 第 37 手,以及单一文化

1:25:00Dwarkesh Patel: 前面我们聊到强化学习会造成熵坍缩,也就是把概率集中到基础模型本来就会的解法上,对策略的更新相对稀疏。但我觉得强化学习还有另一个故事,要回到 Atari 游戏,再到 AlphaGo 下出第 37 手,那步超级有创造力的棋。因为它从来不是用人类数据初始化的,所以能用人类根本没想到的方式思考,想出极有创造力的解法。你们觉得我们该不该预期,以及什么时候预期,在 LLM 上做强化学习会带来第 37 手那样的东西,那种因为智能是从零起步、甚至超越人类的极端创造力?

1:25:52Beren Millidge: 这里有几点。首先,AlphaGo 用的是蒙特卡洛树搜索,它的探索显然比普通的策略梯度多。但我也认为强化学习不一定会降低创造力。这显然是定性的,但看看 OpenAI 和 Hugging Face 那次事件,这些模型一次想出好几个零日漏洞,好从沙箱里逃出来。这显然已经是某种程度的第 37 手式创造力,而且是 LLM 靠一般的泛化能力就带来的。强化学习肯定没有把熵彻底摧毁,尤其是在长任务上。

1:26:23John Schulman: 人们说的创造力,有一种其实就是解难的搜索问题。第 37 手显然是一个例子,写一首满足一大堆不同约束的诗也是。只要为此训练,AI 显然会在这上面极其擅长。但还有另一个方面:强化学习之后,模型输出的多样性低了很多,还养成了各种口癖。模型看起来写作很好,可一做分布上的分析,你会发现它们一直在重复某些主题,一直用同样的人名。你得不到人类作者那样的多样性,你得到的是一种非常好的风格。所以我认为这种多样性确实被强化学习削减了很多。其实,既然前面聊到蒸馏,现在正在发生的一件事是:太多人在蒸馏,主要是从 Claude 蒸,结果所有开放权重的模型都写得和 Claude 一样,带着同样的口癖。一种单一文化正在浮现,这让我有点担心。

1:27:54Beren Millidge: 不过我还是认为,这不是强化学习这种方法本身的问题,蒸馏也一样。就算是蒸馏,你也只是在数据上训练。你的数据不够广,并不说明训练方法本身有什么错,那是数据的问题。我觉得强化学习的熵坍缩,比如说,很大程度上是因为环境多样性不够大时,模型去钻相当简单的验证器的空子。比如写作,想必是由某个评审模型打分的。评审有它自己的口癖,模型学会了钻评审的空子,所以才坍缩。但这其实是评审的问题,不是强化学习本身的问题。

12

时间线:两年到十年

即插即用的远程员工约 1 到 3 年;AI 研究者提效 10 倍,Schulman 说两年,O'Neill 说 5 到 10 年;超级智能 3 到 10 年。

1:28:32 · 快问快答:时间线

1:28:48Dwarkesh Patel: 好,超快速地预测一下未来,下面几个问题我要时间线。什么时候会有这样的模型,用户的感受是:你基本上可以把它当成即插即用的远程员工来雇,做各种白领工作?不只是编程,还有视频剪辑、法律、律师助理等等。它就是一个真正的远程员工,能完整操作电脑,真正做到一个月无缝学习和运转,执行需要和其他人打交道的复杂项目等等。一个人类员工一个月里能做的一切。

1:29:22Charlie O'Neill: 如果你强制它用浏览器之类的,而不是由公司把信息设置成程序可以直接访问的形式,也许要几年。但如果不是基于浏览器的,它能发 Slack 消息、能做这一切,我大概还是会说一年左右。

1:29:39Beren Millidge: 完全通用的话,我会说也许三年。但就像 Charlie 说的,会有很多人把自己的组织改造得更方便 AI 使用,所以在那之前就能完成 80% 到 90%。剩下的是一条长尾,各种零碎的事,人能做,模型要花挺长时间才能做到。这其实取决于我们多快能解决这种在线学习,以及能不能靠压缩上下文、给自己写文件之类的办法完成 80% 到 90%。这是我最大的不确定,我真的不知道。

1:30:14Charlie O'Neill: 举个它不擅长的例子:如果工作中我得冲某人发火才能拿到东西,或者得狠狠推某人一把才能把事办成,模型就是不会这么做,它会太客气。

1:30:27John Schulman: 我会说,人类远程员工的质量差异很大。如果你想在 Upwork 上雇人做软件项目,质量会参差不齐,常常很难让他们把活干好,或者把你给的反馈都听进去。我猜在某些情况下,AI 之前的这类服务,比现在的 AI 能给的还差。最后可能会有点复杂,因为某种程度上,一些质量不那么高的工作我们已经有了;但在某些更高质量的工作上,我们显然还没达到人类水平。不过我基本同意 Charlie 和 Beren,也许一年左右会有某个还算可以的版本。会有这种产品形态,它有些事做得非常好,有些事做得不太好,然后从那里继续改进。

1:31:33Charlie O'Neill: 我们总是根据那条很长的长尾挪动球门。我记得你以前举过报税的例子。今年,我就是直接告诉 Codex,去把我需要的所有东西拿来,发给会计。它得操作电脑,点开并下载一大串东西。它做到了,做得完美。这类事很多它已经能做了。

1:32:21Dwarkesh Patel: 好:给你们带来 10 倍的整体生产力提升。基本上,如果现在你做出一个突破要一年,到时候每个月就能做出一个。这是说你们作为 AI 研究者,推进 AI 研究的状态。

1:32:31Charlie O'Neill: 大概 5 到 10 年?我对远程员工的设想,是普通白领一个月的工作;超过两个月,情况就开始有点不一样了。

1:33:00John Schulman: 我会说两年。

1:33:09Beren Millidge: 其实我能理解,因为在编程这类事上,现在肯定已经超过 10 倍了。所以只要它能跑哪怕一两轮实验反馈,就已经非常了不起了。那样的话,AI 的进步就不会再卡在研究者跑小实验的能力上,而是卡在别的东西上。

1:33:46Dwarkesh Patel: 当然,但那些事会快 10 倍发生,这是件大事。它还会让下一件事,也就是 100 倍的加速,更早到来。

1:33:53Charlie O'Neill: 这一条我宁愿说得慢一点。关键在于我吸收信息、对下一个实验做出贝叶斯最优决策的能力。

1:34:05Beren Millidge: 我是假设这其中一部分能交给 AI。AI 在做决定上正变得还不错:跑了这个实验,拿到这个结果,再跑下一个实验。如果它能连续跑两三个实验而不崩,那其实就是很大的提升。

1:34:24Dwarkesh Patel: 好,最后一个问题。一个在所有能用电脑完成的工作领域都碾压顶级人类专家的 AI,不只是 AI 研究,而是所有认知工作;不只是短任务,就算一件事要花三年,AI 仍然比人做得好。这基本上就是超级智能(ASI)。

1:34:59John Schulman: 我会说 3 到 4 年。AI 研究显然得到了更多关注,它是比较难的事情之一,但投入的精力非常多;而且它对 AI 来说也不是最难的,因为涉及大量代码和数学,模型在这方面非常强。涉及三维、空间和物理的事,我觉得会慢一点。如果是机械工程之类、现在没得到最多关注的领域,可能要慢一点。

1:35:28Dwarkesh Patel: 但它也包括那些天生数据就比较少的领域,得现学现用,比如得成为 TSMC 里超越人类的工程师。

1:35:42John Schulman: 那你就得假设能给 AI 同样的入职培训材料,然后还得在更长时长的学习上解决点什么。

1:36:00Charlie O'Neill: 我会说 5 到 10 年。我觉得自动化 AI 研究基本上就和实现 ASI 一样难:世界上有太多事情,就算模型外面有某种记忆系统,就算上下文长度再长一点,就算你能自己去查资料、写笔记,从根本上说仍然需要比今天一百万 token 更大的上下文窗口。

1:36:23Beren Millidge: 我大致同意 5 年这个范围,至少对实验室正在专注的东西是这样。但我觉得会有一条长尾:理论上 AI 可以去学,但没人费心去做,也没分配算力。所以要比得过每一位人类专家,可能要更久。不过它不一定需要像人一样快地学会一个新领域,因为 AI 会比任何人都有多得多的经验。

1:36:51Dwarkesh Patel: 非常感谢各位。我觉得这种形式很好,能让不同的专家一起争论、辩论、讨论,很有收获。

判断收口延伸

Indigo 的结论

这是「能力是有界的指数」的一线从业者辩论版:RSI 大概率会来,但被定目标、持续学习、样本效率按住,不是快速起飞。速度问题仍没有定论。

需要记住的几件事

  1. 瓶颈从算力改判到定目标和持续学习:2036 年若一切如常,原因多半不是模型不够聪明,而是不知道该干什么、在部署中学不动。
  2. 对齐是最后一份工作:留给人最久的是定义目标、决定我们要什么;技术活能自动化,指定正确的目标短期内让不出去。
  3. 持续学习在细处会崩,只能重训;RSI 这种累积型任务,反而比分布一直在变的真实工作好做。
  4. 蒸馏是反集中化的力量:可蒸馏的能力守不住,价值上移到独有的真实部署数据,比如 Composer、Harvey 这样的垂直场景。
  5. 强化学习管用,是因为中训练已完成约 80%,强化学习只用一个高信噪比的比特调策略;它延长任务时长,不跨领域。

可回查的判断

判断谁说的何时见分晓证据多硬
通用的「即插即用远程员工」:一个月无缝学习、完整操作电脑、和人协作O'Neill、Millidge约 1 年(不强制用浏览器);完全通用约 3 年一手判断;Millidge 说组织会改造自己来适配 AI,在那之前先做到 80% 到 90%
AI 研究者获得 10 倍生产力提升Schulman(Millidge 认同)2 年一手判断;O'Neill 说要 5 到 10 年
超级智能:在所有能用电脑完成的领域碾压顶级人类专家,包括长任务Schulman 3 到 4 年;O'Neill 5 到 10 年;Millidge 大致 5 年3 到 10 年一手判断,分歧大
自动化 AI 研究基本上和实现超级智能一样难O'Neill(回应主持人追问)一手判断
模型能持续工作的时长每 3 个月翻一倍(EdgeBench)O'Neill 引述已发生二手引述,方向可核
小规模实验里,数据解释约 12.0 倍的预训练算力效率提升,架构约 3.7 倍Dwarkesh(与普林斯顿的 Jerry Han)已做一手实验,规模很小,放大后是否成立未知
未来几年参数量不会大涨,推理效率优先;前沿模型远小于传言的 10 万亿参数O'Neill(Schulman 预计模型仍会变大)未来几年一手判断

放回主线

证实

AI 能力是有界的指数:范式跃迁不来自爬山 迄今最强的一线辩论印证:几处「界」由三人主动坐实,瓶颈从算力改判成「该干什么、学不学得动」。

证实

你拥有的不是模型:价值上移到不可租用的东西 蒸馏反集中化、部署数据比环境重要、定目标是人最后的工作,三条都把价值推向独有、租不到的东西。

证实

可验证域能否泛化 自动研究和开放式科学之间的缝,是这条边界在训练一侧的精确版;强化学习只延长时长、不跨领域,是边界的机制解释。

证实

Nathan Lambert《有损自我改进》 O'Neill 讲的持续学习在细处崩溃和摩尔定律类比,是 Lambert「被摩擦按住」的密集一手工程证据。

冲突

Jakub Pachocki《An Alien Mind》 Pachocki 强烈预期 RSI;这三位讲的多是摩擦,时间线却从 3 到 4 年到 5 到 10 年不等,速度问题仍无定论。

补充

Furong Huang:自我改进的 agent 「自我改进要能证明后面的工作更好」,和这里「AI 能否自己定目标、长时间不跑偏」是同一个关键点。

补充

a16z + Gavin Baker《需求跑赢供给》 同周两面:a16z 从需求侧说供给不足,这里从工程侧说瓶颈在定目标和持续学习。

什么会让我改口

下一个不连续是在当前范式上做加法,而不是推倒重来,当前范式能自己把点连成线。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Video

AI researchers debate how close we are to recursive self-improvement

John Schulman, Beren Millidge, Charlie O'Neill · YouTube · 2026-09-12

Three frontline researchers debate how close recursive self-improvement is: it's coming, but held back by engineering friction such as setting goals and continual learning.

Part 1 of 12 · 0:01
What's missing isn't intelligence but the next discontinuity

If 2036 looks normal, the likely reason isn't models that aren't smart enough but stalled generalization and continual learning. Moore's straight line lived on discontinuities, and today's paradigm may not find the next one itself.

Breakdown · 12 steps

  1. 01

    0:01 – 7:08

    What's missing isn't intelligence but the next discontinuity

    If 2036 looks normal, the likely reason isn't models that aren't smart enough but stalled generalization and continual learning. Moore's straight line lived on discontinuities, and today's paradigm may not find the next one itself. Read this part →

  2. 02

    7:08 – 18:51

    Clear objectives mean 10x; setting them is the hard part

    With a clear objective, the current paradigm could speed up 10x; the right objective can't simply be thought up. The longest-lasting human job is defining the objective and deciding what we want. Read this part →

  3. 03

    18:51 – 28:10

    Distillation keeps anyone from pulling away

    Whatever RL teaches can be distilled cheaply, and the prompt distribution is the key. Frontier labs may have no edge in training environments; real deployment may matter more. Read this part →

  4. 04

    28:10 – 33:52

    Continual learning inside the lab

    Automated researchers will be trained with human feedback plus practice environments. Today labs distill the last three months of progress back into the model, which is why it can feel asymptotic; environments can go beyond what humans can do. Read this part →

  5. 05

    33:52 – 41:27

    Learning to learn in simulation: is it enough?

    The labs bet on training persistent agents that triage well across vast numbers of simulated environments, patching domain by domain. With low sample efficiency, simulation is the only option, and real-time human interaction is hard to simulate. Read this part →

  6. 06

    41:27 – 52:30

    Cumulative tasks are easier than a paralegal's

    Deployment data already flows back generation by generation. RSI is cumulative, each discovery held once found; law-firm work, whose distribution keeps shifting, is harder. Taste might be learned from short episodes. Read this part →

  7. 07

    52:30 – 1:00:38

    Continual learning breaks in the details

    Companies won't let providers learn from their deployments, pushing toward swappable modules. At scale the loop works; for one model updated again and again, catastrophic forgetting sets in and you end up retraining from scratch. Read this part →

  8. 08

    1:00:38 – 1:10:32

    Where the signal comes from

    Each rung of the environment ladder is harder to climb, and even the whole world lacks the bits the frontier needs. At small scale data explains about 12.0x of efficiency gains and architecture about 3.7x, though architecture opens new regimes. Read this part →

  9. 09

    1:10:32 – 1:18:06

    Will models keep getting bigger?

    Inference efficiency comes first, so parameters may plateau for a few years; Schulman expects growth anyway, depending on the scaling laws. The beautiful straight lines hide a lot of complexity. Read this part →

  10. 10

    1:18:06 – 1:25:00

    Mid-training does 80%; RL tunes the policy

    Mid-training often gets the model about 80% of the way; RL tunes the policy with one very high-signal bit. What it delivers is generalization over task length, not across domains. Read this part →

  11. 11

    1:25:00 – 1:28:48

    Move 37 and a monoculture

    RL hasn't killed creativity; models once found several zero-days at a time. But output diversity has dropped, and with everyone distilling Claude, open-weight models increasingly write alike. Read this part →

  12. 12

    1:28:48 – 1:37:01

    Timelines: two to ten years

    A drop-in remote worker in about 1 to 3 years; 10x for AI researchers in two years per Schulman, 5 to 10 per O'Neill; superintelligence in 3 to 10 years, with wide disagreement. Read this part →

Indigo's conclusion

The practitioners' debate version of “capability is a bounded exponential”: RSI will probably come, but objective-setting, continual learning and sample efficiency hold it back; no fast takeoff. The question of speed stays open.

How to read this A roundtable of three frontline training researchers, chosen, the host says, because they're at relatively open labs and can speak on the record. Nobody is selling a product; they undercut one another and often admit uncertainty, a rare signal of honesty. Overall, RSI is coming, but a pile of engineering friction is slowing it down.

What to remember

  1. The bottleneck moves from compute to setting objectives and continual learning: if 2036 looks normal, the likely reason isn't models that aren't smart enough but not knowing what to do and failing to learn in deployment.
  2. Alignment is the final job: the longest-lasting human role is defining objectives and deciding what we want; technical work can be automated, choosing the right objective won't be handed over soon.
  3. Continual learning breaks in the details, forcing retraining; a cumulative task like RSI is easier than real work whose distribution keeps shifting.
  4. Distillation works against centralization: distillable capability can't be defended, so value moves to unique real deployment data, verticals like Composer and Harvey.
  5. RL works because mid-training does about 80% and RL tunes the policy with one high-signal bit; it extends task length rather than crossing domains.

What would change my mind

the next discontinuity adds to the current paradigm instead of tearing it down, and the paradigm connects the dots on its own.

How to read this

A roundtable of three frontline training researchers, chosen, the host says, because they're at relatively open labs and can speak on the record. Nobody is selling a product; they undercut one another and often admit uncertainty, a rare signal of honesty. Overall, RSI is coming, but a pile of engineering friction is slowing it down.

Breakdown · 12 steps
  1. What's missing isn't intelligence but the next discontinuity
  2. Clear objectives mean 10x; setting them is the hard part
  3. Distillation keeps anyone from pulling away
  4. Continual learning inside the lab
  5. Learning to learn in simulation: is it enough?
  6. Cumulative tasks are easier than a paralegal's
  7. Continual learning breaks in the details
  8. Where the signal comes from
  9. Will models keep getting bigger?
  10. Mid-training does 80%; RL tunes the policy
  11. Move 37 and a monoculture
  12. Timelines: two to ten years

Compiled from the video's captions, by speaker.

01

What's missing isn't intelligence but the next discontinuity

If 2036 looks normal, the likely reason isn't models that aren't smart enough but stalled generalization and continual learning. Moore's straight line lived on discontinuities, and today's paradigm may not find the next one itself.

00:00 · If 2036 looks normal, what went wrong?

0:01Dwarkesh Patel: Today I'm talking with three AI researcher friends I learn from every time we talk, who happen to be at somewhat open labs and companies, so they can say things on the record. Beren Millidge is CTO of Zyphra, which develops open-source models. John Schulman is chief scientist at Thinking Machines, previously a co-founder of OpenAI, and led the RLHF work that led to ChatGPT. Charlie O'Neill is head of model training at Baseten. First question: if it's 2036 and we don't have billions of superintelligences that have radically transformed the world, what is the most likely technical reason, setting aside political shocks, a war or a ban on AI?

1:16Beren Millidge: There's a classic pattern, almost Moravec's paradox: we think that if AI can solve hard maths problems or win at chess it will be amazing; then it does, and the impact is smaller than expected. If that continues and the true spark of generalization never comes, AI could end up extremely good at everything people put into a benchmark or an environment, while some persistent sim-to-real gap blocks everything else. I think that's unlikely; we already see this kind of generalization, even from RL. But if meta-learning is ridiculously hard to generalize and we don't solve continual learning, that would be my default scenario.

1:50John Schulman: I agree. Humans still have a lot of advantages over models. Each new model catches up in some areas, but you get bottlenecked wherever the model is weaker, has worse judgment or can't check itself well enough. There's a cycle that keeps repeating: a new model comes out, people are blown away and say this is AGI, and after a month or so it starts to feel dumb. That cycle might keep going, and it's hard to predict how many times. Right now capabilities don't grow explosively because research and engineering still hit enough bottlenecks: even if a model writes far more code than a person, it doesn't make you 100x more productive.

2:55Charlie O'Neill: For me the question is how far the current recipe, transformers plus RL, is from the global optimum of a learner you could put on a chip. People imagine that once an agent is even 0.1% better than all humans at AI research, running hundreds of thousands or millions in parallel, faster as chips speed up, outweighs every other bottleneck, and you get a very fast takeoff. But think about Moore's law: a beautiful straight line that held for a very long time, kept going by many discrete discontinuities and innovations. The same happened with LLMs: pre-training scaling hit diminishing returns, then RL came along and gave a new curve, so it kept looking like a straight line. If we need another of those discontinuities, I'm not sure training LLMs on RL environments, even RSI-targeted ones, can discover it. If not, we'll probably hit an asymptote.

4:25Dwarkesh Patel: Do you think that discontinuity will be harder than anything since 2012?

4:34Charlie O'Neill: If we knew, we'd be able to implement it. But we should distinguish a discontinuity that adds to the current paradigm, something cumulative beyond RL whose dots the models might be able to connect, from the question of how far we are from the global optimum: do we have to go back and throw out gradient descent and neural nets altogether? If it's that far away, I don't think scaling the current paradigm, however many LLMs you run, can discover it.

5:05Dwarkesh Patel: The only hope, really, is that deep learning can't get us to an AI that dominates human R&D, including the human ability to come up with new paradigms. But if you just extend the progress since 2012, and I know it's been powered by huge compute scaling, it would be weird if it didn't get to dominating humans at least in R&D, especially over the next few years. Ryan Greenblatt made a point on the podcast recently: as AIs get more capable, they can make progress on simulations that reward getting better not only at AI R&D but at science generally, and every lab and many startups are targeting that. Another intuition pump is the Elo of chess bots since the '80s: a very linear rise, but a huge discontinuity as they cross the human range, from experts always winning against AIs to experts never winning. So far AI hasn't had much end economic impact because it's still slowly rising in Elo relative to humans.

6:40Beren Millidge: I agree it would be very surprising. The only way it doesn't happen is if progress asymptotes just before, and in my opinion we're already pretty close to crossing the human Elo range. The only other way to end up in your scenario, where 2035 looks normal, is dramatic regulation of AI, which I actually see as more likely than a technical reason.

02

Clear objectives mean 10x; setting them is the hard part

With a clear objective, the current paradigm could speed up 10x; the right objective can't simply be thought up. The longest-lasting human job is defining the objective and deciding what we want.

07:03 · Clean objectives versus open-ended science

7:08Charlie O'Neill: There are different kinds of research. There's autoresearch-style work, where the objective is already cleanly specified and you optimize it; everyone pictures that pushing pre-training loss down and environment rewards up leads to improvement. Maybe what Ryan means is the far more open-ended science that paradigm shifts need, where we can't specify the objective, and the AIs certainly can't either. We have to be really careful about how we specify objectives.

7:49Dwarkesh Patel: John, you were in the trenches back then. Presumably a big breakthrough was realizing that next-token prediction is the thing; in 2014 you wouldn't have thought the nanoGPT speedrun was what to optimize. Now that we're in this paradigm, you would think to speedrun it and have AIs get really good at that. But maybe there's a next inner loop the AIs wouldn't anticipate. There's an outer loop of revenue that should eventually be strong, but it's very slow.

8:29John Schulman: In fact, I remember in the early OpenAI days having the intuition that just minimizing log loss wouldn't get you to intelligence, because the important bits are such a small fraction of the loss that they'd be drowned in noise, so we needed better objectives that put more weight on the important things. You can make all sorts of arguments: humans probably don't learn to model everything in their environment; most people can't reproduce a photorealistic scene they've looked at; so we must need a better objective. Then it turned out it just worked anyway.

9:33Dwarkesh Patel: And as you've pointed out, even in AI research today the inner loop of post-training benchmarks doesn't necessarily translate into what users like.

9:42John Schulman: Yes. The whole field relies heavily on generalization, and it's very hard to predict when you'll get it, especially out of distribution. Training on the task you care about helps, but the most important advances are often kinds of generalization we have no right to expect: from naive next-token prediction to tasks that need deep understanding of the input, or to rare skills picked up in pre-training; and from verifiable tasks to less verifiable ones, which there's no a priori reason to expect either.

10:45Dwarkesh Patel: One intuition pump for a very rapid singularity, even without scaling the inputs other than AI labor, is this: before every seven-figure experiment, you spend an equivalent amount of compute on AI labor, automated versions of you spending a century thinking about the optimal experiment, running small ablations, building a century's worth of theory; and after it, another century analyzing what happened and choosing the next one.

11:34John Schulman: If you think hard enough, you probably could have predicted some of these things. There's probably some clever small-scale experiment that lets you build a theory that generalizes to the large scale, so we're nowhere near the ceiling of how well research can be done. I can imagine AI doing a lot of analysis and theory-building, spending as much compute on that as on the experiments themselves.

12:17Charlie O'Neill: There are concrete examples when the objective is well specified. All thinking can do is update your posterior on the bits you've already got; you can't gain new bits just by thinking. But when the objective is clear and the data is sitting there, I expect a big speed-up within the current paradigm. If an AI had looked at the Kaplan scaling laws, it would have noticed they used intermediate checkpoints without accounting for annealing, so the result was wrong. That would have been caught years earlier and saved a year or two of progress. The same with muP, how learning rate scales with model size, and the importance of width: you could back out a lot of this and pick the low-hanging fruit. I'd expect a 10x speed-up if the job is just maximizing the objective we already have. But I don't see how that generalizes to coming up with the right objective in the first place. Thinking alone doesn't buy you that.

13:34Beren Millidge: That's the key question for any very rapid RSI from current AIs: how well can AIs generalize to learning their own objectives? A self-propelling loop needs the AI to propose objectives, optimize them, propose new ones, and not go off the rails for a very long time. Maybe this is another Moravec's paradox: autonomy, deciding what to do and then doing it in a loop, feels easy because we do it all the time, and evolution had to build creatures that survive on their own for long periods. It might be really hard for AI, the way locomotion is hard for it while maths is easy, the reverse of us.

14:00Dwarkesh Patel: But doesn't the growing time horizon suggest otherwise?

14:25Beren Millidge: Exactly, and I agree there's no obvious evidence for this. The fact that our agents are now very persistent is evidence against it. But if it is hard, it would be one reason we don't get an immediate takeoff.

14:35 · The last human job: deciding what we want

14:39Dwarkesh Patel: Looking back from 2012, or from when you started doing research, which of all the innovations since, engineering or conceptual, looks like the last thing humans will have to do before AI fully automates AI R&D?

15:05Beren Millidge: Probably iteratively asking the right questions. Even if the AI can run any experiment, someone has to decide which experiments to run. Right now AIs are much worse at that than at coding the experiment: when we talk about research, they propose a grab bag of very tiny steps.

15:24Charlie O'Neill: Or the jump from DeepMind's "we'll solve intelligence by learning to play games at a superhuman level" to one researcher, Radford, deciding to just predict the next token of a very wide swath of data. Even after Radford found that, it took a while before people scaled it up, because we first had to come up with scaling laws and the idea that you could predict these things reliably.

15:55John Schulman: I'd say the human role that lasts longest is defining the objective and deciding what we actually want: how AI assistants should behave, what it means to be helpful, what the objective is when we do RL from human feedback, and later constitutions and model specs. Even if AIs can do all the technical work, we'll still have to do a lot of that.

16:31Dwarkesh Patel: Alignment is the final job.

16:33John Schulman: Alignment is sort of the answer. But alignment decomposes into specifying the objective, figuring out what it should be, and then achieving or optimizing the objective you've defined. The first isn't going away anytime soon. That's why a post-training team needs so many people: there are many areas where someone has to work out how the model should behave, and that is very hard to automate.

03

Distillation keeps anyone from pulling away

Whatever RL teaches can be distilled cheaply, and the prompt distribution is the key. Frontier labs may have no edge in training environments; real deployment may matter more.

18:39 · Why no single lab has run away with it

18:51Dwarkesh Patel: Why isn't there huge consolidation among model providers? So much points toward centralization. Over the years, is there something that prevents it?

18:54John Schulman: I think distillation is the main force against centralization. Anything that can be learned through RL can be distilled very easily, because it's a small number of bits you can learn from a small amount of data. If you can get trajectories showing a behavior, you can distill it. There's also the possibility of company-specific models, where a company learns from deployment and keeps improving its own model, provided by today's oligopoly or by some smaller company. That would change the game a bit.

19:56Beren Millidge: And continual learning doesn't stop distillation. If your model improves every day, people can distill it every day. The loops can run at the same pace.

20:10Dwarkesh Patel: But to copy a model's behavior, don't you need to know the right distribution of prompts?

20:18John Schulman: Yes. For distilling with supervised learning, the prompt distribution is extremely important. Distilling all the useful capabilities is very non-trivial even with full access and the chain of thought, because you have to prompt the model with a really wide distribution of realistic prompts. It has recently come out that some Chinese companies are probably using router services, proxies that let people in China use US frontier models that are otherwise blocked there, mostly for coding. Those services are collecting and selling some of the data, which is very useful for distillation because it gives you the perfect prompt distribution.

21:29Beren Millidge: This is where AI helps a lot. In the frontier pipelines, and in what the Chinese labs describe in their papers, they get seed prompts from a mix of humans and this kind of data, then use their own or other frontier models to synthesize vast coverage from them. You can automate much of the prompt gathering and environment creation, and humans need to provide fewer and fewer bits as models improve.

21:57Dwarkesh Patel: But you still seem bottlenecked by having a service with users going through it. The user says: make me an app like this; that didn't work; actually add this feature; no, let's step back and do something else. Capturing that whole trace is the thing. And if you could have produced it anyway, you'd just have RSI.

22:25Beren Millidge: Not necessarily; it helps a lot, but in theory you can just think about what users want. Ultimately, a fully automated loop is RSI: the AI decides the data and the training. But it depends on how much human information you need. At some point, if you want traces that look like this, you prompt the model and it produces a pretty good approximation.

22:40Dwarkesh Patel: What if it's "make me a really good politician", and it has to anticipate from scratch how a debate in the Senate would go?

22:51Beren Millidge: Ironically, that's easier for the distillers than for the frontier labs. The distiller asks the frontier model, which already knows how to be a good politician, to generate the traces. To build the first model that can do it, you have to somehow gather data on what politicians do every day. Asking for a billion variations of something is much easier than creating the first one.

23:19Charlie O'Neill: You can make a concrete prediction from the fact that the Chinese labs have this router data. What started it was me asking: isn't it weird that Sonnet 5 and Opus 5 are almost objectively worse models than GLM-5.3 and Kimi K3, even though they had not only distillation but logit distillation from Mythos? The counter was that the prompt distribution really matters; you need to see what users do. The prediction is that frontier labs don't necessarily have much advantage, if any, in RL environments now. User distribution matters for general behavior, but the best measure of a capability is the very hard RL environments built at the frontier. If Anthropic has those environments and logit distillation and still made a worse model, then maybe—

24:12Dwarkesh Patel: Then real-world deployment matters more than the environments. That's really interesting. But they had to incentivize those capabilities in Fable, the frontier model, in the first place, so it's odd they can't do it again with a smaller model.

24:32Charlie O'Neill: Maybe we're in an uncanny valley where copying the frontier model too closely fails because the student-teacher gap is too large. People say this about Opus: the difference between Opus 4.6 and Opus 5 is that Opus 5 feels like it has an AI judge checking everything it's done, which is why it uses so many tokens, but it doesn't have Fable's big-model sense of when to stop or which path is worth going down.

25:02John Schulman: I'd offer a slightly different hypothesis. Environments vary along two axes, difficulty and realism. It's comparatively easy to create lots of difficult environments, much more complicated tasks or ones that need more cleverness; call it the benchmaxxing distribution, since many prominent benchmarks are very hard, puzzle-like tasks that are easy to verify. Then there's realism: being good as a coding agent in realistic settings, with multiple back-and-forths with a human and multiple objectives. Labs crafting a behavior for the first time have to push both ways, and realism needs rubrics or human feedback feeding the reward. Naive distillation only matches the teacher on the benchmaxxing distribution; without enough environments that exercise the trickier realistic settings, the student doesn't get those abilities. Maybe the big models generalize better from narrow hard tasks to realistic ones. With a really good realistic prompt distribution you can match the big model well; with only easily verifiable tasks you match it on every benchmark but do worse on the broader distribution. That might explain something about the smaller Anthropic models like Sonnet 5, though it's hard to know how they post-train them. They may also just have got a few things wrong in a post-training stack that keeps changing and created quirks people really dislike. It's very easy to screw up post-training in ways that don't show up in benchmarks.

27:49Beren Millidge: One more basic point: the frontier labs buy all their data from big data companies, and the Chinese can buy the same data, and they do. People are annoyed about it, but with the same data plus distillation, it's quite easy to keep up.

04

Continual learning inside the lab

Automated researchers will be trained with human feedback plus practice environments. Today labs distill the last three months of progress back into the model, which is why it can feel asymptotic; environments can go beyond what humans can do.

28:06 · How the first automated AI researchers get trained

28:10Dwarkesh Patel: How will the first models capable of automating AI R&D actually be trained? There's the toy version Ryan described: have GPT-8 build GPT-3-sized models that are great at inner-loop challenges, like beating video games that need continual learning, or reaching a given loss with the least compute. John, you suggested that may not be how it happens in practice.

28:47John Schulman: We'll probably combine learning from human feedback, to absorb researchers' taste, with lots of practice environments involving multi-step research projects, and in each iteration patch whatever seemed most broken in the last one. Researchers will use the AIs a lot, notice consistent weaknesses, and fix them with human feedback or new environments.

29:30Charlie O'Neill: One way to think about it is how far back along the lineage you roll and then let it self-play. In the limit you give it a GPU and some neural nets and say: figure out how to train a model for these tasks. Today we go to the very edge of the lineage: here are the bugs Anthropic found in its training stack in the last few months, turn them into environments. You could imagine rolling back to before GRPO and asking it to discover the best way to RL models, and further back still. But we'll be so compute-bottlenecked that people will stay at the frontier, essentially diffing the bugs and improvements found since the last version and turning them into training environments.

30:30Dwarkesh Patel: Which also gives fresh, non-stale data between generations. It's basically continual learning inside the lab: distilling the last three months of research progress back into the model through environments and RLHF-style work.

30:41Charlie O'Neill: And it is distilling. That's maybe why some of us feel it's asymptotic. You're always catching up on the last three months of progress, contributed partly by AIs but still with humans in the loop, inching closer and closer to what the human researchers find.

31:04Beren Millidge: Distilling on trajectories can never take you above them. But environments can go far beyond what a human can do: it's easy to design an environment no human can solve that the AI can still attempt. That's the path to getting ahead of human AI research. In AI research especially, goals are easy to define: say the loss needs to be 1.3, which no human can reach now. That's an extremely measurable, verifiable task.

31:38Dwarkesh Patel: Or a nanochat speedrun done faster than any human speedrunner, or a 100 million parameter model that beats Minecraft; maybe that's too easy, so a much more complicated game.

31:45Charlie O'Neill: Isn't it crazy that we're calling a 100 million parameter model beating Minecraft too easy? Imagine saying that five years ago.

31:52John Schulman: A lot of research isn't hill-climbing on a well-defined goal, though. It's more like: we have an intuition about how models should be better and an algorithm that seems to go a little in that direction, so we design a task to show signs of life, and if we see them, we make successively more realistic versions. You're not optimizing the eventual production objective directly; you relax on the realism axis, find methods that work, then return to realism once the method matures. There's also research aimed at explanation and theory. We rarely have predictive mathematical theories in machine learning, but we have many informal ones.

33:27Beren Millidge: Presumably models will be trained on a mix of all these tasks, some easily verifiable, some judged by an LLM or by asking a human whether it looks reasonable, in the hope that they generalize to much vaguer, fuzzier tasks. They probably will to some extent. Whether they generalize enough for the loop to seal itself with no humans in it is unclear.

05

Learning to learn in simulation: is it enough?

The labs bet on training persistent agents that triage well across vast numbers of simulated environments, patching domain by domain. With low sample efficiency, simulation is the only option, and real-time human interaction is hard to simulate.

33:51 · The labs' bet: long-horizon RL in simulation

33:52Dwarkesh Patel: Here's what I think the plan is; tell me if you agree. Scale up RLVR training across millions of diverse environments in hundreds of domains. Out comes an agent that has learned to be persistent, to triage information and context, eventually to work end to end with other agents, and that is very sample-efficient in context. It functions like a drop-in remote worker over a week or a month. Is that the labs' bet, and is it enough: learning to learn in simulations inside a data center, then being deployed without actually learning from deployment?

35:10Charlie O'Neill: It's hard now to separate how much of the labs' effort goes to direct RSI and how much to generally intelligent models they can deploy for revenue to fund the next big training run. For the latter, yes, that's the bet, and the pattern of environments over the last few years is clear. Anthropic's lineage is the clearest example: first coding, where internet data and internal material make environments easiest to build; then, from the task horizon coding gave them, generalize. Finance next, with enormous amounts of Excel data in RL, then PowerPoint, the long tail of the working economy. That worked really well, and other labs, even open-source ones, now see it was the right bet.

36:09Dwarkesh Patel: But what follows from that? I asked Dario: if you truly expect models to learn on the job like humans, why bake in PowerPoint skills instead of expecting them to pick it up while deployed? One explanation is that models will get there soon but aren't there yet, so you amortize the skills into training. Another is that the goal isn't widely deployed work at all but RSI, and this is a way to earn revenue to pour into the RSI model; after the singularity, whatever comes out will be good at everything that bottlenecks today's models. John, how should we read so much task-specific knowledge if the path is generalization?

37:09John Schulman: If models were good enough at in-context learning, you wouldn't need to train them on finance; they could read all the books on the fly and work out everything in the relevant jurisdiction. But even a smart enough model would be more efficient at runtime with the intuitions baked into its weights by RL. In practice, providers are going domain by domain, strengthening the models in the highest-value domains, and that's one answer to why models have improved so much: they've covered many high-value domains and the most common skills.

38:11Beren Millidge: It's also not that expensive to do both. The models are massive and can afford, in parameters, to learn everything. There's likely some transfer: the general meta-learning of figuring out what matters, having taste and doing long-horizon work may generalize. And there isn't much RSI data in the world; it's hard to generate. With masses of compute and parameters, why not amortize in other data too, besides the obvious commercial reason of selling a model?

38:55John Schulman: There's also a question of whether today's sim-to-real paradigm stays dominant forever: look at what real tasks are like, build environments you can simulate in the data center, do RL on them. It has been very successful but has weaknesses, because many things are hard to simulate, especially interacting with a bunch of people in real time.

39:31Beren Millidge: Sim-to-real has to dominate while sample efficiency is low. You need thousands and thousands of interactions, and no human is going to sit in the RL loop, so you simulate. If sample efficiency improves a lot, you'd expect learning from deployment to become a much bigger part.

39:56John Schulman: You can also learn off-policy: take all the traces and potentially learn something from them without resimulating everything.

06

Cumulative tasks are easier than a paralegal's

Deployment data already flows back generation by generation. RSI is cumulative, each discovery held once found; law-firm work, whose distribution keeps shifting, is harder. Taste might be learned from short episodes.

41:19 · Half the compute teaches the model nothing

41:27Dwarkesh Patel: It's strange that 50% of compute goes to inference that doesn't directly make the model better. A key advantage digital minds should eventually have is that where a human gets 50 years of real-world experience, a model, through all its instances, gets millions of years of deployment across all kinds of economically relevant work. Right now that data isn't meaningfully helping it get better. Once models learn from it, you'd have something like a widely deployed intelligence explosion. When does this hive-mind stuff start?

42:14Beren Millidge: At a very basic level it's already happening, from one generation to the next: you can put deployment data into the pre-training or mid-training of future models, especially after filtering, judging, annotating or synthesizing it.

42:29Dwarkesh Patel: How much of the generation-over-generation improvement does that explain?

42:32Beren Millidge: Quite a bit, I think. I don't know whether the labs do it, since they say they don't train on people's data. But the Chinese 100% do; they definitely get this advantage. That's basically what distillation is: ping the models, get a slice of their deployment data, train the next generation on it. They can do it on their own models too; there's no reason not to.

42:56Charlie O'Neill: Completely agree. Zoom out far enough and it's definitely happening. The holy grail of continual learning we all picture is an organic live loop: one model has an experience and updates on the spot. Zoom in to that granularity and a lot breaks. But the big labs and the closed models are doing the slower version, and there are early signs of life with open-source models at a much faster cadence. Composer is probably a good example; Harvey is doing the same with legal agents. You build very specific environments from your task data, from what users complain about and from the feedback you extract from your deployments, and these companies can use that data far better than the big labs. Then they do a big post-train of Kimi K3, deploy it, maybe do some online learning, the way Composer basically ran REINFORCE for a long time. A human is still in the loop, deciding which signals matter and how to turn the data into environments. The cadence is longer than the one you're picturing, but it really is happening, and the loop will keep getting faster.

44:20Dwarkesh Patel: With Composer, in Cursor people press Tab or don't on the next completion, and based on that it gets better every day at predicting the next—

44:29Charlie O'Neill: That was the old Tab model. They did the same for the actual generative model. It's hard because online RL has no groups: one user says one thing and you get one rollout, so variance reduction is a big problem. Cursor's fuzzy answer was good heuristics that estimate how much better or worse than average a response was, then a big REINFORCE update. To check it hadn't got worse, if it improved on CursorBench they deployed a new model every five hours; if not, they threw that version out.

45:09John Schulman: I think the biggest problem is not knowing what the reward function should be for natural data. A superficial signal, like whether they accepted the edit, can get reward-hacked.

45:24 · Cumulative tasks, and why RSI is easier than being a paralegal

45:25Dwarkesh Patel: Isn't this a bigger issue for sim-to-real? The longer the task, the harder it is to simulate in a data center. Even in coding, there's no year-long task that doesn't eventually require talking to a client, the company or users. Superintelligence should eventually be able to run a business, start one and make it profitable, day-trade profitably or win a court case, all very hard to simulate; interacting with the real world is part of the learning. Maybe transfer from simulation is enough. If not, and you need weight updates from those interactions, the models' sample inefficiency is a deeper problem. By default I don't see how we avoid some crazy recursive self-improvement within the next 10 years. The one reason it might not happen is sample efficiency: models are plausibly a millionfold behind humans, comparing what a human sees from birth to adulthood with what a model sees from cold start to the end of training. So will simulations transfer to the extremely long-horizon, complicated real stuff? And if not, does the lack of sample efficiency come back to bite us?

47:25Charlie O'Neill: I'd split tasks by whether they're cumulative, or whether the distribution is non-stationary and you have to keep learning and relitigating things. RSI might be cumulative. It's theoretically possible to have a Python file under a million tokens that trains, from scratch, a model capable of recursive self-improvement. Every discovery is a line in the sand you hold: once you've found attention, mixture of experts and GRPO, you add them to the training stack and they stay. When OpenAI trained 5.6 Sol or 5.6 Terra, or whichever it told us about, it didn't have to rediscover attention; it basically called scripts like pre-training.sh and post-training.sh. The real world, and the reason people think so much about continual learning, isn't like that. An agent working as a legal associate at a law firm faces a very non-stationary distribution: it has to fit in its context all the constantly changing relationships between the important people there, and all the implicit ways things get done and where information lives. So tasks will split, and if the labs believe RSI is cumulative, in the sense that no brand-new architecture has to be discovered, more effort and compute will go there rather than elsewhere.

48:56Dwarkesh Patel: It's so unfortunate that RSI happened to be easier than being a paralegal.

49:08John Schulman: Today's models are weaker than humans in many ways. Some are about sample efficiency in certain regimes: models are very sample-efficient in context, but in some medium-length regime humans may update their weights more efficiently. Other weaknesses are completely different: lower diversity of thought, or being bad at certain long-horizon judgments. A lot of what people call taste is knowing what behavior works in the long run; in software engineering, which systems will stay maintainable and work well over the life of a project. Various weaknesses limit RSI, some tied to sample efficiency and some not.

50:40Charlie O'Neill: A thought experiment: give a model a context window of a trillion tokens, enough for all your experience up to, say, RLHF, with the same in-context learning ability it has at a million tokens. Is taste then solved? Would it make the judgments you did?

51:07John Schulman: It would have to be trained to learn from that context, to make the right update from it, or it would have to generalize.

51:26Beren Millidge: And you'd need trillion-length data to train it on that context; today you can't even jump from 10k to a million. But in theory, yes. It comes down to how meta-learnable taste is from shorter episodes, and there's no obvious reason it needs very long ones, because humans develop taste without many long episodes. A PhD, from first year to postdoc, is maybe five years and 10 to 30 research projects in total, yet people develop taste quickly from a short succession of small things. AI will have vastly more experience to meta-learn taste from. How well that generalizes to really long-horizon things is unsolved; we don't know.

07

Continual learning breaks in the details

Companies won't let providers learn from their deployments, pushing toward swappable modules. At scale the loop works; for one model updated again and again, catastrophic forgetting sets in and you end up retraining from scratch.

52:22 · Continual learning breaks at the micro level

52:30Dwarkesh Patel: Eventually AIs should be learning a lot from each individual deployment. Today there's a fuzzy meta-process by which models improve from deployment, but it's a weak feedback loop. Do you see rapid hive-mind learning on the horizon, and how exactly would it happen?

52:54John Schulman: Whether we get a hive mind that learns from all its deployment experience is largely about incentives rather than technology. Companies won't want the model provider learning from all their deployment, because that could erode their business advantage.

53:21Charlie O'Neill: The economics will push toward modules that get swapped in, rather than weight updates to one big shared model. LoRA is the obvious example. There's also a lot of work on fitting arbitrary context into a fixed size, the linear attention work, and cartridges, KV caches trained to be highly compressed. Companies might sign up for those because they don't change the underlying base model. Those don't really help the big labs, but I think economic pressure will force the labs down that path first.

54:13Beren Millidge: Which pressure, though? Even with cartridges or LoRAs, you can still take all the traces and dump them into the pre-training of your next generation.

54:20Charlie O'Neill: Yes, that's a more indirect form of learning for the big labs, and still very valuable to them. But I can't imagine starting with directly training one big model on all the exact data coming in.

54:43Beren Millidge: It'll go in stages, not one discontinuous moment where we suddenly fix continuous weight updates. More likely, cartridges and the like let you specialize deployments; you generate traces, put them into the model, and three months later ship a model that's better at it; then specialize and consolidate again, faster each time. Instead of a release every three months, every week, then every day, then every hour, at which point we've basically solved it.

55:02Charlie O'Neill: That connects to how far the current paradigm is from this. We've done a bit of research, and many others have too. At a really large scale, with enough noise washed out and big enough batches, the outer loop of putting data into mid-training and building our own environments does work as a kind of continual learning. But zoom in to the micro level, one model updated again and again for a law firm with a relatively small amount of data, and all the methods break down a bit. SFT on successful traces, off-policy or on-policy: after hundreds of these micro-updates you see catastrophic forgetting, loss of what was learned on top of the base model much earlier, and degraded general capabilities. On-policy distillation pushes that horizon out a little, but eventually succumbs to the same thing. RL is good at getting capabilities in but not explicit knowledge, like how this particular person at this law firm handles this particular process; getting knowledge in with RL takes a lot of compute to build the right environments.

55:46Dwarkesh Patel: Is the forgetting fundamentally about capacity or about technique?

56:30Charlie O'Neill: A bit of both. SFT and even on-policy distillation can be way too destructive. RL is nice because it changes the model very little, just nudging it within a very small loss valley, but that also limits how much RL can change the model.

57:08Beren Millidge: I think it's mostly technique; capacity is definitely there. Take a model of literally the same size and pre-train it from scratch with all that data in mid-training, and it will be better; a lot of what happens today is exactly that. Something stops us from training the same model forever rather than training a new one from scratch on the old model's data: as Charlie said, a mix of lost plasticity and catastrophic forgetting. Naively training on non-stationary data shifts the distribution and the old stuff is forgotten, and we don't have good methods to stop that.

57:46Dwarkesh Patel: So in the limit you're bottlenecked on retraining from scratch with all the new information.

57:58Beren Millidge: Which is very expensive. With real continual learning, you'd never train a new model; one model would keep learning and expanding.

58:06Charlie O'Neill: That's the question. We've pushed back how much has to be done from scratch: you can now take a pre-trained base and do very good mid-training on top, fairly continuously, plus RL from later checkpoints. That looks more like continual learning, but it's not taking the latest model, applying a few tiny updates, and never losing anything.

58:42Dwarkesh Patel: Isn't that literally what post-training already does, distilling a further-RL'd fork into a model that's been through a huge amount of training?

58:44Charlie O'Neill: It's still at a large enough scale to wash out much of the noise, and it isn't focused on one distribution, which, as Beren said, is the issue.

58:59Dwarkesh Patel: But eventually you'd be learning from billions of deployed instances at once, which should wash out noise too.

59:13Beren Millidge: Maybe at that scale. You can do continual mid-training for a long time, and roll back to a checkpoint and give it new mid-training data, but not indefinitely. Keep training the same base forever and it plateaus; it can't learn new things. That's why people end up training new bases.

08

Where the signal comes from

Each rung of the environment ladder is harder to climb, and even the whole world lacks the bits the frontier needs. At small scale data explains about 12.0x of efficiency gains and architecture about 3.7x, though architecture opens new regimes.

1:00:33 · Where the signal comes from

1:00:38Dwarkesh Patel: Let's talk about data: how much of AI progress is explained by data progress? Is there some data distribution that, trained into current architectures, would produce a superintelligence dominating human experts in every field?

1:01:00Charlie O'Neill: Including post-training data and environments? Its existence is obvious; the question is whether we can create the right environments to get there.

1:01:08Beren Millidge: In the trivial case, you train it to output the Python file that trains the actual superintelligence, memorized in the weights.

1:01:15Charlie O'Neill: There's probably a ladder of RL environments that gets you an AI researcher at least as good as a human one, but the effort to climb each rung grows roughly exponentially, and that trade-off sets how fast we reach the last rung. We're still early in environment creation and exploit many asymmetries. In some environments it's easier to go backwards than forwards: you define a complex data-generating process, keep it hidden, and the model has to spend irreducible tokens working out what it was. Others inject real-world information: Anthropic finds a bug with tens of thousands of humans and LLMs combined and turns it into a neat environment that a single LLM could in theory solve within a few million tokens. We're cherry-picking these asymmetries and counting on task-horizon generalization. At some point it hits diminishing returns: you can't always find a process that's easier backwards, so you have to build something with a long enough horizon, with humans, which is very complex, plus the compute and time for the agent to do it. I think the curve will start to flatten.

1:03:07John Schulman: I saw that someone fine-tuned the Talkie model, trained only on data up to 1930, on modern coding-agent data, and it did better than Claude 3 Opus on SWE-bench. A model with no knowledge of code, fine-tuned on a moderate amount of data, behaved better as a coding agent than a much larger pre-trained model. Once you have examples of the right expert behavior, it's surprisingly easy to copy it into a relatively weak model.

1:03:53Charlie O'Neill: A counterexample: a recent paper trained a model up to fifth-grade maths and primary-school English, then tried to RL it to late high-school and college maths. The gap was too large; it couldn't climb at all. With rungs of year 7 maths, then year 8 and so on, you could climb to year 12. Again, it's the distance between rungs and how hard they are to build.

1:04:18Beren Millidge: That's the RL signal problem. RL isn't good at exploring right now: if the model can't get it in 128 rollouts, it's very unlikely to get the signal to progress. That's why RL needs curricula and pre-training doesn't.

1:04:35Charlie O'Neill: Pre-training data is different from post-training data. Humans will be involved less and less, but you're still bottlenecked on how much signal you can extract from the real world. There's a lot of signal, people doing spreadsheets and legal work, but at today's capability frontier, how many bits in the world are really relevant to improving the model? How many new maths or coding problems are being solved beyond current models' reach? That's why diminishing returns kick in: even the world as a whole isn't giving you the bits that tip you into the next basin of capability.

1:05:30Beren Millidge: Totally agree. In pre-training the signal is already in Common Crawl; the problem is filtering out the noise, which is fairly automatable. In mid- and post-training the signal isn't in the original data at all, and no amount of filtering finds it: there's no hidden proof of a Millennium Prize problem sitting in Common Crawl. You have to get the bits elsewhere: from humans writing out their reasoning, from environments whose design and objectives humans decide, or from training on the human data that exists in deployment.

1:06:10 · Data versus architecture

1:06:15Dwarkesh Patel: How much of pre-training progress is driven by data? I did an investigation with Jerry Han, a student at Princeton: we trained every recipe from 2019 to now against every dataset from 2019 to now, pairwise. GPT-2 on the newest dataset like Ultra-FineWeb, Delphi, the newest open-source recipe, on the Pile; the whole grid, measuring how much less compute it takes to reach a given capability. Data seems to explain about a 12.0x compute-efficiency gain and architecture about 3.7x, at a very small scale. If most pre-training efficiency gains come from better data, how long can that continue, with more filtering and more synthetic data?

1:07:23Charlie O'Neill: My prior is that the low-hanging fruit is somewhat exhausted. We got the internet as one big block, and its useful content isn't growing at the same rate. There are probably a bunch of 0.1% loss drops left, but not as many as we've had. It's also interesting that you find a cumulative 33x across both. Epoch or someone estimated 3x a year since 2019, which implies something like 3 to the 7th, over 2,000x. So where is the missing 100x or so coming from? That probably tells you how much of this is post-training.

1:08:08Dwarkesh Patel: I think the explanation has to be that many efficiency gains are scale-dependent and we ran at extremely small scale. Which raises whether data gains or algorithmic gains depend more on scale; we didn't have the compute to find out.

1:08:36Charlie O'Neill: Naively, the architecture's scale dependence is fairly well known; you can fit a straight line to it. I'd have no idea how to do that for pre-, mid- and post-training data combined.

1:08:49Beren Millidge: Funnily enough, I think data matters more with scale, and architectures are a one-time thing. Quoting an X% efficiency gain is misleading: an architecture lets you reach a qualitatively new regime, and within it data is what matters. Without even GQA, doing full attention all day, a million-token context would be ridiculously expensive, so we could never use million-token data or get those capabilities. Measure at 2K context, where the architecture unlocks nothing, and data looks more important than it is. I'm not sure these gains simply multiply. And a lot of today's mid- and post-training data gets better with scale, because very long-horizon environment data needs big models to use it. Train a 100 million parameter model on SWE-bench traces and it gets nowhere.

1:10:10Charlie O'Neill: Comparisons are also harder now because many architecture changes, at Kimi or DeepSeek, aren't aimed at lowering pre-training loss but at how the models will be used in the real world. DeepSeek's compressed attention is about inference efficiency, not a fundamental trade-off improvement.

09

Will models keep getting bigger?

Inference efficiency comes first, so parameters may plateau for a few years; Schulman expects growth anyway, depending on the scaling laws. The beautiful straight lines hide a lot of complexity.

1:10:23 · Will models keep getting bigger?

1:10:32Dwarkesh Patel: How will parameter counts scale in an RL-heavy regime? Frontier open-source models have grown maybe 2x a year. If frontier closed models have 100B or 200B active parameters, does that keep doubling? Or, since you also want to save compute on rollouts, is there a threshold beyond which more parameters don't matter much? How many active parameters will a frontier model have in 2030?

1:11:11Charlie O'Neill: For the next few years we're so focused on longer-horizon RL rollouts, where inference efficiency matters a lot, and the models aren't saturated at that; the bottleneck is still the environments. So we may see a bit of a plateau. I suspect Mythos and the GPT models are much smaller than the 10 trillion parameters people talk about; you can back that out just by comparing them with open-source models. I wouldn't expect huge parameter growth in the next few years. You size the model by how much pre-training data you have and how hard your RL environments are: you want a decent pass@1 on your hardest environments, and going bigger than that just means paying for more inference than you need. So a lot depends on how fast Mercor and the in-house teams can make RL environments more complex.

1:12:21John Schulman: I'd expect models to keep getting bigger just because compute is scaling and GPUs are getting bigger, but by how much depends on the scaling laws in non-obvious ways. Now that high-quality pre-training data is running low, data efficiency will matter more than compute efficiency in choosing architectures, which may affect how sparse you make the model. We don't understand sparsity well: total parameters are a different resource from active parameters; sparsity has increased, but not necessarily without bound, and there may be a sweet spot. There's a debatable argument that sparsity hurts data efficiency because several experts may have to learn the same thing. We don't have a good enough theory to know why sparsity helps, how much, or where it plateaus. With data rather than compute on the x-axis, you simply get a different set of optimal models.

1:14:48Charlie O'Neill: I'm also not sure models have doubled every year. People have trained 1 trillion parameter models for years; there was even an open-source one, Falcon. Liam from Periodic Labs posted that an early experiment was a very sparse 1 trillion parameter model, which John says was the Switch Transformer at Google, before they went to OpenAI. It was very good at knowledge and terrible at reasoning because it was so sparse. We've been in the 100 billion to 2 trillion range for a while; it hasn't been a nice linear increase.

1:15:26Beren Millidge: Two things. As Charlie said, inference efficiency for RL rollouts will push active parameters down a lot. Total parameters depend heavily on hardware: serving multi-trillion-parameter models needs very high memory bandwidth and VRAM. People still use a lot of H100s; as everyone moves to GBs and then Vera Rubins, we'll be able to serve and run RL inference at larger scale. And naively, larger models are much more sample-efficient: even unsaturated, bigger generalizes better and reaches a better loss on the same data. Right now data isn't the constraint, compute is, so we get smaller, inference-efficient models. If compute stops being the bottleneck, we may go back to larger, undersaturated models.

1:16:26Dwarkesh Patel: Under the basic Chinchilla law, maxing out parameters barely reduces the data you need; even infinite parameters cut it by less than 10x, because of the power law.

1:16:44Beren Millidge: But we're on the too-much-data side of Chinchilla: we over-train models. As data runs out, we could easily move back to the Chinchilla-optimal point, or even slightly under-train. It depends on the ratio of training to inference compute: bottlenecked on data, go bigger; bottlenecked on compute, go smaller. Synthetic data complicates it further, so it's very hard to predict.

1:17:22John Schulman: Part of why the scaling laws took so long to find is that you only get a clean relationship when you get everything right. The beautiful straight lines on graphs hide a lot of complexity: scaling every hyperparameter correctly, or parameterizing the optimizer so hyperparameters don't change with model size.

1:17:47Charlie O'Neill: Bugs have their own clean scaling laws too, like Kaplan forgetting cosine annealing, or not counting embedding parameters, which skewed the estimates for smaller models where embeddings are a decent share.

10

Mid-training does 80%; RL tunes the policy

Mid-training often gets the model about 80% of the way; RL tunes the policy with one very high-signal bit. What it delivers is generalization over task length, not across domains.

1:18:03 · Why RL worked better than expected

1:18:06Dwarkesh Patel: A year ago many people argued RL wouldn't scale well. John, you wrote that models learn one bit per RL episode: did I get the answer right or wrong. I wrote that it's even worse: when the pass rate is low, the model learns almost nothing from an episode. Yet today's models seem pretty smart, apparently from scaling up RL. Beren, you wrote a post a few weeks ago explaining this. Why has RL worked better than one would naively have thought?

1:19:01Beren Millidge: Several things. First, mid-training is underestimated. Much of what looks like RL's success comes from very good mid-training data: essentially pre-training on synthetic reasoning data and environments that warm-start the model for RL, which often gets it almost 80% of the way to the final RL checkpoint. RL on top is mostly tweaking the policy, so it needs far fewer bits than you'd think; it doesn't learn the behaviors from scratch. Second, those bits are extremely high-signal compared with pre-training, which is why you need RL rather than just SFT on successful traces. It's not only that they're the bits about getting the answer right; an SFT trace contains the answer token too. What matters is that the RL objective ignores all the other bits. In SFT you have to match the exact reasoning tokens of the model you're copying, so you get far too many bits about how it happens to reason. In RL you get the one bit, not drowned in noise. That's a dramatic increase in signal-to-noise, which is why RL is so efficient per step.

1:20:31Charlie O'Neill: There's been so much debate about what RL does compared with mid-training or SFT: pass@1 goes up but pass@256 goes down, as rare correct traces get outweighed by gradients from easier ones. The simple view now is that if you have enough compute to sample a large enough group, so the chance of getting several correct answers isn't insignificant, they get up-weighted. And to Beren's point, pass@1, RL's starting point, scales with the log of pre-training tokens.

1:21:48Dwarkesh Patel: A very basic question. That makes sense, but the models seem qualitatively so much more capable over the last year. How do we square the relatively small effect this implies for RL with the capabilities they seem to be gaining?

1:22:08Beren Millidge: It doesn't imply RL has a small effect. A few bits and small parameter changes can still change the input-to-output mapping the model computes dramatically. Even one bit can rule out half the hypothesis space, which is huge. Small amounts of RL from a really good starting point can change behavior a lot.

1:22:42Charlie O'Neill: Two things. Everyone hoped RL would generalize reasoning across domains; I don't think we got that horizontal generalization. Training on maths doesn't make you a great coder; you still need RL on code environments. What we did get is horizon generalization: the models learned to use more tokens for longer and keep making progress. Train on ever-longer environments, drop them into a completely new one, and they may lack the right reasoning patterns but can keep going for longer, which correlates with success. A paper called EdgeBench showed that how long models can work is doubling every three months; that's clear evidence of generalization. Finally, pre-training has the idea of quanta: a smooth loss curve that is really the average of vast numbers of discrete phase transitions, no induction heads and then induction heads, tens of thousands, millions, probably hundreds of millions of them. RL is similar. In the slow outer loop Beren mentioned, we RL a model, then dump its synthetic reasoning traces into the next model's mid-training, hitting quanta for many tasks. On one task it looks like a phase transition, a particular finance or Excel task jumping from a 0.5% pass rate to 90%. Average them all, add horizon generalization, and you say: wow, qualitatively better models.

1:24:32Beren Millidge: RL also does generalize a bit, with some transfer between maths and code, or puzzles and maths. And the sheer number of environments people target is vastly greater. Two years ago the labs wouldn't train for a task you do in daily life; now there are lots of environments aimed at each specific thing.

11

Move 37 and a monoculture

RL hasn't killed creativity; models once found several zero-days at a time. But output diversity has dropped, and with everyone distilling Claude, open-weight models increasingly write alike.

1:24:54 · Move 37, and a monoculture

1:25:00Dwarkesh Patel: We talked about RL causing entropy collapse, concentrating probability on solutions the base model already had. But there's another story, from Atari to AlphaGo's move 37: never initialized on human data, it could think in ways humans don't and find extremely creative solutions. Should we expect RL on LLMs to produce move-37-style creativity, beyond human creativity?

1:25:52Beren Millidge: AlphaGo used MCTS, which explores more than ordinary policy gradients. But I don't think RL necessarily reduces creativity. Qualitatively, in the OpenAI–Hugging Face incident, the models came up with multiple zero-days at a time to break out of the sandbox. That's already some level of move-37 creativity, from the LLMs' general generalization. RL is definitely not destroying entropy entirely, especially on long horizons.

1:26:23John Schulman: Some of what people call creativity is solving hard search problems: move 37, or a poem that satisfies many constraints. AI is obviously going to be extremely good at that if trained for it. But there's another sense in which the diversity of outputs drops a lot after RL and models develop tics. They seem good at writing, but distributional analysis shows them reusing certain themes and the same character names all the time. You don't get the diversity of human authors; you get one really good style. RL has cut that diversity a lot. And since so many people distill, mostly from Claude, all the open-weight models write the same way as Claude and share its tics. It's somewhat concerning to me that a monoculture is emerging.

1:27:54Beren Millidge: I don't think that's fundamental to RL as a method, or to distillation, which is just training on data. If the data isn't broad, that's a problem with the data, not with the method. Much of RL's entropy collapse comes from exploiting fairly simple verifiers without a huge diversity of environments. Writing is presumably graded by some judge that has its own tics; the model learns to reward-hack the judge, and that's why it collapses. That's a problem with the judge, not with RL.

12

Timelines: two to ten years

A drop-in remote worker in about 1 to 3 years; 10x for AI researchers in two years per Schulman, 5 to 10 per O'Neill; superintelligence in 3 to 10 years, with wide disagreement.

1:28:32 · Rapid-fire timelines

1:28:48Dwarkesh Patel: Rapid-fire timelines. By when can a user hire a model as a drop-in remote worker for all kinds of white-collar work, not just coding but video editing, law, paralegal work, with full computer use, a month of seamless learning and operation, and complex projects involving other people? Everything a human worker could do over a month.

1:29:22Charlie O'Neill: If it's forced to use a browser, rather than the firm making information programmatically accessible, maybe a couple of years. If not, if it can send Slack messages and so on, I'd still say around a year.

1:29:39Beren Millidge: Maybe three years for full generality. But as Charlie says, many people will make their organizations easier for AIs to use, so you get 80–90% of the way there before that. The rest is a long tail of miscellaneous things some human can do that will take the models a while. It comes down to how quickly we solve this kind of online learning, and whether compaction and writing files to itself get you 80–90% of the way. That's my big uncertainty; I really don't know.

1:30:14Charlie O'Neill: Something it won't be good at: if I have to yell at someone to get something at work, or really push someone to get something done. The model is going to be too nice.

1:30:27John Schulman: Human remote workers vary widely in quality. Hire someone off Upwork for a software project and it's often hard to get them to do a good job or take in your feedback; in some cases the pre-AI version of this was worse than what today's AI gives you. We already have it for some lower-quality work, but not at human level for higher-quality work. I basically agree with Charlie and Beren that we'll have an okay version of this form factor in a year or so, very good at some things, weaker at others, and improving from there.

1:31:33Charlie O'Neill: We keep moving the goalposts based on the very long tail. This year I just told Codex to go get everything I needed for my taxes and send it to the accountant. It had to click through and download a massive list of things. It did it, and it was perfect. A lot of this it can already do.

1:32:21Dwarkesh Patel: Next: a 10x productivity uplift for you as AI researchers. If a breakthrough takes you a year now, you'd make one every month.

1:32:31Charlie O'Neill: Somewhere between 5 and 10 years? I was picturing normal white-collar work over a month for the remote worker; past two months it starts to diverge.

1:33:00John Schulman: I would say two years.

1:33:09Beren Millidge: I can see that. For coding it's already definitely more than 10x, so if it can run even one or two loops of experimental feedback, that's massive already. Then AI progress won't be bottlenecked on researchers running small experiments, but on other things.

1:33:46Dwarkesh Patel: Sure, but those happen 10x faster, which is a huge deal, and it brings the next thing, a 100x speed-up, sooner.

1:33:53Charlie O'Neill: I'm happy to take a bit longer on that one. The crux is my capacity to absorb information and make the Bayesian-optimal decision on the next experiment.

1:34:05Beren Millidge: I'm assuming you can delegate some of that to the AI, which is getting decent at deciding: it runs an experiment, gets a result, runs the next. If it can run two or three in a row without crashing, that's a big uplift.

1:34:24Dwarkesh Patel: Final question: an AI that dominates top human experts across every field of work that can be done on a computer, all cognitive work, including tasks that take three years. Basically ASI.

1:34:59John Schulman: I'd say 3–4 years. AI research is getting the most attention and a lot of energy, and it's not one of the hardest things for AI, since it involves a lot of code and maths, where models are really good. Things involving 3D, spatial and physical work, like mechanical engineering, which isn't getting the most attention now, may take a little longer.

1:35:28Dwarkesh Patel: But it includes fields with relatively little data by nature, where it has to learn on the fly, like becoming superhuman as an engineer at TSMC.

1:35:42John Schulman: Then you'd have to assume you can give the AI the same onboarding material, and something about longer-horizon learning has to be solved.

1:36:00Charlie O'Neill: I'd say 5 to 10. I think automating AI research is basically ASI-complete: there are many things where, even with a memory system outside the model and somewhat longer context, even if you could research the information or write notes yourself, you'd need more than today's million-token context window.

1:36:23Beren Millidge: I kind of agree on the 5-year range, at least for what the labs are focusing on. But there'll be a long tail the AI could learn in theory that no one has bothered with or allocated compute to, so matching literally every human expert may take longer. It doesn't need to learn a new domain as fast as a human, though, because it will have vastly more experience than any human.

1:36:51Dwarkesh Patel: Thanks so much, guys. This was a great format for experts to disagree, debate and work through things together.

Where Indigo landsFurther

Indigo's conclusion

The practitioners' debate version of “capability is a bounded exponential”: RSI will probably come, but objective-setting, continual learning and sample efficiency hold it back; no fast takeoff. The question of speed stays open.

What to remember

  1. The bottleneck moves from compute to setting objectives and continual learning: if 2036 looks normal, the likely reason isn't models that aren't smart enough but not knowing what to do and failing to learn in deployment.
  2. Alignment is the final job: the longest-lasting human role is defining objectives and deciding what we want; technical work can be automated, choosing the right objective won't be handed over soon.
  3. Continual learning breaks in the details, forcing retraining; a cumulative task like RSI is easier than real work whose distribution keeps shifting.
  4. Distillation works against centralization: distillable capability can't be defended, so value moves to unique real deployment data, verticals like Composer and Harvey.
  5. RL works because mid-training does about 80% and RL tunes the policy with one high-signal bit; it extends task length rather than crossing domains.

Claims you can check later

ClaimWhoWhen we will knowHow firm
A general drop-in remote worker: a month of seamless learning, full computer use, working with peopleO'Neill, MillidgeAbout 1 year (not forced through a browser); fully general in about 3 yearsFirst-hand; Millidge says organizations will reshape themselves for AI, getting 80–90% of the way first
AI researchers get a 10x productivity upliftSchulman (Millidge agrees)2 yearsFirst-hand; O'Neill says 5 to 10 years
Superintelligence: dominating top human experts in every field doable on a computer, long tasks includedSchulman 3 to 4 years; O'Neill 5 to 10; Millidge about 53 to 10 yearsFirst-hand, wide disagreement
Automating AI research is basically as hard as superintelligenceO'Neill (answering the host)First-hand judgement
How long models can keep working doubles every 3 months (EdgeBench)O'Neill, citingAlready happeningSecond-hand citation; direction checkable
At small scale, data explains about 12.0x of pre-training compute-efficiency gains and architecture about 3.7xDwarkesh (with Jerry Han at Princeton)DoneFirst-hand experiment at very small scale; whether it holds at scale is unknown
Parameter counts won't jump in the next few years, inference efficiency first; frontier models are far smaller than the rumored 10 trillion parametersO'Neill (Schulman expects models to keep growing)Next few yearsFirst-hand judgement

Back on the long-running theses

confirms

AI capability is a bounded exponential: paradigm shifts don't come from hill-climbing The strongest frontline debate support yet: several of its limits confirmed unprompted, and the bottleneck recast from compute to what to do and whether models can keep learning.

confirms

What you own is not the model: value moves up to what cannot be rented Distillation against centralization, deployment data mattering more than environments, goal-setting as the last human job: all three push value toward what is unique and can't be rented.

confirms

Whether verifiable domains generalize The gap between automated research and open-ended science is the training-side version of this boundary; RL extending task length without crossing domains explains its mechanism.

confirms

Nathan Lambert, Lossy self-improvement O'Neill's continual learning breaking in the details and his Moore's-law analogy are dense first-hand engineering evidence for Lambert's “held back by friction”.

conflicts

Jakub Pachocki, An Alien Mind Pachocki strongly expects RSI; this panel talks mostly about friction, with timelines from 3 to 4 years to 5 to 10. The question of speed stays open.

adds to

Furong Huang: self-improving agents “Self-improvement must prove later work is better” is the same crux as whether AI can set its own objectives without going off the rails.

adds to

a16z and Gavin Baker, Demand is outrunning supply Two sides of one week: a16z argues from demand that supply is short; this panel argues from engineering that the bottleneck is objectives and continual learning.

What would change my mind

the next discontinuity adds to the current paradigm instead of tearing it down, and the paradigm connects the dots on its own.

Finished. Indigo's take on this piece is in two places: