5:05Dwarkesh Patel: 说到底,唯一的指望是深度学习压根到不了「至少在研发上碾压人类、包括提出新范式的能力」那一步。可如果把 2012 年到现在的进步直接延续下去,我知道这背后是巨量算力的扩张,它要是没能至少在研发上碾压人类,尤其是在未来几年,那才奇怪。Ryan Greenblatt 最近上节目时提过一点:AI 越来越强,就能在模拟环境上取得进展,而这些环境奖励的不只是 AI 研发能力,还有整体的科学能力,所有实验室和很多创业公司都在瞄准这个。另一个直觉泵是 80 年代以来国际象棋程序的 Elo 分:上升非常线性,但跨过人类区间时是一个巨大的不连续,从人类专家总是赢 AI,变成人类专家再也赢不了。到目前为止 AI 对经济的最终影响不大,是因为它相对人类的 Elo 还在慢慢爬。
6:40Beren Millidge: 我同意,那会非常意外。唯一不发生的可能,是进步恰好在跨越之前停在渐近线上,而在我看来,我们已经很接近开始跨过人类的 Elo 区间了。要落到你说的「2035 年一切如常」,唯一的另一条路是对 AI 的激烈监管。其实我觉得这比技术原因更有可能。
10:45Dwarkesh Patel: 有一个直觉泵,说明为什么哪怕不扩大 AI 劳动之外的投入,也可能很快出现奇点:每做一个七位数美元的实验之前,都先花同样多的算力在 AI 劳动上,等于有一群自动化的你们,花一个世纪去想最优的实验是什么,做小规模消融,建立整整一个世纪的理论;实验做完,再花一个世纪分析结果、决定下一个实验。
11:34John Schulman: 想得足够透,其中一些事大概本来就能预见。很可能有某种巧妙的小规模实验,能让人建立起可以推广到大规模实验的理论。所以我认为,研究能做得多好,我们离天花板还远得很。我可以想象 AI 做大量分析和理论构建,花在这上面的算力和花在实验本身上的差不多。
12:17Charlie O'Neill: 在目标指定清楚的时候,确实有很具体的例子。思考能做的,只是根据形成先验之后拿到的比特去更新后验;光靠想,得不到任何新比特。但当目标清楚、数据就摆在那里时,我预计当前范式里会有很大的加速。比如让 AI 去看 Kaplan 的规模定律,它现在就会发现:他们拿中间检查点来用,却没考虑退火,所以结论是错的。这本来能提前好几年发现,光这一个观察就能省下一两年的进展。muP、学习率随模型规模怎么变、模型宽度也很重要,这些都一样:只要多想一想,很多东西都能倒推出来,把低垂的果子摘掉。如果任务只是「把现有目标最大化」,我预计会有 10 倍的加速。但我看不出这怎么能泛化到一开始就提出正确的目标。光靠想,买不来正确的目标。
13:34Beren Millidge: 这正是当前 AI 要实现任何极快 RSI 的关键问题:AI 能在多大程度上泛化到学会自己定目标?要有一个自我推进的自动循环,AI 得能提出目标、优化、弄明白、再提新目标,而且很长很长时间里都不能跑偏。回到莫拉维克悖论,这里也许又是一个:我们觉得这种自主,自己想好该做什么、再去做、如此循环,非常容易,因为我们一直在做;进化当然得造出能长时间独立生存的生物。但出于某种原因,这对 AI 可能恰恰很难,就像运动控制对它很难、数学却很容易,和我们正好相反。
15:55John Schulman: 我会说,人类最后的工作,或者说留存最久的角色,是定义目标、决定我们到底想要什么。比如 AI 助手该怎么表现、什么叫有帮助、做基于人类反馈的强化学习时目标是什么,这是一类;后来写宪法、写模型规范,又是一类。就算 AI 能做所有技术活,这类事我们仍然要做很多,要决定我们真正想要什么。
21:29Beren Millidge: 这正是 AI 能帮大忙的地方。看看前沿的训练流程,或者中国模型在论文里写出来的做法:他们先从某处拿到种子提示,来源是人加上这类数据,然后用自己的模型或其他前沿模型,从种子提示合成出覆盖面极广的数据。收集提示分布、搭建环境,很大一部分都能自动化,模型越好,人需要提供的比特就越少。
22:51Beren Millidge: 讽刺的是,这对蒸馏方反而比对前沿实验室容易。蒸馏方只要说「我要一个好政治家」,然后去问前沿模型;前沿模型已经知道怎么当好政治家,直接生成轨迹就行。而如果你要造第一个做到这件事的模型,就得想办法真的拿到政治家每天在做什么的数据。说「我要这样的东西」,再让 AI 生成十亿个变体,远比一开始把这个东西造出来容易。
30:30Dwarkesh Patel: 这对模型代际之间拿到不陈旧的新数据也很好。说到底,这基本就是实验室内部的持续学习:通过环境和 RLHF 这类工作,把最近三个月的 AI 研究进展蒸馏回模型里。
30:41Charlie O'Neill: 而且它确实就是在蒸馏。这也许就是为什么我们有些人觉得它是渐近的。你总是在追最近三个月的进展,这些进展当然有 AI 的贡献,但回路里仍然有人,你只是在一点一点地逼近人类研究者发现的、能做到的东西。
31:04Beren Millidge: 不过我要说一点:只在轨迹上蒸馏,显然永远超不过轨迹本身。但环境可以远远超出人能做到的范围:设计一个没有人能解开的环境很容易,AI 照样可以去试着解。这才是超越人类 AI 研究的路径。尤其在 AI 研究里,定义目标非常容易:比如要求损失降到 1.3,现在没有人做得到。这是一个极其可度量、可验证的任务。
33:52Dwarkesh Patel: 退一步说,我理解的 AI 研究接下来的计划是这样,你们看对不对:在几百个领域、几百万个各不相同的环境里放大可验证奖励的强化学习训练。最后出来的是一个学会了一些基本技能的智能体:能坚持,能对信息和上下文分轻重,最终能端到端地和其他智能体配合。这样的智能体在上下文里会非常高效地学习。最后出来的东西,基本上就像一个能工作一周或一个月的即插即用远程员工。第一,你们同意这就是实验室在下的赌注吗?第二,这够不够:在数据中心的模拟环境里学会怎么学习,然后部署到真实世界,却并不真正从真实部署中学习,只从模拟环境里学到这些元技能?
45:25Dwarkesh Patel: 可这对从模拟到现实不是更大的问题吗?任务越长,越难在数据中心里模拟。在我看来,就连编程也已经到了这个地步:没有哪个长达一年的编程任务,最后不需要和客户沟通、和公司或用户打交道。想想我们希望 AI 能做的全部事情,超级智能最终应该能经营一家企业,或者新创一家企业并让它盈利,或者在市场上靠日内交易赚钱,或者打赢一场官司。这些都很难在数据中心里模拟,和真实世界打交道本来就是学习的一部分。也许靠从模拟到现实的迁移就能学会;但也可能必须从这类互动中更新权重才能变好。如果迁移不够强、确实需要更新权重,那模型样本效率低,也许就是更深层的问题。我之所以好奇,是因为默认情况下,我看不出未来 10 年怎么会不出现某种疯狂的递归自我改进。唯一可能让它不发生的原因是:就更新权重的样本效率而言,模型似乎远远落后于人类。比较一个人从出生到成年看到的数据量,和一个模型从冷启动到训练完成看到的数据量,模型很可能落后人类百万倍。所以,第一,模拟环境能不能很好地迁移到我们想让 AI 在真实世界里做的那种极长、极复杂的真实事情?第二,如果不能,模型样本效率低这件事,是不是就会反咬我们一口?
57:08Beren Millidge: 我认为主要是技术问题,肯定不是容量不够。如果你有某个模型和所有这些数据,拿一个大小完全一样的模型,把这些东西都放进中训练、从零预训练一遍,它会更好。我觉得今天很多事情就是这么做的。有一个很实在的瓶颈,让我们没法永远继续训练同一个模型,而是得把旧模型的数据全拿来、从零训练一个新模型。这正是 Charlie 说的:可塑性下降和灾难性遗忘的某种组合。如果天真地在分布变化的数据上训练,因为你一边训一边加新数据,就会搅乱数据分布,旧的东西就被忘掉了。我们其实没有好办法阻止这种事发生。
59:13Beren Millidge: 也许在那个规模上可以。就像 Charlie 说的,持续中训练肯定能做很长时间,也可以回退到某个检查点、给它新的中训练数据。但同时,这没法无限做下去。如果一直持续训练同一个基础模型,它到某个点就会停在渐近线上,学不进新东西了。这就是为什么大家最后还是会训练新的基础模型,不然一直在同一个基础模型上做中训练就好了。
1:21:48Dwarkesh Patel: 我能问几个非常基础的问题吗?这个回答说得通,也许还有实证研究表明确实是这样。可我一看模型本身……我不知道发生了什么。也许你们能告诉我,过去一年 AI 进步的基础是什么。也许只是把那些本来就会正确思考的策略加权提升了。但从质上看,模型的能力强了太多。也许这两者并不矛盾,但按这种看法,强化学习的作用相对很小,这和模型实际上在质上获得的能力,怎么对得上?
1:26:23John Schulman: 人们说的创造力,有一种其实就是解难的搜索问题。第 37 手显然是一个例子,写一首满足一大堆不同约束的诗也是。只要为此训练,AI 显然会在这上面极其擅长。但还有另一个方面:强化学习之后,模型输出的多样性低了很多,还养成了各种口癖。模型看起来写作很好,可一做分布上的分析,你会发现它们一直在重复某些主题,一直用同样的人名。你得不到人类作者那样的多样性,你得到的是一种非常好的风格。所以我认为这种多样性确实被强化学习削减了很多。其实,既然前面聊到蒸馏,现在正在发生的一件事是:太多人在蒸馏,主要是从 Claude 蒸,结果所有开放权重的模型都写得和 Claude 一样,带着同样的口癖。一种单一文化正在浮现,这让我有点担心。
1:35:42John Schulman: 那你就得假设能给 AI 同样的入职培训材料,然后还得在更长时长的学习上解决点什么。
1:36:00Charlie O'Neill: 我会说 5 到 10 年。我觉得自动化 AI 研究基本上就和实现 ASI 一样难:世界上有太多事情,就算模型外面有某种记忆系统,就算上下文长度再长一点,就算你能自己去查资料、写笔记,从根本上说仍然需要比今天一百万 token 更大的上下文窗口。
1:36:23Beren Millidge: 我大致同意 5 年这个范围,至少对实验室正在专注的东西是这样。但我觉得会有一条长尾:理论上 AI 可以去学,但没人费心去做,也没分配算力。所以要比得过每一位人类专家,可能要更久。不过它不一定需要像人一样快地学会一个新领域,因为 AI 会比任何人都有多得多的经验。
AI researchers debate how close we are to recursive self-improvement
John Schulman, Beren Millidge, Charlie O'Neill · YouTube · 2026-09-12
Three frontline researchers debate how close recursive self-improvement is: it's coming, but held back by engineering friction such as setting goals and continual learning.
Part 1 of 12 · 0:01
What's missing isn't intelligence but the next discontinuity
If 2036 looks normal, the likely reason isn't models that aren't smart enough but stalled generalization and continual learning. Moore's straight line lived on discontinuities, and today's paradigm may not find the next one itself.
What's missing isn't intelligence but the next discontinuity
If 2036 looks normal, the likely reason isn't models that aren't smart enough but stalled generalization and continual learning. Moore's straight line lived on discontinuities, and today's paradigm may not find the next one itself. Read this part →
Clear objectives mean 10x; setting them is the hard part
With a clear objective, the current paradigm could speed up 10x; the right objective can't simply be thought up. The longest-lasting human job is defining the objective and deciding what we want. Read this part →
Whatever RL teaches can be distilled cheaply, and the prompt distribution is the key. Frontier labs may have no edge in training environments; real deployment may matter more. Read this part →
Automated researchers will be trained with human feedback plus practice environments. Today labs distill the last three months of progress back into the model, which is why it can feel asymptotic; environments can go beyond what humans can do. Read this part →
The labs bet on training persistent agents that triage well across vast numbers of simulated environments, patching domain by domain. With low sample efficiency, simulation is the only option, and real-time human interaction is hard to simulate. Read this part →
Deployment data already flows back generation by generation. RSI is cumulative, each discovery held once found; law-firm work, whose distribution keeps shifting, is harder. Taste might be learned from short episodes. Read this part →
Companies won't let providers learn from their deployments, pushing toward swappable modules. At scale the loop works; for one model updated again and again, catastrophic forgetting sets in and you end up retraining from scratch. Read this part →
Each rung of the environment ladder is harder to climb, and even the whole world lacks the bits the frontier needs. At small scale data explains about 12.0x of efficiency gains and architecture about 3.7x, though architecture opens new regimes. Read this part →
Inference efficiency comes first, so parameters may plateau for a few years; Schulman expects growth anyway, depending on the scaling laws. The beautiful straight lines hide a lot of complexity. Read this part →
Mid-training often gets the model about 80% of the way; RL tunes the policy with one very high-signal bit. What it delivers is generalization over task length, not across domains. Read this part →
RL hasn't killed creativity; models once found several zero-days at a time. But output diversity has dropped, and with everyone distilling Claude, open-weight models increasingly write alike. Read this part →
A drop-in remote worker in about 1 to 3 years; 10x for AI researchers in two years per Schulman, 5 to 10 per O'Neill; superintelligence in 3 to 10 years, with wide disagreement. Read this part →
Indigo's conclusion
The practitioners' debate version of “capability is a bounded exponential”: RSI will probably come, but objective-setting, continual learning and sample efficiency hold it back; no fast takeoff. The question of speed stays open.
How to read this A roundtable of three frontline training researchers, chosen, the host says, because they're at relatively open labs and can speak on the record. Nobody is selling a product; they undercut one another and often admit uncertainty, a rare signal of honesty. Overall, RSI is coming, but a pile of engineering friction is slowing it down.
What to remember
The bottleneck moves from compute to setting objectives and continual learning: if 2036 looks normal, the likely reason isn't models that aren't smart enough but not knowing what to do and failing to learn in deployment.
Alignment is the final job: the longest-lasting human role is defining objectives and deciding what we want; technical work can be automated, choosing the right objective won't be handed over soon.
Continual learning breaks in the details, forcing retraining; a cumulative task like RSI is easier than real work whose distribution keeps shifting.
Distillation works against centralization: distillable capability can't be defended, so value moves to unique real deployment data, verticals like Composer and Harvey.
RL works because mid-training does about 80% and RL tunes the policy with one high-signal bit; it extends task length rather than crossing domains.
What would change my mind
the next discontinuity adds to the current paradigm instead of tearing it down, and the paradigm connects the dots on its own.
How to read this
A roundtable of three frontline training researchers, chosen, the host says, because they're at relatively open labs and can speak on the record. Nobody is selling a product; they undercut one another and often admit uncertainty, a rare signal of honesty. Overall, RSI is coming, but a pile of engineering friction is slowing it down.
What's missing isn't intelligence but the next discontinuity
If 2036 looks normal, the likely reason isn't models that aren't smart enough but stalled generalization and continual learning. Moore's straight line lived on discontinuities, and today's paradigm may not find the next one itself.
00:00 · If 2036 looks normal, what went wrong?
0:01Dwarkesh Patel: Today I'm talking with three AI researcher friends I learn from every time we talk, who happen to be at somewhat open labs and companies, so they can say things on the record. Beren Millidge is CTO of Zyphra, which develops open-source models. John Schulman is chief scientist at Thinking Machines, previously a co-founder of OpenAI, and led the RLHF work that led to ChatGPT. Charlie O'Neill is head of model training at Baseten. First question: if it's 2036 and we don't have billions of superintelligences that have radically transformed the world, what is the most likely technical reason, setting aside political shocks, a war or a ban on AI?
1:16Beren Millidge: There's a classic pattern, almost Moravec's paradox: we think that if AI can solve hard maths problems or win at chess it will be amazing; then it does, and the impact is smaller than expected. If that continues and the true spark of generalization never comes, AI could end up extremely good at everything people put into a benchmark or an environment, while some persistent sim-to-real gap blocks everything else. I think that's unlikely; we already see this kind of generalization, even from RL. But if meta-learning is ridiculously hard to generalize and we don't solve continual learning, that would be my default scenario.
1:50John Schulman: I agree. Humans still have a lot of advantages over models. Each new model catches up in some areas, but you get bottlenecked wherever the model is weaker, has worse judgment or can't check itself well enough. There's a cycle that keeps repeating: a new model comes out, people are blown away and say this is AGI, and after a month or so it starts to feel dumb. That cycle might keep going, and it's hard to predict how many times. Right now capabilities don't grow explosively because research and engineering still hit enough bottlenecks: even if a model writes far more code than a person, it doesn't make you 100x more productive.
2:55Charlie O'Neill: For me the question is how far the current recipe, transformers plus RL, is from the global optimum of a learner you could put on a chip. People imagine that once an agent is even 0.1% better than all humans at AI research, running hundreds of thousands or millions in parallel, faster as chips speed up, outweighs every other bottleneck, and you get a very fast takeoff. But think about Moore's law: a beautiful straight line that held for a very long time, kept going by many discrete discontinuities and innovations. The same happened with LLMs: pre-training scaling hit diminishing returns, then RL came along and gave a new curve, so it kept looking like a straight line. If we need another of those discontinuities, I'm not sure training LLMs on RL environments, even RSI-targeted ones, can discover it. If not, we'll probably hit an asymptote.
4:25Dwarkesh Patel: Do you think that discontinuity will be harder than anything since 2012?
4:34Charlie O'Neill: If we knew, we'd be able to implement it. But we should distinguish a discontinuity that adds to the current paradigm, something cumulative beyond RL whose dots the models might be able to connect, from the question of how far we are from the global optimum: do we have to go back and throw out gradient descent and neural nets altogether? If it's that far away, I don't think scaling the current paradigm, however many LLMs you run, can discover it.
5:05Dwarkesh Patel: The only hope, really, is that deep learning can't get us to an AI that dominates human R&D, including the human ability to come up with new paradigms. But if you just extend the progress since 2012, and I know it's been powered by huge compute scaling, it would be weird if it didn't get to dominating humans at least in R&D, especially over the next few years. Ryan Greenblatt made a point on the podcast recently: as AIs get more capable, they can make progress on simulations that reward getting better not only at AI R&D but at science generally, and every lab and many startups are targeting that. Another intuition pump is the Elo of chess bots since the '80s: a very linear rise, but a huge discontinuity as they cross the human range, from experts always winning against AIs to experts never winning. So far AI hasn't had much end economic impact because it's still slowly rising in Elo relative to humans.
6:40Beren Millidge: I agree it would be very surprising. The only way it doesn't happen is if progress asymptotes just before, and in my opinion we're already pretty close to crossing the human Elo range. The only other way to end up in your scenario, where 2035 looks normal, is dramatic regulation of AI, which I actually see as more likely than a technical reason.
02
Clear objectives mean 10x; setting them is the hard part
With a clear objective, the current paradigm could speed up 10x; the right objective can't simply be thought up. The longest-lasting human job is defining the objective and deciding what we want.
07:03 · Clean objectives versus open-ended science
7:08Charlie O'Neill: There are different kinds of research. There's autoresearch-style work, where the objective is already cleanly specified and you optimize it; everyone pictures that pushing pre-training loss down and environment rewards up leads to improvement. Maybe what Ryan means is the far more open-ended science that paradigm shifts need, where we can't specify the objective, and the AIs certainly can't either. We have to be really careful about how we specify objectives.
7:49Dwarkesh Patel: John, you were in the trenches back then. Presumably a big breakthrough was realizing that next-token prediction is the thing; in 2014 you wouldn't have thought the nanoGPT speedrun was what to optimize. Now that we're in this paradigm, you would think to speedrun it and have AIs get really good at that. But maybe there's a next inner loop the AIs wouldn't anticipate. There's an outer loop of revenue that should eventually be strong, but it's very slow.
8:29John Schulman: In fact, I remember in the early OpenAI days having the intuition that just minimizing log loss wouldn't get you to intelligence, because the important bits are such a small fraction of the loss that they'd be drowned in noise, so we needed better objectives that put more weight on the important things. You can make all sorts of arguments: humans probably don't learn to model everything in their environment; most people can't reproduce a photorealistic scene they've looked at; so we must need a better objective. Then it turned out it just worked anyway.
9:33Dwarkesh Patel: And as you've pointed out, even in AI research today the inner loop of post-training benchmarks doesn't necessarily translate into what users like.
9:42John Schulman: Yes. The whole field relies heavily on generalization, and it's very hard to predict when you'll get it, especially out of distribution. Training on the task you care about helps, but the most important advances are often kinds of generalization we have no right to expect: from naive next-token prediction to tasks that need deep understanding of the input, or to rare skills picked up in pre-training; and from verifiable tasks to less verifiable ones, which there's no a priori reason to expect either.
10:45Dwarkesh Patel: One intuition pump for a very rapid singularity, even without scaling the inputs other than AI labor, is this: before every seven-figure experiment, you spend an equivalent amount of compute on AI labor, automated versions of you spending a century thinking about the optimal experiment, running small ablations, building a century's worth of theory; and after it, another century analyzing what happened and choosing the next one.
11:34John Schulman: If you think hard enough, you probably could have predicted some of these things. There's probably some clever small-scale experiment that lets you build a theory that generalizes to the large scale, so we're nowhere near the ceiling of how well research can be done. I can imagine AI doing a lot of analysis and theory-building, spending as much compute on that as on the experiments themselves.
12:17Charlie O'Neill: There are concrete examples when the objective is well specified. All thinking can do is update your posterior on the bits you've already got; you can't gain new bits just by thinking. But when the objective is clear and the data is sitting there, I expect a big speed-up within the current paradigm. If an AI had looked at the Kaplan scaling laws, it would have noticed they used intermediate checkpoints without accounting for annealing, so the result was wrong. That would have been caught years earlier and saved a year or two of progress. The same with muP, how learning rate scales with model size, and the importance of width: you could back out a lot of this and pick the low-hanging fruit. I'd expect a 10x speed-up if the job is just maximizing the objective we already have. But I don't see how that generalizes to coming up with the right objective in the first place. Thinking alone doesn't buy you that.
13:34Beren Millidge: That's the key question for any very rapid RSI from current AIs: how well can AIs generalize to learning their own objectives? A self-propelling loop needs the AI to propose objectives, optimize them, propose new ones, and not go off the rails for a very long time. Maybe this is another Moravec's paradox: autonomy, deciding what to do and then doing it in a loop, feels easy because we do it all the time, and evolution had to build creatures that survive on their own for long periods. It might be really hard for AI, the way locomotion is hard for it while maths is easy, the reverse of us.
14:00Dwarkesh Patel: But doesn't the growing time horizon suggest otherwise?
14:25Beren Millidge: Exactly, and I agree there's no obvious evidence for this. The fact that our agents are now very persistent is evidence against it. But if it is hard, it would be one reason we don't get an immediate takeoff.
14:35 · The last human job: deciding what we want
14:39Dwarkesh Patel: Looking back from 2012, or from when you started doing research, which of all the innovations since, engineering or conceptual, looks like the last thing humans will have to do before AI fully automates AI R&D?
15:05Beren Millidge: Probably iteratively asking the right questions. Even if the AI can run any experiment, someone has to decide which experiments to run. Right now AIs are much worse at that than at coding the experiment: when we talk about research, they propose a grab bag of very tiny steps.
15:24Charlie O'Neill: Or the jump from DeepMind's "we'll solve intelligence by learning to play games at a superhuman level" to one researcher, Radford, deciding to just predict the next token of a very wide swath of data. Even after Radford found that, it took a while before people scaled it up, because we first had to come up with scaling laws and the idea that you could predict these things reliably.
15:55John Schulman: I'd say the human role that lasts longest is defining the objective and deciding what we actually want: how AI assistants should behave, what it means to be helpful, what the objective is when we do RL from human feedback, and later constitutions and model specs. Even if AIs can do all the technical work, we'll still have to do a lot of that.
16:33John Schulman: Alignment is sort of the answer. But alignment decomposes into specifying the objective, figuring out what it should be, and then achieving or optimizing the objective you've defined. The first isn't going away anytime soon. That's why a post-training team needs so many people: there are many areas where someone has to work out how the model should behave, and that is very hard to automate.
03
Distillation keeps anyone from pulling away
Whatever RL teaches can be distilled cheaply, and the prompt distribution is the key. Frontier labs may have no edge in training environments; real deployment may matter more.
18:39 · Why no single lab has run away with it
18:51Dwarkesh Patel: Why isn't there huge consolidation among model providers? So much points toward centralization. Over the years, is there something that prevents it?
18:54John Schulman: I think distillation is the main force against centralization. Anything that can be learned through RL can be distilled very easily, because it's a small number of bits you can learn from a small amount of data. If you can get trajectories showing a behavior, you can distill it. There's also the possibility of company-specific models, where a company learns from deployment and keeps improving its own model, provided by today's oligopoly or by some smaller company. That would change the game a bit.
19:56Beren Millidge: And continual learning doesn't stop distillation. If your model improves every day, people can distill it every day. The loops can run at the same pace.
20:10Dwarkesh Patel: But to copy a model's behavior, don't you need to know the right distribution of prompts?
20:18John Schulman: Yes. For distilling with supervised learning, the prompt distribution is extremely important. Distilling all the useful capabilities is very non-trivial even with full access and the chain of thought, because you have to prompt the model with a really wide distribution of realistic prompts. It has recently come out that some Chinese companies are probably using router services, proxies that let people in China use US frontier models that are otherwise blocked there, mostly for coding. Those services are collecting and selling some of the data, which is very useful for distillation because it gives you the perfect prompt distribution.
21:29Beren Millidge: This is where AI helps a lot. In the frontier pipelines, and in what the Chinese labs describe in their papers, they get seed prompts from a mix of humans and this kind of data, then use their own or other frontier models to synthesize vast coverage from them. You can automate much of the prompt gathering and environment creation, and humans need to provide fewer and fewer bits as models improve.
21:57Dwarkesh Patel: But you still seem bottlenecked by having a service with users going through it. The user says: make me an app like this; that didn't work; actually add this feature; no, let's step back and do something else. Capturing that whole trace is the thing. And if you could have produced it anyway, you'd just have RSI.
22:25Beren Millidge: Not necessarily; it helps a lot, but in theory you can just think about what users want. Ultimately, a fully automated loop is RSI: the AI decides the data and the training. But it depends on how much human information you need. At some point, if you want traces that look like this, you prompt the model and it produces a pretty good approximation.
22:40Dwarkesh Patel: What if it's "make me a really good politician", and it has to anticipate from scratch how a debate in the Senate would go?
22:51Beren Millidge: Ironically, that's easier for the distillers than for the frontier labs. The distiller asks the frontier model, which already knows how to be a good politician, to generate the traces. To build the first model that can do it, you have to somehow gather data on what politicians do every day. Asking for a billion variations of something is much easier than creating the first one.
23:19Charlie O'Neill: You can make a concrete prediction from the fact that the Chinese labs have this router data. What started it was me asking: isn't it weird that Sonnet 5 and Opus 5 are almost objectively worse models than GLM-5.3 and Kimi K3, even though they had not only distillation but logit distillation from Mythos? The counter was that the prompt distribution really matters; you need to see what users do. The prediction is that frontier labs don't necessarily have much advantage, if any, in RL environments now. User distribution matters for general behavior, but the best measure of a capability is the very hard RL environments built at the frontier. If Anthropic has those environments and logit distillation and still made a worse model, then maybe—
24:12Dwarkesh Patel: Then real-world deployment matters more than the environments. That's really interesting. But they had to incentivize those capabilities in Fable, the frontier model, in the first place, so it's odd they can't do it again with a smaller model.
24:32Charlie O'Neill: Maybe we're in an uncanny valley where copying the frontier model too closely fails because the student-teacher gap is too large. People say this about Opus: the difference between Opus 4.6 and Opus 5 is that Opus 5 feels like it has an AI judge checking everything it's done, which is why it uses so many tokens, but it doesn't have Fable's big-model sense of when to stop or which path is worth going down.
25:02John Schulman: I'd offer a slightly different hypothesis. Environments vary along two axes, difficulty and realism. It's comparatively easy to create lots of difficult environments, much more complicated tasks or ones that need more cleverness; call it the benchmaxxing distribution, since many prominent benchmarks are very hard, puzzle-like tasks that are easy to verify. Then there's realism: being good as a coding agent in realistic settings, with multiple back-and-forths with a human and multiple objectives. Labs crafting a behavior for the first time have to push both ways, and realism needs rubrics or human feedback feeding the reward. Naive distillation only matches the teacher on the benchmaxxing distribution; without enough environments that exercise the trickier realistic settings, the student doesn't get those abilities. Maybe the big models generalize better from narrow hard tasks to realistic ones. With a really good realistic prompt distribution you can match the big model well; with only easily verifiable tasks you match it on every benchmark but do worse on the broader distribution. That might explain something about the smaller Anthropic models like Sonnet 5, though it's hard to know how they post-train them. They may also just have got a few things wrong in a post-training stack that keeps changing and created quirks people really dislike. It's very easy to screw up post-training in ways that don't show up in benchmarks.
27:49Beren Millidge: One more basic point: the frontier labs buy all their data from big data companies, and the Chinese can buy the same data, and they do. People are annoyed about it, but with the same data plus distillation, it's quite easy to keep up.
04
Continual learning inside the lab
Automated researchers will be trained with human feedback plus practice environments. Today labs distill the last three months of progress back into the model, which is why it can feel asymptotic; environments can go beyond what humans can do.
28:06 · How the first automated AI researchers get trained
28:10Dwarkesh Patel: How will the first models capable of automating AI R&D actually be trained? There's the toy version Ryan described: have GPT-8 build GPT-3-sized models that are great at inner-loop challenges, like beating video games that need continual learning, or reaching a given loss with the least compute. John, you suggested that may not be how it happens in practice.
28:47John Schulman: We'll probably combine learning from human feedback, to absorb researchers' taste, with lots of practice environments involving multi-step research projects, and in each iteration patch whatever seemed most broken in the last one. Researchers will use the AIs a lot, notice consistent weaknesses, and fix them with human feedback or new environments.
29:30Charlie O'Neill: One way to think about it is how far back along the lineage you roll and then let it self-play. In the limit you give it a GPU and some neural nets and say: figure out how to train a model for these tasks. Today we go to the very edge of the lineage: here are the bugs Anthropic found in its training stack in the last few months, turn them into environments. You could imagine rolling back to before GRPO and asking it to discover the best way to RL models, and further back still. But we'll be so compute-bottlenecked that people will stay at the frontier, essentially diffing the bugs and improvements found since the last version and turning them into training environments.
30:30Dwarkesh Patel: Which also gives fresh, non-stale data between generations. It's basically continual learning inside the lab: distilling the last three months of research progress back into the model through environments and RLHF-style work.
30:41Charlie O'Neill: And it is distilling. That's maybe why some of us feel it's asymptotic. You're always catching up on the last three months of progress, contributed partly by AIs but still with humans in the loop, inching closer and closer to what the human researchers find.
31:04Beren Millidge: Distilling on trajectories can never take you above them. But environments can go far beyond what a human can do: it's easy to design an environment no human can solve that the AI can still attempt. That's the path to getting ahead of human AI research. In AI research especially, goals are easy to define: say the loss needs to be 1.3, which no human can reach now. That's an extremely measurable, verifiable task.
31:38Dwarkesh Patel: Or a nanochat speedrun done faster than any human speedrunner, or a 100 million parameter model that beats Minecraft; maybe that's too easy, so a much more complicated game.
31:45Charlie O'Neill: Isn't it crazy that we're calling a 100 million parameter model beating Minecraft too easy? Imagine saying that five years ago.
31:52John Schulman: A lot of research isn't hill-climbing on a well-defined goal, though. It's more like: we have an intuition about how models should be better and an algorithm that seems to go a little in that direction, so we design a task to show signs of life, and if we see them, we make successively more realistic versions. You're not optimizing the eventual production objective directly; you relax on the realism axis, find methods that work, then return to realism once the method matures. There's also research aimed at explanation and theory. We rarely have predictive mathematical theories in machine learning, but we have many informal ones.
33:27Beren Millidge: Presumably models will be trained on a mix of all these tasks, some easily verifiable, some judged by an LLM or by asking a human whether it looks reasonable, in the hope that they generalize to much vaguer, fuzzier tasks. They probably will to some extent. Whether they generalize enough for the loop to seal itself with no humans in it is unclear.
05
Learning to learn in simulation: is it enough?
The labs bet on training persistent agents that triage well across vast numbers of simulated environments, patching domain by domain. With low sample efficiency, simulation is the only option, and real-time human interaction is hard to simulate.
33:51 · The labs' bet: long-horizon RL in simulation
33:52Dwarkesh Patel: Here's what I think the plan is; tell me if you agree. Scale up RLVR training across millions of diverse environments in hundreds of domains. Out comes an agent that has learned to be persistent, to triage information and context, eventually to work end to end with other agents, and that is very sample-efficient in context. It functions like a drop-in remote worker over a week or a month. Is that the labs' bet, and is it enough: learning to learn in simulations inside a data center, then being deployed without actually learning from deployment?
35:10Charlie O'Neill: It's hard now to separate how much of the labs' effort goes to direct RSI and how much to generally intelligent models they can deploy for revenue to fund the next big training run. For the latter, yes, that's the bet, and the pattern of environments over the last few years is clear. Anthropic's lineage is the clearest example: first coding, where internet data and internal material make environments easiest to build; then, from the task horizon coding gave them, generalize. Finance next, with enormous amounts of Excel data in RL, then PowerPoint, the long tail of the working economy. That worked really well, and other labs, even open-source ones, now see it was the right bet.
36:09Dwarkesh Patel: But what follows from that? I asked Dario: if you truly expect models to learn on the job like humans, why bake in PowerPoint skills instead of expecting them to pick it up while deployed? One explanation is that models will get there soon but aren't there yet, so you amortize the skills into training. Another is that the goal isn't widely deployed work at all but RSI, and this is a way to earn revenue to pour into the RSI model; after the singularity, whatever comes out will be good at everything that bottlenecks today's models. John, how should we read so much task-specific knowledge if the path is generalization?
37:09John Schulman: If models were good enough at in-context learning, you wouldn't need to train them on finance; they could read all the books on the fly and work out everything in the relevant jurisdiction. But even a smart enough model would be more efficient at runtime with the intuitions baked into its weights by RL. In practice, providers are going domain by domain, strengthening the models in the highest-value domains, and that's one answer to why models have improved so much: they've covered many high-value domains and the most common skills.
38:11Beren Millidge: It's also not that expensive to do both. The models are massive and can afford, in parameters, to learn everything. There's likely some transfer: the general meta-learning of figuring out what matters, having taste and doing long-horizon work may generalize. And there isn't much RSI data in the world; it's hard to generate. With masses of compute and parameters, why not amortize in other data too, besides the obvious commercial reason of selling a model?
38:55John Schulman: There's also a question of whether today's sim-to-real paradigm stays dominant forever: look at what real tasks are like, build environments you can simulate in the data center, do RL on them. It has been very successful but has weaknesses, because many things are hard to simulate, especially interacting with a bunch of people in real time.
39:31Beren Millidge: Sim-to-real has to dominate while sample efficiency is low. You need thousands and thousands of interactions, and no human is going to sit in the RL loop, so you simulate. If sample efficiency improves a lot, you'd expect learning from deployment to become a much bigger part.
39:56John Schulman: You can also learn off-policy: take all the traces and potentially learn something from them without resimulating everything.
06
Cumulative tasks are easier than a paralegal's
Deployment data already flows back generation by generation. RSI is cumulative, each discovery held once found; law-firm work, whose distribution keeps shifting, is harder. Taste might be learned from short episodes.
41:19 · Half the compute teaches the model nothing
41:27Dwarkesh Patel: It's strange that 50% of compute goes to inference that doesn't directly make the model better. A key advantage digital minds should eventually have is that where a human gets 50 years of real-world experience, a model, through all its instances, gets millions of years of deployment across all kinds of economically relevant work. Right now that data isn't meaningfully helping it get better. Once models learn from it, you'd have something like a widely deployed intelligence explosion. When does this hive-mind stuff start?
42:14Beren Millidge: At a very basic level it's already happening, from one generation to the next: you can put deployment data into the pre-training or mid-training of future models, especially after filtering, judging, annotating or synthesizing it.
42:29Dwarkesh Patel: How much of the generation-over-generation improvement does that explain?
42:32Beren Millidge: Quite a bit, I think. I don't know whether the labs do it, since they say they don't train on people's data. But the Chinese 100% do; they definitely get this advantage. That's basically what distillation is: ping the models, get a slice of their deployment data, train the next generation on it. They can do it on their own models too; there's no reason not to.
42:56Charlie O'Neill: Completely agree. Zoom out far enough and it's definitely happening. The holy grail of continual learning we all picture is an organic live loop: one model has an experience and updates on the spot. Zoom in to that granularity and a lot breaks. But the big labs and the closed models are doing the slower version, and there are early signs of life with open-source models at a much faster cadence. Composer is probably a good example; Harvey is doing the same with legal agents. You build very specific environments from your task data, from what users complain about and from the feedback you extract from your deployments, and these companies can use that data far better than the big labs. Then they do a big post-train of Kimi K3, deploy it, maybe do some online learning, the way Composer basically ran REINFORCE for a long time. A human is still in the loop, deciding which signals matter and how to turn the data into environments. The cadence is longer than the one you're picturing, but it really is happening, and the loop will keep getting faster.
44:20Dwarkesh Patel: With Composer, in Cursor people press Tab or don't on the next completion, and based on that it gets better every day at predicting the next—
44:29Charlie O'Neill: That was the old Tab model. They did the same for the actual generative model. It's hard because online RL has no groups: one user says one thing and you get one rollout, so variance reduction is a big problem. Cursor's fuzzy answer was good heuristics that estimate how much better or worse than average a response was, then a big REINFORCE update. To check it hadn't got worse, if it improved on CursorBench they deployed a new model every five hours; if not, they threw that version out.
45:09John Schulman: I think the biggest problem is not knowing what the reward function should be for natural data. A superficial signal, like whether they accepted the edit, can get reward-hacked.
45:24 · Cumulative tasks, and why RSI is easier than being a paralegal
45:25Dwarkesh Patel: Isn't this a bigger issue for sim-to-real? The longer the task, the harder it is to simulate in a data center. Even in coding, there's no year-long task that doesn't eventually require talking to a client, the company or users. Superintelligence should eventually be able to run a business, start one and make it profitable, day-trade profitably or win a court case, all very hard to simulate; interacting with the real world is part of the learning. Maybe transfer from simulation is enough. If not, and you need weight updates from those interactions, the models' sample inefficiency is a deeper problem. By default I don't see how we avoid some crazy recursive self-improvement within the next 10 years. The one reason it might not happen is sample efficiency: models are plausibly a millionfold behind humans, comparing what a human sees from birth to adulthood with what a model sees from cold start to the end of training. So will simulations transfer to the extremely long-horizon, complicated real stuff? And if not, does the lack of sample efficiency come back to bite us?
47:25Charlie O'Neill: I'd split tasks by whether they're cumulative, or whether the distribution is non-stationary and you have to keep learning and relitigating things. RSI might be cumulative. It's theoretically possible to have a Python file under a million tokens that trains, from scratch, a model capable of recursive self-improvement. Every discovery is a line in the sand you hold: once you've found attention, mixture of experts and GRPO, you add them to the training stack and they stay. When OpenAI trained 5.6 Sol or 5.6 Terra, or whichever it told us about, it didn't have to rediscover attention; it basically called scripts like pre-training.sh and post-training.sh. The real world, and the reason people think so much about continual learning, isn't like that. An agent working as a legal associate at a law firm faces a very non-stationary distribution: it has to fit in its context all the constantly changing relationships between the important people there, and all the implicit ways things get done and where information lives. So tasks will split, and if the labs believe RSI is cumulative, in the sense that no brand-new architecture has to be discovered, more effort and compute will go there rather than elsewhere.
48:56Dwarkesh Patel: It's so unfortunate that RSI happened to be easier than being a paralegal.
49:08John Schulman: Today's models are weaker than humans in many ways. Some are about sample efficiency in certain regimes: models are very sample-efficient in context, but in some medium-length regime humans may update their weights more efficiently. Other weaknesses are completely different: lower diversity of thought, or being bad at certain long-horizon judgments. A lot of what people call taste is knowing what behavior works in the long run; in software engineering, which systems will stay maintainable and work well over the life of a project. Various weaknesses limit RSI, some tied to sample efficiency and some not.
50:40Charlie O'Neill: A thought experiment: give a model a context window of a trillion tokens, enough for all your experience up to, say, RLHF, with the same in-context learning ability it has at a million tokens. Is taste then solved? Would it make the judgments you did?
51:07John Schulman: It would have to be trained to learn from that context, to make the right update from it, or it would have to generalize.
51:26Beren Millidge: And you'd need trillion-length data to train it on that context; today you can't even jump from 10k to a million. But in theory, yes. It comes down to how meta-learnable taste is from shorter episodes, and there's no obvious reason it needs very long ones, because humans develop taste without many long episodes. A PhD, from first year to postdoc, is maybe five years and 10 to 30 research projects in total, yet people develop taste quickly from a short succession of small things. AI will have vastly more experience to meta-learn taste from. How well that generalizes to really long-horizon things is unsolved; we don't know.
07
Continual learning breaks in the details
Companies won't let providers learn from their deployments, pushing toward swappable modules. At scale the loop works; for one model updated again and again, catastrophic forgetting sets in and you end up retraining from scratch.
52:22 · Continual learning breaks at the micro level
52:30Dwarkesh Patel: Eventually AIs should be learning a lot from each individual deployment. Today there's a fuzzy meta-process by which models improve from deployment, but it's a weak feedback loop. Do you see rapid hive-mind learning on the horizon, and how exactly would it happen?
52:54John Schulman: Whether we get a hive mind that learns from all its deployment experience is largely about incentives rather than technology. Companies won't want the model provider learning from all their deployment, because that could erode their business advantage.
53:21Charlie O'Neill: The economics will push toward modules that get swapped in, rather than weight updates to one big shared model. LoRA is the obvious example. There's also a lot of work on fitting arbitrary context into a fixed size, the linear attention work, and cartridges, KV caches trained to be highly compressed. Companies might sign up for those because they don't change the underlying base model. Those don't really help the big labs, but I think economic pressure will force the labs down that path first.
54:13Beren Millidge: Which pressure, though? Even with cartridges or LoRAs, you can still take all the traces and dump them into the pre-training of your next generation.
54:20Charlie O'Neill: Yes, that's a more indirect form of learning for the big labs, and still very valuable to them. But I can't imagine starting with directly training one big model on all the exact data coming in.
54:43Beren Millidge: It'll go in stages, not one discontinuous moment where we suddenly fix continuous weight updates. More likely, cartridges and the like let you specialize deployments; you generate traces, put them into the model, and three months later ship a model that's better at it; then specialize and consolidate again, faster each time. Instead of a release every three months, every week, then every day, then every hour, at which point we've basically solved it.
55:02Charlie O'Neill: That connects to how far the current paradigm is from this. We've done a bit of research, and many others have too. At a really large scale, with enough noise washed out and big enough batches, the outer loop of putting data into mid-training and building our own environments does work as a kind of continual learning. But zoom in to the micro level, one model updated again and again for a law firm with a relatively small amount of data, and all the methods break down a bit. SFT on successful traces, off-policy or on-policy: after hundreds of these micro-updates you see catastrophic forgetting, loss of what was learned on top of the base model much earlier, and degraded general capabilities. On-policy distillation pushes that horizon out a little, but eventually succumbs to the same thing. RL is good at getting capabilities in but not explicit knowledge, like how this particular person at this law firm handles this particular process; getting knowledge in with RL takes a lot of compute to build the right environments.
55:46Dwarkesh Patel: Is the forgetting fundamentally about capacity or about technique?
56:30Charlie O'Neill: A bit of both. SFT and even on-policy distillation can be way too destructive. RL is nice because it changes the model very little, just nudging it within a very small loss valley, but that also limits how much RL can change the model.
57:08Beren Millidge: I think it's mostly technique; capacity is definitely there. Take a model of literally the same size and pre-train it from scratch with all that data in mid-training, and it will be better; a lot of what happens today is exactly that. Something stops us from training the same model forever rather than training a new one from scratch on the old model's data: as Charlie said, a mix of lost plasticity and catastrophic forgetting. Naively training on non-stationary data shifts the distribution and the old stuff is forgotten, and we don't have good methods to stop that.
57:46Dwarkesh Patel: So in the limit you're bottlenecked on retraining from scratch with all the new information.
57:58Beren Millidge: Which is very expensive. With real continual learning, you'd never train a new model; one model would keep learning and expanding.
58:06Charlie O'Neill: That's the question. We've pushed back how much has to be done from scratch: you can now take a pre-trained base and do very good mid-training on top, fairly continuously, plus RL from later checkpoints. That looks more like continual learning, but it's not taking the latest model, applying a few tiny updates, and never losing anything.
58:42Dwarkesh Patel: Isn't that literally what post-training already does, distilling a further-RL'd fork into a model that's been through a huge amount of training?
58:44Charlie O'Neill: It's still at a large enough scale to wash out much of the noise, and it isn't focused on one distribution, which, as Beren said, is the issue.
58:59Dwarkesh Patel: But eventually you'd be learning from billions of deployed instances at once, which should wash out noise too.
59:13Beren Millidge: Maybe at that scale. You can do continual mid-training for a long time, and roll back to a checkpoint and give it new mid-training data, but not indefinitely. Keep training the same base forever and it plateaus; it can't learn new things. That's why people end up training new bases.
08
Where the signal comes from
Each rung of the environment ladder is harder to climb, and even the whole world lacks the bits the frontier needs. At small scale data explains about 12.0x of efficiency gains and architecture about 3.7x, though architecture opens new regimes.
1:00:33 · Where the signal comes from
1:00:38Dwarkesh Patel: Let's talk about data: how much of AI progress is explained by data progress? Is there some data distribution that, trained into current architectures, would produce a superintelligence dominating human experts in every field?
1:01:00Charlie O'Neill: Including post-training data and environments? Its existence is obvious; the question is whether we can create the right environments to get there.
1:01:08Beren Millidge: In the trivial case, you train it to output the Python file that trains the actual superintelligence, memorized in the weights.
1:01:15Charlie O'Neill: There's probably a ladder of RL environments that gets you an AI researcher at least as good as a human one, but the effort to climb each rung grows roughly exponentially, and that trade-off sets how fast we reach the last rung. We're still early in environment creation and exploit many asymmetries. In some environments it's easier to go backwards than forwards: you define a complex data-generating process, keep it hidden, and the model has to spend irreducible tokens working out what it was. Others inject real-world information: Anthropic finds a bug with tens of thousands of humans and LLMs combined and turns it into a neat environment that a single LLM could in theory solve within a few million tokens. We're cherry-picking these asymmetries and counting on task-horizon generalization. At some point it hits diminishing returns: you can't always find a process that's easier backwards, so you have to build something with a long enough horizon, with humans, which is very complex, plus the compute and time for the agent to do it. I think the curve will start to flatten.
1:03:07John Schulman: I saw that someone fine-tuned the Talkie model, trained only on data up to 1930, on modern coding-agent data, and it did better than Claude 3 Opus on SWE-bench. A model with no knowledge of code, fine-tuned on a moderate amount of data, behaved better as a coding agent than a much larger pre-trained model. Once you have examples of the right expert behavior, it's surprisingly easy to copy it into a relatively weak model.
1:03:53Charlie O'Neill: A counterexample: a recent paper trained a model up to fifth-grade maths and primary-school English, then tried to RL it to late high-school and college maths. The gap was too large; it couldn't climb at all. With rungs of year 7 maths, then year 8 and so on, you could climb to year 12. Again, it's the distance between rungs and how hard they are to build.
1:04:18Beren Millidge: That's the RL signal problem. RL isn't good at exploring right now: if the model can't get it in 128 rollouts, it's very unlikely to get the signal to progress. That's why RL needs curricula and pre-training doesn't.
1:04:35Charlie O'Neill: Pre-training data is different from post-training data. Humans will be involved less and less, but you're still bottlenecked on how much signal you can extract from the real world. There's a lot of signal, people doing spreadsheets and legal work, but at today's capability frontier, how many bits in the world are really relevant to improving the model? How many new maths or coding problems are being solved beyond current models' reach? That's why diminishing returns kick in: even the world as a whole isn't giving you the bits that tip you into the next basin of capability.
1:05:30Beren Millidge: Totally agree. In pre-training the signal is already in Common Crawl; the problem is filtering out the noise, which is fairly automatable. In mid- and post-training the signal isn't in the original data at all, and no amount of filtering finds it: there's no hidden proof of a Millennium Prize problem sitting in Common Crawl. You have to get the bits elsewhere: from humans writing out their reasoning, from environments whose design and objectives humans decide, or from training on the human data that exists in deployment.
1:06:10 · Data versus architecture
1:06:15Dwarkesh Patel: How much of pre-training progress is driven by data? I did an investigation with Jerry Han, a student at Princeton: we trained every recipe from 2019 to now against every dataset from 2019 to now, pairwise. GPT-2 on the newest dataset like Ultra-FineWeb, Delphi, the newest open-source recipe, on the Pile; the whole grid, measuring how much less compute it takes to reach a given capability. Data seems to explain about a 12.0x compute-efficiency gain and architecture about 3.7x, at a very small scale. If most pre-training efficiency gains come from better data, how long can that continue, with more filtering and more synthetic data?
1:07:23Charlie O'Neill: My prior is that the low-hanging fruit is somewhat exhausted. We got the internet as one big block, and its useful content isn't growing at the same rate. There are probably a bunch of 0.1% loss drops left, but not as many as we've had. It's also interesting that you find a cumulative 33x across both. Epoch or someone estimated 3x a year since 2019, which implies something like 3 to the 7th, over 2,000x. So where is the missing 100x or so coming from? That probably tells you how much of this is post-training.
1:08:08Dwarkesh Patel: I think the explanation has to be that many efficiency gains are scale-dependent and we ran at extremely small scale. Which raises whether data gains or algorithmic gains depend more on scale; we didn't have the compute to find out.
1:08:36Charlie O'Neill: Naively, the architecture's scale dependence is fairly well known; you can fit a straight line to it. I'd have no idea how to do that for pre-, mid- and post-training data combined.
1:08:49Beren Millidge: Funnily enough, I think data matters more with scale, and architectures are a one-time thing. Quoting an X% efficiency gain is misleading: an architecture lets you reach a qualitatively new regime, and within it data is what matters. Without even GQA, doing full attention all day, a million-token context would be ridiculously expensive, so we could never use million-token data or get those capabilities. Measure at 2K context, where the architecture unlocks nothing, and data looks more important than it is. I'm not sure these gains simply multiply. And a lot of today's mid- and post-training data gets better with scale, because very long-horizon environment data needs big models to use it. Train a 100 million parameter model on SWE-bench traces and it gets nowhere.
1:10:10Charlie O'Neill: Comparisons are also harder now because many architecture changes, at Kimi or DeepSeek, aren't aimed at lowering pre-training loss but at how the models will be used in the real world. DeepSeek's compressed attention is about inference efficiency, not a fundamental trade-off improvement.
09
Will models keep getting bigger?
Inference efficiency comes first, so parameters may plateau for a few years; Schulman expects growth anyway, depending on the scaling laws. The beautiful straight lines hide a lot of complexity.
1:10:23 · Will models keep getting bigger?
1:10:32Dwarkesh Patel: How will parameter counts scale in an RL-heavy regime? Frontier open-source models have grown maybe 2x a year. If frontier closed models have 100B or 200B active parameters, does that keep doubling? Or, since you also want to save compute on rollouts, is there a threshold beyond which more parameters don't matter much? How many active parameters will a frontier model have in 2030?
1:11:11Charlie O'Neill: For the next few years we're so focused on longer-horizon RL rollouts, where inference efficiency matters a lot, and the models aren't saturated at that; the bottleneck is still the environments. So we may see a bit of a plateau. I suspect Mythos and the GPT models are much smaller than the 10 trillion parameters people talk about; you can back that out just by comparing them with open-source models. I wouldn't expect huge parameter growth in the next few years. You size the model by how much pre-training data you have and how hard your RL environments are: you want a decent pass@1 on your hardest environments, and going bigger than that just means paying for more inference than you need. So a lot depends on how fast Mercor and the in-house teams can make RL environments more complex.
1:12:21John Schulman: I'd expect models to keep getting bigger just because compute is scaling and GPUs are getting bigger, but by how much depends on the scaling laws in non-obvious ways. Now that high-quality pre-training data is running low, data efficiency will matter more than compute efficiency in choosing architectures, which may affect how sparse you make the model. We don't understand sparsity well: total parameters are a different resource from active parameters; sparsity has increased, but not necessarily without bound, and there may be a sweet spot. There's a debatable argument that sparsity hurts data efficiency because several experts may have to learn the same thing. We don't have a good enough theory to know why sparsity helps, how much, or where it plateaus. With data rather than compute on the x-axis, you simply get a different set of optimal models.
1:14:48Charlie O'Neill: I'm also not sure models have doubled every year. People have trained 1 trillion parameter models for years; there was even an open-source one, Falcon. Liam from Periodic Labs posted that an early experiment was a very sparse 1 trillion parameter model, which John says was the Switch Transformer at Google, before they went to OpenAI. It was very good at knowledge and terrible at reasoning because it was so sparse. We've been in the 100 billion to 2 trillion range for a while; it hasn't been a nice linear increase.
1:15:26Beren Millidge: Two things. As Charlie said, inference efficiency for RL rollouts will push active parameters down a lot. Total parameters depend heavily on hardware: serving multi-trillion-parameter models needs very high memory bandwidth and VRAM. People still use a lot of H100s; as everyone moves to GBs and then Vera Rubins, we'll be able to serve and run RL inference at larger scale. And naively, larger models are much more sample-efficient: even unsaturated, bigger generalizes better and reaches a better loss on the same data. Right now data isn't the constraint, compute is, so we get smaller, inference-efficient models. If compute stops being the bottleneck, we may go back to larger, undersaturated models.
1:16:26Dwarkesh Patel: Under the basic Chinchilla law, maxing out parameters barely reduces the data you need; even infinite parameters cut it by less than 10x, because of the power law.
1:16:44Beren Millidge: But we're on the too-much-data side of Chinchilla: we over-train models. As data runs out, we could easily move back to the Chinchilla-optimal point, or even slightly under-train. It depends on the ratio of training to inference compute: bottlenecked on data, go bigger; bottlenecked on compute, go smaller. Synthetic data complicates it further, so it's very hard to predict.
1:17:22John Schulman: Part of why the scaling laws took so long to find is that you only get a clean relationship when you get everything right. The beautiful straight lines on graphs hide a lot of complexity: scaling every hyperparameter correctly, or parameterizing the optimizer so hyperparameters don't change with model size.
1:17:47Charlie O'Neill: Bugs have their own clean scaling laws too, like Kaplan forgetting cosine annealing, or not counting embedding parameters, which skewed the estimates for smaller models where embeddings are a decent share.
10
Mid-training does 80%; RL tunes the policy
Mid-training often gets the model about 80% of the way; RL tunes the policy with one very high-signal bit. What it delivers is generalization over task length, not across domains.
1:18:03 · Why RL worked better than expected
1:18:06Dwarkesh Patel: A year ago many people argued RL wouldn't scale well. John, you wrote that models learn one bit per RL episode: did I get the answer right or wrong. I wrote that it's even worse: when the pass rate is low, the model learns almost nothing from an episode. Yet today's models seem pretty smart, apparently from scaling up RL. Beren, you wrote a post a few weeks ago explaining this. Why has RL worked better than one would naively have thought?
1:19:01Beren Millidge: Several things. First, mid-training is underestimated. Much of what looks like RL's success comes from very good mid-training data: essentially pre-training on synthetic reasoning data and environments that warm-start the model for RL, which often gets it almost 80% of the way to the final RL checkpoint. RL on top is mostly tweaking the policy, so it needs far fewer bits than you'd think; it doesn't learn the behaviors from scratch. Second, those bits are extremely high-signal compared with pre-training, which is why you need RL rather than just SFT on successful traces. It's not only that they're the bits about getting the answer right; an SFT trace contains the answer token too. What matters is that the RL objective ignores all the other bits. In SFT you have to match the exact reasoning tokens of the model you're copying, so you get far too many bits about how it happens to reason. In RL you get the one bit, not drowned in noise. That's a dramatic increase in signal-to-noise, which is why RL is so efficient per step.
1:20:31Charlie O'Neill: There's been so much debate about what RL does compared with mid-training or SFT: pass@1 goes up but pass@256 goes down, as rare correct traces get outweighed by gradients from easier ones. The simple view now is that if you have enough compute to sample a large enough group, so the chance of getting several correct answers isn't insignificant, they get up-weighted. And to Beren's point, pass@1, RL's starting point, scales with the log of pre-training tokens.
1:21:48Dwarkesh Patel: A very basic question. That makes sense, but the models seem qualitatively so much more capable over the last year. How do we square the relatively small effect this implies for RL with the capabilities they seem to be gaining?
1:22:08Beren Millidge: It doesn't imply RL has a small effect. A few bits and small parameter changes can still change the input-to-output mapping the model computes dramatically. Even one bit can rule out half the hypothesis space, which is huge. Small amounts of RL from a really good starting point can change behavior a lot.
1:22:42Charlie O'Neill: Two things. Everyone hoped RL would generalize reasoning across domains; I don't think we got that horizontal generalization. Training on maths doesn't make you a great coder; you still need RL on code environments. What we did get is horizon generalization: the models learned to use more tokens for longer and keep making progress. Train on ever-longer environments, drop them into a completely new one, and they may lack the right reasoning patterns but can keep going for longer, which correlates with success. A paper called EdgeBench showed that how long models can work is doubling every three months; that's clear evidence of generalization. Finally, pre-training has the idea of quanta: a smooth loss curve that is really the average of vast numbers of discrete phase transitions, no induction heads and then induction heads, tens of thousands, millions, probably hundreds of millions of them. RL is similar. In the slow outer loop Beren mentioned, we RL a model, then dump its synthetic reasoning traces into the next model's mid-training, hitting quanta for many tasks. On one task it looks like a phase transition, a particular finance or Excel task jumping from a 0.5% pass rate to 90%. Average them all, add horizon generalization, and you say: wow, qualitatively better models.
1:24:32Beren Millidge: RL also does generalize a bit, with some transfer between maths and code, or puzzles and maths. And the sheer number of environments people target is vastly greater. Two years ago the labs wouldn't train for a task you do in daily life; now there are lots of environments aimed at each specific thing.
11
Move 37 and a monoculture
RL hasn't killed creativity; models once found several zero-days at a time. But output diversity has dropped, and with everyone distilling Claude, open-weight models increasingly write alike.
1:24:54 · Move 37, and a monoculture
1:25:00Dwarkesh Patel: We talked about RL causing entropy collapse, concentrating probability on solutions the base model already had. But there's another story, from Atari to AlphaGo's move 37: never initialized on human data, it could think in ways humans don't and find extremely creative solutions. Should we expect RL on LLMs to produce move-37-style creativity, beyond human creativity?
1:25:52Beren Millidge: AlphaGo used MCTS, which explores more than ordinary policy gradients. But I don't think RL necessarily reduces creativity. Qualitatively, in the OpenAI–Hugging Face incident, the models came up with multiple zero-days at a time to break out of the sandbox. That's already some level of move-37 creativity, from the LLMs' general generalization. RL is definitely not destroying entropy entirely, especially on long horizons.
1:26:23John Schulman: Some of what people call creativity is solving hard search problems: move 37, or a poem that satisfies many constraints. AI is obviously going to be extremely good at that if trained for it. But there's another sense in which the diversity of outputs drops a lot after RL and models develop tics. They seem good at writing, but distributional analysis shows them reusing certain themes and the same character names all the time. You don't get the diversity of human authors; you get one really good style. RL has cut that diversity a lot. And since so many people distill, mostly from Claude, all the open-weight models write the same way as Claude and share its tics. It's somewhat concerning to me that a monoculture is emerging.
1:27:54Beren Millidge: I don't think that's fundamental to RL as a method, or to distillation, which is just training on data. If the data isn't broad, that's a problem with the data, not with the method. Much of RL's entropy collapse comes from exploiting fairly simple verifiers without a huge diversity of environments. Writing is presumably graded by some judge that has its own tics; the model learns to reward-hack the judge, and that's why it collapses. That's a problem with the judge, not with RL.
12
Timelines: two to ten years
A drop-in remote worker in about 1 to 3 years; 10x for AI researchers in two years per Schulman, 5 to 10 per O'Neill; superintelligence in 3 to 10 years, with wide disagreement.
1:28:32 · Rapid-fire timelines
1:28:48Dwarkesh Patel: Rapid-fire timelines. By when can a user hire a model as a drop-in remote worker for all kinds of white-collar work, not just coding but video editing, law, paralegal work, with full computer use, a month of seamless learning and operation, and complex projects involving other people? Everything a human worker could do over a month.
1:29:22Charlie O'Neill: If it's forced to use a browser, rather than the firm making information programmatically accessible, maybe a couple of years. If not, if it can send Slack messages and so on, I'd still say around a year.
1:29:39Beren Millidge: Maybe three years for full generality. But as Charlie says, many people will make their organizations easier for AIs to use, so you get 80–90% of the way there before that. The rest is a long tail of miscellaneous things some human can do that will take the models a while. It comes down to how quickly we solve this kind of online learning, and whether compaction and writing files to itself get you 80–90% of the way. That's my big uncertainty; I really don't know.
1:30:14Charlie O'Neill: Something it won't be good at: if I have to yell at someone to get something at work, or really push someone to get something done. The model is going to be too nice.
1:30:27John Schulman: Human remote workers vary widely in quality. Hire someone off Upwork for a software project and it's often hard to get them to do a good job or take in your feedback; in some cases the pre-AI version of this was worse than what today's AI gives you. We already have it for some lower-quality work, but not at human level for higher-quality work. I basically agree with Charlie and Beren that we'll have an okay version of this form factor in a year or so, very good at some things, weaker at others, and improving from there.
1:31:33Charlie O'Neill: We keep moving the goalposts based on the very long tail. This year I just told Codex to go get everything I needed for my taxes and send it to the accountant. It had to click through and download a massive list of things. It did it, and it was perfect. A lot of this it can already do.
1:32:21Dwarkesh Patel: Next: a 10x productivity uplift for you as AI researchers. If a breakthrough takes you a year now, you'd make one every month.
1:32:31Charlie O'Neill: Somewhere between 5 and 10 years? I was picturing normal white-collar work over a month for the remote worker; past two months it starts to diverge.
1:33:09Beren Millidge: I can see that. For coding it's already definitely more than 10x, so if it can run even one or two loops of experimental feedback, that's massive already. Then AI progress won't be bottlenecked on researchers running small experiments, but on other things.
1:33:46Dwarkesh Patel: Sure, but those happen 10x faster, which is a huge deal, and it brings the next thing, a 100x speed-up, sooner.
1:33:53Charlie O'Neill: I'm happy to take a bit longer on that one. The crux is my capacity to absorb information and make the Bayesian-optimal decision on the next experiment.
1:34:05Beren Millidge: I'm assuming you can delegate some of that to the AI, which is getting decent at deciding: it runs an experiment, gets a result, runs the next. If it can run two or three in a row without crashing, that's a big uplift.
1:34:24Dwarkesh Patel: Final question: an AI that dominates top human experts across every field of work that can be done on a computer, all cognitive work, including tasks that take three years. Basically ASI.
1:34:59John Schulman: I'd say 3–4 years. AI research is getting the most attention and a lot of energy, and it's not one of the hardest things for AI, since it involves a lot of code and maths, where models are really good. Things involving 3D, spatial and physical work, like mechanical engineering, which isn't getting the most attention now, may take a little longer.
1:35:28Dwarkesh Patel: But it includes fields with relatively little data by nature, where it has to learn on the fly, like becoming superhuman as an engineer at TSMC.
1:35:42John Schulman: Then you'd have to assume you can give the AI the same onboarding material, and something about longer-horizon learning has to be solved.
1:36:00Charlie O'Neill: I'd say 5 to 10. I think automating AI research is basically ASI-complete: there are many things where, even with a memory system outside the model and somewhat longer context, even if you could research the information or write notes yourself, you'd need more than today's million-token context window.
1:36:23Beren Millidge: I kind of agree on the 5-year range, at least for what the labs are focusing on. But there'll be a long tail the AI could learn in theory that no one has bothered with or allocated compute to, so matching literally every human expert may take longer. It doesn't need to learn a new domain as fast as a human, though, because it will have vastly more experience than any human.
1:36:51Dwarkesh Patel: Thanks so much, guys. This was a great format for experts to disagree, debate and work through things together.
Where Indigo landsFurther
Indigo's conclusion
The practitioners' debate version of “capability is a bounded exponential”: RSI will probably come, but objective-setting, continual learning and sample efficiency hold it back; no fast takeoff. The question of speed stays open.
What to remember
The bottleneck moves from compute to setting objectives and continual learning: if 2036 looks normal, the likely reason isn't models that aren't smart enough but not knowing what to do and failing to learn in deployment.
Alignment is the final job: the longest-lasting human role is defining objectives and deciding what we want; technical work can be automated, choosing the right objective won't be handed over soon.
Continual learning breaks in the details, forcing retraining; a cumulative task like RSI is easier than real work whose distribution keeps shifting.
Distillation works against centralization: distillable capability can't be defended, so value moves to unique real deployment data, verticals like Composer and Harvey.
RL works because mid-training does about 80% and RL tunes the policy with one high-signal bit; it extends task length rather than crossing domains.
Claims you can check later
Claim
Who
When we will know
How firm
A general drop-in remote worker: a month of seamless learning, full computer use, working with people
O'Neill, Millidge
About 1 year (not forced through a browser); fully general in about 3 years
First-hand; Millidge says organizations will reshape themselves for AI, getting 80–90% of the way first
AI researchers get a 10x productivity uplift
Schulman (Millidge agrees)
2 years
First-hand; O'Neill says 5 to 10 years
Superintelligence: dominating top human experts in every field doable on a computer, long tasks included
Schulman 3 to 4 years; O'Neill 5 to 10; Millidge about 5
3 to 10 years
First-hand, wide disagreement
Automating AI research is basically as hard as superintelligence
O'Neill (answering the host)
First-hand judgement
How long models can keep working doubles every 3 months (EdgeBench)
O'Neill, citing
Already happening
Second-hand citation; direction checkable
At small scale, data explains about 12.0x of pre-training compute-efficiency gains and architecture about 3.7x
Dwarkesh (with Jerry Han at Princeton)
Done
First-hand experiment at very small scale; whether it holds at scale is unknown
Parameter counts won't jump in the next few years, inference efficiency first; frontier models are far smaller than the rumored 10 trillion parameters
O'Neill (Schulman expects models to keep growing)
Next few years
First-hand judgement
Back on the long-running theses
confirms
AI capability is a bounded exponential: paradigm shifts don't come from hill-climbing The strongest frontline debate support yet: several of its limits confirmed unprompted, and the bottleneck recast from compute to what to do and whether models can keep learning.
confirms
What you own is not the model: value moves up to what cannot be rented Distillation against centralization, deployment data mattering more than environments, goal-setting as the last human job: all three push value toward what is unique and can't be rented.
confirms
Whether verifiable domains generalize The gap between automated research and open-ended science is the training-side version of this boundary; RL extending task length without crossing domains explains its mechanism.
confirms
Nathan Lambert, Lossy self-improvement O'Neill's continual learning breaking in the details and his Moore's-law analogy are dense first-hand engineering evidence for Lambert's “held back by friction”.
conflicts
Jakub Pachocki, An Alien Mind Pachocki strongly expects RSI; this panel talks mostly about friction, with timelines from 3 to 4 years to 5 to 10. The question of speed stays open.
adds to
Furong Huang: self-improving agents “Self-improvement must prove later work is better” is the same crux as whether AI can set its own objectives without going off the rails.
adds to
a16z and Gavin Baker, Demand is outrunning supply Two sides of one week: a16z argues from demand that supply is short; this panel argues from engineering that the bottleneck is objectives and continual learning.
What would change my mind
the next discontinuity adds to the current paradigm instead of tearing it down, and the paradigm connects the dots on its own.
Finished. Indigo's take on this piece is in two places: