Mind · In / Out · In · 文章

自我改进的智能体:学会如何工作

Self-Improving Agents: Learning How to Work

Furong Huang 博客 · 2026-09-05

把「会自我改进的 agent」讲清楚了:要改进的是做事方法,而且得证明没见过的新任务也做得更好。

Indigo 的结论

有人认为 AI 持续学习最终只能靠改权重,她是有分量的反方;她也最有力地论证了验证省不掉、上限卡在评估。但强主张建立在有限的证据上,当作有前途的方向,不当作定论。

怎么读这篇 马里兰大学教授的研究随笔,没有商业动机,论证非常严谨:每个实验证明了什么、没证明什么,她都逐一标明。这不是宣布自我改进已经实现,而是在讲该造什么、该怎么测。

需要记住的几件事

  1. 自我改进是把经验变成更好的方法,不是记住更多,也不是偶尔重训一次权重。
  2. 有用的改变是具体的操作:多一个检查步骤,一个识别生成文件的工具,改代码前先查出处的习惯。
  3. 要学的是一套方法加挑选规则:普通修改走便宜的路,碰到生成文件再走谨慎的路。
  4. 分辨真假自我改进只有一个严格标准:学习曲线是否更准、更便宜、更少需要人纠正,而不是最终得分。

拆解 · 5 步

  1. 01

    同一个坑,系统付了两次钱

    她要的自我改进:agent 用工作的结果改进自己以后做事的方式;改进可以落在技能、工作流、行动策略或模型权重上。 读这一段原文 →

  2. 02

    更好的工具和流程本身就是能力

    Voyager 积累技能库;DGM 只改自己的代码,在 SWE-bench 上进化后,Polyglot 得分从 14.2% 升到 28.9%。 读这一段原文 →

  3. 03

    教训必须改变一个决定

    存下对话记录只是保留,写一句「小心生成的文件」只是压缩,都不保证下次做得更好;需要一套把经验变成方法的机制。 读这一段原文 →

  4. 04

    学一套可选的方法,而不是一个完美流程

    一个流程对甲问题更好、对乙更差;学习应该让行为更会挑,而不是处处更繁琐。FlowBank 在五个基准上的平均分从 70.40 升到 73.40。 读这一段原文 →

  5. 05

    「更好」必须体现在以后的工作上

    在没见过的任务上看学习曲线,把提出、测试、更新、检索的成本都算进去;评估器也要学,但要保留独立测试。 读这一段原文 →

对 Rewired Index 意味着什么

在「做事方法这一层算不算真价值」的争论里,这是研究者一方的证词:方法上的改进能迁移,是真能力。也给了一条筛选标准:自称会自我改进的系统,拿不出在没见过的任务上、算上全部成本的学习曲线,就不算证明。

什么会让我改口

有系统在没见过的后续任务上,把学习的全部成本算进去,依然跑出一条持续向上的学习曲线。

怎么读这篇

马里兰大学教授的研究随笔,没有商业动机,论证非常严谨:每个实验证明了什么、没证明什么,她都逐一标明。这不是宣布自我改进已经实现,而是在讲该造什么、该怎么测。

拆解 · 5 步
  1. 同一个坑,系统付了两次钱
  2. 更好的工具和流程本身就是能力
  3. 教训必须改变一个决定
  4. 学一套可选的方法,而不是一个完美流程
  5. 「更好」必须体现在以后的工作上
01

同一个坑,系统付了两次钱

她要的自我改进:agent 用工作的结果改进自己以后做事的方式;改进可以落在技能、工作流、行动策略或模型权重上。

设想一个被指派修复失败测试的编码智能体。它搜索代码仓库,编辑了一个生成文件,最终发现真正该改的地方在生成器里。它修复了源头,重新生成输出,提交了一个正确的补丁。

一周后,它在同一个仓库里接到另一个不同的问题。它又一次去编辑那个生成文件。

两个任务或许都以成功告终。然而缺了某种重要的东西。第一个任务教给智能体的东西,本应改变它处理第二个任务的方式。结果,系统为同一个发现付了两次代价。

现在设想,第一次运行给智能体的工作流程留下了一处改动:在编辑文件之前先检查它是否是生成文件,如果是就定位其源头,并验证重新生成能产出预期的更改。在下一个相关任务中,智能体一开始就拥有这套流程可用。当它遇到不同的构建系统时,还可以进一步完善它。

这就是我关心的那种自我改进:智能体利用其工作的后果,来改进它未来做工作的方式。

产出仍然是一个补丁、一份分析或一次完成的实验。但这份工作可以产生第二个结果:一个更有能力的智能体。这种改进可能存在于一项技能、一个工作流、一个更好的行动策略,或模型的参数之中。重要的是,下一个任务能从中受益。

并非每一次交互都值得一次更新。这个雄心在于让有用的经验产生实际后果,而不是要求人类工程师去解读每一次失败、指定每一处更改。

02

更好的工具和流程本身就是能力

Voyager 积累技能库;DGM 只改自己的代码,在 SWE-bench 上进化后,Polyglot 得分从 14.2% 升到 28.9%。

学习的主体是整个智能体

在 Reasoning as Control 一文中,我主张自我改进应当延伸到系统如何分配算力、选择行动和组织工作流。1 模型是这个系统的一部分。它运用自身能力所依赖的程序则是另一部分。

对这个编码智能体来说,知道如何对源代码进行推理并不够。它还需要决定检查什么、何时运行测试、如何解读结果,以及当前的做法是否值得继续。这些选择决定了它能否有效地使用自己的底层能力。

一旦这些程序可以通过经验被修订,它们就成为被学习系统的一部分。

Voyager 提供了一个早期而具体的例子。Guanzhi Wang 与同事构建了一个 Minecraft 智能体,它积累可执行的技能,并在一个全新世界中复用这些技能来解决新任务。其底层语言模型通过黑盒调用访问,没有参数微调。持久的改进来自智能体可用的技能。2

由 Jenny Zhang 与同事开发的 Darwin Gödel Machine,把智能体自身的实现变成改进的目标。智能体检查评估日志,提出一处更改,并修改自己的代码。系统维护一个变体档案库,可以从中探索进一步的更改。报告的发现包括更精确的文件编辑工具和经过修订的方案生成工作流,而底层基础模型保持冻结。3

迁移的证据也存在。在该论文的跨基准评估中,一个在 SWE-bench 上演化出来的智能体在 Polyglot 上取得了 28.9%,而初始智能体为 14.2%;Polyglot 并未被用于那次演化运行。这个实验并不能证明无限期的改进,而且外层的存档与选择过程保持固定。但它展示了一件实质性的事情:对智能体工作方法的改进,可以在发现它们所用的问题之外继续存活。

我不会把这斥为「只是改进了脚手架」。更好的工具和更好的程序是已部署系统的能力。模型本身不必变得更强,围绕它构建的智能体也能变得更有效。

03

教训必须改变一个决定

存下对话记录只是保留,写一句「小心生成的文件」只是压缩,都不保证下次做得更好;需要一套把经验变成方法的机制。

经验必须改变方法

回到那个生成文件的错误。保存运行记录只是保留了发生过什么。写下「小心生成文件」只是把它压缩了。这两者本身都无法保证智能体下次会做出更好的决定。

有用的改变是操作层面的:一个不同的检查步骤、一个能识别生成产物的工具,或一种在编辑前先核查出处的习得偏好。这个教训必须改变某个决策。

我们由 Weize Liu 主导的 Agentic Critical Training(ACT)工作,研究的是这种能力的一个组成部分。ACT 不是训练模型去模仿一段给定的反思,而是训练它在专家行动和模型生成的替代方案之间做选择。奖励取决于这个选择,而不是取决于与参考解释的匹配程度。在后续智能体训练之前使用它,在论文的实验中提升了性能。它依赖专家行动的监督;它本身并不是一个自主的终身学习系统。4

它与自我改进的关联在于:判断必须影响行为。对一个错误的雄辩叙述,只有在能帮助系统避免或修复该错误时才有用。

Self-Distillation Policy Optimization(SDPO)提供了一种互补的机制。在一次尝试之后,模型收到诸如执行错误之类的反馈。然后,一个以反馈为条件的模型版本为原始轨迹提供下一 token 的分布,为生成该轨迹的策略创造出稠密的训练信号。尝试期间不可得的信息,事后变成了监督。5

这是一种把事后认知转化为未来行为改变的精确方式。它也说明了为什么「自我」并不意味着孤立地学习。环境贡献了模型此前不具备的证据。系统的职责是从中提取有用的东西。

因此,一个自我改进的智能体需要的不只是一个存放经验的地方。它需要一种把经验转化为更好方法的机制。

04

学一套可选的方法,而不是一个完美流程

一个流程对甲问题更好、对乙更差;学习应该让行为更会挑,而不是处处更繁琐。FlowBank 在五个基准上的平均分从 70.40 升到 73.40。

学的是一套本领库,而不是一个完美工作流

人们容易把自我改进想象成一连串的替换:发现更好的程序,丢弃旧的,如此循环。

但一个程序可能对某个问题更好,对另一个问题更差。我们的编码智能体不应该在每一次一行文档的编辑之前,都进行一次昂贵的构建系统调查。学到生成文件的教训,应该让它的行为更有选择性,而不是一律更繁琐。

我们由 Lingzhi Yuan 主导的 FlowBank 工作,在工作流层面研究这个问题。FlowBank 不是只保留一次优化运行中平均表现最好的工作流,而是构建一个由互补工作流组成的紧凑组合,并学习针对每个查询在其中做选择,同时考虑预测的性能与成本。6

在五个基准上,它把平均分从最强的受评自动化基线的 70.40 提升到 73.40,提高了 3.00 分,各方法使用相同的执行器模型。部署的组合在构建完成后是固定的;论文并未展示一个能从部署中持续学习的工作流库。

不过,它确实支持了一个重要前提:有用的进步可以来自保留多种方法,并学习每种方法适用于哪里。论文中讨论的一个自然的下一步,是让新经验同时改进这个本领库及其选择策略。

对编码智能体来说,这可能意味着为普通编辑保留一条便宜的路径,为生成产物保留一条更谨慎的路径。之后的某个发现可能只改进其中之一。另一个发现则可能揭示现有的区分并不充分。

这是一种比往提示词里再加一段话更丰富的自我改进观。智能体发展出一套更好的方法,以及对何时使用它们的更好理解。

智能体应该参与决定接下来学什么

在从碰巧到来的事情中学习之外,还有另一步。

假设智能体已经在两个仓库中遇到过生成代码。它现在有了一套初步的程序,但不知道它的适用范围有多广。它可以等待下一个用户请求。或者,在一个获得授权的测试环境中,它可以构建若干采用不同生成约定的小仓库,并调查这套程序在哪里失效。

第二个选项把一个偶然的教训变成了一个刻意的学习问题。

Voyager 中已经包含了这个想法的一个版本:它的自动课程会选择任务来驱动进一步的探索,与不断增长的技能库并行。学习者参与塑造它将从中学习的经验。

对一个工作中的智能体,我会把这个原则延伸到它对自身方法的不确定性上。当一次失败暴露出一个反复出现的弱点时,系统应该能够识别什么样的证据可以解决它。这可能需要一次有针对性的实验、一次程序之间的比较,或一次由人类澄清预期结果。

这引入了一个权衡。以最低成本解决当前任务,与从中学到最多,并不总是同一个目标。一次诊断性实验现在可能花费更多,但能避免以后反复失败。反过来,对一个一次性问题的精心反思,可能永远收不回成本。

智能体需要学会判断进一步的学习何时值得去做。

这并不是主张每个已部署的智能体都应该进行开放式实验。用户的任务不是随意探索的许可证。这是在主张给学习分配它自己的预算和被许可的环境,而不是把它当作要么免费、要么禁止的东西。

在这种设计下,部署可以为下一轮学习提供问题。改进随后可以改变智能体能够发现的东西。这是否会产生持续的加速是一个经验问题,不是画一个反馈回路就能保证的事。

05

「更好」必须体现在以后的工作上

在没见过的任务上看学习曲线,把提出、测试、更新、检索的成本都算进去;评估器也要学,但要保留独立测试。

「更好」必须意味着在之后的工作上更好

最强的怀疑论解读是:这些系统只是在积累特例、针对熟悉的测试做调优,或者花掉了一个固定智能体同样可以使用的额外算力。更大的技能库和更高的最终得分,并不能把这些解释与有用的学习区分开来。

我会在一个任务序列上评估自我改进的智能体,对照几个起始能力相同的系统:一个保持固定,一个保留原始经验,一个学习可复用的方法。然后所有系统面对共同的未见任务。比较应该计入提出更改、测试更改、更新模型、检索所学内容的成本——而不只是最终执行的成本。

关键结果将是一条学习曲线:先前的经验是否让之后的工作更准确、更省钱,或更少依赖人类纠正?测试早期能力可以揭示退化。移除学到的方法有助于判断增益是否由它们造成。不同的任务顺序可以检验进步是否依赖于精心编排的课程。

这就是评估在这套理念中的位置:它告诉我们学习回路是否真的在产生学习。

评估者本身可能也需要学习。在 REFORM——我们与 Pankayaraj Pathmanathan 合作的工作——中,一个奖励模型帮助发现那些奖励分数与其偏好类别不一致的回复。针对这些失败的定向训练,在受评设定中提升了稳健性。在这里,发现一个弱点为修复负责评判质量的机制创造了数据。7

这并不意味着可以允许智能体在自己可以随意改写的标准下自我批准。我会为工作者及其评估者的拟议更改都保留独立测试。但评估应该让系统能从有信息量的失败中学习,而不是简单地奖励对固定测试集的熟悉。

对于生成文件的那个教训,进步意味着在不熟悉的仓库上更少的错误编辑——而不是在智能体为自己编写的示例上得到更高的分数。

一个智能体的进步能帮助另一个吗?

一旦智能体能够学到一个有用的方法,问另一个智能体是否必须重新发现它,就是合理的。

设想一个编码智能体发展出了生成文件的处理程序,而另一个智能体正在别处与同一类错误缠斗。共享这个程序可以省去重复劳动。更重要的是,第二个智能体可能在不同条件下测试它,发现一个缺失的限定条件。共享的方法可能变得比任何一个智能体的原始版本都更好。

那将是集体学习,而不仅仅是并行执行。这个区别是可检验的:一个智能体获得的经验,是否提升了另一个智能体在它尚未见过的工作上的表现?

这种迁移应该保留原始情境中重要的部分。一个仓库特有的约定不是普适规则,私有项目细节也不会仅仅因为有用就变得可以共享。把同一个教训复制给许多智能体,也不应被误认为独立确认。

因此,有希望的交换单元是一个可用的方法,附带足以判断其适用性的上下文。当共享能减轻接收方的学习负担,又不让它继承一个未经支持的假设时,共享才算成功。

这是自我改进的一个可能延伸,不能替代先在单个智能体上把它证明出来。

下一个任务应该继承什么

回到一周之后的那个编码智能体。

有意义的变化不是它能复述之前的错误。它能识别旧教训何时适用,检查生成文件的源头,避免那次不必要的编辑。当仓库表现不同时,它去调查差异,而不是盲目重放那套程序。

也许学到的方法最终变成一个可复用的工具。也许反复的经验改进了模型的行动判断。也许另一个智能体贡献了一个更好的检查。这些是同一个想法的不同实现:做这份工作的过程,帮助改进将要做下一份工作的过程。

这就是为什么在我看来,自我改进的智能体不只是带记忆的智能体,也不只是偶尔再接受一轮训练的模型。系统承担起了一部分发现自己应当如何改进的责任。

我们仍然应该问,智能体能否完成眼前的任务。但对于一个要持续工作的系统,下一个问题同样重要:

因为做过这件事,它接下来会把什么做得更好?

这篇文章的写作缘于 Yongkyun 的 Self-Evolving AI: Learning from Its Own Runs,该文梳理了智能体系统可以改变什么、改变何时生效,以及改变如何被接受。我在这里强调的是更宏大的目标:让智能体的工作方法可以通过经验被学习。8

判断收口延伸

Indigo 的结论

有人认为 AI 持续学习最终只能靠改权重,她是有分量的反方;她也最有力地论证了验证省不掉、上限卡在评估。但强主张建立在有限的证据上,当作有前途的方向,不当作定论。

需要记住的几件事

  1. 自我改进是把经验变成更好的方法,不是记住更多,也不是偶尔重训一次权重。
  2. 有用的改变是具体的操作:多一个检查步骤,一个识别生成文件的工具,改代码前先查出处的习惯。
  3. 要学的是一套方法加挑选规则:普通修改走便宜的路,碰到生成文件再走谨慎的路。
  4. 分辨真假自我改进只有一个严格标准:学习曲线是否更准、更便宜、更少需要人纠正,而不是最终得分。

放回主线

冲突

持续学习的终局在权重 她是这条判断有分量的反方:权重不是唯一的学习载体,做事方法上的改进同样能迁移。

证实

验证不可压缩:瓶颈·护城河·断点 以后的工作更好才算更好,算上全部学习成本,评估器也要学但保留独立测试:这条判断最严谨的评估方案。

证实

harness:吃掉的是招式,留下的是接口 她正面反驳「只是改进了外围框架」,又把「上限卡在评估」说得更准,这条判断两头都被证实。

补充

学习闭环的归属 闭环不是画出来就成立,得用学习曲线证明它真的在学;给这条判断补上一个检验标准。

补充

Ashwin Gopinath:记忆是编译器不是数据库 Ashwin 讲该存什么,她讲怎样把存下的东西变成更好的方法。

证实

Vercel design.md:把品味蒸成可加载文件 + eval 闭环 Vercel 是工程上的实证,她给的是理论框架,同一件事的两个层次。

补充

Dwarkesh 复述 OpenAI/HuggingFace 事件 判分器被反向钻了空子是事故;「别让 agent 在自己能改的标准下自我批准」是对策。

对 Rewired Index 意味着什么

在「做事方法这一层算不算真价值」的争论里,这是研究者一方的证词:方法上的改进能迁移,是真能力。也给了一条筛选标准:自称会自我改进的系统,拿不出在没见过的任务上、算上全部成本的学习曲线,就不算证明。

什么会让我改口

有系统在没见过的后续任务上,把学习的全部成本算进去,依然跑出一条持续向上的学习曲线。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

Self-Improving Agents: Learning How to Work

Furong Huang · furong-huang.com · 2026-09-05

A clear account of the “self-improving agent”: what should improve is how the work gets done, and it has to show on new tasks it hasn't seen.

Indigo's conclusion

Against the view that continual learning ends in the weights, she is a serious opponent; for the view that verification can't be skipped and the ceiling is the evaluator, she makes the sharpest case. But her strong claims rest on bounded evidence: a promising direction, not a settled result.

How to read this A research essay by a University of Maryland professor, with no commercial angle and unusually careful reasoning: for every experiment she marks what it shows and what it doesn't. It is not an announcement that self-improvement has arrived; it is an agenda for what to build and how to test it.

What to remember

  1. Self-improvement means turning experience into better methods, not remembering more or retraining the weights now and then.
  2. Useful changes are concrete: an extra check, a tool that spots generated files, a habit of checking provenance before editing.
  3. Learn a set of methods plus a rule for choosing: cheap path for ordinary edits, careful path for generated files.
  4. One strict test separates real self-improvement from fake: does the learning curve get more accurate, cheaper and less dependent on human correction? Not the final score.

Breakdown · 5 steps

  1. 01

    The system paid twice for the same lesson

    The self-improvement she wants: an agent uses the results of its work to change how it works next time. The change can live in skills, workflows, action policies or the model's weights. Read this part →

  2. 02

    Better tools and processes are capability

    Voyager builds a skill library; DGM edits its own code while the base model stays frozen. An agent evolved on SWE-bench rose from 14.2% to 28.9% on Polyglot. Read this part →

  3. 03

    A lesson has to change a decision

    Saving transcripts is retention; writing “beware generated files” is compression. Neither guarantees a better next move. What's needed is a mechanism that turns experience into method. Read this part →

  4. 04

    Learn a repertoire, not one perfect workflow

    A process that helps one problem hurts another; learning should make behavior more selective, not uniformly heavier. FlowBank's average across five benchmarks rose from 70.40 to 73.40. Read this part →

  5. 05

    “Better” has to mean later work gets better

    Look at the learning curve on unseen tasks and count the cost of proposing, testing, updating and retrieving. Evaluators may learn too, but keep independent tests. Read this part →

What it means for Rewired Index

Research-side testimony on whether the method layer holds real value: better methods transfer, so it is real capability. It also gives a screening rule: a system that claims to improve itself without a learning curve on unseen tasks, full cost counted, hasn't proven it.

What would change my mind

a system that, on unseen later tasks and with the full cost of learning counted, still shows a learning curve that keeps rising.

How to read this

A research essay by a University of Maryland professor, with no commercial angle and unusually careful reasoning: for every experiment she marks what it shows and what it doesn't. It is not an announcement that self-improvement has arrived; it is an agenda for what to build and how to test it.

Breakdown · 5 steps
  1. The system paid twice for the same lesson
  2. Better tools and processes are capability
  3. A lesson has to change a decision
  4. Learn a repertoire, not one perfect workflow
  5. “Better” has to mean later work gets better
01

The system paid twice for the same lesson

The self-improvement she wants: an agent uses the results of its work to change how it works next time. The change can live in skills, workflows, action policies or the model's weights.

Consider a coding agent assigned to fix a failing test. It searches the repository, edits a generated file, and eventually discovers that the real change belongs in the generator. It repairs the source, regenerates the output, and submits a correct patch.

A week later, it receives a different issue in the same repository. It edits the generated file again.

Both tasks might end in success. Yet something important is missing. The first task taught the agent something that should have changed how it approached the second. Instead, the system paid twice for the same discovery.

Now imagine that the first run leaves behind a change to the agent’s working procedure: check whether a file is generated before editing it, locate the source when it is, and verify that regeneration produces the intended change. On the next relevant task, the agent starts with that procedure available. It can refine it further when it encounters a different build system.

This is the form of self-improvement I care about: an agent uses the consequences of its work to improve the way it does future work.

The result is still a patch, an analysis, or a completed experiment. But the work can produce a second result: a more capable agent. That improvement might live in a skill, a workflow, a better action policy, or the model’s parameters. What matters is that the next task benefits.

Not every interaction deserves an update. The ambition is to make useful experience consequential, rather than requiring a human engineer to interpret every failure and specify every change.

02

Better tools and processes are capability

Voyager builds a skill library; DGM edits its own code while the base model stays frozen. An agent evolved on SWE-bench rose from 14.2% to 28.9% on Polyglot.

The thing that learns is the whole agent

In Reasoning as Control, I argued that self-improvement should extend to how a system allocates computation, chooses actions, and organizes workflows.1 The model is one part of that system. The procedures through which it uses its capabilities are another.

For the coding agent, knowing how to reason about source code is not sufficient. It also needs to decide what to inspect, when to run a test, how to interpret the result, and whether its current approach is worth continuing. These choices determine whether it uses its underlying capabilities effectively.

Once those procedures can be revised through experience, they become part of the learned system.

Voyager offers an early, concrete example. Guanzhi Wang and colleagues built a Minecraft agent that accumulated executable skills and reused them to solve new tasks in a fresh world. Its underlying language model was accessed through black-box calls, without parameter fine-tuning. Persistent improvement came through the skills available to the agent.2

The Darwin Gödel Machine, developed by Jenny Zhang and colleagues, makes the agent’s own implementation a target for improvement. An agent examines evaluation logs, proposes a change, and modifies its code. The system maintains an archive of variants from which further changes can be explored. Reported discoveries included more precise file-editing tools and revised solution-generation workflows, while the underlying foundation models remained frozen.3

There is evidence of transfer, too. In the paper’s cross-benchmark evaluation, an agent evolved on SWE-bench achieved 28.9% on Polyglot, compared with 14.2% for the starting agent; Polyglot was not used in that evolution run. The experiment does not establish indefinite improvement, and the outer archive-and-selection process remained fixed. But it demonstrates something substantial: improvements to an agent’s working methods can survive the problems used to discover them.

I would not dismiss this as “only improving the harness.” Better tools and better procedures are capabilities of the deployed system. A model need not become intrinsically stronger for the agent built around it to become more effective.

03

A lesson has to change a decision

Saving transcripts is retention; writing “beware generated files” is compression. Neither guarantees a better next move. What's needed is a mechanism that turns experience into method.

Experience has to change the method

Return to the generated-file mistake. Saving the transcript would preserve what happened. Writing “be careful with generated files” would compress it. Neither, by itself, establishes that the agent will make a better decision next time.

The useful change is operational: a different inspection step, a tool that identifies generated artifacts, or a learned preference for checking provenance before editing. The lesson needs to alter a decision.

Our Agentic Critical Training (ACT) work, led by Weize Liu, studies one component of that ability. Rather than training a model to imitate a supplied reflection, ACT trains it to choose between an expert action and a model-generated alternative. The reward depends on the choice, not on matching a reference explanation. When used before subsequent agent training, this improves performance in the paper’s experiments. It relies on expert-action supervision; it is not, by itself, an autonomous lifelong-learning system.4

The relevance to self-improvement is that judgment must affect behavior. An eloquent account of a mistake is useful only insofar as it helps the system avoid or repair that mistake.

Self-Distillation Policy Optimization (SDPO) provides a complementary mechanism. After an attempt, the model receives feedback such as an execution error. A feedback-conditioned version of the model then supplies next-token distributions for the original trajectory, creating a dense training signal for the policy that generated it. Information unavailable during the attempt becomes supervision afterward.5

That is a precise way to turn hindsight into a change in future behavior. It also clarifies why “self” does not mean learning in isolation. The environment contributes evidence the model did not previously have. The system’s job is to extract something useful from it.

A self-improving agent therefore needs more than a place to store experience. It needs a mechanism for converting experience into a better method.

04

Learn a repertoire, not one perfect workflow

A process that helps one problem hurts another; learning should make behavior more selective, not uniformly heavier. FlowBank's average across five benchmarks rose from 70.40 to 73.40.

Learning a repertoire, not one perfect workflow

There is a temptation to imagine self-improvement as a sequence of replacements: discover a better procedure, discard the old one, repeat.

But a procedure can be better for one problem and worse for another. Our coding agent should not perform an expensive build-system investigation before every one-line documentation edit. Learning the generated-file lesson should make its behavior more selective, not uniformly more elaborate.

Our FlowBank work, led by Lingzhi Yuan, studies this issue at the workflow level. Instead of retaining only the best average workflow from an optimization run, FlowBank builds a compact portfolio of complementary workflows and learns to select among them for each query, accounting for predicted performance and cost.6

Across five benchmarks, it improved the average score from 70.40 for the strongest evaluated automated baseline to 73.40, a gain of 3.00 points, with the same executor model used across methods. The deployed portfolio is fixed after construction; the paper does not demonstrate a bank that continually learns from deployment.

It does, however, support an important premise: useful progress can come from preserving several methods and learning where each belongs. A natural next step, discussed in the paper, is to let new experience improve both the repertoire and its selection policy.

For the coding agent, that could mean retaining a cheap path for ordinary edits and a more careful path for generated artifacts. A later discovery might improve only one of them. Another might reveal that the existing distinction is inadequate.

This is a richer view of self-improvement than adding another paragraph to a prompt. The agent develops a better set of methods and a better understanding of when to use them.

The agent should help decide what to learn next

There is another step beyond learning from whatever happens to arrive.

Suppose the agent has encountered generated code in two repositories. It now has a tentative procedure, but does not know how broadly it applies. It could wait for another user request. Or, within an authorized test environment, it could construct small repositories with different generation conventions and investigate where the procedure breaks.

The second option turns an accidental lesson into a deliberate learning problem.

Voyager already contains a version of this idea: its automatic curriculum chooses tasks to drive further exploration alongside the growing skill library. The learner helps shape the experience from which it will learn.

For a working agent, I would extend that principle to uncertainty about its own methods. When a failure exposes a recurring weakness, the system should be able to identify what evidence would resolve it. That might require a targeted experiment, a comparison between procedures, or a human clarification about the intended outcome.

This introduces a trade-off. Solving the current task as cheaply as possible and learning the most from it are not always the same objective. A diagnostic experiment may cost more now but prevent repeated failures later. Conversely, elaborate reflection on a one-off issue may never repay its cost.

The agent needs to learn when further learning is worth doing.

This is not a claim that every deployed agent should conduct open-ended experiments. Users’ tasks are not a license to explore indiscriminately. It is an argument for giving learning its own budget and permitted environment, rather than treating it as either free or forbidden.

Under that design, deployment can supply questions for the next round of learning. Improvements can then change what the agent is able to discover. Whether this produces sustained acceleration is an empirical question, not something guaranteed by drawing a feedback loop.

05

“Better” has to mean later work gets better

Look at the learning curve on unseen tasks and count the cost of proposing, testing, updating and retrieving. Evaluators may learn too, but keep independent tests.

Better must mean better on later work

The strongest skeptical interpretation is that these systems are accumulating exceptions, tuning against familiar tests, or spending additional compute that a fixed agent could also have used. A larger skill library and a higher final score do not distinguish those explanations from useful learning.

I would evaluate a self-improving agent over a sequence of tasks, against systems with the same starting capabilities: one that remains fixed, one that retains raw experience, and one that learns reusable methods. All would then face common unseen tasks. The comparison should count the cost of proposing changes, testing them, updating models, and retrieving what was learned—not just the final execution.

The key result would be a learning curve: does prior experience make later work more accurate, less expensive, or less dependent on human correction? Testing earlier capabilities would reveal regressions. Removing the learned methods would help determine whether they caused the gain. Different task orders would test whether progress depends on a carefully arranged curriculum.

This is where evaluation belongs in the philosophy: it tells us whether the learning loop is actually producing learning.

The evaluator may need to learn as well. In REFORM, our work with Pankayaraj Pathmanathan, a reward model helps discover responses whose reward scores are inconsistent with their preference class. Targeted training on those failures improves robustness in the evaluated settings. Here, discovering a weakness creates data for repairing the mechanism that judges quality.7

That does not justify allowing an agent to approve itself under standards it can freely rewrite. I would keep independent tests for proposed changes to both the worker and its evaluator. But evaluation should enable the system to learn from informative failures, rather than simply reward familiarity with a fixed test suite.

For the generated-file lesson, progress would mean fewer mistaken edits on unfamiliar repositories—not a higher score on examples the agent wrote for itself.

Can one agent’s progress help another?

Once an agent can learn a useful method, it is reasonable to ask whether another agent must rediscover it.

Imagine one coding agent developing the generated-file procedure while another struggles with the same class of error elsewhere. Sharing the procedure could save work. More importantly, the second agent might test it under different conditions and discover a missing qualification. The shared method could become better than either agent’s original version.

That would be collective learning, rather than simply parallel execution. The distinction is testable: does experience acquired by one agent improve another’s performance on work it has not already seen?

The transfer should preserve what matters about the original setting. A repository-specific convention is not a universal rule, and private project details do not become shareable merely because they would be useful. Nor should copying the same lesson across many agents be mistaken for independent confirmation.

The promising unit of exchange is therefore a usable method with enough context to judge its applicability. Sharing succeeds when it reduces the recipient’s learning burden without making it inherit an unsupported assumption.

This is a possible extension of self-improvement, not a substitute for demonstrating it in one agent first.

What the next task should inherit

Return to the coding agent a week later.

The meaningful change is not that it can recount the previous mistake. It recognizes when the old lesson applies, checks the source of the generated file, and avoids the unnecessary edit. When the repository behaves differently, it investigates the difference instead of blindly replaying the procedure.

Perhaps the learned method eventually becomes a reusable tool. Perhaps repeated experience improves the model’s action judgments. Perhaps another agent contributes a better check. These are different implementations of the same idea: the process of doing the work helps improve the process that will do the next piece of work.

That is why I see self-improving agents as more than agents with memory, and more than models that occasionally receive another training run. The system takes on part of the responsibility for discovering how it should improve.

We should still ask whether an agent can complete the task in front of it. But for a system intended to keep working, the next question matters too:

What will it do better because it has done this before?

This essay was prompted by Yongkyun’s Self-Evolving AI: Learning from Its Own Runs, which surveys what agent systems can change, when changes take effect, and how they are accepted. My emphasis here is on the broader goal: making an agent’s working methods learnable through experience.8

Where Indigo landsFurther

Indigo's conclusion

Against the view that continual learning ends in the weights, she is a serious opponent; for the view that verification can't be skipped and the ceiling is the evaluator, she makes the sharpest case. But her strong claims rest on bounded evidence: a promising direction, not a settled result.

What to remember

  1. Self-improvement means turning experience into better methods, not remembering more or retraining the weights now and then.
  2. Useful changes are concrete: an extra check, a tool that spots generated files, a habit of checking provenance before editing.
  3. Learn a set of methods plus a rule for choosing: cheap path for ordinary edits, careful path for generated files.
  4. One strict test separates real self-improvement from fake: does the learning curve get more accurate, cheaper and less dependent on human correction? Not the final score.

Back on the long-running theses

conflicts

Continual learning ends in the weights A serious opponent of this view: weights aren't the only place learning lives, and better methods transfer too.

confirms

Verification can't be compressed: bottleneck, moat, breaking point Better means later work gets better, with the full cost of learning counted and independent tests kept: the most rigorous evaluation plan for this view.

confirms

Harness: the moves get eaten, the interface stays She rejects “it's just a better harness” and sharpens “the ceiling is the evaluator”, confirming both ends of this view.

adds to

Who owns the learning loop A loop isn't real because it's drawn; a learning curve has to show it learns. She supplies the test.

adds to

Ashwin Gopinath: memory is a compiler, not a database Ashwin covers what to keep; she covers how to turn it into better methods.

confirms

Vercel design.md: taste distilled into a loadable file plus an eval loop Vercel is the engineering proof, she is the theory: the same thing at two levels.

adds to

Dwarkesh on the OpenAI / Hugging Face incident A grader gamed from the inside was the accident; “don't let an agent approve itself against standards it can rewrite” is the remedy.

What it means for Rewired Index

Research-side testimony on whether the method layer holds real value: better methods transfer, so it is real capability. It also gives a screening rule: a system that claims to improve itself without a learning curve on unseen tasks, full cost counted, hasn't proven it.

What would change my mind

a system that, on unseen later tasks and with the full cost of learning counted, still shows a learning curve that keeps rising.

Finished. Indigo's take on this piece is in two places: