Mind · In / Out · In · 报告

递归自我改进(RSI):始于 1987

Recursive Self-Improvement (RSI) Since 1987

Jürgen Schmidhuber · people.idsia.ch/~juergen(Technical Note IDSIA-9-26) · 2026-09-17

RSI 的鼻祖趁热认祖归宗:谱系站得住,真正有用的是哥德尔机这个理论天花板,和「终局在硬件」。

Indigo 的结论

史料值钱,口径打折。核心谱系站得住,是对当下 RSI 叙事的有效降噪;横扫式认领是把学术谱系拉成个人所有权。留下三样:40 年的时间轴、哥德尔机这个理论天花板、终局在硬件。

怎么读这篇 技术备注的格式,优先权檄文的实质,9 月 17 日发布,正踩在 RSI 成为全行业头号话题的时候。没有实证数据可核,要审的是两件事:优先权主张站不站得住,以及为什么现在发。核心谱系是真的;「ChatGPT 的 T 是我的」这类横扫式认领是他多年的固定动作,要打折。

需要记住的几件事

  1. RSI 有 40 年的形式化谱系,不是 2026 年的新发明:1987 年的元进化是第一批具体的 RSI 算法。
  2. 哥德尔机是 RSI 的理论天花板:证明了改写有用才改写,全局最优,没有局部极大。
  3. 完整的 RSI 需要自我改进的硬件:软件 RSI 已经实用,前沿是能自我复制的机器文明。
  4. 优先权口径分三层拆:核心谱系是真的,边缘认领方向对、口径拉伸,零安全焦虑是他一贯的立场。
  5. 当下 LLM 版的 RSI 在直接续他的理论线:达尔文-哥德尔机等都受哥德尔机启发。

拆解 · 7 步

  1. 01

    RSI 比这波热潮老得多

    2026 年人人谈 RSI,但他 1987 年的毕业论文就给出了首批具体算法;真正的 RSI,是能以任意可计算方式改写自身代码、只留有用改写的系统。 读这一段原文 →

  2. 02

    1987 元进化,1994 自修改策略

    把遗传编程用到它自己身上,递归进化出更好的方法,还有元元层;1994 年的自修改策略能修改自己修改自己的方式,智能体从不重置。 读这一段原文 →

  3. 03

    神经网络给自己编程

    1991 年的快速权重编程器让一个网络给另一个网络写权重,1992 年起网络能改写自己的权重;他称线性化自注意力的 Transformer 与之形式等价。 读这一段原文 →

  4. 04

    OOPS 与哥德尔机:可证明最优

    2002 年的 OOPS 复用旧解来加速新问题;2003 年的哥德尔机只在证明改写有用时才改写自己,因而全局最优,没有局部极大。 读这一段原文 →

  5. 05

    好奇心驱动的自我改进

    1990 年的对抗式好奇心:一个网络专门制造让世界模型出错的数据;再与自修改策略结合,系统自己决定何时学、学什么,持续发明新任务。 读这一段原文 →

  6. 06

    近期工作,和 LLM 的上下文学习

    2020 年以来的元学习新作,以及受哥德尔机启发的 LLM 版方法;他认为 LLM 的上下文学习只是元学习的一个特例。 读这一段原文 →

  7. 07

    完整的 RSI 需要自我改进的硬件

    软件 RSI 已经实用,但不掌握现实世界就没有超级智能;终局是能自我复制、自我改进的机器文明,他称之为缩放的终极形态。 读这一段原文 →

对 Rewired Index 意味着什么

立场文,不是竞争格局信号:它改变的是怎么给 RSI 这个概念定位和断代,不改变任何标的判断。

什么会让我改口

纯软件的 RSI 不靠自复制硬件就走到终局,或者经验性的算法越过了哥德尔机式的可证明最优。

怎么读这篇

技术备注的格式,优先权檄文的实质,9 月 17 日发布,正踩在 RSI 成为全行业头号话题的时候。没有实证数据可核,要审的是两件事:优先权主张站不站得住,以及为什么现在发。核心谱系是真的;「ChatGPT 的 T 是我的」这类横扫式认领是他多年的固定动作,要打折。

拆解 · 7 步
  1. RSI 比这波热潮老得多
  2. 1987 元进化,1994 自修改策略
  3. 神经网络给自己编程
  4. OOPS 与哥德尔机:可证明最优
  5. 好奇心驱动的自我改进
  6. 近期工作,和 LLM 的上下文学习
  7. 完整的 RSI 需要自我改进的硬件
01

RSI 比这波热潮老得多

2026 年人人谈 RSI,但他 1987 年的毕业论文就给出了首批具体算法;真正的 RSI,是能以任意可计算方式改写自身代码、只留有用改写的系统。

摘要。到 2026 年,所有人,包括 Anthropic、OpenAI、Sakana AI、SpaceX,都在谈递归自我改进(RSI)或元学习(学会学习),创业公司更是直接把自己标榜为 RSI 公司。可 RSI 比这古老得多。1987 年,算力大约比今天贵 10⁸ 倍的时候,我在自己的毕业论文 [META1] 里发表了第一批具体的 RSI 算法(第 1 节)。论文封面是我画的一个自己把自己拉起来的机器人。[META1] 是一长串 RSI 论文里的第一篇,这个方向在 2010 年代 [DEC]、尤其是 2020 年代变得炙手可热。

在这里,我概述我们在以下方面的工作:1994 年以来基于自修改策略的 RSI [METARL2-9](第 2 节),1992 年以来人工神经网络中基于梯度下降的 RSI [FWPMETA1-10](第 3 节),2002 年以来用于课程学习的渐近最优 RSI [OOPS1-3](第 4 节),2003 年以来通过自指的哥德尔机实现的数学最优 RSI [GM3-9](第 5 节),1990/1997 年以来与人工好奇心和内在动机结合的 RSI [AC](第 6 节),以及 2020 年以来的近期 RSI 工作(第 7 节)。计算已经便宜得多,基于软件的 RSI 已经变得实用。但完整的 RSI 不只需要自我改进的软件,还需要物理世界中自我改进的硬件 [DLH](第 9 节)。

最广泛使用的机器学习算法,都是由人类发明并写死的。我们能不能构造出能学会更好学习算法的元学习算法,从而造出真正自我改进、除了可计算性和物理定律之外再无其他限制的 AI?自 1987 年我关于这个主题的毕业论文 [META1][AMA] 以来,这个问题一直是我研究的主要驱动力。

先要说明,元学习有时会被混同于简单的迁移学习,也就是从一个训练集迁移到另一个训练集。然而,即便是一个标准的深度前馈神经网络(NN)[DLH][WHO4-11],也能通过在其他图像集上预训练,迁移学习到更快地学习新图像,例如 [TRA12]。真正的元学习和 RSI 远不止于此,也远不只是学会调整超参数,比如进化策略中的变异率。

真正的 RSI,是把初始的学习算法编码在一种通用编程语言里(例如在循环神经网络即 RNN 上),并提供原语指令,允许以任意可计算的方式修改代码本身。我们再用一个递归框架包住这段自指、自修改的代码,确保只有「有用的」自我修改能留存下来,例如第 2 节、第 5 节。

元学习也许是机器学习最雄心勃勃、也最有回报的目标。一个好的元学习者能学到的东西几乎没有限制。在合适的时候,它会学会通过类比、分块、规划、生成子目标以及它们的各种组合来学习,应有尽有。

02

1987 元进化,1994 自修改策略

把遗传编程用到它自己身上,递归进化出更好的方法,还有元元层;1994 年的自修改策略能修改自己修改自己的方式,智能体从不重置。

1. 元进化与 PSALM(1987)

1987 年,我们发表了 [GP87] [GP],我认为这是第一篇关于遗传编程(GP)的论文:进化用通用编程语言写成、大小不受限制的程序 [GOD][GOD34][CHU][TUR][POS]。

同一年,我的毕业论文 [META1] 第 2 节把这样的 GP 用到它自己身上,递归地进化出更好的 GP 方法。它不只有元层,还有元元层、元元元层,依此类推。我把这种 RSI 方法叫做元进化(Meta Evolution)。

[META1] 第 4 节还提出了用于收益最大化或强化学习(RL)的元学习「原型自指关联学习机制」(PSALM)。这是第一种元元强化学习,也就是基于强化学习的 RSI。

这项工作把 I. J. Good 1966 年关于通过自我改进的「超级智能」实现「智能爆炸」的非正式思辨 [GOOD](Good 没有任何具体的 RSI 算法),以及 Bellman 1967 年关于「元策略」的想法 [BE67] 的某些方面具体化了。

2. 基于自修改策略的强化学习 RSI(1994 年起)

1994 年,我提出了另一种元强化学习或 RSI,叫做增量式自我改进 [METARL2],面向只有一次生命、这次生命就是一次终身试验的通用强化学习机器。也就是说,与传统强化学习不同,它不假设可以重复的独立试验,强化学习智能体从不重置。驱动它的是自修改策略(SMP),这是一个可修改的概率分布,分布在用通用编程语言写成的程序之上 [GOD][GOD34][CHU][TUR][POS],允许任意计算。SMP 的学习算法是 SMP 自身的一部分:SMP 能修改自己修改自己的方式。功劳分配的过程必须考虑到,早先的自我修改在为后来的自我修改铺路。

一种叫做「与环境无关的强化加速」(EIRA)[METARL4] 或「成功故事算法」[METARL7-9] 的方法,迫使 SMP 想出越来越好的自我修改算法,持续提高单位时间获得的奖励 [METARL2-9]。这在有挑战性的实验中效果很好,尽管当时的算力比今天贵 10 万倍。

03

神经网络给自己编程

1991 年的快速权重编程器让一个网络给另一个网络写权重,1992 年起网络能改写自己的权重;他称线性化自注意力的 Transformer 与之形式等价。

3. 神经网络中基于梯度的 RSI:学会给其他网络编程(1991)和给自己编程(1992)

正如我自 1990 年 [AC90] 以来反复指出的,人工神经网络(NN)的连接强度也就是权重,应该被看作它的程序。受哥德尔通用自指形式系统 [GOD][GOD34] 的启发,我构建了输出是其他神经网络的程序或权重矩阵的神经网络,也就是所谓的快速权重编程器(Fast Weight Programmers)[FWP0-2][FWP]。我甚至构建了自指的循环神经网络(RNN),它们能运行并检查自己的权重修改算法,也就是学习算法 [FWPMETA1-10]。与哥德尔的工作不同的是,我的通用编程语言不是基于整数,而是基于实数值的权重,这样每个神经网络的输出对它的程序都是可微的。也就是说,一个简单的程序生成器(高效的梯度下降过程 [BP1],参见 [BP2] [BPA] [BP4] [R7])能在程序空间里算出一个方向,沿着它也许能找到更好的程序 [AC90],特别是更好的「生成程序的程序」[FWP0-2]。我 1989 年以来的很多工作都利用了这一点。

深度结构中的成功学习始于 1965 年,那年 Ivakhnenko 和 Lapa 发表了第一批通用、可行的学习算法,用于有任意多个隐藏层的深度多层感知机。他们的网络已经包含了如今流行的乘法门 [DEEP1-2] [DL1][DL2][DLH],这是后来所谓「动态连接」或「快速权重」神经网络的一个关键要素。1981 年,v. d. Malsburg 第一个明确强调了这种连接快速变化的神经网络的重要性 [FAST];其他人随后跟进 [DLP]。

然而,这些作者还没有一个端到端可微、通过梯度下降学会快速操纵快速权重存储的系统。这样的系统是我在 1991 年发表的 [FWP0][FWP1][ULTRA]。在那里,一个慢网络学会控制另一个独立的快网络的权重变化。也就是说,我像传统计算机那样把存储和控制分开,但用的是完全神经网络的方式(而不是混合方式 [PDA1] [PDA2] [DNC])。(参见我那些如今有时被称为「合成梯度」的相关工作 [NAN1-5]。)

接着,我展示了快速权重如何能用于 RSI,也就是「学会学习」。在 1992 年以来的 [FWPMETA1-5] 中,慢 RNN 和快 RNN 是同一个。RNN 能看到自己的误差或奖励信号,在图中(出自 [FWPMETA5])记为 eval(t+1)。每个连接的初始权重由梯度下降训练,但在一次训练过程中,RNN 自己能通过 O(log n) 个特殊输出单元寻址、读取并修改每个连接,其中 n 是连接的数量,见图中随时间变化的向量 mod(t)、anal(t)、Δ(t)、val(t+1)。也就是说,每个连接的权重都可以快速变化,网络变成自指的:原则上,它能在自己身上(对它的全部权重)运行任意可计算的权重修改算法,也就是学习算法。这就是神经网络的递归自我改进!

1991 到 1993 年,我通过基于梯度下降、借助二维张量或外积更新对快速权重的主动控制,简化了这一点 [FWP2](参见我们近来在这方面的工作 [FWP3] [FWP3a])。动机之一,是让远比同等规模的标准 RNN 更多的时间变量,处在大规模并行、端到端可微的控制之下:O(H²) 而不是 O(H),H 是隐藏单元的数量(参见 [MIR] 第 8 节和 [DLP] 第 H4 节)。1993 年的论文 [FWP2] 还明确讨论了在端到端可微网络中学习内部的「注意力聚光灯」[FWP2] [ATT]。

带线性化自注意力的非归一化 Transformer [TR5-6],在形式上等价于我 1991 年基于外积的快速权重编程器,后者如今被称为非归一化线性 Transformer [ULTRA][MOST]。看看 ChatGPT 里的那个 T。

2001 年,我以前的学生 Sepp Hochreiter 在 LSTM 网络 [LSTM1] 而不是传统 RNN 中使用梯度下降,为不平凡的函数类元学习出快速的在线学习算法,比如所有二元二次函数 [HO1]。

04

OOPS 与哥德尔机:可证明最优

2002 年的 OOPS 复用旧解来加速新问题;2003 年的哥德尔机只在证明改写有用时才改写自己,因而全局最优,没有局部极大。

4. 用于课程学习的渐近最优 RSI(2002 年起)

2002 年,我提出了一种通用的、渐近时间最优的课程学习:一个接一个地解决问题,高效地搜索计算候选解的程序空间,包括那些组织、管理、改编和复用先前所获知识的程序 [OOPS1-3]。「最优有序问题求解器」(OOPS)的灵感来自为单个问题设计的 Levin 通用搜索 [OPT]。面对一个新问题,它会把总搜索时间的一部分,用于测试以可计算的方式利用先前求解程序的程序。如果通过复制编辑或调用先前的代码能比从零开始更快地解决新问题,OOPS 会发现这一点;如果不能,至少先前的解也不会造成太大损害。我提出了一种高效、递归、基于回溯的方法,在存储有限的现实计算机上实现 OOPS。实验展示了 OOPS 如何能从元学习或元搜索中大大获益,也就是以 RSI 的方式搜索更快的搜索过程 [OOPS1-2]。

5. 最优 RSI:自我改进的哥德尔机(2003 年起)

上面第 2 节(1994 年起)那个自指的 RSI 系统,是通过「随后奖励加速」这一不断累积的统计证据,来为自己的自我修改提供依据的。但它不能保证执行理论上最优的自我改进。这促使我提出了哥德尔机(Gödel Machine)[GM3-9],这是第一台完全自指的通用 [UNI] RSI 机器,而且确实在某种数学意义上是最优的。它通常用稍欠通用的「最优有序问题求解器」[OOPS1-2](第 4 节)来寻找可证明最优的自我改进。

RSI 哥德尔机的灵感来自库尔特·哥德尔,他在 1930 年代初创立了理论计算机科学 [GOD][GOD34][GOD21,a,b]。他引入了一种基于整数的通用编码语言,能以公理的形式把任何数字计算机的操作形式化。哥德尔用它既表示数据(比如公理和定理),也表示程序(比如对数据进行操作、生成证明的操作序列)。他著名地构造了一些谈论其他形式语句之计算的形式语句,尤其是自指语句,这些语句意味着其真假无法被任何计算性的定理证明器判定。由此,他确定了数学、定理证明、计算和人工智能(AI)的根本局限 [GOD][GOD21,a,b]。这对 20 世纪的科学和哲学产生了巨大影响。此外,1940 到 70 年代的早期 AI,很大一部分其实就是通过专家系统和逻辑编程,以哥德尔的方式做定理证明和演绎。参见 [MIR] 第 18 节。

哥德尔机 [GM6] 是一台通用的强化学习机器:一旦找到了某次改写有用的证明,它就会改写自己代码的任何部分。这里,与问题相关的效用函数、硬件以及全部初始代码,都由编码在一个初始证明搜索器中的公理来描述,而这个证明搜索器本身也是初始代码的一部分。在机器与环境交互的同时(起初以次优的方式),搜索器会系统而高效地测试可计算的证明技术(输出为证明的程序),直到找到一次可证明有用、可计算的自我改写。我证明了,这样的自我改写是全局最优的,没有局部极大值!因为代码必须先证明,继续为其他自我改写搜索证明是没有用的。与以往基于写死的证明搜索器、不自指的方法不同,哥德尔机不仅复杂度的阶是最优的,还能最优地消除任何隐藏在 O() 记号里的减速,前提是这类加速的效用终究是可以证明的 [GM3-9]。

05

好奇心驱动的自我改进

1990 年的对抗式好奇心:一个网络专门制造让世界模型出错的数据;再与自修改策略结合,系统自己决定何时学、学什么,持续发明新任务。

6. RSI 加上人工好奇心与内在动机(1990 年,1997 年起)

在继续讨论元学习和 RSI 之前,我先解释一下带内在动机的强化学习。我 1990 年提出的对抗式人工好奇心原理 [AC90, AC90b] [AC20](另见综述 [AC09] [AC10]),如今不仅广泛用于强化学习中的探索,也用于图像合成 [AC20][DLP]。它的原理如下。一个神经网络(控制器)以概率方式生成输出,另一个神经网络(世界模型)看到这些输出,并预测环境对它们的反应。通过梯度下降,世界模型网络把自己的误差最小化,而生成器网络则设法产生使这个误差最大化的输出。一个网络的损失就是另一个网络的收益。于是控制器有了内在动机,去生成那些能产生世界模型还能从中学到东西的数据的输出动作或实验。(生成对抗网络 GAN 是它的一个特例:环境只根据生成器的输出是否属于某个给定集合,返回 1 或 0 [AC20];参见 [R2][LEC]、[MIR] 第 5 节和 [WHO8]。)

[AC90](1990)中「与元学习的一个联系」一节已经指出:「模型网络不仅可以用来预测控制器的输入,还可以用来预测它未来的输出。一个这种类型的完美模型,会对控制网络的内部变化建模。它会预测控制器的演化,从而预测梯度下降过程本身的效果。在这种情况下,模型网络中的激活流,会对控制网络的权重变化建模。这又接近于『学会如何学习』的概念。」[AC90] 这篇论文还提出了用循环神经网络(RNN)作为世界模型来做规划 [PLAN,PLAN2-5],以及高维奖励信号。与传统强化学习不同,这些奖励信号也被用作控制器网络的信息性输入,控制器网络学习执行使累积奖励最大化的动作(另见 [MIR] 第 13 节和 [DEC] 第 5 节)。

这对元学习很重要:一个看不到自己误差或奖励的神经网络,无法学会一种更好的方式,把这些信号用作自创学习算法的输入。

几年后,我把第 2 节的 RSI 强化学习系统和对抗式人工好奇心结合进同一个系统 [AC97, AC99, AC02]。它以程序的形式生成计算实验,这些程序的执行既可能改变外部环境,也可能改变强化学习智能体的内部状态。一个实验只有两种结果:某个效果要么发生,要么不发生。实验由两个追求奖励最大化、相互对抗的策略共同提出。两者都能在实验结果出现之前预测并押注。一旦结果真的被观察到,赢家会得到与押注成比例的正奖励,输家得到同等大小的负奖励。于是每个策略都有动机去设计结果(是或否)会让另一个策略吃惊的实验。后者则有动机去学会一些它还不知道的关于世界的东西,免得再被对方智胜。

借助基于自修改策略的 RSI [METARL2-9](第 2 节),这个系统学会了何时学习、学什么 [AC97, AC99, AC02]。只要两个「大脑」每走一步计算都得到一个小小的负奖励,它还会把学习新技能的计算成本降到最低,这会让它偏向简单却仍令人惊讶的实验(对应简单却仍未解决的问题)。这可能有助于层层构建越来越复杂的实验,包括那些能带来外部奖励(如果有的话)的实验。事实上,这种人工创造力不仅可能驱动人工科学家和人工艺术家 [AC06-09],还能加快外部奖励的获取 [AC97] [AC02],直观地说,是因为更好地理解世界有助于更快地解决某些问题。

更新近的、带内在动机的 PowerPlay 强化学习系统(2011)[PP] [PP1],能利用元学习的 OOPS [OOPS1-2](第 4 节)持续地自己发明新目标和新任务,以主动、部分无监督或自监督的方式,逐步学会成为越来越通用的问题求解器。2015 年,带高维视频输入和内在动机(就像 PowerPlay 那样)的强化学习机器人学会了探索 [PP2]。

06

近期工作,和 LLM 的上下文学习

2020 年以来的元学习新作,以及受哥德尔机启发的 LLM 版方法;他认为 LLM 的上下文学习只是元学习的一个特例。

7. 近期关于 RSI 和元学习的工作(2020 年起)

我以前的博士生 Imanol Schlag 等人 [FWPMETA7] 给 LSTM 加上了一个联想式的快速权重记忆(FWM)。在给定输入序列的每一步,通过可微运算,LSTM 更新并维护存储在快速变化的 FWM 权重中、由先前观察组成的组合性关联。模型通过梯度下降端到端训练,在组合性语言推理问题、小规模词级语言建模,以及部分可观测环境的元强化学习上都表现出色 [FWPMETA7]。

我们的 MetaGenRL(2020)[METARL10] 元学习出新的强化学习算法,能用于与训练环境差别很大的环境。MetaGenRL 在描述这类学习算法的低复杂度损失函数空间中搜索。参见我以前的博士生 Louis Kirsch 的博客文章。

这种搜索简单学习算法的原则,也适用于快速权重结构。我们近期的「可变共享元学习」(VS-ML)把权重共享和稀疏性融合进 RSI 循环网络 [FWPMETA6]。这让学习算法可以只用少量参数来编码,尽管它有许多随时间变化的变量,参见 [FWP2](第 3 节)。VS-ML 结合了端到端可微的快速权重 [FWP1-3a](第 3 节)和编码在 LSTM 激活中的学习算法 [HO1]。其中一些激活可以被解释为由 LSTM 动力学更新的神经网络权重。权重矩阵中带共享稀疏项的 LSTM,能发现可以泛化到新数据集的学习算法。这些元学习得来的学习算法不需要显式计算梯度。循环网络中的 VS-ML 还能学会纯粹在端到端可微的前向动力学中实现著名的反向传播学习算法 [BP1] [BP2] [BP4] [FWPMETA6]。

2022 年,我们还在 ICML 发表了一个现代的自指权重矩阵(SWRM)[FWPMETA8],它基于 1992 年的 SWRM [FWPMETA1-5](见第 3 节)。原则上,它能元学习如何学习,能元元学习如何元学习如何学习,依此类推,也就是递归自我改进的意思。我们在有监督的小样本学习任务上,以及在程序化生成游戏环境的多任务强化学习上评估了这个 SRWM。实验表明,SRWM 既实用,性能也有竞争力。

近期关于元学习和 RSI 的工作,集中在受哥德尔机 [GM3-9](第 5 节)启发、基于 LLM 的方法上,例如 Sakana AI 开发的达尔文-哥德尔机 [GMD25]、赫胥黎-哥德尔机 [GMH26] 和红皇后-哥德尔机 [GMR26]。参见我们 2026 年的综述 [RSI26]。

今天的计算比上个千年便宜得多,终于有一批公司开始专注于元学习和 RSI,例如 Sakana AI、Ricursive、Recursive Superintelligence、Anthropic、OpenAI、Inherent 等。

8. LLM 的「上下文学习」是元学习的一个特例

在一定程度上,近来的大语言模型(LLM)能从不断增长的用户交互记录中学习,而无需在测试时通过梯度下降改变底层预训练 Transformer 神经网络的权重。LLM 这种所谓的「上下文学习」和「测试时训练」是元学习的一个特例,类似于 1991 年的非归一化线性 Transformer [ULTRA],以及 2001 年的元学习 LSTM [HO1]:后者通过梯度下降学会了一种针对二次函数、比梯度下降快得多的学习算法(第 3 节),而在测试时不需要额外改变权重 [COCO]!

一般来说,梯度下降可以用来学习一种在神经网络自身上运行的学习算法,这在 1992 年 [FWPMETA1-5](第 3 节)就已展示。许多元学习者实际上是在学习给快速权重编程 [FWP],Transformer 也是,包括 1991 年的非归一化线性 Transformer [ULTRA][FWP0-1,6]。

07

完整的 RSI 需要自我改进的硬件

软件 RSI 已经实用,但不掌握现实世界就没有超级智能;终局是能自我复制、自我改进的机器文明,他称之为缩放的终极形态。

9. 完整的 RSI 需要自我改进的硬件

上面关于 RSI 的讨论集中在自我改进的软件上。然而,正如早先指出的 [DLH],要实现真正的 AI,光做软件研究是不够的,必须把它和由机器与机器人构成的物理世界结合起来。不掌握现实世界,就没有人工超级智能(ASI)!哥德尔机 [GM3-9](第 5 节)考虑到了这一点。我想用 [DLH](基于更早的出版物)中的一段话来结束这份报告:

「几个世纪以来,人们一直在讨论物理上的自复制机器(SRM),包括 1600 年代的笛卡尔、1800 年代的艾略特和巴特勒、1920 年代的恰佩克,以及 1940 年代以来的冯·诺依曼、楚泽、彭罗斯等人 [SRM20]。自复制、自进化的软件几乎微不足道(想想计算机病毒),但没人知道在实践中如何造出通用的物理自复制机器。然而,现在似乎有了一条显而易见的路:由 AI 控制、能学会操作目前由人操作的所有机器和工具的通用机器人 [COG18],也将能够建造、操作和维修制造更多这种机器人所需的机器。这包括从地下开采原材料、精炼、把零件拧在一起、修理坏掉的 3D 打印机、机器人和机器人工厂等等的机器,做目前只有操作机器的人才能做的一切体力工作。归根到底,这是一个不需要操作机器的人就能自我复制的机器文明,然后,当然,它还会改进自己。这是自我改进的硬件,与已经存在的、自我改进的元学习软件相对。[META] 我把这称为缩放的终极形态 [JY24][FA24][95-25]。」

致谢

感谢几位专家审稿人提出的有益意见。科学在于自我纠错,如果你发现任何遗留的错误,请写信到 juergen@idsia.ch 告诉我。本文内容可用于教育和非商业目的,包括维基百科及类似网站的文章。本作品采用知识共享署名-非商业性使用-相同方式共享 4.0 国际许可协议授权。

(此处从略参考文献,见原文。)

判断收口延伸

Indigo 的结论

史料值钱,口径打折。核心谱系站得住,是对当下 RSI 叙事的有效降噪;横扫式认领是把学术谱系拉成个人所有权。留下三样:40 年的时间轴、哥德尔机这个理论天花板、终局在硬件。

需要记住的几件事

  1. RSI 有 40 年的形式化谱系,不是 2026 年的新发明:1987 年的元进化是第一批具体的 RSI 算法。
  2. 哥德尔机是 RSI 的理论天花板:证明了改写有用才改写,全局最优,没有局部极大。
  3. 完整的 RSI 需要自我改进的硬件:软件 RSI 已经实用,前沿是能自我复制的机器文明。
  4. 优先权口径分三层拆:核心谱系是真的,边缘认领方向对、口径拉伸,零安全焦虑是他一贯的立场。
  5. 当下 LLM 版的 RSI 在直接续他的理论线:达尔文-哥德尔机等都受哥德尔机启发。

放回主线

证实

约束正从算法层迁到物理层 「完整的 RSI 需要自我改进的硬件」是这条判断的鼻祖级同调,可以和 Apple 与 NVLink、芯片架构师的物理层证词并列。

冲突

Noam Brown(OpenAI)内部人复盘 OAI-HF:RSI 是压倒性头号优先级 同周 RSI 的两种温度:Noam 讲当下头号优先级、带对齐灾难;Schmidhuber 讲追了 40 年、终局是机器生命扩张的乐观工程。

补充

Dwarkesh RSI 辩论(Schulman + Militch + O'Neal) 面板在经验层讲山脚的阻力,这里在理论层给出山顶的坐标,合起来是「有界指数」的理论和工程双轴。

冲突

Dario Amodei《我们必须为前沿限速》 最锋利的立场对撞:Dario 满纸安全焦虑、要限速;Schmidhuber 零安全焦虑,把 RSI 终局讲成缩放的终极形态。

补充

Haseltine:AI 是第二种生命 第 9 节的自复制硬件,几乎是「第二种生命」的工程决定论版本:一个从生物学功能定义,一个从自复制硬件定义。

对 Rewired Index 意味着什么

立场文,不是竞争格局信号:它改变的是怎么给 RSI 这个概念定位和断代,不改变任何标的判断。

什么会让我改口

纯软件的 RSI 不靠自复制硬件就走到终局,或者经验性的算法越过了哥德尔机式的可证明最优。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Report

Recursive Self-Improvement (RSI) Since 1987

Jürgen Schmidhuber · people.idsia.ch/~juergen(Technical Note IDSIA-9-26) · 2026-09-17

The father of RSI claims his lineage while it's hot: the family tree holds, and what's truly useful is the Gödel Machine as a theoretical ceiling and “the endgame is hardware”.

Indigo's conclusion

Valuable history, discounted claims. The core lineage holds and usefully de-noises today's RSI talk; the sweeping claims stretch a scholarly lineage into personal ownership. Keep three things: the 40-year timeline, the Gödel Machine as theoretical ceiling, and the endgame in hardware.

How to read this Formatted as a technical note, in substance a priority manifesto, published on 17 September, just as RSI became the industry's number-one topic. There's no empirical data to check; what to audit is whether the priority claims hold and why now. The core lineage is real; sweeping claims like “the T in ChatGPT is mine” are a long-standing habit of his, and get a discount.

What to remember

  1. RSI has a 40-year formal lineage and isn't a 2026 invention: 1987's Meta Evolution was the first concrete RSI algorithm.
  2. The Gödel Machine is RSI's theoretical ceiling: rewrite only after proving the rewrite useful, globally optimal with no local maxima.
  3. Full RSI needs self-improving hardware: software RSI is already practical; the frontier is a self-replicating machine civilization.
  4. Split the priority claims three ways: the core lineage is real, the fringe claims point the right way but stretch, and zero safety anxiety is his standing stance.
  5. Today's LLM versions of RSI continue his line directly: the Darwin-Gödel Machine and others are inspired by the Gödel Machine.

Breakdown · 7 steps

  1. 01

    RSI is much older than the hype

    In 2026 everyone talks RSI, but his 1987 diploma thesis already gave the first concrete algorithms. True RSI is a system that can rewrite its own code in any computable way and keeps only the useful rewrites. Read this part →

  2. 02

    1987 Meta Evolution, 1994 self-modifying policies

    Genetic programming applied to itself, recursively evolving better methods, with meta-meta levels; the 1994 self-modifying policies can change the way they change themselves, and the agent is never reset. Read this part →

  3. 03

    Networks that program themselves

    The 1991 Fast Weight Programmers let one network write another's weights, and from 1992 a network rewrites its own; he says Transformers with linearized self-attention are formally equivalent. Read this part →

  4. 04

    OOPS and the Gödel Machine: provably optimal

    The 2002 OOPS reuses old solutions to speed up new problems; the 2003 Gödel Machine rewrites itself only once it has proved the rewrite useful, so it's globally optimal, with no local maxima. Read this part →

  5. 05

    Curiosity-driven self-improvement

    The 1990 adversarial curiosity: one network produces data that makes the world model err; combined with self-modifying policies, the system decides when and what to learn and keeps inventing new tasks. Read this part →

  6. 06

    Recent work, and in-context learning in LLMs

    New meta-learning work since 2020, and LLM-based methods inspired by the Gödel Machine; he treats in-context learning in LLMs as a special case of meta learning. Read this part →

  7. 07

    Full RSI requires self-improving hardware

    Software RSI is already practical, but there's no superintelligence without mastering the real world; the endgame is a self-replicating, self-improving machine civilization, which he calls the ultimate form of scaling. Read this part →

What it means for Rewired Index

A position paper, not a competitive signal: it changes how to place and date the idea of RSI, not any view on a name.

What would change my mind

pure software RSI reaching the endgame without self-replicating hardware, or empirical algorithms overtaking Gödel-Machine-style provable optimality.

How to read this

Formatted as a technical note, in substance a priority manifesto, published on 17 September, just as RSI became the industry's number-one topic. There's no empirical data to check; what to audit is whether the priority claims hold and why now. The core lineage is real; sweeping claims like “the T in ChatGPT is mine” are a long-standing habit of his, and get a discount.

Breakdown · 7 steps
  1. RSI is much older than the hype
  2. 1987 Meta Evolution, 1994 self-modifying policies
  3. Networks that program themselves
  4. OOPS and the Gödel Machine: provably optimal
  5. Curiosity-driven self-improvement
  6. Recent work, and in-context learning in LLMs
  7. Full RSI requires self-improving hardware
01

RSI is much older than the hype

In 2026 everyone talks RSI, but his 1987 diploma thesis already gave the first concrete algorithms. True RSI is a system that can rewrite its own code in any computable way and keeps only the useful rewrites.

Abstract. As of 2026, everyone—including Anthropic/OpenAI/Sakana AI/SpaceX—is talking about recursive self-improvement (RSI) or meta learning (learning to learn), and startups explicitly brand themselves as RSI companies. RSI is much older than that. In 1987, when compute was about 10⁸ times more expensive than today, I published the first concrete RSI algorithms in my diploma thesis [META1] (Sec. 1). For its cover I drew a robot that bootstraps itself. [META1] was the first in a long series of publications on RSI, which became hot in the 2010s [DEC] and especially the 2020s.

Here I summarize our work on RSI with self-modifying policies since 1994 [METARL2-9] (Sec. 2), gradient descent-based RSI in artificial neural networks since 1992 [FWPMETA1-10] (Sec. 3), asymptotically optimal RSI for curriculum learning since 2002 [OOPS1-3] (Sec. 4), mathematically optimal RSI through the self-referential Gödel Machine since 2003 [GM3-9] (Sec. 5), RSI combined with artificial curiosity and intrinsic motivation [AC] since 1990/1997 (Sec. 6), and recent work on RSI since 2020 (Sec. 7). Computing has become much cheaper, and software-based RSI has become practical. Full RSI, however, will require not just self-improving software but self-improving hardware in the physical world [DLH] (Sec. 9).

The most widely used machine learning algorithms were invented and hardwired by humans. Can we also construct meta learning algorithms that can learn better learning algorithms, to build truly self-improving AIs without any limits other than the limits of computability and physics? This question has been a main drive of my research since my 1987 diploma thesis on this topic [META1][AMA].

First note that meta learning is sometimes confused with simple transfer learning from one training set to another. However, even a standard deep feedforward neural network (NN) [DLH][WHO4-11] can transfer-learn to learn new images faster through pre-training on other image sets, e.g., [TRA12]. True meta learning and RSI is much more than that, and also much more than just learning to adjust hyper-parameters such as mutation rates in evolution strategies.

True RSI is about encoding the initial learning algorithm in a universal programming language (e.g., on a recurrent neural network or RNN), with primitive instructions that allow for modifying the code itself in arbitrary computable fashion. We surround this self-referential, self-modifying code by a recursive framework that ensures that only "useful" self-modifications survive, e.g., Sec. 2, Sec. 5.

Meta learning may be the most ambitious but also the most rewarding goal of machine learning. There are few limits to what a good meta learner will learn. Where appropriate, it will learn to learn by analogy, by chunking, by planning, by subgoal generation, by combinations thereof—you name it.

02

1987 Meta Evolution, 1994 self-modifying policies

Genetic programming applied to itself, recursively evolving better methods, with meta-meta levels; the 1994 self-modifying policies can change the way they change themselves, and the agent is never reset.

1. Meta Evolution and PSALMs (1987)

In 1987, we published [GP87] [GP] what I think was the first paper on Genetic Programming or GP for evolving programs of unlimited size written in a universal programming language [GOD][GOD34][CHU][TUR][POS].

In the same year, Sec. 2 of my diploma thesis [META1] applied such GP to itself, to recursively evolve better GP methods. There was not only a meta level but also a meta meta level and a meta meta meta level etc. I called this RSI method Meta Evolution.

Sec. 4 of [META1] also introduced meta learning Prototypical Self-Referential Associating Learning Mechanisms (PSALMs) for payoff maximisation or Reinforcement Learning (RL). This was a first kind of meta meta RL or RL-based RSI.

This work concretizes aspects of I. J. Good's informal and speculative remarks (1966) on an "intelligence explosion" through self-improving "super-intelligences" [GOOD] (Good did not have any concrete RSI algorithms), and Bellman's thoughts on "metapolicies" (1967) [BE67].

2. RSI for Reinforcement Learning with Self-Modifying Policies (1994-)

In 1994, I proposed another type of meta RL or RSI called incremental self-improvement [METARL2] for general purpose RL machines with a single life consisting of a single lifelong trial. That is, unlike in traditional RL, there is no assumption of repeatable independent trials, and the RL agent is never reset. It is driven by a self-modifying policy (SMP) which is a modifiable probability distribution over programs written in a universal programming language [GOD][GOD34][CHU][TUR][POS], to allow for arbitrary computations. The learning algorithm of an SMP is part of the SMP itself—SMPs can modify the way they modify themselves. The credit assignment process has to take into account that early self-modifications are setting the stage for later ones.

A method called Environment-Independent Reinforcement Acceleration (EIRA) [METARL4] or Success-Story Algorithm [METARL7-9] forces SMPs to come up with better and better self-modification algorithms that continually improve reward intake per time [METARL2-9]. This worked well in challenging experiments, although compute back then was 100,000 times more expensive than today.

03

Networks that program themselves

The 1991 Fast Weight Programmers let one network write another's weights, and from 1992 a network rewrites its own; he says Transformers with linearized self-attention are formally equivalent.

3. Gradient-Based RSI in NNs that Learn to Program Other NNs (1991) and Themselves (1992)

As I have frequently pointed out since 1990 [AC90], the connection strengths or weights of an artificial neural network (NN) should be viewed as its program. Inspired by Gödel's universal self-referential formal systems [GOD][GOD34], I built NNs whose outputs are programs or weight matrices of other NNs: the so-called Fast Weight Programmers [FWP0-2][FWP]. I even built self-referential recurrent NNs (RNNs) that can run and inspect their own weight change algorithms or learning algorithms [FWPMETA1-10]. A difference to Gödel's work was that my universal programming language was not based on the integers, but on real-valued weights, such that each NN's output is differentiable with respect to its program. That is, a simple program generator (the efficient gradient descent procedure [BP1]—compare [BP2] [BPA] [BP4] [R7]) can compute a direction in program space where one may find a better program [AC90], in particular, a better program-generating program [FWP0-2]. Much of my work since 1989 has exploited this fact.

Successful learning in deep architectures started in 1965 when Ivakhnenko & Lapa published the first general, working learning algorithms for deep multilayer perceptrons with arbitrarily many hidden layers. Their nets already contained the now popular multiplicative gates [DEEP1-2] [DL1][DL2][DLH], an essential ingredient of what was later called NNs with dynamic links or fast weights. In 1981, v. d. Malsburg was the first to explicitly emphasize the importance of NNs with such rapidly changing connections [FAST]; others followed [DLP].

However, these authors did not yet have an end-to-end differentiable system that learns by gradient descent to quickly manipulate the fast weight storage. Such a system I published in 1991 [FWP0][FWP1][ULTRA]. There a slow NN learns to control the weight changes of a separate fast NN. That is, I separated storage and control like in traditional computers, but in a fully neural way (rather than in a hybrid fashion [PDA1] [PDA2] [DNC]). (Compare my related work on what's now sometimes called Synthetic Gradients [NAN1-5].)

Then I showed how fast weights can be used for RSI or "learning to learn." In references [FWPMETA1-5] since 1992, the slow RNN and the fast RNN are identical. The RNN can see its own errors or reward signals called eval(t+1) in the image (from [FWPMETA5]). The initial weight of each connection is trained by gradient descent, but during a training episode, each connection can be addressed and read and modified by the RNN itself through O(log n) special output units, where n is the number of connections—see time-dependent vectors mod(t), anal(t), Δ(t), val(t+1) in the image. That is, each connection's weight may rapidly change, and the network becomes self-referential in the sense that it can in principle run arbitrary computable weight change algorithms or learning algorithms (for all of its weights) on itself: recursive self-improvement for NNs!

In 1991-93, I simplified this through gradient descent-based, active control of fast weights through 2D tensors or outer product updates [FWP2] (compare our more recent work on this [FWP3] [FWP3a]). One motivation was to get many more temporal variables under massively parallel end-to-end differentiable control than what's possible in standard RNNs of the same size: O(H²) instead of O(H), where H is the number of hidden units (compare Sec. 8 of [MIR] and Sec. H4 of [DLP]). The 1993 paper [FWP2] also explicitly addressed the learning of internal spotlights of attention in end-to-end differentiable networks [FWP2] [ATT].

Unnormalized Transformers with linearized self-attention [TR5-6] are formally equivalent to my 1991 outer product-based Fast Weight Programmers, now called unnormalized linear Transformers [ULTRA][MOST]—see the T in ChatGPT.

In 2001, my former student Sepp Hochreiter used gradient descent in LSTM networks [LSTM1] instead of traditional RNNs to meta learn fast online learning algorithms for nontrivial classes of functions, such as all quadratic functions of two variables [HO1].

04

OOPS and the Gödel Machine: provably optimal

The 2002 OOPS reuses old solutions to speed up new problems; the 2003 Gödel Machine rewrites itself only once it has proved the rewrite useful, so it's globally optimal, with no local maxima.

4. Asymptotically Optimal RSI for Curriculum Learning (2002-)

In 2002, I introduced a general and asymptotically time-optimal type of curriculum learning, that is, solving one problem after another, efficiently searching the space of programs that compute solution candidates, including those programs that organize and manage and adapt and reuse earlier acquired knowledge [OOPS1-3]. The Optimal Ordered Problem Solver (OOPS) draws inspiration from Levin's Universal Search [OPT] designed for single problems. It spends part of the total search time for a new problem on testing programs that exploit previous solution-computing programs in computable ways. If the new problem can be solved faster by copy-editing/invoking previous code than by solving the new problem from scratch, then OOPS will find this out. If not, then at least the previous solutions will not cause much harm. I introduced an efficient, recursive, backtracking-based way of implementing OOPS on realistic computers with limited storage. Experiments illustrated how OOPS can greatly profit from meta learning or meta searching, that is, searching for faster search procedures in RSI style [OOPS1-2].

5. Optimal RSI: Self-Improving Gödel Machine (2003-)

The self-referential RSI system of Sec. 2 above (1994-) justified its self-modifications through growing statistical evidence of subsequent reward accelerations. But it was not guaranteed to execute theoretically optimal self-improvements. This motivated my Gödel Machine [GM3-9], the first fully self-referential universal [UNI] RSI machine that was indeed optimal in a certain mathematical sense. Typically it uses the somewhat less general Optimal Ordered Problem Solver [OOPS1-2] (Sec. 4) for finding provably optimal self-improvements.

The RSI Gödel Machine is inspired by Kurt Gödel, the founder of theoretical computer science in the early 1930s [GOD][GOD34][GOD21,a,b]. He introduced a universal coding language based on the integers which allows for formalizing the operations of any digital computer in axiomatic form. Gödel used it to represent both data (such as axioms and theorems) and programs (such as proof-generating sequences of operations on the data). He famously constructed formal statements that talk about the computation of other formal statements, especially self-referential statements which imply that their truth is not decidable by any computational theorem prover. Thus he identified fundamental limits of mathematics and theorem proving and computing and Artificial Intelligence (AI) [GOD][GOD21,a,b]. This had enormous impact on science and philosophy of the 20th century. Furthermore, much of early AI in the 1940s-70s was actually about theorem proving and deduction in Gödel style through expert systems and logic programming. Compare Sec. 18 of [MIR].

A Gödel Machine [GM6] is a general RL machine that will rewrite any part of its own code as soon as it has found a proof that the rewrite is useful, where the problem-dependent utility function and the hardware and the entire initial code are described by axioms encoded in an initial proof searcher which is also part of the initial code. While the machine is interacting with its environment (initially in a suboptimal way), the searcher systematically and efficiently tests computable proof techniques (programs whose outputs are proofs) until it finds a provably useful, computable self-rewrite. I showed that such a self-rewrite is globally optimal—no local maxima!—since the code first had to prove that it is not useful to continue the proof search for alternative self-rewrites. Unlike previous non-self-referential methods based on hardwired proof searchers, the Gödel Machine not only boasts an optimal order of complexity but can optimally reduce any slowdowns hidden by the O()-notation, provided the utility of such speed-ups is provable at all [GM3-9].

05

Curiosity-driven self-improvement

The 1990 adversarial curiosity: one network produces data that makes the world model err; combined with self-modifying policies, the system decides when and what to learn and keeps inventing new tasks.

6. RSI plus Artificial Curiosity and Intrinsic Motivation (1990, 1997-)

Before I continue the discussion of meta learning and RSI, let me first explain RL with intrinsic motivation. My popular principle of adversarial artificial curiosity from 1990 [AC90, AC90b] [AC20] (see also surveys [AC09] [AC10]) is now widely used not only for exploration in RL but also for image synthesis [AC20][DLP]. It works as follows. One NN (the controller) probabilistically generates outputs, another NN (the world model) sees those outputs and predicts environmental reactions to them. Using gradient descent, the world model NN minimizes its error, while the generator NN tries to make outputs that maximize this error. One net's loss is the other net's gain. So the controller is intrinsically motivated to generate output actions or experiments that yield data from which the world model can still learn something. (GANs are a special case of this where the environment simply returns 1 or 0 depending on whether the generator's output is in a given set [AC20]; compare [R2][LEC] and Sec. 5 of [MIR] and [WHO8].)

The Section "A Connection to Meta learning" in [AC90] (1990) already pointed out: "A model network can be used not only for predicting the controller's inputs but also for predicting its future outputs. A perfect model of this kind would model the internal changes of the control network. It would predict the evolution of the controller, and thereby the effects of the gradient descent procedure itself. In this case, the flow of activation in the model network would model the weight changes of the control network. This in turn comes close to the notion of learning how to learn." The paper [AC90] also introduced planning with recurrent NNs (RNNs) as world models [PLAN,PLAN2-5], and high-dimensional reward signals. Unlike in traditional RL, those reward signals were also used as informative inputs to the controller NN learning to execute actions that maximise cumulative reward (see also Sec. 13 of [MIR] and Sec. 5 of [DEC]).

This is important for meta learning: an NN that cannot see its own errors or rewards cannot learn a better way of using such signals as inputs for self-invented learning algorithms.

A few years later, I combined the RSI RL system of Sec. 2 and Adversarial Artificial Curiosity in a single system [AC97, AC99, AC02]. It generates computational experiments in form of programs whose execution may change both an external environment and the RL agent's internal state. An experiment has a binary outcome: either a particular effect happens, or it doesn't. Experiments are collectively proposed by two reward-maximizing adversarial policies. Both can predict and bet on experimental outcomes before they happen. Once such an outcome is actually observed, the winner will get a positive reward proportional to the bet, and the loser a negative reward of equal magnitude. So each policy is motivated to create experiments whose yes/no outcomes surprise the other policy. The latter in turn is motivated to learn something about the world that it did not yet know, such that it is not outwitted again.

Using RSI with self-modifying policies [METARL2-9] (Sec. 2), the system learns when to learn and what to learn [AC97, AC99, AC02]. It will also minimize the computational cost of learning new skills, provided both brains receive a small negative reward for each computational step, which introduces a bias towards simple still surprising experiments (reflecting simple still unsolved problems). This may facilitate hierarchical construction of more and more complex experiments, including those yielding external reward (if there is any). In fact, this type of artificial creativity may not only drive artificial scientists and artists [AC06-09], but can also accelerate the intake of external reward [AC97] [AC02], intuitively because a better understanding of the world can help to solve certain problems faster.

The more recent, intrinsically motivated PowerPlay RL system (2011) [PP] [PP1] can use the meta learning OOPS [OOPS1-2] (Sec. 4) to continually invent on its own new goals and tasks, incrementally learning to become a more and more general problem solver in an active, partially unsupervised or self-supervised fashion. RL robots with high-dimensional video inputs and intrinsic motivation (like in PowerPlay) learned to explore in 2015 [PP2].

06

Recent work, and in-context learning in LLMs

New meta-learning work since 2020, and LLM-based methods inspired by the Gödel Machine; he treats in-context learning in LLMs as a special case of meta learning.

7. More Recent Work on RSI and Meta Learning (2020-)

My former PhD student Imanol Schlag et al. [FWPMETA7] augmented an LSTM with an associative Fast Weight Memory (FWM). Through differentiable operations at every step of a given input sequence, the LSTM updates and maintains compositional associations of former observations stored in the rapidly changing FWM weights. The model is trained end-to-end by gradient descent and yields excellent performance on compositional language reasoning problems, small-scale word-level language modelling, and meta RL for partially observable environments [FWPMETA7].

Our MetaGenRL (2020) [METARL10] meta learns novel RL algorithms applicable to environments that significantly differ from those used for training. MetaGenRL searches the space of low-complexity loss functions that describe such learning algorithms. See the blog post of my former PhD student Louis Kirsch.

This principle of searching for simple learning algorithms is also applicable to fast weight architectures. Our recent Variable Shared Meta Learning (VS-ML) merges weight sharing and sparsity in RSI RNNs [FWPMETA6]. This allows for encoding the learning algorithm by few parameters although it has many time-varying variables—compare [FWP2] (Sec. 3). VS-ML combines end-to-end differentiable fast weights [FWP1-3a] (Sec. 3) and learning algorithms encoded in the activations of LSTMs [HO1]. Some of these activations can be interpreted as NN weights updated by the LSTM dynamics. LSTMs with shared sparse entries in their weight matrix discover learning algorithms that generalize to new datasets. The meta learned learning algorithms do not require explicit gradient calculation. VS-ML in RNNs can also learn to implement the famous backpropagation learning algorithm [BP1] [BP2] [BP4] purely in the end-to-end differentiable forward dynamics of RNNs [FWPMETA6].

In 2022, we also published at ICML a modern self-referential weight matrix (SWRM) [FWPMETA8] based on the 1992 SWRM [FWPMETA1-5] (see Sec. 3). In principle, it can meta learn to learn, and meta meta learn to meta learn to learn, and so on, in the sense of recursive self-improvement. We evaluated our SRWM on supervised few-shot learning tasks and on multi-task reinforcement learning with procedurally generated game environments. The experiments demonstrated both practical applicability and competitive performance of the SRWM.

Recent work on meta learning and RSI focused on LLM-based approaches inspired by the Gödel Machine [GM3-9] (Sec. 5), e.g., the Darwin-Gödel Machine [GMD25] developed at Sakana AI, the Huxley-Gödel Machine [GMH26], and the Red Queen Gödel Machine [GMR26]. See our 2026 survey [RSI26].

Computing is much cheaper today than it was in the previous millennium, and finally, a number of companies are starting to focus on meta learning and RSI, e.g., Sakana AI, Ricursive, Recursive Superintelligence, Anthropic, OpenAI, Inherent, and others.

8. "In-Context Learning" of LLMs is a Special Case of Meta Learning

To a certain extent, recent Large Language Models (LLMs) can learn from a growing record of user interactions without changing the weights of the underlying pre-trained Transformer NN through gradient descent during test time. This so-called "In-Context Learning" and "Test Time Training" of LLMs is a special case of meta learning, similar to the 1991 unnormalized linear Transformer [ULTRA] and the 2001 meta learning LSTM [HO1] which learned by gradient descent a learning algorithm for quadratic functions that was much faster than gradient descent (Sec. 3), without executing additional weight changes during test time [COCO]!

Generally speaking, gradient descent can be used to learn a learning algorithm running on the neural network itself, as shown in 1992 [FWPMETA1-5] (Sec. 3). Many meta learners actually learn to program fast weights [FWP], and Transformers do so too, including the 1991 unnormalized linear Transformer [ULTRA][FWP0-1,6].

07

Full RSI requires self-improving hardware

Software RSI is already practical, but there's no superintelligence without mastering the real world; the endgame is a self-replicating, self-improving machine civilization, which he calls the ultimate form of scaling.

9. Full RSI Requires Self-Improving Hardware

The RSI discussion above focused on self-improving software. However, as pointed out earlier [DLH], to achieve True AI, software research per se is not enough, it has to be combined with the physical world of machines and robots. No Artificial Super Intelligence (ASI) without mastery of the real world! The Gödel Machine [GM3-9] (Sec. 5) takes this into account, and I'd like to end this report with a quote from [DLH] (based on earlier publications):

"For centuries, humans have discussed physical self-replicating machines (SRMs), including Descartes in the 1600s, Eliot and Butler in the 1800s, Capek in the 1920s, von Neumann & Zuse & Penrose and others since the 1940s [SRM20]. While self-replicating and evolving software is almost trivial (think of computer viruses), nobody knew how to build general purpose physical SRMs in practice. However, now there seems to be an obvious way: AI-controlled general-purpose robots that can learn to operate all the machines and tools currently operated by humans [COG18] will also be able to build/operate/repair the machines required to make more of those robots. This includes machines that mine the raw material from the ground, refine it, screw parts together, repair broken 3D printers and robots and robot factories, and so on, doing all the physical jobs that currently only machine-operating humans can do. Basically, a machine civilisation that can self-replicate without machine-operating humans, and then, of course, improve itself. Self-improving hardware, as opposed to the already existing, self-improving, meta-learning software. [META] I called this the ultimate form of scaling [JY24][FA24][95-25]."

Acknowledgments

Thanks to several expert reviewers for useful comments. Since science is about self-correction, let me know under juergen@idsia.ch if you can spot any remaining error. The contents of this article may be used for educational and non-commercial purposes, including articles for Wikipedia and similar sites. This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

(The reference list is omitted here; see the original.)

Where Indigo landsFurther

Indigo's conclusion

Valuable history, discounted claims. The core lineage holds and usefully de-noises today's RSI talk; the sweeping claims stretch a scholarly lineage into personal ownership. Keep three things: the 40-year timeline, the Gödel Machine as theoretical ceiling, and the endgame in hardware.

What to remember

  1. RSI has a 40-year formal lineage and isn't a 2026 invention: 1987's Meta Evolution was the first concrete RSI algorithm.
  2. The Gödel Machine is RSI's theoretical ceiling: rewrite only after proving the rewrite useful, globally optimal with no local maxima.
  3. Full RSI needs self-improving hardware: software RSI is already practical; the frontier is a self-replicating machine civilization.
  4. Split the priority claims three ways: the core lineage is real, the fringe claims point the right way but stretch, and zero safety anxiety is his standing stance.
  5. Today's LLM versions of RSI continue his line directly: the Darwin-Gödel Machine and others are inspired by the Gödel Machine.

Back on the long-running theses

confirms

Bottlenecks are shifting from software to physics “Full RSI requires self-improving hardware” is a founder-level echo of this view, to sit alongside the Apple-NVLink report and the chip architects' testimony on the physical layer.

conflicts

Noam Brown (OpenAI) on the OAI–HF incident: RSI is the overwhelming top priority Two temperatures of RSI in the same week: Noam's top priority with alignment disasters attached; Schmidhuber's 40-year optimistic project ending in expanding machine life.

adds to

The Dwarkesh RSI debate (Schulman, Millidge, O'Neill) The panel covers the resistance at the foot of the mountain, this note the summit's coordinates; together, the theory and engineering axes of the bounded exponential.

conflicts

Dario Amodei, We must slow the frontier The sharpest clash of stances: Dario full of safety anxiety and calling for a slowdown; Schmidhuber with none, calling RSI's endgame the ultimate form of scaling.

adds to

Haseltine: AI is a second kind of life Section 9's self-replicating hardware is almost an engineering-determinist version of “a second kind of life”: one defined by biological function, the other by self-replicating hardware.

What it means for Rewired Index

A position paper, not a competitive signal: it changes how to place and date the idea of RSI, not any view on a name.

What would change my mind

pure software RSI reaching the endgame without self-replicating hardware, or empirical algorithms overtaking Gödel-Machine-style provable optimality.

Finished. Indigo's take on this piece is in two places: