The father of RSI claims his lineage while it's hot: the family tree holds, and what's truly useful is the Gödel Machine as a theoretical ceiling and “the endgame is hardware”.
Indigo's conclusion
Valuable history, discounted claims. The core lineage holds and usefully de-noises today's RSI talk; the sweeping claims stretch a scholarly lineage into personal ownership. Keep three things: the 40-year timeline, the Gödel Machine as theoretical ceiling, and the endgame in hardware.
How to read this Formatted as a technical note, in substance a priority manifesto, published on 17 September, just as RSI became the industry's number-one topic. There's no empirical data to check; what to audit is whether the priority claims hold and why now. The core lineage is real; sweeping claims like “the T in ChatGPT is mine” are a long-standing habit of his, and get a discount.
What to remember
RSI has a 40-year formal lineage and isn't a 2026 invention: 1987's Meta Evolution was the first concrete RSI algorithm.
The Gödel Machine is RSI's theoretical ceiling: rewrite only after proving the rewrite useful, globally optimal with no local maxima.
Full RSI needs self-improving hardware: software RSI is already practical; the frontier is a self-replicating machine civilization.
Split the priority claims three ways: the core lineage is real, the fringe claims point the right way but stretch, and zero safety anxiety is his standing stance.
Today's LLM versions of RSI continue his line directly: the Darwin-Gödel Machine and others are inspired by the Gödel Machine.
Breakdown · 7 steps
01
RSI is much older than the hype
In 2026 everyone talks RSI, but his 1987 diploma thesis already gave the first concrete algorithms. True RSI is a system that can rewrite its own code in any computable way and keeps only the useful rewrites. Read this part →
02
1987 Meta Evolution, 1994 self-modifying policies
Genetic programming applied to itself, recursively evolving better methods, with meta-meta levels; the 1994 self-modifying policies can change the way they change themselves, and the agent is never reset. Read this part →
03
Networks that program themselves
The 1991 Fast Weight Programmers let one network write another's weights, and from 1992 a network rewrites its own; he says Transformers with linearized self-attention are formally equivalent. Read this part →
04
OOPS and the Gödel Machine: provably optimal
The 2002 OOPS reuses old solutions to speed up new problems; the 2003 Gödel Machine rewrites itself only once it has proved the rewrite useful, so it's globally optimal, with no local maxima. Read this part →
05
Curiosity-driven self-improvement
The 1990 adversarial curiosity: one network produces data that makes the world model err; combined with self-modifying policies, the system decides when and what to learn and keeps inventing new tasks. Read this part →
06
Recent work, and in-context learning in LLMs
New meta-learning work since 2020, and LLM-based methods inspired by the Gödel Machine; he treats in-context learning in LLMs as a special case of meta learning. Read this part →
07
Full RSI requires self-improving hardware
Software RSI is already practical, but there's no superintelligence without mastering the real world; the endgame is a self-replicating, self-improving machine civilization, which he calls the ultimate form of scaling. Read this part →
What it means for Rewired Index
A position paper, not a competitive signal: it changes how to place and date the idea of RSI, not any view on a name.
What would change my mind
pure software RSI reaching the endgame without self-replicating hardware, or empirical algorithms overtaking Gödel-Machine-style provable optimality.
How to read this
Formatted as a technical note, in substance a priority manifesto, published on 17 September, just as RSI became the industry's number-one topic. There's no empirical data to check; what to audit is whether the priority claims hold and why now. The core lineage is real; sweeping claims like “the T in ChatGPT is mine” are a long-standing habit of his, and get a discount.
In 2026 everyone talks RSI, but his 1987 diploma thesis already gave the first concrete algorithms. True RSI is a system that can rewrite its own code in any computable way and keeps only the useful rewrites.
Abstract. As of 2026, everyone—including Anthropic/OpenAI/Sakana AI/SpaceX—is talking about recursive self-improvement (RSI) or meta learning (learning to learn), and startups explicitly brand themselves as RSI companies. RSI is much older than that. In 1987, when compute was about 10⁸ times more expensive than today, I published the first concrete RSI algorithms in my diploma thesis [META1] (Sec. 1). For its cover I drew a robot that bootstraps itself. [META1] was the first in a long series of publications on RSI, which became hot in the 2010s [DEC] and especially the 2020s.
Here I summarize our work on RSI with self-modifying policies since 1994 [METARL2-9] (Sec. 2), gradient descent-based RSI in artificial neural networks since 1992 [FWPMETA1-10] (Sec. 3), asymptotically optimal RSI for curriculum learning since 2002 [OOPS1-3] (Sec. 4), mathematically optimal RSI through the self-referential Gödel Machine since 2003 [GM3-9] (Sec. 5), RSI combined with artificial curiosity and intrinsic motivation [AC] since 1990/1997 (Sec. 6), and recent work on RSI since 2020 (Sec. 7). Computing has become much cheaper, and software-based RSI has become practical. Full RSI, however, will require not just self-improving software but self-improving hardware in the physical world [DLH] (Sec. 9).
The most widely used machine learning algorithms were invented and hardwired by humans. Can we also construct meta learning algorithms that can learn better learning algorithms, to build truly self-improving AIs without any limits other than the limits of computability and physics? This question has been a main drive of my research since my 1987 diploma thesis on this topic [META1][AMA].
First note that meta learning is sometimes confused with simple transfer learning from one training set to another. However, even a standard deep feedforward neural network (NN) [DLH][WHO4-11] can transfer-learn to learn new images faster through pre-training on other image sets, e.g., [TRA12]. True meta learning and RSI is much more than that, and also much more than just learning to adjust hyper-parameters such as mutation rates in evolution strategies.
True RSI is about encoding the initial learning algorithm in a universal programming language (e.g., on a recurrent neural network or RNN), with primitive instructions that allow for modifying the code itself in arbitrary computable fashion. We surround this self-referential, self-modifying code by a recursive framework that ensures that only "useful" self-modifications survive, e.g., Sec. 2, Sec. 5.
Meta learning may be the most ambitious but also the most rewarding goal of machine learning. There are few limits to what a good meta learner will learn. Where appropriate, it will learn to learn by analogy, by chunking, by planning, by subgoal generation, by combinations thereof—you name it.
02
1987 Meta Evolution, 1994 self-modifying policies
Genetic programming applied to itself, recursively evolving better methods, with meta-meta levels; the 1994 self-modifying policies can change the way they change themselves, and the agent is never reset.
1. Meta Evolution and PSALMs (1987)
In 1987, we published [GP87] [GP] what I think was the first paper on Genetic Programming or GP for evolving programs of unlimited size written in a universal programming language [GOD][GOD34][CHU][TUR][POS].
In the same year, Sec. 2 of my diploma thesis [META1] applied such GP to itself, to recursively evolve better GP methods. There was not only a meta level but also a meta meta level and a meta meta meta level etc. I called this RSI method Meta Evolution.
Sec. 4 of [META1] also introduced meta learning Prototypical Self-Referential Associating Learning Mechanisms (PSALMs) for payoff maximisation or Reinforcement Learning (RL). This was a first kind of meta meta RL or RL-based RSI.
This work concretizes aspects of I. J. Good's informal and speculative remarks (1966) on an "intelligence explosion" through self-improving "super-intelligences" [GOOD] (Good did not have any concrete RSI algorithms), and Bellman's thoughts on "metapolicies" (1967) [BE67].
2. RSI for Reinforcement Learning with Self-Modifying Policies (1994-)
In 1994, I proposed another type of meta RL or RSI called incremental self-improvement [METARL2] for general purpose RL machines with a single life consisting of a single lifelong trial. That is, unlike in traditional RL, there is no assumption of repeatable independent trials, and the RL agent is never reset. It is driven by a self-modifying policy (SMP) which is a modifiable probability distribution over programs written in a universal programming language [GOD][GOD34][CHU][TUR][POS], to allow for arbitrary computations. The learning algorithm of an SMP is part of the SMP itself—SMPs can modify the way they modify themselves. The credit assignment process has to take into account that early self-modifications are setting the stage for later ones.
A method called Environment-Independent Reinforcement Acceleration (EIRA) [METARL4] or Success-Story Algorithm [METARL7-9] forces SMPs to come up with better and better self-modification algorithms that continually improve reward intake per time [METARL2-9]. This worked well in challenging experiments, although compute back then was 100,000 times more expensive than today.
03
Networks that program themselves
The 1991 Fast Weight Programmers let one network write another's weights, and from 1992 a network rewrites its own; he says Transformers with linearized self-attention are formally equivalent.
3. Gradient-Based RSI in NNs that Learn to Program Other NNs (1991) and Themselves (1992)
As I have frequently pointed out since 1990 [AC90], the connection strengths or weights of an artificial neural network (NN) should be viewed as its program. Inspired by Gödel's universal self-referential formal systems [GOD][GOD34], I built NNs whose outputs are programs or weight matrices of other NNs: the so-called Fast Weight Programmers [FWP0-2][FWP]. I even built self-referential recurrent NNs (RNNs) that can run and inspect their own weight change algorithms or learning algorithms [FWPMETA1-10]. A difference to Gödel's work was that my universal programming language was not based on the integers, but on real-valued weights, such that each NN's output is differentiable with respect to its program. That is, a simple program generator (the efficient gradient descent procedure [BP1]—compare [BP2] [BPA] [BP4] [R7]) can compute a direction in program space where one may find a better program [AC90], in particular, a better program-generating program [FWP0-2]. Much of my work since 1989 has exploited this fact.
Successful learning in deep architectures started in 1965 when Ivakhnenko & Lapa published the first general, working learning algorithms for deep multilayer perceptrons with arbitrarily many hidden layers. Their nets already contained the now popular multiplicative gates [DEEP1-2] [DL1][DL2][DLH], an essential ingredient of what was later called NNs with dynamic links or fast weights. In 1981, v. d. Malsburg was the first to explicitly emphasize the importance of NNs with such rapidly changing connections [FAST]; others followed [DLP].
However, these authors did not yet have an end-to-end differentiable system that learns by gradient descent to quickly manipulate the fast weight storage. Such a system I published in 1991 [FWP0][FWP1][ULTRA]. There a slow NN learns to control the weight changes of a separate fast NN. That is, I separated storage and control like in traditional computers, but in a fully neural way (rather than in a hybrid fashion [PDA1] [PDA2] [DNC]). (Compare my related work on what's now sometimes called Synthetic Gradients [NAN1-5].)
Then I showed how fast weights can be used for RSI or "learning to learn." In references [FWPMETA1-5] since 1992, the slow RNN and the fast RNN are identical. The RNN can see its own errors or reward signals called eval(t+1) in the image (from [FWPMETA5]). The initial weight of each connection is trained by gradient descent, but during a training episode, each connection can be addressed and read and modified by the RNN itself through O(log n) special output units, where n is the number of connections—see time-dependent vectors mod(t), anal(t), Δ(t), val(t+1) in the image. That is, each connection's weight may rapidly change, and the network becomes self-referential in the sense that it can in principle run arbitrary computable weight change algorithms or learning algorithms (for all of its weights) on itself: recursive self-improvement for NNs!
In 1991-93, I simplified this through gradient descent-based, active control of fast weights through 2D tensors or outer product updates [FWP2] (compare our more recent work on this [FWP3] [FWP3a]). One motivation was to get many more temporal variables under massively parallel end-to-end differentiable control than what's possible in standard RNNs of the same size: O(H²) instead of O(H), where H is the number of hidden units (compare Sec. 8 of [MIR] and Sec. H4 of [DLP]). The 1993 paper [FWP2] also explicitly addressed the learning of internal spotlights of attention in end-to-end differentiable networks [FWP2] [ATT].
Unnormalized Transformers with linearized self-attention [TR5-6] are formally equivalent to my 1991 outer product-based Fast Weight Programmers, now called unnormalized linear Transformers [ULTRA][MOST]—see the T in ChatGPT.
In 2001, my former student Sepp Hochreiter used gradient descent in LSTM networks [LSTM1] instead of traditional RNNs to meta learn fast online learning algorithms for nontrivial classes of functions, such as all quadratic functions of two variables [HO1].
04
OOPS and the Gödel Machine: provably optimal
The 2002 OOPS reuses old solutions to speed up new problems; the 2003 Gödel Machine rewrites itself only once it has proved the rewrite useful, so it's globally optimal, with no local maxima.
4. Asymptotically Optimal RSI for Curriculum Learning (2002-)
In 2002, I introduced a general and asymptotically time-optimal type of curriculum learning, that is, solving one problem after another, efficiently searching the space of programs that compute solution candidates, including those programs that organize and manage and adapt and reuse earlier acquired knowledge [OOPS1-3]. The Optimal Ordered Problem Solver (OOPS) draws inspiration from Levin's Universal Search [OPT] designed for single problems. It spends part of the total search time for a new problem on testing programs that exploit previous solution-computing programs in computable ways. If the new problem can be solved faster by copy-editing/invoking previous code than by solving the new problem from scratch, then OOPS will find this out. If not, then at least the previous solutions will not cause much harm. I introduced an efficient, recursive, backtracking-based way of implementing OOPS on realistic computers with limited storage. Experiments illustrated how OOPS can greatly profit from meta learning or meta searching, that is, searching for faster search procedures in RSI style [OOPS1-2].
The self-referential RSI system of Sec. 2 above (1994-) justified its self-modifications through growing statistical evidence of subsequent reward accelerations. But it was not guaranteed to execute theoretically optimal self-improvements. This motivated my Gödel Machine [GM3-9], the first fully self-referential universal [UNI] RSI machine that was indeed optimal in a certain mathematical sense. Typically it uses the somewhat less general Optimal Ordered Problem Solver [OOPS1-2] (Sec. 4) for finding provably optimal self-improvements.
The RSI Gödel Machine is inspired by Kurt Gödel, the founder of theoretical computer science in the early 1930s [GOD][GOD34][GOD21,a,b]. He introduced a universal coding language based on the integers which allows for formalizing the operations of any digital computer in axiomatic form. Gödel used it to represent both data (such as axioms and theorems) and programs (such as proof-generating sequences of operations on the data). He famously constructed formal statements that talk about the computation of other formal statements, especially self-referential statements which imply that their truth is not decidable by any computational theorem prover. Thus he identified fundamental limits of mathematics and theorem proving and computing and Artificial Intelligence (AI) [GOD][GOD21,a,b]. This had enormous impact on science and philosophy of the 20th century. Furthermore, much of early AI in the 1940s-70s was actually about theorem proving and deduction in Gödel style through expert systems and logic programming. Compare Sec. 18 of [MIR].
A Gödel Machine [GM6] is a general RL machine that will rewrite any part of its own code as soon as it has found a proof that the rewrite is useful, where the problem-dependent utility function and the hardware and the entire initial code are described by axioms encoded in an initial proof searcher which is also part of the initial code. While the machine is interacting with its environment (initially in a suboptimal way), the searcher systematically and efficiently tests computable proof techniques (programs whose outputs are proofs) until it finds a provably useful, computable self-rewrite. I showed that such a self-rewrite is globally optimal—no local maxima!—since the code first had to prove that it is not useful to continue the proof search for alternative self-rewrites. Unlike previous non-self-referential methods based on hardwired proof searchers, the Gödel Machine not only boasts an optimal order of complexity but can optimally reduce any slowdowns hidden by the O()-notation, provided the utility of such speed-ups is provable at all [GM3-9].
05
Curiosity-driven self-improvement
The 1990 adversarial curiosity: one network produces data that makes the world model err; combined with self-modifying policies, the system decides when and what to learn and keeps inventing new tasks.
6. RSI plus Artificial Curiosity and Intrinsic Motivation (1990, 1997-)
Before I continue the discussion of meta learning and RSI, let me first explain RL with intrinsic motivation. My popular principle of adversarial artificial curiosity from 1990 [AC90, AC90b] [AC20] (see also surveys [AC09] [AC10]) is now widely used not only for exploration in RL but also for image synthesis [AC20][DLP]. It works as follows. One NN (the controller) probabilistically generates outputs, another NN (the world model) sees those outputs and predicts environmental reactions to them. Using gradient descent, the world model NN minimizes its error, while the generator NN tries to make outputs that maximize this error. One net's loss is the other net's gain. So the controller is intrinsically motivated to generate output actions or experiments that yield data from which the world model can still learn something. (GANs are a special case of this where the environment simply returns 1 or 0 depending on whether the generator's output is in a given set [AC20]; compare [R2][LEC] and Sec. 5 of [MIR] and [WHO8].)
The Section "A Connection to Meta learning" in [AC90] (1990) already pointed out: "A model network can be used not only for predicting the controller's inputs but also for predicting its future outputs. A perfect model of this kind would model the internal changes of the control network. It would predict the evolution of the controller, and thereby the effects of the gradient descent procedure itself. In this case, the flow of activation in the model network would model the weight changes of the control network. This in turn comes close to the notion of learning how to learn." The paper [AC90] also introduced planning with recurrent NNs (RNNs) as world models [PLAN,PLAN2-5], and high-dimensional reward signals. Unlike in traditional RL, those reward signals were also used as informative inputs to the controller NN learning to execute actions that maximise cumulative reward (see also Sec. 13 of [MIR] and Sec. 5 of [DEC]).
This is important for meta learning: an NN that cannot see its own errors or rewards cannot learn a better way of using such signals as inputs for self-invented learning algorithms.
A few years later, I combined the RSI RL system of Sec. 2 and Adversarial Artificial Curiosity in a single system [AC97, AC99, AC02]. It generates computational experiments in form of programs whose execution may change both an external environment and the RL agent's internal state. An experiment has a binary outcome: either a particular effect happens, or it doesn't. Experiments are collectively proposed by two reward-maximizing adversarial policies. Both can predict and bet on experimental outcomes before they happen. Once such an outcome is actually observed, the winner will get a positive reward proportional to the bet, and the loser a negative reward of equal magnitude. So each policy is motivated to create experiments whose yes/no outcomes surprise the other policy. The latter in turn is motivated to learn something about the world that it did not yet know, such that it is not outwitted again.
Using RSI with self-modifying policies [METARL2-9] (Sec. 2), the system learns when to learn and what to learn [AC97, AC99, AC02]. It will also minimize the computational cost of learning new skills, provided both brains receive a small negative reward for each computational step, which introduces a bias towards simple still surprising experiments (reflecting simple still unsolved problems). This may facilitate hierarchical construction of more and more complex experiments, including those yielding external reward (if there is any). In fact, this type of artificial creativity may not only drive artificial scientists and artists [AC06-09], but can also accelerate the intake of external reward [AC97] [AC02], intuitively because a better understanding of the world can help to solve certain problems faster.
The more recent, intrinsically motivated PowerPlay RL system (2011) [PP] [PP1] can use the meta learning OOPS [OOPS1-2] (Sec. 4) to continually invent on its own new goals and tasks, incrementally learning to become a more and more general problem solver in an active, partially unsupervised or self-supervised fashion. RL robots with high-dimensional video inputs and intrinsic motivation (like in PowerPlay) learned to explore in 2015 [PP2].
06
Recent work, and in-context learning in LLMs
New meta-learning work since 2020, and LLM-based methods inspired by the Gödel Machine; he treats in-context learning in LLMs as a special case of meta learning.
7. More Recent Work on RSI and Meta Learning (2020-)
My former PhD student Imanol Schlag et al. [FWPMETA7] augmented an LSTM with an associative Fast Weight Memory (FWM). Through differentiable operations at every step of a given input sequence, the LSTM updates and maintains compositional associations of former observations stored in the rapidly changing FWM weights. The model is trained end-to-end by gradient descent and yields excellent performance on compositional language reasoning problems, small-scale word-level language modelling, and meta RL for partially observable environments [FWPMETA7].
Our MetaGenRL (2020) [METARL10] meta learns novel RL algorithms applicable to environments that significantly differ from those used for training. MetaGenRL searches the space of low-complexity loss functions that describe such learning algorithms. See the blog post of my former PhD student Louis Kirsch.
This principle of searching for simple learning algorithms is also applicable to fast weight architectures. Our recent Variable Shared Meta Learning (VS-ML) merges weight sharing and sparsity in RSI RNNs [FWPMETA6]. This allows for encoding the learning algorithm by few parameters although it has many time-varying variables—compare [FWP2] (Sec. 3). VS-ML combines end-to-end differentiable fast weights [FWP1-3a] (Sec. 3) and learning algorithms encoded in the activations of LSTMs [HO1]. Some of these activations can be interpreted as NN weights updated by the LSTM dynamics. LSTMs with shared sparse entries in their weight matrix discover learning algorithms that generalize to new datasets. The meta learned learning algorithms do not require explicit gradient calculation. VS-ML in RNNs can also learn to implement the famous backpropagation learning algorithm [BP1] [BP2] [BP4] purely in the end-to-end differentiable forward dynamics of RNNs [FWPMETA6].
In 2022, we also published at ICML a modern self-referential weight matrix (SWRM) [FWPMETA8] based on the 1992 SWRM [FWPMETA1-5] (see Sec. 3). In principle, it can meta learn to learn, and meta meta learn to meta learn to learn, and so on, in the sense of recursive self-improvement. We evaluated our SRWM on supervised few-shot learning tasks and on multi-task reinforcement learning with procedurally generated game environments. The experiments demonstrated both practical applicability and competitive performance of the SRWM.
Recent work on meta learning and RSI focused on LLM-based approaches inspired by the Gödel Machine [GM3-9] (Sec. 5), e.g., the Darwin-Gödel Machine [GMD25] developed at Sakana AI, the Huxley-Gödel Machine [GMH26], and the Red Queen Gödel Machine [GMR26]. See our 2026 survey [RSI26].
Computing is much cheaper today than it was in the previous millennium, and finally, a number of companies are starting to focus on meta learning and RSI, e.g., Sakana AI, Ricursive, Recursive Superintelligence, Anthropic, OpenAI, Inherent, and others.
8. "In-Context Learning" of LLMs is a Special Case of Meta Learning
To a certain extent, recent Large Language Models (LLMs) can learn from a growing record of user interactions without changing the weights of the underlying pre-trained Transformer NN through gradient descent during test time. This so-called "In-Context Learning" and "Test Time Training" of LLMs is a special case of meta learning, similar to the 1991 unnormalized linear Transformer [ULTRA] and the 2001 meta learning LSTM [HO1] which learned by gradient descent a learning algorithm for quadratic functions that was much faster than gradient descent (Sec. 3), without executing additional weight changes during test time [COCO]!
Generally speaking, gradient descent can be used to learn a learning algorithm running on the neural network itself, as shown in 1992 [FWPMETA1-5] (Sec. 3). Many meta learners actually learn to program fast weights [FWP], and Transformers do so too, including the 1991 unnormalized linear Transformer [ULTRA][FWP0-1,6].
07
Full RSI requires self-improving hardware
Software RSI is already practical, but there's no superintelligence without mastering the real world; the endgame is a self-replicating, self-improving machine civilization, which he calls the ultimate form of scaling.
9. Full RSI Requires Self-Improving Hardware
The RSI discussion above focused on self-improving software. However, as pointed out earlier [DLH], to achieve True AI, software research per se is not enough, it has to be combined with the physical world of machines and robots. No Artificial Super Intelligence (ASI) without mastery of the real world! The Gödel Machine [GM3-9] (Sec. 5) takes this into account, and I'd like to end this report with a quote from [DLH] (based on earlier publications):
"For centuries, humans have discussed physical self-replicating machines (SRMs), including Descartes in the 1600s, Eliot and Butler in the 1800s, Capek in the 1920s, von Neumann & Zuse & Penrose and others since the 1940s [SRM20]. While self-replicating and evolving software is almost trivial (think of computer viruses), nobody knew how to build general purpose physical SRMs in practice. However, now there seems to be an obvious way: AI-controlled general-purpose robots that can learn to operate all the machines and tools currently operated by humans [COG18] will also be able to build/operate/repair the machines required to make more of those robots. This includes machines that mine the raw material from the ground, refine it, screw parts together, repair broken 3D printers and robots and robot factories, and so on, doing all the physical jobs that currently only machine-operating humans can do. Basically, a machine civilisation that can self-replicate without machine-operating humans, and then, of course, improve itself. Self-improving hardware, as opposed to the already existing, self-improving, meta-learning software. [META] I called this the ultimate form of scaling [JY24][FA24][95-25]."
Acknowledgments
Thanks to several expert reviewers for useful comments. Since science is about self-correction, let me know under juergen@idsia.ch if you can spot any remaining error. The contents of this article may be used for educational and non-commercial purposes, including articles for Wikipedia and similar sites. This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
(The reference list is omitted here; see the original.)
Where Indigo landsFurther
Indigo's conclusion
Valuable history, discounted claims. The core lineage holds and usefully de-noises today's RSI talk; the sweeping claims stretch a scholarly lineage into personal ownership. Keep three things: the 40-year timeline, the Gödel Machine as theoretical ceiling, and the endgame in hardware.
What to remember
RSI has a 40-year formal lineage and isn't a 2026 invention: 1987's Meta Evolution was the first concrete RSI algorithm.
The Gödel Machine is RSI's theoretical ceiling: rewrite only after proving the rewrite useful, globally optimal with no local maxima.
Full RSI needs self-improving hardware: software RSI is already practical; the frontier is a self-replicating machine civilization.
Split the priority claims three ways: the core lineage is real, the fringe claims point the right way but stretch, and zero safety anxiety is his standing stance.
Today's LLM versions of RSI continue his line directly: the Darwin-Gödel Machine and others are inspired by the Gödel Machine.
Back on the long-running theses
confirms
Bottlenecks are shifting from software to physics “Full RSI requires self-improving hardware” is a founder-level echo of this view, to sit alongside the Apple-NVLink report and the chip architects' testimony on the physical layer.
The Dwarkesh RSI debate (Schulman, Millidge, O'Neill) The panel covers the resistance at the foot of the mountain, this note the summit's coordinates; together, the theory and engineering axes of the bounded exponential.
conflicts
Dario Amodei, We must slow the frontier The sharpest clash of stances: Dario full of safety anxiety and calling for a slowdown; Schmidhuber with none, calling RSI's endgame the ultimate form of scaling.
adds to
Haseltine: AI is a second kind of life Section 9's self-replicating hardware is almost an engineering-determinist version of “a second kind of life”: one defined by biological function, the other by self-replicating hardware.
What it means for Rewired Index
A position paper, not a competitive signal: it changes how to place and date the idea of RSI, not any view on a name.
What would change my mind
pure software RSI reaching the endgame without self-replicating hardware, or empirical algorithms overtaking Gödel-Machine-style provable optimality.
Finished. Indigo's take on this piece is in two places: