DEV Community

Cover image for Recursive self-improvement: what Google's Dream-RSI paper really does

Recursive self-improvement: what Google's Dream-RSI paper really does

Recursive self-improvement is the idea that an AI makes itself smarter, then uses the smarter version to do it again. This week Google researchers published Dream-RSI: Recursive Self-Improvement through Evolving Worlds, and X announced that Google had "cracked" it. I read the 36 pages. The loop is real and the savings are measured, but in the whole paper the model's weights never move.

TL;DR

  • The phrase: since I. J. Good in 1965, recursive self-improvement has meant a system that improves its own intelligence, pass after pass. In AI safety it means a model rewriting its own code or weights.
  • The paper: Dream-RSI improves a search policy, a small Python program that decides which ideas a frozen Gemini coding agent explores next. The coder, the evaluator and the policy rewriter are all fixed models.
  • The headline number: on a Lasso solver task, 317 Gemini calls instead of 550 for a faster result (2.93 s against 3.59 s). That is 42 % fewer calls. A discount; nothing here looks like takeoff.
  • The tie: on circle packing it scored 2.635983, exactly what AlphaEvolve V2 scored in 2025, to six decimals.
  • How to read the next "RSI" paper: ask what object changes, who is frozen, and what the headline number is a unit of.

What is recursive self-improvement?

The idea is older than the field's hardware. In 1965 I. J. Good, a statistician who had worked with Turing at Bletchley Park, wrote in "Speculations Concerning the First Ultraintelligent Machine":

an ultraintelligent machine could design even better machines; there would then unquestionably be an 'intelligence explosion'… Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.

Most people quote that up to the comma. The second half is the whole AI-safety field in one clause.

In 2008 Eliezer Yudkowsky's LessWrong essay "Recursive Self-Improvement" fixed the meaning the safety community still uses: a seed AI that rewrites its own source or weights and gets smarter with each pass, so each improvement makes the next one easier. For forty years nobody had a machine to test the definition on, which is the ideal condition for a definition.

The part that matters for everything below: in that sense, the thing that improves is the thing doing the improving. That is what makes it recursive.

From AlphaEvolve to Dream-RSI

In May 2025 Google DeepMind announced AlphaEvolve, a Gemini-powered coding agent in an evolutionary loop: propose a program, score it with an automatic evaluator, keep the best, mutate them again. The results were real. It found a way to multiply 4×4 complex matrices with 48 scalar multiplications, beating Strassen's 1969 algorithm. A scheduling heuristic it found for Borg, Google's cluster manager, recovers on average 0.7 % of Google's worldwide compute. It sped up a kernel in Gemini's own training by 23 %, cutting total training time by 1 %.

That last one is where the phrase started to move. Gemini helped make Gemini's training cheaper, and "self-improvement" crossed from safety blogs into press releases. But in AlphaEvolve the model is still fixed. What evolves is the programs it writes.

Dream-RSI, from 17 authors at Google, Google DeepMind, the University of Maryland and the University of Virginia, adds one layer on top of that kind of loop. The abstract says what it is: "A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged."

How Dream-RSI works, under the hood

There are three models in the loop, and all three are frozen:

  1. A discovery agent (Gemini, through Gemini CLI) writes candidate code for a task, for example a faster Lasso solver.
  2. A fixed evaluator scores each candidate.
  3. A policy-development agent, another frozen LLM, rewrites the exploration policy.

The only thing that changes is the exploration policy: an executable Python program that decides which branches of the search tree to expand, how many in parallel, how deep to go, and when to stop. Here is its minimal structure, verbatim from Appendix B.2 of the paper:

def solve(self, question, budget=None):
    question.reset()
    res, closed = SimResult(), set()
    while not _budget_done(question, budget):
        prefix = question.observed()
        update_closed(closed, prefix, question)
        batch = select_batch(prefix, question, closed)
        if not batch:
            break
        question.probe_batch(batch, on_reveal=lambda _: _record_curve(res, question))
    return finalize_result(question, res)
Enter fullscreen mode Exit fullscreen mode

Read it as a budgeted search: look at what has been revealed so far, mark branches as closed, pick the next batch to probe, repeat until the budget runs out. The policy never writes the solution. It only chooses where the coding agent looks.

The loop around it has four steps (Fig. 1 and §3):

Step What happens Cost
1. Online exploration The policy guides the frozen coder; the evaluator scores candidates; the result is a recorded discovery tree Real agent calls
2. Replay world The recorded tree becomes a simulator: every node's outcome is already known None
3. Dreaming The policy agent rewrites the policy and tests each version against the recorded outcomes Near zero
4. Redeploy The best dreamed policy runs the next real round; the pool of recorded worlds grows Real agent calls

Step 3 is the idea. In the paper's words, "a single costly online run enables thousands of rapid, zero-execution-cost off-policy evaluations." You pay once to explore for real, then replay that exploration thousands of times to find a better way to explore. The name "dreaming" comes from there.

There is a rule that keeps the dreaming honest, the prefix-only rule. A policy may use only what has been revealed so far: "Never use unrevealed scores, a true optimum, hardcoded winning cell ids". Otherwise a policy tested in a replay world could just memorise where the good nodes were.

And the prompt that drives the rewriter spells out the scope. This is the part the "Google cracked RSI" posts did not quote:

Appendix B.2, Listing 2, lines 1 to 3:

"You are improving one prefix-only exploration policy. Edit only {method_file} and implement OptimalPolicy.solve(self, question, budget=None). Do not solve the scientific task and do not edit any other program." The self-improving AI is told in writing to touch one file.

What Dream-RSI measured

The paper tests eight discovery tasks in three domains: algorithm engineering (a Lasso regularization path), mathematical optimisation (sum-difference, autocorrelation, circle packing) and GPU kernels (four KernelBench tasks). Budgets were held equal per round: with Gemini 3.1 Pro, 10 parallel workspaces times 11 refinement steps, 110 agent calls per round.

Lasso. With Gemini 3.1 Pro, fixed exploration used 550 agent calls to reach an average runtime of 3,587 ms on six held-out datasets. Dream-RSI used 317 calls to reach 2,931 ms. With Gemini 3.7 Flash: 3,200 calls for 2,517 ms against 1,879 calls for 2,351 ms. The discovered solvers beat scikit-learn and glmnet on all six datasets; glmnet averaged 13,768 ms.

Lasso task: average runtime on six held-out datasets for glmnet, fixed exploration (550 calls) and Dream-RSI (317 calls)

Kernels. On VGG16 and LayerNorm it reached comparable kernel performance with 2.43× and 1.79× fewer generations, and on two convolution tasks up to 2.09× faster kernels for the same budget.

Maths. Against SimpleTES, which used GPT-OSS-120B and 51,200 generations, the paper reports "roughly two orders of magnitude fewer discovery-agent calls", up to 162×. On autocorrelation SimpleTES still holds the state of the art, which the paper says. On circle packing Dream-RSI scored 2.635983. AlphaEvolve V2 scored 2.635983 in 2025.

Circle packing, Table 1: AlphaEvolve V2 (2025) and Dream-RSI (2026) both at 2.635983

The paper's own summary is careful: "competitive or improved discovery quality while substantially reducing discovery cost in several settings". Same answers, fewer calls. The improvement is in the bill.

Is Dream-RSI really recursive self-improvement?

Put the three meanings next to each other:

What changes What stays frozen Recursive in Good's sense?
Classic RSI (Good, Yudkowsky) The model's own code or weights Nothing Yes, by definition
AlphaEvolve (2025) Candidate programs for a task The Gemini model, the evaluator No, the model does not improve
Dream-RSI (2026) The search policy that guides the coder Coder, evaluator, policy rewriter Partly: the search improves the search

There is a loop, and the loop feeds on its own history: each round's recorded tree trains the next round's policy. That is a fair use of "recursive". What is missing is the "self". The model that does the work comes out of the paper exactly as it went in.

Hacker News read the PDF (thread, 169 points). rybosworld: "Unless I'm misunderstanding, calling this RSI seems misleading? This looks like an optimization of current training methods, and a good one, but not 'RSI' in the sense of a system that can perpetually improve itself forever." jephs called it "much closer to the concept of continual learning".

X read the title:

@thesupermannx on X:

And the same week, on September 12, Anthropic's CEO Dario Amodei wrote in "We Must Pace the Frontier" that recursive self-improvement has been under way "since roughly this summer… including at Anthropic", and that "We must slow the pace at which we improve the capabilities of AI models." Three sentences with the same two words, about three different machines. One of them has a paper. The paper says the weights are frozen.

How to read the next RSI paper

This is the Monday takeaway from the episode, as a checklist:

  1. What object changes? Weights, the model's own code, or the search settings around a fixed model? Dream-RSI changes the third.
  2. Who is frozen? Here the coder, the judge and the rewriter, all three. If everything that thinks is frozen, the intelligence is not what is improving.
  3. What is the headline number a unit of? Agent calls saved, generations saved and milliseconds of runtime are optimisation. The takeoff would be a result the model could not reach before at any budget, and it would have to keep compounding. Nobody has published that one yet.

If you build agent systems, the paper is still worth your time for a practical reason: step 3 is a cheap way to tune an agent's search. Record your agent's real exploration once, replay it as a simulator, and test new scheduling policies against it before you spend another GPU hour. The code is on GitHub (293 stars on September 17, and no license file yet, so check before you reuse it) with a project page.

Verdict: NEEDS REVIEW

I stamped it NEEDS REVIEW. The loop is real, the savings are measured (317 against 550 calls on Lasso, 2.43× fewer generations on a kernel), and the replay trick is a good idea. But the two words on the cover describe a machine that is not in the paper. The weights never move.

FAQ

What is recursive self-improvement in AI?
A system that improves its own intelligence, then uses the improved version to improve itself again. The term goes back to I. J. Good's 1965 "intelligence explosion"; in AI safety it usually means a model rewriting its own code or weights.

Did Google achieve recursive self-improvement with Dream-RSI?
Not in that sense. Dream-RSI improves a search policy around a frozen Gemini coding agent. The paper itself says the agent stays "unchanged".

What is the difference between AlphaEvolve and Dream-RSI?
AlphaEvolve evolves candidate programs for a task. Dream-RSI adds a layer that evolves the policy deciding which candidates to explore, and tests new policies against recorded runs instead of live ones.

What is an intelligence explosion?
I. J. Good's 1965 term for a chain of machines each designing a better one, ending in intelligence far above human. No published system has shown it.

Sources


This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.

Top comments (0)