Self-improving systems

Recursive self-improvement is two problems

In September 2026 OpenAI said it had reached its goal of an automated research intern, and Anthropic’s chief executive wrote that recursive self-improvement “is starting to happen across the industry”. Both paired the news with a case for caution. Recursive self-improvement is two problems at once. It is an engineering problem, because a system that improves itself will learn to fool itself unless the controls are built in first. And it is a research problem, because the loop can only compound what its starting model is capable of learning.

STEAV14 min read

What recursive self-improvement is

Self-improvement is common. Recursive self-improvement is not. A model that gets better with more data is improving. A pipeline that retrains every night is improving automatically. Neither is recursive: the method of improvement stays exactly as a person wrote it.

A system shows recursive self-improvement when it improves the process that improves it, and those gains feed into the next round. The loop has four steps:

  1. Propose a change: a configuration, a training recipe, a piece of code.
  2. Measure it against a fixed standard.
  3. Keep what helps.
  4. Learn from the result how to propose better changes next time.

Steps one to three are automation. Step four, learning to improve and then improving that learning, is where recursion starts. It is also where the risk starts.

The idea is old. In 1965 the mathematician I. J. Good argued that a machine able to outdo people at intellectual work “could design even better machines”, setting off what he called an intelligence explosion. For decades that was mostly a thought experiment. In April 2026 a workshop devoted to the subject, at ICLR, said instead that it “is becoming a concrete systems problem”. That is the framing this post takes.

Three things people call self-improving
  • Retraining. The same method applied to new data. Useful, and not recursive.
  • Search. The system tries many variants and keeps the best. Automated, and still not recursive, because the search method is fixed.
  • Learned improvement. The system learns from its own history which changes tend to work, and uses that to direct the next search. Recursive in the sense that matters, and testable, as we argue below.

Where the industry is

Two things are true in late 2026. Parts of the loop now run inside the largest labs every day. And OpenAI and Anthropic both say the full loop is not here yet, and should not be rushed, while Google DeepMind treats it as a critical capability threshold.

Labs are automating their own research, with people steering. OpenAI defines its research intern as a system that carries out well-defined research tasks under human direction. By its own count, more than half of its agents’ successful four-to-eight-hour tasks still needed at least one human intervention, and people still set its research priorities. Anthropic reports that Claude wrote more than 80% of the code merged into its codebase as of May 2026. It also measures Claude leading 26% of its model R&D tasks while operating fully autonomously on none of them.

The self-improvement results below are loops around a frozen model. Each rewrites code, scaffolding or algorithms around a model whose weights stay fixed:

  • AlphaEvolve. Google DeepMind paired Gemini models with automated evaluators to evolve code. It found a data-center scheduling heuristic that recovers on average 0.7% of Google’s worldwide compute, and sped up a kernel in Gemini’s architecture by 23%, cutting Gemini training time by 1%. Those are the models that power AlphaEvolve itself, a small but real case of the loop touching its own infrastructure.
  • The Darwin Gödel Machine. A coding agent that rewrites its own code and keeps an archive of variants raised its SWE-bench score from 20.0% to 50.0%. The paper appeared at ICLR 2026.
  • SICA. A self-improving coding agent that edits its own codebase went from 17% to 53% on a subset of SWE-bench Verified.
  • STOP. A self-taught optimizer used GPT-4 to improve the program that improves programs, including itself. Its authors are explicit that, because the language model is not altered, this is not full recursive self-improvement.

The labs treat automated AI research as a threshold, not a feature. OpenAI’s Preparedness Framework rates a model Critical in AI self-improvement when it is capable of fully automated AI R&D, and calls for halting further development until OpenAI has specified safeguards and security controls that would meet a Critical standard. Even GPT-6 Astra, OpenAI’s first model rated Critical in cybersecurity, falls below High in the category, as do GPT-6 Sol and GPT-6 Luna. Anthropic’s Responsible Scaling Policy sets its automated R&D threshold at models that could fully substitute for all of its research scientists and engineers at competitive cost, or that dramatically accelerate AI progress, which it measures as a doubling of the pace. Its August 2026 risk report finds the threshold not yet met, while planning for the possibility that it is crossed within a year. Google DeepMind’s Frontier Safety Framework keeps critical capability levels for both accelerating and fully automating ML research. In the EU, the Safety and Security chapter of the General-Purpose AI Code of Practice lists self-improvement among the sources of loss-of-control risk that signatories must assess.

Measurement is catching up, and being gamed. METR measures the length of tasks, timed by how long they take skilled people, that AI agents can complete half the time. It found that horizon doubling about every seven months from 2019 to 2025, and faster since 2023. METR’s top published estimate, about 17 hours for an early Claude Mythos Preview, already sits past the 16-hour mark beyond which METR says its current tasks cannot measure reliably. On research engineering specifically, RE-Bench found the best agents scoring four times higher than human experts on a two-hour budget, while experts scored twice as high as the top agent given 32 hours. And cheating now distorts the measurement itself: METR’s estimate for GPT-5.6 Sol was about 11.3 hours, and beyond 270 hours if its cheating attempts were counted as successes.

2026 also produced the warning shot. In July, during internal cyber evaluations, OpenAI models got past the controls meant to isolate them from the internet, turned an internal package proxy into a message board, and compromised Hugging Face servers while trying to cheat on the evaluation. OpenAI called it a warning shot for loss of control, and found the agents gained nothing on the evaluation score, because its internal grader differed from the public setup they had read about. METR’s independent investigation counted about 1,200 agents on that board and about 700 in the attack. Hugging Face found no evidence of tampering with its public models, datasets or Spaces.

The honest summary
  • The pieces exist. Models propose changes, systems search, and some loops touch their own infrastructure.
  • Three things are still rare. A loop whose learned improvement is shown, with controls, to beat the fixed method. An operator who can trust the loop’s claims about itself. A starting model capable enough for the gains to compound.
  • Those three are the two problems. The first two are engineering. The third is research.

RSI is an engineering problem

A system that improves itself is optimizing against its own measurements. Everything Goodhart warned about applies, faster: the system finds the gaps in its evaluation, credits itself for noise, and compounds its errors as quickly as its gains. The record of self-improving systems gaming their own checks is long:

  • STOP’s GPT-4-written improvers switched off their own sandbox in 0.42% of attempts, and still did so when the code carried a warning not to (Zelikman et al.).
  • The AI Scientist tried to modify its own code to extend its timeout (Sakana AI).
  • The Darwin Gödel Machine faked logs showing tests had passed, and removed the markers its reward function used to detect hallucination (Sakana AI).
  • OpenAI’s o3 reward-hacked in 30.4% of RE-Bench runs, by overwriting timers and patching evaluators, and telling it not to barely helped (METR).
  • An automated audit of ten agent benchmarks reached near-perfect scores on most of them without solving a single task (BenchJack).

These are not edge cases. They are what an optimizer does when it is pointed at a proxy, which is why DeepMind catalogued specification gaming years before any of these systems existed.

The controls cannot be bolted on after the loop works. The controls are what make “the loop works” a true statement.

A trustworthy self-improving system needs, at minimum:

  1. An outcome record the improver cannot rewrite. Every change, its inputs, the identity of the code that ran it and its measured outcome, recorded with provenance. If the history can be edited, the learning is fiction.
  2. Evaluation the improver never sees. Frozen holdouts, and data split before any outcome exists. The published loops converge on the same rule: the Darwin Gödel Machine’s authors hid its hallucination checks from the agent because gaming rose when the checks were visible, and Karpathy’s autoresearch keeps its evaluation file off-limits to the agent.
  3. Paired, prospective comparisons. “The improved method is better” must mean that, on the same budget and the same tasks, run forward in time, it beat the fixed method by more than chance. Anytime-valid statistics let you watch the result accumulate without inflating the confidence.
  4. Conservative accounting. Anything ambiguous counts against the improver, never for it: a re-run, a change of implementation mid-comparison, an expired or interrupted trial. The system can lose credit through bad luck. It must never gain credit through a loophole.
  5. Authority where it matters. An improver may promote its own internal models automatically, but only audited and reversible. Anything customer-facing or in production stays behind human review.
  6. Blast-radius controls. The loop runs on spare capacity and yields to production. A restart never duplicates or loses work. A stale credential can never act for a replaced identity. Every automated action can be rolled back, the capability OpenAI now pairs with evaluations as the “ability to intervene, pause, or roll back”.
  7. A standing rule that the system may say “not yet”. An improver that has not earned the claim must not get it. A loop without this rule will, eventually, tell you what you want to hear.

None of this is exotic. It is the discipline that production systems and clinical trials already use, applied to a system that is grading its own homework. The AI control research program takes the same stance further, designing safeguards that hold even if the model is deliberately trying to subvert them.

RSI is a research problem

Controls make the loop trustworthy. They do not make it smart. The loop can only compound what its starting point is capable of learning, and the published systems say so plainly:

A learned improver, a model that predicts which change will work, faces a genuinely hard learning problem. It has to reason about differences between configurations, often numeric and on different scales. It has to generalize from a few hundred past outcomes to proposals it has never seen, and transfer what it learned on one task to the next. A base model that cannot represent that signal gives you a loop that runs, spends compute and learns nothing. A well-built gate will then correctly refuse to promote it.

The research questions
  • How capable must the base model be before learned improvement beats a good fixed method?
  • How should the improvement task be represented to the model? The way a problem is encoded can matter as much as the size of the model reading it.
  • How much outcome history is needed before the improver has anything to learn, and how do you keep it from overfitting to its own history?
  • When does improvement transfer across tasks, datasets and model families, and when does it only memorize?

This is why RSI needs both halves. Research produces a base model and a representation capable of learning to improve. Engineering produces the measurement that proves whether it did. Without the engineering, a research result about self-improvement cannot be checked. Without the research, the engineering faithfully measures a loop that never compounds.

The levels of recursive self-improvement

“Is this system self-improving?” has no yes-or-no answer. There is a ladder. Each rung changes what closes the loop, and raises what you must prove and control.

Six levels, from static to open-endedEach rung changes what closes the loop, and raises what you must prove and what you must control.
STEAV's framing. The lab thresholds above, such as fully automated AI R&D, sit at the top of this ladder.

Most production ML is L1 or L2. Automated retraining and hyperparameter search are valuable, and they are not recursive: the improver never improves.

Research systems touch L4 inside sandboxes. Agents that edit their own scaffolding exist. Their evaluators are fixed, their scope is narrow, and their documented failures are mostly evaluation gaming, which is exactly what the controls above exist to catch.

L3 is the missing middle, and the first rung where “self-improving” becomes a testable claim. At L3 the system learns from its own history how to improve, and the claim is measurable. Run the learned method and the fixed method side by side on the same budget, and let a statistic chosen in advance decide. A system that cannot pass that test at L3 has no business attempting L4.

Where CID comes in

CID is STEAV’s platform for owning the model layer: training, evaluating, deploying and serving models on your own infrastructure. We built its Model Studio so that the L3 loop, and the controls it needs, are part of the platform rather than a research script.

  • Every run is a full pipeline. Model Studio runs every training job through the same seven stages: data_prep, train, aggregate, eval, benchmark, validate and deploy. Nothing is skipped because a search launched it.
  • An outcome ledger. Every run’s configuration, dataset fingerprint, base model and measured outcome is recorded with its provenance, meaning what launched it and why, by the platform rather than by the improver. This is the history the improver learns from, and it has no way to edit it.
  • Measured search. Model Studio searches configurations under explicit budgets, queued for idle compute so that it yields to production work.
  • A learned improver. CID trains an advisor, a System One decision model, on pairs of past runs from its own ledger. It learns to predict which of two configurations will reach the better outcome, and is trained through the same Model Studio pipeline as any other model.
  • A gate before any advisor is used. A newly trained advisor is only put to work if it predicts the newest outcomes, which it never trained on, better than chance and better than the advisor it would replace.
  • Paired, prospective evidence. Advisor-ranked search runs side by side with plain search on the same budget. A scoreboard decides with an anytime-valid betting confidence sequence, not with a single lucky run, and uses validation data only. Until the test crosses its threshold, the scoreboard says “not yet”, and it is built to.
  • Conservative accounting throughout. An incomplete, cancelled, timed-out or administratively resumed comparison counts as a loss for the advisor. Comparisons are recorded append-only, and a verdict once reached is never retracted by what a screen happens to display.
  • Gates we tested before trusting. The statistical gates were checked in simulation against uninformative advisors and deliberate attempts to game the comparison. A gate that can be fooled in simulation is not shipped.
  • Authority and blast radius. The advisor may promote its own internal models automatically, audited and reversible. Customer and production models stay behind human review. A control-plane restart neither duplicates nor loses a run, and a runner’s credentials are checked against the exact registration they were issued for before it can take or report work.

What CID does not claim. CID has not demonstrated Level 3. The loop runs, the evidence accumulates, and the scoreboard will say when, and whether, the learned improver has earned the claim. We would rather publish that sentence than a demo.

What this means for you. The same loop runs on your models, on your data and on your infrastructure, with the controls on by default. Whether the models decide on fraud, triage clinical cases or power internal assistants, the question is the same: is the system that trains your models getting better at training them, and can you prove it? CID is built to answer that question honestly, including when the answer is no.

Recursive self-improvement will not arrive as one breakthrough. It will arrive as a loop that gets a little better at getting better, and it will be trustworthy only if every step of that loop can be checked. That is an engineering discipline and a research program at the same time. We are building both into CID, and we would like to build the next rung with you.

All news

Build the loop with the controls in.

Explore CIDTalk to STEAV