When AI Builds AI: Anthropic’s Recursive Self-Improvement Evidence and Limits

AI already writes code and analyzes experiments used to build the next generation of AI. Anthropic’s recursive self-improvement study does not declare a completed intelligence explosion. It examines a narrower but consequential loop: AI makes an AI lab more productive, the lab builds a stronger model, and that model further accelerates the next development cycle.
Anthropic surveyed 130 research staff and reports a median estimate of roughly four times as much output with Mythos Preview in March 2026. It also says code shipped per engineer per quarter rose about eightfold from 2021 through 2025, while an internal Claude Code success measure on open-ended work increased from roughly 26% to 91%. The numbers are striking, but they are largely internal, and some Claude work was evaluated by Claude.
Three takeaways
- Today’s recursive self-improvement is AI accelerating code, experiments and evaluation inside a human-run organization—not a system independently controlling every goal and resource.
- Anthropic’s fourfold-output and eightfold-code measures show a strong internal trend, but do not isolate model effects from organizational growth, tooling changes and evaluator bias.
- Bottlenecks are moving from producing candidate work to verification, research direction, compute allocation, safety judgment and real-world experiments.
Anthropic’s reported indicators

What recursive self-improvement means
In the broad sense, it is a feedback loop in which AI contributes to AI research. A model finds a training-code bug, filters data, proposes an experiment or evaluates another model. If those contributions help produce a stronger system, the next research cycle receives a stronger assistant.
The stronger version is an AI system that independently controls goals, training data, architecture, compute and deployment while repeatedly improving its own capabilities. Anthropic does not claim that has happened. Humans still choose the research questions, allocate expensive training resources, assess hazards and authorize release.
The evidence therefore supports “automation inside AI R&D is rising quickly,” not “AI now evolves itself without people.” Keeping that distinction prevents productivity research from turning into an unsupported story about a conscious system taking over its own development.

How Anthropic arrived at the fourfold estimate
Anthropic asked research staff how much work they produced with Mythos Preview relative to working without it. The median response was about four times as much. The measure captures perceived gains across analysis, coding, experiment setup and documentation.
Self-reporting is not a controlled time study. Heavy AI users may be more likely to respond, and “output” can mean different things for a bug fix and a new research hypothesis. If AI multiplies attempted experiments but the organization cannot validate them, scientific productivity does not rise at the same rate.
The eightfold code-shipment increase also cannot be assigned entirely to Claude. Between 2021 and 2025, Anthropic changed in headcount, internal platforms, testing automation and product scope. The result is evidence of direction and magnitude, not a formula that one engineer plus Claude always equals eight engineers.
What the 26% to 91% success measure shows
Anthropic says Claude Code’s success on internal open-ended tasks rose from about 26% to 91%. The claim points to improvement on ambiguous repository work that requires persistence, rather than only constrained benchmark questions.
The task distribution and grading remain internal, and some evaluation uses Claude to judge Claude-produced work. A model family can prefer its own style or miss the same subtle errors. Independent validation would require task definitions, human review, failure distributions and evaluation by systems that do not share the same training biases.
A 91% internal task score also does not mean Claude autonomously performs 91% of AI research. Completing a repository task is different from choosing a valuable scientific question, rejecting a fashionable but false premise or deciding whether a capability should be deployed.
Where the human bottleneck moves
Faster coding makes the next constraint visible: GPUs for experiments, reliable evaluation data, expert interpretation, physical-lab throughput and safety approval. A model can produce thousands of hypotheses in a day while making it harder—not easier—for people to identify which results deserve trust.

Human work shifts from manually producing every artifact toward setting research goals and evaluation criteria, demanding independent checks and deciding acceptable failure cost. That does not eliminate people. It lets fewer decisions move much larger computational resources, increasing the importance of authority and accountability.
Potential benefits for science and medicine
Models that summarize literature, write experiment code and compare results can let scientists discard weak hypotheses faster and deepen promising ones. Large search spaces in protein design, drug discovery, materials and mathematics are natural beneficiaries.
Biology and medicine are not validated by code tests alone. Laboratory replication, clinical safety, population diversity and regulatory review remain necessary. As AI raises development speed, organizations need better provenance, preservation of negative results and independent replication—not less.
Why safety becomes harder
If development accelerates faster than alignment research and evaluation design, the organization may build the next model before it understands the current one. When AI also writes safety filters and evaluation code, a shared model-family error can enter both the system and its tests. Deliberate deception is not required; correlated assumptions across an automated organization are enough to create risk.
Safety for recursive development is therefore broader than an answer filter. It covers training environments, data, evaluator independence, sandboxes, outside audits, release authority and incident reporting. AI-written research infrastructure needs cross-model and human review, while consequential holdout evaluations should remain separated from the model development loop.
What U.S. companies should do
Companies should map their actual bottleneck before asking how many developers AI replaces. Coding may not be the slow step; test environments, security approval, data access or product review may dominate. Multiplying generated work can simply create a larger review backlog.
A responsible pilot has three parts. First, measure baseline success and human review time on sanitized internal tasks. Second, require automated tests, static analysis and dependency checks for AI-written code. Third, separate authority to develop, evaluate and approve a consequential model or agent.
Proprietary research, personal data and export-controlled information require explicit rules for outside API transmission and log retention. The business case depends less on a vendor’s fourfold survey result than on the cost of one undetected error in the company’s own environment.
Frequently asked questions
Is AI already building the next AI by itself?
It is contributing materially, but the loop is not fully autonomous. AI accelerates code, experiments and evaluation while humans control direction, compute, verification and release.
Should the 4x and 8x figures be taken literally?
They are strong internal trend signals, not independent causal estimates. Self-reporting, organizational change and evaluator bias must be included in the interpretation.
What is the central risk?
The development loop could move faster than human oversight and safety evaluation, while common errors propagate across models, research code and the evaluation system.
Bottom line
Recursive self-improvement is both a future concept and a present workplace change. Fully autonomous AI development has not arrived, but AI already helps create the code, experiments and evaluations that produce stronger successors.
Anthropic’s 4x, 8x and 91% measures show speed, not an independently verified universal law. The durable advantage may belong less to organizations that generate the most work than to those that can verify what deserves to be trusted and released.
Sources and rights notice
Anthropic, Claude, OpenAI and related names and marks belong to their respective owners. This independent editorial analysis is not sponsored, endorsed or approved by either company. It distinguishes internal company measurements from independently established findings. Its images are editorial concepts, not actual research facilities or product interfaces.



