Artificial intelligence may be taking another significant step toward becoming capable of improving itself.
A new study from Anthropic has found early evidence that advanced AI systems can help train and improve other AI models, including reducing certain undesirable behaviors without damaging their overall performance.
The findings have attracted attention because the ability of AI to meaningfully improve AI is often considered an important step toward recursive self-improvement — a concept frequently associated with the long-term path toward artificial general intelligence, or AGI.
AGI generally refers to a hypothetical form of AI capable of performing or surpassing humans across most economically valuable tasks.
While Anthropic’s research does not claim that AGI has been achieved, it suggests that automated AI researchers could increasingly take on parts of the work currently performed by human AI scientists.
Anthropic tests AI systems as automated researchers
The research was detailed in a paper titled Automated Researchers Can Reliably Mitigate Alignment Failures, published on August 28.
The study examined whether automated AI systems could identify ways to improve the safety and alignment of another AI model.
Researchers built what they called automated alignment researchers, or AARs, powered by Claude Opus 4.8.
Each automated researcher was assigned a specific alignment problem and asked to work on improving the target model’s behavior.
The AI systems were designed to search relevant research literature, propose a possible training method and then use that method to train the target model.
The process was repeated over several iterations.
Effective training approaches were retained, while methods that did not produce useful results were discarded.
According to the study, this approach allowed the automated systems to work rapidly and test multiple ideas at a scale that could become increasingly difficult for human researchers to match.
AI improved performance across all 10 tests
Anthropic evaluated the automated researchers using 10 benchmarks designed to test specific forms of misaligned AI behavior.
The study found that the AAR systems were able to improve performance on every benchmark tested without causing an overall decline in the model’s general capabilities.
That result is particularly important in AI safety research.
Improving one aspect of an AI system can sometimes create unintended trade-offs elsewhere. Anthropic’s experiment found that the automated training approaches could improve targeted behaviors while maintaining overall performance.
The researchers described the findings as early evidence that automated post-training for AI alignment could become practical in the near future.
However, the results remain limited to the specific benchmarks and experimental environment used in the study.
Automated AI researchers beat human proposals in the experiment
Anthropic also compared the performance of its automated systems with proposals made by experienced human AI researchers.
According to the study, the best automated method performed better, on average, than the approaches proposed by experienced human researchers within a six-hour period.
The potential cost difference was also significant.
The researchers said the automated systems cost roughly $4 per hour in API inference, compared with approximately $150 per hour for human AI researchers.
If automated systems become capable of reliably conducting more AI research tasks, the economic implications could be substantial.
Research that currently requires large teams of highly specialized scientists could potentially be accelerated through AI systems capable of reviewing research, generating hypotheses, testing training methods and evaluating results.
That possibility is one reason the concept of AI improving AI has become such an important area of research.
Why recursive self-improvement matters
Recursive self-improvement refers to the idea of an AI system contributing to the development of increasingly capable AI systems.
In its most ambitious form, an AI could help researchers build a better model, which could then help build an even more capable system.
Such a feedback loop could potentially accelerate technological progress.
The Anthropic study does not demonstrate unrestricted recursive self-improvement.
Instead, it shows that AI can successfully automate a specific part of the AI research process: finding and testing methods to address alignment failures.
Even so, the results could represent an early step toward a future in which AI systems become increasingly involved in improving other AI systems.
For now, human researchers remain responsible for setting goals, designing evaluation systems and determining whether improvements are genuinely useful and safe.
The research arrives amid growing AGI competition
Anthropic’s findings come as competition among major AI companies increasingly centers on the possibility of reaching AGI.
OpenAI has reportedly made progress toward developing increasingly autonomous and capable systems, with CEO Sam Altman previously expressing optimism about the pace of progress.
OpenAI’s own definition describes AGI as highly autonomous systems capable of outperforming humans at most economically valuable work.
However, the path toward AGI remains heavily debated.
Some researchers believe increasingly powerful models could eventually reach that level through continued improvements in computing power, training techniques and AI-assisted research.
Others argue that simply scaling current AI systems will not be enough and that major scientific breakthroughs may still be required.
The debate reflects a larger uncertainty surrounding the future of AI: models are becoming more capable at assisting with research, but it remains unclear how close those capabilities are to truly general intelligence.
AI researchers are not becoming obsolete yet
Anthropic’s paper also highlighted several important limitations.
The automated researchers can only work effectively when the benchmarks used to evaluate them accurately represent the real-world goals researchers want to achieve.
In other words, an AI system may become extremely good at improving a score on a benchmark without necessarily solving the broader problem the benchmark was designed to represent.
Creating, maintaining and improving those benchmarks remains a major challenge.
Human researchers are also still needed to expand the body of scientific literature and knowledge that automated systems can draw upon.
For that reason, the research does not suggest that AI scientists are about to be replaced.
Instead, it points toward a future in which AI could increasingly function as a powerful research assistant — capable of testing ideas faster, evaluating more possibilities and reducing the time required for certain technical tasks.
A small step with potentially enormous implications
Anthropic’s experiment was focused on a narrow but important problem: improving AI alignment through automated research and training.
Still, the broader implications are difficult to ignore.
If AI systems become increasingly capable of conducting meaningful research into how to improve other AI systems, the speed of AI development could accelerate significantly.
The key question will be whether researchers can maintain reliable human oversight while automated systems take on a larger role in the development process.
For now, Anthropic’s findings are best viewed as an early demonstration rather than proof that self-improving AI has arrived.
But the study offers another sign of how quickly the relationship between AI and scientific research is changing.
The machines are not yet independently creating the next generation of artificial intelligence.
They are, however, beginning to show that they can help researchers figure out how to build it.
