Skip to main content
AI researchers discuss concerns as advanced AI models display increasing capabilities, raising ethical questions.

Editorial illustration for Top AI Lab Researchers' Warnings Gain Credence as AI Achievements Mount

AI Self-Improvement Loop: Lab Researchers Sound Alarm

Top AI Lab Researchers' Warnings Gain Credence as AI Achievements Mount

4 min read

Severin Field spent late summer 2025 interviewing 25 researchers at OpenAI, Anthropic, Google DeepMind, Meta, and several US universities about a narrow but consequential question: can an AI system build a better version of itself, then keep doing that in a loop? Twenty of the 25 people he talked to rated this kind of automated AI research among the most severe and urgent risks the field faces. Field wrote up the results for his newsletter, The Attack Surface, and the timing matters. Some of the milestones his interviewees named as markers of recursive self-improvement have already been reached in the months since he conducted the interviews.

The researchers Field spoke with kept returning to one yardstick: METR's Task Horizon benchmark, which tracks how long a task an AI agent can complete on its own before it needs help. That number has been doubling roughly every six months since 2019, and by some estimates the doubling time has shrunk to four months since 2024. Nobody Field interviewed disputed that AI systems are already improving AI systems in some sense. The argument was over whether those gains compound on their own, without a human in the loop, and what would have to break first for that to happen.

IAPS fellow Severin Field interviewed 25 researchers from OpenAI, Anthropic, Google Deepmind, Meta, and US universities about recursive self-improvement. In a new blog post, he takes stock. Several of the milestones those researchers named have already been hit.

Why this matters

Field's interviews weren't meant as forecasts anyone had to sweat over on a quarterly basis, but the speed at which the named milestones fell should recalibrate how we read lab safety statements. When gold-medal Math Olympiad results, a peer-reviewed AI-generated paper, and Claude writing 80 percent of Anthropic's own code all land within roughly a year of researchers naming them as checkpoints, the gap between "warning" and "observed fact" is shrinking fast. For developers and founders building on top of these systems, that's a signal to stop treating recursive self-improvement as a distant governance problem and start asking concrete questions about what their tools are already doing unsupervised.

For researchers, it's a prompt to publish sharper, checkable predictions rather than broad concern, since Field's value here came from specificity, not alarm. The labs making these systems are also the ones grading their own homework, so outside tracking like this matters more, not less, as the milestones keep landing on schedule.

Common Questions Answered

What specific concern did Severin Field investigate by interviewing 25 AI researchers?

Severin Field interviewed 25 researchers from leading AI labs including OpenAI, Anthropic, Google DeepMind, and Meta to explore whether an AI system could build a better version of itself and repeat this process in a loop, known as recursive self-improvement. Twenty of the twenty-five researchers rated this form of automated AI research as among the most severe and urgent risks facing the field.

What are some of the AI milestones that researchers predicted and have already been achieved?

Researchers named several checkpoints including gold-medal Math Olympiad results, peer-reviewed AI-generated papers, and AI systems writing significant portions of code for their own developers. According to the article, all of these milestones have already been hit within roughly a year of researchers naming them, demonstrating the rapid pace of AI capability advancement.

How does the speed of achieving predicted AI milestones affect our interpretation of lab safety statements?

The rapid achievement of named milestones suggests that the gap between researchers' warnings and observed facts is shrinking significantly faster than anticipated. This accelerated timeline should recalibrate how developers and the public read official lab safety statements, as the predictions are being validated much more quickly than expected.

Which AI systems demonstrated the capability to write code for their own developers?

Claude, Anthropic's AI system, demonstrated the ability to write approximately 80 percent of Anthropic's own code, representing a significant milestone in automated AI research capabilities. This achievement occurred within the timeframe that researchers had identified as a potential checkpoint for concerning AI capabilities.

LIVE13:48Anthropic Will Invisibly Watermark All New AI Models’ Text Outputs