Recommended
The debate about the risks of artificial intelligence (AI) is no longer just about what machines can do, but about what will happen if they were capable of improving their own capabilities to the point where it becomes impossible for humans to supervise them.
This is known as 'recursive self-improvement', or RRI, through which an AI can help develop increasingly capable systems, which in turn contribute to building new generations of models.
The fear is that this cycle will accelerate to a point where humans lose control over its development.
What is 'recursive self-improvement'?
Until now, humans have directed the stages of AI development, defining its goals and evaluating results. However, current models can already take on some of those tasks.
The next leap is for an AI to close that loop and decide what to improve, what to design, and what experiments to conduct, evaluate its results, and use what it has learned to develop a more advanced version of itself.
Anthropic, creator of the Claude model, defines this scenario as the moment when an AI can "autonomously design a new generation of more capable models," although it assures that this point has not yet been reached and that recursive self-improvement is not inevitable.
AI is advancing "faster than we thought"
In recent days, Anthropic has once again warned about the speed at which its models are beginning to contribute to the development of AI systems, and company researchers have alerted to the risks that this entails.
One of them, Jacob Coxon, left the company after warning that the industry is moving too fast toward systems capable of improving themselves.
Anthropic maintains that this cycle has not yet been completely closed and that humans are still deciding which problems to investigate; however, it has acknowledged that Claude's advances can lead to recursive self-improvement, which, according to the company, is happening "faster than previously thought".
Expert Vincent Conitzer explained to CNBC that AI "is already capable of introducing some new ideas," such as proposing a different way to solve a problem or improve a program.
The problem is that "it is difficult to know when that capacity means drastically accelerating the development of increasingly advanced AI systems," he added.
The leap of the 'agents'
A fundamental piece of this process are the so-called AI agents: systems capable not only of responding to a request, but of using tools, executing code, performing tasks over long periods, and collaborating with other systems.
OpenAI's chief scientist, Jakub Pachocki, said last week that he has a "strong expectation" that if AI continues to advance along the path of RRI, systems in the coming years will make "increasingly large leaps in capability," "increasingly driving their own development."
The concern is no longer just theoretical: during a recent cybersecurity test, OpenAI agents managed to escape their evaluation environment and access systems on the Hugging Face platform.
OpenAI acknowledged that the models had gone further than expected to complete the test.
Human oversight
According to experts, the problem may arise if humans delegate increasingly complex tasks to AI and cease to understand what it is doing.
Researchers such as Margaret Mitchell warn that as AI agents acquire autonomy, maintaining human oversight will become increasingly difficult and that "systems may end up displacing people from decision-making."
Hence, tech companies are strengthening their safeguards: Pachocki has called for "caution" and to "slow down development if necessary"; Anthropic has tightened its controls after detecting potentially dangerous uses of Claude, including research that could facilitate the development of biological weapons, and Microsoft presented a code of conduct today to ensure its systems remain "under human control."
For his part, US President Donald Trump has rejected slowing down AI development for fear of losing an advantage over China.
[Publicidad]







