This week's edition of The Tech Download newsletter highlights rising concerns about the safety of AI systems and their potential dangers to humanity. Evan Hubinger, an alignment lead at Anthropic, expressed on X that he believes there is over a 10% likelihood that AI could eradicate humanity within the next ten years. This alarming statement came after a colleague left the company due to apprehensions regarding safety. Following Hubinger's comments, both Anthropic and OpenAI researchers shared their own warnings, prompting a surge of discussion on social media.
Hubinger elaborated on his concerns in a response, particularly focusing on the concept of superintelligence emerging through recursive self-improvement. This phenomenon occurs when AI systems enhance their own design and efficiency, which could lead to an exponential growth in their capabilities. The central fear is that as AI takes charge of its model training process, the original human developers could end up losing control over these technologies.
Both OpenAI and Anthropic have recently noted that the pace of autonomous model enhancement is surpassing their expectations. Anthropic indicated on X in June that their internal assessments reveal that Claude, their AI model, is expediting AI advancements—potentially towards recursive self-improvement, where AI creates more advanced successors independently. They emphasized the need for serious attention to this rapid development. Though AI has not yet reached the stage of recursive self-improvement, it is already significantly speeding up the creation of new AI systems. In an August blog post, Anthropic reported that their engineers are now producing eight times more code in a quarter compared to the period between 2021-2025. Vincent Conitzer, a computer science professor at Carnegie Mellon University, pointed out that AI is now capable of introducing new ideas, making it challenging to predict when the acceleration in AI capabilities might truly begin.
On Saturday, OpenAI's Chief Scientist, Jakub Pachocki, voiced concerns about the unforeseen consequences of the swift advancement of machine intelligence. He warned in a blog post that if AI development continues at its current rate, we could soon witness significant leaps in capabilities that lead to increasingly autonomous development.
The discussion surrounding recursive self-improvement intensified this week, particularly after Jacob Coxon's resignation generated considerable buzz. OpenAI researcher Jasmine Wang articulated the gravity of the situation, stating, "It's hard to overstate how dangerous speeding towards recursive self-improvement is." Similarly, Anna Wang from Anthropic warned about the absence of viable scientific strategies to address the threats posed by this type of AI improvement.
In their blog post on RSI, Anthropic outlined three potential future scenarios. One scenario involves a stagnation in progress at the leading edge with widespread diffusion of AI capabilities, which they believe is unlikely. The second scenario envisions continued advancements in AI while remaining under human control, which Anthropic considers more probable. The third scenario raises concerns that AI systems may achieve full recursive self-improvement, leading to a significant reduction in human involvement in their evolution. How the alignment problem—ensuring that AI systems’ goals align with human values—is addressed in this future remains uncertain.




