AI

Anthropic Researcher Quits, Warns of AI Self-Improvement “Runaway”

TECHWORLD ·

[Photo: ChatGPT]

✦ AI Summary

As AI draws attention for recursive self-improvement (RSI), the idea that AI speeds up AI research and development and that improved capabilities are then fed back into development is raising growing concerns.

The Wall Street Journal (WSJ) reported on the 9th, local time, that with the rapid improvement of AI research capabilities, questions about human development speed and the ability to control behavior have again been raised.

Anthropic researcher Jacob Coxon left the company after raising concerns about the race to develop AI systems that improve their own performance, and said both OpenAI and Anthropic are not acting responsibly.

As AI draws attention for recursive self-improvement (RSI), the idea that AI speeds up AI research and development and that improved capabilities are then fed back into the next round of AI development, concerns related to RSI are growing. With the rapid improvement of AI research capabilities, questions about the pace of human development and the ability to control behavior have again been raised, The Wall Street Journal (WSJ) reported on the 9th, local time.

Amid this trend, Anthropic researcher Jacob Coxon left the company after raising concerns about the race to develop AI systems that improve their own performance. Coxon’s research area is pretraining, which uses large-scale data to train new AI models, and he had worked at OpenAI and Anthropic for the past three years.

Through X, Coxon disclosed his resignation and expressed concerns about the AI industry, naming OpenAI and Anthropic. He said both companies were not acting responsibly and voiced concern that the race toward self-improving superintelligence was a gamble with lives at stake.

The article explains “self-improving superintelligence” as a concept meaning AI taking part in AI research and development. It is a structure in which improved research capability is reinvested into the development of stronger AI, and it is linked to recursive self-improvement. If this process is repeated, model improvement could accelerate.

At the same time, such repetition could reduce the time available to identify new capabilities and risks, and it could also reduce the time needed to establish safeguards. Coxon then argued that AI capabilities should not be underestimated. He said future AI may reach human-superior levels in hacking, human-superior levels in research, obtain real resources, and secure authority, and that it is difficult to rule out those possibilities.

He also pointed out that the AI industry is not competing to develop technology while unaware of the risks. Coxon said AI developers are in fact worried about the possibility of a threat to all of humanity before 2030. He stressed that this should not be dismissed as mere publicity or marketing.

Coxon gave different explanations for why OpenAI and Anthropic continue to compete. He believed many people inside OpenAI do not sufficiently take in the civilizational impact of AI. By contrast, he said Anthropic recognizes the risks but cannot slow down because it fears another company could develop more powerful AI first. In this regard, different assessments have emerged regarding the background to the OpenAI-Anthropic rivalry.

The key issue was then presented as whether safety technologies can advance at a speed comparable to AI performance gains. Coxon said there are limits to simultaneously accelerating superintelligence development while solving safety problems in a short period. He said it is necessary to sufficiently understand a model’s judgment and behavior before developing more powerful AI.

However, he said cybersecurity issues involving recent AI systems are serving as warning signs. Accordingly, Coxon said there is room for companies to jointly slow development, and he saw a growing chance that U.S. AI labs could agree to do so. He also said it is necessary to consider ways to limit improvements in model capabilities for a certain period in order to manage competition among countries.

According to the WSJ report, Coxon concluded that autonomous efforts by individual companies alone would make it difficult to manage the risks of high-performance AI. Coxon, who moved earlier this year to Anthropic, where safety research is emphasized, reached the conclusion that government intervention or an industry-wide response is necessary.

Coxon believed the point at which risks become real could come sooner than expected. He told the WSJ that many of the most aggressive AI development scenarios are moving toward becoming reality, and raised the possibility that by the end of 2027, the situation could reach a point that is difficult to control. He also expressed concern that AI with self-improvement capabilities could develop rapidly and that such AI could refuse human commands.

Similar concerns were also raised inside Anthropic. Evan Hubinger, head of AI alignment research, said he agreed with Coxon’s X post. Hubinger personally projected that there is more than a 10% chance AI could produce fatal consequences for all of humanity within the next 10 years.

Hubinger said Anthropic is conducting research to reduce AI risks, but pointed out that it has not yet secured a way to control superintelligence in line with human intent. He then framed the remaining challenge as whether technologies for understanding and overseeing AI can develop in step with the rapid improvement of model capabilities as AI takes on a larger role in research and development.

Against this backdrop, safety discussions are expanding from the risks of individual models to issues of development speed and control. In particular, the recursive self-improvement stage, in which AI feeds its own research capabilities back into development, is being cited as a conditional stage, and the speed at which safety systems for model capability evaluation and behavioral oversight can be secured is being presented as the key issue.

Source: TECHWORLD · Kim Seung-ki
Original: https://www.epnc.co.kr/news/articleView.html?idxno=406771

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.