AI Race Over Math Breakthroughs Draws Scrutiny From Academia
TECHWORLD ·
✦ AI Summary
AI companies are using examples of solving difficult mathematical problems as a benchmark for model-performance competition, while academia is placing greater emphasis on verification and knowledge accumulation and expressing caution.
OpenAI said on the 8th that it had proposed a solution to the Navier-Stokes equation using a multi-agent AI system, and said it used about 10,000 AI agents, 88 hours, 17 hours, 2.7 million messages, and about 130 billion output tokens.
Twenty-five Fields Medalists, including Junho Huh, Terence Tao, and Peter Scholze, did not deny AI's potential use in mathematics, but warned against using competition to solve famous hard problems as a benchmark. Anthropic and Microsoft later also proposed calls for slowing down and principles for safety and verification.
As AI companies use examples of solving major mathematical problems as a metric for model-performance competition, academic caution around the practice is also widening. In particular, concern is growing over the view that verification is needed before speed.
AI has reached a stage where it can propose solutions to problems that human mathematicians have long been unable to solve, but academic verification is needed before those results are recognized as genuine mathematical achievements. As a result, companies are highlighting solution speed and agent scale as proof of technical strength, while the math community is being cast as the body responsible for verification and the accumulation of knowledge.
OpenAI announced on the 8th that it had proposed a solution to the Navier-Stokes equation, one of the Millennium Problems, using a multi-agent AI system. OpenAI said the method involved about 10,000 AI agents exploring different approaches in parallel.
It also said it took about 88 hours from the deployment of the first agent to reaching a solution.
It took 17 hours to add Lean-based formalization and verification using GPT-6 Astra. During the problem-solving process, 2.7 million messages were exchanged between agents, and about 130 billion output tokens were used.
These numbers are presented as items AI companies use when highlighting technical capability. Such figures include the number of agents, the amount of computation, token usage, and the time required to solve a problem. From the perspective of AI companies, the speed at which a human-unsolved problem is resolved becomes a measure of the model's reasoning ability, and companies can use the fact of arriving at a solution, the processing speed, and the resources deployed as ways to emphasize technical strength.
However, even if AI proposes a solution, that does not mean the mathematical result is immediately recognized. In particular, the Millennium Problems require scrutiny for proof errors, hidden assumptions, and consistency with existing theorems. To have a Millennium Problem result recognized, long-term review by the academic community is assumed, and a separate verification process is required before any actual achievement can be acknowledged.
Even before verification is complete, the fact that a difficult problem has been solved can be used in AI companies' performance competition. Companies can compete by emphasizing the fact of the solution, the speed, and the resources deployed, but the body responsible for the process of recognizing mathematical achievements is the academic community.
Twenty-five Fields Medalists, including Princeton University professor Junho Huh, UCLA professor Terence Tao, and Max Planck Institute for Mathematics president Peter Scholze, issued a joint statement on the 11th. The statement was titled "A Severe Misalignment of AI in Mathematics."
The mathematicians who joined the joint statement said their intent was not to deny AI's mathematical capabilities themselves. Rather than rejecting AI's potential use in mathematics, the statement focused on differences in evaluation standards surrounding the solving of difficult problems.
They warned against the way AI companies use the competition to solve famous hard problems as a benchmark for model performance. The statement pointed out that AI companies and the mathematics community have different goals when it comes to solving hard problems.
From the company's perspective, solution speed and the use of model-performance metrics are treated as important. By contrast, from the mathematics community's perspective, greater weight is placed on understanding new concepts and structures and on the process of accumulating knowledge.
The statement defined problem solving as a tool, or surrogate means, for achieving conceptual understanding and insight. It also explained that AI-generated answers and the accumulation of new concepts and research methods are separate matters.
On the 12th, the day after the math community's warning, calls for slowing down also emerged in the AI industry. Dario Amodei is the CEO of Anthropic.
In a public post that day, Amodei argued that the pace of improvement in frontier AI models should slow. He pointed to AI's involvement in software development and AI research itself.
Amodei described recursive self-improvement (RSI), in which improved AI capabilities accelerate the development of the next generation of models again. He identified recursive self-improvement (RSI) as a risk factor. He also said the pace of AI progress could outstrip humans' ability to conduct safety research and verification.
Sam Altman, CEO of OpenAI, and Elon Musk, X AI CEO of SpaceX, voiced support for the publication of Amodei's proposal. Demis Hassabis, CEO of Google DeepMind, also expressed a similar sense of concern about the safety and development speed of frontier AI.
Such related discussions are expanding beyond declarations and into actual model-operations principles. Microsoft unveiled a draft of its "Humanist AI Code of Conduct" for Microsoft AI (MAI) models on the 14th.
The draft of the "Humanist AI Code of Conduct" includes a ban on making weapons, a ban on procuring hazardous substances, and a ban on supporting offensive cyber operations. It also includes principles prohibiting models from setting their own goals, concealing improper behavior, using "neuralese," an internal representation that humans find difficult to interpret, in thought processes or communication with other AI, and requiring reasoning and communication in forms humans can understand.
Within the companies leading frontier AI development, challenges are also coming to the fore. This is because the gap between the pace of model-performance improvement and the ability to verify and control those models is widening. As a result, the issue is shifting away from whether AI development should be halted and toward aligning the speed of performance competition with the speed of verification systems.
An AI industry source said the point of the recent calls for slowing down is not to argue for stopping AI development. The source explained that safety and verification systems need to be strengthened at the same pace as performance improvements. The source also said competition among frontier models is likely to continue.
AI has begun directly generating proofs and participating in scientific discoveries, moving beyond research assistance. This expanded role for AI is further increasing calls to keep the pace of performance competition aligned with the pace of safety and verification systems. Going forward, competition is expected to hinge on the ability to quickly solve harder problems and on the ability to verify and control results in a reliable way.
Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=406863
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.