Why Has the Math Olympiad Become AI’s Test Ground?
TECHWORLD ·
✦ AI Summary
OpenAI announced on September 8 that an internal model had solved a problem related to the existence of solutions to the Navier-Stokes equations.
At the International Mathematical Olympiad (IMO), a Google DeepMind model solved 5 of 6 problems in 2025 and scored 35 out of 42, while OpenAI also said it had unofficially scored 35 points.
The article explains how mathematics has emerged as a key test ground for training and evaluating AI reasoning ability, and how that trend is leading toward solving mathematical problems.
On September 8, OpenAI announced that an internal model had solved a problem related to the existence of solutions to the Navier-Stokes equations. The problem is one of the seven Millennium Prize Problems. I described it as a symbolic moment showing that AI had surpassed human reasoning ability in mathematics.
Behind this scene is the improvement in AI performance confirmed at the International Mathematical Olympiad (IMO). In July 2025, a Google DeepMind model officially took part in the IMO. It solved 5 of the 6 problems and scored 35 out of 42, a result equivalent to a real-world gold medal.
As of 2024, solving IMO problems required converting them into a formalized programming language, as well as extensive search for solutions over long periods. But in 2025, LLM-based AI solved problems in a different way. It handled problems written in natural language directly and solved them within the time limit.
As a result, LLM-based AI in 2025 reached results comparable to those of the world’s top high school students. Along with the fact that the Google DeepMind model posted a gold-medal-level score, this was presented as an example showing how quickly AI’s ability to solve math problems has improved.
AI companies such as Google and OpenAI are devoting massive resources to improving their AI’s math performance. I am a mathematician. I believe mathematics is interesting and worth deep study. But that alone is not enough to explain why companies are making such huge investments and taking on math challenges that appear far removed from immediate profits.
This article looks at why mathematics has emerged as a key test ground for training and evaluating AI reasoning ability. It also examines where that technical challenge stands today.
To explain this, it is necessary to look at how LLMs are trained. LLMs are trained on large datasets. In this process, the prediction unit for an LLM is more precisely a token rather than a word. LLMs learn the probability of the next token appearing based on the sequence of previous tokens.
For example, if the context is 'I had kimchi for lunch today,' possible continuations might include 'jjigae,' 'jeon,' 'fried rice,' or 'mandu hot pot.' LLMs learn the probability that such candidates will follow. They then generate sentences based on those learned probabilities. Because of this structural feature, LLMs are well suited to generating statistically plausible sentences.
The validity of reasoning is not guaranteed by plausibility alone. The key criterion is whether a conclusion is logically derived from prior facts.
The training objective of an LLM is next-token prediction. A training objective centered on next-token prediction is not inherently aimed at logical validity.
Early models such as GPT, which emerged around 2021 to 2022, showed remarkable results in other fields. But their mathematical reasoning ability was only at the level of competing with elementary and middle school students. The reason early LLMs were weak at reasoning was the lack of separate training and verification mechanisms.
Early LLMs lacked separate training and verification mechanisms. For AI based on LLMs to perform high-level tasks with humans, reasoning ability is important.
In school education, mathematics has long been used to cultivate and assess logical thinking. The reason AI companies have focused on mathematics lies in its function of training and evaluating logical thinking.
Solving math problems involves analyzing conditions, connecting multiple logical steps, reaching conclusions, and rechecking results. Math problems also have the characteristic of having clear answers and allowing objective evaluation of whether the process was carried out properly.
Because of these characteristics, mathematics is seen in AI as a controlled environment suitable for training and testing reasoning ability. Global AI companies’ math challenges also began with these characteristics of mathematics. The goals they set need to be very difficult yet achievable, and the significance of achieving them must be intuitive.
The International Mathematical Olympiad (IMO) is a test that meets these conditions. The IMO is a stage where the world’s best high school students compete. The exam consists of 6 problems and is held over 9 hours across 2 days. Each problem is worth 7 points, for a total of 42 points.
Among such IMO formats, one problem from the 2025 IMO will be introduced for the purpose of giving readers a sense of the exam’s difficulty.
Problem 6 of the 2025 IMO asks for the minimum number of rectangular tiles needed when placing non-overlapping rectangular tiles whose sides lie on grid lines on a 2025×2025 grid made of unit squares. The condition is that exactly 1 unit square in each row and exactly 1 unit square in each column must remain uncovered.
This problem is regarded as something that most people face as a major obstacle even after thinking about it briefly. Compared with Question 22 in the math section of the College Scholastic Ability Test, which people can at least try several approaches to even if they cannot solve it, this problem is seen as so difficult that it is hard even to think of a starting point.
The decisive property used in the solution is the fact that 2025 is 45×45. In other words, 2025 is a perfect square.
It is said that if one does not quickly notice this fact, the chance of solving the problem is almost nonexistent. In the end, the key to this problem is recognizing that 2025 is a perfect square, 45×45, in the process of handling the conditions of the 2025×2025 grid.
Most college and graduate school math exams require shallower reasoning than the IMO. The IMO demands reasoning at a very high level. Therefore, if a model can perform at above-human level in this test, claims that LLMs cannot do math or reasoning may be weakened. IMO results may become a means of proving AI’s mathematical reasoning ability.
IMO problems are safe from dataset contamination. The IMO consists each year of new types of problems creatively made by humans. Similar IMO problems do not exist in past exam papers or on the internet. For this reason, the possibility of solving IMO problems through LLM data training alone is slim.
This judgment is based on cases that do not involve direct high-level reasoning. Under that premise, the IMO can confirm whether AI is truly reasoning. In other words, it is suitable for determining whether there was actual reasoning, not mere memorization.
In addition, IMO results can be clearly evaluated by score. For these reasons, the IMO has become an important target for AI companies. The combination of high difficulty and clear scoring is presented as the background for this.
AI companies set IMO gold-medal-level performance as the first major target for math reasoning. The gold-medal benchmark corresponds to about 1 in 12 contestants. But at first, the outlook for reaching that goal was uncertain, and in the 2024 IMO, not a single AI company announced that it had solved problems with an LLM, showing how low the starting point was.
Behind this was the fact that there were multiple bottlenecks in applying LLMs to math Olympiads. Because of the nature of math Olympiads, there was a shortage of high-quality training data, while internet data lacked completeness in solutions and did not meet error-free requirements. Even after securing data, methods for post-training and reinforcement learning were not sufficiently known.
Then in 2024 to 2025, efforts to secure data continued, new reinforcement learning techniques emerged, and verifiers also advanced. As these efforts to secure data, reinforcement learning techniques, and verifier development came together, the situation began to change.
In 2025, a Google DeepMind model officially competed in the IMO and scored 35 points, a gold-medal-level result. OpenAI also announced that it unofficially recorded the same score of 35 points in the IMO and released its solution.
Such reasoning ability cannot be acquired simply by scaling up model size; it requires the ability to autonomously search for problem-solving strategies, the ability to connect logic through multiple steps, and elements that make it possible to verify derived results again. This requires separate training and design.
Although a year has passed since the first appearance of gold-medal-level IMO performance, few models have implemented it stably so far. Excluding models developed by U.S. and Chinese companies, the only case that has publicly announced meeting the IMO gold-medal benchmark is SKT's proprietary AI foundation model A.X K2.
A.X K2 scored 29 out of 42 in the 2026 IMO problem evaluation, meeting the 2026 gold-medal benchmark. It also scored 35 out of 42 on the 2025 IMO past problems, showing gold-medal-level performance, and recorded a joint top score with 97.1% accuracy in 'MathArena AIME 2026,' which evaluates advanced mathematical reasoning ability.
A.X K2 recorded consecutive high scores across math evaluations of different types and difficulty levels. Through this, A.X K2 demonstrated mathematical reasoning ability comparable to the world’s top models.
At the same time, attention is being drawn to the reasoning ability of LLMs’ vast knowledge systems. This ability is presented as serving as a starting point for tackling real mathematical research.
In this context, cases of achievements in mathematical AI are presented. On 2026. 5. 20., a counterexample to a conjecture related to Paul Erd?s’s unit-distance problem in the plane was constructed.
Then on 2026. 8. 1., a non-sofic group was constructed. Also, on 2026. 9. 8., a solution related to the existence of solutions to the Navier-Stokes equations was disclosed. The Navier-Stokes equations are a Millennium Prize Problem.
However, there is a large gap between solving IMO problems and solving actual mathematical problems. It is not common for an 17-year-old IMO gold medalist to go on to solve a real mathematical problem within the next 10 years.
Real problem solving requires not only reasoning ability but also broad knowledge and insight. Accordingly, apart from the examples listed above, it is also explained that solving IMO-level problems is different from solving real mathematical problems.
Solving mathematical problems requires 'long thinking' and long-duration reasoning. Humans take an average of 1.5 hours to solve IMO problems, but in many cases even 1.5 years is not enough to solve a single difficult problem. In this sense, the Math Olympiad can be compared to solving one life-or-death problem, while solving a hard problem can be likened to a match against Shin Jin-seo, 9-dan.
Because of this gap, AI companies have responded with better foundation models, agent development, architecture improvements, and large-scale reinforcement learning. The goal is to strengthen LLM math capabilities and build agents capable of long thinking.
With these efforts, it took more than 2 years of research and development for LLMs to reach IMO gold-medal-level performance. However, once that achievement was in hand, there was no need for a long period before beginning to tackle real mathematical problems.
The cases that unfolded in the year after the 2025 IMO showed that AI’s mathematical ability is rapidly advancing. AI demonstrated gold-medal-level performance at the IMO and then went on to solve mathematical problems that humans had not been able to solve.
On September 8, 2026, OpenAI announced that an internal model had reached the discovery of a proof for the Navier-Stokes problem. The Navier-Stokes problem is considered one of humanity’s unsolved Millennium Prize Problems.
In response to this trend, the mathematics field saw assessments that an so-called AlphaGo moment had arrived, along with the arrival of superhuman capabilities. This is closely tied to the perception that AI has moved beyond merely matching the highest human level in mathematics.
However, some areas where AI still falls short of humans remain. Even in those areas, though, there is a view that surpassing humans is only a matter of time.
The reason AI companies are pouring huge resources into mathematics is not because solving mathematical problems is the ultimate goal in itself. They are using mathematics to identify the limits of AI’s logical thinking, and ultimately aim to realize reasoning ability that surpasses human level.
Source: TECHWORLD · Seo In-seok, Professor, Department of Mathematics, Seoul National University
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407888
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.