Upstage, SKT, and LG AI Research Advance to Final Three in 2nd Independence AI Model Review; Motif Technologies Eliminated
IT DAILY · · 1 views

✦ AI Summary
The Ministry of Science and ICT and NIPA held the second-stage evaluation of the independent AI foundation model project.
The three teams that passed the second evaluation were Upstage, SKT, and LG AI Research, while Motif Technologies was eliminated.
The evaluation was based on the combined scores of benchmarks, experts, and users, and the user evaluation had a decisive impact on the final ranking.
The second-stage evaluation was held for the
This project proceeds through a step-by-step compression competition to foster a national flagship independent model. At the outset, five teams were selected: Naver Cloud, NC AI, SKT, LG AI Research, and Upstage.
The government said at a briefing on the 18th that the three teams that passed the second-stage evaluation were Upstage, SKT, and LG AI Research. Motif Technologies was eliminated in the second-stage evaluation.
An initial stage evaluation was held earlier this year, and Naver Cloud and NC AI were eliminated for failing to meet the criteria for verifying technological independence and for the comprehensive evaluation. The field was then reduced to three teams after the eliminations, and the Ministry of Science and ICT conducted an additional call for proposals to fill the vacancy and strengthen competitiveness. As a result, Motif Technologies joined as a new elite team, reshaping the contest into a final-four lineup.
The second evaluation was conducted on a 100-point scale, consisting of 40 points for benchmark evaluation, 35 points for expert evaluation, and 25 points for user evaluation. The second evaluation ranked the teams by total score. The 40 points for benchmark evaluation were made up of 25 points for global AAII indicators and 15 points for NIA indicators, with the NIA indicators covering Korean language and safety.
The results showed that the participating companies were so closely matched technologically that it was difficult to determine clear superiority. Perceived effectiveness in actual service environments and usefulness in industrial settings emerged as the deciding factors in the final outcome. In the benchmark evaluation, the four teams averaged 22.5 points, and the gap between first and fourth place was 4.0 points. No team exceeded 30 points on the converted score scale, given the difficulty of the evaluation.
According to the Ministry of Science and ICT, the expert evaluation, in which 10 AI experts participated, was conducted with a weighting of 35 points, and all four teams demonstrated independence in the expert review. The overall average in the expert evaluation was 28.8 points, and the gap between first and fourth place was 2.4 points.
The result showed that the technological differentiation was relatively small. The expert evaluation alone was not enough to widen the technological gap among the four teams, and the score differences were limited.
By contrast, the factor that drove the actual ranking changes was the 25-point user evaluation. The user evaluation consisted of a 15-point assessment by 49 AI expert users and a 10-point assessment by 185 members of the general public. The overall average for the user evaluation was 17.6 points, and the gap between first and fourth place was 5.0 points.
The deciding factors were the service-experience evaluations by expert users and the ecosystem spillover effect in the expert evaluation, rather than the general public assessment. The service-experience assessment by expert users carried 15 points, and the ecosystem spillover effect in the expert evaluation also carried 15 points. The items with a large weighting for real-world usability and applicability determined the final outcome.
Three teams advanced to the next stage. Those three teams received high marks from the expert committee for technological completeness, real-world service application, and ecosystem diffusion demonstration elements.
The evaluation compared not only each company’s technical level but also real-world service integration, industrial application, reliability management, and ecosystem usability.
Upstage pushed for service integration with the portal Daum and the Timely platform, and also carried out demonstrations based on a FuriosaAI NPU. It also drew favorable reviews for highlighting efforts to reduce dependence on foreign hardware.
SKT moved to deploy its own model in large-scale commercial services, provided the A.X K1 model for defense use, and carried out demonstrations of agents specialized for the manufacturing industry, as well as legal and tax demonstrations. Through this, it proved its practical value in industrial settings and its applicability. LG AI Research, meanwhile, was positively evaluated for possessing differentiated Agentic AI technology, having established a model reliability and safety management system that includes Hallucination prevention, and presenting a strategy for collaboration with global international organizations.
Motif Technologies pursued in-house full-stack development of its architecture, tokenizer, and optimizer, and demonstrated technological excellence in global benchmarks (AAII). However, it showed relative weakness in real-world usability and ecosystem utilization, and was ultimately eliminated.
Upstage said it would join efforts to build an ecosystem foundation so that everyone in Korea can benefit from AI. It also described participating in the process of bringing Korean AI to the center of the global stage as a historic moment, and said it plans to advance user-friendly real-world service models through Daum and consortium participants while expanding supported languages into Asia. It added that it would also pursue the creation of a foundation for global expansion.
Lim Woo-hyeong, head of LG AI Research, said that it is difficult to narrow the technological gap with global big tech using small models alone. He said the company would continue challenging itself to achieve its goal of delivering equal or better performance at the same scale. He also explained that it optimized training for an ultra-large model over a limited period and accumulated technical assets and know-how in the process. He added that the company aims to achieve its target performance based on those accumulated assets and know-how.
SKT said it attaches significance to advancing to the next stage. It also said the result was thanks to the efforts of the companies and institutions participating in the consortium.
SKT said it plans to concentrate its capabilities on producing outcomes for shared use across Korea’s AI ecosystem, based on the results of expanding its own AI model into industry and everyday life.
Motif Technologies CEO Lim Jeong-hwan said the project was meaningful because it demonstrated the possibility of developing AI models at the level of Korea’s global frontier. He also thanked those who made the challenge possible and said the company would devote all its efforts to developing a global frontier model going forward.
A Q&A session was also held at the briefing that day. Questions were raised about the reason Motif Technologies was eliminated and whether the so-called benchmarkmaxing controversy had any effect.
In response, officials said the benchmarkmaxing controversy was not reflected in the evaluation at all, and that the official review by the evaluation agency confirmed there were no signs of systematic memorization or overfitting. They also said Motif’s elimination was due to differences in the point-allocation structure, and that although Motif demonstrated world-class technological capability in benchmarks, it was eliminated in the comprehensive evaluation because real-world usability, industrial applicability, and ecosystem contribution were relatively undervalued in the remaining areas, which accounted for 75 out of 100 points in total.
In the second evaluation, no company ranked first across all categories, and the top performer differed by item. No team exceeded 30 points on the converted benchmark score. The gap in expert evaluation scores was 2.4 points, and the maximum gap in user evaluation was 5.0 points.
The user evaluation had a decisive impact on the final ranking. The number of general public participants was 185, and the weighting for the general public assessment was 10 points. The AI expert user evaluation carried 15 points, and its practical differentiating power was confirmed.
After that, an objection process will take place before the launch of the third stage. The number of teams advancing to the third stage is 3, and the support period for the third stage is 6 months. The supporting infrastructure is at the level of a total of 1,000 NVIDIA B200 units, and the total infrastructure cost is about KRW 120 billion.
The government is discussing a revision of its support framework with relevant ministries and companies. The purpose of the revision is to narrow the gap with global big tech. The existing approach was a sequential elimination competition, while the revised direction is a model of concentrated national support. The revised framework will be called the 'new frontier AI model support framework,' and the government plans to make an official announcement after the budget is finalized.
Source: IT DAILY · Kwon Yeong-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=241048
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.