AI

Motif Lodges Formal Objection to 'Independent AI Foundation Model' Project, Says Evaluation Criteria and Scores Must Be Disclosed

IT DAILY · · 1 views

독자 AI 파운데이션 모델 프로젝트 1차 발표회. (사진: 과기정통부)

✦ AI Summary

Four consortiums — Motif, Upstage, SK Telecom, and LG AI Research — took part in the second-stage evaluation of the 'Independent AI Foundation Model' project, and Motif was eliminated with the lowest score.

Motif Technologies formally objected to the selection results on the 27th, demanding disclosure of its AAII score, expert and user evaluations, and the evaluation criteria and scores.

The Ministry of Science and ICT said Motif 3 had strong benchmark scores but lacked real-world inference performance, usability, and utility, while Motif argued that there were problems with the evaluation and review process.

In the second-stage evaluation for the 'Independent AI Foundation Model' project, 4 consortiums competed. The participating consortiums were Motif, Upstage, SK Telecom, and LG AI Research. In the final results released on the 18th, Motif received the lowest score and was eliminated.

Accordingly, Motif Technologies filed a formal objection on the 27th to the selection results of the second stage of the 'Independent AI Foundation Model' project.

Motif said it recorded 47 points in the Intelligence Index (AAII) benchmark. Based on AAII, Motif ranked first among the participating companies, with Upstage at 37 points, SKT at 35 points, and LG at 31 points.

Motif questioned how it could receive an unsatisfactory rating in expert and user evaluations despite ranking first on the global overall performance indicator. It also demanded disclosure of the evaluation criteria and scores.

The Ministry of Science and ICT said that although Motif's 'Motif 3' had an advantage in benchmark scores, its real-world inference performance, usability, and utility were lacking. In response, Motif argued that the value of a foundation model lies not in simple chatbot usability, but in base performance that can expand across diverse domains.

Motif pointed out that the Ministry of Science and ICT emphasized the achievement of 'Independent AI Foundation Model 47 points, top 3 globally' in its presidential report at a Cabinet meeting on August 25, yet excluded the very team that delivered those scores, calling it a policy contradiction.

Motif explained that benchmarks are standardized, objective evaluation criteria that drive the development of the computer science industry and technology. It also said AAII is a credible composite indicator that aggregates 9 representative benchmarks.

Motif said that while AAII and similar benchmarks are not absolute standards, there is a significant performance gap between 47 points and 31 points. It added that the expert evaluation results went in the opposite direction of that performance gap, and urged the government to provide specific explanations for the reasons and background behind the expert evaluation results.

Motif raised issues with both the evaluation results and the review process. It asked for the basis on which LG ranked first in the expert evaluation. LG received 31 points even after expanding the model size by more than 2 times.

Motif also sought an explanation for the background to the inclusion of non-technical questions, such as foreign capital ownership ratios, during the review process. Such non-technical questions, including foreign capital ownership ratios, were included in the review process. Motif said that, amid continuing criticism over usability, it sees no sustainability in usability for a model whose base performance is lacking.

Motif judged that the performance of domestic models is currently too low for complacency. With the global frontier model scoring 63 points on AAII and Motif scoring 47 points, Motif stood at 75% of the frontier's level. The lowest-ranked model scored 31 points, or 49% of the frontier's level. The original project target was to secure performance at 95% or more of the frontier level, and Motif said the current performance remains far from that goal.

Motif also rejected arguments from some within the government that large companies should handle frontier models while startups focus on specialized models. Motif countered that startups are the ones leading the global market, and said the AAII benchmarks of 2 startup companies also outperformed those of 2 large companies in the second-stage evaluation.

Motif grouped together the items it believed required disclosure and explanation, presenting 4 points of objection related to the evaluation. It demanded the disclosure of the detailed scores and criteria for each evaluation item and each company, the basis for assigning benchmark points, the detailed criteria and results of the expert evaluation, and the methodology and detailed criteria of the user evaluation.

Motif said the purpose of this objection is not to overturn the results. It also said that regardless of the outcome of the re-review, it will not participate in the third phase of the project and will continue developing frontier models with its own capabilities.

Source: IT DAILY · Kwon Young-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=241232

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.