Policy

Even If a Contender Falls in the Independent AI Foundation Model Race, the Data Remains... 1.56 Trillion Tokens Unlocked

IT DAILY ·

(AI가 생성한 이미지)

✦ AI Summary

The government will open AI training data built through the independent AI foundation model project and convert it into a shared resource for the domestic AI ecosystem.

The Ministry of Science and ICT and NIA said on the 27th that they will release 29 types of AI training data secured by the five elite teams that participated in the first-stage evaluation, amounting to about 35.44 million items, 11.3TB, and about 1.56 trillion tokens.

The open data consists of text for LLM pretraining, multimodal data, AI agent training data, Korean-style red-teaming data, and physical AI data based on household environments, and can be accessed in the 'Independent AI Model Data' section of AI Hub.

The government will turn the AI training data built through the

The scale of the release is 1.56 trillion tokens. The open data consists of text for LLM pretraining, multimodal data based on video, audio, and images, AI agent training data, Korean-style red-teaming data, and physical AI data based on household environments.

Small and midsize AI companies, universities, and research institutions are in a situation where it is difficult to build large-scale data on their own. Accordingly, the open data is expected to serve as a foundation for reducing the initial cost and time required for model development.

The Ministry of Science and ICT and the National Information Society Agency (NIA) announced on the 27th that they will make public 29 types of AI training data secured by the five elite teams that participated in the first-stage evaluation of the independent AI foundation model project on AI Hub. This release aims to spread the training data built with KRW 15 billion beyond the elite teams.

The organizations that built the publicly released data are Naver Cloud, Upstage, SK Telecom, NC AI, and LG AI Research. The total scale of the open data is about 35.44 million items and 11.3TB, and the estimated amount converted into text tokens is about 1.56 trillion tokens.

The government said the scale of the data could, in theory, be used to train large AI models with about 70 billion to 80 billion parameters. However, it added that whether a model can actually be built and what its performance will be depend not only on the amount of data, but also on factors such as deduplication, the level of cleaning, training objectives, data composition, and computing resources.

The data released this time are the result of the elite teams in the independent AI foundation model project using the government’s 2025 data-building and processing budget of KRW 15 billion. The data planning criteria were each team’s model development strategy, and the purpose of the independent AI foundation model project is to support domestic companies in developing their own AI foundation models.

The independent AI foundation model project has been carried out as a competition in which the number of participating teams is reduced through phased evaluations. For this reason, as possible criteria for judging business efficiency, the use of the budget invested in teams that fail the evaluation and the use of the development results of those teams are being discussed together.

Accordingly, this release means that only the winning model in the competition will not remain as a government project achievement. Data built during the development process can be reused by other companies and researchers, which has structural significance in that it does not stop at supporting specific elite teams but returns assets formed in the support process to the domestic AI ecosystem.

As a condition of participation in the project, there is an obligation to disclose more than 50% of the data secured with government data-building and processing funds. After the first-stage evaluation, the five teams submitted their data to NIA and the Telecommunications Technology Association (TTA), and NIA and TTA are responsible for checking the quality, privacy, harmfulness, and license restrictions of the submitted data so that other companies and researchers can reuse it.

The participating organizations showed differences from the data disclosure method itself. Naver Cloud, Upstage, SK Telecom, and NC AI set a policy of fully disclosing the data that passed quality verification. By contrast, LG AI Research decided to disclose samples extracted at set intervals by dataset to meet the mandatory disclosure threshold of at least 50%.

The samples disclosed by LG AI Research undergo a statistical selection process, and the verification body for that process is TTA. While disclosure methods differ between full disclosure and sample disclosure by institution, the open data allows a look at the AI capabilities that each elite team prioritized developing.

This strategic presentation was made for five teams, including video, agent, and physical AI teams. The composition and uses of the disclosed data were presented in a way that reveals which fields each team first built capabilities in.

Naver Cloud built 2.34 million publicly released videos, 520,000 broadcast videos, 15.5 million text items based on video clips, and 12 million audio question-and-answer items. The data includes scene-level descriptive captions and is used for AI video comprehension training and image generation model training.

Upstage disclosed 1 trillion tokens’ worth of pretraining data and 500,000 items of post-training data. The pretraining data is used for the model’s initial training, while the post-training data focused on developing AI agents with reasoning, judgment, and execution abilities beyond question and answer.

The data was designed with the goal of developing models that understand user intent and decide on subsequent actions, rather than simply producing natural-sounding language. The data set consists of 300,000 items of knowledge and intelligence data, 100,000 items of user-preference data, 100,000 items of agent behavior-capability data, and 2,084 benchmark items.

To this end, SK Telecom built difficult multi-step reasoning data in specialized fields such as math, science, and law, and also built post-training data based on voice and images. It also included 10,000 items of Korean-style red-teaming data reflecting domestic laws and the common values of Korean society.

Red-teaming data is used to present AI with aggressive and risky questions to identify vulnerabilities and verify the possibility of harmful responses. Overseas safety data has the limitation of insufficient reflection of Korean law and Korean cultural context, and Korean-style red-teaming data may be used as a resource to address those limitations.

NC AI built training data based on real-world industrial data such as manufacturing technical documents and voice recordings from customer service consultations. The training goals are long-context understanding, step-by-step reasoning, multi-turn dialogue, question and answer, and multimodal capabilities. The release includes high-quality multi-turn data, synthetic and instruction data, chain-of-thought training and verification data, multilingual pretraining and machine reading comprehension data, as well as image and voice data in manufacturing, medical, and civil complaint domains.

This composition is focused on expanding specialized AI for industrial settings, beyond improving general-purpose conversational performance.

LG AI Research built a multimodal physical AI dataset that reflects real Korean home environments. To do this, it filmed 50 households in South Korea and secured more than 170,000 video clips covering more than 50 types of household activities. The labeling items are objects, region segmentation, human posture, and surrounding context, and the labeling method has a multilayer structure.

The release includes data for training humanoid robot scene comprehension, and about 152,000 images and text data are being disclosed. The data is intended for use in understanding the structure of Korean living spaces and living environments, and serves to address the difficulty of learning Korean living spaces using only overseas data.

Those who may use the data include domestic companies, researchers, and students who are Korean nationals, and the access route is the 'Independent AI Model Data' section of AI Hub. The method of use is free download and utilization. However, some data are subject to separate usage restrictions, and the reason for the restrictions is the need for security and privacy protection. The location for using these restricted data is 'Safe Zone,' the closed online development environment provided by AI Hub, and the procedure for using the restricted data requires a separate application.

It was also noted that releasing the data alone would not immediately improve AI model competitiveness. The terms of use may differ by dataset, and large-scale data cleaning and training require substantial storage space and computing resources. The criteria for judging business performance are the actual number of downloads after release, the number of companies using the data, and the models and service cases developed using the data. The point made was that release alone cannot determine results, and follow-up usage is necessary.

The achievements in data building should be judged not by the total volume of 1.56 trillion tokens alone, but by the degree to which duplication and errors are reduced, and by the extent to which the data include Korean-language and industry-specific information that is difficult to secure with existing open data, as well as the range in which they can be used for commercial model development. In addition, even after passing quality verification, the usefulness of each data set needs to be continuously checked and supplemented during the actual model training stage.

The Ministry of Science and ICT said that additional data built with government funds in the second-stage evaluation of the independent AI foundation model project will also be disclosed after quality verification. As the project proceeds, the common training data accumulated in AI Hub is also expected to expand.

Kim Kyung-man, director of the AI Policy Bureau at the Ministry of Science and ICT, said the independent AI foundation model project has significance because it contributes to the qualitative growth of the entire AI ecosystem through the experience and know-how accumulated during development, beyond AI model development itself. He also said the open data is the foundation for the growth and self-sufficiency of the AI ecosystem.

Source: IT DAILY · Lee Jae-young
Original: https://www.itdaily.kr/news/articleView.html?idxno=241247

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.