Software

Two Key Challenges in the Open Source AI Era: 'Agent Standards' and 'Data Supply Chain Risk'

IT DAILY ·

On-site at Open Source Tech Day 2026. [Photo: Kwon Young-seok]

✦ AI Summary

'Open Source Tech Day 2026' was hosted by the Ministry of Science and ICT and the National Research Council of Science & Technology (NST), and its keynote sessions covered the standards and compliance challenges arising from the real-world deployment of agentic AI.

Mazin Gilbert argued that open standards such as MCP and A2A, as well as an L4 platform standard, are needed to build the agent internet.

Executive Vice President Lee Hwayoung presented an AI-BOM and a data compliance framework to prevent license chain contamination in an environment where 99% of data depends on open data and synthetic data.

A photo from the scene at 'Open Source Tech Day 2026' was provided by reporter Kwon Young-seok. As agentic AI is being deployed in the real world, attention in the open source community is shifting beyond the technology itself to operational standards and compliance frameworks. 'Open Source Tech Day 2026' was hosted by the Ministry of Science and ICT and the National Research Council of Science & Technology (NST), and concerns from the industry surfaced in the keynote session. The speaker lineup consisted of global open source foundations and domestic AI researchers.

The event identified two core practical tasks for the open source community: interoperability standards and data supply chain compliance. On one hand, securing open protocols was cited as a way to counter dominant big tech players. On the other, in an environment where 99% of training data depends on external open data, license chain contamination was flagged as a major risk. As a result, securing open protocols and preventing license chain contamination were presented together as survival variables for AI companies.

Against this backdrop, two strategies were presented together in the keynote session. Strategy 1 was an 'L4 platform standard' for connecting agents, and Strategy 2 was an 'AI Bill of Materials (AI-BOM)' for proving data lineage. This showed that in the phase of real-world agentic AI deployment, both connection standards and data source and license management are required.

Mazin Gilbert, executive director of the Agentic AI Foundation (AIF) under the Linux Foundation, delivered the first keynote address via video and framed the dawn of the agent internet era as his central message. The photo was taken by reporter Kwon Young-seok.

Gilbert said AI has already moved beyond the single-turn question-and-answer stage. He explained that AI is evolving into agentic systems that use tools on their own to accomplish goals.

Based on a macro-level market analysis, Gilbert argued for the need for open standards to build the 'Internet of Agents (IoA).' Citing how the web ecosystem grew on the basis of open protocols such as TCP/IP and HTTP and neutral governance, he said the same foundation is essential for the agent internet.

Gilbert pointed out that if standards are absent, the dominance of specific big tech firms could intensify. He characterized the lack of standards as a systemic risk and explained that interoperation between agents and tools could become recklessly fragmented into an N×M connection structure, creating massive integration costs.

Over the past year, token prices fell 90%, but usage rose 50 times over the same period. This was presented as a Jevons Paradox phenomenon. He also presented the view that the performance gap between open models and closed models has shrunk to just 4 months, indicating that model-level differentiation is gradually leveling out.

In line with this trend, he identified the 'agentic platform' among the five layers of the agentic AI stack as the key battleground in the market going forward. Director Gilbert said the point where real value is created lies in the L4 layer. The L4 layer handles runtime, authentication, observability, and coordination of communication between agents. Gilbert said open standards such as MCP and A2A must be secured first, explaining that doing so would reduce integration costs and maximize network effects.

According to an AIF survey, 81% of member companies were operating in production environments. The same survey also found that 89% of member companies were evaluating or adopting MCP. These figures show that the shift in competition toward layers above the model is already under way at the actual adoption stage.

A photo taken by reporter Kwon Young-seok showed Lee Hwayoung, executive vice president at LG AI Research. Lee delivered the second keynote address, presenting the need for an 'AI-BOM' to prevent license chain contamination.

Lee diagnosed legal risks across the large-scale training data supply chain and offered a solution by sharing a practical compliance framework independently developed by LG AI Research. He explained that the scale of training data has surged to 40 trillion to 400 trillion tokens, and that 99% of the data relies on open data and synthetic data such as web crawling.

Against that backdrop, he explained copyright infringement risks at the collection stage, the processing stage, the pretraining and deployment stages, and the inference stage. He then noted that copyright infringement risks arise in a chain across the entire supply chain.

In particular, while there are processes for merging and redistributing dozens of sub-datasets, he explained that a condition from the original author, such as 'Non-Commercial,' causes license change problems as it passes through open source portals. He said cases of distortion occur in which the 'Non-Commercial' condition turns into a commercial license such as Apache or MIT.

Lee said this phenomenon is linked to copyright risks across the entire supply chain and identified 'license chain contamination' as a serious bottleneck. He also shared solutions for responding to this problem.

To prevent issues with data usage eligibility, LG AI Research has introduced both a data compliance framework and an agentic AI-based automated verification solution. Data is classified into A, B, and C grades based on 18 criteria, including whether commercial use is permitted, with weights assigned to each criterion. Grade A means commercial deployment is allowed, Grade B means limited internal use, and Grade C means in-house research only.

The system applies the Worst-Case superset principle. If even one sub-dataset is graded C, the entire model is downgraded. LG AI Research said reviewing one dataset previously took weeks, with manual review by lawyers creating the bottleneck, and set as its goal building agents that autonomously trace sources and evaluate licenses.

Accordingly, the review framework was designed so that agents handle the first review and lawyers perform the final cross-check, and the system is scheduled to be opened to the public in November. Lee said institutionalizing the AI-BOM, which can transparently track base models, fine-tuning data, and license grades, is necessary, and that such institutionalization is a condition for responding to global regulatory barriers such as the European Union (EU) AI Act.

During the on-site Q&A, a speaker said black-box models with unclear provenance are not risk-free, and that companies bear the entire risk themselves. The speaker added that transparent data lineage must be secured and that transparent data lineage strengthens the sustainability of the open source AI ecosystem.

Source: IT DAILY · Kwon Young-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=242046

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.