The Mirage of Open Models: What Does “Open” Really Mean?
IT DAILY ·
✦ AI Summary
Open-source AI must provide weights, the sources and composition and processing of training data, training and inference code, and freedom to use, study, modify, and redistribute, but many models in the market are closer to open weights.
Open weights are more open than closed models, but nondisclosure of training data and training processes makes it difficult to verify and reproduce the causes of bias and limitations.
As AI expands into autonomous complex work, the limits of a model-weights-only approach have become clear, and if the core of system connection and control is dependent on a specific big tech platform, it is difficult to secure technological leadership.
In the global AI market of late, competition among big tech companies over large-scale models and platforms has become a major focus. But during this month's on-the-ground reporting in the industry, it became clear that technical power is dispersed and that the open-source ecosystem remains the core driver behind system operation. After checking the situation in the field, questions arose over whether the industry's idea of “open” is really faithful to the values of open source.
That sense of concern was repeatedly confirmed this month on the ground at “Open Source Summit Korea 2026” and the “MCP Developer Summit Seoul” in Seoul. According to the “Open Source AI Definition 1.0” presented by the Open Source Initiative (OSI), open-source AI must provide weights, information on the sources of training data, information on how training data is assembled and processed, training and inference code, and freedom to use, study, modify, and redistribute.
However, many models described in the market as “open source” were found to be closer in reality to open weights, with only completed weights and some code disclosed. In many models, the contents of the training data, the training method, safety evaluation standards and procedures, and bias evaluation standards and procedures are either not disclosed or disclosed only insufficiently. This information is essential for judging model reliability.
Open weights are more open than closed models because they can be downloaded, run on in-house infrastructure, and fine-tuned. They are also similar to past software in that they disclose the output but not the production process, and they reduce dependence on application programming interface (API) providers while giving users the option to run systems themselves. As a result, operational independence is possible.
However, when the composition of the training data is not disclosed, it is difficult to verify the causes of model bias. When the training process is not disclosed, it is difficult to verify the causes of a model's limitations. In addition, when training data and the training process are not disclosed, reproducing results through retraining under the same conditions is restricted. And when the training process, licenses, and connected platforms are closed, autonomous control over the model supply chain and the path of technological development becomes difficult. This leaves room for the continued possibility of limits on technical reproducibility and long-term choice.
As AI evolves beyond simple question-and-answer tasks into a stage where it can carry out complex work autonomously, critics say that simply opening the model itself is not enough. An AI model is only one component that serves as the brain of the overall system, and actual business operations require linking vast amounts of internal corporate data as well as seamless integration with existing business systems.
For that reason, the limits of a state in which only the model weights are disclosed are becoming even clearer, and the limits of half-hearted openness are deepening, according to the assessment. In particular, if the key channel for system connection and control becomes dependent on a specific big tech platform, it becomes impossible to secure technological leadership, which is being pointed to as a problem.
An analysis says the background to this normalization of unverifiable, half-open models lies in the convenience-driven practices of the domestic industry. South Korea's industry tends to see open source as something to be taken for free and used as an end product, and it has been criticized for aggressively consuming open source while being stingy in contributing to the ecosystem, as well as for failing to properly value the intellectual labor of developers who work with code.
The reality that SW value assessment remains centered on the convention of counting input labor and time, and that the reason for using AI tools is reduced to an attempt to cut development costs, exposes a distorted view that treats developers as mere functional workers. This also conflicts with the common refrain from the government and industry of emphasizing “technological sovereignty” and “sovereign AI.” At the same time, there are also criticisms that the government and industry have neglected fostering the open-source ecosystem and treating developers properly.
In this regard, comments by Democratic Party lawmaker Im Moon-young on sovereign AI point out that its essence is not exclusivity in the sense of “only using our own technology.” Im explained that the key is securing an “alternative” that makes it possible to control and choose technology without relying on a specific big tech company. In the process, it is also argued that there is a need to move beyond the illusion that open weights are simply “free outputs.”
To that end, companies must go beyond the role of simple consumers and take the lead in establishing technical standards. The public sector is expected to introduce open source proactively and pay fair compensation. In addition, the intellectual labor of developers in the field must be properly recognized, and recognizing the value of developers will lead to the completion of sustainable AI sovereignty that goes beyond technological dependence.
Source: IT DAILY · Kwon Young-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=241136
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.