AI

[AI-Ready Data 3] As AI Makes Direct Decisions and Executes Tasks, Data Management Must Become a Constant System

IT DAILY ·

Image created by AI

✦ AI Summary

As the use of generative AI and AI agents expands, data management needs AI Ready Data that goes beyond collection, cleansing, and analysis to manage meaning, freshness, and access rights.

DataStreams, WiseIT, IBM, and others proposed automated data quality checks, continuous operating systems, control over sensitive data and usage paths, and the need for real-time policy application.

Gartner said AI Ready Data must expand to Agent-Ready Data, and real-world cases are advancing continuous quality management, GraphRAG, and AI agent deployment.

As the use of generative AI and AI agents spreads across business operations, new standards are being demanded for data management. The data systems companies have built so far have centered on collection, cleansing, and analysis, but existing data may be ill-suited for AI use. That is because data is dispersed across multiple systems, and even the same information can have different meanings and management standards depending on the task.

In this environment, AI is taking on the role of directly exploring, interpreting, and using data for work, which is expanding the elements required for data management. The need to manage not only data quality but also meaning, freshness, and access rights in a consistent way is growing. As a response concept, “AI Ready Data” is drawing attention. The core of AI Ready Data goes beyond simple collection and cleansing, with the goal of enabling AI to understand data meaning and context and to trust the data.

This article is AI Ready Data 3, one part of a three-part series. The series examines changes in data management requirements driven by the expansion of AI use and covers the challenges and technical approaches involved in building AI Ready Data, as well as real-world applications.

This article covers [AI Ready Data 1] Changes in the criteria for “good data” in the AI era, [AI Ready Data 2] AI needs links to “business context” to understand corporate data, and [AI Ready Data 3] The expansion of AI’s direct decisions and execution, and the need for a “constant system” for data management.

The image was produced with AI.

The article says one-time maintenance has limits, so continuous management based on automated data quality is needed.

It says that simply building a real-time data supply system for AI cannot ensure reliability at all times, citing the possibility of changes in business systems and the addition of new data as reasons.

As environments that deliver data to AI immediately after it is created spread, the flow is toward the importance of a continuous management system to detect and improve errors early in such environments.

DataStreams emphasized an operational system for automatically checking quality throughout the data creation and change process for that reason.

DataStreams said this operational system should include data collection failures, data collection delays, changes in existing data structures, and the occurrence of outliers as management targets.

DataStreams said quality auto-checks are needed in data pipeline operations. It also identified delayed updates, collection failures, outliers, and schema changes as detection targets. In addition, it explained that an operating system is needed that can issue warnings or restrict AI use if quality falls below the standard.

DataStreams said data changes, as well as the management information that describes them, need to be updated together. It then explained that it is expanding its technology in that direction. The company already has its existing quality management solution, “QualityStream.”

DataStreams presented the “AI Quality Engine” for its AI data platform. The AI Quality Engine has functions for generating quality verification rules based on data characteristics, detecting errors, and suggesting improvement measures. Through this, it proposed an approach that automates some of the quality management tasks that were previously configured individually by people.

WiseIT is focusing on a continuous operating system that extends beyond one-time evaluation or concentrating on the point in time when a system is built, linking quality diagnosis to improvement and re-verification. A WiseIT official said data quality management needs to shift from a model centered on evaluation and system-building time to a continuous operating system, and explained that improvement and re-verification must continue after quality diagnosis. The official also said that when important data changes, related quality standards and AI-use datasets need to be checked together.

To support this, WiseIT uses its data quality management solution, “WiseDQ.” WiseDQ provides data-characteristics analysis profiling, quality diagnosis based on business rules, inspection schedule management, and error-improvement management. It also has a function that recommends suitable verification rules based on AI analysis of data characteristics, and it is focused on reducing the repetitive work of setting up diagnostics for staff in charge.

Errors discovered while operating AI services need to be linked to improvements in source data. If an error is corrected only in an individual AI service but remains in the source data, the same problem may recur in other services. Accordingly, the scope of management is not limited to source data.

In this flow, the scope of data governance management is expanding. The background to this is the connection between AI agents and enterprise systems.

Some say that in AI agent environments, it is necessary to review the purpose of data use and the tasks that can be executed. In this environment, the first management target is the information delivered to AI. The approach is also being presented in a way that distinguishes not just whether a system can be accessed, but the actual scope of use by data unit.

DataStreams proposed a plan to discover and classify sensitive data, apply policies and access rights by purpose of use, and track usage history. This aligns with a direction of managing not just the data itself, but also the purpose of use, access, and usage history.

DataStreams said the core of sensitive data management is not simply access control at login, but identifying sensitive data and controlling data units by user purpose and level of use. It also explained that it does not adopt a method of handing all sensitive data to AI and then controlling it in the LLM; instead, it manages control from the data platform stage before data is delivered to AI.

However, it pointed out that access restrictions at the data platform alone cannot manage the full process of AI agent use. Because agents can not only generate answers based on search results but also remember previous work information and call external system tools, it said a key issue is whether the permissions applied to source data continue to hold throughout later information-use processes.

GTOne concluded that in an AI environment, the objects of permission management need to expand to the entire AI use path. A GTOne official said there is a need to control the entire usage path, including prompts, vector DB, RAG context, agent memory, and tools, as well as source data and DB permissions. The official also cited purpose, authority, and regulation as the criteria for control.

GTOne also identified both permission management and tracking the AI agent’s decision-making process as challenges. Existing data lineage management has focused on identifying data creation, transformation, and usage paths, but as the agent environment changes, the tracking range needs to expand to the process after data use as well. GTOne provides the data lineage management solution “DataHawk,” which supports functions for tracing data creation points, processes passed through, and usage paths. In addition, it said a management system is also needed that connects the results of AI agent decisions and the history of tool execution.

These records are used to explain the basis for work performed by AI. They are also used to reproduce the decision-making process when problems arise.

In data governance, the timing of policy application is becoming more important. Separate from checking access rights before data is delivered to AI, it is also necessary to apply policies at the moment AI actually retrieves and makes decisions on information. Kim Yong-min of IBM Korea CSE said policies must be applied at the moment AI looks up data and makes decisions.

In this regard, IBM proposed an approach that manages data from the original point of creation. This approach assumes that metadata, lineage, and policies move along with the data. Kim Yong-min of IBM Korea CSE explained that in a real-time environment, when data is created, AI agents read, decide, and act on it immediately, so policies must operate in line with that flow.

Kim Yong-min of IBM Korea CSE said governance needs to move beyond post hoc checks and shift to a continuous system that moves at the same speed as data. He explained that governance must break away from after-the-fact review and become a continuous system that keeps pace with data.

This flow also aligns with the direction of moving from AI Ready to Agent-Ready. Gartner released a report in August titled “AI-Ready Data Needs to Expand to Agent-Ready Data.” Gartner argued that enterprise data management systems need to expand in step with the spread of AI agents.

Gartner presented the difference between existing AI Ready Data and Agent-Ready Data as a flow of expanding the objects of verification and the scope of application. As criteria for explaining the difference between AI Ready Data and Agent-Ready Data, Gartner presented three items: data alignment, continuous verification, and governance.

Existing AI Ready Data is a concept that continuously verifies suitability for a specific AI use purpose. By contrast, Agent-Ready Data is a concept that expands the scope of verification in a way that checks data reliability throughout the entire interaction process among multiple agents.

Gartner said that in an agent environment, data exchanges among multiple agents require matching meaning and context. It also explained that data suitability needs to be checked at each agent’s decision and task stage.

To this end, Gartner recommended that the agents using data clearly define the conditions required. It also recommended that the data provider present machine-verifiable evidence that those conditions have been met. Data contracts and metadata are tools that support this verification.

Gartner said context and meaning management are important in agentic AI use. However, it viewed context and meaning management alone as insufficient to prove data readiness. Gartner said data requirements need to be defined and verified dynamically at every stage of the agent workflow.

AI Ready Data is being built in a variety of ways in actual businesses. Approaches include turning field experts’ experience into data and connecting business records scattered across multiple systems. Examples include improving the search accuracy of existing AI services and reorganizing data management systems.

A representative case is the high pathogenic avian influenza (HPAI) risk prediction project promoted by BigValue and the government. Based on the epidemiological experience of on-site quarantine experts, BigValue defined conditions related to infection risk. It then turned those conditions into measurable variables such as environment, quarantine, and vehicle movement.

BigValue used past outbreak cases. It also verified how well those variables explained actual outbreak patterns. The case showed a flow of converting field quarantine experience into measurable variables and validating their effectiveness through past cases.

To reduce inconsistencies between farmland information and actual land-use conditions, supplementation is being made with external spatial information. If the prediction results do not match the field judgment of quarantine officers, the relevant data and variables are reviewed again. Work is also under way to continuously reflect new outbreak cases and on-site verification results.

A BigValue official said the system collects new data every day during operation and runs quality checks. The official also said the model is retrained to reflect new outbreak cases, and after performance is verified against past outbreak cases, it is applied in the field. The results of on-site verification by quarantine officers are then reflected in the next round of training.

At the same time, there have been cases in which the existing data quality management methods of public institutions have been improved. The Seoul Facilities Corporation adopted WiseIT’s data quality management solution, WiseDQ. Through this, it built a continuous data quality management system.

The Seoul Facilities Corporation case is presented as one in which data quality management for public data was shifted from a process centered on quality diagnosis and evaluation timing to a daily diagnosis and improvement system. Photo credit: WiseIT.

WiseIT explained that it has meaning because it shifted quality management from a method carried out only at specific evaluation times to a continuous management system. It also said that through such a system, it can be used as a foundation to continuously secure the reliability of source data needed for future AI use.

Next, as a case of using data management technology in the process of advancing generative AI services, a public institution AI search service upgrade project carried out by DataStreams was presented. The institution had been operating an existing RAG-based service, but it faced difficulties maintaining the freshness and accuracy of search results because of frequent changes in laws and regulations and the complexity of internal data structures, and it also had limitations in connecting results to actual work beyond simple question-and-answer interactions.

Accordingly, DataStreams reorganized data standards and data meaning, and improved the quality management system and metadata management system. It then applied GraphRAG combining vector DB and a knowledge graph to use relationships among data, and introduced a LangGraph-based AI agent to expand the functional scope to query analysis, information retrieval, document drafting, and linkage with existing business systems.

A DataStreams official said existing one-off, question-and-answer type AI is evolving into AI that maintains business context and supports real outcomes. The official also explained that the direction of AI advancement lies not only in improving the AI model itself, but in connecting data fabric, data quality, governance, GraphRAG, and agentic AI.

Flitto is using AI translation services in industrial settings and in the financial sector based on language data construction and verification technology. Flitto supplied Hanwha Ocean with the AI translation solution “Otalk,” which includes noise-canceling functionality. “Otalk” supports high-noise environments and up to 42 languages, including Thai, Indonesian, and Vietnamese. It also supports work instructions and reporting, and Flitto is also providing AI translation solutions to 10 major domestic and international financial firms, including Hana Bank and Woori Bank.

Encore is carrying out AI Ready Data build PoC and projects across multiple industries. It is also refining the direction of strategic cooperation with Hanwha Systems to combine data governance and ontology with enterprise AI business. At the same time, it is pursuing AI Ready Data building and collaboration by first diagnosing customers’ data readiness, then structuring the meaning and relationships of the data needed for actual work, and finally applying the structured data to AI agents and others for verification.

An Encore official explained that expanding ontology itself is not the goal, and that it is necessary to verify first from the minimum scope capable of solving clear business problems in areas such as customer consultation, financial analysis, manufacturing quality, and equipment management.

GTOne is evolving its product lineup in a direction that connects existing data management technology to AI use environments. “DQMiner” diagnoses and improves erroneous data, “MetaCatalog” supports integrated search and management of structured and unstructured data, and “MetaMiner” manages standard terms and data structures. DataHawk traces data flows.

A GTOne official explained that the company is not stopping at providing individual data management functions but is pursuing the connection of quality, context, identification, and traceability. The official added that the company is advancing its technology in a direction that builds an AI Ready Data governance foundation for the use of trustworthy enterprise data in AI.

IBM watsonx.data presented an outline that supports AI use through integration of data fabric, metadata, governance, open-source data formats, and hybrid infrastructure.

IBM Korea introduced a case in which Lockheed Martin replaced multiple data lakes and 46 data management, analytics, and business intelligence (BI) systems with a single integrated platform. In the process, it reduced data and AI tools by 50% and built an AI factory capable of AI development and deployment for about 10,000 engineers. Kim Yong-min of IBM Korea CSE said this established a company-wide foundation of trustworthy data and improved AI response accuracy by more than 20%.

IBM Korea explained that simply possessing the same data has limits for AI use. It said AI use is constrained if there is no connection to business decision standards, and also constrained if changed on-site conditions are not reflected.

Accordingly, companies need to clarify the work for which they will use AI, and it is important to verify whether the necessary data is actually ready. It also said the level of AI Ready Data depends less on how much data is held than on the ability to connect it to actual work, and also on the ability to use it stably across a variety of AI services.

Source: IT DAILY · Yang Seung-gap
Original: https://www.itdaily.kr/news/articleView.html?idxno=242092

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.