AI

[AI-Ready Data II] For AI to Understand Enterprise Data, It Must Be Connected to Business Context

IT DAILY ·

Image created with AI

✦ AI Summary

As generative AI and AI agent use spreads in enterprise operations, the need for AI Ready Data and consistent management of data meaning, freshness, and access rights is growing.

GT One and ENCOA proposed data systems AI can understand by structuring business context with metadata, ontology, data catalogs, Context Maps, and Semantic Models.

WISEIT, DeltaStream, and IBM introduced ways to supply data for AI use through unstructured document management, CDC, streaming, Data Fabric, real-time data processing, and governance.

As generative AI and AI agents spread across enterprise operations, a new standard is being demanded for corporate data management. Companies have already established systems for data collection, cleansing, and analysis, but existing data may be unsuitable for AI use. Data is scattered across multiple systems, and even the same information can have different meanings and management standards depending on the business context.

In environments where AI directly searches, interprets, and uses data for work, the need is growing not only for consistent data quality management, but also for consistent management of data meaning, freshness, and access rights. Amid these changes, the concept drawing attention is "AI Ready Data." Its core lies in preparing data not just through collection and cleansing, but in a state where AI can understand its meaning and context, and in making data trustworthy for AI.

IT Daily examines the data management requirements that come with the expanding use of enterprise AI through AI Ready Data II, highlights the challenges in building AI Ready Data, and covers technical approaches and real-world applications. This feature is divided into three parts.

In the AI era, the standard for good data is changing, and GT One believes that AI must be connected to business context in order to understand enterprise data.

Accordingly, the target of context structuring for AI goes beyond data definitions to include business relationships, and after establishing the scope of data and verification standards, structuring business knowledge is proposed as the next step.

The purpose of this structuring is to transform the knowledge held by a company into a form that AI services can use in common.

A data catalog was presented as a means to achieve this, and the data catalog is responsible for finding the data needed and checking where that data is used in business.

However, to operate a data catalog continuously and consistently, metadata management is necessary, and metadata includes the data name, data structure, and definitions used in business.

GT One believes that such metadata can establish common standards between internal business terms and data used for AI.

A GT One official said the role of metadata is to prevent AI from interpreting data arbitrarily, to standardize the meaning, structure, attributes, and relationships of data, and to function as a common language. The official added that metadata connects business terms with logical and physical data, and provides a foundation in which the business meaning and the meaning understood by actual data and AI align.

However, it has been pointed out that metadata alone has limits in expressing the complex relationships and decision rules of real-world operations. ENCOA is focusing on ontology as a way to address these limitations.

An ENCOA official said that metadata alone makes it difficult to fully explain the business meaning of data and its relationships with other data. The official then presented ontology as a way to supplement those limitations.

Ontology is a knowledge system that defines the attributes, relationships, and rules of business entities such as customers, orders, and contracts, and it is structured so that both people and machines can understand it. Implemented as a knowledge graph, it can ease the limitations of AI's simple synonym search and allow AI to search for needed evidence by following relationships and rules among business entities.

ENCOA is building a Context Map to connect the data structure and common concepts across the enterprise based on common corporate data definitions. Based on this, it creates business context by progressively combining the Semantic Model needed for each specific business area and the relationship between that model and actual data, while reflecting decision criteria and relationship-linking methods for each business area.

ENCOA is considering the possibility of ensuring that the business context it has built is not consumed by a single project but can also be used by other AI agents. It is also focusing on managing the built business context as a shared asset.

An ENCOA official said the key is not stopping at the success of the first AI project. The official added that the data meanings and decision criteria confirmed in the first task must be left as the starting point for the next task, so that AI investment can be accumulated as a long-term corporate data asset rather than repeated PoCs.

The photo was provided by ENCOA.

The methodology for building AI Ready Data assumes that a company's business knowledge is accumulated not only in databases, but also in unstructured documents such as contracts, business reports, and manuals. To use these documents in AI services, it is necessary to analyze the content and extract the required information.

A GT One official explained that with unstructured data, AI cannot easily understand business meaning and its relationship with other data simply because the document exists. The official added that key terms, classifications, summaries, and relationships must be extracted from the document content, and that the extracted information must be managed in connection with existing structured data.

In the actual build process, materials in various formats such as PDF, Word, and Hangul documents are analyzed and key information is extracted. The build is then carried out by linking the extracted content with existing structured data.

In this process, Chunking, which splits documents into small units to support AI search, and Embedding, which converts text into vectors, are used. The processed information is provided in the form of a document corpus for RAG and a vector DB.

There is concern that search materials need to have the validity of the target documents managed even after document processing. Even documents for the same business can differ depending on when they were written, and revised regulations and contract terms can remain alongside older materials. Accordingly, it is necessary to understand the contents of the documents being searched, and to manage information that can confirm whether the materials are currently valid.

A WISEIT official explained that for unstructured data, it is difficult to identify the information needed for use based only on the file name, storage location, and extension. The official added that users need to be able to confirm the content, related business, source, person in charge, and whether it is currently valid.

In response, WISEIT is preparing to launch its data asset management solution, "WiseGrandata." "WiseGrandata" is a product that collects and catalogs metadata from structured and unstructured data dispersed across an organization, and supports search and discovery of the needed data. WISEIT is planning automatic classification features for the type and business area of collected data, as well as features that recommend tags and content descriptions, and is focusing on managing legacy documents as assets that business users and AI developers can find and use.

AI use raises not only the challenge of accurately searching and using information, but also the challenge of supplying data in a timely manner. When using frequently changing information such as orders and inventory, existing stored data alone has limits in accurately reflecting the current business situation. Accordingly, a system is needed to deliver changes in source systems to AI at the needed time.

Kim Yong-min of IBM Korea CSE said companies expect AI, analytics, and automation to reflect current data, not data from hours ago or days ago. However, the reality for many companies is that they rely on previous-day data loaded through overnight batch processing, and as a result, they analyze and make decisions based on data that is not real-time, he explained.

To address these limitations, DeltaStream is combining Data Fabric with existing data warehouses and data lakes. Data Fabric is a technical architecture that supports integrated connection and use of data across different systems. Based on this, DeltaStream is pursuing expansion into an AI data platform.

In its operational approach, it maintains the existing periodic data collection system. It is adding real-time data processing and data virtualization as additional elements. DeltaStream is responding by adding these elements to its existing storage-centered system.

Core technologies used for data supply include CDC and streaming. CDC is the concept of capturing changes in the source database and delivering them to other systems. Streaming is the concept of collecting and processing continuously generated data so that AI can use the latest information.

DeltaStream has a configuration that combines its existing data integration solution, "TeraStream," with its CDC solution, "DeltaStream." It also provides a way to combine "TeraStream" with the streaming processing product "TeraStream BASS."

When data is stored in multiple systems, "TeraOne SuperQuery," which supports distributed queries and data virtualization, is used. According to a DeltaStream official, "TeraOne SuperQuery" uses connection, virtualization, and metadata-based methods instead of data movement and duplication. It also enables needed data to be found and connected in time.

IBM is expanding its real-time data supply system by combining its own data management framework with real-time data processing technology. Through the combination of data streaming technology and existing data management platforms, it is pursuing an IBM strategy that connects data collection and integration to actual AI use in a single flow.

In this flow, IBM uses Confluent technology to collect and process data in real time, and uses the open hybrid lakehouse, "watsonx.data," to integrate on-premises and cloud data. It then links this with "watsonx.data intelligence" to carry out data quality management, data lineage management, and data access policy management, creating a structure that extends through to AI use.

Kim Yong-min of IBM Korea CSE said that combining IBM watsonx.data and Confluent can turn real-time business events into governed AI Ready Data. He also explained that AI agents can use data quickly and with confidence, and that the entire process from the time data is created to the time it is used by AI can be connected in a real-time single flow.

Source: IT DAILY · Yang Seung-gap
Original: https://www.itdaily.kr/news/articleView.html?idxno=242051

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.