From an AI That Summarizes and Answers to One That Classifies and Processes
IT DAILY ·
✦ AI Summary
ACNS Executive Vice President Kook Hyun-woo said that as generative AI spreads, electronic document use is shifting from document explanation to document-processing judgment.
He explained that N2SF-based C/S/O classification requires analysis of the entire document, searching for applicable criteria and original-text evidence, and presenting candidate grades and the basis for judgment.
He also said ACNS prefers directly judging the current document according to criteria rather than accumulating sensitive and nonpublic documents as separate vector search data.
In an interview-format article, ACNS Executive Vice President Kook Hyun-woo said that electronic document usage is changing as generative AI spreads across companies and public institutions. He said the main functions of existing LLMs for document use include summarizing content, searching for needed information, and answering user questions. He added that recently, RAG-based work support systems have been spreading, with the main focus of these systems centered on easy search and use of documents held by organizations.
Kook said that requirements are also changing as generative AI enters actual work processes. He said electronic document use is now moving from a stage of explaining documents to a stage of judging how documents should be processed.
He explained that at this judgment stage, it is necessary to determine whether a document contains important information, whether it can be disclosed externally, what level of internal protection it requires, and whether it can be used with AI services. He added that such judgments are made based on a comprehensive consideration of the document content and institutional standards.
The National Network Security System (N2SF) requires the classification of work information into three grades: C (Classified, confidential), S (Sensitive, sensitive), and O (Open, public). The classification standard is importance and sensitivity. As a result, institutions are now faced with the challenge of how to actually identify and manage vast amounts of electronic documents.
ACNS, a company developing electronic document processing technologies, sees this situation as a process in which the role of document AI is expanding from an information utilization tool to a work-judgment tool. This interview was conducted with ACNS Executive Vice President Kook Hyun-woo and focused on the N2SF C/S/O document classification function of the product under development, Docu AI, an electronic document classification and processing engine.
The interview also compared AI that judges and processes documents with existing document search and RAG. It also raised the question of whether C/S/O classification under N2SF is a representative example.
Within institutions, documents are being created and stored in large volumes, and document content must be checked to separate work information into the grades required by the N2SF framework. Both newly created documents and existing documents will need to be reviewed in the future, but it is effectively impossible for staff members to inspect each document individually and determine C/S/O one by one.
Accordingly, ACNS is reviewing an approach in which AI analyzes the entire document, searches for applicable criteria and original-text evidence, and suggests a grade. This method does not simply ask AI to assign one of C/S/O, and it requires prior structuring of institutional tasks and applicable criteria. It also requires AI to analyze documents based on the criteria, and the goal of the feature currently under development is not simple classification but automatic classification that includes C/S/O grade, applicable criteria, and an explanation of the original-text evidence.
Among these, clearly defined documents become targets for automatic classification. By contrast, documents in which criteria conflict or for which necessary fact-checking is not possible become targets for staff review.
Regarding the topic of the question, which concerned a vector DB built from sensitive and nonpublic documents and a similarity-comparison method, the speaker said the company’s approach is substantially different and pointed to two major differences. First, he noted the issue of whether protected information needs to be reaccumulated as separate search data. He explained that the vector DB approach assumes similarity comparisons between existing nonpublic or sensitive documents and new documents, and that this requires embedding institutional documents and managing them in a searchable form.
He went on to say that even when the original text is not stored, embeddings and search indexes are still derived from nonpublic and sensitive information. Accordingly, he said additional management is also required for embeddings and search indexes, including access control, retention, deletion, and backup. He added that this creates additional items requiring management.
For that reason, he said that from the perspective of public institutions, a structure that gathers sensitive documents again in order to detect sensitive documents can be burdensome. His point was that the structure itself, in which protected information must be recreated and managed as a separate search asset for sensitive-document detection, is a burden.
Finally, he said Docu AI does not use, as its basic structure for C/S/O classification, a method of accumulating past nonpublic and sensitive documents as separate vector search data. He presented this as a difference between the company’s approach and a vector DB-based similarity-comparison method.
The image description shows the Docu AI-N2SF C/S/O classification processing structure, and the photo source is ACNS. In security-grade classification, document similarity and security-grade identity do not match. Even documents containing similar content may already be public materials, or they may be internal review materials before release, making their status different. Bidding materials, too, can be judged differently depending on whether the work is ongoing or has ended. In addition, public body text and nonpublic attachments can exist together within the same document.
In response, ACNS believes that judging the current document according to criteria is more important than finding similar documents. Accordingly, it chose a method that directly determines which requirements of the institution’s standards the current document meets, rather than checking how similar it is to past nonpublic documents. In this process, the content, structure, work context, and existing management information of the current document are used as analytical factors, while institutional policies and judgment criteria are reflected as application factors.
The output of this approach presents candidate C/S/O grades and their basis. Only documents requiring additional judgment are separated out for staff review.
Regarding the LLM processing method used in actual document analysis and grade judgment, the speaker explained that while the LLM plays an important role, the system does not leave grade decisions entirely to a single LLM. He said that security-grade judgment cannot be made from just a few sentences and that the body text, title, tables, document structure, surrounding context, existing management information, work status, and attachments must all be examined together.
ACNS said it avoids an approach that splits documents and makes judgments based on only some sentences, and instead analyzes documents from a perspective that comprehensively checks the necessary information while preserving the overall structure and context. The point is that document security-grade judgment is not carried out by reading only a few sentences, but by checking the various factors surrounding the document together.
In this process, the LLM is responsible for analyzing sentence meaning and areas that require contextual judgment, such as before-and-after relationships. By contrast, information with a clear format, such as a resident registration number, or existing management information within a document, which does not require LLM inference, is checked in an appropriate way, and the LLM is used only for parts that require interpretation of context and meaning.
In the final processing stage, the institution’s C/S/O judgment criteria and the results are synthesized. Based on this, the final output structure presents candidate grades and the basis for judgment.
The speaker explained that the C/S/O classification function is not a method of directly querying an LLM for a document grade. He said the development direction of the C/S/O classification function is to check document structure and work context, analyze applicable criteria and evidence, and examine conflicting information and exceptions across the document.
He then said the key to LLM use is analyzing documents without omissions and accurately searching for evidence that meets institutional criteria. He described this as C/S/O classification not merely as a grade decision, but as a method that checks document structure and context, criteria and evidence, and exceptions and conflicts together.
When asked how different classification criteria by institution would be reflected in AI, he said common criteria and institution-specific criteria need to be separated. He said the common application criteria of N2SF and related laws can be built into the basic policy, while at the actual application stage, institution policies reflecting each institution’s work characteristics, detailed internal criteria, and document types need to be separately designed.
He also explained that when reflecting institutional criteria, greater importance is placed on adjusting policies and judgment guidelines rather than on accumulating large volumes of institutional documents and retraining on them. He said Docu AI separates common policies from institution-specific policies, and its development direction is also focused on policy adjustment that reflects institutional criteria and guidelines at the actual application stage.
The nature of grade classification is a starting point, not an end in itself. The important factor is how classification decisions are linked to follow-up work. The document’s C/S/O grade can be checked, and the presence of protected information can also be confirmed. The confirmed classification result can then be used in the system.
Uses for the classification result include reviewing whether external release is allowed, applying differentiated DRM policies, applying differentiated DLP policies, masking areas that contain personal information, determining whether a document can be used for internal RAG, and determining whether a document can be used with generative AI. Grades and protected information are designed to serve as the basis for subsequent processing after security and utilization decisions are made.
However, the principle is that classification and actual measures must be separated. An example of an inappropriate structure would be automatically approving external release based only on AI’s O recommendation. The institution’s approval procedures must be applied separately, and actual security policies must also be applied separately.
The information provided by the C/S/O classification function includes the document grade, judgment status, applicable criteria, and original-text evidence. Judgment status items include automatic classification completed and staff review required. Presenting this information prevents the classification result from immediately leading to execution measures while allowing it to be used as a basis for subsequent judgment.
If necessary, it can be linked with Docu AD, an electronic document personal information masking solution. The purpose of the linkage is to connect follow-up processing such as masking at the output stage. The development direction is to build a structure that links the C/S/O classification function with subsequent processing.
Executive Vice President Kook said ACNS’s perspective on document AI is not a way of simply adding AI onto documents. He explained that ACNS applies AI with documents at the center.
Kook said ACNS focuses on this area because it recognizes two important factors in document AI. As the first factor, he pointed to implementing AI’s understanding of documents as closely as possible to the way humans understand them.
He explained that a document is not just a collection of text. A document consists of a title, body text, tables, attachments, document structure, and the context of creation, and these elements are interconnected. He also said that the same sentence can mean different things depending on its position in the document or the work status. He added that proper use of a document goes beyond the level of text extraction or reading only some sentences. For this reason, he said document AI needs the ability to understand the entire structure and context of a document at the same time.
As the second factor, Kook pointed to connecting the results of document understanding to actual work-processing methods. He explained that people do not stop at simply reading documents. After checking a document, they carry out tasks such as classification, approval, determining whether it can be disclosed, and assigning a security grade, and if needed, they protect certain information or pass it to another system.
The aim of document AI is to go beyond explaining document content and to support actual work processing. The reason ACNS has taken interest in this field is also rooted in this orientation of document AI toward supporting work processing.
Since its founding, ACNS has continuously developed and supplied electronic document-based solutions. Its scope of electronic document processing covers conversion, verification, viewing, and protection.
In this process, ACNS’s ongoing challenge has been understanding the structure of electronic documents and using information within documents accurately in work settings. ACNS said it sees AI not as an independent technology but as a new technical element of electronic document processing.
The company also said it believes that more important than competing on model performance is applying AI in ways suited to the structure and characteristics of electronic documents and linking that to actual document-work processing methods. It added that it aims not to add AI onto documents, but to apply it with documents at the center, and that this is its perspective on document AI.
The speaker said document AI is expected to evolve beyond simple Q&A and toward broader incorporation into the judgment processes of existing work systems. He explained that the early use of generative AI centered on answering human questions.
He then said that future functions of document AI will include reading documents, detecting what they contain, determining applicable criteria, suggesting management grades, and guiding policies to apply in the next system. He said the role of document AI is also changing from an AI that “summarizes and answers” to an AI that “supports judgment and connects processing.”
He explained that model size and type alone are insufficient to reach this stage. The factors that determine real-world application include the criteria assigned to AI, the level of analysis without omissions, the method for handling exceptions and conflicts, and the ability for humans to verify the judgment results.
He said that while the key question in corporate and institutional document AI adoption has so far been document search and explanation performance, it is now expanding to whether documents can be processed according to criteria. Accordingly, he said the focus of adoption is shifting from the ability to find and explain documents well to the ability to judge how to process documents according to set criteria and connect that judgment to the next steps.
N2SF-based C/S/O classification is a representative example of how document AI is changing. The system needed for document AI is not sufficiently covered by searching only for past nonpublic documents and similar documents; it must also analyze the entire current document, make judgments according to clear criteria and exceptions, and explain the basis for those judgments.
The first stage of document AI was document summarization and search, and the next stage is document judgment and linkage to actual work processing.
Source: IT DAILY · Kim Ho
Original: https://www.itdaily.kr/news/articleView.html?idxno=241898
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.