AI

Tnaps Says Accuracy Alone Is Not Enough in Financial AI Validation

IT DAILY ·

강민승 티냅스 대표가 지난 26일 서울 강남구 한 공유오피스에서 인터뷰하고 있다. (사진: 스노우플레이크)

✦ AI Summary

Kang Min-seung, CEO of Tnaps, said that apart from AI performance improvements, it is important for companies to define the scope of work they assign to AI.

He explained that in finance, even small errors can be fatal, so fact-checking and compliance with internal policies and operational standards must be verified together.

Tnaps controls the scope of AI delegation with rules, domain-specific AI, human review, and its 'domain judge' technology.

Separately from the rapid improvement in AI performance, it has been pointed out that when companies deploy AI in real work, setting the scope of what can be delegated is important. In an interview with IT Daily on financial AI validation methods, Kang Min-seung, CEO of Tnaps, said during a meeting on the 26th at a shared office in Gangnam-gu, Seoul, that while AI intelligence continues to advance, the range of tasks companies can entrust to AI remains limited. He added that there is a need to define the boundary between the scope of AI automation and the point at which work is handed over to people.

In particular, Kang emphasized that in finance, even a small error can have a fatal impact on business outcomes, making it important to set standards for the scope of AI judgment. This shows that improving AI performance itself is not the same as determining how much real work can be delegated.

Tnaps is an AI reliability infrastructure startup founded last year. The company is developing technology that verifies the answers and judgments of enterprise AI and blocks or controls problematic AI outputs from being reflected in actual operations. Its main focus is the financial sector, which it chose because of the complexity of regulation and operational standards.

The photo shows Tnaps CEO Kang Min-seung being interviewed on the 26th at a shared office in Gangnam-gu, Seoul. The photo is courtesy of Snowflake.

The direction of AI advancement is shifting from chatbots to agents. As agent capabilities expand, AI is taking on not only answer generation but also judgment and execution, and the scope of reliability that companies must secure is also expanding accordingly.

Kang explained that while AI chatbots in the past simply provided answers, the age of agents expands the range of decision-making. He said the ripple effects in the event of an accident also become larger.

Kang said agents invoke a variety of tools. He added that multiple tool calls in the execution and decision-making process can reduce reliability.

For that reason, it has been pointed out that attaching sources for answers alone is insufficient as a way to improve agent reliability. There can be inappropriate retrieval cases, such as materials that do not match the question, similar but different products, or outdated materials. Even if the retrieved material is factual, it is hard to regard it as an appropriate answer if it is not relevant to the question.

Kang said that in finance, whether an answer is backed by evidence and whether it can be provided to a customer are separate matters, adding that answer validation in finance should consider fact-checking, alignment with internal policy, and compliance with operational standards together.

He also explained that while a company can predefine procedures for an agent’s work, predefining the process alone cannot control the outcome. He said AI decision-making is unavoidable even in the control process, and if errors occur in that decision-making, controlling the path cannot guarantee the correct answer.

Kang said that for this reason, it is difficult to decide whether to use AI based only on accuracy figures. He said it is ambiguous to judge whether AI with 95% accuracy can be used, and also ambiguous to judge whether AI with 99% accuracy can be used, adding that even a 1% to 5% mix of wrong answers can distort business results.

In the end, he meant that accuracy alone is not enough and that the impact of wrong answers on work must also be considered. He also said that the boundary between AI and human work needs to be set.

Tnaps addresses the issue of setting the scope of AI delegation by applying rules, domain-specific AI, and human review step by step. It applies predefined rules to areas that can be handled under clear standards, while domain-specific AI reviews areas that are difficult to judge by rules alone. Matters that AI also finds difficult to judge are handed over to frontline staff.

Tnaps uses its in-house 'domain judge' technology in this process. 'Domain judge' learns finance terminology and judgment criteria by task, and is responsible for verifying whether AI answers are appropriate for the relevant work. In this flow, 'domain judge' plays the role of learning standards in the financial context and determining whether AI judgments fit the business task.

Kang compared 'domain judge' to 'AI's compliance officer.' He explained that a compliance officer may not be an expert in every task, but has the domain knowledge needed to tell right from wrong. He also said the company collects AI error cases in finance, analyzes the causes of those errors by type, and repeatedly corrects areas where mistakes occur.

Validation is necessary for AI to be used in real operations. If staff must recheck everything after AI produces results, the automation effect is limited. Kang said that if human recheck is assumed, questions arise about the need to use AI, and that leaving 1% to 2% uncertainty becomes a bottleneck to wider adoption.

For these reasons, Tnaps is discussing technology deployment in the financial sector, with corporate lending as the main area of discussion. Compared with household loans for individuals, corporate lending involves many review documents, many judgment factors, and complex work that includes corporate analysis and report writing. As a result, demand for automation in corporate lending was presented as significant.

Tnaps is also expanding deployment discussions to retirement pensions, foreign exchange, and deposits. In addition, Tnaps is discussing cooperation with KB Kookmin Bank on applying AI reliability technology.

Tnaps designed an AI reliability validation system based on Snowflake. The photo is courtesy of Snowflake, and Kang Min-seung is the CEO of Tnaps.

In AI operations, even after the scope of AI delegation is set, stable operation becomes difficult if judgment criteria change every time the model changes. Kang Min-seung said that continuous model changes are inevitable.

Kang Min-seung explained that a company’s answers and standards should not change every time the model changes. He added that, assuming model changes, permanent accumulation of corporate assets is necessary.

In this context, the importance of a data management foundation becomes more pronounced. Kang Min-seung said data is needed to set AI boundaries and determine whether answers meet the standards, and explained that the level of management within the governance system for that data is important.

The Tnaps-Snowflake collaboration emerged from the same recognition of the problem. Tnaps combined its validation technology with Snowflake’s data and AI environment.

This structure is designed to verify AI answers and judgments. It also blocks or rewrites results deemed problematic.

Audit logs and frontline feedback accumulated during the validation process are fed back into the data. Kang said that when data changes or retraining becomes necessary, audit logs are re-entered into Snowflake data.

Kang said that as this circular structure operates, it becomes possible to continuously accumulate corporate assets in the form of data. He explained that records generated during the validation stage and frontline opinions are fed back into the data, creating a cycle within the same framework.

Kang highlighted Snowflake’s ease of deployment and integration as an advantage. He explained that understanding and customizing the different system environments of each financial institution takes a great deal of time, but Snowflake has a consistent AI layer and data layer structure, making it easy to plug in technology within a standardized framework.

Tnaps plans to expand its business into overseas markets after building technology and references in Korea’s financial sector. The largest overseas market it is evaluating is the United States, and Tnaps is seeking business opportunities in Japan and Singapore, while also recently discussing entry into the Middle East.

Tnaps is leveraging its Snowflake collaboration in the process of global expansion, and that collaboration is serving as a channel to expand contact points with overseas customers. Tnaps was selected as a global top 10 finalist in the '2026 Snowflake Startup Challenge' and was the only Asian company to enter the global top 10.

Kang said that after being selected, inquiries from global venture capitalists (VC) and customers increased, and the level of interest was clearly noticeable. However, he said additional proof of performance is needed and more references must be accumulated, adding that expansion could accelerate if the company is connected with customers already using Snowflake.

Source: IT DAILY · Yang Seung-gab
Original: https://www.itdaily.kr/news/articleView.html?idxno=241278

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.