Crowdworks Wins AI Agent Safety and Reliability Verification Project
TECHWORLD ·
✦ AI Summary
Crowdworks said on the 3rd that it will participate in the Ministry of Science and ICT and the National Information Society Agency (NIA)-commissioned "AI Agent Safety and Reliability Verification System Support" project.
The project aims to verify the safety and reliability of AI agents during autonomous execution and to establish a system that can evaluate the appropriateness of judgment and tool calls while carrying out multi-step tasks such as reservations, payments, and dispatching.
Crowdworks will be responsible for building multi-step reasoning scenarios and a Korean-style verification dataset, and the project aims to complete the development and construction of the verification system by the end of this year.
Crowdworks said on the 3rd that it will participate in a government-commissioned AI agent verification project. The company announced its participation in the Ministry of Science and ICT and the National Information Society Agency (NIA)-commissioned "AI Agent Safety and Reliability Verification System Support" project.
The newly awarded project is designed to verify the safety and reliability of AI agents during autonomous execution. Its goal is to establish a verification system that can evaluate the appropriateness of judgment and tool calls while carrying out multi-step tasks such as reservations, payments, and dispatching.
Existing AI evaluations have focused on the accuracy of final answers. However, because AI agents independently formulate plans and complete tasks by calling external tools and APIs, there has been a growing need for evaluation criteria that can assess not only outcomes but also the validity of the execution process.
Crowdworks said it has become a participant in the government-commissioned "AI Agent Safety and Reliability Verification System Support" project.
The project is being carried out with the goal of reflecting international standards and establishing evaluation criteria suited to domestic service environments. The project process is intended to lead to private-sector AI reliability certification through efforts to promote real-world deployment.
Suresofttech is leading the project, with Crowdworks, KAIST, and the Korea Intelligent Information Society Agency participating in the consortium. Completion is scheduled for the end of this year, and the scope to be completed covers the development and construction of the verification system.
Crowdworks will be responsible for building multi-step reasoning scenarios and a Korean-style verification dataset. The scenario design will reflect the use of various MCP tools and external APIs, and the scenario conditions will include requirements for sequential or parallel calls, with the scenarios structured by difficulty level.
The verification dataset will be built to include data that can assess final answer accuracy, data that can evaluate the agent's judgment process, and data that can evaluate whether tools are called and how they are called. The dataset will reflect the domestic service environment and frequently changing dynamic data, and its scale will exceed 7,000 cases. The evaluation target is the agent's judgment ability in response to changing circumstances.
A Crowdworks official said that in the era of AI agents, the importance of a verification system for the safety and reliability of execution processes is expanding, in addition to model performance. The official added that Crowdworks has data-building and evaluation capabilities and expressed the company's commitment to contributing to the establishment of AI agent verification standards suited to domestic conditions.
Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=406493
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.