AI

Select Star Has 5 AI Reliability Papers Accepted to EMNLP 2026

TECHWORLD ·

[Photo: SelectStar]

✦ AI Summary

Five papers on AI reliability and safety involving Select Star researchers were accepted to EMNLP 2026.

The accepted papers consist of 1 main conference paper, 1 Findings paper, and 3 workshop papers.

The research expanded into human-LLM response alignment, guard model confidence, and safety evaluations across financial and cultural contexts.

Five papers on AI reliability and safety involving Select Star researchers have been accepted to EMNLP 2026. Select Star said on the 28th that a total of 5 papers involving its researchers were accepted.

The accepted studies broadened their scope to human-LLM response alignment, the confidence of guard model judgments, and safety evaluations across financial and cultural contexts.

The accepted papers consist of 1 main conference paper, 1 Findings paper, and 3 workshop papers. EMNLP is a major international conference in natural language processing organized by the Association for Computational Linguistics (ACL).

This year's event will be held next month from the 24th to the 29th in Budapest, Hungary. Select Star said about 18,000 papers were submitted to the main conference this year. The acceptance rate for the main conference is 15.4%.

"When Accuracy Aligns and Fails: Diagnosing Human–LLM Response Alignment at Population Scale," with Select Star researcher Choi Hye-ji as a co-first author, was accepted to the main conference. The study went beyond simply analyzing whether AI answers were right or wrong, and focused on whether it also chooses human-like wrong answers when it gets things wrong.

The researchers compared the distribution of correct and incorrect choices made by people and the response distribution of the LLM on standardized test items. Based on this, they proposed a framework for quantitatively diagnosing the degree to which AI demonstrates human-like success and failure, evaluating not only correct and incorrect outcomes but also how humans and LLMs succeed and fail.

The core of this evaluation framework is that it adds human response similarity to conventional accuracy-centered evaluation. Even when both a person and an AI answer incorrectly on the same item, the human's main wrong answer choice and the AI's wrong answer choice may differ, and the evaluation makes it possible to check the extent to which the model reflects the judgment patterns of specific groups.

Select Star plans to combine the research results with its 19 AI safety taxonomy and harmfulness evaluation technologies. It is also considering expanding the work into an AI safety benchmark that reflects users' risk perceptions and cultural characteristics by country.

"Guard Models Are Overconfident Where Base Models Are Uncertain" was accepted to Findings. The paper analyzes the confidence levels of guard models that determine whether harmful or dangerous requests should be blocked. The findings showed that some guard models display high confidence even when the base model's judgment is uncertain. It was also analyzed that some guard models may make incorrect judgments despite high confidence.

This study revealed a gap between guard models' confidence and their actual accuracy. Based on this, it was suggested that AI safety evaluations should examine not only simple accuracy but also whether a model's confidence level aligns with its actual judgment accuracy.

The key issue connected to real-world guardrail operations is that if a benign request is mistakenly judged harmful, normal inputs can also be blocked. As a result, service usability may decline.

As a follow-up, Select Star plans to push ahead with research-based technological development. The technology reflects both the model's judgment results and the possibility of misjudgment, and it will be applied to adjusting the blocking standards of sLM-based guardrails. Researcher Hong Jong-hyeon participated as first author, researcher Jeong Min-jae as co-author, and Kim Min-woo, leader of the AI Safety team, as corresponding author.

In the same vein, 3 joint studies on AI safety across financial and cultural contexts were accepted to workshops. Kim Min-woo, leader of the AI Safety team, participated as corresponding author in all 3. Of these, 2 joint studies with the Korea Financial Security Institute were accepted to FinNLP, a workshop specializing in financial NLP, and 1 joint study with Ewha Womans University was accepted to PANDORA, which focuses on responsible AI alignment.

"Regression-Aware Statistically Gated Policy Updating for Korean Financial LLM Input Guardrails" was accepted to FinNLP. The study targeted red-team attacks in the financial sector and proposed iteratively improving defense policies by reflecting previously unblocked cases in training. The policy update process was also designed to automatically verify whether existing defense performance had deteriorated.

"Hidden Below the Landing Page: Tracking Linked Pages and Relation-Centered DOM Summaries for Financial Phishing Detection" was also accepted to the same workshop. The study expanded the analysis scope beyond the first screen of phishing sites to include linked pages, and incorporated features that jointly analyze relationships between pages. Through this, it addressed LLM identification of actual phishing behaviors such as entering and transmitting personal information.

The PANDORA workshop paper "LatAm-RT: Toward Culturally Adaptive Red Teaming for AI Safety in Latin America" took Mexico, El Salvador, Costa Rica, and Chile as its research scope and reflected each country's language, institutions, and social context. The study was conducted by localizing red-team attack data and analyzing differences in safety evaluation by country and culture. In doing so, it went beyond simply translating safety data and reflected cultural context itself in the evaluation data.

The analysis found that the same AI model can produce different safety evaluation results depending on the country and culture. This approach led not only to acceptance at the PANDORA workshop but also to selection for an oral presentation.

Select Star plans to apply these research achievements to the Datumo Platform. The Datumo Platform currently provides AI reliability evaluation consulting, automated red teaming, and sLM-based guardrail technology. The company plans to upgrade its evaluation frameworks by country and industry by reflecting the research findings.

Kim Se-yup, CEO of Select Star, said the acceptance of the EMNLP papers signifies that the company's AI reliability evaluation and safety research capabilities have been recognized on an international academic stage. He added that he plans to link the research results to evaluation, red teaming, and guardrail technologies, with the goal of developing them into an AI reliability verification framework that can be used in real-world enterprise service environments.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407422

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.