AI

SelectStar Has 5 AI Reliability and Safety Papers Accepted to NLP Society's EMNLP 2026

IT DAILY ·

[Photo: SelectStar]

✦ AI Summary

SelectStar announced on the 28th that 5 AI reliability and safety papers involving its researchers had been accepted to EMNLP 2026.

The accepted papers consist of 1 main conference paper, 1 Findings paper, and 3 workshop papers.

The main conference paper presented an evaluation framework that compares human correct and incorrect choice distributions with LLM response distributions to quantitatively diagnose AI's success and failure similarities to humans.

SelectStar announced on the 28th that 5 AI reliability and safety papers involving its researchers had been accepted to EMNLP 2026.

The accepted papers consist of 1 main conference paper, 1 Findings paper, and 3 workshop papers. About 18,000 papers were submitted to this year's EMNLP main conference, and the acceptance rate was 15.4%.

The main conference acceptance was "When Accuracy Aligns and Fails: Diagnosing Human–LLM Response Alignment at Population Scale." The study addressed a topic that expands AI evaluation methods.

While existing evaluation methods have centered on whether AI arrives at the correct answer, the study examined whether it chooses human-like wrong answers in failure cases. SelectStar researcher Choi Hye-ji participated as a co-first author on the paper.

In the study, the researchers compared the distribution of human correct and incorrect choices for standardized test questions with the distribution of LLM responses, and presented an evaluation framework that quantitatively diagnoses AI's success and failure similarities to humans. SelectStar said the framework can assess both a model's accuracy and human likeness to identify how much it reflects the judgment patterns of specific groups and to detect cases in which errors occur in ways different from those of people.

SelectStar said it plans to combine this research with its 19 types of AI safety classification system and harmfulness evaluation technology to expand it into an AI safety benchmark that reflects users' risk perceptions and cultural characteristics by country.

It also mentioned the Findings paper, "Guard Models Are Overconfident Where Base Models Are Uncertain." The study analyzed the confidence of guard models that determine harmful or risky requests, and examined how they evaluate the likelihood that their judgments are correct when deciding whether a request is harmful or safe. The results showed that some guard models may make wrong decisions while showing high confidence even when the base model's judgment is uncertain.

SelectStar said it plans to expand the technology based on this research. The expansion direction is to reflect both model judgment results and the risk of misjudgment, and the application target is guardrails based on customer-environment-customized small language models (sLM). The adjustment target is the guardrail blocking criterion.

Kim Se-yop, CEO of SelectStar, said the significance of the EMNLP paper acceptances lies in the recognition of SelectStar's capabilities in AI reliability evaluation and safety research on the international academic stage. He added that the company plans to link the research outcomes with SelectStar's evaluation, red teaming, and guardrail technologies, with the goal of developing them into an AI reliability verification system that can be used in real enterprise service environments.

Source: IT DAILY · Yang Seung-gab
Original: https://www.itdaily.kr/news/articleView.html?idxno=241848

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.