AI

Xenon Adds Computer Control Capabilities to 397 Billion-Parameter Mega Model, Unveils “Hunmin VLM”

TECHWORLD ·

[사진=제논]

✦ AI Summary

Xenon made “Hunmin VLM 397B” open source on the 18th.

The model is based on “Qwen3.5-397B,” with 397 billion parameters, and strengthened screen recognition and computer control capabilities.

Xenon said it confirmed performance gains in benchmarks and real-task execution, while keeping the decline in Korean-language performance within 2 points.

Xenon made “Hunmin VLM 397B” open source on the 18th. “Hunmin VLM 397B” is a vision-language model (VLM) capable of carrying out tasks through computer screen recognition and direct control.

“Hunmin VLM 397B” is based on “Qwen3.5-397B,” a model with 397 billion parameters. Xenon said it strengthened screen recognition and computer control capabilities compared with the base model while preserving existing language performance, and focused on adding execution capabilities for real-world task completion.

To that end, it reinforced grounding functions for detecting objects to manipulate on the screen, such as buttons and input fields. It also improved computer control performance so the model can move between operating systems and applications to complete given tasks.

Xenon unveiled “Hunmin 32B” last year and “Hunmin VLM 235B” earlier this year. With this model, Xenon expanded the mega-model tier of the Hunmin series and is advancing the series in the direction of preserving language performance while adding execution capabilities such as computer use.

Xenon said it verified the improvements in Hunmin VLM 397B through screen recognition and computer control benchmarks. According to Xenon, Hunmin VLM 397B outperformed both Qwen3.5-397B and Qwen-CUA across five benchmarks for identifying the location of objects to manipulate. Its score on ScreenSpot-Pro was an accuracy of 75.6.

In the ScreenSpot-Pro results, Hunmin VLM 397B ranked second on the official Hugging Face leaderboard. At the time, there were 48 models listed on the official Hugging Face leaderboard. The company said Hunmin VLM 397B was the only model among the 48 developed by a South Korean company.

Its real-task performance also improved. On OSWorld, it scored 70.5 under conditions requiring it to complete 360 tasks on a Linux desktop, outperforming the base model by 22.3 points. On WindowAgentArena, it also outperformed the base model by 9.1 points in a Windows environment.

Xenon said it strengthened computer control capabilities while minimizing any decline in Korean-language performance. The Korean-language benchmarks covered eight tests, including KMMLU and K-MMBench. The Korean-language benchmark results kept the performance gap within 2 points compared with the base model.

The model incorporates post-training techniques for advancing mega models with limited computing resources. Xenon transferred the capabilities of a public computer control model using a low-rank approach, and carried out supervised fine-tuning (SFT) and reinforcement learning (RL) in FP8 quantized form. For total post-training resources, it used eight NVIDIA B200 GPUs.

Xenon released a technical report in July on a training method for a 743 billion-parameter model in a small-GPU environment and applied that methodology to developing Hunmin VLM 397B. This release focused on explaining how it advanced a mega model through post-training even with limited computing resources, and the configuration behind it.

Xenon plans to use the screen recognition and computer control technologies secured in Hunmin VLM 397B to further advance its execution-oriented AI, “OneAgent.” It also plans to inherit the computer use and browser use capabilities strengthened in Hunmin VLM 235B. Based on this, Xenon intends to expand the range of tasks that AI can directly perform in enterprise work.

This release was structured to disclose not only the model itself but also the methods used to develop and validate it. The company disclosed its development and evaluation methodology, applied the Apache 2.0 license, and provided the original model. It also offered an FP8 version at about half the size. In addition, it re-evaluated comparison models in the same environment and disclosed the re-evaluation conditions and methodology documents, presenting an effort to ensure fairness in performance comparisons.

Myung Dae-woo, vice president and CTO at Xenon, said the center of competition in AI models is shifting from increasing parameters to adding capabilities efficiently and advancing them to real-world use levels. He said Hunmin VLM 397B strengthened computer control capabilities with limited computing resources while maintaining existing Korean-language performance.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407133

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.