Nota Unveils Multimodal AI Compression Technology Selected for NeurIPS 2026 Main Track
TECHWORLD ·
✦ AI Summary
Nota announced the development of a multimodal AI model optimization technology.
The technology works by separately evaluating the importance of expert modules for text and image processing in MoE-based multimodal models and removing relatively low-importance modules.
In a model with about 31 billion parameters, it retained 98% of original performance with 25% of expert modules removed and 91% with 50% removed, while reducing GPU memory usage from 58GB to 31GB.
Nota announced on the 1st that it had developed an optimization technology for multimodal AI models. The technology is designed to reduce GPU memory usage while minimizing performance degradation. The related paper, "Modality-Aware Expert Pruning for MoE-Based Multimodal Large Language Models," has been accepted to the main track of the AI conference NeurIPS 2026.
The technology evaluates the importance of expert modules used for text and image processing separately in MoE-based multimodal models, then removes relatively low-importance modules. Nota's approach focuses on determining the necessity of expert modules for text and images in MoE-based multimodal models and pruning the less important ones.
In an MoE architecture, only some of the multiple expert modules needed for a given input are selected for computation. However, even though only a subset of experts participate in actual computation, memory is still required to store all modules. As model sizes grow, GPU memory pressure increases, and this technology aims to reduce that burden.
Expert pruning reduces overhead by removing low-utilization modules. However, if pruning methods for conventional language models are applied directly to multimodal models, there is a risk that experts for image understanding may also be removed. In that case, performance could decline in tasks where visual information is important, such as chart analysis.
Nota therefore applied a method that evaluates expert importance separately based on text and image criteria. It preserved experts important for text processing while setting a criterion to remove experts whose importance was relatively low for images. The approach takes both text and image performance into account at the same time.
Whereas conventional methods removed the same number of experts at each layer, Nota changed the approach by comparing expert importance across the entire model and selecting which ones to remove. It also applied a principle of leaving more experts in the layers that need them. This was an adjustment aimed at reducing performance degradation.
Nota evaluated the performance of a multimodal model with about 31 billion parameters. As a result, when 25% of expert modules were removed, it retained 98% of the original performance, and when 50% were removed, it achieved 91% of the original performance. The company said that even with 50% removal, the model recorded 91% of the original performance.
GPU memory usage for the model fell from 58GB to 31GB, a reduction of about 47%. As a result, the model became capable of running on a single 40GB memory GPU.
This compression can be applied in a single pruning process without separate additional training. Users can set the expert removal ratio, and optimization can be carried out even after a new foundation model is unveiled, without additional retraining.
Nota plans to reflect the technology in its AI optimization platform NetsPresso. Support will expand from text-centered LLMs to VLMs, and the company also plans to extend its future application scope to other input types such as video and audio.
The company has accumulated experience applying expert pruning to large-scale models such as "Kimi K3" and "Solar Open 2." This year, it has presented optimization-related research results at NeurIPS, following ICLR, ICML, and EMNLP.
Chae Myung-soo, CEO of Nota, said that AI model development is shifting from text-centered LLMs to multimodal AI that can understand images, video, and audio simultaneously, and added that GPU memory burden increases as multimodal AI advances. He said the company plans to expand the optimization capabilities it has built up in large language models to multimodal AI, and that it will further develop related technologies with the goal of helping companies use their GPU and AI infrastructure more efficiently.
Source: TECHWORLD · Kim Seung-ki
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407674
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.