Kakao Predicts Large AI Learning Rates Through Small-Scale Experiments, Applies It to Kanana
IT DAILY ·
✦ AI Summary
Kakao said it has developed a technology that predicts appropriate learning rates for large AI models based on experiments with small models, and applied it to its in-house language model, Kanana.
The research also disclosed a low-computation method for finding learning rates for pretraining MoE models and a lightweighting method that reduces the number of parameters by changing the input-output structure.
Kakao said it plans to present the findings at COLM 2026 and TokShop, and also said Kanana-2.6-155b-a17b trained stably and may reduce memory burdens in on-device environments.
Kakao said it has developed a technology that predicts appropriate learning rates for large AI models based on experiments with small models, reducing the burden of repeatedly testing training settings for large-scale AI models. The technology has been applied to its in-house language model, Kanana.
Kakao's research on training optimization, which it disclosed, focused on finding the learning rate for pretraining mixture-of-experts (MoE) models with less computation. A learning rate refers to the extent to which internal parameters are adjusted during training. If the learning rate is too high, training may become unstable; if it is too low, training may slow down.
Kakao also disclosed research on lightweighting that reduces the number of parameters by changing the input-output structure. The study was presented in two directions: improving training efficiency and reducing the number of parameters.
Kakao plans to present these training optimization and model lightweighting results on the 8th at COLM 2026. COLM 2026 will be held from the 6th to the 9th local time in San Francisco, the United States.
The main conference paper is titled "Let's Scale Step by Step." The paper proposed an approach to reduce the burden of repeatedly finding the learning rate whenever the model size and data volume change in large-scale MoE training. It introduced a method for transferring settings from a small model to a large model and for analyzing changes by training data volume, with the goal of predicting the learning rate needed for long-term training.
The MoE structure computes by selecting only some modules from among multiple expert modules for each input. This structure can expand the overall model scale and has the characteristic of limiting the amount of computation used for each input. However, it also has the limitation of making training settings more complex through expert selection and allocation.
For this reason, the appropriate learning rate also changes depending on model scale and training volume. As a result, repeated experiments to find the right learning rate require substantial computation.
Kakao applied the μ-parameterization (μP) technique to the MoE structure. The μ-parameterization (μP) technique transfers the training settings of a small model to a large model. Kakao further analyzed changes in the appropriate learning rate by training data volume through small-scale experiments, and used the results to predict the learning rate needed for the target training volume.
The paper explained that the total computation used in small-scale experiments to find the learning rate before full training is about 1% of the target model's full training computation. This 1% figure is based on a comparison of the computation spent on setting exploration before full training.
Kakao said it applied the method to pretraining for "Kanana-2.6-155b-a17b." Kakao said "Kanana-2.6-155b-a17b" was trained stably on 10 trillion tokens.
"Kanana-2.6-155b-a17b" has a structure with 155 billion total parameters. Of those, 17 billion are parameters used for input-specific computation.
Kakao also disclosed a separate study, "BBT: BPE-Guided Byte Transformer." The study aims to reduce the number of parameters by changing the text input-output structure while preserving text segmentation. "BBT: BPE-Guided Byte Transformer" was discussed at the tokenization workshop TokShop on the 9th, local time.
A standard BPE model splits text into tokens at the word or word-piece level. It also stores information corresponding to each token in the input and output layers. As a result, when the vocabulary grows, the parameters related to the input and output layers increase.
BBT preserves BPE-based segmentation boundaries while representing and generating each segment at the byte level, leveraging the advantages of processing short segments. It also replaces input and output layers whose size scales with the vocabulary.
The researchers conducted comparative experiments on English text and multilingual experiments. In the English text comparative experiments, the number of parameters decreased by 20.2% to 40.4% depending on model scale, while bytes per byte (BPB), a language modeling performance metric, fell by 2.7% to 6.6%. A decline in BPB indicates improved text prediction. The amount of computation for input processing and output calculation was broadly similar to that of the comparison model, and the drop in performance under input perturbations such as character deletion and space insertion was smaller than in the comparison model. In the multilingual experiments, prediction performance also improved for languages that were not included in training but shared a writing system with the training language.
Kakao said these research characteristics are expected to ease memory burdens when running AI in on-device environments.
Kakao said the studies are meaningful in that they present practical solutions aimed at both maintaining and improving AI model performance while reducing training and operating costs.
Kakao plans to apply the research methods to a variety of model architectures. It also plans follow-up research in multimodal environments that process multiple data types such as text and images at the same time, and will continue follow-up research in multilingual environments as well.
Source: IT DAILY · Kim Byeong-ju
Original: https://www.itdaily.kr/news/articleView.html?idxno=242080
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.