AI Bottlenecks Are Not Just a GPU Problem...Model, Chip, and Infrastructure Integration Is the Competitive Edge
TECHWORLD ·
✦ AI Summary
As AI services spread, securing high-performance models and GPUs alone is no longer enough to ensure competitiveness, and the ability to connect models, AI semiconductors, serving software, and applications, along with a balance between cost and performance, is becoming increasingly important.
At the panel talk in Seoul's COEX on the 29th, inference costs, infrastructure efficiency, and responses to diverse service environments amid the spread of agents were raised as major challenges.
The panelists highlighted the need for feedback structures linking model development and real-world services, operating standards for agent environments that run multiple models and tools together, and infrastructure optimization as hardware and workload combinations increase.
As AI services spread, securing high-performance models and GPUs alone is no longer enough to ensure competitiveness. As a result, the ability to connect models, AI semiconductors, serving software, and applications in ways suited to each use environment is becoming increasingly important, along with the need to secure a balance between cost and performance.
This trend was also raised during the panel talk, "Inference Optimization That Solves AI Bottlenecks," held on the 29th at COEX in Seoul as part of "lab | up > /conf/6." The session was moderated by No Jeong-seok, CEO of Bfactory, and panelists included Kim Jun-ki, CTO of Label Up; Kim Tae-ho, CTO of Nota; Lee Jin-won, CTO of HyperAccel; and Lee Hwal-seok, CTO of Upstage.
The panelists said AI use is expanding from the training stage into actual services, with inference costs and infrastructure efficiency emerging as major challenges. They also said that as agents spread, model-calling methods and service environments have become more diverse, limiting the effectiveness of improving only one side, either the model or the semiconductor.
Lee Jin-won, CTO of HyperAccel, said that as the inference market expands, cost efficiency will become the benchmark for judging semiconductor competitiveness. He explained that the market is shifting from training to inference, making cost reduction the biggest challenge. He also said that if cost efficiency improves, barriers to market entry could fall, and if barriers to market entry fall, overall demand could expand.
Lee said requirements differ by service. Some services need fast responses, while others require high-volume processing at low cost, making it difficult to improve all performance metrics at once. He cited email summarization as an example of a task that does not require an immediate response, and said a market that handles large volumes of requests based on low cost could grow.
In line with this trend, HyperAccel is using DDR-series memory instead of HBM. HyperAccel is pursuing an inference semiconductor strategy centered on cost and capacity.
Upstage emphasized the importance of creating feedback loops between model development and real-world services. Lee Hwal-seok, CTO of Upstage, said performance gains alone are not enough and pointed to the need to feed data and reactions generated during actual use back into model improvement. He also identified securing feedback from real users as the biggest challenge for model developers.
Lee Hwal-seok, CTO of Upstage, pointed to an agent-connected environment that goes beyond the B2B and B2C divide, saying there needs to be room to accumulate continuous experiments and feedback. Upstage said that rather than model performance itself, what matters is a structure that reconnects data and reactions obtained from actual services back into development.
At the same time, the panel said that as model performance becomes more advanced, the difficulty of training data and evaluation systems rises. Lee Hwal-seok, CTO of Upstage, noted that training models that surpass expert level requires the parallel use of synthetic data, reinforcement learning (RL), and evaluation systems. He added that there is a shortage of workers who possess all three capabilities: operating synthetic data, reinforcement learning (RL), and evaluation systems.
Kim Tae-ho, CTO of Nota, said the focus is shifting toward optimization. He explained that the center of optimization is moving from the model and kernel stages to the agent operation stage, whereas in the past the focus was on lightweighting a single model and improving serving efficiency. He also said that in an agent environment, multiple models and tools are used in parallel, making it necessary to review the entire execution flow.
With the spread of open source, lightweighting and serving technologies for individual models were said to have reached a certain level. Kim said lightweighting and serving technologies for single models have been leveled up through open source.
By contrast, in agent environments where multiple models are operated together, operating standards are still lacking. Kim said efficient operating methods for agent environments that run multiple models in combination have yet to be standardized.
In this situation, differences in data security, costs, and business characteristics by company are adding to the need for model-semiconductor optimization technology. It was also pointed out that technologies capable of responding to diverse environments are needed rather than a one-model, one-semiconductor customization approach.
At the infrastructure level, the growing combination of hardware and upper-layer workloads was identified as a new bottleneck. As accelerators such as GPUs and NPUs become more diverse and the number of models and applications running on them surges, the number of test cases for middleware is increasing. Kim Jun-ki, CTO of Label Up, said that testing various accelerator chip and application combinations consumes significant resources, and explained that when multiple workloads share a single GPU, actual resource usage must be measured precisely. He added that cooperation from chip makers and interface standardization are needed for this.
Corporate requirements are changing. In the past, the main challenge was managing GPU training resources, but recently demand has increased for training automation, model serving, AI transformation (AX), and agent operations.
In this regard, Kim said infrastructure deployment alone cannot meet customer needs. He added that it is necessary to respond to operating methods, applications, and model utilization as well.
Amid these changes, the panel talk projected that AI coding tools and model performance will improve rapidly. As a result, the center of corporate competitiveness may shift from the code itself to hands-on verification experience in real environments and domain knowledge.
Kim said AI can generate code quickly. However, he said long-term verification, maintenance experience in diverse corporate environments, and trust are difficult to replace in a short period of time.
Lee Hwal-seok said improvements in general-purpose model performance are advancing rapidly, led by big tech companies. At the same time, he added that there is room for differentiation in document formats by country and industry, as well as in legal domains and internal business procedures.
As competition in AI inference intensifies, the individual performance of models and semiconductors is becoming more important, while the importance of the ability to tailor actual service connections across each layer is also growing. Accordingly, connecting model development, semiconductors, infrastructure, and applications to secure a balance between cost and performance is emerging as a key challenge in AI service competition.
Source: TECHWORLD · Kim Seung-ki
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407510
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.