AI

LabUp Moves Beyond GPU Management to an AI Operating System

TECHWORLD ·

Shin Jeong-gyu, CEO of Lablup. [Photo: Kim Seung-ki]

✦ AI Summary

LabUp is expanding its business scope beyond GPU and AI cluster management to include model training and inference, token usage management, and AI agent operations.

At the opening keynote of 'lab | up > /conf/6' held on the 29th at COEX in Seoul, CEO Shin Jeong-gyu said there is a need to connect the entire process from infrastructure to task execution into a single operating system.

LabUp is envisioning a structure that ties together training, inference, request management, and agent operations through 'Token Factory,' 'Continuum,' 'NeuVM,' an agent platform, and 'AI:GO.'

LabUp is expanding its business scope beyond GPU and AI cluster management to cover the full spectrum of AI use. Its previous scope was GPU and AI cluster management, while the expanded scope includes model training and inference, token usage management, and AI agent operations.

This expansion aligns with the stage in which AI adoption is moving beyond research and development into industry and services. LabUp plans to connect the entire process, from infrastructure to actual work execution, into a single operating system.

Shin Jeong-gyu, CEO of LabUp, delivered the opening keynote at the annual technology conference 'lab | up > /conf/6' held on the 29th at COEX in Seoul. The keynote was titled 'About End-to-End in the AI Industry Era.'

Shin said that there is a need to move beyond separately managing infrastructure building, model training, inference, and service operations, and instead connect the entire process. As a benchmark for this, he proposed 'tokens, power, and intelligence.' He explained that AI usage tokens need to be measured, the power required for computation needs to be measured, and the level of problems that can be solved with the same amount of tokens needs to be measured.

He stressed that, in AI use, expanding intelligence compression based on the same token benchmark is important. He also said it is necessary to calculate both the required level of intelligence and the amount of tokens invested in problem-solving at the same time.

He then compared AI infrastructure to a manufacturing plant. He explained that plant operations produce products, inspect them, then calculate quantity and cost before shipping them to their destination.

Shin said AI operations also need a corresponding system. He explained that such a system would train and serve models, measure usage, and supply the services needed.

LabUp said it is concretizing this system as 'Token Factory.' 'Token Factory' is a structure in which computing resources, model execution, request distribution, token usage, and agent services are connected step by step.

The base platform for this structure is the AI infrastructure platform 'Backend.AI.' 'Backend.AI' supports management of AI accelerator resources such as GPU and NPU, and is expanding its support to large-model training and inference as well as heterogeneous cluster operations.

LabUp has completed support for NVIDIA GB200 NVL72 in line with the latest infrastructure changes and is also preparing to support the next-generation Vera Rubin architecture. The GB200 NVL72 structure connects 72 Blackwell GPUs in a single NVLink domain. The GB200 NVL72 requirements call for an approach different from traditional server-level resource management.

LabUp is focusing on improving fault recovery speed in large-scale training environments. During long-term operation of hundreds of GPUs, it is difficult to avoid failures in some equipment, and recovery time directly affects the overall training period.

LabUp carried out training for Upstage's 'Solar Open 100B.' In the process, it combined LabUp's infrastructure automation with Upstage's training code optimization, reducing the expected training period from about 120 days to 66 days.

LabUp has expanded the scope of its monitoring and control. Based on 'all-smi,' it supports checking GPU, NPU, and CPU status, and it also supports integrated monitoring of power and heat information. It also supports upper-layer integrated operations in Kubernetes and Slurm environments.

LabUp is broadening the range of accelerators it can support beyond NVIDIA to include AMD, Intel, and others. Support for AI accelerators is also being expanded step by step, and the targets include FuriosaAI and Rebellions.

At the event, LabUp unveiled 'Continuum.' Continuum has functions for managing model requests and managing token usage.

Continuum consists of a request-distribution router and a hub for checking overall usage status. The router connects external AI services such as OpenAI, Anthropic, Amazon Bedrock, and Google Gemini with enterprise in-house models.

Integration with in-house serving environments is also possible. Supported environments include vLLM, SGLang, MLXCell, Ollama, and llama.cpp.

Continuum is designed to distribute requests based on response speed, cost, and service status. If an external cloud API fails, requests can be switched to in-house models, and it is also possible to choose models in line with service-level objectives.

Continuum plays a role in quantifying AI usage by organization. It can show token consumption by team and service, and it can also show costs by team and service. In addition, it can track cost changes after model changes and after cache application.

Shin emphasized the need for quantification, comparing token usage to electricity usage. He explained that it is necessary to verify the scale of savings and to check token usage by team.

LabUp proposed using Continuum alongside NVIDIA inference software Dynamo. Dynamo is responsible for inference optimization inside the cluster, while Continuum handles request distribution between external clouds and in-house models.

The Continuum router is also planned to be applied to the local AI platform AI:GO in the future. It provides a standalone router and an enterprise hub, and based on this supports an environment that integrates cloud and in-house models.

As agents spread, CPU and I/O also become management targets.

With the spread of AI agents, the scope of infrastructure management is expanding. GPU is the resource responsible for model inference, but agents also perform file reading, code building, testing, and external information gathering. The resources used in this process include CPU, memory, storage, and networks.

When multiple agents work simultaneously, resource interference can occur. A particular agent's disk input/output (I/O) and network usage can affect other tasks. Existing container environments have limitations in that they make it difficult to finely isolate CPU, memory, storage, and network resources.

As a solution, LabUp is developing 'NeuVM.' 'NeuVM' is a technology that combines micro virtual machines (VMs) and containers to separate resources by agent while lowering virtualization overhead compared with ordinary VMs. The purpose of 'NeuVM' is resource isolation by agent, and its characteristic is reduced virtualization overhead compared with ordinary VMs.

'NeuVM' operates by deploying microVMs on an agent-by-agent basis and running containers inside them. LabUp said it is preparing to unveil a beta version by the end of the year.

LabUp has developed a platform for direct AI agent operations. The platform is designed to connect agent creation, execution, external tool invocation, permission management, and results evaluation on the infrastructure side into a single workflow.

Its components are divided into 'Agent Factory,' which handles creation and deployment; 'Agent Runtime,' which is responsible for execution stability and isolation; and 'Agent Gateway,' which centralizes external calls and control.

'Agent Factory' supports the creation of agents based on natural-language task descriptions. It also supports execution in Kubernetes environments and parallel operation of agents built with external frameworks.

'Agent Runtime' supports parallel processing of multiple tasks. It also preserves sessions and state after failures and restarts. Generated code runs in a sandbox based on containers and microVMs, which is intended to separate work areas.

'Agent Gateway' is responsible for external service calls and permission management. It also centrally controls model requests, policies, guardrails, human approval, team budgets, and audit logs.

LabUp presented a system that repeats the cycle from agent creation to operation and evaluation. The system includes post-operation assessment items such as execution results, tool usage, safety, and cost evaluation, as well as a revalidation process for improvement measures.

LabUp said it uses the agent platform API. It also implemented 'Space,' a space where people and multiple AI agents exchange tasks, and presented 'Space' as an example of an applied collaboration environment based on its own platform API.

The company also announced additional features planned for the future. Planned additions include 'Tool Studio' and 'Warden.' 'Tool Studio' is a function that connects internal enterprise systems and agents without code, while 'Warden' is a function that decides whether to run automatically or require human approval based on task risk level.

LabUp's end-to-end strategy extends beyond large data centers to local AI environments as well. The company is not limiting the scope of this strategy to data centers, but is extending it to local AI environments as well.

As an example, it pointed to 'AI:GO.' 'AI:GO' is a desktop AI platform officially launched in January this year that can run AI models on individual PCs, and it also has the ability to connect the resources of multiple devices.

Multiple PCs can be connected through 'GO Mesh,' and distributed model execution across different devices is also possible. Its agent features provide file editing, web search, and code execution, supporting agents in carrying out actual work.

LabUp plans to operate training and inference infrastructure with Backend.AI, and entrust model request management and token usage management to Continuum. It is also envisioning a structure that connects actual services and work through the agent platform and AI:GO.

As AI spreads into industrial sites and enterprise work, the objects of management are expanding to GPU and model performance, token costs, CPU resources, network resources, agent permissions, and agent behavior. Accordingly, LabUp is making the single-system operation of infrastructure, models, and agents its core focus, with the goal of reducing operational complexity.

Shin said that infrastructure remains necessary even after AI use begins. He also said the company will push ahead with a direction that intuitively simplifies complex AI infrastructure problems, and that it intends to make intelligence measurable and available for use whenever needed.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407488

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.