RebelUp Unveils “AI End-to-End” Roadmap, Zeroing In on Agent Infrastructure
IT DAILY · · 1 views
✦ AI Summary
RebelUp held the “6th RebelUp Conference” on the 29th and delivered the opening keynote.
CEO Shin Jeong-gyu explained last year’s achievements and development plans for the agent era under the theme “About End-to-End in the AI Industry Era.”
RebelUp introduced Backend.AI Core, Continuum, Backend.AI:GO, microVM-and-container integration, an on-premises agent operations platform, and diffusion language model and statistical model technologies.
As the center of AI technology shifts to “AI agents” that autonomously carry out work, compute demand is surging and infrastructure complexity is also increasing. RebelUp unveiled a roadmap that addresses this surging demand and complex infrastructure from a single perspective.
RebelUp held the “6th RebelUp Conference” on the 29th. On the day, RebelUp CEO Shin Jeong-gyu delivered the opening keynote.
The keynote was titled “About End-to-End in the AI Industry Era.” At the event, Shin Jeong-gyu introduced the company’s achievements over the past year and explained development plans aimed at the agent era.
The roadmap’s key rationale was presented as RebelUp’s core orchestrator, “Backend.AI Core.” Backend.AI Core has supported NVIDIA's latest rack-scale architecture since earlier this year.
This architecture is characterized by configuring 72 GPUs into a single NVLink domain. RebelUp said it has also completed a proof of concept with a Japanese company for the architecture.
While participating in Upstage’s model development process in March, the company reduced failure recovery time by about 50% and shortened the pretraining period from an expected roughly 120 days to about 70 days. In addition, GPU virtualization targets have expanded from NVIDIA to AMD and Intel, and alpha versions for domestic NPUs from FuriosaAI and Rebellions are also underway.
Continuum began as a fallback tool that switches to local operation when the cloud fails. It later expanded into a smart router that supports multi-cloud APIs and local serving, and it also serves as a meter by displaying team-by-team token usage and team-by-team savings.
CEO Shin introduced it as the first result of the measurable intelligence initiative proposed last year. The optimization benchmark is “Goodput,” which refers to the maximum number of requests that can be processed while meeting service objectives. The service objective items include time to first token and generation time per token, and a standalone version is scheduled for release next month.
RebelUp open-sourced its in-house inference engine last May to simplify the model serving environment for machine learning. Backend.AI:GO, which federates office PCs to create a single AI environment, is being prepared for version 2.0.
Regarding the combination of agent isolation solutions and microVMs, CEO Shin said there are bottlenecks and resource isolation issues when operating agents at scale. He cited an increase in the number of CPU cores in servers as the cause, saying that this sharply reduces memory bandwidth per core.
Shin said containers such as Kubernetes cannot isolate disk I/O and network bandwidth, so the load of one agent affects other agents. By contrast, VMs have the advantage of being easier to control, but he said it is difficult to run dozens of them on a single server.
Accordingly, RebelUp is developing technology that combines microVMs and containers in a two-stage structure. In this setup, the VM limits the total resources available to each agent, while containers inside the VM flexibly run sub-agents. The technology is compatible with Docker, and the target release timing is by the end of this year.
RebelUp is developing an on-premises agent operations platform that includes an agent gateway and agent-specific ID issuance functions. With only the platform's API, the company built “Space,” a Slack-like collaboration space. Space is designed so people and agents can work together in a single channel.
RebelUp determined that diffusion models are better suited for small-scale compute environments. Accordingly, it is developing and scaling up a diffusion language model, with a target of releasing it within this year. It also unveiled a Korean draft model for accelerating GLM 5.2 serving. The draft model uses the diffusion method.
For areas that require fast, intuitive judgment, RebelUp also introduced technology that uses statistical models instead of large neural networks. The technology’s real-data-based training time is under 1 minute. Based on Apple M-series CPUs, processing performance is about 75,000 requests per second, and the decision time per request is about 13 microseconds.
In this technology, most judgments end at the statistical model stage. Only highly difficult tasks are passed on to upper-level models and LLMs. Through this, RebelUp presented a flow in which immediate decisions are first handled by statistical models, while only difficult tasks are handed off to upper-level models and LLMs.
CEO Shin said that as the complexity of AI and the AI infrastructure supporting it increases, RebelUp’s mission becomes even clearer. He went on to stress the company’s determination to democratize intelligence for humanity by enabling anyone to measure intelligence and by pushing ahead with a direction that allows intelligence to be used whenever needed.
Source: IT DAILY · Kwon Young-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=241876
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.