[Edge Infrastructure 1] From Edge to AIDC: Expanding the AI Infrastructure Ecosystem
IT DAILY ·
✦ AI Summary
Agentic AI repeatedly sets goals, plans, calls tools, and checks results, consuming large amounts of tokens, so companies are reviewing local inference and hybrid AI architectures instead of relying on cloud APIs.
In an Nvidia survey, 64% of respondents said they had adopted an AI operating environment inside their organizations, and in a Dell-IDC survey, 89% said AI features were a key criterion when replacing PCs.
Desk-side and edge infrastructure are tied to issues of cost, Data Sovereignty, Latency, and on-site stability, and the industry is emphasizing CPU, GPU, NPU, Unified Memory, and an open software ecosystem.
With the emergence of agentic AI, generative AI has moved beyond the stage of a question-and-answer support tool. Agentic AI now goes beyond chatbots to lead complex workflows, set goals, and autonomously execute workflows. As a result, changes are under way in enterprise computing environments, and the expansion of the AI infrastructure ecosystem from the edge to AIDC has emerged as the topic of this article.
The center of gravity of traditional enterprise AI was in-house on-premises servers and public cloud data centers. But the enterprise infrastructure landscape is shifting to a locally based distributed computing model, and the current trend is to move computing resources forward to desk-side and industrial edge locations. This is intended to respond to the actual points where data is generated and practical decisions are made, and the shift toward infrastructure distribution is accelerating.
Companies are moving routine, high-frequency inference workloads to the edge, citing the unpredictability of cloud API token costs and the risk of intellectual property leakage. The goal of this response is to optimize TCO. In line with this, major vendors are accelerating the establishment of end-to-end hardware ecosystems spanning the desk to the data center.
The shift in AI adoption can be summed up as a structural transition driven by the combination of business-site demand for Economics, Data Sovereignty, and lower Latency. In a survey by Nvidia of more than 3,200 people in finance, manufacturing, retail, healthcare, and telecommunications, 64% of respondents said they had adopted an AI operating environment within their organizations. Another 42% identified the optimization of existing AI workflows and production cycles as a top investment priority. This suggests that corporate interest is moving from the pilot validation stage to the production stage centered on cost and efficiency optimization for everyday work.
This trend is also confirmed by a joint survey by Dell Technologies and IDC. Among corporate respondents, 89% said they considered AI features a key criterion when replacing PCs, and 78% rated the security advantages of local endpoint devices as very important. In response to this shift in demand, global semiconductor and finished-system companies are presenting hybrid architectures that encompass endpoints, workstations, and distributed servers, and are competing to secure leadership in “desk-side AI infrastructure.”
The photo shows the Lenovo SE30n.1. One driver behind the expansion of desk-side infrastructure is the limits of AI Tokenomics. Whereas chatbots handle one-off prompts, agentic AI operates autonomously.
Agentic AI repeatedly plans, searches for data, calls tools, and checks results in the process of achieving goals. As a result, it consumes massive amounts of tokens.
For this reason, relying on public cloud APIs for every request makes it difficult to control cost volatility. A Dell Technologies official said the market is moving toward a “hybrid AI architecture” that allocates the optimal execution location for each workload type.
According to analysis, if high-frequency and repetitive tasks that accumulate cloud API usage are handled locally on a high-performance AI PC on the desk, operating stability can be secured within a predictable cost structure.
Local accelerated environments have a cost advantage over cloud API calls, and analysis suggests that the initial hardware investment can be recouped within a few months. According to proof-of-concept analysis in the infrastructure industry, switching repetitive, high-frequency inference tasks to local-only infrastructure can reduce 3-year cumulative TCO to one-tenth of cloud pay-as-you-go spending. Analysis has also shown that distributed processing of repetitive and predictable enterprise inference workloads at the edge can reduce operating costs per token compared with cloud IaaS, broadening the persuasive power of cost-per-token reduction analysis.
Vinay Awasthi, vice president at AMD, said employees without a cloud AI license can use local inference without adding separate licenses. He said local inference can be used without adding token usage to cloud costs, and cited document writing, summarization, analysis, and test code generation as examples of applicable tasks. He added that for organizations seeking broader AI use among more members, local inference becomes a way to expand access without increasing costs.
Against this backdrop, infrastructure and software companies are expanding the adoption of “intelligent hybrid routing” technology. This approach preemptively processes high-frequency and sensitive workloads locally, and calls the cloud only when complex inference or external data retrieval is needed. As a result, companies are seeking to improve both cost and accessibility by sending simple, repetitive, and sensitive tasks to local infrastructure with a strong cost advantage, while using the cloud only when more complex inference or external data use is required.
The product name shown in the photo is Dell Pro Max with GB10. The adoption of desk-side and edge infrastructure is presented not only as a matter of cost efficiency but also as an issue connected to Data Sovereignty and Latency. One factor that determines infrastructure distribution is on-site stability.
There is also a background that makes it difficult to send data to external public cloud networks. If manufacturing design data, financial transaction records, healthcare research data, public secrets, or internal corporate source code are sent to external public cloud networks, there is a risk of violating security regulations or leaking intellectual property. Accordingly, keeping internal data on local endpoints is explained as a strategy to physically shrink the security attack surface.
Latency is presented as a factor directly tied to practical productivity. Today's knowledge worker environment is described as simultaneously running dozens of web browser tabs, live meeting caption generation and translation, multiple AI agent calls, and code verification on a single device, making it highly demanding.
In such a high-load environment, cloud round-trip time, which varies depending on network conditions, is cited as a factor that creates bottlenecks in continuous work processing. In the end, the adoption of desk-side and edge infrastructure is described not as a simple cost issue but as a matter intertwined with Data Sovereignty, Latency, and on-site stability.
In the field, there is a call for criteria to determine how cloud and local deployment should be divided. AMD presented four major criteria for evaluating workloads: Latency, Data Sensitivity, Economics, and Operational Efficiency. According to these criteria, large-scale model training is suitable for the cloud, and highly variable workloads with extreme demand swings are also suitable for elastic cloud infrastructure. By contrast, real-time tasks sensitive to Latency are best handled in a local environment, as are tasks involving sensitive data that cannot be moved outside. Always-on, routine productivity workloads are also best handled locally.
Yoon Seok-jun, vice president at Lenovo, said inference workloads occur repeatedly and continuously, unlike training. He added that tasks such as quality inspection on manufacturing floors, customer behavior analysis in stores, real-time video analysis at logistics hubs, and network optimization in telecommunications edge environments are more efficiently processed close to the site than by sending them to the cloud for batch processing, because of Latency requirements and the sensitivity of the data. In line with this trend, a hybrid AI model in which the cloud, on-premises systems, the edge, and devices each take on a role is expected to spread instead of a model that processes all AI in the cloud.
In practice, companies prioritize system stability, reduced management complexity, and interoperability with existing systems over simple hardware performance figures. Before adopting AI and digital workloads, companies demand reduced platform management burdens and continuity of service.
These needs translate into the key criteria for selecting desk-side and edge AI infrastructure. The key criteria for selecting desk-side and edge AI infrastructure are reduced platform management burdens and continuity of service.
Within this trend, the competition among semiconductor makers is converging on leadership in “desk-side infrastructure.” The point where architectural competition converges is three technical directions.
The three technical directions are high-capacity Unified Memory for accommodating large parameter counts on local endpoints, CPU-centered heterogeneous computing for directing agent workflows, and an open software ecosystem for breaking away from dependence on a specific vendor. The item labeled in the photo is Dell Pro Max with GB300, and the photo credit goes to Dell Technologies.
In the age of agentic AI, the role of the CPU is being reevaluated. The work performed by agents includes process execution, file analysis, command execution, and coordination of multiple sub-agents, and the engine responsible for this base execution and orchestration is the CPU. Even if accelerator computation is fast, bottlenecks emerge if the CPU falls short in system pipeline and data supply processing speed.
Accordingly, the industry is focusing on heterogeneous computing designs that organically combine CPU, GPU, and NPU. In this structure, the CPU handles general-purpose control, the GPU handles parallel computation, and the NPU handles always-on, low-power processing. At the same time, on-device use cases now require tens of billions to hundreds of billions of parameters and long-context memory loading, making Unified Memory architectures, which enable ultra-high-speed memory sharing between CPU and GPU, increasingly important. Unified Memory has emerged as a new standard for desk-side systems, with the effect of overcoming the physical limits of existing discrete card VRAM capacity and enabling large-model inference on slim PCs and small form factors.
The interview headline is “AI Infrastructure, Workload-Centered Distributed Architecture.” The interviewee is Vinay Awasthi, vice president at AMD. The interview focused on evaluating the distributed and expanding trend of AI infrastructure.
Jo Min-seong, managing director at Intel Korea, said that power, space, cost, and real-time requirements differ by edge environment. He explained that a single-accelerator-centered approach alone has limits in meeting diverse on-site demands.
As an alternative, Jo recommended a method of gradually adding AI functions to PCs, workstations, and servers rather than completely replacing existing systems. This was presented as an approach that aims for phased expansion instead of a full infrastructure overhaul.
He cited reduced initial investment burden as an effect of the gradual-addition approach. He also presented reduced system integration burden as another effect of this method.
This approach is combining with the spread of open programming models and is also combining with the spread of software optimization tools. This combination has been presented as leading to reduced dependence on a specific chipset.
In this way, the combination of an approach that aims for phased expansion rather than a full infrastructure replacement, the spread of open programming models, and the spread of software optimization tools is drawing attention as a key means of reducing dependence on specific chipsets.
Agentic AI is characterized not by one-off inference but by continuous inference, information retrieval, task coordination, and execution across multiple systems. Accordingly, companies must review both the model and the entire surrounding infrastructure at the same time, which also means there is no single absolute optimal location for AI execution. Depending on the workload, some tasks are suited to large cloud and data center infrastructure, but depending on Latency, Data Governance, efficiency, and token economics, on-premises systems, the edge, and AI PCs may be more suitable.
Therefore, the key principle is to choose the appropriate scale and environment for each workload and secure the ability to move flexibly between environments when requirements change. The question raised in this context is why the importance of the CPU is expanding in the age of agentic AI. The gist of the answer is that agentic AI involves substantial orchestration work beyond model inference.
The CPU handles base execution and orchestration in the process of executing an agent's processes, analyzing files, carrying out commands, and coordinating multiple workers. It also coordinates system resources during AI inference on the GPU and NPU, connects workloads with data and services, and continuously supplies the data needed by the GPU and NPU. As a result, the CPU supports the smooth progression of the overall workflow, and as AI agents become more autonomous, the importance of CPU performance also increases. It was also noted that no matter how fast an AI model may be, its practical value is limited if the speed of the surrounding workflow processing does not keep up.
CPU, GPU, and NPU each serve as complementary computing engines in AI systems. In an AI PC, the CPU handles general-purpose workloads and is responsible for orchestration and data movement management. The GPU handles model inference and content generation, and takes charge of workloads requiring highly parallel processing. The NPU performs dedicated AI tasks and is responsible for low-power continuous processing and efficient execution of dedicated AI tasks.
AMD said it sees AI as a system-level challenge, and that AI performance is determined by the operating efficiency of the CPU, GPU, NPU, memory, networking, and software. Accordingly, AMD said its goal is to support customers in building AI by providing an open full-stack AI platform, while securing greater flexibility, performance, and choice. It also said it aims to reduce unnecessary complexity and dependence on specific vendors while supporting the stable deployment and expansion of AI from data centers to AI PCs.
As lightweight models and “brownfield” infrastructure combinations draw attention, advances in AI model slimming technology are becoming the decisive catalyst for the spread of desk-side and on-site edge infrastructure. As knowledge distillation, weight quantization, and training data refinement techniques advance, small models with 30 billion parameters or less have begun to deliver performance comparable to that of past large models. As a result, practical business logic processing is now possible on desk-side equipment consuming only tens of watts, without large GPU clusters.
A Mobinnet official said that when AI moves from screen-based services into robots, vehicles, equipment, and living spaces, the requirements change fundamentally. The official said decisions must be made in milliseconds, data collected by cameras and sensors must not be taken outside, and systems must operate continuously in places physically connected to people’s everyday lives.
The same official added that these three conditions are areas that central servers beyond the network struggle to satisfy structurally. Accordingly, as AI use cases expand into physical environments, issues of Latency, data export, and always-on operation emerge as key conditions, and the need for infrastructure closer to the site than a central server is becoming more pronounced.
As a solution for adoption in industrial settings, a brownfield strategy is emerging over a greenfield one. The core of the brownfield strategy is to add AI acceleration functions while maintaining existing large-scale infrastructure.
This approach is consistent with the characteristics of control-room and manufacturing environments, where many general cameras and legacy servers are already in operation. That is because fully replacing control-room and manufacturing environments would entail an astronomical cost burden. Accordingly, for existing server-slot installation methods, low-power accelerator cards are proposed, and for endpoint expansion methods, USB dongles and modules are applied. USB dongles and modules have the effect of instantly giving endpoints AI inference performance, and this flexible hardware approach lowers the barrier for companies to enter AI. The product in the photo is Mobinnet’s “MLD-R1 AI USB.”
Companies are reaching a common conclusion that the spread of desk-side and edge AI will not replace the cloud. They view desk-side and edge AI and the cloud as not being in a zero-sum relationship.
The division of roles is also clear. Large-scale model training and tasks requiring momentary large-scale computation are handled by data centers and the cloud, while everyday tasks sensitive to Latency, internal security data retrieval, and on-site equipment control are handled by the desk and endpoints. Accordingly, a hybrid distributed structure is becoming established as the standard architecture.
Against the backdrop of the rapid evolution of AI models and workloads, demand is growing for flexibility that avoids dependence on specific hardware and closed architectures. Since dependence on a single vendor creates the burden of redesigning the entire system with each technological shift, the industry is turning its attention to open ecosystem standards such as ROCm, UALink, Ultra Ethernet, OpenVINO, and oneAPI.
At the same time, companies face the task of setting criteria for evaluating the results of adopting agentic AI. AMD said it expects the criteria for measuring AI agent ROI to move away from traditional measures such as cost savings and time savings from software adoption and toward business-outcome-centered measures. Accordingly, the key metrics under discussion include cost per completed task, cost per successful outcome, task completion rate, quality of output, and cost per execution. The ultimate goal is not simply to increase the number of AI agents, and AMD said the core of evaluation lies in the efficiency, stability, economics, and degree of goal completion achieved by agents in a distributed hybrid environment.
While the center of the enterprise AI paradigm is shifting from central cloud systems to user endpoints and on-site environments, the key to future corporate AI competitiveness will not be securing large, expensive accelerators themselves, but rather deploying compute resources in the optimal locations and linking them organically based on workload characteristics, Data Sensitivity analysis, and Tokenomics analysis. The outlook is that hybrid infrastructure strategy will determine the success or failure of enterprise AI deployment in practice.
Source: IT DAILY · Kwon Young-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=241932
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.