Nvidia Boosts AI Data Center Power Efficiency With Vera Rubin and Groq 3 LPX
TECHWORLD ·
✦ AI Summary
Nvidia unveiled a full-stack AI factory strategy based on Vera Rubin NVL72 and Groq 3 LPX to improve power efficiency in AI data centers.
The strategy integrates systems, networking, software, and power management, and said DSX MaxLPS can dynamically redistribute rack-level power to increase GPU capacity by up to 40% within the same facility power capacity and improve token throughput by up to 35%.
To support long-context and multi-step reasoning, it combined Groq 3 LPX with the Vera Rubin platform, and said token throughput per MW can improve by up to 35 times compared with GB200 NVL72 for long-context environments and models with more than 2 trillion parameters.
Nvidia is pushing to improve power efficiency to address a key challenge facing AI data centers. Nvidia unveiled its full-stack AI factory strategy built on Vera Rubin NVL72 and Groq 3 LPX. On the 16th, Nvidia also announced technology to improve token throughput per watt based on Vera Rubin NVL72 and Groq 3 LPX.
This full-stack AI factory consists of an integration of systems, networking, software, and power management. The goal of the strategy is to maximize computing performance within limited power infrastructure. Power supply is the key constraint for AI factories.
AI models are getting larger, and inference workloads are increasing. As a result, the number of tokens that can be processed per MW of power has emerged as a key metric that determines data center operating efficiency.
DSX MaxLPS is a technology that dynamically reallocates power in response to changes in rack-level power demand within an AI factory. It raises the efficiency of power capacity utilization across the facility when the power demand of a particular rack rises or falls.
Nvidia said that combining a Vera Rubin NVL72-based system with DSX MaxLPS can increase GPU capacity by up to 40% within the same facility power capacity. It also said token throughput can improve by up to 35% without adding new power lines.
Each rack is equipped with Intelligent Power Smoothing software and expanded energy buffering capabilities. Intelligent Power Smoothing and the energy buffering functions absorb momentary power spikes, and this approach helps the system operate closer to continuous power demand. Nvidia explained that spare power capacity needed to keep existing data centers running stably can be converted into actual GPU computing resources.
Agentic AI refers to the continuous execution of AI agents' multi-step reasoning and tool-calling structure. As a result, longer context lengths may lead to higher latency and greater computing burden.
To support long-context and multi-step reasoning in agentic AI, Nvidia combined Groq 3 LPX with the Vera Rubin platform. Groq 3 LPX adds deterministic ultra-low-latency inference capabilities to the Vera Rubin platform and also enhances DSX MaxLPS power management functions.
For long-context environments and models with more than 2 trillion parameters, Nvidia said token throughput per MW can improve by up to 35 times compared with GB200 NVL72.
Measured performance was presented for a Qwen 3.8 27B workload with a 100,000-token context. In this configuration, Groq 3 LPX delivered 2,529 output tokens per second per user.
Nvidia said that if performance headroom is secured, agents can carry out more tasks within the processing range even as AI workloads expand, and added that this would allow agents to increase reasoning steps and tool calls within the same response handling range.
Source: TECHWORLD · Park Gyu-chan
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407022
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.