Insight

[Tech Report] The Real Competitive Edge of the AI Era Datacenter Is Power Efficiency

TECHWORLD ·

Choi Tae-seok, Senior Executive Director of Technical Support at MPS [Photo: MPS]

✦ AI Summary

As AI spreads, both GPU and AI accelerator performance improvements and datacenter power problems are growing.

Datacenter standards are shifting from a focus on computing speed to throughput per unit of power.

The text explains that to improve efficiency, not only computational performance but also minimizing data movement and power measurement and control are important.

As changes in the computing industry accelerate alongside advances in AI, year-by-year improvements in GPU and AI accelerator performance continue. However, as AI spreads, rising computational performance is also accompanied by a growing power problem.

As a result, the standards for datacenter competitiveness are also changing. The benchmark, once centered on computing speed, is shifting toward throughput per unit of power.

The background to this shift is the sharp increase in datacenter power demand driven by the large-scale computing needs of AI and cryptocurrency. The push to restart the Three Mile Island nuclear plant in the U.S. is cited as an example showing this rise in power demand.

The estimated worldwide energy consumption of AI and cryptocurrency is put at more than 1.5%. The projected share of energy used for related data processing is 3% in 2030 and 4.4% in 2035.

A significant share of a computer's energy consumption comes not from computation, but from data movement. Data moves along the storage-to-memory-to-processor path, passing through multiple interfaces in the process.

Each time data moves along the data bus, power is consumed and heat is generated. Because of this structure, data movement rather than computation itself is identified as a major cause of power consumption.

Accordingly, to improve the efficiency of AI systems, it is necessary to consider not only computational performance but also minimizing data movement. Only by weighing computational performance and data movement minimization together can the conditions for improved efficiency be met.

Memory is an important variable in this process. Memory can account for about 22% of a server's total power consumption.

HBM (high-bandwidth memory) is widely used in AI accelerators. The interface efficiency of HBM is about 1.5 pJ/bit, while the interface efficiency of a typical DDR memory module is about 15 pJ/bit, making HBM more efficient than a standard DDR memory module.

However, an HBM stack can consume 75 to 100 W. Because the HBM stack is placed on the same board as a high-power processor, applying HBM requires addressing heat generation and cooling at the same time.

A typical processor cache line is 64 bytes, but when DRAM is read, internal data much larger than 64 bytes is activated, increasing the inefficiency of moving far more data than is actually needed inside the memory.

In one example, the amount of data moved in processing 64 bytes is about 20 kB, and in that case the share of the data actually needed is only 0.3%.

For this reason, an architecture that reduces unnecessary data movement has become more important than simply improving memory performance.

One improvement method is placing data appropriately in the memory hierarchy, with the principle that 'Hot Data' should be placed in a layer close to the processor and 'Cold Data' in a relatively distant layer.

Along with this, NUMA and CXL are cited as technologies that improve flexibility in using memory resources, though if data passes through multiple paths, power consumption may increase.

The more likely it is that multiple hops will be required to access memory, the more power consumption rises, by more than 3 times. Accordingly, the key factor is said to be close coupling between computation and data, rather than the amount of memory secured itself. The focus of importance is shifting toward reducing the distance between computation and data rather than the absolute amount of memory.

This power efficiency is directly tied to datacenter economics. According to an estimate quoted in the original text, cooling costs account for about 43% of datacenter operating expenses. Cooling costs were presented as being on par with the operating costs of the actual computing equipment.

As power efficiency improves, both electricity bills and cooling costs can be reduced. Along with this, datacenter performance standards are also changing. The shift is from simple Petaflops per second to Petaflops per Wh, which takes power consumption into account.

A prerequisite for such efficiency improvements is the ability to measure where power is being consumed. The point is that efficiency gains are only possible if it is possible to identify where power is used. It was emphasized that securing measurability comes first.

In this process, the role of power semiconductors in the AI era is also expanding. Power semiconductors are broadening their function beyond simple power supply to include real-time power measurement and system control. The importance of measurement and control for improving efficiency is also growing.

MPS's voltage regulator for memory modules is a PMIC. This PMIC is a product designed to address that need.

The system management interface is based on I²C, I³C, and Sideband Bus. Through this, it is possible to read current on each voltage rail while the system is running, as well as understand the power consumption of the memory module.

It is also possible to log states in which set thresholds are exceeded. The host system can read telemetry information and, if necessary, adjust application performance.

At the same time, the importance of power conversion efficiency was highlighted. MPS said its PMIC can improve power conversion efficiency by about 4% compared with competing solutions, and that this could translate into about 2% power savings across the datacenter as a whole.

In large-scale datacenters, even improvements in the single-digit percent range are meaningful. They can lead to lower power costs, lower cooling costs, and reduced carbon emissions.

As software is no exception, the discussion of power efficiency is not limited to hardware. Factors affecting power consumption for the same computation include data type and programming method. Using a data type larger than necessary dilutes the effect of hardware efficiency improvements, and repeatedly moving unnecessary data also dilutes the effect of hardware efficiency improvements.

For this reason, improving hardware alone has limited effect on efficiency gains. The need for an integrated perspective on processors, memory, software, and power systems has been raised. The trend is to consider implementation methods and data processing methods that directly affect power consumption as well.

The limits of simple performance evaluation standards in the AI era are also becoming clear. It is becoming harder to judge performance based only on the amount of computation processed, and the important performance benchmark going forward will shift to the amount of effective computation performed per 1 Wh of power.

Behind this change is the rising data demand of AI. As AI's data demand increases, memory expansion continues, and data movement also increases. On top of that, the complexity of power design rises as AI's data demand increases.

Accordingly, the level of power semiconductor capability needed in the future rises to a stage where stable voltage supply alone is not enough. The conditions for future power semiconductors are presented as high conversion efficiency, real-time measurement of power consumption points, and support for autonomous system optimization based on measurement data.

As AI performance improvements continue to be pursued, power efficiency has changed in status from an auxiliary design element to a core technology that determines both system performance and TCO.

In a field closely aligned with this change, Choi Tae-seok, Executive Vice President of Technical Support at MPS (Monolithic Power Systems), is responsible for technical support.

Engineer Choi Tae-seok, Executive Vice President, has focused on technical support for power semiconductor technologies needed for memory, and joined MPS in 2004.

He has accumulated broad experience through technical support for consumer products (TV, STB, Mobile) and for the automotive sector, and has contributed to the development of power solutions through his technical expertise.

Source: TECHWORLD · Lee Gwang-jae
Original: https://www.epnc.co.kr/news/articleView.html?idxno=406428

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.