Hardware

Huawei Unveils Inference Storage Aimed at KV Cache Bottlenecks

AI TIMES ·

According to AI TIMES, Huawei on September 17 unveiled "context memory storage" to relieve the KV cache bottleneck in AI inference, saying it supports up to 64PB in a single cluster. The OceanStor M900 emphasizes an architecture that connects the NPU and SSD in one hop, cutting latency to as low as 60 microseconds. Its core idea is to extend the KV cache, which has stayed in on-chip memory and DRAM, to SSDs so that the data needed for ultra-long context and multi-turn inference can be stored and reused in a larger shared space. Huawei said this architecture shows a shift from existing compute-centric AI infrastructure toward a model that combines storage and networks. It added that the design raises inference performance by improving cache access efficiency while also managing storage endurance, which has been cited as an issue in large-scale operations. In the end, this unveiling shows that beyond competition in the model itself, the memory architecture and data movement methods used in the inference stage are emerging as key factors in AI infrastructure competitiveness.

Perspective

The significance of this issue is that the battleground for AI infrastructure is shifting from simple computing performance to how generated data during inference is stored and retrieved. In particular, as long contexts and complex interactions become more common, bottlenecks are more likely to arise across the entire system rather than inside the chip, and this announcement can be read as a signal that Huawei intends to extend its solution to storage as well. For the industry, the competitive standard is likely to become not just how to run larger models, but how stably and economically inference services can be operated with the same resources.

This perspective is BizCrush's own commentary and is not part of the reporting by AI TIMES.

This article was produced with the help of an automated content generation algorithm.