Hardware
[On Site] Overnight for the KRW 1.99 Million iPhone 18 Pro... Its Variable Aperture Wins Over the First Buyer
Overnight at Apple Myeongdong Since the Previous Evening, Concern Over the KRW 200,000 Price Increase and Disappointment With AI
Hardware
Overnight at Apple Myeongdong Since the Previous Evening, Concern Over the KRW 200,000 Price Increase and Disappointment With AI
Hardware
In-House Developed 'IGRIS-C' Reaches 36 Units in Cumulative Production, Supporting Supply to Manufacturing and Logistics Sites and Data-Learning Training
Hardware
5.77 Times the Vector Search Throughput of a CPU Server... ASIC-Based Evaluation Scheduled for the Fourth Quarter
Hardware
Up to 5 Times Higher Power Density, 4-Month Shorter Development Cycle Silicon Wafer Packaging Technology Addresses Space, Thermal, and Efficiency Challenges at Once
AI TIMES ·
According to AI TIMES, Huawei on September 17 unveiled "context memory storage" to relieve the KV cache bottleneck in AI inference, saying it supports up to 64PB in a single cluster. The OceanStor M900 emphasizes an architecture that connects the NPU and SSD in one hop, cutting latency to as low as 60 microseconds. Its core idea is to extend the KV cache, which has stayed in on-chip memory and DRAM, to SSDs so that the data needed for ultra-long context and multi-turn inference can be stored and reused in a larger shared space. Huawei said this architecture shows a shift from existing compute-centric AI infrastructure toward a model that combines storage and networks. It added that the design raises inference performance by improving cache access efficiency while also managing storage endurance, which has been cited as an issue in large-scale operations. In the end, this unveiling shows that beyond competition in the model itself, the memory architecture and data movement methods used in the inference stage are emerging as key factors in AI infrastructure competitiveness.
Perspective
The significance of this issue is that the battleground for AI infrastructure is shifting from simple computing performance to how generated data during inference is stored and retrieved. In particular, as long contexts and complex interactions become more common, bottlenecks are more likely to arise across the entire system rather than inside the chip, and this announcement can be read as a signal that Huawei intends to extend its solution to storage as well. For the industry, the competitive standard is likely to become not just how to run larger models, but how stably and economically inference services can be operated with the same resources.
This perspective is BizCrush's own commentary and is not part of the reporting by AI TIMES.
This article was produced with the help of an automated content generation algorithm.
Hardware
Intel B70 Processes 32 Video Channels at Once on One Card... Reducing Hardware Selection and Optimization Burdens
Hardware
Woostbin CTO: “Data Is King”... The Shift Must Be From Applications to Data
Hardware
Lattice Semiconductor announced a new FPGA family, Mach-N2. Based on the Nexus 2 small FPGA platform, Mach-N2 includes integrated flash, Roo…
Hardware
Integrated Flash and Tamper Response Applied... AI Design Tool Shortens FPGA Development Time
Hardware
Official Response to Reuters Report on Intel Talks: No Finalized Cooperation or Production Plans, Company Reviewing Various Options to Strengthen Global Competitiveness
Hardware
At Climate Industry International Expo, Heating, Cooling, and Hot Water Products Unveiled; Samsung Expands Production and Installation Training, LG Verifies Apartment Operation
Hardware
Integrating Ultra-Thin p-Type and n-Type Semiconductor Electrodes With a Single SnSe₂ Material
Hardware
Deal Approved After First Agreement Was Rejected; After 2 Weeks of Intensive Talks, 57.08% of Members Vote in Favor; Cash 50%, Stock 50% Adjusted ... Expanding Individual Choice
Hardware
Infineon Technologies has launched the dual-phase smart power stages TDA235E5 and TDA235E0. The product family is designed to address the po…
Hardware
AI Carries Out App Actions by Recognizing Personal Context and Screens; Model Built on Google Gemini Collaboration
Hardware
Optimized for AI, Industrial, Automotive, and Data Center Applications
Hardware
Mouser Electronics has announced the availability of onsemi's GaNEXUS GaN FETs. The GaNEXUS GaN FETs feature high-speed switching, low gate…
Hardware
ADI to Add Alif Processor to Portfolio and Expand TAM
Hardware
Semiconductor Exports Rise 13.8% to Set New Record as Share of Total Exports Jumps From 54% to 61%
Hardware
Up to 26% Power Savings and Lower Heat Through In-House High-Efficiency NPU Architecture That Minimizes Data Movement