WD: AI Infrastructure Economics Are Shaped by Data, Not Just Compute
TECHWORLD ·
✦ AI Summary
As AI spreads, data generation volumes and retention periods are increasing, expanding the key economics of large-scale AI infrastructure from compute to data storage.
WD released IDC global research findings on the 10th, and IDC said data storage demand is structurally increasing as AI adoption expands.
In the IDC survey, 94.7% of respondents from companies and institutions said the amount of data requiring storage increased, and 98.2% said TCO per TB is important in storage decisions.
As AI spreads, the amount of data being generated is rising and data retention periods are getting longer, expanding the key factors behind the economics of large-scale AI infrastructure from compute to data storage.
In this regard, WD on the 10th released global research findings sponsored by WD and IDC. The IDC white paper is titled “Built for Scale: The Enduring Role of HDDs in the AI Era.”
IDC said the demand for data storage from companies and institutions is structurally increasing as AI adoption expands. It also said there are more cases of long-term retention of AI-generated data and more cases of repurposing archived data for AI workloads. As a result, it said the importance of storage design that takes the entire data lifecycle into account is growing.
As AI adoption spreads, companies are generating more data and placing greater value on that data. Accordingly, data retention periods are trending longer. At the same time, more previously stored data is being brought back online. These reactivations are being carried out to use the data for new AI workloads.
According to the IDC survey, 94.7% of respondents from companies and institutions said the amount of data requiring storage increased after adopting AI and generative AI over the past 12 months. 61% of respondents said data increased by at least 25% through AI over the past year. 74% of respondents expected data to increase by at least 25% over the next 3 years.
The trend toward larger data lakes was also confirmed. 85.4% of respondents said the volume of data stored in data lakes increased over the past 12 months. This shows that data lakes are continuing to expand in line with the recent increase in data that needs to be stored.
AI-generated data was cited as the biggest factor behind the expansion of data lakes, accounting for 59.4%. Examples of AI-generated data included synthetic data, inference results, and model logs. This shows that growth in data lakes is being driven largely by the spread of AI-related data.
The structure of AI use is taking shape as data is used and new data is generated at the same time, with the generated data then being accumulated again. In this process, AI does not simply consume data; it creates new data, and as a result the storage and accumulation structure is strengthening. However, the volatility of compute tasks may vary depending on specific workloads and investment cycles.
A substantial portion of data generated during AI processes remains stored and retained even after the tasks are completed. This trend ties into the subtitle’s point that data value rises and retention periods lengthen after AI adoption.
The impact of AI adoption is affecting not only the amount of data but also the value of existing data. According to IDC survey results, about 95% of the companies and institutions surveyed said the value of the data they hold rose after adopting AI and generative AI.
In the same survey, 74.3% said data retention periods lengthened after AI adoption. This shows that AI adoption is linked not only to the creation and accumulation of data but also to longer retention periods.
At the same time, the use of archived cold data from the past is also increasing. 75.9% of respondents said there has been a rise in cases of bringing existing archived cold data back online for use in AI workloads.
As demand for using archived data is tied to supporting AI inference and RAG applications, the need to improve access speed is growing. 96% of all respondents said they expect faster access to archived data will be needed to support AI inference and RAG applications. Accordingly, the importance of faster access to archived data is also increasing.
At the same time, in actual data environments, low-frequency access tiers accounted for a large share. According to IDC, 74.6% of enterprise data at the companies and institutions surveyed was stored in warm, cool, and cold storage tiers. In addition, more than 60% of data lakes consisted of cold data or infrequently accessed data.
In this structure, the importance of building the right storage tiers according to data characteristics and access frequency is growing in large-scale AI environments. The analysis also indicated that AI infrastructure storage design needs to consider the entire data lifecycle rather than simply expanding high-performance storage capacity.
In particular, as data volume and retention periods increase, it is necessary to consider not only performance but also storage capacity, accessibility, and cost in a comprehensive way. Accordingly, AI infrastructure requires storage design that takes lifecycle and tier-specific characteristics into account, not just expansion of high performance.
Storage costs have emerged as a major factor determining economics in AI infrastructure. In the survey, 98.2% of respondents said total cost of ownership, or TCO, per TB was important or very important in storage decisions. In connection with this trend, Irving Tan, WD CEO, said that AI infrastructure discussions over the past several years have centered on compute, but AI is built on data. He said companies are expanding data generation and extending retention periods while also seeking ways to create new value from data they already hold.
Irving Tan, WD CEO, said compute demand changes over time. By contrast, he explained that demand for storing, managing, and accessing data when needed in large-scale environments continues to rise.
He added that this data foundation is a key factor determining how far AI can scale.
In AI environments, various types of generated data are created during work processes, including training datasets, model checkpoints, inference results, logs, and synthetic data. A significant portion of this generated data continues to be retained even after compute tasks end.
Retained data is used as input data for future AI applications. As a result, the boundary between current-use data and archived data is changing, data stored in the past is becoming core input data for AI workloads, and long-unused data is also being converted into material for reuse.
As these changes take place, the considerations for AI infrastructure design are also expanding. It is necessary to consider the entire data lifecycle together, including not only compute and storage performance at a given point in time but also data generation, storage, long-term retention, and reuse when needed.
Based on the IDC research results, WD said that cost-efficient data storage and management capabilities in large-scale AI environments are emerging as a core factor for AI expansion, beyond a simple infrastructure element.
Source: TECHWORLD · Park Gyu-chan
Original: https://www.epnc.co.kr/news/articleView.html?idxno=406778
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.