Featured image of post Huawei Repositions Storage as AI Inference Scales Up

Huawei Repositions Storage as AI Inference Scales Up

AI inference is reshaping storage.

Storage Moves Into the AI Runtime

Storage Moves Into the AI Runtime

Huawei used its 2026 Data Storage User Elite Forum and OceanClub Carnival in Wuxi on August 6 to explain a broader AI data center infrastructure plan built around five layers: AI data lake, AI data platform, compute, model, and agent.

The message was not centered on a single storage appliance. Instead, Huawei focused on how enterprise data should be stored, governed, and repeatedly accessed once companies begin to deploy many models and agents. Yuan Yuan, Huawei vice president and president of the company’s data storage product line, said the focus of AI development is extending from compute and models to data. In Huawei’s reference architecture, security and resilience run across all layers.

Why Inference Creates a New Storage Problem

Enterprise data is often scattered across data centers, servers, and business systems. It may include text, images, video, and other formats. For AI use, this data must be cleaned, labeled, retrieved, and transformed before it can become training material, a knowledge base, or context for an agent. Huawei positions products such as OceanStor Pacific and Omni-Dataverse as ways to form a unified data space.

The more notable shift is at the AI data platform layer. As inference scales up, agents generate and consume far more context. Relying only on HBM and DRAM can create pressure in both capacity and cost. HBM is high-bandwidth memory used by accelerators such as GPUs; it is fast, but expensive and limited in capacity. KV Cache refers to intermediate attention data stored during large language model inference to avoid repeated computation.

Huawei introduced CMS, a context memory storage system for large-scale inference and heterogeneous compute, and described it as a G3.5 storage layer. The common hierarchy in AI systems is often described as G1 HBM, G2 DRAM, G3 local SSD, and G4 shared storage. G3.5 sits between local SSD and shared storage, aiming to provide a faster, larger, shareable buffer outside accelerator memory and system memory.

Key points include:

  • It is not simply replacing HBM with SSDs; it is meant to move data that does not always need to stay in accelerator memory.
  • The benefit becomes clearer at very large cluster scale, including supernodes and thousand-card or ten-thousand-card deployments.
  • Huawei links the idea to PB-level shared KV Cache pools that can reduce repeated computation during inference.

Full Stack Means Integration, Not a Strategy Reset

Huawei has already released technologies such as AI data lake, UCM inference memory data management, and AI data platform capabilities over the past year. The stronger emphasis on “full stack” does not appear to signal a new direction. Wu Junjie, a Huawei data storage product line executive, said the central goal remains the same: helping AI land in real enterprise scenarios.

The difference is integration. Earlier work focused on point solutions such as KV Cache, knowledge bases, and joint industry innovation. Huawei is now packaging these capabilities into a more unified architecture. This also pushes the storage product line beyond the traditional boundary of storage. ModelEngine now covers model deployment and GPU/NPU resource scheduling, while Nexent reaches into agent development and runtime environments. An NPU is a processor optimized for neural network workloads, commonly used for AI training and inference.

For enterprises, migration matters as much as architecture. Huawei describes two deployment paths: customers building a new AI data platform can use OceanStor A800, while customers that already run OceanStor Dorado can add data engine nodes to bring in AI capabilities while preserving existing storage investment. This reflects a broader reality: traditional databases, ERP systems, and AI workloads will coexist for a long time, so AI data centers are likely to evolve gradually rather than through one-time replacement.

Data Becomes the Next AI Bottleneck

A recurring view at the forum was that the second half of AI will depend heavily on data. Compute remains foundational, but as enterprises complete early infrastructure buildouts and models mature, outcomes will increasingly depend on data quality, data governance, and retrieval efficiency.

Healthcare illustrates the shift. A hospital may previously have stored CT images and pathology slides inside departmental servers. To train a pathology model, however, that historical data must be gathered, governed, and converted into datasets usable by models. Similar issues exist in research, automotive, and manufacturing. Data that was once merely archived now has to become model knowledge, agent context, and long-term memory.

That is why storage vendors are paying more attention to vector search, KV Cache, and agent memory, instead of discussing only capacity and IOPS. Vector search turns text, images, or other content into mathematical vectors and finds similar items, a common technique for knowledge-base question answering and retrieval-augmented generation.

Market Cycle and Industry Direction

AI demand is also affecting storage supply and pricing. Wu said global storage vendors are being influenced by upstream supply changes. Based on current demand, he expects this storage cycle could last until around the first half of 2028, while stressing that this is not certain. If AI enthusiasm cools or infrastructure demand returns to normal, prices may also stabilize.

AI may not remove storage’s cyclical nature, but it is creating new categories of demand: training corpora, vector data, KV Cache, agent context, and long-term memory all increase the amount of data AI systems must manage. The deeper change is that storage is moving from a background persistence layer into the inference path. For future AI infrastructure, enterprises will likely evaluate not only compute scale, but also how fast data can be found, how long context can be retained, and how much repeated computation can be avoided through memory and caching.