Huawei announced the OceanStor M900 context memory storage during the HUAWEI CONNECT 2026 event in Shanghai. David Wang, Vice Chairman of the Board and Rotating Chairman at Huawei, presented the hardware during his keynote speech. The system supports AI inference workloads in hyperscale data centers by providing SuperPoD setups with a shared memory space reaching petabyte-level capacity and terabytes-per-second throughput.
Hardware Architecture and Network Integration
The system combines a central processing unit, network controller, and NAND flash controller into a single unified hardware architecture. This setup provides native key-value semantics that allow direct single-hop communication between neural processing units and solid-state drives. As a result, the design eliminates protocol conversion and data forwarding steps through intermediate host processors. In benchmark evaluations, internal testing showed a 90 percent drop in data latency, bringing response times down from milliseconds to 60 microseconds.
Memory Capacity and SuperPoD Systems
As artificial intelligence models evolve toward complex multi-agent workflows and extended context processing, data center storage infrastructure faces substantial performance demands. Modern large language models frequently operate with context windows exceeding one million tokens, generating large amounts of cache data during multi-round inference tasks. To address these memory constraints on computers and server clusters, the platform expands cache tiers across solid-state drives using the UnifiedBus high-speed network. Consequently, a single deployment provides up to 64 petabytes of shared capacity while increasing available cache per processor to terabyte levels.
Performance Features of OceanStor M900
A single deployment of the OceanStor M900 delivers a total access bandwidth of 40 terabytes per second. In standard software development and coding benchmarks, the infrastructure doubled the token output rate across inference clusters. Furthermore, the architecture reduced the duration required to generate the initial output token by 50 percent, improving operational throughput for large-scale enterprise services.
Durability and Operational Cost Reduction
To support data center economy metrics, the OceanStor M900 incorporates adaptive storage management that monitors cache lifecycles and places data across relevant media tiers. The technology sustains up to 24 drive writes per day, extending solid-state drive endurance by 16 times over three years of operation. This efficiency lowers hardware replacement frequencies, minimizes hardware wear, and curbs long-term maintenance expenses for enterprise operators managing intensive inference clusters.