Featured image of post Ascend 950 Super-Node Commercially Launched: Unified Memory for Trillion-Parameter-Scale Models, Training and Inference Integrated

Ascend 950 Super-Node Commercially Launched: Unified Memory for Trillion-Parameter-Scale Models, Training and Inference Integrated

Huawei launches Ascend 950 super-node in dual cooling variants, enabling efficient training/inference of trillion-parameter-scale models with unified memory architecture.

New post

Ascend 950 Super-Node Commercially Launched

Ascend 950 Super-Node Commercially Launched
Ascend 950 Super-Node Commercially Launched|News screenshot

Huawei officially launched the Ascend 950 super-node at HUAWEI CONNECT 2026, marking the industry’s largest-scale commercial super-node deployment. The product family includes two variants:

  • Atlas 950 SuperPoD liquid-cooled super-node: 1,024 NPUs, 1,300+ Lingqu interconnection chips, 4,000+ optical modules for liquid-cooled data centers
  • Atlas 850E air-cooled super-node: Supports 32 to 768 cards in standard air-cooled racks, requiring no facility modification
  • Key capabilities: sub-10ms inference latency, training and inference integrated for trillion-parameter models, the only vendor commercially supporting mxFP8 low-precision training and full-stack mxFP4 inference in China
  • Availability: Commercial availability launched at HUAWEI CONNECT 2026, now available for commercial adoption by internet, finance, manufacturing, healthcare, and transportation sectors

Breaking the “Compute-Storage-Communication” Three Walls

![Breaking the “Compute-Storage-Communication” Three Walls](/images/ascend-950-super-node-commercially-launched-unified-memory-for-trillion-02.png “Breaking the “Compute-Storage-Communication” Three Walls|News screenshot”)

The core breakthrough lies in hardware-software co-designed architecture enabling thousands of chips to act as one supercomputer:

  • Chip-level innovation: Hardware now handles data transfer, format conversion, communication synchronization, and cache management—simplifying developer workflows
  • Lingqu high-speed interconnect: Enables 1,024-card unified memory addressing; remote memory access semantics match local access, with 2.1x higher inter-chip bandwidth vs. previous generation
  • System-level reliability: Three-layer reliability (device, network, system) supports month-long stable training on 10,000-card clusters

Engineering milestones make deployment feasible:

  • Orthogonal architecture eliminates internal cabinet cabling, reducing 6,000+ cables per cabinet; 50 km copper saved per 1,024-card node
  • TPSU power supply modules provide energy storage for training/inference power peaks
  • Multi-channel Lingbrake cold plate dissipates 2,000W per unit
  • Link-level retransmission and 2+2 optical module protection solve industry flash-failure issues

Unexpected data point: Optical module liquid cooling saves power equivalent to thousands of households’ annual consumption—exceeding industry expectations for energy efficiency targets.

Product VariantCooling MethodTypical ScaleDeployment RequirementTarget Use Case
Atlas 950 SuperPoDLiquid1,024 cardsLiquid-cooled data centerLarge model training, high-throughput inference
Atlas 850EAir32-768 cardsStandard air-cooled rackEnterprise inference, Agent multi-turn conversations

Software System: The “Brain” for Unified Resource Management

![Software System: The “Brain” for Unified Resource Management](/images/ascend-950-super-node-commercially-launched-unified-memory-for-trillion-03.png “Software System: The “Brain” for Unified Resource Management|News screenshot”)

The software layer enables holistic resource coordination:

  • Lingqu system services: Support cross-node unified memory addressing with Load/Store semantics consistent with single-system operation
  • Collective communication performance: 1.6x speedup on 1,024-card domains; 1K packet latency drops from 70μs to 8.6μs (8x improvement)
  • Ecosystem: CANN and Mind series fully open-sourced; CANN community reaches 5,200+ monthly active developers, ranking #1 in China

Implementation Recommendations

Implementation Recommendations
Implementation Recommendations|News screenshot

  • Deploy now if:

    • You have liquid-cooled infrastructure and need integrated train-inference for billion-trillion parameter models (Atlas 950 SuperPoD)
    • Your facility uses air-cooling and requires sub-10ms Agent inference (Atlas 850E)
    • Full-stack mxFP4 inference and mxFP8 training are mandatory requirements
  • Wait if:

    • You depend on GPU ecosystem compatibility (昇腾 ecosystem still maturing)
    • Budget is tight and model size is under 100B parameters (node unit cost remains opaque)

Final Thoughts

昇腾 super-node demonstrates that commercial viability hinges on integrating hardware architecture, systems engineering, and software ecosystem. Its dual-cooling approach balances peak performance with deployment flexibility, establishing a reference architecture for trillion-parameter model infrastructure.

The super-node competition has shifted beyond specification sheets to real-world reliability—能否实现"千卡如一机" now defines infrastructure maturity.