Featured image of post Huawei Unveils Ascend 960DT Chip and Lingqu Ultra-Node Cluster at HCC 2026, Doubling AI Performance

Huawei Unveils Ascend 960DT Chip and Lingqu Ultra-Node Cluster at HCC 2026, Doubling AI Performance

Huawei accelerates Ascend 960DT launch with 2 PFLOPS FP8 computing power and introduces Lingqu UnifiedBus for million-chip clusters.

Huawei Unveils Next-Generation AI Compute infrastructure at HCC 2026

Huawei Unveils Next-Generation AI Compute infrastructure at HCC 2026
Huawei Unveils Next-Generation AI Compute infrastructure at HCC 2026|News screenshot

Huawei launched a suite of AI infrastructure products at the Huawei Connect Conference 2026 in Shanghai on September 16. Key highlights include:

  • Ascend 960DT chip: Supports 2 PFLOPS (FP8) and 4 PFLOPS (FP4) compute, launched ahead of schedule, available Q1 2027
  • Ascend 960 Ultra-Node: World’s first mass-produced NPO (Near-Package Optics) based node, supports 4096 cards per node, up to 8 EFLOPS FP8 total
  • Lingqu UnifiedBus protocol: Unified interconnect standard, fully open-sourced for public access
  • Kunpeng 950 Ultra-Node upgrade: Supports up to 4096 nodes, forming a 256TB unified memory pool
  • OceanStor M900 Memory Storage: Features Lingqu UnifiedBus interface, over 50% KV Cache latency reduction
  • Lingqu All-Optical Switching Equipment: Supports two-tier CLOS flexible networking architecture

The Lingqu UnifiedBus eliminates protocol conversion overhead across computing, storage, and communication components, enabling multi-node clusters to operate as a single high-efficiency system.

11 chips power the ultra-node ecosystem

11 chips power the ultra-node ecosystem
11 chips power the ultra-node ecosystem|News screenshot

Based on Lingqu UnifiedBus, Huawei has developed 11 key chips for the ultra-node ecosystem, categorized into four types:

  • Compute chips: Kunpeng CPUs and Ascend NPUs for core processing
  • Interconnect chips: Scale-out/Scale-up switches and high-speed optical interconnect chips
  • Storage chips: Medium controllers and Smart Storage Unit for AI KV Cache optimization
  • Management chips: System coordination and control functions

The Ascend 960PR variant extends the performance trajectory with 8 PFLOPS FP4 support, scheduled for Q3 2027, one quarter earlier than planned. The Ascend 970 and 980 will follow in 2028 and 2029 respectively, leveraging the “Tao Law” innovation framework for continued doubling of specifications.

An important indicator of adoption momentum: Ascend 910C ultra-nodes have deployed over 1,000 units, while Ascend 950 nodes have entered large-scale commercial deployment—proving the commercial viability of Huawei’s prior generation in actual production environments.

Architecture breakthrough: breaking the communication barrier

The industry’s fundamental challenge with large-scale AI clusters is communication overhead limiting Model Floating-Point Utilization (MFU). Conventional 100K-card clusters typically achieve only ~20% MFU, while Huawei’s Lingqu-based architecture elevates MFU to 35%.

For multi-rack scalability, Huawei introduced Hi-ONE, the world’s first mass-produced NPO engine delivering 7.2 Tbps bandwidth with integrated light source for balanced power efficiency, reliability, and latency. The Ascend 960 Ultra-Node integrates “Lingqu UnifiedBus + Hi-ONE” to double mean time between failures and achieve 99.8% system availability.

Storage innovations are equally significant: OceanStor M900’s direct Lingqu connection to Layer 3.5 KV Cache pools slashes data movement steps from five (traditional PCIe+RoCE) to one, cutting I/O and first-token latency by over 50%, while extending SSD lifespan 16x for consistent high cache hit rates.

Kunpeng ultra-nodes and edge AI strides

Kunpeng ultra-nodes and edge AI strides
Kunpeng ultra-nodes and edge AI strides|News screenshot

General-purpose computing ultra-nodes target Agent workloads. The Kunpeng 950 upgrade enables 4096-node interconnection with a unified 256TB memory pool, accelerating 100K sandbox startups 30x over traditional servers and improving sandbox density by 25%. Vector retrieval reaches 300K TPS (transactions per second), doubling performance.

Edge AI initiatives follow a quadruple-pillar approach: “Kirin + Ascend” self-reliant compute, “Pangu + third-party models” intelligent logic, HarmonyOS as an Agent-native OS, and open ecosystem expansion. The Mate X2 exceptional master edition already ships with a 30B-parameter native Pangu MoE model. OpenHarmony ecosystem exceeds 1.3 billion deployments, and openEuler has surpassed 20 million installations—the #1 in China’s server OS market.

Ecosystem milestones and compute delivery models

Ecosystem milestones and compute delivery models
Ecosystem milestones and compute delivery models|News screenshot

Huawei’s Kunpeng ecosystem now counts over 4.16 million global developers, while the Ascend CANN community attracts more than 5,000 monthly active users—with external developers representing 61%, surpassing internal contributors for the first time.

The Ascend NPU platform appears directly on PyTorch’s official installation page, joining NVIDIA, AMD, and Intel as one of only four supported backends. Over 40 native Ascend training models exist across 90+ mainstream open-source communities.

Huawei will meet diverse customer needs via two core capabilities: infrastructure (hardware-first) and cloud services (flexible consumption). The ultra-node architecture specifically optimizes hot/warm/cold KV Cache tiering with single-hop access for maximum resource efficiency.

Recommendations for Adoption

Early adopters should evaluate: Cloud providers planning 10K+ card AI training deployments benefit from the 35% MFU gain, translating to lower per-FLOP costs; Agent developers can leverage the Kunpeng ultra-node’s sandbox startup speed and memory pool benefits.

Consider waiting if: Your workloads demand FP16/BF16 precision—this announcement lacks published specifications for these formats; production stability concerns may favor waiting for the Q2 2027 air-cooled Ascend 960 Ultra-Node availability.

写在最后:The contest is shifting from standalone chip sales to integrated cluster solutions, where communication compute co-design defines competitive advantage. Huawei’s “silicon-based black soil” vision transforms legacy hardware strengths into systematic engineering capability for the era of 10-trillion-parameter models.