Huawei Ascend 960 AI chips will arrive in two versions during 2027, with the 960DT training processor scheduled for the first quarter and the 960PR inference processor planned for the third quarter. The timetable sharpens Huawei’s challenge to Nvidia in China’s fast-growing AI infrastructure market.

 

Huawei rotating and acting chairman David Wang disclosed the schedule at Huawei Connect 2026 in Shanghai on September 17. The company also reported broader deployment of systems built around UnifiedBus, its interconnect technology for combining many processors into a single logical computing platform.

 

Huawei’s announcement adds four material details:

  • The 960DT training chip is due in the first quarter of 2027
  • The 960PR inference chip is scheduled for the third quarter
  • UnifiedBus systems have reached more than 370 customers
  • Huawei says its largest clusters can link one million processors

 

Huawei Ascend 960 AI Chips Get a Split Release

Huawei previously placed the Ascend 960 generation on a 2027 roadmap without detailing this two-chip release sequence. The new plan separates a DT product for model training from a PR product aimed at inference, giving customers a clearer view of how the family will address different workloads.

 

The South China Morning Post reported that the 960DT schedule moved forward by nine months, while the 960PR is now expected one quarter earlier. Huawei also intends to maintain an annual upgrade cycle, with later Ascend generations planned for 2028 and 2029.

 

Huawei described the 960DT as a performance-doubling generation, but it did not publish a complete specification sheet at the event. Manufacturing process, power consumption, memory capacity, software performance and independent benchmark results remain necessary before direct comparisons with Nvidia accelerators can be made.

 

The staged launch reflects how AI chips are increasingly specialized. Training systems prioritize throughput, memory bandwidth and fast communication during large distributed jobs, while inference hardware must balance latency, efficiency and cost as trained models answer requests in production.

 

UnifiedBus Becomes Huawei’s Scaling Strategy

Huawei’s competitive argument extends beyond the performance of a single processor. UnifiedBus is designed to connect large numbers of CPUs, NPUs and other components so they share data efficiently and operate as a larger computing system.

 

Wang said Huawei has developed 11 semiconductors for systems using the technology. The company’s largest linked configurations, which it calls superclusters, are designed to support as many as one million AI processors.

 

Huawei also said it has shipped more than 1,000 smaller linked systems, known as supernodes, to over 370 customers. It did not identify those customers or provide a model-by-model shipment breakdown, limiting independent assessment of how the deployments are being used.

 

The shipment figure nevertheless supplies an adoption signal that a product roadmap alone cannot. Huawei said in September 2025 that more than 300 Atlas 900 A3 SuperPoDs using UnifiedBus 1.0 had been deployed, providing a technical and commercial base for the larger 2027 systems.

 

Related Research

 

Huawei Targets System-Level Competition With Nvidia

Nvidia’s advantage combines leading accelerators, high-speed networking and a mature software ecosystem. Huawei is answering with a full-stack strategy that uses Ascend processors, UnifiedBus interconnects and large system designs to compensate for constraints on access to advanced foreign chips.

 

United States export controls restrict China’s access to certain high-end processors and semiconductor manufacturing equipment. Those controls have increased the strategic value of domestic alternatives for Chinese cloud providers, telecommunications groups, research institutions and government-backed computing projects.

 

Scaling more processors together is not a simple substitute for stronger individual chips. Large clusters must manage communication overhead, power consumption, cooling, reliability and software scheduling. Efficiency losses can erase much of the theoretical performance gained by adding hardware.

 

UnifiedBus therefore carries much of Huawei’s claim. If it can keep thousands of processors fed with data and coordinated during demanding workloads, Huawei can compete at the system level even when its individual accelerators face manufacturing disadvantages.

 

Benchmarks Will Decide the Ascend 960’s Position

The 2027 dates give customers a planning window, but the decisive evidence will come from shipping hardware. Buyers will need independently reproducible results for model training speed, inference throughput, energy efficiency, memory behavior and performance across commonly used AI frameworks.

 

Software compatibility will be equally important. Nvidia’s CUDA ecosystem benefits from years of developer adoption, optimized libraries and tooling. Huawei must show that models can move onto Ascend systems without expensive rewrites or persistent performance gaps.

 

The first-quarter 960DT launch is now the earliest major milestone. A successful release would give Huawei most of 2027 to prove the training platform before the 960PR broadens the family toward large-scale inference later in the year.

 

Huawei has moved from a broad annual roadmap to named products, dated launch windows and deployment figures. That makes the Ascend 960 program easier to evaluate—and raises expectations that the company will support its claims with detailed specifications and real-world performance data.