When Nvidia's CEO Jensen Huang acknowledged that "Huawei has been moving quickly" and cited Huawei's CloudMatrix as evidence, the semiconductor world took notice. ByteDance has ordered $5.6 billion in Huawei chips for 2026. Hyperscaler demand could reach $12-15 billion. But can Huawei Ascend actually replace Nvidia for the most demanding AI workloads? The answer is nuanced—but increasingly optimistic for China's domestic chip industry.

77%
H100 Performance
$5.6B
ByteDance Order
600K
910C Units Planned 2026
1/3-1/5
Cost vs H100

The Ascend Family: Current Generation

Huawei's HiSilicon division designs AI accelerator chips under the Ascend brand. The current production flagship is the Ascend 910C, built on SMIC's enhanced 7nm process. Here's how it stacks up against Nvidia's offering:

Specification Ascend 910C Nvidia H100 Analysis
FP16 Compute 256 TFLOPS ~330 TFLOPS ~77% of H100
Memory Bandwidth 1.2 TB/s 3.35 TB/s Gap exists
Interconnect 600 GB/s 900 GB/s (NVLink) Below H100
Price $25K-35K est. $25K-40K (original) Similar list, but H100 costs more in China
Availability Domestic supply Export restricted Huawei wins by being available

The Ascend 910C delivers approximately 77% of the H100's compute performance at roughly one-third to one-fifth the effective cost when you factor in export premiums and availability constraints. For inference workloads where cost-per-compute matters more than raw speed, Ascend chips are already competitive.

The Software Challenge: CANN vs CUDA

Hardware specs tell only part of the story. The real moat Nvidia has built isn't just silicon—it's CUDA, the software ecosystem that has accumulated over 15 years of tooling, libraries, and developer expertise.

Huawei's answer is CANN (Compute Architecture for Neural Networks). But CANN is younger and less mature:

  • CUDA advantage: Decades of optimization, community, and compatibility
  • CANN maturity: Improving but still gaps in tooling
  • Porting effort: Developers must rewrite code to move off CUDA
  • PyTorch support: ~80% of standard inference workloads can run with minor adjustments

Huawei is aware of this limitation and is taking steps to address it. In September 2025, the company announced plans to open-source CANN, releasing all operators on GitCode. This could significantly accelerate ecosystem development.

"NVIDIA's CUDA has decades of ecosystem support, while Huawei's CANN is newer and less stable. Developers must rewrite code to move off CUDA, which slows adoption." — Industry Analysis

System-Level Innovation: CloudMatrix 384

Rather than competing chip-to-chip, Huawei has found an alternative strategy: system-level innovation. In July 2025, the company unveiled CloudMatrix 384—a rack-scale supernode containing 384 Ascend 910C chips.

According to SemiAnalysis, this system outperforms Nvidia's GB200 NVL72 on some metrics despite using weaker individual chips. The improvement comes from Huawei's supernode architecture, which allows chips to interconnect at very high speeds.

This is a significant insight: Huawei isn't trying to beat Nvidia chip-for-chip. It's achieving parity at the system level by connecting more chips and innovating at the cluster architecture level.

The 2026 Roadmap: Ascend 950 Series

Huawei published its full three-year Ascend roadmap at the Full Connect Conference in September 2025:

Q1 2026

Ascend 950PR Released

1.56 petaflops FP4 compute, optimized for prefill inference and recommendation workloads. 112 GB HiBL 1.0 memory, 1.4 TB/s bandwidth.

Q4 2026

Ascend 950DT

Optimized for decoding and training workloads. The missing piece for full training capability.

Q4 2027

Ascend 960

Next-generation architecture with improved performance across all metrics.

Q4 2028

Ascend 970

Targeting performance parity with Nvidia's then-current offerings.

With the 950 series, Huawei introduces self-developed HBM memory—branded HiBL 1.0 and HiZQ 2.0. This is a critical milestone: breaking dependency on SK Hynix, Samsung, and Micron for high-bandwidth memory.

The DeepSeek Effect on Huawei

Nothing accelerated Huawei Ascend adoption more than DeepSeek's V4 release in April 2026. When DeepSeek demonstrated that their latest model was optimized for Huawei hardware, major Chinese firms scrambled to secure Ascend chips.

Reuters reported that Alibaba, Tencent, and ByteDance all placed significant orders. ByteDance alone committed to $5.6 billion in Huawei chip purchases for 2026—representing the largest single procurement of domestic AI chips to date.

💡 Why DeepSeek V4 Changed Everything

Before DeepSeek V4, most Chinese companies used a hybrid approach: training on Nvidia hardware, inference on Huawei. V4's Huawei optimization proved that domestic hardware could handle demanding workloads. This shifted perception from "Huawei as backup" to "Huawei as primary."

Market Adoption: Who's Using Ascend

Adoption of Huawei AI chips is rising across China's AI sector:

iFlytek

Using Ascend for training and inference of AI models

SenseTime

Computer vision and AI model training on Ascend

China Mobile

Telecom AI infrastructure on domestic chips

DeepSeek

R1 and V4 inference running on Ascend clusters

ByteDance

$5.6B order for TikTok/AI workloads

Alibaba

Large orders for Qwen model deployment

Total hyperscaler demand could reach $12-15 billion in 2026 according to industry estimates—a market that barely existed two years ago.

The H20 Problem: Why Export-Restricted Chips Failed

To understand why Huawei won, you need to understand the H20 failure. The US government forced Nvidia to create export-compliant chips—downgraded versions designed to be "just slow enough" to satisfy export regulations but "fast enough" to keep Chinese cloud giants buying.

But the H20 was a technical compromise that no one actually wanted:

  • H20 performance: Roughly 15% of full H100 in certain critical workloads
  • H20 price: Still expensive for what you get
  • Math stopped working: Companies were paying premiums for intentionally crippled hardware

This created a vacuum. The H20 was supposed to be the bridge between restricted and unrestricted—instead, it accelerated the move to domestic alternatives. Companies realized: if we're going to accept significant performance penalties, we might as well invest in our own supply chain.

Supply Chain Independence: The Memory Question

High-bandwidth memory (HBM) was the targeted chokepoint in US export controls. SK Hynix, Samsung, and Micron were prohibited from selling advanced HBM to China. This was supposed to cripple Chinese AI chip development.

But China has been building domestic HBM capacity:

  • CXMT: China's primary domestic memory manufacturer
  • HBM3 samples: Shipped in early 2026—faster than expected
  • Timeline: Just two years after HBM export restrictions were tightened

Huawei's 950 series uses self-developed HiBL 1.0 memory. While still behind the leading edge (HBM3e and HBM4), this represents genuine progress toward supply chain independence.

The Production Scale

Huawei's production plans reveal the ambition of China's domestic chip push:

Metric 2025 2026 Plan Growth
Ascend 910C units ~300,000 ~600,000 2x
Total Ascend dies Not disclosed ~1.6 million
Cluster deployments Multiple 1,000+ card clusters 10,000+ card clusters 10x scale

This scale creates its own advantages: volume discounts, learning curve improvements, and supply chain leverage that smaller competitors cannot match.

The Geopolitical Dimension

The semiconductor competition has become explicitly geopolitical. Nvidia's loss of China isn't just a business problem—it's a strategic one. A year ago, China accounted for nearly 25% of Nvidia's total data center revenue. That figure has cratered to around 4%.

This creates a dangerous feedback loop for Nvidia:

  1. Less China revenue: Fewer sales in the world's largest AI market
  2. Less R&D capital: Revenue funds next-generation development
  3. Closing the gap: Huawei uses saved R&D money more efficiently
  4. Wider adoption: More users means more optimization for Huawei hardware

The US is caught in a bind. Tighter restrictions accelerate Chinese innovation. Looser restrictions restore Nvidia's market. There's no obvious path that both maintains restrictions and prevents Huawei from closing the gap.

What's Huawei Still Missing?

Despite impressive progress, honest analysis requires acknowledging where Huawei still trails:

  • Training performance: Memory bandwidth limitations affect training throughput more than inference
  • Software ecosystem: CUDA's 15-year head start cannot be closed overnight
  • EUV lithography: SMIC's 7nm process lags TSMC and Samsung at 3nm/2nm
  • Frontier model training: The most advanced training still happens on Nvidia hardware

The gap is real but narrowing—and narrowing faster than US policy assumed it would.

Atlas SuperPoD: The Cluster Story

In September 2025, Huawei unveiled the Atlas 950 SuperPoD—a full AI compute cluster system scheduled for Q4 2026. The specifications are designed to compete with Nvidia's DGX SuperPOD at the cluster level:

  • 15,488 Ascend cards per SuperPoD
  • High-speed interconnect: LingQu protocol enables efficient multi-card communication
  • Scalability: Multiple SuperPoDs can be clustered together

This is the strategy: not individual chip superiority, but cluster-level competitiveness through superior system design.

Conclusion: The Answer Is Complicated

So is Huawei Ascend really China's answer to Nvidia? The honest answer: partially, and increasingly.

For inference workloads—running AI models rather than training them—the answer is increasingly yes. The cost-performance ratio is compelling, availability is guaranteed, and software support is adequate for most production use cases.

For training workloads—especially frontier model training—the answer is more complicated. Huawei is closing the gap, but memory bandwidth and ecosystem maturity still favor Nvidia.

For system-level deployment—the actual infrastructure that Chinese companies run—the answer may already be yes. CloudMatrix and Atlas SuperPoD demonstrate that China can build competitive AI infrastructure without depending on American hardware.

The trajectory is clear: three years ago, Huawei was a distant second choice. Today, it's the primary compute provider for many Chinese AI companies. In another three years, the question may not be "can Huawei replace Nvidia" but "when will Huawei surpass Nvidia in the markets where it matters."

Nvidia CEO Jensen Huang noticed. Everyone else should too.