Is Huawei Ascend Really China's Answer to Nvidia?
When Nvidia's CEO Jensen Huang acknowledged that "Huawei has been moving quickly" and cited Huawei's CloudMatrix as evidence, the semiconductor world took notice. ByteDance has ordered $5.6 billion in Huawei chips for 2026. Hyperscaler demand could reach $12-15 billion. But can Huawei Ascend actually replace Nvidia for the most demanding AI workloads? The answer is nuanced—but increasingly optimistic for China's domestic chip industry.
The Ascend Family: Current Generation
Huawei's HiSilicon division designs AI accelerator chips under the Ascend brand. The current production flagship is the Ascend 910C, built on SMIC's enhanced 7nm process. Here's how it stacks up against Nvidia's offering:
| Specification | Ascend 910C | Nvidia H100 | Analysis |
|---|---|---|---|
| FP16 Compute | 256 TFLOPS | ~330 TFLOPS | ~77% of H100 |
| Memory Bandwidth | 1.2 TB/s | 3.35 TB/s | Gap exists |
| Interconnect | 600 GB/s | 900 GB/s (NVLink) | Below H100 |
| Price | $25K-35K est. | $25K-40K (original) | Similar list, but H100 costs more in China |
| Availability | Domestic supply | Export restricted | Huawei wins by being available |
The Ascend 910C delivers approximately 77% of the H100's compute performance at roughly one-third to one-fifth the effective cost when you factor in export premiums and availability constraints. For inference workloads where cost-per-compute matters more than raw speed, Ascend chips are already competitive.
The Software Challenge: CANN vs CUDA
Hardware specs tell only part of the story. The real moat Nvidia has built isn't just silicon—it's CUDA, the software ecosystem that has accumulated over 15 years of tooling, libraries, and developer expertise.
Huawei's answer is CANN (Compute Architecture for Neural Networks). But CANN is younger and less mature:
- CUDA advantage: Decades of optimization, community, and compatibility
- CANN maturity: Improving but still gaps in tooling
- Porting effort: Developers must rewrite code to move off CUDA
- PyTorch support: ~80% of standard inference workloads can run with minor adjustments
Huawei is aware of this limitation and is taking steps to address it. In September 2025, the company announced plans to open-source CANN, releasing all operators on GitCode. This could significantly accelerate ecosystem development.
"NVIDIA's CUDA has decades of ecosystem support, while Huawei's CANN is newer and less stable. Developers must rewrite code to move off CUDA, which slows adoption." — Industry Analysis
System-Level Innovation: CloudMatrix 384
Rather than competing chip-to-chip, Huawei has found an alternative strategy: system-level innovation. In July 2025, the company unveiled CloudMatrix 384—a rack-scale supernode containing 384 Ascend 910C chips.
According to SemiAnalysis, this system outperforms Nvidia's GB200 NVL72 on some metrics despite using weaker individual chips. The improvement comes from Huawei's supernode architecture, which allows chips to interconnect at very high speeds.
This is a significant insight: Huawei isn't trying to beat Nvidia chip-for-chip. It's achieving parity at the system level by connecting more chips and innovating at the cluster architecture level.
The 2026 Roadmap: Ascend 950 Series
Huawei published its full three-year Ascend roadmap at the Full Connect Conference in September 2025:
Ascend 950PR Released
1.56 petaflops FP4 compute, optimized for prefill inference and recommendation workloads. 112 GB HiBL 1.0 memory, 1.4 TB/s bandwidth.
Ascend 950DT
Optimized for decoding and training workloads. The missing piece for full training capability.
Ascend 960
Next-generation architecture with improved performance across all metrics.
Ascend 970
Targeting performance parity with Nvidia's then-current offerings.
With the 950 series, Huawei introduces self-developed HBM memory—branded HiBL 1.0 and HiZQ 2.0. This is a critical milestone: breaking dependency on SK Hynix, Samsung, and Micron for high-bandwidth memory.
The DeepSeek Effect on Huawei
Nothing accelerated Huawei Ascend adoption more than DeepSeek's V4 release in April 2026. When DeepSeek demonstrated that their latest model was optimized for Huawei hardware, major Chinese firms scrambled to secure Ascend chips.
Reuters reported that Alibaba, Tencent, and ByteDance all placed significant orders. ByteDance alone committed to $5.6 billion in Huawei chip purchases for 2026—representing the largest single procurement of domestic AI chips to date.
💡 Why DeepSeek V4 Changed Everything
Before DeepSeek V4, most Chinese companies used a hybrid approach: training on Nvidia hardware, inference on Huawei. V4's Huawei optimization proved that domestic hardware could handle demanding workloads. This shifted perception from "Huawei as backup" to "Huawei as primary."
Market Adoption: Who's Using Ascend
Adoption of Huawei AI chips is rising across China's AI sector:
iFlytek
SenseTime
China Mobile
DeepSeek
ByteDance
Alibaba
Total hyperscaler demand could reach $12-15 billion in 2026 according to industry estimates—a market that barely existed two years ago.
The H20 Problem: Why Export-Restricted Chips Failed
To understand why Huawei won, you need to understand the H20 failure. The US government forced Nvidia to create export-compliant chips—downgraded versions designed to be "just slow enough" to satisfy export regulations but "fast enough" to keep Chinese cloud giants buying.
But the H20 was a technical compromise that no one actually wanted:
- H20 performance: Roughly 15% of full H100 in certain critical workloads
- H20 price: Still expensive for what you get
- Math stopped working: Companies were paying premiums for intentionally crippled hardware
This created a vacuum. The H20 was supposed to be the bridge between restricted and unrestricted—instead, it accelerated the move to domestic alternatives. Companies realized: if we're going to accept significant performance penalties, we might as well invest in our own supply chain.
Supply Chain Independence: The Memory Question
High-bandwidth memory (HBM) was the targeted chokepoint in US export controls. SK Hynix, Samsung, and Micron were prohibited from selling advanced HBM to China. This was supposed to cripple Chinese AI chip development.
But China has been building domestic HBM capacity:
- CXMT: China's primary domestic memory manufacturer
- HBM3 samples: Shipped in early 2026—faster than expected
- Timeline: Just two years after HBM export restrictions were tightened
Huawei's 950 series uses self-developed HiBL 1.0 memory. While still behind the leading edge (HBM3e and HBM4), this represents genuine progress toward supply chain independence.
The Production Scale
Huawei's production plans reveal the ambition of China's domestic chip push:
| Metric | 2025 | 2026 Plan | Growth |
|---|---|---|---|
| Ascend 910C units | ~300,000 | ~600,000 | 2x |
| Total Ascend dies | Not disclosed | ~1.6 million | — |
| Cluster deployments | Multiple 1,000+ card clusters | 10,000+ card clusters | 10x scale |
This scale creates its own advantages: volume discounts, learning curve improvements, and supply chain leverage that smaller competitors cannot match.
The Geopolitical Dimension
The semiconductor competition has become explicitly geopolitical. Nvidia's loss of China isn't just a business problem—it's a strategic one. A year ago, China accounted for nearly 25% of Nvidia's total data center revenue. That figure has cratered to around 4%.
This creates a dangerous feedback loop for Nvidia:
- Less China revenue: Fewer sales in the world's largest AI market
- Less R&D capital: Revenue funds next-generation development
- Closing the gap: Huawei uses saved R&D money more efficiently
- Wider adoption: More users means more optimization for Huawei hardware
The US is caught in a bind. Tighter restrictions accelerate Chinese innovation. Looser restrictions restore Nvidia's market. There's no obvious path that both maintains restrictions and prevents Huawei from closing the gap.
What's Huawei Still Missing?
Despite impressive progress, honest analysis requires acknowledging where Huawei still trails:
- Training performance: Memory bandwidth limitations affect training throughput more than inference
- Software ecosystem: CUDA's 15-year head start cannot be closed overnight
- EUV lithography: SMIC's 7nm process lags TSMC and Samsung at 3nm/2nm
- Frontier model training: The most advanced training still happens on Nvidia hardware
The gap is real but narrowing—and narrowing faster than US policy assumed it would.
Atlas SuperPoD: The Cluster Story
In September 2025, Huawei unveiled the Atlas 950 SuperPoD—a full AI compute cluster system scheduled for Q4 2026. The specifications are designed to compete with Nvidia's DGX SuperPOD at the cluster level:
- 15,488 Ascend cards per SuperPoD
- High-speed interconnect: LingQu protocol enables efficient multi-card communication
- Scalability: Multiple SuperPoDs can be clustered together
This is the strategy: not individual chip superiority, but cluster-level competitiveness through superior system design.
Conclusion: The Answer Is Complicated
So is Huawei Ascend really China's answer to Nvidia? The honest answer: partially, and increasingly.
For inference workloads—running AI models rather than training them—the answer is increasingly yes. The cost-performance ratio is compelling, availability is guaranteed, and software support is adequate for most production use cases.
For training workloads—especially frontier model training—the answer is more complicated. Huawei is closing the gap, but memory bandwidth and ecosystem maturity still favor Nvidia.
For system-level deployment—the actual infrastructure that Chinese companies run—the answer may already be yes. CloudMatrix and Atlas SuperPoD demonstrate that China can build competitive AI infrastructure without depending on American hardware.
The trajectory is clear: three years ago, Huawei was a distant second choice. Today, it's the primary compute provider for many Chinese AI companies. In another three years, the question may not be "can Huawei replace Nvidia" but "when will Huawei surpass Nvidia in the markets where it matters."
Nvidia CEO Jensen Huang noticed. Everyone else should too.