On June 30, 2026, Meituan's LongCat-2.0—a 1.6-trillion-parameter Mixture-of-Experts model—became the first publicly confirmed trillion-parameter AI model to complete full-process training and inference entirely on a 50,000-chip domestic Chinese cluster. Not a single Nvidia A100 or H100 was involved. Two months later, on August 8, HeteroFlow v2 launched with the ability to unify nine domestic GPU brands under a single API. The question is no longer whether China can build competitive AI chips—it is whether those chips can scale to meet the demands of the next generation of models that are already arriving.

56%
Domestic Share of China AI Server Chips (2026E)
90%
Projected Domestic AI Chip Market Share (2026)
50K
Domestic Chips in LongCat-2.0 Cluster
108%
Cambricon H1 2026 Revenue Growth YoY

The Landscape: Who's Building What

China's domestic AI chip ecosystem has grown from a handful of government-funded projects into a genuinely competitive industrial landscape. Five major players now define the space, each with a distinct strategy and market position.

Huawei Ascend: The 800-Pound Gorilla

Huawei's Ascend series commands roughly half of China's domestic AI chip market by shipment volume. The Ascend 910C, released in 2025, delivers approximately 60% of Nvidia's H100 inference performance according to DeepSeek's own internal testing—not a like-for-like replacement, but workable for a large class of inference workloads. Huawei is reportedly targeting 600,000 units of the 910C in 2026 alone.

The real story, however, is the next-generation Ascend 950 series, launched in Q1 2026. Huawei has split the lineup into two specialized variants: the 950PR, optimized for inference prefill and recommendation workloads, and the 950DT, optimized for inference decode and model training. Each is paired with Huawei's own proprietary HBM—HiBL 1.0 and HiZQ2.0 respectively—representing a level of vertical integration that no other Chinese chipmaker can match. At the system level, Huawei's CloudMatrix 384 supercomputer, built around 384 Ascend 910C processors with all-to-all optical interconnects, reportedly delivers 300 PFLOPS of BF16 compute—higher than Nvidia's GB200 NVL72 platform on key metrics, according to reports citing SemiAnalysis.

Cambricon: The Volume Player

Cambricon Technologies has emerged as the most commercially aggressive domestic AI chipmaker. The company reported first-half 2026 revenue of RMB 5.996 billion (~$830 million), up 108% year-over-year, with net profit surging 123%. Its Siyuan 590 processor, built on SMIC's N+2 (7nm-class) process, shipped nearly 100,000 units in 2025. The newer Siyuan 690, which entered mass production in Q1 2026, features a dual-die chiplet package delivering over 700 TFLOPS of FP16 compute with 196 GB of HBM3 and 890 Gbps+ of inter-die bandwidth.

The customer list tells the story: ByteDance has deployed over 100,000 Siyuan 590 and 690 cards across its infrastructure. Alibaba and Tencent have moved from testing to batch procurement. Cambricon is targeting roughly 500,000 total accelerator deliveries in 2026, including up to 300,000 Siyuan-series units—more than triple its 2025 output.

Alibaba T-Head: The Dark Horse

Alibaba's semiconductor subsidiary T-Head secured state procurement certification for its Zhenwu processors in May 2026, a milestone that signals Beijing's formal endorsement. At the 2026 World Artificial Intelligence Conference, T-Head demonstrated the Zhenwu M890 × Panjiu AL128 supernode, pairing the in-house Zhenwu M890 chip with the ICN Switch 1.0 interconnect. The system is designed to support clusters of up to 100,000 cards—a scale that puts it in direct competition with Nvidia's DGX SuperPOD architecture. T-Head also released the SAIL open-source software stack alongside the hardware, signaling a vertically integrated approach: custom silicon, custom interconnect, and a software layer built to match.

Behind the scenes, T-Head has been working with SMIC since December 2025 on a new 5nm-class chip aimed specifically at AI inference—a sign that the company is betting on the foundry's ability to push beyond 7nm, even without EUV lithography.

Moore Threads and Biren: The Wildcards

Moore Threads' MTT S5000 has been positioned as a GPU capable of running trillion-parameter models at inference scale, targeting the data center market where Nvidia's supply constraints have left the largest gap. Biren Technology, meanwhile, reported first-half 2026 revenue of RMB 1.15–1.3 billion—an increase of roughly 1,852% to 2,107% year-over-year—driven by demand for GPGPU computing across AI coding, long-horizon AI agents, and generative AI workloads. Both companies have secured state procurement certification, giving them access to government-funded data center projects that represent a growing share of China's AI infrastructure spend.

Huawei Ascend 950PR/950DT

Dual-variant architecture: prefill/recommendation (950PR) and decode/training (950DT). Proprietary HiBL 1.0 and HiZQ2.0 HBM. CloudMatrix 384: 300 PFLOPS BF16 across 384 chips via all-to-all optical interconnect.

Cambricon Siyuan 690

Dual-die chiplet on SMIC N+2 (7nm-class). 700+ TFLOPS FP16, 196 GB HBM3, 890 Gbps+ inter-die bandwidth. Targeting 300,000 units in 2026. Deployed at ByteDance, Alibaba, Tencent.

Alibaba Zhenwu M890

State-certified May 2026. Panjiu AL128 supernode with ICN Switch 1.0 interconnect. Designed for 100,000-card clusters. SAIL open-source software stack. 5nm-class inference chip in development with SMIC.

Moore Threads MTT S5000

Targeting trillion-parameter model inference. State procurement certified. Positioned for data center GPU workloads where Nvidia supply is constrained.

The Proof Point: LongCat-2.0 and the 50,000-Chip Question

Meituan's LongCat-2.0 is the single most important data point in the domestic chip debate. Here is why: it is not a lab experiment. It is a 1.6-trillion-parameter MoE model, trained on over 30 trillion tokens, that completed full-process pre-training and inference on a cluster of 50,000 domestic chips—with zero Nvidia hardware involved. The team spent three years working through operator adaptation, communication optimization, and distributed stability, scaling from a thousand cards to fifty thousand. The results speak for themselves:

  • Monthly failure rate reduced by over 70% through HCCL exception handling, elastic card scaling, and automatic fault recovery
  • Training MFU (Model FLOPs Utilization) improved 1.5x through pipeline scheduling, memory optimization, and operator-level core control
  • Steady-state daily throughput exceeded 1 trillion tokens per day
  • No irreversible loss spikes or rollbacks across the entire training run

On benchmarks, LongCat-2.0 scored 59.5 on SWE-bench Pro—beating GPT-5.5 (58.6) and Claude Opus 4.6 (57.3). On SWE-bench Multilingual, it scored 77.3, nearly matching Claude Opus 4.6's 77.8. On Terminal-Bench 2.1, it scored 70.8, demonstrating real-world agentic capability in production environments. As Maybank Investment Bank noted in a research report published August 18, 2026: "The milestone challenges the view that Chinese-made chips remain suitable mainly for inference rather than large-scale AI training."

The Maybank Verdict

"China's domestic AI ecosystem is gaining momentum as major models demonstrate the ability to train and run on locally developed chips, while Nvidia supply constraints and Beijing's restrictions on foreign processors accelerate the shift towards domestic hardware." The report also noted that Baidu's Ernie-5.1 achieved a 97% effective training rate on a fully domestic Kunlunxin cluster, and DeepSeek-V4-Pro completed full-parameter post-training on at least 1,000 Huawei Ascend 910C chips.

The Software Layer: HeteroFlow v2 and the CUDA Problem

Hardware is only half the story. The reason Nvidia has dominated AI computing for a decade is not just its chips—it is the CUDA software ecosystem, which represents roughly 20 years and billions of dollars of accumulated developer tooling, libraries, and optimization. Any Chinese chip that wants to compete must solve the software problem, and the solution emerging is not to replicate CUDA but to abstract above it.

HeteroFlow v2, launched on August 8, 2026, is the clearest expression of this strategy. It unifies nine domestic GPU brands—Huawei Ascend, Hygon DCU, Cambricon, Moore Threads, Biren, AMD, Kunlunxin, Enflame, and others—through a single OpenAI-compatible API. The system uses an open-source, multi-level intermediate representation-based compiler framework, allowing domestic chips to leverage existing community-developed tools rather than rebuilding Nvidia's software ecosystem from scratch.

The architecture works in three layers. At the hardware abstraction layer, an agent deployed on each compute node auto-detects the GPU type, model, memory capacity, and driver version—standardizing everything from nvidia-smi to npu-smi to cnmon into a unified resource model. At the engine routing layer, the system automatically selects the optimal inference engine for each chip: vLLM for Nvidia CC≥8.0, MINDIE for Huawei Ascend, vLLM-MTT for Moore Threads, and a Transformers fallback for everything else. At the scheduling layer, a five-stage pipeline (PreScore → Filter → Score → PostScore → Bind) supports rule-based routing that lets operators specify, for example, that low-latency small-batch requests go to Nvidia GPUs while high-throughput batch inference runs on domestic chips.

The practical impact is significant: AI service providers no longer need to maintain separate deployment environments for each chip brand. One codebase, one configuration, and traffic is intelligently distributed across a heterogeneous compute pool. This is the kind of infrastructure that makes the transition from "Nvidia-only" to "mixed deployment" economically viable.

Training vs. Inference: The Great Bifurcation

One of the clearest findings from the Maybank report is that the transition to domestic chips is splitting along workload lines. Pre-training remains largely dependent on Nvidia hardware because of interconnect limitations and weaker compute utilization at scale. The massive all-to-all communication patterns required during large-model pre-training demand high-bandwidth, low-latency interconnects—an area where Nvidia's NVLink and NVSwitch still hold a commanding advantage. China's domestic alternatives, such as Huawei's HCCS (Huawei Cache Coherence System), are improving rapidly but have not yet closed the gap at the scale of tens of thousands of chips.

Inference, however, is a different story. Large-scale inference, standard workloads, and edge and automotive applications are gradually shifting towards domestic chips. The reason is straightforward: inference workloads are more easily partitioned, less sensitive to interconnect bandwidth, and more cost-sensitive—all factors that favor domestic alternatives. When a Cambricon Siyuan 690 delivers 80% of the performance of a comparable Nvidia chip at 60-70% of the cost, the economics for inference become compelling.

This bifurcation has strategic implications. If China can handle the vast majority of inference workloads domestically—and that is the direction the data points—the remaining Nvidia dependency narrows to a specific, albeit critical, bottleneck: frontier model pre-training. And even there, LongCat-2.0 has demonstrated that the gap is closing.

Huawei's Tau Scaling Law: Rethinking Moore's Law

At the 2026 ISCAS international conference, Huawei unveiled the Tau (τ) Scaling Law, a new framework for chip performance improvement that explicitly departs from traditional Moore's Law transistor scaling. The core insight: instead of relying on ever-smaller transistor geometries—a path blocked by the EUV embargo—optimize for signal propagation delay (τ) through architectural innovation.

The Tau Law operates through three technical pathways: Logic Folding, which restructures compute logic to reduce critical path delays without requiring new process nodes; 3D vertical stacking with hybrid bonding to increase transistor density within the same footprint; and on-chip interconnect optimization that reduces wasted compute cycles and idle latency. The approach is designed to extract performance gains from mature process nodes (7nm, 14nm) that would traditionally require a node shrink to achieve.

Logic Folding is scheduled for commercial deployment in Huawei's next-generation chips in the fall of 2026, with the Kirin 2026 Pro mobile processor—expected in the Mate 90 Pro and RS models—being the first consumer-facing implementation. The implications extend beyond smartphones: the same architectural principles apply to AI accelerators, and Huawei has confirmed that Tau Law optimization is being applied across its Ascend roadmap.

"The Tau Law establishes a new paradigm for post-Moore's Law chip evolution: system-level compute efficiency through architectural reconstruction, rather than single-dimensional dependence on process node shrinks." — Huawei ISCAS 2026 Presentation

The Manufacturing Ceiling: SMIC, 7nm, and the EUV Wall

For all the progress in chip design, the manufacturing story is where the self-sufficiency narrative gets complicated. SMIC is China's designated foundry for strategic AI hardware, and its N+2 (7nm-class) process underpins both the Huawei Ascend 910C and Cambricon's Siyuan 590. But SMIC's 7nm is not TSMC's 7nm. It is a DUV-based multi-patterning process that achieves roughly comparable transistor density at significantly higher cost and lower yield. The deeper structural constraint is EUV lithography—ASML's extreme ultraviolet systems, the only viable tool for economically manufacturing below 7nm at volume, remain inaccessible to SMIC under current export controls.

Maybank's report cited Epoch's assessment that China could still be at least a decade behind in hardware, with frontier AI training feasible but materially more expensive. The "migration tax" involved in moving AI workloads from Nvidia's CUDA platform to domestic alternatives adds at least 50% to the time and cost for each engineering team. Pre-training remains the hardest problem: interconnect limitations and weaker compute utilization at scale mean that even when domestic chips are technically capable of training large models, the total cost of ownership can be significantly higher than an equivalent Nvidia deployment.

Yet the gap is narrowing faster than most Western observers predicted. Nvidia's share of China's AI chip market has collapsed from roughly 40% to an estimated 8% in 2026, while domestic vendors are forecast to account for 56% of the market—and some projections suggest the figure could reach close to 90% for high-end AI accelerators. Beijing has banned foreign AI chips in state-funded data centers and ordered projects less than 30% complete to remove installed foreign silicon. The policy trajectory is clear: the market is being structurally redirected toward domestic suppliers, regardless of the performance gap.

DimensionCurrent StatusGap to Global Frontier
AI Chip DesignCompetitive at 60-80% of Nvidia mid-range1-2 years
Inference Performance80%+ of H100 at 60-70% cost6-12 months
Pre-Training at ScaleProven at 50K-chip scale (LongCat-2.0)2-3 years
Interconnect (NVLink-equivalent)HCCS improving but not yet at NVLink scale3-4 years
HBM MemoryHiBL 1.0/HiZQ2.0 shipping; HBM3 still scarce3-4 years
Manufacturing (EUV Access)Locked at DUV multi-patterning 7nm5-10 years
Software Ecosystem (CUDA-equivalent)HeteroFlow v2 abstraction; growing but fragmented5+ years

The Self-Reinforcing Cycle

There is a dynamic at work in China's domestic chip ecosystem that Maybank's analysts identified as potentially transformative: AI-generated code is being used to accelerate the migration to domestic hardware. Automated AI coding and porting tools can now achieve portability rates as high as 90% in some cases, dramatically reducing the "migration tax" that has historically made CUDA-to-CANN (or CUDA-to-MagicMind) transitions prohibitively expensive. This creates a self-reinforcing cycle: as more developers target domestic chips, the software ecosystem improves; as the software ecosystem improves, more developers target domestic chips.

The cycle is already visible in the numbers. Cambricon's Siyuan 690 has moved from POC (proof of concept) to batch deployment at all three major Chinese cloud providers. Biren's revenue grew twenty-fold in a single year. Huawei's Ascend 910C is the de facto standard for any Chinese AI company that cannot reliably source Nvidia hardware—which, in 2026, is most of them. The Nvidia H100 one-year rental rate has risen to $2.35 per GPU-hour in March 2026, up 40% from $1.70 in October 2025, while China-specific H100 rates have climbed to RMB 80,000–90,000 (~$12,000–$13,350) per month. Nvidia's Blackwell capacity through August-September 2026 is fully booked, against a reported backlog of approximately $1 trillion through 2027. The supply constraints are structural, not cyclical, and they are accelerating the domestic transition.

The Supply Squeeze in Numbers

Nvidia H100 rental rates: $2.35/GPU-hour (March 2026), up from $1.70 (October 2025). China-specific H100 rates: RMB 80,000–90,000/month. Nvidia Blackwell backlog: ~$1 trillion through 2027. Some H100 renewal contracts extending to four years. Nvidia's China AI chip market share: collapsed from ~40% to ~8% in 2026.

What the Next Two Years Will Decide

The question posed in the title—can China's domestic AI chips handle the next generation of AI models?—does not have a binary answer. The evidence suggests a more nuanced picture: yes for inference, increasingly yes for post-training, and conditionally yes for pre-training at the trillion-parameter scale—but with higher costs, more engineering effort, and meaningful gaps in interconnect and memory bandwidth.

Several developments will determine how quickly the remaining gaps close:

First, the Tau Law's commercial deployment. If Huawei's Logic Folding and 3D stacking deliver meaningful performance improvements on mature nodes, the EUV constraint becomes less binding. The Tau Law is not a replacement for process node advancement, but it could extend the useful life of 7nm-class manufacturing for an additional generation or two—buying time for SMIC to develop alternative lithography approaches or for the geopolitical landscape to shift.

Second, the software abstraction layer's maturity. HeteroFlow v2 is a promising start, but unifying nine GPU brands under a single API is a fundamentally different challenge from making those nine brands perform optimally on every workload. The next 12-18 months will reveal whether the abstraction approach can deliver consistent, production-grade performance across the full diversity of domestic hardware, or whether fragmentation remains a drag on adoption.

Third, the HBM bottleneck. High-bandwidth memory is essential for AI accelerators, and China's domestic HBM production—led by Huawei's HiBL and HiZQ2 lines—is still ramping. The gap between domestic HBM3 supply and the explosive demand from Cambricon, Biren, and Huawei's own Ascend series is one of the most underappreciated constraints in the entire ecosystem. If HBM supply cannot keep pace with chip production, the domestic chip buildout will hit a ceiling that has nothing to do with design capability.

Fourth, the economics of scale. Cambricon's 108% revenue growth and Biren's twenty-fold increase are impressive, but they come from a small base. The question is whether these companies can sustain growth rates that support the R&D investment needed to close the gap with Nvidia's relentless product cadence. Huawei has the scale and the balance sheet to compete indefinitely. For the smaller players, the path from "promising" to "sustainable" requires navigating a brutal combination of technology risk, supply chain constraints, and the ever-present possibility that export controls could tighten further—or loosen.

Conclusion: The Gap Is Closing, but the Hardest Part Is Ahead

China's domestic AI chip industry has cleared a threshold that many Western observers—and some Chinese engineers—thought would take another five years. Meituan's LongCat-2.0 proved that a 50,000-chip domestic cluster can train a trillion-parameter model without catastrophic failure. Cambricon proved that a Chinese AI chip company can generate meaningful commercial revenue. Huawei proved that system-level integration—the CloudMatrix 384 approach—can compensate for per-chip performance gaps. HeteroFlow v2 proved that the software problem is solvable with the right abstraction layer.

But the hardest part is ahead. The gap between "training a trillion-parameter model on domestic chips" and "training frontier models at competitive cost and speed" is substantial. The EUV embargo means SMIC's process technology will remain at least a generation behind TSMC and Samsung for the foreseeable future. The CUDA ecosystem's 20-year head start in developer tooling, library optimization, and framework integration is not something that can be replicated in two or three years, no matter how many engineers are assigned to the task.

What has changed—and what makes the current moment genuinely different from previous cycles of China tech optimism—is that the domestic chip transition is no longer driven primarily by government policy. It is driven by market reality. When Nvidia's H100 costs $2.35 per GPU-hour and has a multi-year backlog, and a Cambricon Siyuan 690 delivers 80% of the performance at 60% of the cost with available supply, the economics become persuasive regardless of what any government mandates. The policy push from Beijing reinforces the market dynamics, but it did not create them.

For the global AI industry, China's domestic chip progress carries implications that extend beyond the immediate question of Nvidia's market share. A world in which China can independently manufacture, deploy, and scale competitive AI hardware is a world in which the AI supply chain is structurally bifurcated—one ecosystem built around TSMC, Nvidia, and ASML; another built around SMIC, Huawei, and a growing constellation of domestic chip designers. Whether that bifurcation leads to parallel innovation or mutual isolation is one of the defining technology questions of the decade. The answer is being written right now, on 50,000 domestic chips in a data center somewhere in China.