China Tech

Huawei's New 4,096-Chip AI Supercomputer Takes On Nvidia — But Not Chip-to-Chip

September 19, 20268 min read
Rows of servers in a data center

At HUAWEI CONNECT 2026 in Shanghai on September 17, Huawei's deputy chairman and rotating chairman Wang Tao took the stage and made an admission that most companies would have buried: Huawei's next-generation AI (Artificial Intelligence) chips will still be slower than Nvidia's best. Then he showed how Huawei plans to win anyway. The headline launch was the Atlas 960E SuperPoD, the world's first AI "supernode" built with NPO (Near-Packaged Optics) optical interconnect. It ties 4,096 of Huawei's upcoming Ascend 960 chips into a single system delivering 8 EFLOPS (exaFLOPS, or roughly eight billion billion calculations per second) of FP8 (8-bit floating-point) compute — a machine purpose-built to train and run the 10-trillion-parameter models now coming out of Chinese labs. The chips don't ship until 2027, and on raw single-chip performance they sit somewhere between Nvidia's H200 and B300 — roughly two to three years behind.

So the real question isn't "can Huawei match Nvidia on a chip?" It's a different one: does matching chip-to-chip even matter anymore?

What Huawei Actually Announced

Let's separate the hardware from the roadmap. The Ascend 960 chips come in two variants:

ChipJobReadiness
Ascend 960DTTrainingQ1 2027 (three quarters ahead of plan)
Ascend 960PRInferenceQ3 2027 (one quarter ahead of plan)

Both have been in lab testing for several months, according to Chinese financial outlet Caijing. On paper, both outperform the Nvidia H200 SXM but trail the Nvidia B300. The H200 entered mass shipment in the second half of 2024; the B300 followed a year later. That's the 2–3 year gap Huawei is candidly conceding.

The Atlas 960E SuperPoD is the more interesting announcement. Its key specs: 4,096 NPUs (Neural Processing Units, Huawei's term for its dedicated AI accelerators) in a single supernode, up from 1,024 cards in the previous Ascend 950 generation; 8 EFLOPS FP8 and 16 EFLOPS FP4 (4-bit floating-point) of peak compute; 1 PB (petabyte) of HBM (High Bandwidth Memory) pooled across the system; a fully liquid-cooled, orthogonal rack architecture; and 99.8% system availability, with fault-free operating time doubled versus a conventional build.

The enabling technology is Hi-ONE, Huawei's new NPO optical engine. Here's why that matters: connecting 4,096 chips the traditional way would require roughly 48,000 individual 800G optical modules. Huawei replaces those with 5,500 Hi-ONE units — 7.2 Tbit/s per engine — cutting over 550 kilowatts of power in the process. Huawei calls Hi-ONE the industry's first mass-produced NPO product and the only one with a built-in light source.

NPO stands for Near-Packaged Optics. The idea is to convert electrical signals to light as close to the chip as possible, shortening the lossy copper path, while keeping the optical engines pluggable and individually serviceable — unlike the more tightly integrated CPO (Co-Packaged Optics) approach, where a failed optical component can mean scrapping the whole package.

Multiple Atlas 960E SuperPoDs can be linked into a SuperCluster: up to 512,000 NPUs via a two-tier Clos network (a layered, non-blocking data-center topology), and a claimed 1 million NPUs with a multi-rail design.

The Supernode Bet: Systems Over Silicon

This is the strategic core of Huawei's answer to Nvidia, and it's worth understanding plainly. Training a frontier model today means wiring together around 100,000 accelerators. In a conventional cluster of 8-GPU (Graphics Processing Unit) servers, Huawei says, inter-server communication eats more than 40% of training time, crushing MFU (Model FLOPs Utilization) — the share of theoretical compute that actually gets used.

Huawei's fix is architectural rather than lithographic: bind thousands of chips into one logical machine with unified memory addressing over its in-house UnifiedBus ("灵衢") interconnect, so the software sees a single computer instead of 12,500 separate servers. The number Huawei leaned on all keynote: in its Markov Lab simulation, a 100,000-card cluster built from 4,096-card supernodes achieves 2.75x the MFU of the same 100,000 cards arranged as traditional 8-card servers.

Note the word simulation. That's a Huawei-modeled figure for shipping product, not an independently benchmarked one. But the direction of travel is real: when communication is the bottleneck, a tighter interconnect can compensate for weaker individual chips — the same logic behind Nvidia's own NVLink rack-scale systems.

Huawei Atlas 960E SuperPoDNvidia GB200 NVL72Nvidia planned NVL576
Accelerators per domain4,096 NPUs72 GPUs per rack576 GPUs across 8 racks
InterconnectUnifiedBus + NPO opticsNVLink 5Next-gen NVLink
StatusAnnounced, 2027 shippingShippingRoadmap

Huawei is effectively claiming a single interconnect domain roughly 57x larger than Nvidia's current NVL72 — and 7x larger than even Nvidia's planned NVL576. Whether the software stack can exploit that domain size efficiently is the open technical question. Catching up in a lab simulation and sustaining MFU on a live 10-trillion-parameter training run are different achievements.

Why Huawei Is Choosing This Fight Now

The context is export controls — and a domestic market that no longer has a choice to make. Since the second half of 2025, there has been no compliant route for Nvidia's most advanced AI chips into China. That didn't kill Chinese AI demand; it redirected it. IDC (International Data Corporation) data from June 2026 tells the scale story: China shipped roughly 4 million AI accelerator cards in 2025; domestic chips accounted for 41% — over 1.6 million units; and Huawei Ascend shipped several hundred thousand cards, the most among Chinese vendors.

The customer list now includes ByteDance, Tencent, Ant Group, and Meituan. ByteDance's trillion-parameter LongCat 2.0 model was trained on Ascend. Wang Tao told reporters that Ascend 910C supernodes have already passed 1,000 deployments, and that he expects some Chinese open-source models to begin native pre-training — not just porting — on Ascend 950/960 from 2026 onward. That last point is the ecosystem milestone Huawei cares about most. A chip that only runs models built for CUDA (Compute Unified Device Architecture, Nvidia's dominant software platform) is a translation device; a chip that models are natively trained on is a platform. Huawei also used the conference to announce that CANN (Compute Architecture for Neural Networks), its compute software stack, has moved to sustained community-driven open source — its pitch to developers weighing the cost of leaving Nvidia's ecosystem.

Wang Tao also committed Huawei to an annual chip cadence: Ascend 960 in 2027, Ascend 970 in 2028, Ascend 980 in 2029, with compute doubling each generation. Rhythm matters here — one of Nvidia's deepest advantages has been partners' confidence that a faster chip arrives every year like clockwork.

Where the Skepticism Belongs

For foreign readers trying to read past the keynote, three caveats deserve equal billing:

  1. Everything 960 is forward-looking. The chips arrive in 2027. The Atlas 860 air-cooled SuperPoD ships Q2 2027 and the liquid-cooled Atlas 960 in Q3. The 2.75x MFU figure is simulated. The 1-million-NPU cluster is an architecture claim, not a deployment.
  2. The manufacturing constraint hasn't gone away. Huawei designs Ascend chips, but advanced Chinese fabs still lack access to ASML's EUV (Extreme Ultraviolet) lithography. Systems engineering can route around some single-chip weakness; it can't manufacture a process node. Yield and capacity at scale remain the unspoken variables behind every "ahead of schedule" claim.
  3. The gap is measured against 2024–2025 Nvidia products. Landing between H200 and B300 in 2027 means Huawei competes with where Nvidia was — while Nvidia's Rubin generation and its Vera CPU platform move the frontier again. Supernode scale narrows the practical gap for Chinese buyers, but the technological race isn't a lap; it's a relay.

What This Means Outside China

The story for international readers isn't that Huawei is about to outsell Nvidia in Frankfurt or Virginia. Under current controls, it can't, and Ascend's ecosystem remains China-centered. The story is that the world's second-largest AI market is building a complete, independent compute stack on a different architectural philosophy: slightly weaker chips, dramatically larger interconnect domains, native domestic software, and a captive customer base of the world's largest internet companies.

That creates something the global AI market hasn't had before: a genuine second systems tradition. Until now, "AI infrastructure" effectively meant "Nvidia's roadmap." Huawei is betting that as models hit 10 trillion parameters and communication dominates training, the optimal design point shifts from the strongest individual GPU toward the most coherent system — a bet that, if its 2027 deployments validate the MFU claims, would matter well beyond China's borders.

Huawei isn't trying to build a better B300. It's trying to make the B300 the wrong unit of comparison. Whether that's engineering realism or a clever reframing of a weakness, the industry gets its answer in 2027.

Sources: Huawei official press release · Xinhua · Caijing Magazine · The Paper · National Business Daily (via company disclosures at HUAWEI CONNECT 2026, Shanghai, September 17–19, 2026). Specs and timeline are as disclosed by Huawei at launch; MFU figures are company simulations, not independent benchmarks. Shipping dates remain vendor guidance.