China’s directive restricting major internet platforms from buying Nvidia’s China-compliant chips has shifted the center of gravity from single-chip supremacy to whole-rack engineering. Instead of trying to match Nvidia’s flagship parts one-to-one, Chinese providers are racing to reach system-level parity by scaling out Huawei Ascend 910C clusters, tightening data movement, and squeezing extra efficiency out of 800-volt HVDC power stages built with GaN.
For readers, that means the near-term competition is less about raw FLOPS on one processor and more about how fast, power-dense racks move and train models under real constraints such as memory bandwidth and availability.
At the same time, high bandwidth memory (HBM) has become the binding constraint for everyone building large AI clusters in China. Analysts have detailed how accelerator output is gated by available HBM stacks rather than by logic die alone, pushing vendors toward smarter software and storage tiering.
Huawei’s Unified Cache Manager (UCM) is a prominent example that routes key-value caches across HBM, DRAM, and even fast SSD, reducing pressure on limited HBM. Pair those memory tactics with China’s accelerating shift to 800 V data-center power and gallium-nitride (GaN) conversion, as outlined in NVIDIA’s 800 V HVDC design guidance and Innoscience’s full-link GaN announcement, and you get a credible path to rack-level gains that narrow the gap created by export rules.

Key Facts: China’s Nvidia Ban, Ascend AI Clusters, 800 V GaN Power, And the HBM Squeeze
- What changed: China’s internet regulator instructed major platforms to stop buying Nvidia’s China-market chips, disrupting upgrade plans that relied on H20 and RTX 6000D.
- Which Nvidia parts: H20 is a Hopper-based, export-compliant accelerator with reduced interconnect and 96 GB HBM3, while RTX 6000D is an Ada workstation card with 48 GB GDDR6 for inference and creative workloads.
- China’s near-term answer: Scale-out clusters based on Huawei Ascend 910C and new system fabrics such as CloudMatrix 384 are positioned to reach parity on selected throughput metrics at the rack level, even if single-chip performance trails Nvidia’s global flagships.
- Power efficiency assist: 800 V DC distribution and GaN converters cut copper losses and conversion steps compared with legacy 54 V systems, improving PUE and rack density.
- The hard limit: HBM availability remains the towering bottleneck for domestic accelerators; output is capped by memory stacks more than logic wafers.
- Software pressure valves: Huawei’s UCM and related caching strategies move KV caches among HBM, DRAM, and SSD, improving long-sequence throughput without more HBM.
What the Ban Actually Covers and Why It Matters
Who and What are Affected?
In early September, China’s regulator notified major platforms to halt new purchases of Nvidia’s China-market parts, including H20 and RTX 6000D. The directive effectively transforms what had been a policy preference for domestic chips into a concrete procurement rule. This move lands after a year of stop-start approvals and constrained shipments that already made planning with H20 difficult.
H20 And RTX 6000D In Brief
- H20: Hopper-generation accelerator adapted for China with toned-down interconnect and HBM3 around 96 GB, targeting training and inference at smaller scales than H100/H200.
- RTX 6000D: Ada workstation card with 48 GB GDDR6 for inference or creative work; uptake has been tepid according to market coverage.
Why this Policy Matters for Performance Stories
The ban narrows access to Nvidia’s tuned China SKUs, but it does not end AI progress inside China. It changes the comparison unit. Instead of measuring per-chip performance against Nvidia’s top devices, providers will increasingly optimize for per-rack throughput, networking, and energy, where Ascend-based clusters and power architecture choices can close the gap. It also accelerates demand for domestic alternatives and spurs investment in memory, packaging, and software that better utilize scarce HBM.

How China Targets AI Parity by Scale with Ascend Clusters
Ascend 910C And CloudMatrix 384 Explained
Huawei’s Ascend 910C is the backbone of new domestic clusters. The CloudMatrix 384 configuration interconnects 384 accelerators in a fabric pitched as a rival to Nvidia’s NVL72 rack. Public demos and analyses argue that on selected cluster-level metrics, CloudMatrix can match or outperform Nvidia’s China-market stacks by relying on more nodes, denser optics, and new topologies. The tradeoff is straightforward: you compensate for a weaker single device by adding more devices and improving interconnect and scheduling.
Throughput Versus Efficiency
Matching throughput by scale can increase power draw and cooling needs, which is why power-delivery innovations matter. Readers should keep two lenses in mind:
- Per-chip metrics such as peak TFLOPS and memory bandwidth.
- Per-rack metrics such as tokens-per-second for a specific LLM at a fixed batch and sequence length.
Software and Precision Choices that Stretch Hardware
System-level parity depends on software that keeps accelerators busy and moves less data per token. China’s stacks emphasize MindSpore and CANN, and model-side choices like FP8 can shrink bandwidth and energy per operation without a large accuracy hit.
FP8 and KV Cache Placement
Long-sequence inference is often bottlenecked by key-value caches that live in HBM. Huawei’s UCM spreads those caches across HBM, DRAM, and SSD to raise tokens-per-second on long prompts. The mechanism is well documented in vendor material and independent trade coverage.
Power Stack Gains from 800 V HVDC and GaN
Traditional AI racks distribute 54 V DC, which incurs heavier copper and more conversion stages as rack power climbs into the multi-kilowatt range. Nvidia’s own guidance now details 800 V direct-current distribution with GaN-based converters that reduce conversion steps and losses, improving PUE and power density. China has strong GaN manufacturing capacity in firms such as Innoscience, which announced a full-link GaN solution for AI data centers.
Why 54 V Runs Into Losses
As currents increase, I²R losses and cable mass scale poorly at 54 V. Higher voltage distribution reduces current for the same power, enabling thinner conductors and less conversion hardware inside the rack.
What 800 V Changes Inside a Rack
Fewer conversion steps, smaller bus bars, and denser power shelves translate into more usable compute per rack footprint and lower cooling overhead. The effect is not magic; it simply improves electrical efficiency and frees space for accelerators.

The 800-Volt Shortcut: GaN Power for Greener AI Racks
Why 54-Volt Power Architectures Hit a Wall
Most legacy AI racks distribute power at 54 V DC, which looks simple until current climbs into the hundreds of amps. High current means higher I²R losses in cables and bus bars, bulkier copper, extra conversion stages, and more heat. As GPU trays scale toward megawatt-class racks, these losses erode efficiency and rack density. Nvidia’s power engineers now recommend moving the distribution spine to 800 V DC so the same rack power flows at far lower current, cutting resistive loss and freeing space for compute.
How 800 V HVDC Changes Rack Efficiency
800 V HVDC shifts most conversion to the data-center perimeter, then steps from 800 V to board and GPU rails inside the rack using high-frequency converters. The result is fewer conversion hops, smaller copper, and higher power density for the same footprint. Nvidia and partners describe this strategy as the path to practical 1 MW+ racks, which traditional 54 V chains struggle to support without severe cabling overhead.
Fewer Conversions, Lower Loss
Each conversion stage wastes energy. Converting 13.8 kV AC to 800 V DC at the perimeter and then stepping down near the load removes intermediate AC/DC and DC/DC stages, trimming losses that accumulate across a big facility.
Higher Power Density, Smaller Copper
Higher voltage means lower current for a given wattage. That allows slimmer conductors, smaller bus bars, and denser power shelves, which translates into more GPUs per rack at a given thermal budget. STMicro, a launch partner in the 800 V ecosystem, shows board-level converters delivering kilowatt-class power in smartphone-sized footprints.
China’s GaN Advantage And Vendor Ecosystem
China’s power-electronics stack leans on gallium nitride (GaN) devices that switch efficiently at high voltage and high frequency. Innoscience says it is supplying a full-link GaN solution that pairs with Nvidia’s 800 V architecture from the rack input down to GPU rails, while industry coverage tracks broader partner momentum around the new power spine. Together, these GaN stages raise rack efficiency and reduce copper weight, which helps close the energy gap with more efficient foreign accelerators.
The Real Choke Point: HBM Supply In China’s AI Data Centers
Why HBM Governs AI Throughput
For LLM training and long-sequence inference, HBM (High-Bandwidth Memory) is usually the real limiter, not raw FLOPS. Each accelerator package needs multiple HBM stacks to feed the compute die. When HBM output is scarce, you simply cannot ship complete accelerator modules, no matter how many logic wafers you print. This is why Chinese accelerators that rely on HBM2E or HBM3 hit a wall even as local fabs ramp up device production.
Where Supply Lags Demand
Industry analyses describe HBM and fab capacity as towering bottlenecks for domestic AI accelerators in 2025. Even optimistic production ramps run into memory constraints first, then packaging throughput. That dynamic shapes deployment timelines more than front-end lithography.
Practical Workarounds: Cache Tiering and AI SSDs
Vendors are responding with software that reduces the amount of hot data trapped in HBM. Huawei’s Unified Cache Manager (UCM) spreads the key-value (KV) cache across HBM, DRAM, and fast NVMe SSD, lifting tokens-per-second on long prompts while easing pressure on limited HBM capacity. Trade coverage also points to a Huawei AI SSD tuned for UCM offload, which keeps accelerator pipelines busy when sequences expand beyond HBM limits. These are workarounds, not replacements for more HBM, but they can meaningfully raise throughput per rack.

What Performance Parity Looks Like and Where It Falls Short
Rack-Level Parity Through Scale
China’s near-term strategy is parity by scale. CloudMatrix 384 connects 384 Ascend 910C accelerators in a supernode pitched as a peer to Nvidia’s GB200 NVL72. Reporting from WAIC and independent analysis argue that on selected cluster-level metrics, the Huawei rack can meet or beat Nvidia’s, despite per-chip disadvantages, by throwing more devices at the problem and optimizing the interconnect and scheduler. This is a rack vs. rack comparison, not a one-to-one chip duel.
Where Nvidia Still Leads
Nvidia’s global flagships still dominate per-chip performance and NVLink fabric maturity. China’s public yardsticks often reference Nvidia’s China-specific parts instead, such as H20 and RTX 6000D, which are constrained by export rules and aimed at different workloads. That distinction matters when readers see headlines about “matching Nvidia.”
The Broader Domestic Ecosystem
Beyond Huawei, China’s accelerator bench includes designs like Biren BR100 and Enflame L600. BR100 integrates 64 GB of HBM2E across four stacks with multi-die packaging and heavy interposer I/O. Enflame’s L600 is reported with 144 GB on-package memory, 3.6 TB/s bandwidth, and FP8 support, a sign that the ecosystem is aligning with precision formats that lift performance per watt. Public, audited LLM benchmarks remain limited, so treat these as capability snapshots rather than final verdicts.
What To Watch Next In China’s AI Compute Race
- Huawei’s Roadmap And Supernodes: Reuters reports a stepped roadmap for Ascend 950/960/970 and Atlas 950/960 supernodes supporting 8,192 to 15,488 accelerators, with debuts targeted from Q4 2025 onward. Track whether real deployments show sustained throughput and efficiency at scale.
- China’s Xiaohong Quantum Milestones: giving compute paradigms suggest that quantum-class research may influence future AI compute scaling.
- Domestic HBM And Advanced Packaging: Watch for credible HBM pilots and capacity ramps, as well as higher-volume 2.5D/3D packaging. This is the real unlock for sustained system growth.
- Domestic Engineering Talent Pool: Growth is also driven by people and Chinese scientists have been returning to China or building at home, highlighting how China’s investment is drawing back researchers critical for both hardware and software progress.
- 800 V HVDC Production Rollouts: Look for the first production-grade 800 V racks with GaN converters in Chinese hyperscalers and how that shifts PUE and rack density.
- Software Maturity: MindSpore, CANN, and cache-tiering tools such as UCM need time in the field. Expect incremental gains in scheduler efficiency, KV placement, and precision auto-tuning.

How China Seeks AI Parity by Scale and Power Efficiency
China’s post-ban strategy is clear. Instead of chasing per-chip wins against Nvidia’s global flagships, providers are pushing rack-level parity by scaling Ascend 910C clusters, optimizing interconnects, and shifting the power spine to 800 V HVDC with GaN. This combination raises tokens per second per rack and trims electrical loss so more of the facility’s wattage becomes useful compute. The approach does not erase export limits, but it does convert engineering choices into tangible throughput for training and inference workloads in China’s AI data centers.
The decisive variable is HBM. Memory supply and advanced packaging will dictate how fast domestic accelerators ship and how large clusters can grow. Until that constraint eases, software-first tactics such as UCM cache tiering and FP8 precision will continue to stretch limited memory bandwidth and capacity. Readers who want the energy-per-token view should keep an eye on 800 V deployments and GaN converter updates, while those tracking compute density should watch Huawei’s Atlas supernodes and any credible news on domestic HBM.
