Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Compare / Chinese AI superpods

Superpod War: scaleX640 vs Huawei CloudMatrix 384 vs NVIDIA NVL72

Written and maintained by Haink's infrastructure team · Compiled from vendor documentation and published third-party analysis, 5 September 2026

Three systems, three accelerator counts, and a comparison that is quoted constantly and almost never set up correctly. This page puts the published figures side by side, separates what each vendor states from what analysts have estimated, and is explicit about the two rows that cannot honestly be filled in at all.

It also corrects something that has propagated widely: the CloudMatrix comparison in circulation is against the previous NVIDIA rack, not the current one.

What is actually being compared

A superpod (超节点) is an enclosure in which many accelerators share a single high-bandwidth, memory-coherent interconnect, rather than being split across servers stitched together by a network. The accelerator count is the number every vendor publishes, and on its own it says very little — the bandwidth holding those accelerators together, and what happens at the enclosure boundary, decide what the machine can actually do. Our explainer on which interconnect layer you are actually comparing works through why.

SystemAcceleratorsAcceleratorSpecification status
NVIDIA GB300 NVL7272Blackwell UltraFull published datasheet
Huawei CloudMatrix 384384Ascend 910CPer-accelerator figures published; system totals from third-party analysis
Sugon scaleX640640Multi-vendor compatibleAnnounced; no specification published

The layer that decides it

This is the table worth having, and as far as we can establish nobody publishes it. Each figure is the vendor's own, per accelerator, on the scale-up interconnect — the layer that determines whether a group of accelerators behaves like one large device.

Scale-up busBandwidth per acceleratorAccelerators in the domainBasis
NVLink — GB300 NVL721,800 GB/s72NVIDIA datasheet
HSL — Sugon scaleX40448 GB/s40Sugon product page
UnifiedBus — CloudMatrix 384392 GB/s384Derived: 14 UB links × 28 GB/s, both figures published by Huawei
Sugon scaleX640not published640 (domain structure unknown)

Two things fall out of it immediately.

The two Chinese buses land within 15% of each other — 448 and 392 GB/s — and both sit roughly a quarter of NVLink's per-accelerator bandwidth on the current NVIDIA rack. That is a coherent picture: the domestic designs are not competing on per-link speed, they are competing on how many accelerators they can hold in one domain.

And that is where the difference is real. 384 accelerators in one coherent domain against 72 is not a marginal advantage; it is a different class of machine for any job whose working set does not fit in the smaller domain. The scaleX40's 40 sits below both. Sugon's scaleX640 would sit above all of them, if it published a bus figure — which it does not, and until it does, 640 is a count rather than a domain.

The comparison everyone is making is against the wrong rack

The widely repeated line that CloudMatrix 384 carries "roughly 3.5 times the HBM" of NVIDIA's rack traces back to analysis published against the GB200 NVL72. NVIDIA's current rack-scale product is the GB300 NVL72, and the two differ substantially in exactly the dimension being compared.

 Memory per acceleratorTotal accelerator memory
GB200 NVL72192 GB13.8 TB
GB300 NVL72288 GB20.7 TB
CloudMatrix 384128 GB49.2 TB

Against GB200 the ratio is 3.6×. Against the rack NVIDIA actually sells now it is 2.4×. Still a large advantage in pooled capacity, and a materially different number from the one in circulation — worth checking which generation any comparison you are shown is built on before it informs a decision.

The per-accelerator column tells the other half of the story, and it runs the other way: 128 GB against 288 GB. CloudMatrix reaches its total by having many more accelerators, not larger ones. That distinction matters for how a model is partitioned, and it is invisible in the headline.

Published per-accelerator figures

Vendor figures only. Where a vendor has not published a number, the cell says so rather than carrying an estimate.

 GB300 NVL72CloudMatrix 384scaleX640
Accelerators72384640
Memory per accelerator288 GB128 GBnot published
Memory bandwidth per accelerator8 TB/s3.2 TB/s (2 × 1.6 TB/s per die)not published
Scale-up bandwidth per accelerator1,800 GB/s392 GB/snot published
16-bit compute per accelerator≈3.5 PFLOPS dense/sparse basis unstated752 TFLOPS BF16, stated densenot published
Enclosure1 rack16 racks — 12 compute, 4 networking1 cabinet
CoolingLiquidLiquidImmersion phase-change, PUE 1.04

The enclosure row deserves attention because it is routinely dropped. CloudMatrix 384 is sixteen racks, not one. Comparing it to a single NVL72 rack on footprint, on facility integration or on deployment effort is not a like-for-like comparison, whatever the accelerator counts say. Sugon's claim for the scaleX640 — 640 accelerators in a single cabinet — is a claim about density that CloudMatrix does not make.

System totals, and who produced them

Huawei publishes per-accelerator figures. The system-level totals that circulate for CloudMatrix 384 — including its power draw — come from third-party analysis, principally SemiAnalysis, rather than from Huawei. We reproduce them labelled as such, because they are the only figures available and because their provenance changes how much weight they should carry.

CloudMatrix 384, system levelFigureSource
Total accelerator memory49.2 TBAnalyst estimate — reconciles exactly with Huawei's 128 GB × 384
Total memory bandwidth1,229 TB/sAnalyst estimate — reconciles exactly with 3.2 TB/s × 384
BF16 dense compute300 PFLOPSAnalyst estimate — does not reconcile: 752 TFLOPS × 384 = 288.8 PFLOPS
System power≈559 kWAnalyst estimate; no Huawei figure published

Two of the four totals reconcile exactly with Huawei's own per-accelerator numbers, which is a good sign for the arithmetic behind them. The compute figure does not: multiplying the published per-accelerator throughput gives 288.8 PFLOPS against the 300 quoted, a gap of about 4%. Small, but it means at least one of the two numbers rests on something other than simple multiplication, and neither is a vendor system-level statement.

Power per accelerator, which reorders the ranking

System power figures are hard to compare when the systems hold 72, 384 and 640 accelerators. Per accelerator they become legible, and the result is not the one the headline numbers suggest.

SystemSystem powerPer accelerator
GB300 NVL72≈140 kW1.94 kW
CloudMatrix 384≈559 kW (estimate)1.46 kW
Sugon scaleX40<45 kW1.12 kW

The 559 kW figure is what draws attention to CloudMatrix, and at a facility level it is the right number to plan against — very few halls can host it. But per accelerator it is lower than NVIDIA's current rack, because Blackwell Ultra is a large, hot part. The Sugon pod is lower still, which is a useful corrective to the reflex that 45 kW in 16U sounds extreme: it is dense, not power-hungry per unit of silicon.

None of this says anything about performance per watt, which is a different question and one the next section explains we cannot answer.

The two rows we will not fill in

Compute, across all three. Sugon publishes a single compute figure for the scaleX40 and it is at FP8. Huawei publishes at BF16. NVIDIA publishes at both, with the dense-or-sparse basis unstated for its current generation. Setting an FP8 number against a BF16 number produces a comparison that looks precise and means nothing, and doing it is the single most common way these systems get misranked. Until the vendors publish at a common precision with a stated dense-or-sparse basis, the row stays empty.

Anything at all for the scaleX640. Memory, compute, interconnect bandwidth, power and price are all unpublished ten months after announcement. It appears in this comparison as an accelerator count and a cooling design, which is what has been disclosed. Our page on what is and is not published about the scaleX640 sets out the specific questions that would fill it in.

There is also no independent measurement of any of the three doing the same work. No Chinese accelerator has been submitted to MLPerf — the April 2026 Inference v6.0 round had twenty-four submitting organisations, all running NVIDIA, AMD or Intel silicon — so every cross-system performance claim in circulation is vendor material or analysis built on vendor material.

What each system is actually for

GB300 NVL72 is the highest-bandwidth coherent domain available: 1,800 GB/s per accelerator across 72, with 288 GB each. Where a job fits inside 72 accelerators and the parts are obtainable, nothing here competes on the interconnect. One rack, roughly 140 kW.

CloudMatrix 384 is a capacity and domain-size play. 384 accelerators in one domain and 49.2 TB pooled is a working set that neither of the others approaches, at a per-accelerator bus roughly a quarter of NVLink's. The cost is sixteen racks and a facility that can deliver and remove half a megawatt. It is a machine for organisations building a data hall around it, not for one dropped into an existing room.

scaleX640 cannot be assessed. What is published — 640 accelerators in one cabinet, immersion phase-change cooling at a claimed PUE of 1.04, compatibility with accelerators from multiple vendors — describes an ambition and a thermal design rather than a system specification.

And the system that is missing from the headline but present in most real evaluations: the scaleX40, 40 accelerators and 5.62 TB in 16U of a standard rack, deployable in an air-cooled room with a CDU, with a full published datasheet. It is a smaller machine than any of the three above and the only one an enterprise can specify today without a new building.

When none of this decides your purchase

When the job fits in eight accelerators. Most fine-tuning and a great deal of inference never leaves a single conventional server. Superpod comparisons are irrelevant to that decision, and the money is better spent on memory capacity and storage.

When the facility sets the ceiling. A 140 kW rack, a 559 kW installation and a cabinet requiring an immersion loop are three different construction projects. Establish what the building accepts first; it usually eliminates most of the table before any technical argument starts.

When the software stack is not portable. These three systems run three different toolchains. A workload tied to hand-optimised CUDA is not choosing between them on hardware grounds, and the porting cost will dominate any specification difference on this page.

Weighing a superpod against a real workload?

Send us the workload profile and whatever configurations are on the table. We will tell you which figures each vendor has actually committed to, which are analyst estimates, where the domain boundary falls for your job, and what the facility will have to provide. Within one business day — and we will say plainly when the honest answer is that a smaller machine does the work.

Get the options compared   Prefer email? sales@haink.org

Frequently asked questions

How many accelerators does each superpod hold?

NVIDIA's GB300 NVL72 holds 72 in one rack. Huawei's CloudMatrix 384 holds 384 across sixteen racks — twelve compute and four networking. Sugon's scaleX640 is stated as 640 in a single cabinet. Counts alone rank them in one order; interconnect bandwidth and enclosure footprint rank them differently.

Does CloudMatrix 384 really have 3.5 times the memory of an NVL72?

Against the GB200 NVL72, yes — 49.2 TB against 13.8 TB is 3.6×. But NVIDIA's current rack is the GB300 NVL72 at 20.7 TB, and against that the ratio is 2.4×. The widely quoted figure is built on the previous generation. Per accelerator the comparison runs the other way: 128 GB on an Ascend 910C against 288 GB on a Blackwell Ultra.

How does UnifiedBus compare with NVLink and HSL?

Huawei publishes 14 UnifiedBus links per accelerator at 28 GB/s each, which is 392 GB/s per accelerator. Sugon's HSL provides 448 GB/s in the scaleX40, and NVLink provides 1,800 GB/s on the GB300 NVL72. The two Chinese buses are within 15% of each other and both sit at roughly a quarter of NVLink's per-accelerator bandwidth — their argument is domain size rather than link speed.

Which system has the highest compute?

We do not publish a ranking, because the vendors do not publish at a common precision. Sugon quotes the scaleX40 at FP8 only, Huawei quotes CloudMatrix at BF16, and NVIDIA quotes its current generation without stating whether the figures are dense or sparse. Comparing across those bases produces a number that looks precise and means nothing.

How much power does CloudMatrix 384 use?

Approximately 559 kW by third-party estimate; Huawei has not published a system figure. That is roughly four times a GB300 NVL72's ~140 kW, but it holds more than five times the accelerators — per accelerator it works out at about 1.46 kW against NVIDIA's 1.94 kW. Both are above the Sugon scaleX40's 1.12 kW per accelerator.

Why is there no performance comparison here?

Because no independent measurement of these systems on the same workload exists. No Chinese accelerator has been submitted to MLPerf — the April 2026 Inference v6.0 round had twenty-four submitters, all on NVIDIA, AMD or Intel silicon — so every cross-system performance claim available is vendor material or analysis built on it. A proof of concept on the actual workload is the only route to a number worth acting on.

Can the scaleX640 be compared with the other two?

Not yet. Ten months after announcement, its memory, compute, interconnect bandwidth, power and price are all unpublished. It enters this comparison as an accelerator count and a cooling design. The Sugon system that can be compared today is the scaleX40, which has a full published datasheet.

Related

Sources

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsHong Kong · Dubai · Singapore · Mainland China · Delaware (USA)