Superpod War: scaleX640 vs Huawei CloudMatrix 384 vs NVIDIA NVL72
Written and maintained by Haink's infrastructure team · Compiled from vendor documentation and published third-party analysis, 5 September 2026
Three systems, three accelerator counts, and a comparison that is quoted constantly and almost never set up correctly. This page puts the published figures side by side, separates what each vendor states from what analysts have estimated, and is explicit about the two rows that cannot honestly be filled in at all.
It also corrects something that has propagated widely: the CloudMatrix comparison in circulation is against the previous NVIDIA rack, not the current one.
What is actually being compared
A superpod (超节点) is an enclosure in which many accelerators share a single high-bandwidth, memory-coherent interconnect, rather than being split across servers stitched together by a network. The accelerator count is the number every vendor publishes, and on its own it says very little — the bandwidth holding those accelerators together, and what happens at the enclosure boundary, decide what the machine can actually do. Our explainer on which interconnect layer you are actually comparing works through why.
| System | Accelerators | Accelerator | Specification status |
|---|---|---|---|
| NVIDIA GB300 NVL72 | 72 | Blackwell Ultra | Full published datasheet |
| Huawei CloudMatrix 384 | 384 | Ascend 910C | Per-accelerator figures published; system totals from third-party analysis |
| Sugon scaleX640 | 640 | Multi-vendor compatible | Announced; no specification published |
The layer that decides it
This is the table worth having, and as far as we can establish nobody publishes it. Each figure is the vendor's own, per accelerator, on the scale-up interconnect — the layer that determines whether a group of accelerators behaves like one large device.
| Scale-up bus | Bandwidth per accelerator | Accelerators in the domain | Basis |
|---|---|---|---|
| NVLink — GB300 NVL72 | 1,800 GB/s | 72 | NVIDIA datasheet |
| HSL — Sugon scaleX40 | 448 GB/s | 40 | Sugon product page |
| UnifiedBus — CloudMatrix 384 | 392 GB/s | 384 | Derived: 14 UB links × 28 GB/s, both figures published by Huawei |
| Sugon scaleX640 | not published | 640 (domain structure unknown) | — |
Two things fall out of it immediately.
The two Chinese buses land within 15% of each other — 448 and 392 GB/s — and both sit roughly a quarter of NVLink's per-accelerator bandwidth on the current NVIDIA rack. That is a coherent picture: the domestic designs are not competing on per-link speed, they are competing on how many accelerators they can hold in one domain.
And that is where the difference is real. 384 accelerators in one coherent domain against 72 is not a marginal advantage; it is a different class of machine for any job whose working set does not fit in the smaller domain. The scaleX40's 40 sits below both. Sugon's scaleX640 would sit above all of them, if it published a bus figure — which it does not, and until it does, 640 is a count rather than a domain.
The comparison everyone is making is against the wrong rack
The widely repeated line that CloudMatrix 384 carries "roughly 3.5 times the HBM" of NVIDIA's rack traces back to analysis published against the GB200 NVL72. NVIDIA's current rack-scale product is the GB300 NVL72, and the two differ substantially in exactly the dimension being compared.
| Memory per accelerator | Total accelerator memory | |
|---|---|---|
| GB200 NVL72 | 192 GB | 13.8 TB |
| GB300 NVL72 | 288 GB | 20.7 TB |
| CloudMatrix 384 | 128 GB | 49.2 TB |
Against GB200 the ratio is 3.6×. Against the rack NVIDIA actually sells now it is 2.4×. Still a large advantage in pooled capacity, and a materially different number from the one in circulation — worth checking which generation any comparison you are shown is built on before it informs a decision.
The per-accelerator column tells the other half of the story, and it runs the other way: 128 GB against 288 GB. CloudMatrix reaches its total by having many more accelerators, not larger ones. That distinction matters for how a model is partitioned, and it is invisible in the headline.
Published per-accelerator figures
Vendor figures only. Where a vendor has not published a number, the cell says so rather than carrying an estimate.
| GB300 NVL72 | CloudMatrix 384 | scaleX640 | |
|---|---|---|---|
| Accelerators | 72 | 384 | 640 |
| Memory per accelerator | 288 GB | 128 GB | not published |
| Memory bandwidth per accelerator | 8 TB/s | 3.2 TB/s (2 × 1.6 TB/s per die) | not published |
| Scale-up bandwidth per accelerator | 1,800 GB/s | 392 GB/s | not published |
| 16-bit compute per accelerator | ≈3.5 PFLOPS dense/sparse basis unstated | 752 TFLOPS BF16, stated dense | not published |
| Enclosure | 1 rack | 16 racks — 12 compute, 4 networking | 1 cabinet |
| Cooling | Liquid | Liquid | Immersion phase-change, PUE 1.04 |
The enclosure row deserves attention because it is routinely dropped. CloudMatrix 384 is sixteen racks, not one. Comparing it to a single NVL72 rack on footprint, on facility integration or on deployment effort is not a like-for-like comparison, whatever the accelerator counts say. Sugon's claim for the scaleX640 — 640 accelerators in a single cabinet — is a claim about density that CloudMatrix does not make.
System totals, and who produced them
Huawei publishes per-accelerator figures. The system-level totals that circulate for CloudMatrix 384 — including its power draw — come from third-party analysis, principally SemiAnalysis, rather than from Huawei. We reproduce them labelled as such, because they are the only figures available and because their provenance changes how much weight they should carry.
| CloudMatrix 384, system level | Figure | Source |
|---|---|---|
| Total accelerator memory | 49.2 TB | Analyst estimate — reconciles exactly with Huawei's 128 GB × 384 |
| Total memory bandwidth | 1,229 TB/s | Analyst estimate — reconciles exactly with 3.2 TB/s × 384 |
| BF16 dense compute | 300 PFLOPS | Analyst estimate — does not reconcile: 752 TFLOPS × 384 = 288.8 PFLOPS |
| System power | ≈559 kW | Analyst estimate; no Huawei figure published |
Two of the four totals reconcile exactly with Huawei's own per-accelerator numbers, which is a good sign for the arithmetic behind them. The compute figure does not: multiplying the published per-accelerator throughput gives 288.8 PFLOPS against the 300 quoted, a gap of about 4%. Small, but it means at least one of the two numbers rests on something other than simple multiplication, and neither is a vendor system-level statement.
Power per accelerator, which reorders the ranking
System power figures are hard to compare when the systems hold 72, 384 and 640 accelerators. Per accelerator they become legible, and the result is not the one the headline numbers suggest.
| System | System power | Per accelerator |
|---|---|---|
| GB300 NVL72 | ≈140 kW | 1.94 kW |
| CloudMatrix 384 | ≈559 kW (estimate) | 1.46 kW |
| Sugon scaleX40 | <45 kW | 1.12 kW |
The 559 kW figure is what draws attention to CloudMatrix, and at a facility level it is the right number to plan against — very few halls can host it. But per accelerator it is lower than NVIDIA's current rack, because Blackwell Ultra is a large, hot part. The Sugon pod is lower still, which is a useful corrective to the reflex that 45 kW in 16U sounds extreme: it is dense, not power-hungry per unit of silicon.
None of this says anything about performance per watt, which is a different question and one the next section explains we cannot answer.
The two rows we will not fill in
Compute, across all three. Sugon publishes a single compute figure for the scaleX40 and it is at FP8. Huawei publishes at BF16. NVIDIA publishes at both, with the dense-or-sparse basis unstated for its current generation. Setting an FP8 number against a BF16 number produces a comparison that looks precise and means nothing, and doing it is the single most common way these systems get misranked. Until the vendors publish at a common precision with a stated dense-or-sparse basis, the row stays empty.
Anything at all for the scaleX640. Memory, compute, interconnect bandwidth, power and price are all unpublished ten months after announcement. It appears in this comparison as an accelerator count and a cooling design, which is what has been disclosed. Our page on what is and is not published about the scaleX640 sets out the specific questions that would fill it in.
There is also no independent measurement of any of the three doing the same work. No Chinese accelerator has been submitted to MLPerf — the April 2026 Inference v6.0 round had twenty-four submitting organisations, all running NVIDIA, AMD or Intel silicon — so every cross-system performance claim in circulation is vendor material or analysis built on vendor material.
What each system is actually for
GB300 NVL72 is the highest-bandwidth coherent domain available: 1,800 GB/s per accelerator across 72, with 288 GB each. Where a job fits inside 72 accelerators and the parts are obtainable, nothing here competes on the interconnect. One rack, roughly 140 kW.
CloudMatrix 384 is a capacity and domain-size play. 384 accelerators in one domain and 49.2 TB pooled is a working set that neither of the others approaches, at a per-accelerator bus roughly a quarter of NVLink's. The cost is sixteen racks and a facility that can deliver and remove half a megawatt. It is a machine for organisations building a data hall around it, not for one dropped into an existing room.
scaleX640 cannot be assessed. What is published — 640 accelerators in one cabinet, immersion phase-change cooling at a claimed PUE of 1.04, compatibility with accelerators from multiple vendors — describes an ambition and a thermal design rather than a system specification.
And the system that is missing from the headline but present in most real evaluations: the scaleX40, 40 accelerators and 5.62 TB in 16U of a standard rack, deployable in an air-cooled room with a CDU, with a full published datasheet. It is a smaller machine than any of the three above and the only one an enterprise can specify today without a new building.
When none of this decides your purchase
When the job fits in eight accelerators. Most fine-tuning and a great deal of inference never leaves a single conventional server. Superpod comparisons are irrelevant to that decision, and the money is better spent on memory capacity and storage.
When the facility sets the ceiling. A 140 kW rack, a 559 kW installation and a cabinet requiring an immersion loop are three different construction projects. Establish what the building accepts first; it usually eliminates most of the table before any technical argument starts.
When the software stack is not portable. These three systems run three different toolchains. A workload tied to hand-optimised CUDA is not choosing between them on hardware grounds, and the porting cost will dominate any specification difference on this page.
Weighing a superpod against a real workload?
Send us the workload profile and whatever configurations are on the table. We will tell you which figures each vendor has actually committed to, which are analyst estimates, where the domain boundary falls for your job, and what the facility will have to provide. Within one business day — and we will say plainly when the honest answer is that a smaller machine does the work.
Frequently asked questions
How many accelerators does each superpod hold?
NVIDIA's GB300 NVL72 holds 72 in one rack. Huawei's CloudMatrix 384 holds 384 across sixteen racks — twelve compute and four networking. Sugon's scaleX640 is stated as 640 in a single cabinet. Counts alone rank them in one order; interconnect bandwidth and enclosure footprint rank them differently.
Does CloudMatrix 384 really have 3.5 times the memory of an NVL72?
Against the GB200 NVL72, yes — 49.2 TB against 13.8 TB is 3.6×. But NVIDIA's current rack is the GB300 NVL72 at 20.7 TB, and against that the ratio is 2.4×. The widely quoted figure is built on the previous generation. Per accelerator the comparison runs the other way: 128 GB on an Ascend 910C against 288 GB on a Blackwell Ultra.
How does UnifiedBus compare with NVLink and HSL?
Huawei publishes 14 UnifiedBus links per accelerator at 28 GB/s each, which is 392 GB/s per accelerator. Sugon's HSL provides 448 GB/s in the scaleX40, and NVLink provides 1,800 GB/s on the GB300 NVL72. The two Chinese buses are within 15% of each other and both sit at roughly a quarter of NVLink's per-accelerator bandwidth — their argument is domain size rather than link speed.
Which system has the highest compute?
We do not publish a ranking, because the vendors do not publish at a common precision. Sugon quotes the scaleX40 at FP8 only, Huawei quotes CloudMatrix at BF16, and NVIDIA quotes its current generation without stating whether the figures are dense or sparse. Comparing across those bases produces a number that looks precise and means nothing.
How much power does CloudMatrix 384 use?
Approximately 559 kW by third-party estimate; Huawei has not published a system figure. That is roughly four times a GB300 NVL72's ~140 kW, but it holds more than five times the accelerators — per accelerator it works out at about 1.46 kW against NVIDIA's 1.94 kW. Both are above the Sugon scaleX40's 1.12 kW per accelerator.
Why is there no performance comparison here?
Because no independent measurement of these systems on the same workload exists. No Chinese accelerator has been submitted to MLPerf — the April 2026 Inference v6.0 round had twenty-four submitters, all on NVIDIA, AMD or Intel silicon — so every cross-system performance claim available is vendor material or analysis built on it. A proof of concept on the actual workload is the only route to a number worth acting on.
Can the scaleX640 be compared with the other two?
Not yet. Ten months after announcement, its memory, compute, interconnect bandwidth, power and price are all unpublished. It enters this comparison as an accelerator count and a cooling design. The Sugon system that can be compared today is the scaleX40, which has a full published datasheet.
Related
- HSL, NVLink and InfiniBand — why domain size, not bandwidth, usually decides
- Sugon scaleX640 — what is published and what is not
- Sugon scaleX40-3G — the specified Sugon system, in full
- Huawei enterprise IT · Sugon (中科曙光)
- Liquid cooling for AI servers · AI cluster architecture
Sources
- NVIDIA — GB300 NVL72 (72 accelerators, 288 GB and 8 TB/s per GPU, 1,800 GB/s NVLink, rack power)
- Huawei — CloudMatrix 384 published specifications (384 Ascend 910C, 128 GB HBM and 752 TFLOPS BF16 per accelerator, 14 UnifiedBus links at 28 GB/s, 540 GB/s die-to-die, 16-rack configuration)
- The Register — CloudMatrix 384 against NVIDIA's rack-scale systems (per-accelerator figures and rack composition, with vendor-published and estimated figures distinguished)
- Tom's Hardware — CloudMatrix 384 system totals (49.2 TB, 1,229 TB/s, 300 PFLOPS BF16 dense, 559 kW, all attributed to SemiAnalysis)
- Sugon — scaleX40 product page (448 GB/s HSL peer-to-peer, 40 accelerators, system power)
- Sugon — scaleX640 announcement (640 accelerators per cabinet, multi-vendor accelerator compatibility, immersion phase-change cooling, PUE 1.04)
- MLCommons — MLPerf Inference v6.0 results, April 2026 (submitter list)
