Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Compare / scaleX40 vs DGX H200 vs DGX B300

Sugon scaleX40 vs NVIDIA DGX H200 vs DGX B300: An Honest Comparison

Written and maintained by Haink's infrastructure team · Compiled from vendor datasheets and primary product documentation, 5 September 2026

These three systems are compared as boxes, not as cards. That is the only comparison that means anything here: a Sugon scaleX40-3G holds forty accelerators and a DGX holds eight, so a per-card table would answer a question nobody is buying against.

The conclusion, stated first: the scaleX40 is a memory-forward system. It carries roughly five times the accelerator memory of a DGX H200 and 2.4 times a DGX B300, with compute between the two NVIDIA systems and about three times the power draw of either. Whether that is the right trade is decided entirely by whether your workload is bounded by memory or by arithmetic — and by two figures Sugon has not published.

The three systems

 DGX H200 (8 GPU)DGX B300 (8 GPU)scaleX40-3G (40 DCU)
Accelerators8 × H200 SXM8 × B30040 × DeepComputing 3
Memory per accelerator141 GB HBM3e288 GB144 GB
Accelerator memory, total1.13 TB2.30 TB5.62 TB
Memory bandwidth per accelerator4.8 TB/s8 TB/snot published
FP8 compute, total15.8 PFLOPS dense56 PFLOPS basis unstated28 PFLOPS
Scale-up bandwidth per accelerator900 GB/s1,800 GB/s448 GB/s
Scale-up domain8 accelerators8 accelerators40 accelerators
FP64, total268 TFLOPS9.6 TFLOPSsupported, not quantified
System power10.2 kW14 kW45 kW typical, 50 kW max

Three notes on reading this table honestly.

The H200 figures are dense values from NVIDIA's datasheet. NVIDIA quotes two numbers for most precisions, one with structured sparsity applied, and the sparse figure is double. Comparing one vendor's dense number against another's sparse one doubles the apparent gap for free, so every figure here is labelled.

The B300 FP8 figure is a vendor number whose dense-or-sparse basis NVIDIA has not stated. It is used as published and flagged wherever it appears, because any conclusion resting on it is provisional.

The scaleX40's per-card memory of 144 GB is derived rather than published: Sugon states 5.62 TB across 40 accelerators, and 144 × 40 = 5,760 GiB ÷ 1,024 = 5.625 TiB. Its FP64 throughput and HBM bandwidth are not published in any Sugon document we have reviewed — the product page, the brochure, the 45-page solution manual or the 207-page user manual.

Memory-bound inference

This is the case the scaleX40 was built for, and Sugon says so: its own brochure scopes the system to inference and fine-tuning rather than to training from scratch.

Serving a large model is first a capacity question. The weights, the activations and the attention cache have to fit, and 5.62 TB in one coherent domain is a working set that neither NVIDIA system in this table can hold — 2.4 times a DGX B300 and roughly five times a DGX H200. A model or a context that does not fit in 2.3 TB has an argument here that no eight-GPU box in the same footprint can answer.

But capacity is only half of it, and the other half is missing. Token generation is bounded by memory bandwidth, not by arithmetic: the decisive figure is how fast weights stream out of HBM, and Sugon has not published it. The H200 moves 4.8 TB/s per accelerator and the B300 8 TB/s. Until the DeepComputing 3 figure exists, throughput on the scaleX40 cannot be estimated at all — for precisely the workload the vendor scopes the product to.

The 40-card domain does real work here too. A model sharded across forty accelerators on 448 GB/s with unified addressing behaves differently from the same model sharded across five DGX chassis, where most peer traffic drops onto a cluster network at roughly 50 GB/s per card. Our explainer on which interconnect layer you are actually comparing works through where that crossover sits.

Training

Shorter, because the answer is clearer. On dense FP8 the scaleX40 delivers 28 PFLOPS against a DGX B300's 56 — half — at roughly three times the power. Against a DGX H200's 15.8 PFLOPS dense it delivers 1.77 times as much, at 4.4 times the power.

Sugon does not contest this. The company scopes the product to inference and fine-tuning, and a vendor narrowing its own claims is worth taking at face value. For training a large model from scratch where current NVIDIA silicon is available, the B300 is the machine.

Sugon does publish two performance claims, stated in its solution manual against a baseline of five mainstream eight-card OAM servers: training performance improved by up to 120% and inference by up to 340%. Both are unaudited, the comparison generation is not named, and no standardised benchmark result exists for any Hygon accelerator — no Chinese accelerator appeared in MLPerf Inference v6.0 in April 2026, whose twenty-four submitters were all running NVIDIA, AMD or Intel silicon. Developer-published measurements do exist, but in promotional write-ups and in non-matching configurations. Treat the vendor figures as positioning.

Double precision, where the table inverts

The most counterintuitive row. NVIDIA's B300 carries 1.2 TFLOPS of FP64 per accelerator, against the H200's 33.5 — a 97% reduction from the B200's 37. At system level that is 9.6 TFLOPS in a DGX B300 against 268 TFLOPS in a DGX H200.

For a buyer with double-precision workloads — cryo-EM reconstruction, meteorological modelling, genomics, computational chemistry — the current NVIDIA flagship is a step backwards of nearly thirty times against the previous generation, and the competitive field is unusually empty as a result.

Sugon markets the scaleX40 into exactly those workloads, and its solution manual describes cryo-EM reconstruction as carrying a rigid requirement for double precision. The system supports FP64. How much of it, Sugon does not say — and it is the one absent figure that could turn a niche into a segment. It should be the first question on the list.

Cost per unit of capability, without any prices

We publish no price for any of these systems: Sugon does not disclose one, and quoted street prices for NVIDIA systems vary enough by region, configuration and channel that publishing them would mislead more than it helps.

The method that survives that is a threshold. Take the capability ratios from the table above, and they tell you the price multiple at which each comparison flips. Then apply your own quoted numbers.

One scaleX40 is cheaper per……if it costs less thanArithmetic
TB of accelerator memory4.97 DGX H2005.62 ÷ 1.13
PFLOPS of FP81.77 DGX H20028 ÷ 15.8 dense
TB of accelerator memory2.44 DGX B3005.62 ÷ 2.30
PFLOPS of FP80.50 DGX B30028 ÷ 56, basis unstated

Read the last row carefully: on FP8 throughput against a B300, a scaleX40 has to cost less than half a DGX B300 to be the cheaper machine. On memory against an H200 it can cost nearly five DGX H200s and still win. The same box is a bargain and an extravagance depending only on which column you are buying.

What Sugon's own price signal implies

Sugon does not publish a price, but it does publish a comparison, and it is more specific than it first appears. Its solution manual states that against the industry-mainstream five eight-card OAM server configuration, the scaleX40 delivers its performance at broadly equivalent procurement cost — and elsewhere in the same document it names NVLink 5.0 as the bus those eight-card OAM servers use.

An eight-accelerator OAM server on NVLink is a DGX- or HGX-class box. Sugon does not name a generation, so the illustration below uses the DGX H200; substitute your own comparison system and the method holds. This is Sugon's own claim about relative cost, not a price for anything.

 5 × DGX H2001 × scaleX40
Accelerator memory5.65 TB5.62 TB
FP8 compute79 PFLOPS dense28 PFLOPS
System power51 kW45 kW
Scale-up domains5 separate domains of 81 domain of 40
Rack space5 chassis16U

At the vendor's own implied parity point you get essentially the same accelerator memory — 5.62 TB against 5.65 — for about a third of the dense FP8 compute, in slightly less power and a fifth of the rack space.

Which makes the proposition unusually legible. What is actually being sold is the domain. Five DGX H200 give you five islands of eight accelerators with a cluster network between them; the scaleX40 gives you one island of forty with unified addressing across all of it. If your job needs thirty accelerators in one coherent memory space, that is worth paying compute for. If your job fits comfortably on eight, you are paying for something you will not use, and 79 PFLOPS beats 28.

Capability per kilowatt

Worth its own table, because power is the binding constraint in more data halls than budget is, and the two rows point in opposite directions.

Per kilowattDGX H200DGX B300scaleX40
Accelerator memory (TB/kW)0.1110.1640.125
FP8 compute (PFLOPS/kW)1.554.000.62

On memory per kilowatt the scaleX40 is 13% better than a DGX H200 and 24% worse than a DGX B300 — competitive, and a more favourable result than the raw 45 kW figure suggests. On compute per kilowatt it is 2.5 times worse than an H200 and 6.4 times worse than a B300, which is the number that decides the matter in a power-constrained hall running compute-bound work.

What is still missing

Three absences prevent a complete verdict, and all three are answerable by the vendor.

Ask for all three in writing, and specify dense figures for any performance number, from every vendor.

Which to choose

If your binding constraint is…ChooseBecause
Training throughputDGX B30056 PFLOPS FP8 in 14 kW; nothing else in this table is close on compute per watt
A working set above 2.3 TB in one domainscaleX405.62 TB across 40 accelerators with unified addressing — subject to the unpublished bandwidth figure
Double precisionDGX H200, today268 TFLOPS FP64 against the B300's 9.6; the scaleX40 is a candidate only once Sugon quantifies FP64
Kilowatts per rackDGX H200 or B30010.2 or 14 kW against 45; a 45 kW rack position is not universally available
A job spanning 9–40 accelerators coherentlyscaleX40One 40-card domain rather than five domains of eight stitched by a cluster network
A stack built on hand-tuned CUDAEither NVIDIA systemThe port to Hygon's DTK is real engineering work; cost it before committing

The framing to avoid is the one that gets used most: this is not a replacement for an NVIDIA system, and treating it as one produces a bad purchase in either direction. It is a differently shaped machine — more memory, less arithmetic, a larger coherent domain, considerably more power — and the question is which shape your workload has.

Comparing a real quote?

Send us the configurations and the numbers you have been given. We will run them through the thresholds on this page, tell you which figures each vendor has actually committed to, which are missing, and where the crossover sits for the job you intend to run. Within one business day — and we will say plainly when the answer is that a different machine fits the work better.

Get the quotes compared   Prefer email? sales@haink.org

Frequently asked questions

Is the Sugon scaleX40 faster than a DGX H200?

On dense FP8 compute, yes by 1.77 times — 28 PFLOPS against 15.8 — but at 45 kW against 10.2 kW, so 2.5 times worse per kilowatt. On accelerator memory it carries 5.62 TB against 1.13 TB. On scale-up bandwidth per accelerator it is slower, 448 GB/s against 900 GB/s. There is no single answer: they are differently shaped machines and the workload decides.

How much memory does each system have?

A DGX H200 has 1.13 TB of accelerator memory across eight GPUs at 141 GB each. A DGX B300 has 2.30 TB across eight at 288 GB. A scaleX40 has 5.62 TB across forty accelerators at 144 GB each — roughly five times the H200 system and 2.4 times the B300.

Which is better for inference?

It depends on whether the working set fits. Above 2.3 TB, only the scaleX40 holds the model in one coherent domain. Below it, the NVIDIA systems have far higher memory bandwidth per accelerator — 4.8 TB/s on H200 and 8 TB/s on B300 — and token generation is bounded by bandwidth rather than by capacity. Sugon has not published the DeepComputing 3 bandwidth figure, which makes a complete inference comparison impossible today.

Which is better for training?

The DGX B300, where current NVIDIA silicon is available: 56 PFLOPS of FP8 in 14 kW against the scaleX40's 28 PFLOPS in 45 kW. Sugon itself scopes the scaleX40 to inference and fine-tuning rather than training from scratch.

What about FP64 and scientific computing?

This is where the table inverts. A DGX H200 delivers 268 TFLOPS of FP64 against a DGX B300's 9.6 — NVIDIA's current generation cut FP64 by 97% from the B200. Sugon states that FP64 is supported on the scaleX40 and markets it into cryo-EM and meteorology, but publishes no figure, so it cannot be placed in the comparison. That single number is the difference between a niche and a segment.

How do I compare cost without published prices?

Use capability thresholds. One scaleX40 is cheaper per TB of memory than a DGX H200 if it costs less than 4.97 of them, and cheaper per PFLOPS of FP8 if it costs less than 1.77 of them. Against a DGX B300 the same thresholds are 2.44 and 0.50. Apply your own quoted numbers to those multiples rather than to a headline price.

What does Sugon's "five eight-card OAM servers" comparison actually mean?

Sugon's solution manual states that a scaleX40 delivers its performance at broadly equivalent procurement cost to the mainstream five eight-card OAM server configuration, and names NVLink 5.0 as the bus in those servers — which makes them DGX- or HGX-class boxes. Taking the DGX H200 as the illustration, five of them give 5.65 TB of memory and 79 PFLOPS of dense FP8 in 51 kW, against the scaleX40's 5.62 TB and 28 PFLOPS in 45 kW. Same memory, about a third of the compute, slightly less power — and five separate eight-card domains instead of one domain of forty. What is being sold at that cost point is the coherent domain.

Related

Sources

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsHong Kong · Dubai · Singapore · Mainland China · Delaware (USA)