Building AI Infrastructure Without NVIDIA in 2026
Written and maintained by Haink's infrastructure team · Compiled from vendor documentation and primary datasheets, 5 September 2026
An AI cluster is seven layers, not one. Most conversations about alternatives to NVIDIA stop at the accelerator, which is why they usually end inconclusively — the accelerator is the layer with the most substitutes and the least leverage over whether the system works.
This page maps the whole stack, names who can supply each layer without NVIDIA silicon, and is specific about the three places where the alternative stack is genuinely weaker. It is the overview page for our Sugon and Chinese-compute coverage; each layer links to the detail.
The seven layers
| Layer | What NVIDIA supplies | The alternative |
|---|---|---|
| Host processor | Grace, or x86 from Intel/AMD | Hygon C86 (x86-64), Huawei Kunpeng (Arm) |
| Accelerator | Blackwell / Hopper | Hygon DCU, Huawei Ascend, Cambricon |
| Scale-up bus | NVLink | Sugon HSL, Huawei UnifiedBus |
| Scale-out fabric | InfiniBand / Spectrum-X | Sugon scaleFabric, RoCE over commodity Ethernet |
| Storage | Third-party (DDN, VAST, WEKA) | Sugon ParaStor, Huawei OceanStor |
| Cooling and power | Reference designs, third-party build | Sugon liquid cooling, and a wide field |
| Software | CUDA, cuDNN, NCCL, NIM | Hygon DTK (ROCm-derived), Huawei CANN, AMD ROCm |
The accelerator row is the one everyone argues about. The software row is the one that decides most projects.
Who can supply the whole stack
Two groups can deliver every layer without NVIDIA silicon anywhere in the system.
Sugon is the only vendor with all seven inside one corporate group. Hygon C86 processors and DCU accelerators, HSL as the scale-up bus, scaleFabric as the cluster network, ParaStor for storage, its own liquid cooling business, and the DTK and SothisAI software stack. Sugon co-founded Hygon and remains its largest shareholder at 27.96%, so the silicon and the system share a roadmap rather than a supplier relationship. Our page on the company sets out that structure.
Huawei is the second, and structurally similar. Kunpeng processors, Ascend accelerators, UnifiedBus, its own fabric, OceanStor storage, and CANN as the software stack. The architectural difference that matters is at the accelerator: Ascend is a domain-specific design with its own programming model, where Hygon's DCU is a general-purpose GPU with a ROCm-derived toolkit. That single choice changes what porting costs on each.
Everyone else supplies a layer, not a stack. That is not a disqualification — a cluster assembled from best-of-breed parts is a normal and often better outcome — but it means integration is the buyer's problem rather than the vendor's, and the single-throat-to-choke argument does not apply.
A worked example of the complete stack
The most fully documented non-NVIDIA system currently available is Sugon's scaleX40-3G, and it is useful precisely because every layer is visible:
| Host | 10 × Hygon C86-4G/H, 12 DDR5-6400 slots each |
| Accelerators | 40 × DeepComputing 3, 5.62 TB HBM total, ≥28 PFLOPS FP8 |
| Scale-up | HSL, 448 GB/s between any two cards, unified memory addressing across all 40 |
| Scale-out | 4 × HHHL PCIe 5.0 slots per node — scaleFabric or another cluster network |
| Storage | ParaStor, with dedicated slots for the storage network |
| Cooling and power | Cold plate and air hybrid, under 45 kW in 16U, deployable in an air-cooled room via CDU |
| Software | DTK toolchain, SothisAI platform, 800+ models pre-adapted |
Sixteen rack units, a standard 19-inch cabinet, no NVIDIA part anywhere. Whether it is the right machine is a separate question — our honest comparison against DGX H200 and B300 works that through — but as a demonstration that the full stack exists and ships, it is the clearest available.
Where the alternative stack is genuinely weaker
Three gaps, and none of them is the accelerator.
Software maturity is the real cost. Framework-level code moves with a recompile. Hand-optimised CUDA — PTX-level work, warp primitives, CUDA-only libraries — is a rewrite, and it lands on the scarcest engineers you have. Beyond the port there is a running cost that nobody budgets: version discipline. Framework packages in these environments carry vendor build tags, and installing a stock build over them breaks the device in ways that surface far from the cause. Plan for an engineer who owns the environment, not just for the migration.
There are no standardised measurements. No Chinese accelerator has been submitted to MLPerf — the April 2026 Inference v6.0 round had twenty-four submitting organisations, all on NVIDIA, AMD or Intel silicon. Developer-published tests exist and some are methodologically sound, but they are single-configuration and mostly appear in promotional material. Every cross-vendor performance claim available is vendor material or built on it, which means a proof of concept on your own workload is not optional diligence — it is the only source of a number you can act on.
The memory supply chain is the constraint nobody controls. Every accelerator in this stack depends on HBM, and HBM production is concentrated in a very small number of suppliers worldwide. Capacity per card is where the domestic designs compete hardest, and it is also the input least within their vendors' control. When evaluating roadmap commitments, that is the dependency to ask about.
A fourth, commercial rather than technical: support outside the domestic market. Warranty holder, spares location, mean time to repair at the destination and continued firmware availability are answerable questions that no datasheet addresses, and they should be settled in writing before a deployment abroad.
What the alternative stack does better
Two things, and both are real.
Pooled memory per enclosure. The design centre of the domestic superpods is more accelerators in one coherent domain rather than faster individual accelerators. A scaleX40 holds 5.62 TB across 40 cards in 16U; Huawei's CloudMatrix 384 holds 49.2 TB across 384. For a workload bounded by working-set size rather than arithmetic, that is an argument no eight-GPU box answers, and our superpod comparison sets out the numbers.
Availability and delivery. A machine you can install this quarter beats a faster machine you can install next year, and this is the argument that actually closes deals in this segment. It is a supply argument rather than a technical one, and it should be made as such.
How to decide
- Audit the software first. How much of the workload is hand-optimised CUDA? If the answer is "a lot", the hardware comparison is academic — cost the port before anything else.
- Establish the facility envelope. Kilowatts per rack position, floor loading, coolant availability. This eliminates most options before a technical argument starts.
- Decide whether you are memory-bound or compute-bound. It determines which side of the trade you want, and it is answerable from your own telemetry.
- Insist on dense figures at a named precision from every vendor, and on the two numbers vendors most often omit — memory bandwidth per accelerator and board power.
- Run a proof of concept. There is no benchmark to substitute for one. Several of these accelerators are available on public compute platforms, so it can start before a purchase order exists.
When the answer is to stay with NVIDIA
When the work is frontier training and current parts are obtainable. The compute-per-watt gap is real and no configuration closes it. If arithmetic throughput binds and the alternative is available on your timeline, take it.
When the stack is CUDA-native and the team is small. The port is not a weekend. Organisations that succeed at this have an engineer whose job it is; organisations that fail assume it is a procurement decision.
When procurement requires a benchmarked result. No standardised number exists for these systems, and no vendor material substitutes for one.
Where the alternative stack deserves serious consideration: memory-bound inference at scale, environments where current NVIDIA parts are not obtainable on an acceptable timeline, double-precision scientific work where the current NVIDIA generation has retreated, and workloads already portable to AMD Instinct, where the migration is closer to a recompile than a rewrite.
Weighing a non-NVIDIA build?
Send us the workload profile, the facility constraints and whatever configurations are on the table. We will tell you which layers actually need to change, what the software port is likely to cost, which vendor figures are committed and which are absent, and how the delivered system compares against the NVIDIA alternative doing the same work. Within one business day — and we will say plainly when the answer is to stay where you are.
Frequently asked questions
Can a complete AI cluster be built without any NVIDIA hardware?
Yes. Two vendor groups supply every layer — host processor, accelerator, scale-up bus, cluster fabric, storage, cooling and software. Sugon does it with Hygon C86 processors and DCU accelerators, HSL, scaleFabric, ParaStor and DTK; Huawei with Kunpeng, Ascend, UnifiedBus, OceanStor and CANN. Sugon's scaleX40-3G is the most fully documented example currently shipping.
What is the hardest part of moving off NVIDIA?
The software, not the accelerator. Framework-level code moves with a recompile, but hand-optimised CUDA with PTX-level work or CUDA-only library dependencies is a rewrite. There is also an ongoing cost that gets overlooked: these environments use vendor-tagged framework builds, and installing stock packages over them breaks device detection in ways that are hard to diagnose.
How do the alternatives compare on performance?
No standardised comparison exists. No Chinese accelerator has been submitted to MLPerf — the April 2026 Inference v6.0 round had twenty-four submitters, all on NVIDIA, AMD or Intel silicon. Developer measurements exist but are single-configuration and mostly published in promotional material. A proof of concept on the actual workload is the only reliable route to a number.
Where does the alternative stack actually win?
Pooled accelerator memory per enclosure, and availability. The domestic superpods are designed around more accelerators in one coherent domain rather than faster individual accelerators — 5.62 TB across 40 cards in a Sugon scaleX40, 49.2 TB across 384 in a Huawei CloudMatrix 384. For workloads bounded by working-set size rather than arithmetic, that is an argument no eight-GPU server answers.
What should I ask every vendor in this segment?
Memory bandwidth per accelerator and board power — the two figures most often omitted. Compute at a named precision with the dense-or-sparse basis stated. The size of the coherent domain and what carries traffic beyond it. And, for deployments outside the domestic market, who holds the warranty, where spares sit and what the expected time to repair is.
Is HBM supply a risk?
It is the dependency least within these vendors' control. Every accelerator in the stack needs HBM, production is concentrated among very few suppliers globally, and memory capacity per card is exactly where the domestic designs compete hardest. It is the right thing to ask about when a vendor commits to a roadmap.
Related
- Sugon (中科曙光) — the only vendor with the complete stack in one group
- Sugon scaleX40-3G · scaleX40 vs DGX H200 vs B300
- Hygon DCU · Hygon C86 · scaleFabric · ParaStor
- HSL, NVLink and InfiniBand · Superpod comparison
- Huawei enterprise IT · AI cluster architecture
Sources
- Sugon — scaleX40 product page and solution manual (full-stack composition, per-layer specifications)
- 模力方舟 — Hygon DCU platform specification (DTK ecosystem, package version conventions)
- MLCommons — MLPerf Inference v6.0 results, April 2026 (submitter list)
- The Register — Huawei CloudMatrix 384 (per-accelerator figures for the Huawei stack)
- Hygon Information Technology (corporate structure and lineage)
