Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Technology / AI Cluster Architecture

AI Cluster Architecture — B300 and GB300 NVL72 Designs, Compute to Storage

A production AI training cluster is not a collection of GPU servers — it is a tightly integrated system where compute, networking, storage, and management planes must be designed together to achieve the GPU utilization and throughput that justifies the hardware investment. In 2026 new clusters are built on NVIDIA B300: either eight-GPU HGX B300 servers joined by an 800G fabric, or GB300 NVL72 racks where 72 GPUs share one NVLink domain. This page describes both, and the same designs on H200 and B200 where a lower entry cost or an existing fleet makes them the right choice, from a single node to a 2 MW hall.

The Four Planes of AI Cluster Architecture

Every AI training cluster has four distinct network and functional planes:

Two Ways to Build on B300: HGX Servers or GB300 NVL72 Racks

HGX B300 serversGB300 NVL72 racks
Unit of scaleOne 8-GPU serverOne 72-GPU rack, delivered as a system
NVLink domain8 GPUs72 GPUs, 130 TB/s
Between units800G InfiniBand or Spectrum-X, one ConnectX-8 per GPU800G InfiniBand or Spectrum-X between racks
Rack power~55–60 kW with four servers~132 kW; up to 192 kW busway advised
CoolingLiquid at cluster density; air-cooled servers existLiquid only
Best forMost training, fine-tuning and inference clusters; scales in steps of one serverTrillion-parameter training and serving the largest reasoning models at high concurrency

The rest of this page follows the HGX path, where the fabric is the part you design. On H200, H100 and B200 the same designs apply with 400G NDR InfiniBand (NVIDIA QM9700 switches) instead of 800G, NVLink 4.0 instead of 5.0 on Hopper, and air cooling at 20–35 kW per rack for H200. Diagrams and bills of materials for each design are on reference architectures.

Single-Node Architecture (8 GPUs)

A single 8-GPU server (HGX B300 or DGX B300; on Hopper, HGX H200 systems such as Supermicro SYS-821GE or Dell XE9680) contains:

For a single-node deployment, there is no external cluster fabric to design. The NVLink fabric is fixed hardware. A management switch (1GbE) and a connection to external storage (if not using local NVMe only) complete the single-node architecture.

Small Cluster Architecture: 4–8 Nodes (32–64 GPUs)

On B300 this is one or two liquid-cooled racks at about 55–60 kW each; on H200, two to three air-cooled racks at 20–35 kW. The fabric design below is the same for both.

InfiniBand Layer

A single leaf switch connects all nodes: on B300, an 800G Quantum-X800 InfiniBand or Spectrum-X Ethernet switch; on H200/H100, an NVIDIA QM9700 with 64 ports of NDR 400G. Each GPU server has 8 fabric ports (one per GPU), so each server occupies 8 switch ports. A 4-node cluster uses 32 ports on the leaf switch — leaving 32 ports available for future expansion or uplinks. The leaf switch provides full bisection bandwidth: every GPU can communicate with every other GPU at full link rate simultaneously.

NVIDIA SHARP (in-network computing) on the QM9790 variant offloads allreduce operations to the switch ASIC, reducing GPU compute load during gradient synchronization by 40–60% for large clusters. For 4–8 node clusters, the benefit is modest; for 16+ node clusters, SHARP provides meaningful training throughput improvement.

Storage Layer (Small Cluster)

A small cluster can use a dedicated all-flash storage appliance (NetApp AFF, Pure Storage FlashArray) connected via 100GbE or 400GbE NFS/NVMe-oF. The storage network can share the InfiniBand fabric (storage over InfiniBand using NVMe-oF/IB) or use a separate 100GbE Ethernet storage network. For 4-node clusters, a single storage appliance with 100 TB+ all-flash capacity and 20–40 GB/s throughput is typically sufficient.

Management Layer

A separate 1GbE management switch provides BMC/IPMI access to all servers and the InfiniBand switch management port. This network is used for OS deployment, firmware updates, power management, and health monitoring — completely isolated from the training fabric. Cisco Catalyst 1000 or Aruba 2530 series switches are common for the management plane.

Medium Cluster Architecture: 16–32 Nodes (128–256 GPUs)

Two-Tier Fat-Tree InfiniBand

A 16-node cluster (128 GPUs, 128 NDR 400G ports from servers) requires two leaf switches (2 × 64-port QM9700 = 128 ports for servers) with uplinks to one or two spine switches for inter-leaf communication. A non-blocking two-tier fat-tree provides full bisection bandwidth across all 128 GPUs: at full utilization, any GPU can communicate with any other GPU at the full 400G link rate without bandwidth contention at the spine.

Cable topology: each server's 8 NICs are distributed across both leaf switches (4 NICs to leaf-1, 4 NICs to leaf-2) to avoid single-switch failure taking out half each server's bandwidth. Each leaf switch has 48 server ports and 16 uplink ports to the spine layer. Spine switch: 1 × 64-port QM9700 with 32 downlink ports to leaf switches and 32 ports available for future expansion.

Parallel Storage (Medium Cluster)

128-GPU clusters require aggregate storage throughput of 30–100 GB/s to keep GPUs fed. A single all-flash NAS appliance cannot sustain this. Parallel file systems (WEKA Data Platform, IBM Spectrum Scale GPFS, Lustre) distribute data across multiple NVMe storage nodes, aggregating bandwidth linearly. Typical medium cluster storage: 4–8 WEKA nodes (each with 12× NVMe U.2 drives, 2× 100GbE uplinks) providing 60–200 GB/s aggregate read throughput to the compute nodes.

Large Cluster Architecture: 64+ Nodes (512+ GPUs)

Three-Tier Fat-Tree

At 64 nodes (512 GPUs, 512 NDR 400G ports), a two-tier fat-tree runs out of non-blocking port capacity. Three-tier fat-tree topology is required: compute nodes connect to leaf switches; leaf switches connect to aggregation/spine switches; aggregation switches connect to core switches. Maintaining full bisection bandwidth at this scale requires careful oversubscription analysis — some 512-GPU clusters use 2:1 oversubscription at the spine layer if training workloads tolerate occasional bandwidth contention during allreduce.

Rail-Optimized Topology

NVIDIA's recommended topology for large AI training clusters is "rail-optimized": each of the 8 GPUs in a server connects to a different leaf switch (8 switches per "rail" across the cluster). All servers' GPU-0 ports connect to switch-1, all GPU-1 ports to switch-2, etc. This topology maximizes allreduce performance for data-parallel training: gradient synchronization between all instances of GPU-0 across all servers goes through a single switch with no inter-switch hop, minimizing latency.

GB300 NVL72 Rack-Scale Architecture

NVIDIA GB300 NVL72 replaces the traditional cluster architecture for the largest deployments. Each rack contains 72 B300 GPUs and 36 Grace CPUs in 18 compute trays, joined by 9 NVLink switch trays into a single NVLink 5.0 domain with 130 TB/s of bandwidth and about 20 TB of HBM3e — the whole rack operates as one very large GPU for model parallelism. Racks connect to each other and to storage over 800G InfiniBand or Spectrum-X with ConnectX-8. This removes the InfiniBand bottleneck inside a rack while keeping it between racks. The site must supply about 132 kW per rack with liquid cooling designed in from the start. The earlier GB200 NVL72 uses the same design with B200 GPUs.

Reference Architecture Summary

ScaleB300 buildGPUsFabricStoragePowerH200 alternative
Entry1 × HGX B3008NVLink onlyLocal NVMe or NFS~14.5 kW1 × HGX H200, ~10.2 kW, air
Small4–8 × HGX B30032–641 leaf switch, 800GAll-flash NAS1–2 liquid racks400G NDR, 2–3 air racks
Medium16–32 × HGX B300128–2562-tier fat-tree, 800GParallel FS (WEKA / GPFS)~0.25–0.5 MW400G NDR fat-tree
Large64+ × HGX B300 or GB300 NVL72 racks512+3-tier fat-tree / rail-optimizedLarge-scale parallel FSup to ~2 MW per hall—

For scale: 2 MW holds roughly 120 HGX B300 servers (about 960 GPUs) or about 13 GB300 NVL72 racks, with headroom for networking and storage. Costs for each build are in AI server cost.

Designing a B300 or GB300 NVL72 cluster?

Tell us the workload, where it should run and when. We come back with the architecture, a costed bill of materials, a site with the power and cooling it needs and a realistic timeline.

Scope my cluster  

Related Resources

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsDelaware (USA) · Hong Kong · Dubai · Singapore · Mainland China