AI Cluster Architecture — B300 and GB300 NVL72 Designs, Compute to Storage
A production AI training cluster is not a collection of GPU servers — it is a tightly integrated system where compute, networking, storage, and management planes must be designed together to achieve the GPU utilization and throughput that justifies the hardware investment. In 2026 new clusters are built on NVIDIA B300: either eight-GPU HGX B300 servers joined by an 800G fabric, or GB300 NVL72 racks where 72 GPUs share one NVLink domain. This page describes both, and the same designs on H200 and B200 where a lower entry cost or an existing fleet makes them the right choice, from a single node to a 2 MW hall.
The Four Planes of AI Cluster Architecture
Every AI training cluster has four distinct network and functional planes:
- Compute plane (NVLink): GPU-to-GPU communication within a single server, provided by NVLink fabric (NVLink 4.0 on H100/H200, NVLink 5.0 on B200/B300). Non-programmable — this is hardware fabric, not a network you design.
- High-speed interconnect plane (InfiniBand or RoCE): GPU-to-GPU communication between servers — the distributed training fabric. This is the most performance-critical plane you design. Typically 400G InfiniBand NDR or 400GbE RoCEv2.
- Storage plane: Connecting GPU servers to shared dataset storage and checkpoint storage. Typically 100GbE or InfiniBand shared with the interconnect plane depending on cluster design.
- Management plane: Out-of-band BMC/IPMI access, OS management, monitoring. Typically 1GbE or 10GbE on a completely separate physical network from the training fabric.
Two Ways to Build on B300: HGX Servers or GB300 NVL72 Racks
| HGX B300 servers | GB300 NVL72 racks | |
|---|---|---|
| Unit of scale | One 8-GPU server | One 72-GPU rack, delivered as a system |
| NVLink domain | 8 GPUs | 72 GPUs, 130 TB/s |
| Between units | 800G InfiniBand or Spectrum-X, one ConnectX-8 per GPU | 800G InfiniBand or Spectrum-X between racks |
| Rack power | ~55–60 kW with four servers | ~132 kW; up to 192 kW busway advised |
| Cooling | Liquid at cluster density; air-cooled servers exist | Liquid only |
| Best for | Most training, fine-tuning and inference clusters; scales in steps of one server | Trillion-parameter training and serving the largest reasoning models at high concurrency |
The rest of this page follows the HGX path, where the fabric is the part you design. On H200, H100 and B200 the same designs apply with 400G NDR InfiniBand (NVIDIA QM9700 switches) instead of 800G, NVLink 4.0 instead of 5.0 on Hopper, and air cooling at 20–35 kW per rack for H200. Diagrams and bills of materials for each design are on reference architectures.
Single-Node Architecture (8 GPUs)
A single 8-GPU server (HGX B300 or DGX B300; on Hopper, HGX H200 systems such as Supermicro SYS-821GE or Dell XE9680) contains:
- 8 GPU dies connected via NVLink 4.0 (H100/H200) or NVLink 5.0 (B200/B300) — 900 GB/s or 1,800 GB/s bidirectional between any two GPUs
- 2 CPU sockets connected to GPUs via PCIe Gen5 (for system memory access and host-device data transfers)
- 8 InfiniBand or Ethernet NICs — each GPU has a dedicated NIC for external cluster communication: 800G ConnectX-8 on B300, 400G ConnectX-7 on H200/H100
- Local NVMe storage (30 TB in DGX) for dataset staging
For a single-node deployment, there is no external cluster fabric to design. The NVLink fabric is fixed hardware. A management switch (1GbE) and a connection to external storage (if not using local NVMe only) complete the single-node architecture.
Small Cluster Architecture: 4–8 Nodes (32–64 GPUs)
On B300 this is one or two liquid-cooled racks at about 55–60 kW each; on H200, two to three air-cooled racks at 20–35 kW. The fabric design below is the same for both.
InfiniBand Layer
A single leaf switch connects all nodes: on B300, an 800G Quantum-X800 InfiniBand or Spectrum-X Ethernet switch; on H200/H100, an NVIDIA QM9700 with 64 ports of NDR 400G. Each GPU server has 8 fabric ports (one per GPU), so each server occupies 8 switch ports. A 4-node cluster uses 32 ports on the leaf switch — leaving 32 ports available for future expansion or uplinks. The leaf switch provides full bisection bandwidth: every GPU can communicate with every other GPU at full link rate simultaneously.
NVIDIA SHARP (in-network computing) on the QM9790 variant offloads allreduce operations to the switch ASIC, reducing GPU compute load during gradient synchronization by 40–60% for large clusters. For 4–8 node clusters, the benefit is modest; for 16+ node clusters, SHARP provides meaningful training throughput improvement.
Storage Layer (Small Cluster)
A small cluster can use a dedicated all-flash storage appliance (NetApp AFF, Pure Storage FlashArray) connected via 100GbE or 400GbE NFS/NVMe-oF. The storage network can share the InfiniBand fabric (storage over InfiniBand using NVMe-oF/IB) or use a separate 100GbE Ethernet storage network. For 4-node clusters, a single storage appliance with 100 TB+ all-flash capacity and 20–40 GB/s throughput is typically sufficient.
Management Layer
A separate 1GbE management switch provides BMC/IPMI access to all servers and the InfiniBand switch management port. This network is used for OS deployment, firmware updates, power management, and health monitoring — completely isolated from the training fabric. Cisco Catalyst 1000 or Aruba 2530 series switches are common for the management plane.
Medium Cluster Architecture: 16–32 Nodes (128–256 GPUs)
Two-Tier Fat-Tree InfiniBand
A 16-node cluster (128 GPUs, 128 NDR 400G ports from servers) requires two leaf switches (2 × 64-port QM9700 = 128 ports for servers) with uplinks to one or two spine switches for inter-leaf communication. A non-blocking two-tier fat-tree provides full bisection bandwidth across all 128 GPUs: at full utilization, any GPU can communicate with any other GPU at the full 400G link rate without bandwidth contention at the spine.
Cable topology: each server's 8 NICs are distributed across both leaf switches (4 NICs to leaf-1, 4 NICs to leaf-2) to avoid single-switch failure taking out half each server's bandwidth. Each leaf switch has 48 server ports and 16 uplink ports to the spine layer. Spine switch: 1 × 64-port QM9700 with 32 downlink ports to leaf switches and 32 ports available for future expansion.
Parallel Storage (Medium Cluster)
128-GPU clusters require aggregate storage throughput of 30–100 GB/s to keep GPUs fed. A single all-flash NAS appliance cannot sustain this. Parallel file systems (WEKA Data Platform, IBM Spectrum Scale GPFS, Lustre) distribute data across multiple NVMe storage nodes, aggregating bandwidth linearly. Typical medium cluster storage: 4–8 WEKA nodes (each with 12× NVMe U.2 drives, 2× 100GbE uplinks) providing 60–200 GB/s aggregate read throughput to the compute nodes.
Large Cluster Architecture: 64+ Nodes (512+ GPUs)
Three-Tier Fat-Tree
At 64 nodes (512 GPUs, 512 NDR 400G ports), a two-tier fat-tree runs out of non-blocking port capacity. Three-tier fat-tree topology is required: compute nodes connect to leaf switches; leaf switches connect to aggregation/spine switches; aggregation switches connect to core switches. Maintaining full bisection bandwidth at this scale requires careful oversubscription analysis — some 512-GPU clusters use 2:1 oversubscription at the spine layer if training workloads tolerate occasional bandwidth contention during allreduce.
Rail-Optimized Topology
NVIDIA's recommended topology for large AI training clusters is "rail-optimized": each of the 8 GPUs in a server connects to a different leaf switch (8 switches per "rail" across the cluster). All servers' GPU-0 ports connect to switch-1, all GPU-1 ports to switch-2, etc. This topology maximizes allreduce performance for data-parallel training: gradient synchronization between all instances of GPU-0 across all servers goes through a single switch with no inter-switch hop, minimizing latency.
GB300 NVL72 Rack-Scale Architecture
NVIDIA GB300 NVL72 replaces the traditional cluster architecture for the largest deployments. Each rack contains 72 B300 GPUs and 36 Grace CPUs in 18 compute trays, joined by 9 NVLink switch trays into a single NVLink 5.0 domain with 130 TB/s of bandwidth and about 20 TB of HBM3e — the whole rack operates as one very large GPU for model parallelism. Racks connect to each other and to storage over 800G InfiniBand or Spectrum-X with ConnectX-8. This removes the InfiniBand bottleneck inside a rack while keeping it between racks. The site must supply about 132 kW per rack with liquid cooling designed in from the start. The earlier GB200 NVL72 uses the same design with B200 GPUs.
Reference Architecture Summary
| Scale | B300 build | GPUs | Fabric | Storage | Power | H200 alternative |
|---|---|---|---|---|---|---|
| Entry | 1 × HGX B300 | 8 | NVLink only | Local NVMe or NFS | ~14.5 kW | 1 × HGX H200, ~10.2 kW, air |
| Small | 4–8 × HGX B300 | 32–64 | 1 leaf switch, 800G | All-flash NAS | 1–2 liquid racks | 400G NDR, 2–3 air racks |
| Medium | 16–32 × HGX B300 | 128–256 | 2-tier fat-tree, 800G | Parallel FS (WEKA / GPFS) | ~0.25–0.5 MW | 400G NDR fat-tree |
| Large | 64+ × HGX B300 or GB300 NVL72 racks | 512+ | 3-tier fat-tree / rail-optimized | Large-scale parallel FS | up to ~2 MW per hall | — |
For scale: 2 MW holds roughly 120 HGX B300 servers (about 960 GPUs) or about 13 GB300 NVL72 racks, with headroom for networking and storage. Costs for each build are in AI server cost.
Designing a B300 or GB300 NVL72 cluster?
Tell us the workload, where it should run and when. We come back with the architecture, a costed bill of materials, a site with the power and cooling it needs and a realistic timeline.
