Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Brands / NVIDIA

How to Choose NVIDIA GPUs for AI — a Practical Buying Guide

Written and maintained by Haink's AI infrastructure team · Updated July 2026 · every GPU order export-screened

Haink supplies NVIDIA AI and data center hardware to enterprises, research institutions, cloud providers, and AI teams in Hong Kong, Dubai, and Mainland China. The NVIDIA portfolio available through Haink includes data center GPU platforms (H100, H200, B200, B300), NVIDIA L40S inference GPUs, NVIDIA RTX Ada Generation professional GPUs for workstations, NVIDIA DGX personal and enterprise AI systems, and NVIDIA InfiniBand and Ethernet networking for high-performance AI clusters.

NVIDIA is the dominant supplier of AI training and inference compute globally, with H100 and H200 installed in the majority of the world's largest AI data centers and H100 SXM5 being the most widely deployed GPU for large language model training. NVIDIA's CUDA ecosystem — the software platform that runs PyTorch, TensorFlow, and all major AI frameworks — has a 10+ year head start over competing GPU platforms, making NVIDIA GPUs the default choice for AI infrastructure in most enterprise and research deployments.

Choosing an NVIDIA GPU by Workload

Start from the workload, not the model number. This picks the right class of GPU; the catalog below has the exact platforms.

WorkloadRecommended NVIDIA GPUWhy
Large-scale training & fine-tuningH100 / H200 SXM5, or B200 / B300 SXM (8-GPU)Maximum FP8/FP4, NVLink, HBM bandwidth
Memory-bound inference (70B+)H200 (141 GB) or B200 (192 GB)Large HBM per GPU serves big models
Cost-efficient inference / servingL40S (48 GB) or RTX PRO 6000 Blackwell Server Ed. (96 GB)PCIe, air-cooled, best cost-per-token
Frontier / trillion-parameterGB200 / GB300 NVL72 (rack-scale)Single NVLink domain, liquid-cooled
On-desk model developmentDGX Spark (GB10, 128 GB unified)Personal AI supercomputer
Professional workstationRTX PRO 6000 / 5000 / 4500 / 4000 BlackwellPro visualization + local AI

What ships vs what waits (as of July 2026): Hopper (H100/H200), L40S and DGX Spark are generally available from stock; Blackwell (B200/B300, GB200/GB300 NVL72) is allocation-constrained — order early. The next architecture, Vera Rubin (VR200 NVL72), reaches volume availability in H2 2026. Prior generations (A100, V100) are legacy — and A100-class GPUs are export-restricted to some destinations.

How to Decide — a 5-Step Framework

  1. Model size & precision — estimate the memory footprint (parameters × bytes/param, plus KV cache). A 70B model at FP8 needs ~70 GB for weights alone; long-context serving adds tens of GB of KV cache. This sets the minimum GPU memory and how many GPUs per replica.
  2. Training or inference — training and fine-tuning want SXM (NVLink, HBM bandwidth); inference is usually cheaper on fewer PCIe GPUs (L40S, RTX PRO 6000 Blackwell) unless the model is memory-bound, where H200/B200 earn their price.
  3. Scale — one node (≤8 GPUs) or multi-node? Multi-node changes the build: you now need an InfiniBand or Spectrum-X fabric matched to the GPU generation, and the network — not the GPU — often becomes the bottleneck.
  4. Facility — can the data center actually power and cool it? Blackwell density usually forces direct-liquid cooling (see below). This constrains the GPU choice more often than budget does.
  5. Own vs rent — high, sustained utilization favours owning; short or spiky demand favours renting. We model the break-even against your real workload rather than a list price.

How Much GPU Do You Actually Need? The Memory Math

Most oversizing comes from guessing. The dominant cost is GPU memory, and it is largely predictable. Rule of thumb: weight memory (GB) ≈ parameters (billions) × bytes per parameter — 1 byte at FP8/INT8, 0.5 at FP4/INT4, 2 at FP16. Then add the KV cache for inference: roughly 30–60% more at long context and high concurrency.

Worked example — serving Llama-3-70B at FP8, 8K context, ~50 concurrent sessions: weights ≈ 70 GB; KV cache ≈ 40–55 GB → ~110–125 GB total. That is tight on a single H200 (141 GB), comfortable on one B200 (192 GB), or about 4× L40S (48 GB) with 4-bit weights. The same model for a handful of users fits on a single H200 or two L40S.

GPUHBM / VRAMRough largest model (FP8 inference)Best role
L40S48 GB GDDR6~30B (or ~70B at 4-bit across 2)Cost-efficient serving
RTX PRO 6000 Blackwell96 GB GDDR7~65BDense inference nodes
H100 SXM580 GB HBM3~55–60BProven training workhorse
H200 SXM5141 GB HBM3e~100BMemory-bound inference & training
B200 SXM192 GB HBM3e~130BBlackwell training
B300 SXM288 GB HBM3e~200BFrontier pre-training

Figures are planning rules of thumb, not guarantees — real footprint depends on context length, batch size, quantization and framework. We size it precisely for your model and target throughput.

Facility Reality: Power, Cooling & Fabric

The GPUs are the easy part; the building is where projects stall. Rough numbers to plan around (as of July 2026):

Haink sizes power, cooling and fabric with the GPUs, not after — see our AI infrastructure cost guide and reference architectures.

Own vs Rent — and When NVIDIA Isn't the Answer

We stay neutral on the verdict, because getting it wrong is expensive. Honestly:

See cloud vs private AI for the full own-vs-rent break-even logic.

Data Center GPUs — Hopper Generation

NVIDIA H100

NVIDIA H200

Data Center GPUs — Blackwell Generation

NVIDIA B200

NVIDIA B300 (Blackwell Ultra)

NVIDIA GB200 NVL72

NVIDIA L40S — Inference and Visualization

NVIDIA RTX PRO Blackwell — Inference and Professional GPUs

The RTX PRO Blackwell generation is the current professional line. The RTX PRO 6000 Blackwell Server Edition (96 GB GDDR7, passively cooled for servers) is the new standard for high-density inference nodes, while the RTX PRO 6000 / 5000 / 4500 / 4000 Blackwell Workstation Edition cards cover on-desk AI development and visualization. The prior RTX Ada generation (below) remains available as a value / from-stock option.

NVIDIA RTX Ada Generation — Professional Workstation GPUs (Value / From Stock)

NVIDIA DGX Systems

NVIDIA DGX Spark

NVIDIA DGX H100

NVIDIA DGX H200

NVIDIA DGX B200

NVIDIA DGX GB200 NVL72

NVIDIA InfiniBand Networking

NVIDIA InfiniBand is the dominant interconnect for AI training clusters, providing GPU-to-GPU communication bandwidth for distributed training across nodes. InfiniBand's RDMA (Remote Direct Memory Access) capability allows GPUs in different servers to communicate directly without CPU involvement, reducing communication overhead during all-reduce operations in distributed LLM training.

NVIDIA Software Stack

NVIDIA hardware value is inseparable from the CUDA software ecosystem — the primary reason NVIDIA maintains its AI infrastructure dominance:

Where Haink Supplies NVIDIA Hardware

Need pricing on NVIDIA hardware?

Get firm pricing, availability and lead times on NVIDIA — export-screened, OEM-warranted.

Get a quote   Prefer email? sales@haink.org

Related Resources

Five Expensive Mistakes We See in GPU Procurement

  1. Buying SXM for inference. Eight-GPU SXM training servers are wasted on serving workloads — L40S or RTX PRO 6000 Blackwell deliver better cost-per-token and do not need liquid cooling.
  2. Undersizing the fabric. A cluster is only as fast as its slowest link; NDR-class networking on a Blackwell build throttles the GPUs you paid for. Match the fabric to the generation.
  3. No power or cooling plan. Ordering B200/GB200 into an air-cooled hall with 10 kW racks stalls the project on day one. Confirm DLC and power budget before you order, not after.
  4. Buying at peak allocation. Committing to the newest GPU at the tightest allocation means paying the most for the longest wait. Often the prior generation ships from stock this month at a better price.
  5. Chasing gray-market “bargains.” Re-marked or non-warranted boards are the most expensive way to save. Verify serials and channel before payment (next).

Authorized Channel & Verifying Your Hardware Before Payment

GPUs are a prime target for gray-market and counterfeit supply. Haink sources NVIDIA hardware through authorized distribution only, and every unit is serial-verifiable before you pay. On request we provide the GPU and system serial numbers so you can confirm — before payment — that the hardware is genuine, factory-new (or explicitly identified as refurbished), carries valid manufacturer warranty for your region, and matches the ordered SKU. Combined with per-order export screening (end-user, end-use and destination), this protects you from counterfeit boards, re-marked GPUs, and non-warrantable gray-market stock — the most common risks in fast-moving GPU procurement.

Frequently Asked Questions

Which NVIDIA GPU should I choose — training or inference?

For large-scale training and fine-tuning, use SXM systems: H100/H200 SXM5 (Hopper) or B200/B300 SXM (Blackwell) in 8-GPU servers such as the Supermicro SYS-821GE-TNHR (8× H100/H200) or ARS-821GL-NHR (B200). For inference and model serving, the L40S (48 GB Ada) or the newer RTX PRO 6000 Blackwell Server Edition (96 GB) in standard PCIe servers are more cost-effective and don't need liquid cooling. H200's 141 GB is the memory-bound sweet spot for serving 70B+ models. Not sure? Describe the model and target throughput and we'll size it.

What are realistic lead times, and how does GPU allocation work?

In-stock items (H200 NVL cards, L40S, DGX Spark) ship within days from Hong Kong or Dubai stock. SXM systems are typically 4–10 weeks for H100, 6–12 for H200, 12–20 for B200 and 14–24 for B300 — Blackwell is allocation-constrained, so order early to lock a slot. We confirm current allocation and a delivered lead time for your exact configuration, shipping via Hong Kong (free port, no duty) or the Dubai free trade zone, and every GPU order is export-screened.

What is the difference between NVIDIA H100 and B200?

H100 (Hopper) delivers 3,958 TFLOPS FP8 with 80 GB HBM3 and NVLink 4.0. B200 (Blackwell) delivers 9,000 TFLOPS FP8 / 18,000 TFLOPS FP4 with 192 GB HBM3e and NVLink 5.0 — 2.3× more FP8 compute, 2.4× more memory, 2× more NVLink bandwidth. B200 also introduces FP4 precision and requires direct liquid cooling at full utilization. See the full H100 vs H200 vs B200 comparison.

What is NVIDIA L40S and when should I use it instead of H100?

NVIDIA L40S is a 48 GB GDDR6 Ada Lovelace GPU for AI inference and professional visualization in standard PCIe rackmount servers. L40S does not require liquid cooling and costs substantially less than H100 per GPU. It is the right choice for AI inference serving (deploying trained models to users) rather than AI training. A 1U server with two L40S GPUs (96 GB total VRAM) can serve 70B models at 4-bit quantization to small-to-medium teams at lower infrastructure cost than H100 SXM servers. For large-scale training, H100 or B200 SXM is required.

What is NVIDIA DGX Spark and who is it for?

NVIDIA DGX Spark is a personal AI supercomputer powered by the GB10 Grace Blackwell Superchip with 128 GB unified memory, capable of running 70B parameter models at full FP16 precision in a compact desktop form factor. It is designed for AI researchers, developers, and small teams who need serious local AI compute without data center infrastructure. DGX Spark plugs into a standard power outlet and ships with the complete NVIDIA AI software stack pre-installed.

Can NVIDIA GPUs be exported to Mainland China?

US export control regulations restrict export of certain high-performance NVIDIA GPUs to Mainland China, including H100, H200, A100, and similar data center GPUs above specific performance thresholds. NVIDIA has developed China-specific variants (H20, L20, L2) with reduced performance to comply with export regulations. Haink advises on currently compliant GPU server configurations available for Mainland China delivery on a per-inquiry basis, as regulations and available configurations change.

What InfiniBand switches does Haink supply for AI clusters?

Haink supplies NVIDIA QM9700 and QM9790 NDR 400G InfiniBand switches for H100 and B200 AI training cluster fabrics, and NVIDIA QM8790 HDR 200G switches for existing H100 HDR cluster deployments. InfiniBand switch procurement is coordinated alongside GPU server platform procurement for complete AI training cluster builds.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsHong Kong · Dubai · Singapore · Mainland China · Delaware (USA)