Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Allocation / Minimum order and phased delivery

GPU Minimum Order Quantity and Phased Delivery

Written and maintained by Haink's procurement and allocation advisory team · Updated September 2026

There is no published minimum order quantity for NVIDIA data-center GPUs. NVIDIA's product pages and DGX documentation don't state one (checked 19 September 2026). The practical minimum is set by the unit in which the hardware is built and supplied: an eight-GPU server for the B200 and B300, and a 72-GPU liquid-cooled rack for the GB200 NVL72 and GB300 NVL72. At the other end of the scale the question changes. Thousands of GPUs rarely arrive on one date, because supply is committed in portions and a site can take only as much hardware as it has energised, cooled rack positions for. Phased delivery ties each tranche to that capacity, and each tranche is fixed on its own terms.

Is there a minimum order quantity for NVIDIA data-center GPUs?

Not as a published number. What exists instead is a smallest unit of supply for each product. A buyer can't order a fraction of it, so it works as the floor.

ProductSmallest unit of supplyGPUs in one unitSource (checked 19 September 2026)
B200Eight-GPU server from an OEM, or NVIDIA DGX B2008DGX B200 user guide: "8 x NVIDIA B200 GPUs"
B300Eight-GPU server from an OEM, or NVIDIA DGX B3008DGX B300 user guide: "8 x NVIDIA B300 Blackwell Ultra GPUs"
GB200 NVL72Liquid-cooled rack, accepted as one system72 (with 36 Grace CPUs)NVIDIA GB200 NVL72 product page
GB300 NVL72Liquid-cooled rack, accepted as one system72 (with 36 Grace CPUs)NVIDIA GB300 NVL72 product page

For B200 and B300 the floor is low: one server. For NVL72 systems the floor is a whole rack, and with it a liquid-cooling loop and a rack-level power feed agreed with the data-center operator. The NVL72 product pages give no rack power figure (checked 19 September 2026), so that figure has to come from the OEM for the exact configuration.

Beyond the unit of supply, individual sellers may set their own minimums, and a supplier with constrained supply may decline orders it considers too small to plan around. Those are commercial decisions of that seller, not a rule of the product. A figure presented as "NVIDIA's MOQ" should be treated as a claim to check, not as a published fact.

What sets the practical minimum for a cluster?

For a single server, only the unit of supply. For a cluster that has to run as one system, the design of the network does. NVIDIA's DGX SuperPOD reference architectures build clusters out of scalable units (SUs), each sized to fill a block of the network fabric:

Both figures are from NVIDIA's DGX SuperPOD reference architecture documents, checked 19 September 2026. They are reference designs, not purchase rules, and OEM cluster designs can use different block sizes. They explain a pattern: large clusters are ordered and delivered in blocks that the fabric was designed to accept, not in arbitrary counts. A phase that stops halfway through a network block leaves switch ports and cabling paid for but idle.

How cluster architecture and fabric sizing work in general is covered in AI cluster architecture.

Does a larger order get priority?

Not because it is larger. When supply is constrained, the source of allocation judges a request as one deployment, and volume is one of seven factors it reads together with the end user, the intended use and the site (see what decides whether an allocation request is approved). A volume that matches a documented workload and a site with contracted power is credible. A round number with no link to either is not, however large it is.

Size changes the kind of decision rather than its likelihood. A few servers are a question of whether units can be found. Thousands of GPUs are a planning decision for the channel: how much constrained supply can be committed to one project, in what order, and against which site dates. That is why large requests are expected to come with a phasing plan.

Why doesn't a large GPU order arrive all at once?

Four constraints act on a large order, and each of them usually releases capacity in portions rather than in one block.

  1. Supply is committed in portions. When demand runs ahead of production, the channel commits what it can for a given period. A first portion can often be committed sooner than the whole volume.
  2. The site is energised in stages. Data halls bring power and cooling online hall by hall or block by block. Hardware delivered ahead of energised capacity waits in storage, where it is insured, depreciating and not running.
  3. Integration and acceptance have a throughput. Racking, cabling, burn-in and acceptance testing proceed at a rate set by the site and the integrator. For NVL72 systems each rack is accepted as a system.
  4. Review and logistics run per shipment. Shipping, customs clearance and import formalities are handled shipment by shipment, and each shipment has to match the same disclosed end user, site and use.

Phasing is therefore not a concession by the supplier. It is usually the plan that fits all four constraints at once, and it is the plan a legitimate supplier would propose even if the whole volume existed today.

How is each phase tied to site energisation?

A phase should be sized to the rack positions that will be energised, cooled and accepted by the data-center operator by the time the phase arrives. The arithmetic starts from the per-unit power in the manufacturer's documentation.

A worked example with NVIDIA's own figures: the DGX B200 user guide gives 14.3 kW maximum system power per server (checked 19 September 2026). One DGX B200 scalable unit of 32 servers therefore needs up to 457.6 kW for the servers alone (32 × 14.3), before networking, storage, cooling overhead and redundancy. A phase of one scalable unit needs that much contracted, energised capacity in place on its delivery date. A phase of two needs twice as much.

For NVL72 systems the unit of planning is the liquid-cooled rack position: a rack power feed agreed with the operator and a working liquid-cooling loop. At rack scale, a phased delivery plan is in practice a schedule of how many such positions the site can bring online, and when. The operator's commissioning dates become the delivery dates.

Two consequences follow. A phase shouldn't be scheduled ahead of the site's commissioning date for its positions, because it will arrive before it can run. And a slip in energisation moves every later phase, whatever the state of supply. How liquid cooling changes site requirements is covered in liquid cooling for AI servers.

What is fixed in each tranche?

Each tranche is a delivery in its own right. The terms that are usually fixed for it:

ItemWhy it is fixed per tranche
Configuration and firmware baselineA cluster that runs as one system needs matching hardware and firmware. A change is effectively a new order with a new date.
Quantity and unitStated in servers or racks, and aligned to the network block the tranche fills.
Site and rack positionsThe tranche is delivered to capacity that exists, not to the project in general.
End user and intended useThe same as disclosed for the whole project. A tranche can't change them.
Site readiness conditionPower and cooling for the tranche's positions energised and confirmed by the operator before shipment.
Delivery point, Incoterm and importer of recordRisk, insurance and customs responsibilities pass per shipment (see Incoterms for IT hardware).
Acceptance tests and criteriaWhat counts as delivered: arrival, installation or a passed acceptance test.
Date windowIndicative until the supplier confirms that tranche. A confirmed first tranche doesn't confirm the later ones.
Payment milestoneTied to the tranche and its acceptance, not to the whole volume.
Change rules for later tranchesWhat can still change (quantity, configuration, generation) and by when, before the next tranche is committed.

The date window deserves the most attention. Only the tranche the supplier has confirmed has a date that can be relied on. Later tranches are a plan. How to read a quoted date, and why dates move, is covered in NVIDIA GPU lead times.

Can different phases use different configurations or generations?

Across separate workloads or separate network blocks, yes. Within one training fabric, generally not: a cluster built to run as one system is designed around uniform hardware, and mixing generations inside it complicates scheduling, networking and support. So configurations and generations change between phases, not within a phase. A first phase can also come from a different supply route from the main volume, for example from existing hardware to start a workload while the rest follows as a factory order (see combining supply routes).

Planning a later phase on the next generation is a reasonable choice when the site and the workload allow it. It is also a change to the order, with its own review and its own date.

What if the configuration isn't available in the quantity needed?

The usual alternatives, in no fixed order:

Reducing the first tranche is often the quickest of these, because a smaller volume can be committed sooner and the site rarely needs the full volume on day one. A change of geography is the most constrained. The installation address is part of what was reviewed, and moving it means a new review, never a way around one.

Frequently asked questions

What is the minimum order quantity for NVIDIA B200 GPUs?

NVIDIA doesn't publish one. B200 GPUs are supplied in eight-GPU servers, from OEMs or as NVIDIA DGX B200, so the practical minimum is one server of eight GPUs. Individual sellers may set their own minimums; those are commercial terms of that seller, not a rule of the product.

Can I buy a single GB200 NVL72?

The unit of supply is one rack of 72 Blackwell GPUs and 36 Grace CPUs, liquid-cooled and accepted as one system. Whether a single rack can be supplied depends on supply at the time and on the site: a liquid-cooling loop and a rack power feed agreed with the data-center operator.

Why won't a supplier deliver all my GPUs at once?

Supply is committed in portions, sites are energised in stages, integration and acceptance have a throughput, and each shipment is reviewed and cleared on its own. Hardware that arrives before energised capacity waits in storage. Phasing is usually the plan that fits all four constraints.

Does ordering more GPUs get me allocation faster?

No. Volume is judged together with the end user, the intended use and the site. A volume that matches a documented workload and contracted power is credible; a large round number with neither is not. Large orders are expected to arrive with a phasing plan tied to the site.

How big should each delivery phase be?

As big as the rack positions that will be energised, cooled and accepted by the time the phase arrives, and aligned to the network block the phase fills. NVIDIA's reference designs use scalable units of 32 DGX B200 servers or 8 DGX GB200 racks (checked 19 September 2026).

Is the date for later phases confirmed when the first phase is?

No. Only a tranche the supplier has confirmed has a date that can be relied on. Later tranches remain indicative and subject to supplier confirmation until each is confirmed in turn.

How Haink takes a phased request from qualification to executable supply and a purchase order: GPU procurement process →

Final allocation and hardware availability remain subject to manufacturer/OEM/supplier approval, applicable compliance requirements and supply availability.

This page describes general market practice. It does not state availability, lead times, prices or minimum quantities for any product or order. Any date for a specific tranche is indicative until confirmed by the supplier.

Sources (checked 19 September 2026): DGX B200 user guide (8 × B200, 14.3 kW max) · DGX B300 user guide (8 × B300) · NVIDIA GB200 NVL72 (72 GPUs, 36 Grace CPUs, liquid-cooled; no rack power figure) · NVIDIA GB300 NVL72 (72 GPUs, 36 Grace CPUs, fully liquid-cooled; no rack power figure) · DGX SuperPOD B200 reference architecture (SU of 32 DGX B200) · DGX SuperPOD GB200 reference architecture (SU of 8 DGX GB200 racks).

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsDelaware (USA) · Hong Kong · Dubai · Singapore · Mainland China