Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Allocation

How NVIDIA GPU Allocation Works

Written and maintained by Haink's procurement and allocation advisory team · Updated September 2026

NVIDIA GPU allocation divides a limited production volume of current-generation data-center GPUs among supply channels and, through them, among buyers. It is not a warehouse of stock waiting to be sold. The decision doesn't rest with the seller who answers your email. It is made upstream, by the manufacturer and OEM channel, and it depends on the project rather than on the price offered. Before that decision can be made, the buyer has to disclose who the end user is, what the hardware will run and where it will be installed. The site also has to be able to power and cool the hardware. A request that gives a quantity, a price and a date but leaves out the end user and the site hasn't yet described anything that can be approved.

What allocation is — and what it is not

The market uses one word, "availability", for four different states of supply. Buyers who treat them as one thing usually find the gap between them late, often after a deposit has been paid.

Allocation sits in the middle. It is what turns a request into supply that can be executed. It is also the step a buyer can't see directly, because it happens inside the channel. The four terms, and the purchase order that follows, are taken apart one by one in allocation vs stock vs executable supply.

Allocation only matters when supply is constrained. When a product is freely available, buyers purchase it from stock or order it to build, and no one divides it up. That happens for mature and prior-generation products, and at times in a generation's life when supply catches up with demand. Allocation becomes the deciding mechanism when demand for a generation runs ahead of what can be built, packaged and integrated in a given period.

Allocation is also not the only route to compute. A workload that fits an earlier GPU generation available from stock doesn't need an allocation route at all; it still goes through the usual end-user screening. A pilot, a proof of concept or a short fine-tuning run is usually faster and cheaper on rented cloud capacity (see cloud vs private AI). And a project without a contracted, energised site isn't ready for allocation, whatever its budget. The first task for such a project is data-center capacity, not GPUs.

How does the NVIDIA GPU supply chain work?

A data-center GPU passes through six kinds of company before it runs a workload. Each one decides some things and not others. Knowing which link decides what is the fastest way to judge whether a seller's claim is plausible.

Foundry, advanced packaging and memory

The GPU die is manufactured at the foundry. It is joined to high-bandwidth memory (HBM) through advanced packaging, CoWoS in NVIDIA's case. Packaging capacity and HBM supply are contracted in advance, and together they cap how many accelerators can be completed in a given period. This level decides the total volume. It doesn't decide who gets any of it. Buyers of finished systems never deal with it directly, and no seller of systems can offer you "capacity at the foundry".

ODM: the contract manufacturers

Original design manufacturers such as Foxconn, Quanta and Wistron build GPU boards and complete systems to order. They mostly build for OEMs and for the largest cloud operators, against committed orders. An ODM decides build schedules and yields. Enterprise buyers generally can't reach one without going through a brand, and an ODM isn't an alternative channel for an ordinary purchase.

OEM: the system brands

Dell, Supermicro, HPE, Lenovo and other OEMs sell complete, supported systems under their own name, with warranty, firmware and service. This is where the decision most buyers care about is effectively made: the OEM receives a share of GPU supply and spreads it across its own customers and channel partners. The OEM also runs its own compliance review of the end user and destination. Volume at this level is committed to specific orders. Once a unit is built into a configuration for a named customer, with end-user documentation attached, it is no longer a free-floating product.

Partner tiers, distributors and NVIDIA Cloud Partners

Below the OEMs sit authorised partners at different tiers, distributors, and NVIDIA Cloud Partners (NCPs), which operate GPU capacity and sell it as a service. Partners pass buyer requests up to the OEM and pass supply down. How much volume a partner can obtain depends on its standing with the OEM and on the quality of the projects it brings. This is also the level where the same supply starts to be offered more than once. Supply is most fungible here, before a unit is committed to a named order. A buyer can check whether a company is a listed member of the NVIDIA Partner Network in NVIDIA's partner locator. A listing shows a relationship. It doesn't show that the company holds allocation for a particular order.

Integrator

The integrator designs the cluster, racks and cables it, installs it, runs acceptance tests and hands over a working system. An integrator decides whether the hardware becomes a functioning cluster on the date promised. It usually doesn't control supply. Its value lies in knowing what the site needs and making sure the order matches it.

End user

The end user is the organisation that will operate the hardware at a known site for a stated purpose. Everything upstream is ultimately a judgement about this party. The end user may be the buyer or a separate entity: a subsidiary, a customer of a cloud operator, or a public body buying through a contractor. When the two are different, both are reviewed.

What role does NVIDIA play?

NVIDIA isn't necessarily a party to the transaction. It may take part in a technical review, an ecosystem review, or an allocation and end-user review. Whether it does depends on the product, the volume, the destination country and the supply route. There is no rule that NVIDIA takes part in every order, and most buyers reach supply through an OEM or an authorised partner. NVIDIA says as much on its startup programme page: "NVIDIA doesn't sell directly to Inception members, and we don't maintain pricing control over our suppliers" (nvidia.com/startups, checked 18 September 2026).

Any claim that NVIDIA has signed off on a particular order is a statement about a specific party's decision. It should be backed by written confirmation from that party. Without the document, the claim carries no weight.

Why do ten companies offer the same GPUs?

Send one request for Blackwell systems into the market and a dozen similar replies come back within days, with similar dates and similar prices. It looks like a deep market. Usually it is the same supply seen through several windows.

The mechanism is simple. An OEM commits a share of supply to a partner. The partner mentions it to its contacts. Brokers who have never held an allocation pass the offer on, each adding a margin and presenting it as their own. None of them controls the units. The buyer sees ten sources, where there may be one source, or a claim of one.

Each extra link removes something the buyer needs. A seller three steps from the vendor of record can't prove where the units came from. It can't confirm the warranty or commit to a date. Above all, it can't tell you what the upstream party will ask about the end user, because it has never seen those questions. These chains usually break when a document is requested and the seller can't obtain it, since it isn't the one that holds the relationship. By then weeks have passed.

Comparing the prices doesn't help. The useful question is how many links separate the seller from the vendor of record, and whether the seller can name its source, state its own role and explain where your end-user documentation will go. How to test those answers before any money moves is covered in how to verify a GPU supplier. The risks of hardware sold outside authorised channels are covered in gray-market channel risks.

What decides whether an allocation request is approved?

The source of allocation judges a request as one deployment. It reads seven factors together, and they have to agree with each other. For each factor there is a question, a reason for asking it and a way to fail it.

1. Intended use

What is asked: what the hardware will do. That might be training, inference, research, a hosted service, or a mix. Why: export rules and vendor policy both turn on end use, and a civilian workload with a clear purpose is the simplest case to approve. What disqualifies: a military, intelligence or weapons-related use. A description too vague to assess, such as "AI" or "compute for clients", doesn't disqualify on its own, but it stops the request until it's clarified.

2. End-user status: is the buyer also the user?

What is asked: who will own and operate the hardware. That means the legal entity, its registration, its ultimate parent and, if the buyer isn't the operator, how the two are connected. Why: the review is of the party that ends up with the hardware, not the party that pays. A buyer purchasing for itself is the simplest case. A buyer purchasing for a named, documented end user is also workable. What disqualifies: an end user that can't be named, an opaque ultimate parent, or a party on a restricted list such as the Entity List. Requests "for a client to be identified later" can't be assessed.

3. Deployment site: address, operator, right to deploy

What is asked: the country and exact address of installation, the data-center operator or hosting provider, and whether the end user actually has the right to place hardware there, such as a signed colocation agreement or a letter of intent for capacity. Why: the address is how every other answer gets tested. Export treatment depends on the destination, and deployability depends on the site. What disqualifies: an unknown or changing destination, or a site the end user has no agreement to use.

4. Available power and cooling

What is asked: whether the site can energise and cool the requested configuration by the requested date. That covers power per rack, total available capacity, and for rack-scale systems, liquid cooling. Why: scarce supply isn't committed to hardware that will sit in a warehouse waiting for a transformer. Delivery isn't deployment. What disqualifies: a configuration the site physically can't host. That includes a liquid-cooled rack for an air-cooled hall, or a volume larger than the contracted power. Power-requirement arithmetic is covered in liquid cooling for AI servers.

5. Volume and delivery structure

What is asked: how many systems, in what configuration, and whether they can arrive in phases. Why: a volume that matches a documented workload and a real site is credible. A round number with no link to either is not. Large volumes are commonly phased, with each phase tied to capacity the site can bring online. What disqualifies: a quantity that the stated workload and site can't explain. An insistence that everything must arrive at once, whatever the site's readiness, weakens a request even if it doesn't disqualify it.

6. Geography: the end user's country and the installation country

What is asked: where the end user is based and where the hardware will be installed. These are two separate questions, because they can differ. Why: both export regulation and the vendor's own compliance apply to destinations and parties (see below). What disqualifies: any sign that the stated destination isn't the real one. Legitimate procurement never uses fictitious end users, substitute destination countries or transit arrangements designed to avoid restrictions.

7. Timing and whether it is realistic

What is asked: the target date, and what it depends on. Why: a date is a claim about the whole plan: supply, logistics, site power and acceptance. What disqualifies: nothing on its own. But a date that ignores site readiness shows the plan hasn't been checked, and that lowers confidence in everything else the request says.

These seven factors also explain why a legitimate supplier's first reply to a GPU request is a list of questions rather than a quote. A request made up of a quantity, a target price and a delivery date has answered none of the seven. The same screening of end user and destination applies to hardware from stock. Being asked these questions doesn't mean you've reached an allocation route. Not being asked them doesn't mean you've found a shortcut.

Which projects clear most easily?

Projects clear most easily when they have a clearly defined owner and operator, a specific data center, fixed end users and an understandable civilian workload. The further a project moves from that, the more review it attracts.

Type of projectClearanceWhat reviewers look at
Universities and research institutes✅ Most straightforwardCivilian AI research, science and education at a known institution and site
Enterprise internal AI: banks, healthcare, industry, telecom, large corporates✅ Most straightforwardOwn models and inference, on infrastructure the company controls
Civilian government: sovereign AI, public services, research✅ Most straightforwardThat there is no military or intelligence component
Dedicated AI / HPC data center✅ Most straightforwardEasiest when the operator, the customers and the workloads are known in advance
Neocloud / GPU-as-a-Service🟡 Possible, with controlsThe ultimate parent and end customers; customer KYC and access control; a prohibition on providing compute to restricted users; the risk of diversion and misuse
Open marketplace; resale to unknown customers🔴 HardestNo identifiable end user exists to review
Military or intelligence; weapons- or WMD-related use🔴 HardestRestricted as a category of end use
Restricted or Entity List parties; opaque ultimate parent🔴 HardestThe end user can't be cleared, or can't be identified

Why passing BIS is not the same as being shipped

Two separate layers apply to every request, and they shouldn't be confused. The first is the regulator. US export controls, administered by the Bureau of Industry and Security, set country groups, export classifications (ECCNs), licence requirements and the Entity List. The country groups are defined in Supplement No. 1 to Part 740 of the EAR (eCFR text up to date as of 16 September 2026, checked 18 September 2026). NVIDIA publishes ECCN and HS classifications by part number on its export regulations page. That page states that they are "provided for informational purposes only and should not be construed as a representation or warranty of any kind" (checked 18 September 2026). So it is a classification lookup, not evidence that a shipment will be approved.

The second layer is the vendor. Manufacturers and OEMs apply their own commercial compliance on top of the regulation, covering allocation, end users and commercial restrictions. A country can be formally unrestricted under BIS rules, and a vendor may still limit supply there or require separate approval. This layer isn't published. It reaches a buyer only through the channel, which is why a buyer can't verify a supply route alone. The exporter or supplier determines the applicable export requirements. How the two layers play out by destination is covered in export controls and dual-use IT hardware.

How does allocation work for Blackwell systems?

Blackwell changed what gets allocated. For most enterprise buyers of earlier generations, the unit received was an eight-GPU server. In the Blackwell generation it may be a server or a whole liquid-cooled rack, and the difference changes the questions above more than any benchmark does. Figures below come from NVIDIA documentation checked on 18 September 2026. Specifications and performance are compared on H100 vs H200 vs B200.

B200

The B200 is supplied the way buyers are used to: as an eight-GPU server from an OEM, or as NVIDIA's own DGX B200. The class of product is familiar, so the review is familiar too: end user, workload, site, volume. What changes is the site. NVIDIA's user guides give a maximum system power of 14.3 kW for the air-cooled, 10U DGX B200, against 10.2 kW for the DGX H100/H200. That's about 40% more power per server (14.3 ÷ 10.2 = 1.40). A hall sized for Hopper servers holds fewer Blackwell servers at the same power budget. A site plan written for the previous generation has to be redone before the request is credible.

B300

The B300, Blackwell Ultra, is also supplied as an eight-GPU server. NVIDIA's DGX B300 user guide gives 8 × 288 GB of GPU memory and a system power of 14.5 kW. For allocation, the B300 behaves much like the B200. The practical difference is timing within the generation. When a newer part and an older part of the same generation are both in the channel, the question isn't which is faster. It is which one the channel can commit for the date the site will be ready. That has to be checked at the time of the request rather than assumed.

GB200 NVL72

The GB200 NVL72 is a different class of product. The unit of supply is a rack. NVIDIA describes it as connecting "36 Grace CPUs and 72 Blackwell GPUs in a rack-scale, liquid-cooled design" with a single 72-GPU NVLink domain (product page, checked 18 September 2026). The product page gives no rack power figure. This changes how approval works. Because the hardware can't run without liquid cooling and a rack-level power feed agreed with the data-center operator, site readiness moves from a supporting detail to the first thing checked. A request for NVL72 racks at a site without liquid-cooling infrastructure fails whatever else it gets right. Integration also gets heavier: the rack arrives as a system and is accepted as a system.

GB300 NVL72

The GB300 NVL72 brings Blackwell Ultra to the same rack-scale form: 72 Blackwell Ultra GPUs and 36 Grace CPUs "in a fully liquid-cooled, rack-scale design" (product page, checked 18 September 2026). Again, no rack power figure is published there. Everything said about the GB200 NVL72 applies, with one addition. At rack scale, a phased delivery plan is effectively a plan for how many liquid-cooled rack positions the site can energise, and when. That makes the site operator almost as much a party to the request as the buyer.

What moves the delivery date?

A specific date can't be stated honestly until the supplier channel has been checked for the exact configuration, volume, destination and end user. A number given before that check is an estimate, even when it comes from a sincere seller. Any date offered earlier should be read as indicative and subject to supplier confirmation. What can be said in advance is which factors move the date, and in which direction.

How these factors combine, and why quoted dates move after they're given, is covered in NVIDIA GPU lead times.

In what order does approval happen?

The order is the same across legitimate supply routes. Qualification of the project comes first: the buyer discloses the end user, the intended use and the deployment site, and the site's power and cooling are confirmed as able to host the configuration. Detailed work with the supply channel starts only after that. This includes supplier and OEM coordination, allocation review and commercial terms. Firm dates and terms come at the end of this sequence, not at the start, which is why a request that begins by asking for them gets questions in reply.

Frequently asked questions

How does NVIDIA GPU allocation work?

Limited production of current-generation data-center GPUs is divided among supply channels and, through them, among specific buyers and deployments. The decision is made upstream by the manufacturer and OEM channel, not by the reseller. It rests on the end user, the intended use, the installation site and its power, the volume, the geography and whether all of these are consistent with each other.

Why can't I just order 1,000 B200s?

At that scale a purchase is an allocation decision about a deployment. Before committing a large share of constrained output, the channel needs to know who will operate the systems, where, for what workload, and whether the site can power them. A request with only a quantity, a price and a date answers none of that, so it can't be approved yet.

What is the difference between allocation and stock?

Stock is hardware that physically exists and can be inspected. Allocation is a committed share of future production, assigned to a particular order for a particular deployment. Stock can be verified by serial number. Allocation is confirmed only by the party that assigned it. A seller's offer is neither until the named supplier confirms it can fulfil it.

Who decides whether I get an allocation?

The manufacturer and OEM channel decide, according to the product, the volume, the destination and the supply route. NVIDIA may take part in a technical, ecosystem or end-user review, but it doesn't take part in every order. Resellers and brokers don't decide. At most they pass a request upstream, and the further they are from the vendor of record, the less they can confirm.

Why does a GPU supplier need my deployment address?

Export requirements depend on the destination, the vendor's review depends on who operates the hardware and where, and whether it can be deployed depends on the site's power and cooling. Without the address none of these can be checked. The same screening applies to hardware sold from stock, so the question is standard practice.

Does NVIDIA approve every GPU order?

No. NVIDIA may take part in a technical, ecosystem or allocation and end-user review, depending on the product, the volume, the country and the route, but there is no rule that it takes part in every order. A claim that NVIDIA signed off on a specific order should be backed by written confirmation from the party concerned.

What disqualifies an allocation request?

An end user that can't be named or is restricted, an opaque ultimate parent, or a military, intelligence or weapons-related use. So does an unknown or inconsistent destination, a site the end user has no right to use, or a configuration the site can't power or cool. Documents that contradict each other stop a legitimate process as well.

How long does the allocation process take?

There is no honest single answer before the supplier channel has been checked for the exact configuration, volume, destination and end user. Configuration, phasing, export and vendor review, the completeness of the documents and site readiness all move the date. For rack-scale systems, site energisation and acceptance often set the real timeline rather than the hardware.

How Haink runs this sequence, step by step: GPU procurement process →

Final allocation and hardware availability remain subject to manufacturer/OEM/supplier approval, applicable compliance requirements and supply availability.

Sources (all checked 18 September 2026): NVIDIA GB200 NVL72 (72 Blackwell GPUs, 36 Grace CPUs, liquid-cooled, 72-GPU NVLink domain; no rack power figure) · NVIDIA GB300 NVL72 (72 Blackwell Ultra GPUs, 36 Grace CPUs, fully liquid-cooled; no rack power figure) · DGX B200 user guide (10U, air-cooled, 14.3 kW max) · DGX B300 user guide (8 × 288 GB, 14.5 kW) · DGX H100/H200 user guide (10.2 kW max) · NVIDIA export regulations · eCFR, 15 CFR Part 740 Supp. No. 1 (up to date as of 16 September 2026) · NVIDIA Inception · NVIDIA partner locator.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsDelaware (USA) · Hong Kong · Dubai · Singapore · Mainland China