Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Technology / GPU cluster deployment timeline

GPU Cluster Deployment Timeline — From Order to Production

Written and maintained by Haink's procurement and allocation advisory team · Updated September 2026

Once the hardware is on site and the data hall is energised, a GPU cluster typically takes 1–3 weeks to reach production at 8–64 GPUs, 4–8 weeks at around 256 GPUs, 2–4 months at around 1,000 GPUs and 4–9 months, in phases, at 4,000–8,000 GPUs (authors' estimate of typical ranges, September 2026, indicative). Those figures cover installation, cabling, burn-in, acceptance and bring-up of the software stack. They exclude two things that usually matter more. The first is the hardware lead time, which only the supplier channel can confirm for a specific order. The second is site readiness: fitting out powered space for liquid cooling typically adds 3–6 months, and a new utility power connection is measured in years. For most large projects the date of production is set by energised capacity, not by the GPUs.

Every duration on this page is the authors' estimate of a typical range as of September 2026. The ranges are indicative, describe general practice rather than any specific project, and are not a delivery or deployment commitment for any order.

What are the stages from order to production?

A GPU cluster passes through seven stages between the decision to buy and the first production workload. Some run in sequence, some can overlap, and one of them usually has to start before anything is ordered.

StageTypical duration
authors' estimate, Sep 2026, indicative
What sets the durationCan it overlap?
1. End-user documentation and supplier review1–2 weeksCompleteness of the end-user package, the country of installation, the intended use and the supplier's own compliance reviewWith site planning, yes. With ordering, no: it comes first
2. Order to shipmentNot estimated hereConfiguration, volume, route and supply at the time; confirmed only by the supplier channel for a specific orderWith site fit-out, yes
3. Freight, export and import clearance1–3 weeks by air, excluding any export-licence reviewMode of transport, export formalities, the importer of record and customs in the destination countryWith site fit-out, yes
4. Site readiness0 if the hall is energised and contracted; 3–6 months to fit out powered space; years for a new grid connectionContracted power, cooling type, the operator's commissioning scheduleShould start before stage 2
5. Installation and cablingDays (a few servers) to several months (thousands of GPUs)Rack count, air or liquid cooling, fabric size, crew sizePhase by phase
6. Burn-in and acceptance2–4 weeks for a cluster; 2–5 days for a few serversAcceptance criteria, failure rate on arrival, time to replace faulty partsPhase by phase
7. Software stack and first production workload1–3 weeksScheduler, storage, drivers and the readiness of the workload itselfCan be prepared during stage 6

The table deliberately gives no figure for stage 2. A duration from order to shipment is a statement about a specific supply route at a specific time, and it holds only once the supplier has confirmed it in writing. What moves that stage is covered in what determines NVIDIA GPU lead times.

Which stage usually sets the production date?

For clusters above a few hundred GPUs, site readiness is usually the critical path, not the hardware. Hardware is committed and shipped in weeks to months. Power is brought online in months to years. A project that starts looking for data-center capacity after ordering GPUs typically finds that its hardware arrives before anywhere exists to run it.

Three situations differ sharply:

The practical test of available capacity is a signed contract with a data-center operator for a stated number of kilowatts, at a stated date, in rack positions that can take the chosen cooling type.

How does cluster size change the timeline?

The duration from hardware on site to production grows with scale, but not in proportion to GPU count. Larger clusters are phased, so the first capacity reaches production long before the last.

ScaleTypical formHardware on site → production
authors' estimate, Sep 2026, indicative; energised site assumed
What dominates
8–64 GPUs1–8 eight-GPU servers, often air-cooled1–3 weeksBurn-in and software stack
~256 GPUs32 eight-GPU servers or 3–4 NVL72 racks4–8 weeksFabric cabling and cluster-wide acceptance
~1,000 GPUs~128 servers or ~14 NVL72 racks2–4 monthsInstallation throughput, liquid-cooling commissioning, acceptance
4,000–8,000 GPUs500–1,000 eight-GPU servers or ~55–110 NVL72 racks, in several phases4–9 months to full production; the first phase earlierEnergisation schedule of the site, phase by phase

At the largest scale the question "how long does deployment take" splits in two: when the first phase serves production, and when the last one does. A credible plan gives both, and ties each phase to rack positions the operator will have commissioned by then. How phases are sized is covered in phased delivery and site energisation.

Why do liquid-cooled rack-scale systems take longer to bring up?

A rack-scale system such as the GB200 NVL72 or GB300 NVL72 is 72 GPUs in one liquid-cooled rack, accepted as a single system. It arrives more integrated than a set of servers, but it moves work from the rack to the facility:

As a rule of thumb, per rack, a liquid-cooled rack-scale deployment takes several times longer to bring into production than an air-cooled eight-GPU server rack (authors' estimate, September 2026, indicative). The physics of the cooling itself is covered in liquid cooling for AI servers.

What happens between delivery and production?

Delivery is not deployment. A cluster that has arrived still has to pass through four steps before it runs a production workload.

  1. Receiving and inspection. Shipments are checked against the packing list and serial numbers, and inspected for transport damage before signing for them. Damage found after signature is harder to claim.
  2. Installation. Racking, power, liquid connections where applicable, and cabling of the compute fabric, the storage network and management. At cluster scale the fabric is the largest single piece of work: thousands of optical links, each of which has to be tested.
  3. Burn-in. Every GPU, node and link runs under sustained load, at temperature, to surface the components that fail early in life. Cutting burn-in short moves those failures into production, where they interrupt training runs.
  4. Acceptance. The cluster is measured against agreed criteria: per-GPU health, interconnect bandwidth, collective-communication performance across the fabric, storage throughput, and a representative training or inference job. The date the criteria are met is the date that matters contractually.

Acceptance criteria are worth fixing in writing before the order, not after delivery. Without them, "delivered" and "working" can be argued to be the same thing, and the buyer carries the gap.

What can run in parallel, and what can't?

The shortest realistic timeline comes from overlapping stages that are independent and not overlapping those that depend on each other.

Can run in parallelHas to run in sequence
Site fit-out and hardware supplyEnd-user documentation → supplier review → order
Software stack preparation and burn-inEnergisation of rack positions → installation in them
Installation of phase 2 and acceptance of phase 1Burn-in → acceptance → production
Import formalities and freight bookingConfirmed site address → export and import paperwork

The first sequence is the one most often underestimated. The end user, the intended use and the address of installation have to be known and documented before a supplier reviews the order, so a project that has not settled its site cannot shorten its timeline by ordering early. Why those questions come first is explained in NVIDIA end-user requirements.

What typically delays a GPU cluster deployment?

What does it cost when hardware arrives before the site?

Hardware that arrives before energised capacity is not free to hold. Warranty periods often start at shipment or delivery, not at production, so months of stored hardware are months of warranty used up. The asset depreciates from the day it is capitalised, in a market where each GPU generation is superseded quickly. Storage has to meet the manufacturer's environmental conditions and has to be insured. And the capital is committed without producing anything. For a large cluster, a few months of idle hardware can cost more than the difference between two supply routes, which is why sequencing the site ahead of the order usually matters more than finding the fastest supplier.

Frequently asked questions

How long does it take to deploy a GPU cluster?

From hardware on an energised site to production: typically 1–3 weeks for 8–64 GPUs, 4–8 weeks for around 256, 2–4 months for around 1,000, and 4–9 months in phases for 4,000–8,000 (authors' estimate, September 2026, indicative). The hardware lead time and any site fit-out come on top, and site readiness is usually the longer of the two.

How long does a GPU cluster burn-in take?

Typically 2–4 weeks for a cluster and 2–5 days for a few servers (authors' estimate, September 2026, indicative). Burn-in runs every GPU, node and link under sustained load to surface early-life failures before production rather than during it.

Is the hardware or the data center the bottleneck?

Above a few hundred GPUs, usually the data center. Hardware is shipped in weeks to months; fitting out powered space typically takes 3–6 months, and a new utility connection takes years. A project with energised, contracted capacity removes the longest stage from its timeline.

Can I order GPUs before I have a data center?

No. The address of installation is part of the end-user information a supplier reviews, and it has to be known before the order enters the channel. A contracted site is also what gives a delivery date any meaning: hardware that arrives before energised capacity uses up warranty and depreciates while stored.

Why does a liquid-cooled NVL72 rack take longer to deploy than air-cooled servers?

The cooling loop has to be commissioned with the data-center operator before the rack can run, each rack is accepted as one 72-GPU system, and experienced installation crews are scarcer. Per rack, bring-up typically takes several times longer than for an air-cooled server rack (authors' estimate, September 2026, indicative).

When is a GPU cluster considered in production?

When it has passed the acceptance criteria agreed between buyer and supplier: GPU health, interconnect and fabric performance, storage throughput and a representative workload. Fixing those criteria before the order is what makes the production date verifiable.

What Haink asks before a supply request is opened, and why each item is needed: What we ask before we open a supply request →

Final allocation and hardware availability remain subject to manufacturer/OEM/supplier approval, applicable compliance requirements and supply availability.

This page describes general market practice. Durations are the authors' estimates of typical ranges as of September 2026, indicative only. The page does not state availability, lead times or prices for any product or order, and no duration on it is a delivery or deployment commitment.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsDelaware (USA) · Hong Kong · Dubai · Singapore · Mainland China