Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Technology / AI server cost

AI Server Cost in 2026: What a GPU Server and a GPU Cluster Cost, and Why Quotes Differ

Written and maintained by Haink's procurement and allocation advisory team · Updated September 2026

Hardware ranges reviewed 28 September 2026. GPU rental rates checked 24–27 September 2026. This page is reviewed quarterly, next review December 2026. All figures are indicative ranges, not quotes.

In September 2026 an eight-GPU NVIDIA HGX B300 server, the current platform for new clusters, costs roughly $550,000–$750,000; an HGX H200 server, the previous generation that is still a working choice for inference, costs $300,000–$400,000. A 32-GPU B300 cluster with fabric and storage runs about $2.5–3.8 million, a 32-GPU H200 cluster about $1.4–2.1 million, and keeping one eight-GPU server powered and hosted adds roughly $35,000–$145,000 a year. The ranges are wide on purpose. Two buyers asking for the same server get different numbers because the final figure is set by the project, not the part number: how many units, the country where the hardware is installed, who the end user is and what it will run, how soon it is needed, what the site can power and cool, and how it is paid for. The rest of this page shows how much each of those moves the cost and the date, so a budget can be built before anyone asks for a quote.

These are planning ranges drawn from our own quoting across the Hopper and Blackwell generations and from public market data. A firm figure needs the project details listed in what we ask before we open a supply request.

How much does an AI server cost in 2026?

The GPU generation sets the band. Within a band, the platform (OEM HGX server or NVIDIA DGX), memory, storage and network cards move the figure by another 10–20%.

ServerGPUsPower per serverIndicative cost, USDWhat to know
HGX B300 (OEM)8 × B300 288 GB~14.5 kW$550k–750kThe forward platform for new training and inference clusters: 288 GB per GPU. Supplied through allocation to documented projects; liquid cooling at cluster density. Widest spread between quotes of any platform
HGX H200 (OEM)8 × H200 SXM 141 GB~10.2 kW$300k–400kPrevious generation, still a working choice for large-model inference at a lower entry cost; air cooling possible at 20–35 kW per rack
DGX H2008 × H200 SXM~10.2 kWabout 10–15% above OEM HGXFactory-integrated, NVIDIA support; pays off when the software stack and support matter more than cost
PCIe inference node4 × H200 NVL 141 GB~4–5 kW$155k–210kH200 NVL cards from about $31k each. Fits standard air-cooled racks; the quickest path for inference. PCIe vs SXM
Also supplied
HGX B200 (OEM)8 × B200 192 GB~14.3 kW$420k–560kSupplied on allocation; most new projects now go straight to B300
HGX H100 (OEM)8 × H100 SXM 80 GB~10.2 kW$260k–360kOlder generation, supplied on request, mainly for expanding existing H100 clusters
Workstation-class inference server4–8 × RTX PRO 6000 or L40S2–5 kW$60k–250kLower cost per GPU; suited to smaller models, multi-model serving and fine-tuning

These ranges are for hardware ordered for a project, not bought off a shelf. In September 2026 the two are different markets. An H200 card ordered for a documented project starts at about $31k; the same card from stock, where any can be found, costs roughly 1.4–1.7 times as much, and stock in cluster quantities is close to impossible to source. The dependable route to a cluster is a project order supplied through allocation, which is why the end-user and site information is worth preparing early: it is what the order waits for, not the factory. How the two routes differ is set out in stock vs built-to-order and allocation vs stock vs executable supply.

Power figures are NVIDIA's system maximums for the equivalent DGX, as collected on our GPU cluster power requirements page. For how the generations differ in memory and throughput, see H100 vs H200 vs B200 vs B300.

What makes up the cost of one AI server?

On an eight-GPU server the accelerators are most of the bill, which is why the GPU generation matters more than the brand of the chassis. Typical shares for an HGX B300 configuration:

Two consequences for a budget. First, trimming memory or storage saves little; choosing H200, or PCIe H200 NVL, instead of B300 for an inference workload that fits in their memory saves a lot. Second, a quote that is far below the band is almost always a quote for something else: a different GPU variant, used or refurbished parts, a server without GPUs, or a price that has not yet been checked against a supplier. How to verify a GPU supplier covers the checks.

Why do two quotes for the same AI server differ?

Because the price belongs to an order, not to a product. These are the six things that move it, in the order they usually matter:

FactorEffect on costEffect on the dateWhat settles it
1. GPU generation and platformSets the band: a B300 server costs roughly twice an H200 serverB300 follows allocation to documented projects; H200 and PCIe orders usually move fasterThe workload: model size, context length, training or inference
2. QuantityA single server carries the full cost of sourcing, freight and integration; from roughly 16 servers upward, per-unit cost falls and allocation is requested per projectLarger orders are reviewed more closely and may ship in phasesA deployment plan: how many GPUs now, how many later
3. Country of installationFreight, insurance, import duty and VAT or GST, local installation. Together often 3–12% on top of hardwareWhere a US export licence is required, weeks to months; Hong Kong, for example, is treated like mainland China for advanced GPUsThe installation address, fixed before a supplier commits
4. End user and intended useIndirect: a well-documented project gets supplier pricing, an unclear one gets a risk premium or no quoteEnd-user documentation and supplier review typically take 1–2 weeksThe end-user statement and a plain description of the workload
5. TimingThe largest single swing: stock costs roughly 1.4–1.7 times a project order for the same GPU, and in cluster quantities is rarely available at allStock, when it exists: days to two weeks. Built to order: depends on the channel's commitmentWhether the site is ready; hardware that arrives before power is idle capital
6. Site and coolingLiquid cooling adds coolant distribution units and manifolds; a site without high-density power adds colocation or a retrofitThe site, not the hardware, sets the go-live date in most cluster projectsRack density in kW, cooling type and the date the hall is energized

Payment terms are a seventh factor that sits on top: an order paid on delivery, one with a deposit, and one that is financed are priced differently because they carry different costs of capital.

How much does a GPU cluster cost?

A cluster adds three things to the servers: the fabric that connects them, shared storage for datasets and checkpoints, and the racks, power distribution and cabling. Together they usually add 12–25% to the server cost; the share falls as clusters grow. Five reference builds, current platform first:

BuildWhat is in itIT powerIndicative hardware costPer GPU, all-in
32-GPU B300 cluster4 × HGX B300, 800G InfiniBand or Spectrum-X, ~200 TB all-flash, liquid-cooled rack~65 kW$2.5–3.8M$79–117k
128-GPU B300 cluster16 × HGX B300, non-blocking leaf-spine fabric, ~1 PB parallel storage~255 kW$10–15M$77–117k
One 2 MW data hall~120 HGX B300 servers (~960 GPUs), full fabric, storage, management~2 MW$75–110M$78–115k
32-GPU H200 cluster4 × HGX H200, InfiniBand NDR (1 switch), ~200 TB all-flash, air-cooled~45 kW$1.4–2.1M$44–66k
H200 inference pod4 PCIe nodes, 8 × H200 NVL, 400G Ethernet, local NVMe~10 kWfrom ~$350k~$45k

Two patterns are worth noticing. A B300 GPU costs roughly 1.8 times an H200 GPU all-in, but carries twice the memory and, with FP4, delivers several times the inference throughput on current models, so cost per unit of work usually favors B300 for new capacity. And per-GPU cost stays roughly flat from 32 GPUs to a full hall: the fabric and storage grow with the cluster, while volume pricing offsets them. Fabric choices are covered in leaf-spine networking and NVLink vs InfiniBand; the full layer-by-layer bill of materials is on GPU cluster deployment.

For a real example of the smallest end of this table, see the H200 inference server that went from first call to first token in 9 days; for the larger end, the eight-node H100 cluster for a Gulf sovereign AI program.

Where the hardware lives: colocation and energy by region

After the hardware, space and power are the largest cost, and in 2026 they are the harder part to secure: high-density capacity in the major markets is scarce and new halls take a year or more to build. Haink places clusters in partner data centers in the United States, Latin America and Asia-Pacific. Capacity available for new projects as of September 2026:

LocationCapacity availableCoolingReadySpace and power costEnergy cost
United StatesSized per projectAir and liquidOn requestMidLower
El SalvadorSized per projectOn requestOn requestMidMid-high
Hong Kong2 MWLiquid-cooling readyNowHigherMid-high
Johor, Malaysia2 MW air + 2 MW liquidAir and liquidJanuary 2027LowerLower
Kuala Lumpur, Malaysia2 MW air + 2 MW liquidAir and liquidNowLower-midLower
Jakarta, Indonesia2 MWLiquid-cooling readyNowLower-midLower
Manila, Philippines2 MWConfirmed per projectNowMidHigher
Tokyo, Japan3 MWLiquid-cooling readyNowHigherHigher
Osaka, Japan2 MWConfirmed per projectNowMid-highHigher
Sydney, Australia3 MWConfirmed per projectNowHigherMid-high

As a planning figure, high-density colocation across these locations runs roughly $130–380 per kW per month for space, contracted power and cooling, with energy billed on top at roughly $0.08–0.25 per kWh depending on the country. The calculator below uses the ranges per location; the firm figure comes from the facility's quote for a specific project.

Three things decide the location before cost does. Cooling: B200 and B300 servers are planned for liquid cooling at scale, which today points to the United States, Hong Kong, Kuala Lumpur, Jakarta or Tokyo, or to Johor from January 2027 (why liquid cooling). Export and import rules: each destination has its own regime, and it decides whether a controlled GPU can be installed there at all and how long the paperwork takes. Data location: where the data must legally stay usually narrows the list further. For a US company, a US site keeps the data at home and a domestic installation needs no export licence, which removes the longest variable from the timeline; see cloud vs private AI.

No site with the power and cooling a B300 cluster needs? Hosting with our Tier 1 data-center partner is an option: GPU cluster hosting.

For scale: one 2 MW hall holds roughly 120 eight-GPU Blackwell servers, about 1,000 GPUs, once networking, storage and headroom are allowed for.

How much does it cost to run an AI server per year?

Per eight-GPU server, hosted in a partner facility, at 70% average utilization and a facility PUE of 1.3:

Annual costHGX B300 (14.5 kW)HGX H200 (10.2 kW)How it is calculated
Energy$9k–25k$6.5k–18kAverage draw × PUE × 8,760 h × $0.08–0.22 per kWh
Colocation (space, contracted power, cooling)$23k–66k$16k–46kContracted kW × $130–380 per kW per month
Hardware support after warranty$17k–45k$9k–24kAbout 3–6% of hardware cost a year, by OEM and service level
Remote hands, monitoring$3k–10k$3k–10kFacility or managed-service contract
Total per server~$50k–145k~$35k–100kStaff and software licences not included

Colocation is contracted on reserved power, not on what the servers actually draw, so an idle cluster still pays most of this line. That is the main reason utilization decides whether owning pays (see own vs rent).

What does it cost to upgrade your own data center for GPU clusters?

Most enterprise server rooms were built for 5–15 kW per rack. GPU clusters need far more:

What goes in the rackRack densityWhat the room needs
4 × HGX B30055–60 kWDirect liquid cooling with coolant distribution units and a facility water loop
2–3 × HGX H20020–35 kWUpgraded power distribution; rear-door heat exchangers or contained hot aisles
GB200 / GB300 NVL72 rack120–140 kWLiquid cooling designed in from the start; HPE advises provisioning up to 192 kW of busway per rack

An upgrade typically touches switchgear, UPS and busway; cooling (rear-door units at the lower end, coolant distribution units and a water loop at the higher end); floor loading, since a fully populated liquid-cooled rack can weigh well over a tonne; and the cabling plant for the fabric. Retrofit costs vary too much between buildings to give a useful single figure, and new AI-ready capacity is commonly estimated at over $10 million per MW of IT load. The larger cost is usually time: 6–18 months, most of it waiting for utility power and long-lead electrical equipment.

If the cluster is needed sooner, or the data does not have to sit in your own building, a partner hall that is already energized is almost always cheaper for the first three years. Our infrastructure audit checks what an existing room can take before any money is committed.

How long does an AI server or cluster project take?

As with cost, there is no single answer: the timeline belongs to the project. The stages, and what stretches each one:

StageTypical durationWhat makes it longer
Scoping and sizingDays to 2 weeksWorkload not yet measured; several sites under consideration
End-user documentation and supplier review1–2 weeksIncomplete end-user package; intended use that does not match the quantity; complex ownership
Export licence, where requiredWeeks to monthsThe destination and the end user; decided by the authority, not the supplier
Hardware supplyWeeks to several months once the order is approved; stock, when it exists, daysBlackwell generation, large quantities, specific OEM configurations
Freight, customs, delivery1–2 weeksImport permits, duty assessment, site access rules
Site ready and energizedNow (partner halls) to 18 months (new build)Utility power, cooling retrofit, liquid-cooling commissioning
Install, burn-in, acceptance1–3 weeks at 8–64 GPUs; months at 1,000+Fabric size, liquid-cooled rack-scale systems

Several stages run in parallel, so the total is shorter than the sum. The detailed breakdown by cluster size is on GPU cluster deployment timeline, and why a quoted delivery date needs reading carefully is on GPU lead times.

Is owning AI servers cheaper than renting cloud GPUs?

For a cluster that is used steadily, usually yes, and in 2026 the question on the rental side is as much whether capacity can be contracted at all as what it costs. The clearest comparison is the cost of one owned GPU-hour against what the market charges for one:

GPUOwned, 3 years, 70% usedOwned, 3 years, 90% usedOwned, 5 years, 70% usedMarket rental, September 2026
B300$5.69$4.43$4.59$6.80 blended index
H200$3.35$2.60$2.75$3.30 blended index · $4.62 settled trades
B200$4.66$3.63$3.83$3.68 (April composite) · $5.87 blended index

Owned cost per GPU-hour, all-in: the server at the middle of its range (B300 $650k, H200 $350k, B200 $490k), plus 15% for fabric and storage, plus running cost of $95k, $65k and $90k a year, less resale value, taken conservatively at 30% of the server price after three years and 15% after five. Market rates: the Silicon Data rental index of 24 September 2026, a blend of neocloud, hyperscaler and rental-platform prices; the Ornn Data settled-transaction benchmark for H200 of 27 September 2026, which ranged from $3.57 to $5.62 over the previous three months; and the SemiAnalysis B200 composite of April 2026.

What the table says:

The comparison is not only about cost. Owning requires a named end user, an installation address and a documented use; renting does not. Data that must stay in one jurisdiction or network, or capacity beyond what providers will rent on acceptable terms, can make owning the only option at any price. For a company already paying a large cloud bill, a cloud exit assessment runs this comparison on its own numbers.

Paying for it: purchase or financing

A cluster does not have to be paid for up front. Haink can arrange financing for projects that are clear: the end user is known, the hardware has a defined use, and it is going to a defined site. The structure is agreed per project and can cover the hardware and the colocation together, so the cost becomes a predictable monthly figure instead of a capital outlay.

What we look at when assessing a project for financing:

Financing does not shorten the compliance steps. A lender or lessor may hold title, but the operator is still the end user, and the same end-user review applies as for a purchase. It does mean that an approved project does not wait for a budget cycle. For startups presenting a GPU plan to investors, AI startup compute budget covers what they will ask.

Estimate your project

A first-order estimate using the ranges on this page. It is indicative, not a quote.

Get a firm cost for this project

Hardware: project-order ranges above (stock costs more and is rarely available in quantity), plus 12–25% for fabric, storage and racks on multi-server builds. Running: energy at the chosen utilization and PUE 1.3, colocation on contracted power plus 10% for network and storage, support at 4% a year and remote hands, less resale value of 30% of the servers after three years or 15% after five. Rental comparison: H200 from the Silicon Data blended index ($3.30, 24 September 2026) to settled trades (Ornn Data, $4.62, 27 September 2026); B200 from the SemiAnalysis April composite ($3.68) to the Silicon Data index ($5.87); H100 and B300 at the Silicon Data index. Duties, taxes, freight, staff and software are not included.

Frequently asked questions

How much does an AI server cost in 2026?

An eight-GPU NVIDIA HGX B300 server, the current platform for new clusters, costs roughly $550,000–$750,000. An HGX H200 server costs $300,000–$400,000, and a four-GPU PCIe inference node with H200 NVL about $155,000–$210,000. B200 servers run about $420,000–$560,000; older H100 servers are supplied on request. The final figure depends on quantity, destination country, end use, timing and the site.

How much does a GPU cluster cost?

A 32-GPU B300 cluster with fabric and storage costs about $2.5–3.8 million, a 128-GPU B300 cluster about $10–15 million, and a 2 MW hall of roughly 960 B300 GPUs about $75–110 million in hardware. On the previous generation, a 32-GPU H200 cluster costs about $1.4–2.1 million and an eight-GPU H200 NVL inference pod starts at about $350,000.

How much does it cost to run an AI server per year?

Hosted in a colocation facility at 70% utilization, an eight-GPU B300 server costs about $50,000–$145,000 a year to run and an H200 server about $35,000–$100,000, covering energy, colocation, support and remote hands. Staff and software are extra.

How much does a single AI rack cost?

A liquid-cooled rack of four HGX B300 servers is about $2.2–3.0 million in hardware at 55–60 kW. A rack of two to three HGX H200 servers is about $0.6–1.2 million at 20–35 kW. Hosting it costs roughly $130–380 per kW per month plus energy.

Is an NVIDIA DGX more expensive than an OEM HGX server?

Usually by about 10–15% for the same GPUs. A DGX is factory-integrated and supported by NVIDIA as a system; an OEM HGX server from Dell, HPE, Lenovo or Supermicro offers more configuration choice and the OEM's own support.

Why do quotes for the same AI server differ so much?

Because the price belongs to the order, not the product. Quantity, the country of installation, the end user and intended use, whether the hardware is in stock (roughly 1.4–1.7 times the price of a project order, and scarce) or ordered to allocation, the site's power and cooling, and the payment terms all move the figure. A quote far below the market range usually describes a different or unverified product.

Can Haink finance a GPU cluster?

Yes, for clear projects: a known end user, a defined use, and a defined site. The structure is agreed per project and can cover hardware and colocation together. The same end-user review applies as for a purchase.

Related

All hardware and running-cost figures on this page are indicative planning ranges as of 28 September 2026, not offers or quotes. Final pricing, availability and delivery dates remain subject to manufacturer, OEM and supplier approval, applicable export and import requirements, end-user screening and supply availability. Colocation capacity is shown as available at the date above and is not reserved until contracted. Financing is subject to assessment and approval for each project. Rental rates are third-party values (Silicon Data index, 24 September 2026; Ornn Data H200 benchmark, 27 September 2026; SemiAnalysis GPU index, April 2026) and public list prices, not Haink prices. System power figures from NVIDIA DGX user guides and product pages and the HPE GB300 NVL72 site-planning guidance, as listed on our power requirements page. This page is not financial or legal advice.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsDelaware (USA) · Hong Kong · Dubai · Singapore · Mainland China