AI Server Cost in 2026: What a GPU Server and a GPU Cluster Cost, and Why Quotes Differ
Hardware ranges reviewed 28 September 2026. GPU rental rates checked 24–27 September 2026. This page is reviewed quarterly, next review December 2026. All figures are indicative ranges, not quotes.
In September 2026 an eight-GPU NVIDIA HGX B300 server, the current platform for new clusters, costs roughly $550,000–$750,000; an HGX H200 server, the previous generation that is still a working choice for inference, costs $300,000–$400,000. A 32-GPU B300 cluster with fabric and storage runs about $2.5–3.8 million, a 32-GPU H200 cluster about $1.4–2.1 million, and keeping one eight-GPU server powered and hosted adds roughly $35,000–$145,000 a year. The ranges are wide on purpose. Two buyers asking for the same server get different numbers because the final figure is set by the project, not the part number: how many units, the country where the hardware is installed, who the end user is and what it will run, how soon it is needed, what the site can power and cool, and how it is paid for. The rest of this page shows how much each of those moves the cost and the date, so a budget can be built before anyone asks for a quote.
These are planning ranges drawn from our own quoting across the Hopper and Blackwell generations and from public market data. A firm figure needs the project details listed in what we ask before we open a supply request.
How much does an AI server cost in 2026?
The GPU generation sets the band. Within a band, the platform (OEM HGX server or NVIDIA DGX), memory, storage and network cards move the figure by another 10–20%.
| Server | GPUs | Power per server | Indicative cost, USD | What to know |
|---|---|---|---|---|
| HGX B300 (OEM) | 8 × B300 288 GB | ~14.5 kW | $550k–750k | The forward platform for new training and inference clusters: 288 GB per GPU. Supplied through allocation to documented projects; liquid cooling at cluster density. Widest spread between quotes of any platform |
| HGX H200 (OEM) | 8 × H200 SXM 141 GB | ~10.2 kW | $300k–400k | Previous generation, still a working choice for large-model inference at a lower entry cost; air cooling possible at 20–35 kW per rack |
| DGX H200 | 8 × H200 SXM | ~10.2 kW | about 10–15% above OEM HGX | Factory-integrated, NVIDIA support; pays off when the software stack and support matter more than cost |
| PCIe inference node | 4 × H200 NVL 141 GB | ~4–5 kW | $155k–210k | H200 NVL cards from about $31k each. Fits standard air-cooled racks; the quickest path for inference. PCIe vs SXM |
| Also supplied | ||||
| HGX B200 (OEM) | 8 × B200 192 GB | ~14.3 kW | $420k–560k | Supplied on allocation; most new projects now go straight to B300 |
| HGX H100 (OEM) | 8 × H100 SXM 80 GB | ~10.2 kW | $260k–360k | Older generation, supplied on request, mainly for expanding existing H100 clusters |
| Workstation-class inference server | 4–8 × RTX PRO 6000 or L40S | 2–5 kW | $60k–250k | Lower cost per GPU; suited to smaller models, multi-model serving and fine-tuning |
These ranges are for hardware ordered for a project, not bought off a shelf. In September 2026 the two are different markets. An H200 card ordered for a documented project starts at about $31k; the same card from stock, where any can be found, costs roughly 1.4–1.7 times as much, and stock in cluster quantities is close to impossible to source. The dependable route to a cluster is a project order supplied through allocation, which is why the end-user and site information is worth preparing early: it is what the order waits for, not the factory. How the two routes differ is set out in stock vs built-to-order and allocation vs stock vs executable supply.
Power figures are NVIDIA's system maximums for the equivalent DGX, as collected on our GPU cluster power requirements page. For how the generations differ in memory and throughput, see H100 vs H200 vs B200 vs B300.
What makes up the cost of one AI server?
On an eight-GPU server the accelerators are most of the bill, which is why the GPU generation matters more than the brand of the chassis. Typical shares for an HGX B300 configuration:
Two consequences for a budget. First, trimming memory or storage saves little; choosing H200, or PCIe H200 NVL, instead of B300 for an inference workload that fits in their memory saves a lot. Second, a quote that is far below the band is almost always a quote for something else: a different GPU variant, used or refurbished parts, a server without GPUs, or a price that has not yet been checked against a supplier. How to verify a GPU supplier covers the checks.
Why do two quotes for the same AI server differ?
Because the price belongs to an order, not to a product. These are the six things that move it, in the order they usually matter:
| Factor | Effect on cost | Effect on the date | What settles it |
|---|---|---|---|
| 1. GPU generation and platform | Sets the band: a B300 server costs roughly twice an H200 server | B300 follows allocation to documented projects; H200 and PCIe orders usually move faster | The workload: model size, context length, training or inference |
| 2. Quantity | A single server carries the full cost of sourcing, freight and integration; from roughly 16 servers upward, per-unit cost falls and allocation is requested per project | Larger orders are reviewed more closely and may ship in phases | A deployment plan: how many GPUs now, how many later |
| 3. Country of installation | Freight, insurance, import duty and VAT or GST, local installation. Together often 3–12% on top of hardware | Where a US export licence is required, weeks to months; Hong Kong, for example, is treated like mainland China for advanced GPUs | The installation address, fixed before a supplier commits |
| 4. End user and intended use | Indirect: a well-documented project gets supplier pricing, an unclear one gets a risk premium or no quote | End-user documentation and supplier review typically take 1–2 weeks | The end-user statement and a plain description of the workload |
| 5. Timing | The largest single swing: stock costs roughly 1.4–1.7 times a project order for the same GPU, and in cluster quantities is rarely available at all | Stock, when it exists: days to two weeks. Built to order: depends on the channel's commitment | Whether the site is ready; hardware that arrives before power is idle capital |
| 6. Site and cooling | Liquid cooling adds coolant distribution units and manifolds; a site without high-density power adds colocation or a retrofit | The site, not the hardware, sets the go-live date in most cluster projects | Rack density in kW, cooling type and the date the hall is energized |
Payment terms are a seventh factor that sits on top: an order paid on delivery, one with a deposit, and one that is financed are priced differently because they carry different costs of capital.
How much does a GPU cluster cost?
A cluster adds three things to the servers: the fabric that connects them, shared storage for datasets and checkpoints, and the racks, power distribution and cabling. Together they usually add 12–25% to the server cost; the share falls as clusters grow. Five reference builds, current platform first:
| Build | What is in it | IT power | Indicative hardware cost | Per GPU, all-in |
|---|---|---|---|---|
| 32-GPU B300 cluster | 4 × HGX B300, 800G InfiniBand or Spectrum-X, ~200 TB all-flash, liquid-cooled rack | ~65 kW | $2.5–3.8M | $79–117k |
| 128-GPU B300 cluster | 16 × HGX B300, non-blocking leaf-spine fabric, ~1 PB parallel storage | ~255 kW | $10–15M | $77–117k |
| One 2 MW data hall | ~120 HGX B300 servers (~960 GPUs), full fabric, storage, management | ~2 MW | $75–110M | $78–115k |
| 32-GPU H200 cluster | 4 × HGX H200, InfiniBand NDR (1 switch), ~200 TB all-flash, air-cooled | ~45 kW | $1.4–2.1M | $44–66k |
| H200 inference pod | 4 PCIe nodes, 8 × H200 NVL, 400G Ethernet, local NVMe | ~10 kW | from ~$350k | ~$45k |
Two patterns are worth noticing. A B300 GPU costs roughly 1.8 times an H200 GPU all-in, but carries twice the memory and, with FP4, delivers several times the inference throughput on current models, so cost per unit of work usually favors B300 for new capacity. And per-GPU cost stays roughly flat from 32 GPUs to a full hall: the fabric and storage grow with the cluster, while volume pricing offsets them. Fabric choices are covered in leaf-spine networking and NVLink vs InfiniBand; the full layer-by-layer bill of materials is on GPU cluster deployment.
For a real example of the smallest end of this table, see the H200 inference server that went from first call to first token in 9 days; for the larger end, the eight-node H100 cluster for a Gulf sovereign AI program.
Where the hardware lives: colocation and energy by region
After the hardware, space and power are the largest cost, and in 2026 they are the harder part to secure: high-density capacity in the major markets is scarce and new halls take a year or more to build. Haink places clusters in partner data centers in the United States, Latin America and Asia-Pacific. Capacity available for new projects as of September 2026:
| Location | Capacity available | Cooling | Ready | Space and power cost | Energy cost |
|---|---|---|---|---|---|
| United States | Sized per project | Air and liquid | On request | Mid | Lower |
| El Salvador | Sized per project | On request | On request | Mid | Mid-high |
| Hong Kong | 2 MW | Liquid-cooling ready | Now | Higher | Mid-high |
| Johor, Malaysia | 2 MW air + 2 MW liquid | Air and liquid | January 2027 | Lower | Lower |
| Kuala Lumpur, Malaysia | 2 MW air + 2 MW liquid | Air and liquid | Now | Lower-mid | Lower |
| Jakarta, Indonesia | 2 MW | Liquid-cooling ready | Now | Lower-mid | Lower |
| Manila, Philippines | 2 MW | Confirmed per project | Now | Mid | Higher |
| Tokyo, Japan | 3 MW | Liquid-cooling ready | Now | Higher | Higher |
| Osaka, Japan | 2 MW | Confirmed per project | Now | Mid-high | Higher |
| Sydney, Australia | 3 MW | Confirmed per project | Now | Higher | Mid-high |
As a planning figure, high-density colocation across these locations runs roughly $130–380 per kW per month for space, contracted power and cooling, with energy billed on top at roughly $0.08–0.25 per kWh depending on the country. The calculator below uses the ranges per location; the firm figure comes from the facility's quote for a specific project.
Three things decide the location before cost does. Cooling: B200 and B300 servers are planned for liquid cooling at scale, which today points to the United States, Hong Kong, Kuala Lumpur, Jakarta or Tokyo, or to Johor from January 2027 (why liquid cooling). Export and import rules: each destination has its own regime, and it decides whether a controlled GPU can be installed there at all and how long the paperwork takes. Data location: where the data must legally stay usually narrows the list further. For a US company, a US site keeps the data at home and a domestic installation needs no export licence, which removes the longest variable from the timeline; see cloud vs private AI.
No site with the power and cooling a B300 cluster needs? Hosting with our Tier 1 data-center partner is an option: GPU cluster hosting.
For scale: one 2 MW hall holds roughly 120 eight-GPU Blackwell servers, about 1,000 GPUs, once networking, storage and headroom are allowed for.
How much does it cost to run an AI server per year?
Per eight-GPU server, hosted in a partner facility, at 70% average utilization and a facility PUE of 1.3:
| Annual cost | HGX B300 (14.5 kW) | HGX H200 (10.2 kW) | How it is calculated |
|---|---|---|---|
| Energy | $9k–25k | $6.5k–18k | Average draw × PUE × 8,760 h × $0.08–0.22 per kWh |
| Colocation (space, contracted power, cooling) | $23k–66k | $16k–46k | Contracted kW × $130–380 per kW per month |
| Hardware support after warranty | $17k–45k | $9k–24k | About 3–6% of hardware cost a year, by OEM and service level |
| Remote hands, monitoring | $3k–10k | $3k–10k | Facility or managed-service contract |
| Total per server | ~$50k–145k | ~$35k–100k | Staff and software licences not included |
Colocation is contracted on reserved power, not on what the servers actually draw, so an idle cluster still pays most of this line. That is the main reason utilization decides whether owning pays (see own vs rent).
What does it cost to upgrade your own data center for GPU clusters?
Most enterprise server rooms were built for 5–15 kW per rack. GPU clusters need far more:
| What goes in the rack | Rack density | What the room needs |
|---|---|---|
| 4 × HGX B300 | 55–60 kW | Direct liquid cooling with coolant distribution units and a facility water loop |
| 2–3 × HGX H200 | 20–35 kW | Upgraded power distribution; rear-door heat exchangers or contained hot aisles |
| GB200 / GB300 NVL72 rack | 120–140 kW | Liquid cooling designed in from the start; HPE advises provisioning up to 192 kW of busway per rack |
An upgrade typically touches switchgear, UPS and busway; cooling (rear-door units at the lower end, coolant distribution units and a water loop at the higher end); floor loading, since a fully populated liquid-cooled rack can weigh well over a tonne; and the cabling plant for the fabric. Retrofit costs vary too much between buildings to give a useful single figure, and new AI-ready capacity is commonly estimated at over $10 million per MW of IT load. The larger cost is usually time: 6–18 months, most of it waiting for utility power and long-lead electrical equipment.
If the cluster is needed sooner, or the data does not have to sit in your own building, a partner hall that is already energized is almost always cheaper for the first three years. Our infrastructure audit checks what an existing room can take before any money is committed.
How long does an AI server or cluster project take?
As with cost, there is no single answer: the timeline belongs to the project. The stages, and what stretches each one:
| Stage | Typical duration | What makes it longer |
|---|---|---|
| Scoping and sizing | Days to 2 weeks | Workload not yet measured; several sites under consideration |
| End-user documentation and supplier review | 1–2 weeks | Incomplete end-user package; intended use that does not match the quantity; complex ownership |
| Export licence, where required | Weeks to months | The destination and the end user; decided by the authority, not the supplier |
| Hardware supply | Weeks to several months once the order is approved; stock, when it exists, days | Blackwell generation, large quantities, specific OEM configurations |
| Freight, customs, delivery | 1–2 weeks | Import permits, duty assessment, site access rules |
| Site ready and energized | Now (partner halls) to 18 months (new build) | Utility power, cooling retrofit, liquid-cooling commissioning |
| Install, burn-in, acceptance | 1–3 weeks at 8–64 GPUs; months at 1,000+ | Fabric size, liquid-cooled rack-scale systems |
Several stages run in parallel, so the total is shorter than the sum. The detailed breakdown by cluster size is on GPU cluster deployment timeline, and why a quoted delivery date needs reading carefully is on GPU lead times.
Is owning AI servers cheaper than renting cloud GPUs?
For a cluster that is used steadily, usually yes, and in 2026 the question on the rental side is as much whether capacity can be contracted at all as what it costs. The clearest comparison is the cost of one owned GPU-hour against what the market charges for one:
| GPU | Owned, 3 years, 70% used | Owned, 3 years, 90% used | Owned, 5 years, 70% used | Market rental, September 2026 |
|---|---|---|---|---|
| B300 | $5.69 | $4.43 | $4.59 | $6.80 blended index |
| H200 | $3.35 | $2.60 | $2.75 | $3.30 blended index · $4.62 settled trades |
| B200 | $4.66 | $3.63 | $3.83 | $3.68 (April composite) · $5.87 blended index |
Owned cost per GPU-hour, all-in: the server at the middle of its range (B300 $650k, H200 $350k, B200 $490k), plus 15% for fabric and storage, plus running cost of $95k, $65k and $90k a year, less resale value, taken conservatively at 30% of the server price after three years and 15% after five. Market rates: the Silicon Data rental index of 24 September 2026, a blend of neocloud, hyperscaler and rental-platform prices; the Ornn Data settled-transaction benchmark for H200 of 27 September 2026, which ranged from $3.57 to $5.62 over the previous three months; and the SemiAnalysis B200 composite of April 2026.
What the table says:
- B300 is cheaper to own at every utilization shown, by roughly 15–35% over three years and more over five. For new capacity that will be used steadily, owning B300 is the clearest case on this page.
- H200 still pays back. At 70% utilization over three years an owned H200 costs about the same as the lowest blended rental rate and roughly 25–30% less than the rate H200 capacity has actually been trading at. Over five years, or at 90% utilization, owning is cheaper than any rental rate in the table.
- B200 depends on the rate that can be contracted: against the higher current index owning wins clearly; against the lower April composite it only wins at sustained utilization near 90%.
- Capacity is the other half of the comparison, on both sides. GPU stock for purchase is as scarce as rental capacity, so the realistic comparison is a project order against a rental contract, not stock against on-demand. SemiAnalysis reported in April 2026 that on-demand GPU capacity was sold out across GPU types, that one-year H100 contract prices had risen almost 40% since October 2025, and that capacity coming online through September 2026 was already booked. A rental rate that cannot be contracted for the size and term needed is not an alternative.
- Resale value is real. Silicon Data has reported H100 and A100 trading well above what three- or five-year straight-line depreciation implies, and B200 above its launch price in September 2026. The table assumes far less; higher resale makes owning cheaper still.
- Rent when utilization is below about 50%, the project has an end date, or the model and GPU generation are still changing. Owned hardware that sits idle still pays colocation and depreciation. The stage-by-stage view is in buy vs rent GPUs.
The comparison is not only about cost. Owning requires a named end user, an installation address and a documented use; renting does not. Data that must stay in one jurisdiction or network, or capacity beyond what providers will rent on acceptable terms, can make owning the only option at any price. For a company already paying a large cloud bill, a cloud exit assessment runs this comparison on its own numbers.
Paying for it: purchase or financing
A cluster does not have to be paid for up front. Haink can arrange financing for projects that are clear: the end user is known, the hardware has a defined use, and it is going to a defined site. The structure is agreed per project and can cover the hardware and the colocation together, so the cost becomes a predictable monthly figure instead of a capital outlay.
What we look at when assessing a project for financing:
- Who the end user is: the operating company, its ownership and its track record.
- What the hardware will do and who pays for its use: your own product with revenue, internal workloads with a budget, or capacity contracted to a customer.
- Where it will run: a partner hall or your own site, with power and cooling confirmed.
Financing does not shorten the compliance steps. A lender or lessor may hold title, but the operator is still the end user, and the same end-user review applies as for a purchase. It does mean that an approved project does not wait for a budget cycle. For startups presenting a GPU plan to investors, AI startup compute budget covers what they will ask.
Estimate your project
A first-order estimate using the ranges on this page. It is indicative, not a quote.
Get a firm cost for this project
Hardware: project-order ranges above (stock costs more and is rarely available in quantity), plus 12–25% for fabric, storage and racks on multi-server builds. Running: energy at the chosen utilization and PUE 1.3, colocation on contracted power plus 10% for network and storage, support at 4% a year and remote hands, less resale value of 30% of the servers after three years or 15% after five. Rental comparison: H200 from the Silicon Data blended index ($3.30, 24 September 2026) to settled trades (Ornn Data, $4.62, 27 September 2026); B200 from the SemiAnalysis April composite ($3.68) to the Silicon Data index ($5.87); H100 and B300 at the Silicon Data index. Duties, taxes, freight, staff and software are not included.
Frequently asked questions
How much does an AI server cost in 2026?
An eight-GPU NVIDIA HGX B300 server, the current platform for new clusters, costs roughly $550,000–$750,000. An HGX H200 server costs $300,000–$400,000, and a four-GPU PCIe inference node with H200 NVL about $155,000–$210,000. B200 servers run about $420,000–$560,000; older H100 servers are supplied on request. The final figure depends on quantity, destination country, end use, timing and the site.
How much does a GPU cluster cost?
A 32-GPU B300 cluster with fabric and storage costs about $2.5–3.8 million, a 128-GPU B300 cluster about $10–15 million, and a 2 MW hall of roughly 960 B300 GPUs about $75–110 million in hardware. On the previous generation, a 32-GPU H200 cluster costs about $1.4–2.1 million and an eight-GPU H200 NVL inference pod starts at about $350,000.
How much does it cost to run an AI server per year?
Hosted in a colocation facility at 70% utilization, an eight-GPU B300 server costs about $50,000–$145,000 a year to run and an H200 server about $35,000–$100,000, covering energy, colocation, support and remote hands. Staff and software are extra.
How much does a single AI rack cost?
A liquid-cooled rack of four HGX B300 servers is about $2.2–3.0 million in hardware at 55–60 kW. A rack of two to three HGX H200 servers is about $0.6–1.2 million at 20–35 kW. Hosting it costs roughly $130–380 per kW per month plus energy.
Is an NVIDIA DGX more expensive than an OEM HGX server?
Usually by about 10–15% for the same GPUs. A DGX is factory-integrated and supported by NVIDIA as a system; an OEM HGX server from Dell, HPE, Lenovo or Supermicro offers more configuration choice and the OEM's own support.
Why do quotes for the same AI server differ so much?
Because the price belongs to the order, not the product. Quantity, the country of installation, the end user and intended use, whether the hardware is in stock (roughly 1.4–1.7 times the price of a project order, and scarce) or ordered to allocation, the site's power and cooling, and the payment terms all move the figure. A quote far below the market range usually describes a different or unverified product.
Can Haink finance a GPU cluster?
Yes, for clear projects: a known end user, a defined use, and a defined site. The structure is agreed per project and can cover hardware and colocation together. The same end-user review applies as for a purchase.
Related
- Buy vs rent GPUs — cost comparison by stage
- H100 vs H200 vs B200 vs B300
- GPU server buying guide
- AI cluster architecture
- How NVIDIA GPU allocation works
- Private AI infrastructure
- Reference architectures
- GPU procurement process
All hardware and running-cost figures on this page are indicative planning ranges as of 28 September 2026, not offers or quotes. Final pricing, availability and delivery dates remain subject to manufacturer, OEM and supplier approval, applicable export and import requirements, end-user screening and supply availability. Colocation capacity is shown as available at the date above and is not reserved until contracted. Financing is subject to assessment and approval for each project. Rental rates are third-party values (Silicon Data index, 24 September 2026; Ornn Data H200 benchmark, 27 September 2026; SemiAnalysis GPU index, April 2026) and public list prices, not Haink prices. System power figures from NVIDIA DGX user guides and product pages and the HPE GB300 NVL72 site-planning guidance, as listed on our power requirements page. This page is not financial or legal advice.
