Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Insights / AI startup compute budget

AI Startup Compute Budget — What Investors Ask About Your GPU Plan

Written and maintained by Haink's procurement and allocation advisory team · Updated September 2026

Hardware figures verified as of 19 September 2026 against nvidia.com, hpe.com and the sources named in each section. NVIDIA Inception terms verified as of 18 September 2026 against nvidia.com/en-us/startups/. Reviewed with each new GPU generation.

An AI startup's compute budget is the number of GPU-hours its plan needs over the next 12–24 months, tied to milestones, multiplied by a rate that a named provider has actually confirmed. Investors test it with a short list of questions: what each block of compute buys, how the GPU-hours were derived, whether the company will rent, reserve or own hardware and why, what document stands behind any claim of "secured" capacity, where owned hardware will be installed and powered, and what happens if utilisation or model size turns out differently. A plan survives diligence when every GPU count traces back to arithmetic and every date traces back to a named party. Most early-stage plans should rent: owning puts end-user screening, a site and a power contract on the critical path of the round.

What do investors ask about a startup's GPU plan?

The questions are consistent across stages because they test the same thing: whether the compute line is a derived number or a negotiated one. Each has an answer that holds up and an answer that stalls the conversation.

QuestionAn answer that holds upAn answer that weakens the plan
What does this compute buy?Each block of GPU-hours is tied to a milestone: a trained model, an evaluation result, a product launch, a customer volumeOne GPU total for the whole period, with no link to what it produces
How was the number derived?Training FLOPs from model size and token count, a stated utilisation assumption with its source, and an explicit allowance for experiments and failures"We need 512 H100s", with the count as the starting point rather than the result
Rent, reserve or own — and why now?A utilisation forecast and the point at which ownership would pay for itself, with the conditions that trigger the switchOwnership justified by "GPUs are scarce" or by a hardware discount
What stands behind "secured"?A named supplier or provider, a written confirmation, and the type of document it isA reseller's email, a price sheet or a verbal assurance
Where will it run?For owned hardware: a named site with contracted, energised power at specific rack positions and a date for itA hardware delivery date presented as the date the cluster goes into production
Can the company complete the purchase?Legal entity, end user, installation address and use case documented and consistentInstallation site "to be decided after the round"
What if the assumptions are wrong?A range for utilisation, model size and inference demand, and which milestones move if the low case happensA single-point estimate with no sensitivity

How do you calculate an AI startup's compute budget for 12–24 months?

A defensible budget has four lines for rented capacity and a fifth for owned hardware. They are estimated separately because they grow for different reasons, and an investor will ask about each one.

Budget lineWhat drives itHow to estimate it
Final training runsParameter count and training tokensTraining FLOPs ≈ 6 × parameters × tokens, converted to GPU-hours at a stated utilisation
ExperimentationAblations, hyperparameter searches, restarts, runs that are abandonedThe company's own ratio of total to final-run compute from past work, stated as an assumption
InferenceRequests, tokens per request, latency targetPeak tokens per second required ÷ tokens per second one GPU sustains at that latency, measured on the company's own model and serving stack
Reliability and idle timeHardware faults, checkpoint restarts, reserved hours not usedA stated percentage on top of productive GPU-hours
Owned hardware onlyPower, hosting, network, storage, spares, operations staffPriced separately from the GPUs; see the section on owning below

The training estimate uses a relationship that has been standard since Kaplan et al., Scaling Laws for Neural Language Models (OpenAI, January 2020): a dense transformer spends about 6 floating-point operations per parameter per training token. Converted to time:

GPU-hours = (6 × parameters × tokens) ÷ (peak FLOPS per GPU × utilisation) ÷ 3,600

Utilisation here is model FLOPs utilisation (MFU): the share of the GPU's peak arithmetic that the training job actually uses. It is not the same as the share of time the GPU is busy, and confusing the two is one of the most common errors in compute plans. For scale, Meta reported 38–43% BF16 MFU when training Llama 3 405B on 8,192 to 16,384 H100 GPUs (The Llama 3 Herd of Models, July 2024, Table 4). A small team on a less tuned software stack has no reason to assume it will do better.

Worked example: the final run for a 7B-parameter model

Take a 7-billion-parameter model trained on 140 billion tokens, the 20-tokens-per-parameter ratio of the Chinchilla paper (Hoffmann et al., DeepMind, March 2022, where Chinchilla itself was 70B parameters on 1.4 trillion tokens). Training compute is 6 × 7×10⁹ × 1.4×10¹¹ = 5.88×10²¹ FLOPs. NVIDIA lists the H100 SXM at 1,979 TFLOPS BF16 "with sparsity"; the dense figure used for training estimates is half, about 989.5 TFLOPS (nvidia.com, checked 19 September 2026).

Assumed MFUGPU-hours, final runWall-clock on 64 GPUs
30%≈ 5,500≈ 86 hours (3.6 days)
40%≈ 4,130≈ 65 hours (2.7 days)
50%≈ 3,300≈ 52 hours (2.1 days)

Calculated from the formula above. The MFU values are a sensitivity range chosen for illustration, not measured results. Excludes experimentation, evaluation, restarts and inference.

Two things follow. First, the final run of a model this size is days on a few dozen GPUs, not months on hundreds. A plan that asks for hundreds of GPUs for eighteen months has to be explained by something else — larger models, many runs, or inference — and the investor will ask which. Second, the final run is usually the smallest line. Experimentation and inference decide most budgets, which is why the ratio of total compute to final-run compute is the assumption most worth stating and defending.

Cost is the GPU-hours multiplied by a rate. The rate belongs in the plan only as a figure quoted by a named provider on a stated date, because GPU rates move and a number copied from a public list page is not an offer.

Why the reliability line is not optional

Large jobs fail and restart. Over a 54-day snapshot of Llama 3 405B pre-training, Meta recorded 466 job interruptions, 419 of them unexpected, while keeping effective training time above 90% (The Llama 3 Herd of Models, July 2024). That was a frontier team with dedicated reliability engineering. A startup budget that assumes every reserved GPU-hour turns into useful training is assuming a better outcome than the best-resourced operators report.

Should an AI startup rent, reserve or own GPUs?

The choice follows demand certainty, not ambition. The table sets out the usual pattern by stage.

StageWhat is known about demandUsual structure
Research, pre-productModel direction still changing; usage unpredictableOn-demand or short commitments with a cloud or GPU provider
Early productA measurable baseline load, with peaksReserved capacity for the baseline, on-demand for peaks
Sustained scaleHigh, steady utilisation of a known workload; a site and an operations team exist or are fundedOwnership becomes a question worth modelling against long-term reservation

For most seed and Series A companies, renting is the right answer, and a plan that says so is stronger than one that buys hardware to look serious. Owned GPUs lose value whether or not they are used, and a model direction that changes in month six can leave the company with capacity sized for the wrong workload.

The difference is not only financial. Buying makes the startup the named end user in the purchase, with its legal entity, installation address, data center and intended use disclosed and screened; renting leaves the hardware purchase, the site and supplier screening with the provider. For a company that has not yet chosen where it will operate, that alone can make an ownership plan unexecutable within the round's timeline. The stage-by-stage comparison and the break-even arithmetic are in buy vs rent GPUs.

What does "secured compute" mean to an investor?

"Secured" is the word investors test hardest, because the GPU market uses it for things that are not equivalent. Five different states are often described with the same word:

A compute slide should say which of these it is describing and name the document. "Capacity secured" backed by a supplier offer is a quotation, and a diligent investor will read it as one. For rented capacity the equivalent is a signed reservation with a provider, stating the GPU type, quantity, term and start date. The distinctions are set out in full in allocation vs stock vs executable supply.

Delivery dates follow the same rule. A date belongs in the plan only as the supplier's written date, labelled as such. Market lead times shift with configuration, volume, destination and end use, which is explained in what determines GPU lead times.

What does owning GPUs require beyond the purchase?

Delivery is not deployment. An owned cluster produces nothing until it is installed at a site with enough contracted, energised power and cooling, and those are usually the longer part of the schedule.

The power numbers are larger than most first plans assume. NVIDIA rates the eight-GPU DGX B200 at about 14.3 kW maximum, and HPE rates a 72-GPU GB300 NVL72 rack at 132 kW nominal, advising data centers to provision 192 kW of busway per rack (both checked 19 September 2026). By contrast, 82% of operators in the Uptime Institute Global Data Center Survey 2025 (July 2025) reported no racks above 30 kW. Two Blackwell eight-GPU servers already reach that line. The figures and their scaling to full clusters are in GPU cluster power requirements.

An investor looking at an ownership plan will therefore ask for the same things a supplier asks for before it can process the order: the buyer's legal entity, the end user, the country and address of installation, the data center or hosting provider, the configuration and quantity, the use case and the target date. If those fields cannot be filled in consistently, the hardware cannot be ordered on the timeline the plan shows, whatever the budget. The full list, and why each item is asked for, is in what is asked before a supply request is opened. The end-user side is covered in NVIDIA end-user requirements.

How do investors treat cloud credits and NVIDIA Inception?

Credits reduce cost for a period. They are not a supply plan, and a budget that counts them as runway without saying when they run out and what replaces them will be discounted.

NVIDIA Inception is a common part of startup compute plans. As of 18 September 2026, NVIDIA states that the program is free — "no application fees, membership fees, or equity requirements" — and open to companies with at least one developer, a working website, official registration and less than ten years of operation; revenue is not required. Consulting and outsourced development firms, crypto companies, cloud service providers, resellers and distributors, and public companies are excluded. Benefits listed include self-paced training, SDK and model access, preferred pricing on selected products, cloud credits from NVIDIA and partners, go-to-market support, and Inception Capital Connect for introductions to investors (nvidia.com/en-us/startups).

For an investor, the practical points are narrow. Membership is decided by NVIDIA and is not guaranteed. It does not assign GPU allocation or create any supply commitment. Capital Connect is an introduction mechanism; any funding decision stays with the investor. A plan that presents Inception as access to hardware or to capital is overstating what the program offers.

Which mistakes stall a round on the compute slide?

  1. A GPU count with no derivation. The number of GPUs should be the output of the arithmetic, not the input.
  2. Time-busy presented as MFU. A cluster that is 95% occupied can still run at 35% model FLOPs utilisation. Budgets built on the first number are short by the difference.
  3. No allowance for experiments and failures. The final run alone is rarely the budget.
  4. "Secured" without a document. An offer described as an allocation, or an allocation with no named party that assigned it.
  5. Delivery date used as go-live date. For owned hardware, energised power at the site sets the start date.
  6. Ownership before the demand is known. Buying to signal seriousness, then carrying depreciating capacity sized for a workload that has changed.
  7. Credits counted as runway. Without an end date and a replacement plan.

Frequently asked questions

How much should an AI startup budget for compute?

There is no universal figure. A defensible budget is derived: training FLOPs from model size and tokens, converted to GPU-hours at a stated utilisation, plus experimentation, inference and reliability allowances, multiplied by a rate quoted by a named provider on a stated date. Investors test the derivation more than the total.

What utilisation should a compute plan assume?

State the assumption and its source. For reference, Meta reported 38–43% BF16 model FLOPs utilisation when training Llama 3 405B on 8,192 to 16,384 H100 GPUs (July 2024). Smaller teams should not assume higher. Model FLOPs utilisation is different from the share of time GPUs are busy.

Should a seed-stage AI startup buy its own GPUs?

Usually not. Ownership pays when utilisation is high and steady, the workload is known, and a site and operations team exist. Buying also makes the startup the named end user in the purchase, with end-user screening, an installation address and a powered site on the critical path. Renting leaves those with the provider.

What does an investor mean by "secured" GPU capacity?

A named supplier or provider has confirmed the capacity in writing, on stated terms, and the document says what it is. For rented capacity that is a signed reservation. For purchased hardware it is executable supply or a purchase order, not a supplier offer or a price sheet.

Does NVIDIA Inception give a startup access to GPUs or investors?

No allocation comes with it. As of 18 September 2026 the program is free, membership is decided by NVIDIA, and benefits include cloud credits, preferred pricing on selected products and Inception Capital Connect, which provides introductions to investors. Any funding decision remains with the investor.

If the plan includes buying hardware, the next question is whether supply can be obtained for the project at all: How NVIDIA GPU allocation works →

Final allocation and hardware availability remain subject to manufacturer/OEM/supplier approval, applicable compliance requirements and supply availability.

This page describes general market practice and publicly documented figures. It does not state availability, lead times or prices for any product or order, and it is not investment advice.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsDelaware (USA) · Hong Kong · Dubai · Singapore · Mainland China