Other

What's the hourly and monthly price to rent a GB200 NVL72 rack in the cloud, and is it cheaper to reserve or use on-demand?

July 24, 2026

Teams planning frontier-scale training or very large model serving eventually price out a GB200 NVL72 rack and hit two surprises: the hourly number is quoted per GPU rather than per rack, and the "reserve vs on-demand" answer flips depending on how busy the hardware actually stays. A GB200 NVL72 is billed as a per-GPU-hour rate that you multiply across the rack, so the real decision is monthly burn at your utilization, not a single sticker price. This guide shows how to convert an hourly quote into a monthly number, and when reserving beats on-demand.

GB200 NVL72 pricing is quoted per GPU-hour, then rolled up to the rack

The GB200 NVL72 is a liquid-cooled, rack-scale system that links 72 Blackwell GPUs over NVLink into one tightly coupled domain for large-scale training and high-density inference. Cloud providers rarely publish one price per rack. Instead they publish a per-GPU-hour rate and let you compose the quantity you need, from a few superchips up to a full 72-GPU rack.

We currently list NVIDIA GB200 NVL72 capacity at from $8.00 per GPU-hour in our current published listings, alongside H100 at from $2.00, H200 at from $2.60, and B200 at from $4.00 per GPU-hour. That starting-list rate is the anchor to reason from, because it is verifiable on our pricing page rather than hidden behind a sales-only quote.

Getting from an hourly rate to a rack number is simple arithmetic worth doing explicitly: rack hourly is the per-GPU-hour rate times 72 GPUs, and always-on monthly is that rack hourly times roughly 730 hours. One boundary before you multiply: a per-GPU-hour rate is not the same commercial object as a full-rack monthly commitment. The per-GPU number is the unit input; the monthly figure depends on how many GPUs you actually hold, for how many hours, and under which billing model. Two teams quoting "the same GB200 rate" can end up with very different invoices because one runs a partial allocation at spiky utilization and the other holds a full reserved rack.

What an hourly rate turns into per month

Use the table as a planning frame. Our row uses the public starting list rate; market rows reflect cross-provider index behavior that drifts with stock and should be re-sampled before purchase.

ScenarioRate basis ($/GPU-hr)Read this as
GMI GB200 NVL72, single GPU, always-onfrom $8.00 (~$5,840/GPU/mo)Unit building block for scaling up
GMI GB200 NVL72, full 72-GPU rack, always-onfrom $8.00 (~$420k/mo, indicative)Full-rack ceiling before commit discounts
Market GB200 on-demand (aggregated index)often ~$10-$21+ rangeScarcity and premium bundles inflate on-demand
Market GB200 reserved / custom (aggregated)usually lower, frequently "custom"Where large racks actually get priced

Three honest caveats on those numbers. Full-rack list math is indicative, not a checkout price: rack-scale deployments almost always move to committed contracts with volume pricing, so the "72 times list" figure is an upper reference, not what a serious buyer pays. Public GB200 indexes show a wide, rising band, with on-demand medians well above single-B200 levels and many rows marked "custom," so treat aggregated ranges as an envelope. And utilization is the real cost driver: a rack at 25 percent utilization on-demand can cost more per useful hour than a reserved rack at 80 percent, even when the reserved unit rate is lower.

Reserve vs on-demand: which is cheaper and when

This is the part the question turns on. The short answer: on-demand is cheaper only when utilization is low and bursty, while reserved capacity is cheaper once the rack runs enough hours to amortize the commitment.

On-demand fits when you are still validating whether an NVL72-class rack is the right target, when usage is intermittent (short training runs, evaluation windows, proof-of-concept), when you cannot yet forecast monthly hours with confidence, or when you value walking away with no commitment more than the lowest unit rate.

Reserved or committed capacity fits when the rack will run sustained training or steady production inference, when you can forecast utilization above the break-even point where the discount more than offsets idle hours, when latency and availability matter enough that you want warm, dedicated capacity instead of fighting for on-demand stock, or when you need predictable monthly spend for budgeting.

We frame this the same way in its commercial model: commitment-based savings reduce unit GPU cost through reserved capacity and sustained deployments, while usage-adaptive pricing lets you start on-demand and transition to dedicated or committed capacity as workloads stabilize. In practice, most production teams land on a reserved baseline plus burst rather than pure on-demand, because a rack that must serve real traffic cannot tolerate a stock-out at peak. A simple break-even instinct: estimate your expected monthly hours, price the same quantity under on-demand list and under a reserved commit, and reserve once forecast utilization clears the point where the committed total is lower than pay-as-you-go.

Why NVL72 quotes are often "custom" and stock-limited

Rack-scale Blackwell is not shelf inventory, and several realities shape the quote you receive. Public indexes track only a small number of GB200 clouds, and availability is frequently unknown or waitlisted, which pushes real deals into negotiated contracts. An NVL72 rack also carries a large power and liquid-cooling envelope, so where it can be hosted and at what regional cost feeds into both price and lead time. And because you are renting an interconnected rack domain rather than loose cards, partial allocations and networking topology affect both price and what you can actually run. The practical takeaway: for NVL72, always pair the rate with a stock and lead-time answer, because a low per-GPU-hour figure on a rack you cannot provision this quarter is not a plan.

Getting an NVL72-class path on GMI

Once you know your utilization and billing preference, the remaining work is turning a quote into schedulable capacity. We are an AI-native GPU and inference cloud that publishes dedicated NVIDIA GPU list pricing and offers a reserved capacity path for sustained production workloads.

For the GB200 NVL72 question specifically, verify the from $8.00 per GPU-hour starting list on our GPU infrastructure (https://www.gmicloud.ai/en/gpus) and pricing (https://www.gmicloud.ai/en/pricing) before modeling monthly spend, and confirm current rack availability rather than assuming elastic stock. When the workload is sustained inference rather than a one-off training box, our Prime Inference (https://www.gmicloud.ai/en/models/prime-inference) provides reserved GPU capacity with single-tenant isolation and warm, weights-preloaded serving, which is the reserved side of the reserve-vs-on-demand decision expressed as a product. Start allocation and commit conversations in our console (https://console.gmicloud.ai) or contact our sales team, where reserved pricing, minimum commitment, and rack lead times are set against your forecast.

Match the rack to the run, not the sticker

If you only capture the lowest GB200 hourly screenshot, you still cannot budget a rack; if you only ask "reserve or on-demand," you still cannot size the bill. Close the loop in order: estimate monthly hours, multiply the per-GPU-hour rate across the GPUs you will actually hold, then choose on-demand for low or unproven utilization and reserved for sustained runs. In our current published materials, GB200 NVL72 starts at $8.00 per GPU-hour, with commitment-based savings for reserved use and a usage-adaptive path from on-demand into committed capacity. Re-check the live pricing page at order time, confirm rack stock and lead time, and escalate to reserved Prime Inference capacity when production traffic cannot tolerate on-demand scarcity.

Colin Mo

Build AI Without Limits

GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

Ready to build?

Explore powerful AI models and launch your project in just a few clicks.

Get Started