How much does it cost to rent an NVIDIA B200 GPU in the cloud per hour, and which providers have it available right now?
July 24, 2026
Teams shopping for NVIDIA B200 capacity hit the same two blockers first: the hourly number on a pricing page, and whether that GPU can actually be provisioned this week. Public B200 cloud pricing is a band set by billing model and stock, not a single sticker that means the same thing on every vendor. This guide shows how to read that band, how to filter for real availability, and where current list prices sit so you can shortlist providers without guessing.
Cloud B200 pricing is a band, not a single sticker
NVIDIA B200 cloud rental is commonly sold as a per-GPU-hour charge for dedicated Blackwell capacity, with public on-demand list levels often landing in the mid-single-digit dollars per GPU-hour depending on provider, region, and commit. That is the practical answer to "how much per hour," but it is incomplete until you pin the commercial container around the number. Four variables move the rate before a contract is even signed:
- Billing model: On-demand, reserved or committed capacity, spot or interruptible inventory, and custom enterprise commits produce different unit rates for the same chip.
- SKU shape: Single-GPU listings, 2/4-GPU nodes, and 8-GPU SXM systems rarely invoice like a one-line "B200" tile. Per-GPU math looks attractive until you must take a full HGX-style node.
- What the hour includes: Some rate cards are GPU-only; others bundle host CPU/RAM, disk, or fabric differently. Egress, shared filesystems, and managed Kubernetes add lines that never appear in the headline figure.
- Stock class: Limited, waitlist, and region-locked inventory is often priced as if it were elastic. A low number on a sold-out SKU is not a sourcing plan.
Cross-provider indexes that scrape public listings currently show B200 quotes spanning from the low-$3 per GPU-hour range on aggressive reserved inventory up through mid-single digits on typical on-demand rows, with scarce bundles climbing higher. Treat those indexes as a market envelope, not an SLA. The real decision is which point on that envelope is buyable for your region, isolation level, and run length.
We publish dedicated NVIDIA GPU list pricing for production AI, including B200 at from $4.00 per GPU-hour under Limited Availability, alongside H100 at from $2.00 and H200 at from $2.60 per GPU-hour. That row is useful because it is checkable on our pricing page rather than implied by a private quote.
One boundary to keep clear: renting B200 GPUs by the hour is not the same commercial object as renting a GB200 NVL72 rack. B200 answers "per-GPU Blackwell capacity for training and inference nodes." NVL72 answers "liquid-cooled, rack-scale system with a different interconnect and power envelope." Collapse those into one spreadsheet cell and your per-hour and lead-time comparisons will be wrong.
How to read availability, not just the rate card
"Which providers have B200 right now?" is an operations question wearing a procurement costume. The status badge on a pricing page rarely tells the whole story, so read it against what it usually means before you budget.
| Status you see | What it usually means | What to confirm first |
|---|---|---|
| In stock / available now | Same-week self-serve provisioning is claimed | Region, max concurrent GPUs, form factor |
| Limited availability | Capacity exists but is quota'd | Allocation size, expand path, reserved baseline |
| Waitlist / coming soon | Marketing ahead of inventory | Expected window, deposit and cancel rules |
| Out of stock on list page | Public pool empty, enterprise stock may exist | Whether private-offer SKUs remain open |
Beyond the badge, validate four operational facts before you commit: region and residency, so a B200 in an irrelevant geography does not add latency or compliance cost that outweighs a small hourly gap; isolation, because shared multi-tenant slices and dedicated single-tenant GPUs are not interchangeable for audit-heavy inference; interconnect, since multi-GPU training and large-model serving depend on NVLink/NVSwitch topology; and path to reserved capacity, because production inference usually cares more about always-on warm capacity than the cheapest burst hour. If a provider cannot state rate, region, isolation, and soonest installable quantity in one conversation, it is not yet a "have it right now" answer for your workload.
Where B200 rental fits next to H100 and H200
A B200 shortlist should always be scored against the generation you can already buy in volume. In our current published listings the dedicated starting ladder is deliberately simple:
| NVIDIA GPU | GMI Cloud list rate | Availability |
|---|---|---|
| H100 | from $2.00 / GPU-hour | Available now |
| H200 | from $2.60 / GPU-hour | Available now |
| B200 | from $4.00 / GPU-hour | Limited Availability |
| GB200 NVL72 | from $8.00 / GPU-hour | Available now (rack-scale) |
That ladder does not by itself prove B200 is worth it; it frames the premium you pay for Blackwell memory and bandwidth versus staying on Hopper. Stay on H100 or H200 when the model and batch plan already clear your latency and throughput targets, when depth of stock this week matters more than peak chip performance, or when your stack is tuned on Hopper and migration cost dominates chip savings. Move toward B200 when you are memory- or bandwidth-bound on H200 for the target model size, when utilization will be high enough that throughput-per-dollar absorbs the higher sticker rate, or when you need a vendor that will discuss limited B200 allocation honestly instead of advertising infinite supply.
Getting a production B200 path without guessing inventory
Once the rate band and filters are clear, the remaining work is operational: turn a candidate SKU into an endpoint or cluster you can schedule against. Fix the commercial shape first (on-demand trial hours versus a reserved baseline for production), prove stock in writing (region, GPU count, form factor, earliest deliverable date), then match the product surface to the job.
We are an AI-native inference and GPU cloud that sells dedicated NVIDIA GPU infrastructure with public list pricing and a reserved inference path for production traffic. For this B200 question specifically, use our pricing (https://www.gmicloud.ai/en/pricing) and GPU infrastructure (https://www.gmicloud.ai/en/gpus) to verify the B200 from $4.00 per GPU-hour starting list and Limited Availability label before you model spend. Use our Prime Inference (https://www.gmicloud.ai/en/models/prime-inference) when the goal is reserved, single-tenant serving on H100, H200, or Blackwell rather than only spinning up a cold training box. Start capacity conversations in our console (https://console.gmicloud.ai) or contact our sales team when you need allocation against limited B200 stock.
Lock price and stock in the same conversation
If you only bookmark the lowest B200 hourly screenshot, you still do not know whether you can ship; if you only ask "do you have Blackwell," you still do not know the burn rate. Close the loop in one pass: ask for a unit rate, a stock status, a region, and an isolation mode together. In our current published materials, B200 begins at $4.00 per GPU-hour with Limited Availability, while H100 and H200 remain the deeper stock rungs for workloads that do not yet pay for Blackwell. Re-check the live pricing page at order time, convert hours into monthly burn at your utilization, and escalate to reserved Prime Inference capacity when production latency cannot tolerate cold starts.
Colin Mo
Build AI Without Limits
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
