July 24, 2026
The GB200 NVL72 is often described as a rack of 72 GPUs, which is accurate but misses the point that makes it interesting. What matters is not that it holds 72 GPUs, but that they behave as one large accelerator rather than 72 separate cards in a network. The GB200 NVL72 is a rack-scale system that connects 72 Blackwell GPUs and 36 Grace CPUs through a single high-bandwidth NVLink domain, so the GPUs can share memory and exchange data as if they were one very large GPU, which is what lets it serve models too big for any single card without a network penalty. This guide explains what the system is and how those 72 GPUs actually work together.
Most multi-GPU setups are separate servers connected over a network, where each server holds a handful of GPUs and talks to the others over comparatively slow links. The GB200 NVL72 is designed differently: it is a single rack-scale unit where the GPUs are wired together by NVLink at bandwidth far above ordinary networking.
That design choice is the whole idea. In a GB200 NVL72, all 72 Blackwell GPUs connect through a unified NVLink fabric rather than a server-to-server network, so a model sharded across them exchanges data over high-bandwidth links instead of a slow interconnect, which is the difference between 72 GPUs acting as one system and 72 GPUs acting as a cluster. The rack also pairs those GPUs with 36 Grace CPUs, coupling CPU and GPU memory tightly, and it is liquid-cooled because packing this much compute into one rack generates heat that air cooling cannot handle. The result is a system engineered to be a single large accelerator, not an assembly of independent nodes.
The mechanism that makes the 72 GPUs act as one is the shared NVLink domain, and it matters because of what large models need.
A model too big for one GPU has to be split across many, and every time the pieces need to exchange activations, they communicate. On a networked cluster, that communication crosses slow inter-node links and becomes a bottleneck that wastes GPU time. Inside the NVL72, that same communication travels over the NVLink fabric at much higher bandwidth, so the GPUs can pool their memory and pass data with far less overhead. In effect, the 72 GPUs present a single large memory space that a model can occupy, which is why the system can hold and serve models that no single GPU or small cluster could. The Grace CPUs add to this by feeding data into the GPU domain efficiently, keeping the accelerators supplied rather than waiting on a slower host connection.
Understanding the design tells you what the system suits. Use the frame below.
| Characteristic | What it means | Best suited for |
|---|---|---|
| 72 GPUs in one NVLink domain | Acts as one large accelerator | Models too big for one GPU or node |
| 36 Grace CPUs coupled to GPUs | Fast host-to-GPU data feeding | Data-heavy training and inference |
| Liquid-cooled rack unit | High compute density in one system | Large-scale, sustained workloads |
| Single high-bandwidth fabric | Low communication overhead | Sharded models needing tight coupling |
The pattern is that the NVL72 is built for problems that need many GPUs to act as one: very large models, tightly coupled training, and high-density inference. It is not built to be split among many small independent jobs, which gain nothing from the shared domain and would use it inefficiently. The system is a single large accelerator, and it pays off on workloads that need exactly that.
Since the NVL72 is a specialized rack-scale system, you can access one through our platform, where we operates it and publishes its pricing. We are an AI-native GPU and inference cloud that lists GB200 NVL72 with dedicated NVIDIA GPU pricing.
We currently list GB200 NVL72 at from $8.00 per GPU-hour Available Now, giving access to the full 72-GPU NVLink domain for models that need many GPUs to act as one system rather than a networked cluster. Verify the current rate and rack availability on our GPU infrastructure (https://www.gmicloud.ai/en/gpus) and pricing (https://www.gmicloud.ai/en/pricing), since Blackwell pricing and stock move quickly, and confirm rack quantity and region rather than assuming elastic supply. Match the system to the workload: an NVL72 pays off when a single model needs the pooled domain, so benchmark a model large enough to use it. When the workload is sustained large-model serving, our Prime Inference (https://www.gmicloud.ai/en/models/prime-inference) provides reserved GPU capacity with single-tenant isolation and warm serving on rack-scale hardware. Start a conversation in our console (https://console.gmicloud.ai) or contact our sales team.
The GB200 NVL72 is best understood not as 72 GPUs but as one large accelerator built from them, joined by a single NVLink domain so they share memory and pass data without a network penalty. That design is what lets it serve models too big for any single card, and it is why the system suits very large models and tightly coupled training rather than many small independent jobs. If your workload needs many GPUs to act as one, the NVL72 is the system built for it, so match it to a model large enough to use the domain it was designed around.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
