
H100
Hopper · baseline
- Memory
- 80 GB HBM3
- Inference perf
- 1.0× (baseline)
The workhorse. General LLM and multimodal inference. Where most production workloads start.
Prime Inference provides dedicated single-tenant GPUs, runtimes tuned to open-source or your model, and GMI engineering team that gets you from prototype to production SLA.
H200 · B300 · Blackwell
NVIDIA-validated hardware
99.9% uptime*
Production SLA

We benchmarked the latest open models on Prime Inference reserved GPUs — tuned per-model runtimes on B200, B300, and GB200 NVL72. At sustained load, a dedicated endpoint delivers the same tokens at less than half the serverless (MaaS) list price — with single-tenant isolation, no cold starts, and predictable latency.
~35%
Sustained utilization where dedicated breaks even with serverless — above it, every token costs less
Up to 5.6×
More tokens per dollar vs serverless list price at full utilization (DeepSeek V4 Pro)
Effective cost at full GPU utilization, 8K input / 1K output, FP4
Serverless list price
$1.13
Dedicated · B200
$0.40
Dynamo + vLLM
-56%Dedicated · GB200 NVL72
$0.20
Dynamo + SGLang, multi-node
-82%Serverless list price
$0.98
Dedicated · B200
$0.43
SGLang, FP4
-56%Dedicated · GB200 NVL72
$0.20
Dynamo + SGLang, multi-node
-80%Serverless list price*
$1.01
Dedicated · H200 node**
$0.46
Serverless (MaaS) list price
Prime Inference — single node
Prime Inference — GB200 NVL72 scale-out
Methodology: 8K input / 1K output tokens per request, FP4 precision (FP8 on H200), fully-loaded reserved GPUs, idle time excluded. Dedicated cost = effective $ per 1M tokens at sustained full utilization on Prime Inference reserved capacity. Serverless reference = GMI Cloud MaaS list price for the same model; for Kimi K2.7 Code, the public serverless list price ($0.70/M input, $3.50/M output) blended for the 8K/1K workload. Kimi K2.7 Code figures are measured from a live production deployment on a single 8× H200 node. Your break-even depends on traffic shape — talk to us and we'll model it against your actual token volume.
Reserved capacity rewards real production traffic — and per-model runtime tuning compounds the advantage over time.
Per-model kernel, scheduling, and routing optimization — not a generic stack. Pick your model, we handle the engine.
Reserved GPUs stay warm with weights pre-loaded. Every call lands hot — no cold-start delay, no first-token jitter.
GPUs reserved only for your workload. No noisy neighbors, no contention under load, no shared-tier surprises.
Any open-source, fine-tuned, or proprietary weights. Load from Hugging Face, S3, or your own storage — onto a runtime built to serve it well.
Our inference engineers continuously tune the runtimes behind the most-deployed open-source models — so when you pick one, the kernel work is already done.
vLLM, TensorRT-LLM, and SGLang pre-tuned per GPU class. Quantization configurable. Multi-GPU orchestration handled.

Region-pin endpoints for first-token latency, or region-lock them for data residency.

Tokyo · Singapore · Taiwan — serving the fastest-growing AI markets.
U.S. West, East, Central, and South — high-throughput production traffic.
EU partner data centers — residency and compliance-sensitive workloads.
Reserved capacity when you need guaranteed performance. Burst capacity when demand spikes. Drain when it doesn't. Pay only for what you actually use.

Spikes get absorbed automatically. No queueing, no manual scaling, no failed requests during demos or launches.
Quiet hours cost less. Capacity scales down gracefully without dropping in-flight calls.
When your home region hits capacity, traffic borrows from the next-closest region to keep latency low and service continuous.
Pick a model, pick the hardware, deploy. The platform handles model loading, resource orchestration, and routing — so you go from selection to a live API in minutes.
Any open-source model, anything from Hugging Face, or upload your own weights.
GPU type, GPU count per replica, replica count, and target region.
Launch from console, CLI, or API. Endpoint is live in minutes, not days.
Monitor latency and throughput. Burst when traffic spikes, drain when it doesn't.

One-click deployment for the leading open-source models — DeepSeek, Kimi, GLM, Llama, NVIDIA, and more. From frontier LLMs to vision, voice, and multimodal — pick a model, get a production endpoint.
deepseek-ai
moonshot-ai
z-ai
meta-llama
nvidia
Production traffic patterns where predictability, throughput, and engineering partnership turn a working prototype into a reliable product.

Many short calls per task. First-call latency dominates user perception. Tool-use needs to be reliable, not just fast.
Stable endpoint per agent fleet · warm capacity · no cold-start during demos or launches.

Voice doesn't tolerate variability. Persistent WebSocket sessions on warm capacity. Region-pinned for short round-trips.
Sub-second first-byte TTS · streaming endpoints · no shared-tier jitter.

Sustain millions of daily queries with hardware-bounded throughput. Consistent tail latency on long-context workloads.
Optimized KV-cache · bounded P95/P99 · no shared-pool contention.

Isolated runtime, audit logs, zero-retention serving. Region-locked for finance, healthcare, public sector.
EU residency available · single-tenant isolation · enterprise SLAs.

Hopper, Hopper-refresh, Blackwell, and Blackwell Ultra — choose by memory footprint, context length, or frontier performance need.

Hopper · baseline
The workhorse. General LLM and multimodal inference. Where most production workloads start.

Hopper refresh
Memory-heavy workloads — long context, large KV-cache, big batch sizes.

Blackwell · frontier
Frontier models, FP4 inference, max throughput. For performance-critical workloads.

Blackwell Ultra · frontier+
Maximum memory per GPU. Built for trillion-parameter models, ultra-long context, and the highest-throughput FP4 inference — when B200 isn't enough.
PRIME INFERENCE · DEDICATED
Reserved single-tenant GPUs — no noisy neighbors, no cold starts
Per-model tuned runtimes: up to 500K TPM (tokens per minute) per GPU
Cost per token drops as utilization rises — up to 5.6× more tokens per dollar
Predictable latency with a 99.9%* uptime SLA
Custom and fine-tuned model deployment
Multi-region**: APAC, North America, Europe
SERVERLESS (MAAS)
Pay per token, zero commitment
Instant start — good for prototyping and evaluation
Shared capacity: throughput and latency vary with platform load
List price stays flat no matter how much you scale
Above ~35-45% sustained utilization, you're overpaying vs dedicated
Frequently asked questions

Spin up a Prime Inference endpoint from the console — or contact sales about reserved capacity, custom tuning, and trial credits.