Ship faster. Scale further.One Prime Inference endpoint.

Prime Inference provides dedicated single-tenant GPUs, runtimes tuned to open-source or your model, and GMI engineering team that gets you from prototype to production SLA.

Start in Console

H200 · B300 · Blackwell

NVIDIA-validated hardware

99.9% uptime*

Production SLA

Serverless pricing is a ceiling.Dedicated is how you break it.

We benchmarked the latest open models on Prime Inference reserved GPUs — tuned per-model runtimes on B200, B300, and GB200 NVL72. At sustained load, a dedicated endpoint delivers the same tokens at less than half the serverless (MaaS) list price — with single-tenant isolation, no cold starts, and predictable latency.

~35%

Sustained utilization where dedicated breaks even with serverless — above it, every token costs less

Up to 5.6×

More tokens per dollar vs serverless list price at full utilization (DeepSeek V4 Pro)

Cost per 1M output tokens — Prime Inference vs. Serverless

Effective cost at full GPU utilization, 8K input / 1K output, FP4

DeepSeek V4 Pro (1.6T)

Serverless list price

$1.13

Dedicated · B200

$0.40

Dynamo + vLLM

-56%

Dedicated · GB200 NVL72

$0.20

Dynamo + SGLang, multi-node

-82%

GLM-5.1 (744B)

Serverless list price

$0.98

Dedicated · B200

$0.43

SGLang, FP4

-56%

Dedicated · GB200 NVL72

$0.20

Dynamo + SGLang, multi-node

-80%

Kimi K2.7 Code

Serverless list price*

$1.01

Dedicated · H200 node**

$0.46

-54%
*
Blended, 8K in / 1K out
**
720K TPM · 43 tok/s per user

Serverless (MaaS) list price

Prime Inference — single node

Prime Inference — GB200 NVL72 scale-out

Methodology: 8K input / 1K output tokens per request, FP4 precision (FP8 on H200), fully-loaded reserved GPUs, idle time excluded. Dedicated cost = effective $ per 1M tokens at sustained full utilization on Prime Inference reserved capacity. Serverless reference = GMI Cloud MaaS list price for the same model; for Kimi K2.7 Code, the public serverless list price ($0.70/M input, $3.50/M output) blended for the 8K/1K workload. Kimi K2.7 Code figures are measured from a live production deployment on a single 8× H200 node. Your break-even depends on traffic shape — talk to us and we'll model it against your actual token volume.

Lease the GPU. Own the throughput.

Reserved capacity rewards real production traffic — and per-model runtime tuning compounds the advantage over time.

Tuned runtimes

Per-model kernel, scheduling, and routing optimization — not a generic stack. Pick your model, we handle the engine.

Warm by default

Reserved GPUs stay warm with weights pre-loaded. Every call lands hot — no cold-start delay, no first-token jitter.

Single-tenant isolation

GPUs reserved only for your workload. No noisy neighbors, no contention under load, no shared-tier surprises.

Bring your own model

Any open-source, fine-tuned, or proprietary weights. Load from Hugging Face, S3, or your own storage — onto a runtime built to serve it well.

Optimized for the models you use

Our inference engineers continuously tune the runtimes behind the most-deployed open-source models — so when you pick one, the kernel work is already done.

Production-grade engines

vLLM, TensorRT-LLM, and SGLang pre-tuned per GPU class. Quantization configurable. Multi-GPU orchestration handled.

Deploy close to your users.

Region-pin endpoints for first-token latency, or region-lock them for data residency.

Asia-Pacific

Tokyo · Singapore · Taiwan — serving the fastest-growing AI markets.

North America

U.S. West, East, Central, and South — high-throughput production traffic.

Europe

EU partner data centers — residency and compliance-sensitive workloads.

Scale with your traffic.

Reserved capacity when you need guaranteed performance. Burst capacity when demand spikes. Drain when it doesn't. Pay only for what you actually use.

Burstable capacity

Spikes get absorbed automatically. No queueing, no manual scaling, no failed requests during demos or launches.

Pay-as-you-rest

Quiet hours cost less. Capacity scales down gracefully without dropping in-flight calls.

One global pool

When your home region hits capacity, traffic borrows from the next-closest region to keep latency low and service continuous.

From idea to live endpoint, in four steps.

Pick a model, pick the hardware, deploy. The platform handles model loading, resource orchestration, and routing — so you go from selection to a live API in minutes.

1

Pick a model

Any open-source model, anything from Hugging Face, or upload your own weights.

2

Choose your setup

GPU type, GPU count per replica, replica count, and target region.

3

Deploy

Launch from console, CLI, or API. Endpoint is live in minutes, not days.

4

Operate & scale

Monitor latency and throughput. Burst when traffic spikes, drain when it doesn't.

Access the model you want.

One-click deployment for the leading open-source models — DeepSeek, Kimi, GLM, Llama, NVIDIA, and more. From frontier LLMs to vision, voice, and multimodal — pick a model, get a production endpoint.

DeepSeek

DeepSeek V4

deepseek-ai

Reasoning · Code
MoonshotAI

Kimi K3

moonshot-ai

1M+ Context
Z.ai

GLM 5.2

z-ai

Agentic · Tool-use
Meta

Llama 4

meta-llama

General LLM
Nvidia

Nemotron Omni

nvidia

Vision · Audio

Workloads where shared inference falls short.

Production traffic patterns where predictability, throughput, and engineering partnership turn a working prototype into a reliable product.

Coding agents & developer tools

Agents & copilots

Many short calls per task. First-call latency dominates user perception. Tool-use needs to be reliable, not just fast.

Stable endpoint per agent fleet · warm capacity · no cold-start during demos or launches.

TTS, transcription, conversation

Real-time voice

Voice doesn't tolerate variability. Persistent WebSocket sessions on warm capacity. Region-pinned for short round-trips.

Sub-second first-byte TTS · streaming endpoints · no shared-tier jitter.

RAG & chat at scale

High-throughput

Sustain millions of daily queries with hardware-bounded throughput. Consistent tail latency on long-context workloads.

Optimized KV-cache · bounded P95/P99 · no shared-pool contention.

Private & compliant deployments

Regulated

Isolated runtime, audit logs, zero-retention serving. Region-locked for finance, healthcare, public sector.

EU residency available · single-tenant isolation · enterprise SLAs.

Pick the right GPU for the job.

Hopper, Hopper-refresh, Blackwell, and Blackwell Ultra — choose by memory footprint, context length, or frontier performance need.

H100

H100

Hopper · baseline

Memory
80 GB HBM3
Inference perf
1.0× (baseline)

The workhorse. General LLM and multimodal inference. Where most production workloads start.

H200

H200

Hopper refresh

Memory
141 GB HBM3e
Inference perf
~1.4× memory & bandwidth

Memory-heavy workloads — long context, large KV-cache, big batch sizes.

B200

B200

Blackwell · frontier

Memory
192 GB HBM3e
Inference perf
Up to ~2.5× on FP4

Frontier models, FP4 inference, max throughput. For performance-critical workloads.

B300

B300

Blackwell Ultra · frontier+

Memory
288 GB HBM3e
Inference perf
Up to ~4× on FP4 (≈1.5× B200)

Maximum memory per GPU. Built for trillion-parameter models, ultra-long context, and the highest-throughput FP4 inference — when B200 isn't enough.

Trusted by Leading AI Teams

PRIME INFERENCE · DEDICATED

For sustained, production-scale traffic

  • Reserved single-tenant GPUs — no noisy neighbors, no cold starts

  • Per-model tuned runtimes: up to 500K TPM (tokens per minute) per GPU

  • Cost per token drops as utilization rises — up to 5.6× more tokens per dollar

  • Predictable latency with a 99.9%* uptime SLA

  • Custom and fine-tuned model deployment

  • Multi-region**: APAC, North America, Europe

SERVERLESS (MAAS)

For spiky or early-stage workloads

  • Pay per token, zero commitment

  • Instant start — good for prototyping and evaluation

  • Shared capacity: throughput and latency vary with platform load

  • List price stays flat no matter how much you scale

  • Above ~35-45% sustained utilization, you're overpaying vs dedicated

  • * Uptime SLA is determined by GPU series and deployment configuration. Actual committed SLAs vary by contract — most production deployments fall within the 99.9% range.
  • ** Regions currently offered. If your workload requires a specific region, we can accommodate requests case by case, subject to capacity availability.

FAQ

Frequently asked questions

Ready when you are.

Spin up a Prime Inference endpoint from the console — or contact sales about reserved capacity, custom tuning, and trial credits.

Start in Console