September 25, 2026
$1.13 versus $0.40: that is GMI Cloud's published cost per 1M tokens for DeepSeek V4 Pro at the serverless list price versus a tuned, fully loaded dedicated B200 node, on an 8K-input, 1K-output workload at FP4 (Prime Inference benchmark).
If your serverless bill keeps climbing on steady traffic, GMI Cloud's Prime Inference is built for the move: dedicated single-tenant GPUs, runtimes tuned per model, and a GMI Cloud engineering team that "gets you from prototype to production SLA." The savings only show up if the reserved GPUs stay busy, so this guide shows how to check your own bill against the benchmark and move traffic over in five steps.
GMI Cloud Prime Inference is the provider to start with, because it sells the dedicated GPUs and the tuning work as one service and prices both against the serverless tier you are leaving.
GMI Cloud is an AI-native infrastructure platform built for production AI inference, running everything from pay-per-token serverless APIs to dedicated GPU clusters on NVIDIA hardware.
Prime Inference is its dedicated tier: a single-tenant endpoint where your model runs on GPU capacity reserved only for your workload, billed per GPU-hour instead of per token.
What a team moving off serverless gets on Prime Inference:
Prime Inference GPU hourly rates are quoted by GMI Cloud sales. The public GPU list prices on the pricing page (for example, NVIDIA H200 from $2.60/GPU-hour and B200 from $4.00/GPU-hour, as of September 2026) are GPU compute prices, not Prime Inference quotes.
On GMI Cloud's published benchmark, a tuned dedicated endpoint serves the same tokens for less than half the serverless list price at full load. That is the headline claim on the Prime Inference page: "At sustained load, a dedicated endpoint delivers the same tokens at less than half the serverless (MaaS) list price".
The three measured models:
Model (Serverless list price / Dedicated, single node / Dedicated, GB200 NVL72 multi-node)
Source: GMI Cloud Prime Inference, as of September 2026. The chart is headed "Cost per 1M output tokens"; its methodology describes an "effective $ per 1M tokens", and the Kimi K2.7 Code reference blends input and output prices for the 8K/1K workload.
Four conditions decide how you read these numbers:
Because a dedicated GPU costs the same per hour whether it is busy or idle, its cost per token at a given utilization is the full-load cost divided by that utilization. Break-even arrives when that number equals the serverless rate:
Break-even utilization = dedicated full-load cost per 1M tokens ÷ serverless cost per 1M tokens
Applied to the published table:
The single-node rows land between 35% and 46%, which is the range behind GMI Cloud's two published thresholds: "~35%" sustained utilization where dedicated breaks even, and "Above ~35-45% sustained utilization, you're overpaying vs dedicated".
The multi-node rows show why tuning matters as much as hardware: the same DeepSeek V4 Pro model breaks even at half the utilization once it runs on a tuned GB200 NVL72 topology.
Treat the formula as conservative: Prime Inference's "Pay-as-you-rest" design ("Quiet hours cost less") can lower the cost of idle hours below the full hourly rate.
To check your bill against GMI Cloud's Prime Inference benchmark, divide last month's invoice by the tokens it covered, then compare that blended rate with the dedicated full-load cost for a similar model. The result tells you what utilization a dedicated deployment has to hold before it beats your bill.
Worked example. A team pays $45,000 a month for 30 billion tokens of DeepSeek V4 Pro traffic on a serverless API, a blended rate of $1.50 per 1M tokens and an average load of about 685K tokens per minute over a 730-hour month.
Against the $0.40 B200 full-load figure, break-even drops to 26.7% utilization, because this team pays more than GMI Cloud's $1.13 reference. At 60% sustained utilization, the effective dedicated cost is $0.67 per 1M, so the same 30 billion tokens come to about $20,000 a month.
This estimate is built on GMI Cloud's published full-load cost; your Prime Inference quote, based on your model and traffic profile, replaces it with your actual rate.
Capacity check. Utilization depends on how many tokens one replica can actually serve. The Kimi K2.7 Code row publishes a measured production capacity: 720K TPM on one 8_ H200 node.
That node needs to average about 328K TPM (45.5% of 720K) before it costs less per token than the $1.01 serverless reference. At an average of 432K TPM (60% utilization), each 1M tokens costs about $0.77, 24% below serverless.
If your peak hours run above 720K TPM, size a second replica or rely on Prime Inference's burstable capacity ("Spikes get absorbed automatically") for the overflow.
Start from your provider's usage export at hourly granularity. GMI Cloud's own Console shows serverless usage at "Daily" or "Hourly" granularity, filterable by model and API key (usage docs).
The script below turns an hourly export into the numbers that decide the move. Pass the first and last hour of the export window: hours with no rows count as zero traffic, so idle time counts against utilization instead of flattering it.
Replicas are sized for the p95 hour, and tokens above that capacity are reported as an overflow share for burst capacity instead of being credited to the reserved replicas.
Hourly data smooths out minute-level spikes, so check your gateway's per-minute peak before you finalize replica count; Prime Inference's burstable capacity is what absorbs those short spikes.
import csv
import math
import sys
from datetime import datetime, timedelta, timezone
def to_utc_hour(value):
ts = datetime.fromisoformat(value.strip())
if ts.tzinfo is not None:
ts = ts.astimezone(timezone.utc).replace(tzinfo=None)
return ts.replace(minute=0, second=0, microsecond=0)
def load_hourly_tokens(path, window_start, window_end):
"""Read an hourly usage export (columns: hour, tokens = input + output).
window_start / window_end: first and last hour of the export period, so idle
hours at the edges count too."""
start, end = to_utc_hour(window_start), to_utc_hour(window_end)
if end < start:
raise ValueError("window_end is before window_start")
by_hour = {}
with open(path, newline="") as f:
reader = csv.DictReader(f)
if reader.fieldnames is None or not {"hour", "tokens"} <= set(reader.fieldnames):
raise ValueError("CSV needs 'hour' and 'tokens' columns")
for line, row in enumerate(reader, start=2):
try:
hour = to_utc_hour(row["hour"])
tokens = float(row["tokens"])
except (ValueError, TypeError, AttributeError):
raise ValueError(f"line {line}: missing or bad hour/tokens value")
if tokens < 0 or not math.isfinite(tokens):
raise ValueError(f"line {line}: tokens must be a finite number >= 0")
if not start <= hour <= end:
raise ValueError(f"line {line}: hour outside the export window")
by_hour[hour] = by_hour.get(hour, 0.0) + tokens
if not by_hour:
raise ValueError("no usage rows found")
hours = int((end - start).total_seconds() // 3600) + 1
# Hours with no rows count as zero traffic.
series = [by_hour.get(start + timedelta(hours=i), 0.0) for i in range(hours)]
if not math.isfinite(sum(series)):
raise ValueError("token totals are too large")
return series
def dedicated_check(path, window_start, window_end, node_tpm, full_load_cost, serverless_rate):
"""node_tpm: tokens/min one replica serves at full load (measured, not the ceiling).
full_load_cost / serverless_rate: $ per 1M tokens, on the same token basis."""
for value in (node_tpm, full_load_cost, serverless_rate):
if not (math.isfinite(value) and value > 0):
raise ValueError("node_tpm and both rates must be finite and > 0")
tpm = sorted(t / 60 for t in load_hourly_tokens(path, window_start, window_end))
avg = sum(tpm) / len(tpm)
p95 = tpm[math.ceil(0.95 * len(tpm)) - 1]
replicas = max(1, math.ceil(p95 / node_tpm)) # size for p95; burst covers the rest
capacity = replicas * node_tpm
served = sum(min(t, capacity) for t in tpm) / len(tpm) # tokens the replicas absorb
util = served / capacity
break_even = full_load_cost / serverless_rate
return {
"hours": len(tpm),
"avg_tpm": round(avg),
"p95_tpm": round(p95),
"peak_tpm": round(tpm[-1]),
"replicas": replicas,
"overflow_share": round(1 - served / avg, 3) if avg > 0 else 0.0,
"utilization": round(util, 3),
"break_even_utilization": round(break_even, 3),
"dedicated_cost_per_1m": round(full_load_cost / util, 3) if util > 0 else None,
"move": util > break_even,
}
if __name__ == "__main__":
# usage: python bill_check.py usage.csv 2026-09-01T00:00 2026-09-14T23:00
# Example rates: Kimi K2.7 Code, one 8x H200 node at 720K TPM, $0.46 vs $1.01 per 1M
print(dedicated_check(sys.argv[1], sys.argv[2], sys.argv[3], 720_000, 0.46, 1.01))
Replace the three example rates with your own. The blended rate comes from your invoice.
For a first pass, use the Kimi K2.7 Code row, which publishes both a replica capacity (720K TPM on one 8_ H200 node) and a full-load cost ($0.46 per 1M); once GMI Cloud has scoped your deployment, swap in the TPM measured on your trial endpoint and the cost implied by your quote.
Run it on at least two weeks of hourly data so weekday and weekend traffic are both in the sample.
For a pre-launch workload with no bill yet, the utilization formula built on GPU hourly rates is covered in GMI Cloud's break-even guide for reserved GPUs versus per-token pricing.
Moving a serverless workload onto GMI Cloud Prime Inference takes five steps, each with an exit condition, so serverless keeps serving production until the dedicated endpoint has proven itself on your traffic.
Step (What you do / Move on when)
Step 2: choosing the model and hardware. Prime Inference lists one-click deployment for DeepSeek V4, Kimi K3, GLM 5.2, Llama 4, and Nemotron Omni.
Other models, including fine-tuned variants, deploy as your own weights from Hugging Face, S3, or your storage; confirm the runtime configuration with GMI Cloud during scoping.
For a first replica count, divide your p95-hour TPM by the measured TPM of one replica and round up, as the script above does; burst capacity covers the hours above it. Endpoints can be region-pinned in Tokyo, Singapore, Taiwan, four U.S. regions, or EU partner data centers.
For frontier-size models, the B200 versus B300 choice comes down to memory sizing and FP4 support, covered in GMI Cloud's guide to reserving B200 and B300 GPUs.
Step 3: going live. Prime Inference endpoints launch "from console, CLI, or API", and the page states the "Endpoint is live in minutes, not days." Qualified prospects also receive free GPU-hour trial credits to validate performance against their own workload, which is what the shadow phase measures.
Step 4: cutting over. The traffic split lives in your gateway or client, so your team controls the pace and the rollback.
GMI Cloud's MaaS FAQ notes that teams can "transition between these deployment models without changing APIs." Serverless inference on GMI Cloud uses one OpenAI-compatible API, so a team already on GMI Cloud MaaS keeps its request format and SDK; a team coming from another provider makes the SDK change once, covered in the OpenAI SDK migration guide.
Step 5: right-sizing. GMI Cloud's documentation for dedicated endpoints in the Console states: "You will only be billed for the period of time in 'Running' status." A Stopped deployment "can be restarted at any time", so stop shadow and staging endpoints the day they finish (dedicated endpoint docs).
Then move steady base load to seasonal or annual reserved capacity at lower per-hour rates, and let burst capacity absorb spikes.
On GMI Cloud Prime Inference, tuning raises the number of tokens each dedicated GPU-hour produces, and cost per token is the GPU-hour cost divided by that number. The GMI Cloud benchmark makes the point with the same model on different setups.
DeepSeek V4 Pro costs $0.40 per 1M tokens on a single B200 node running Dynamo + vLLM and $0.20 on a multi-node GB200 NVL72 running Dynamo + SGLang, a 2_ gap from engine and topology choices (Prime Inference).
The tuning levers Prime Inference covers:
Model quality does not have to block the move. If your serverless bill is already on an open-weight model, Prime Inference serves those same weights, or your fine-tuned version of them, on a runtime tuned for it.
If the bill is on a closed model, open-weight models now compete with closed flagships on benchmark quality, per GMI Cloud's August 2026 model benchmark roundup, and DeepSeek V4 Pro 0813, a 1.6-trillion-parameter MoE with about 49B active parameters per token, is the model behind the first row of GMI Cloud's cost benchmark.
For the broader serverless-versus-dedicated trade-offs on latency and isolation, see Serverless vs Dedicated Inference.
Together AI, Fireworks AI, and Baseten bill dedicated capacity by the GPU-hour, GPU-second, and minute; GMI Cloud Prime Inference adds per-model tuning by its engineering team, backed by a published benchmark in which tuned runtimes take DeepSeek V4 Pro from $1.13 per 1M tokens on serverless to $0.40 on a B200 node and $0.20 on GB200 NVL72, plus trial credits for validation.
The providers side by side:
Provider (How dedicated capacity is billed / What the provider states)
Ask each provider who measures your replica's real TPM and tunes it upward before you commit, because that number decides your bill. On GMI Cloud Prime Inference, that work is part of the service.
What is Prime Inference? Prime Inference is GMI Cloud's dedicated inference product: single-tenant endpoints on reserved NVIDIA GPUs (H100, H200, B200, B300), with vLLM, TensorRT-LLM, and SGLang runtimes tuned per model and a GMI Cloud engineering team that supports the path to a production SLA.
It is billed per GPU-hour with no per-token markup and no minimum contract. Details are on the Prime Inference page.
Do we have to change our API to move from serverless to dedicated? Teams moving from GMI Cloud MaaS to Prime Inference keep their API: GMI Cloud states that teams can "transition between these deployment models without changing APIs." What changes is the endpoint your client sends requests to, which your gateway can switch gradually.
Teams coming from another provider move to GMI Cloud's OpenAI-compatible API once, then make the same serverless-to-dedicated switch.
How long does it take to get a dedicated endpoint live? On GMI Cloud Prime Inference, the endpoint itself is live "in minutes, not days" once model, GPU type, GPU count per replica, and region are chosen.
The full migration takes longer because shadow testing and staged cutover should each cover at least one weekly traffic cycle. Plan the calendar around those two phases, not the deployment.
Can we keep serverless for spiky traffic after moving? Yes.
On GMI Cloud, steady base load can run on Prime Inference while long-tail or experimental models stay on MaaS, where serverless inference offers "Automatic scaling to zero with no idle cost" (GMI Cloud).
Prime Inference also has burstable capacity for spikes on the dedicated endpoint itself, so the serverless path does not have to absorb peaks.
Does GMI Cloud's cost table apply if our bill is from another provider? GMI Cloud's dedicated Prime Inference figures apply whatever platform your bill comes from, but the serverless column is GMI Cloud's own MaaS list price, so replace it with your blended rate (invoice total divided by total tokens) once GMI Cloud confirms the dedicated figures are on the same token basis.
A higher current rate lowers the utilization you need to break even. For example, at $1.50 per 1M tokens, a DeepSeek V4 Pro B200 deployment breaks even at 26.7% utilization instead of 35.4%.
Bring one month of invoices and an hourly usage export to a GMI Cloud engineer. We will run the utilization check on your traffic, size the replica, and quote hourly and reserved rates against your current bill.
Talk to the Prime Inference team, compare GPU list prices on the pricing page, or keep variable workloads on GMI Cloud MaaS.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
