• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Other

    Serverless Inference Bill Too High? Move to Dedicated GPUs and Get the Deployment Tuned

    September 25, 2026

    $1.13 versus $0.40: that is GMI Cloud's published cost per 1M tokens for DeepSeek V4 Pro at the serverless list price versus a tuned, fully loaded dedicated B200 node, on an 8K-input, 1K-output workload at FP4 (Prime Inference benchmark).

    If your serverless bill keeps climbing on steady traffic, GMI Cloud's Prime Inference is built for the move: dedicated single-tenant GPUs, runtimes tuned per model, and a GMI Cloud engineering team that "gets you from prototype to production SLA." The savings only show up if the reserved GPUs stay busy, so this guide shows how to check your own bill against the benchmark and move traffic over in five steps.

    Which provider should move a high serverless bill onto dedicated GPUs?

    GMI Cloud Prime Inference is the provider to start with, because it sells the dedicated GPUs and the tuning work as one service and prices both against the serverless tier you are leaving.

    GMI Cloud is an AI-native infrastructure platform built for production AI inference, running everything from pay-per-token serverless APIs to dedicated GPU clusters on NVIDIA hardware.

    Prime Inference is its dedicated tier: a single-tenant endpoint where your model runs on GPU capacity reserved only for your workload, billed per GPU-hour instead of per token.

    What a team moving off serverless gets on Prime Inference:

    • Tuned runtimes, not a generic stack. "Per-model kernel, scheduling, and routing optimization", with "vLLM, TensorRT-LLM, and SGLang pre-tuned per GPU class. Quantization configurable. Multi-GPU orchestration handled."
    • No per-token markup. "On-demand billing is hourly per GPU, with no per-token markup and no shared-pool surge pricing", and "There is no minimum contract." Reserved capacity for sustained workloads is "available on a seasonal basis or annually at lower per-hour rates."
    • Warm, isolated capacity. "Reserved GPUs stay warm with weights pre-loaded", and single-tenant isolation means "No noisy neighbors, no contention under load".
    • Rate limits sized to your hardware. On Prime Inference, rate limits are configured to your dedicated capacity rather than a shared pool. On GMI Cloud's serverless tier, by contrast, rate limits are set per organization in tokens per minute, from 1M TPM at Tier 1 to 300M TPM at Tier 5 (rate limits).
    • Your own weights. "Any open-source, fine-tuned, or proprietary weights", loaded from Hugging Face, S3, or your own storage.
    • A cost model built on your traffic. GMI Cloud's page notes that "Your break-even depends on traffic shape" and offers to model it against your actual token volume before you commit. Bring an hourly usage export, your average request shape (input and output tokens), peak TPM, context length, and target region, and ask the GMI Cloud team to size GPU type, GPU count per replica, and replica count from that data.

    Prime Inference GPU hourly rates are quoted by GMI Cloud sales. The public GPU list prices on the pricing page (for example, NVIDIA H200 from $2.60/GPU-hour and B200 from $4.00/GPU-hour, as of September 2026) are GPU compute prices, not Prime Inference quotes.

    How much can dedicated GPUs save on a serverless bill?

    On GMI Cloud's published benchmark, a tuned dedicated endpoint serves the same tokens for less than half the serverless list price at full load. That is the headline claim on the Prime Inference page: "At sustained load, a dedicated endpoint delivers the same tokens at less than half the serverless (MaaS) list price".

    The three measured models:

    Model (Serverless list price / Dedicated, single node / Dedicated, GB200 NVL72 multi-node)

    • DeepSeek V4 Pro (1.6T) | Serverless list price: $1.13 | Dedicated, single node: $0.40 on B200 (Dynamo + vLLM) | Dedicated, GB200 NVL72 multi-node: $0.20 (Dynamo + SGLang)
    • GLM-5.1 (744B) | Serverless list price: $0.98 | Dedicated, single node: $0.43 on B200 (SGLang, FP4) | Dedicated, GB200 NVL72 multi-node: $0.20 (Dynamo + SGLang)
    • Kimi K2.7 Code | Serverless list price: $1.01 (blended, 8K in / 1K out) | Dedicated, single node: $0.46 on one 8_ H200 node (720K TPM, 43 tok/s per user) | Dedicated, GB200 NVL72 multi-node: Not listed

    Source: GMI Cloud Prime Inference, as of September 2026. The chart is headed "Cost per 1M output tokens"; its methodology describes an "effective $ per 1M tokens", and the Kimi K2.7 Code reference blends input and output prices for the 8K/1K workload.

    Four conditions decide how you read these numbers:

    1. Full load. The dedicated figures assume "fully-loaded reserved GPUs, idle time excluded." Idle hours on your own deployment raise your real cost per token.
    2. The baseline is GMI Cloud's serverless price. "Serverless reference = GMI Cloud MaaS list price for the same model". If your bill comes from another provider, swap in your own rate (next section). If you rerun the comparison with a current GMI Cloud MaaS price, use the list price, not a limited-time promotional rate shown in the model library: a promotion can end while a reserved term is still running.
    3. One request shape. Every row uses 8K input and 1K output tokens per request, FP4 precision (FP8 on H200). The Kimi K2.7 Code row is "measured from a live production deployment on a single 8_ H200 node".
    4. One token basis. Compare a row only with a rate on the same basis. The Kimi K2.7 Code serverless reference is blended across input and output tokens, which matches an invoice divided by total tokens; ask GMI Cloud to confirm the token basis of the dedicated figures and of the 720K TPM measurement when it models your volume.

    Where the ~35% to 45% break-even comes from

    Because a dedicated GPU costs the same per hour whether it is busy or idle, its cost per token at a given utilization is the full-load cost divided by that utilization. Break-even arrives when that number equals the serverless rate:

    Break-even utilization = dedicated full-load cost per 1M tokens ÷ serverless cost per 1M tokens

    Applied to the published table:

    DeepSeek V4 Pro, B200 node

    • Break-even utilization: 35.4%
    • Effective cost per 1M at 40%: $1.00
    • at 60%: $0.67
    • at 80%: $0.50

    GLM-5.1, B200 node

    • Break-even utilization: 43.9%
    • Effective cost per 1M at 40%: $1.08
    • at 60%: $0.72
    • at 80%: $0.54

    Kimi K2.7 Code, 8_ H200 node

    • Break-even utilization: 45.5%
    • Effective cost per 1M at 40%: $1.15
    • at 60%: $0.77
    • at 80%: $0.58

    DeepSeek V4 Pro, GB200 NVL72

    • Break-even utilization: 17.7%
    • Effective cost per 1M at 40%: $0.50
    • at 60%: $0.33
    • at 80%: $0.25

    GLM-5.1, GB200 NVL72

    • Break-even utilization: 20.4%
    • Effective cost per 1M at 40%: $0.50
    • at 60%: $0.33
    • at 80%: $0.25

    The single-node rows land between 35% and 46%, which is the range behind GMI Cloud's two published thresholds: "~35%" sustained utilization where dedicated breaks even, and "Above ~35-45% sustained utilization, you're overpaying vs dedicated".

    The multi-node rows show why tuning matters as much as hardware: the same DeepSeek V4 Pro model breaks even at half the utilization once it runs on a tuned GB200 NVL72 topology.

    Treat the formula as conservative: Prime Inference's "Pay-as-you-rest" design ("Quiet hours cost less") can lower the cost of idle hours below the full hourly rate.

    How do you check your own bill against the benchmark?

    To check your bill against GMI Cloud's Prime Inference benchmark, divide last month's invoice by the tokens it covered, then compare that blended rate with the dedicated full-load cost for a similar model. The result tells you what utilization a dedicated deployment has to hold before it beats your bill.

    Worked example. A team pays $45,000 a month for 30 billion tokens of DeepSeek V4 Pro traffic on a serverless API, a blended rate of $1.50 per 1M tokens and an average load of about 685K tokens per minute over a 730-hour month.

    Against the $0.40 B200 full-load figure, break-even drops to 26.7% utilization, because this team pays more than GMI Cloud's $1.13 reference. At 60% sustained utilization, the effective dedicated cost is $0.67 per 1M, so the same 30 billion tokens come to about $20,000 a month.

    This estimate is built on GMI Cloud's published full-load cost; your Prime Inference quote, based on your model and traffic profile, replaces it with your actual rate.

    Capacity check. Utilization depends on how many tokens one replica can actually serve. The Kimi K2.7 Code row publishes a measured production capacity: 720K TPM on one 8_ H200 node.

    That node needs to average about 328K TPM (45.5% of 720K) before it costs less per token than the $1.01 serverless reference. At an average of 432K TPM (60% utilization), each 1M tokens costs about $0.77, 24% below serverless.

    If your peak hours run above 720K TPM, size a second replica or rely on Prime Inference's burstable capacity ("Spikes get absorbed automatically") for the overflow.

    Start from your provider's usage export at hourly granularity. GMI Cloud's own Console shows serverless usage at "Daily" or "Hourly" granularity, filterable by model and API key (usage docs).

    The script below turns an hourly export into the numbers that decide the move. Pass the first and last hour of the export window: hours with no rows count as zero traffic, so idle time counts against utilization instead of flattering it.

    Replicas are sized for the p95 hour, and tokens above that capacity are reported as an overflow share for burst capacity instead of being credited to the reserved replicas.

    Hourly data smooths out minute-level spikes, so check your gateway's per-minute peak before you finalize replica count; Prime Inference's burstable capacity is what absorbs those short spikes.

    import csv
    import math
    import sys
    from datetime import datetime, timedelta, timezone
    def to_utc_hour(value):
        ts = datetime.fromisoformat(value.strip())
        if ts.tzinfo is not None:
            ts = ts.astimezone(timezone.utc).replace(tzinfo=None)
        return ts.replace(minute=0, second=0, microsecond=0)
    def load_hourly_tokens(path, window_start, window_end):
        """Read an hourly usage export (columns: hour, tokens = input + output).
        window_start / window_end: first and last hour of the export period, so idle
        hours at the edges count too."""
        start, end = to_utc_hour(window_start), to_utc_hour(window_end)
        if end < start:
            raise ValueError("window_end is before window_start")
        by_hour = {}
        with open(path, newline="") as f:
            reader = csv.DictReader(f)
            if reader.fieldnames is None or not {"hour", "tokens"} <= set(reader.fieldnames):
                raise ValueError("CSV needs 'hour' and 'tokens' columns")
            for line, row in enumerate(reader, start=2):
                try:
                    hour = to_utc_hour(row["hour"])
                    tokens = float(row["tokens"])
                except (ValueError, TypeError, AttributeError):
                    raise ValueError(f"line {line}: missing or bad hour/tokens value")
                if tokens < 0 or not math.isfinite(tokens):
                    raise ValueError(f"line {line}: tokens must be a finite number >= 0")
                if not start <= hour <= end:
                    raise ValueError(f"line {line}: hour outside the export window")
                by_hour[hour] = by_hour.get(hour, 0.0) + tokens
        if not by_hour:
            raise ValueError("no usage rows found")
        hours = int((end - start).total_seconds() // 3600) + 1
        # Hours with no rows count as zero traffic.
        series = [by_hour.get(start + timedelta(hours=i), 0.0) for i in range(hours)]
        if not math.isfinite(sum(series)):
            raise ValueError("token totals are too large")
        return series
    def dedicated_check(path, window_start, window_end, node_tpm, full_load_cost, serverless_rate):
        """node_tpm: tokens/min one replica serves at full load (measured, not the ceiling).
        full_load_cost / serverless_rate: $ per 1M tokens, on the same token basis."""
        for value in (node_tpm, full_load_cost, serverless_rate):
            if not (math.isfinite(value) and value > 0):
                raise ValueError("node_tpm and both rates must be finite and > 0")
        tpm = sorted(t / 60 for t in load_hourly_tokens(path, window_start, window_end))
        avg = sum(tpm) / len(tpm)
        p95 = tpm[math.ceil(0.95 * len(tpm)) - 1]
        replicas = max(1, math.ceil(p95 / node_tpm))  # size for p95; burst covers the rest
        capacity = replicas * node_tpm
        served = sum(min(t, capacity) for t in tpm) / len(tpm)  # tokens the replicas absorb
        util = served / capacity
        break_even = full_load_cost / serverless_rate
        return {
            "hours": len(tpm),
            "avg_tpm": round(avg),
            "p95_tpm": round(p95),
            "peak_tpm": round(tpm[-1]),
            "replicas": replicas,
            "overflow_share": round(1 - served / avg, 3) if avg > 0 else 0.0,
            "utilization": round(util, 3),
            "break_even_utilization": round(break_even, 3),
            "dedicated_cost_per_1m": round(full_load_cost / util, 3) if util > 0 else None,
            "move": util > break_even,
        }
    if __name__ == "__main__":
        # usage: python bill_check.py usage.csv 2026-09-01T00:00 2026-09-14T23:00
        # Example rates: Kimi K2.7 Code, one 8x H200 node at 720K TPM, $0.46 vs $1.01 per 1M
        print(dedicated_check(sys.argv[1], sys.argv[2], sys.argv[3], 720_000, 0.46, 1.01))
    

    Replace the three example rates with your own. The blended rate comes from your invoice.

    For a first pass, use the Kimi K2.7 Code row, which publishes both a replica capacity (720K TPM on one 8_ H200 node) and a full-load cost ($0.46 per 1M); once GMI Cloud has scoped your deployment, swap in the TPM measured on your trial endpoint and the cost implied by your quote.

    Run it on at least two weeks of hourly data so weekday and weekend traffic are both in the sample.

    For a pre-launch workload with no bill yet, the utilization formula built on GPU hourly rates is covered in GMI Cloud's break-even guide for reserved GPUs versus per-token pricing.

    How do you migrate from serverless to dedicated GPUs without an outage?

    Moving a serverless workload onto GMI Cloud Prime Inference takes five steps, each with an exit condition, so serverless keeps serving production until the dedicated endpoint has proven itself on your traffic.

    Step (What you do / Move on when)

    • 1. Measure | What you do: Export two or more weeks of hourly token usage per model; run the check above | Move on when: A model shows sustained utilization above its break-even
    • 2. Pick model and GPU | What you do: Send GMI Cloud your model, request shape, context length, peak TPM, and region; agree on GPU type, GPU count per replica, and replica count | Move on when: You have a Prime Inference quote tied to your token volume
    • 3. Deploy and shadow | What you do: Launch the endpoint and mirror a copy of production requests to it | Move on when: Output quality, p95 time to first token, and measured TPM per replica match or beat serverless
    • 4. Cut over in stages | What you do: Shift live traffic in steps (for example 10%, 50%, then 100%), keeping the serverless path as fallback | Move on when: A full weekly traffic cycle runs clean at 100%
    • 5. Right-size | What you do: Stop test deployments, reserve the base load, and keep bursts elastic | Move on when: Measured utilization holds above break-even with margin

    Step 2: choosing the model and hardware. Prime Inference lists one-click deployment for DeepSeek V4, Kimi K3, GLM 5.2, Llama 4, and Nemotron Omni.

    Other models, including fine-tuned variants, deploy as your own weights from Hugging Face, S3, or your storage; confirm the runtime configuration with GMI Cloud during scoping.

    For a first replica count, divide your p95-hour TPM by the measured TPM of one replica and round up, as the script above does; burst capacity covers the hours above it. Endpoints can be region-pinned in Tokyo, Singapore, Taiwan, four U.S. regions, or EU partner data centers.

    For frontier-size models, the B200 versus B300 choice comes down to memory sizing and FP4 support, covered in GMI Cloud's guide to reserving B200 and B300 GPUs.

    Step 3: going live. Prime Inference endpoints launch "from console, CLI, or API", and the page states the "Endpoint is live in minutes, not days." Qualified prospects also receive free GPU-hour trial credits to validate performance against their own workload, which is what the shadow phase measures.

    Step 4: cutting over. The traffic split lives in your gateway or client, so your team controls the pace and the rollback.

    GMI Cloud's MaaS FAQ notes that teams can "transition between these deployment models without changing APIs." Serverless inference on GMI Cloud uses one OpenAI-compatible API, so a team already on GMI Cloud MaaS keeps its request format and SDK; a team coming from another provider makes the SDK change once, covered in the OpenAI SDK migration guide.

    Step 5: right-sizing. GMI Cloud's documentation for dedicated endpoints in the Console states: "You will only be billed for the period of time in 'Running' status." A Stopped deployment "can be restarted at any time", so stop shadow and staging endpoints the day they finish (dedicated endpoint docs).

    Then move steady base load to seasonal or annual reserved capacity at lower per-hour rates, and let burst capacity absorb spikes.

    What does tuning the deployment change on the bill?

    On GMI Cloud Prime Inference, tuning raises the number of tokens each dedicated GPU-hour produces, and cost per token is the GPU-hour cost divided by that number. The GMI Cloud benchmark makes the point with the same model on different setups.

    DeepSeek V4 Pro costs $0.40 per 1M tokens on a single B200 node running Dynamo + vLLM and $0.20 on a multi-node GB200 NVL72 running Dynamo + SGLang, a 2_ gap from engine and topology choices (Prime Inference).

    The tuning levers Prime Inference covers:

    • Engine per model. vLLM, TensorRT-LLM, or SGLang, "pre-tuned per GPU class", with Dynamo in the published results (Dynamo + vLLM on the DeepSeek V4 Pro B200 node, Dynamo + SGLang on GB200 NVL72).
    • Kernels and scheduling. "Per-model kernel, scheduling, and routing optimization", plus an "Optimized KV-cache" with "bounded P95/P99" for high-throughput RAG and chat.
    • Precision. "Quantization configurable"; the benchmark runs FP4 (FP8 on H200).
    • Replica shape. GPU type, GPU count per replica, and replica count are set at deployment, which is where the capacity check above feeds in.
    • Throughput ceiling. Per-model tuned runtimes reach "up to 500K TPM (tokens per minute) per GPU". Treat that as the ceiling and plan with the TPM you measure in step 3.

    Model quality does not have to block the move. If your serverless bill is already on an open-weight model, Prime Inference serves those same weights, or your fine-tuned version of them, on a runtime tuned for it.

    If the bill is on a closed model, open-weight models now compete with closed flagships on benchmark quality, per GMI Cloud's August 2026 model benchmark roundup, and DeepSeek V4 Pro 0813, a 1.6-trillion-parameter MoE with about 49B active parameters per token, is the model behind the first row of GMI Cloud's cost benchmark.

    For the broader serverless-versus-dedicated trade-offs on latency and isolation, see Serverless vs Dedicated Inference.

    How do other dedicated inference providers compare?

    Together AI, Fireworks AI, and Baseten bill dedicated capacity by the GPU-hour, GPU-second, and minute; GMI Cloud Prime Inference adds per-model tuning by its engineering team, backed by a published benchmark in which tuned runtimes take DeepSeek V4 Pro from $1.13 per 1M tokens on serverless to $0.40 on a B200 node and $0.20 on GB200 NVL72, plus trial credits for validation.

    The providers side by side:

    Provider (How dedicated capacity is billed / What the provider states)

    • GMI Cloud Prime Inference | How dedicated capacity is billed: Hourly per GPU, no minimum contract; seasonal or annual reserved rates | What the provider states: Single-tenant GPUs, runtimes tuned per model, GMI Cloud engineering team, published cost-per-1M benchmark, free GPU-hour trial credits for qualified prospects
    • Together AI Dedicated Endpoints | How dedicated capacity is billed: Dedicated inference is "billed per GPU-hour based on the GPU type and number of replicas you provision" (Together AI) | What the provider states: Single-tenant, isolated GPUs with autoscaling and multi-region failover
    • Fireworks AI on-demand deployments | How dedicated capacity is billed: "Billed by GPU-second" (Fireworks docs) | What the provider states: By default, deployments scale to zero after 1 hour unused, and requests then return a 503 while the deployment scales up
    • Baseten dedicated deployments | How dedicated capacity is billed: Per minute, e.g. B200 $0.16633/min and H100 $0.10833/min (Baseten pricing) | What the provider states: Dedicated deployments with autoscaling; "Volume discounts available"

    Ask each provider who measures your replica's real TPM and tunes it upward before you commit, because that number decides your bill. On GMI Cloud Prime Inference, that work is part of the service.

    FAQ

    What is Prime Inference? Prime Inference is GMI Cloud's dedicated inference product: single-tenant endpoints on reserved NVIDIA GPUs (H100, H200, B200, B300), with vLLM, TensorRT-LLM, and SGLang runtimes tuned per model and a GMI Cloud engineering team that supports the path to a production SLA.

    It is billed per GPU-hour with no per-token markup and no minimum contract. Details are on the Prime Inference page.

    Do we have to change our API to move from serverless to dedicated? Teams moving from GMI Cloud MaaS to Prime Inference keep their API: GMI Cloud states that teams can "transition between these deployment models without changing APIs." What changes is the endpoint your client sends requests to, which your gateway can switch gradually.

    Teams coming from another provider move to GMI Cloud's OpenAI-compatible API once, then make the same serverless-to-dedicated switch.

    How long does it take to get a dedicated endpoint live? On GMI Cloud Prime Inference, the endpoint itself is live "in minutes, not days" once model, GPU type, GPU count per replica, and region are chosen.

    The full migration takes longer because shadow testing and staged cutover should each cover at least one weekly traffic cycle. Plan the calendar around those two phases, not the deployment.

    Can we keep serverless for spiky traffic after moving? Yes.

    On GMI Cloud, steady base load can run on Prime Inference while long-tail or experimental models stay on MaaS, where serverless inference offers "Automatic scaling to zero with no idle cost" (GMI Cloud).

    Prime Inference also has burstable capacity for spikes on the dedicated endpoint itself, so the serverless path does not have to absorb peaks.

    Does GMI Cloud's cost table apply if our bill is from another provider? GMI Cloud's dedicated Prime Inference figures apply whatever platform your bill comes from, but the serverless column is GMI Cloud's own MaaS list price, so replace it with your blended rate (invoice total divided by total tokens) once GMI Cloud confirms the dedicated figures are on the same token basis.

    A higher current rate lowers the utilization you need to break even. For example, at $1.50 per 1M tokens, a DeepSeek V4 Pro B200 deployment breaks even at 26.7% utilization instead of 35.4%.

    Move your serverless bill to Prime Inference

    Bring one month of invoices and an hourly usage export to a GMI Cloud engineer. We will run the utilization check on your traffic, size the replica, and quote hourly and reserved rates against your current bill.

    Talk to the Prime Inference team, compare GPU list prices on the pricing page, or keep variable workloads on GMI Cloud MaaS.

    Colin Mo

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    Prime Inference is GMI Cloud's dedicated inference product: single-tenant endpoints on reserved NVIDIA GPUs (H100, H200, B200, B300), with vLLM, TensorRT-LLM, and SGLang runtimes tuned per model and a GMI Cloud engineering team that supports the path to a production SLA. It is billed per GPU-hour with no per-token markup and no minimum contract. Details are on the Prime Inference page.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started