• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Engineering

    FinOps for AI Infrastructure: Unit Economics, Chargeback, and Cost Forecasting

    AI costs are rising even as per-token prices fall, making traditional cloud FinOps approaches insufficient. This article explains how to manage AI infrastructure costs through accurate attribution, unit economics, showback and chargeback, GPU utilisation, and driver-based forecasting.

    October 01, 2026

    Per-token prices have collapsed and continue to collapse, with Gartner projecting another 90 percent drop in per-inference cost by 2030, and total AI bills keep climbing anyway. That is the defining paradox of AI cost management, and it explains why the State of FinOps 2026 report found 98 percent of organisations now managing AI spend, up from 63 percent a year earlier. The cause is not pricing. It is that cheaper inference invites more of it: longer contexts, more agentic steps, more features, more users, and an infrastructure layer where the dominant cost driver is not the token bill at all. Cast AI’s 2026 analysis across 23,000 Kubernetes clusters found enterprise GPU utilisation averaging around 5 percent, which means 95 percent of paid GPU time is producing nothing.

    • Token-billed API usage cannot idle. A GPU fleet can, and mostly does. This single asymmetry explains why the largest cost lever in most AI infrastructure is utilisation rather than price negotiation, and why the two deployment models need completely different FinOps treatment.

    • The token invoice is one of nine cost buckets. FinOps X 2026 named this explicitly: a forecast built on the token bill alone omits eight of the nine costs, which is why token-only budgeting consistently underestimates.

    • GMI Cloud returns per-request routing metadata including the model that served each call, which addresses the foundational attribution problem: model providers do not natively support the tagging structures FinOps teams rely on for allocation.

    • Showback comes before chargeback, and the sequencing is not optional. Showback gives teams visibility into what they spend without debiting their budget. Chargeback debits it. Introducing chargeback before attribution is trustworthy produces disputes rather than savings.

    • The unit that matters to leadership is not the token. A CEO or CFO thinks in cost per transaction, per resolved ticket, per document processed, and per dollar of margin. AI unit economics means the full quality-adjusted cost of completing one meaningful unit of business work, and the token is a denominator on the way there.

    • The measured upside is substantial. An analysis of 84 Bedrock deployments found cost per answer dropping from $0.41 to $0.07, an 83 percent reduction, once routing, caching, and right-sizing were in place.

    Why AI Breaks Traditional Cloud FinOps

    Traditional cloud FinOps measures resource utilisation against provisioned capacity. The practice matured around instances that run continuously, bill by the hour, and can be tagged at provisioning time. AI workloads violate three of those assumptions at once.

    Billing unit. Cloud FinOps budgets in instance-hours and storage-gigabytes, both of which finance teams have handled for a decade. AI introduces the token, a unit most finance and business stakeholders have never encountered and that resists traditional budgeting tools. A team cannot reason about whether 40 million tokens per month is reasonable without a denominator that connects it to business activity.

    Two cost structures in one practice. API-billed inference has zero idle cost and variable spend that scales with usage. Self-hosted GPU infrastructure has fixed cost and variable utilisation. These require opposite optimisations: for API spend the lever is reducing consumption, and for GPU spend the lever is increasing utilisation of capacity already paid for. A FinOps practice applying one playbook to both will optimise the wrong half.

    Attribution. An EC2 instance carries tags from the moment it is provisioned, and the billing export carries them through. An LLM API call carries whatever the request headers happen to include, which for most providers is nothing usable. The FinOps Foundation’s token economics working group identifies this as the foundational challenge: model providers do not natively support the tagging structures FinOps teams rely on for allocation.

    The Attribution Problem and How to Solve It

    Allocation is where an AI FinOps practice either becomes real or stays theoretical, and the mechanics differ entirely from cloud tagging.

    What is missing. There is no provider-side tag on an inference request. The invoice shows aggregate token consumption per model per period. It does not show which application, which team, which customer, or which feature produced it.

    Where the attribution has to happen. In the application layer, at request time, recorded alongside the request. Every call should carry an identifier for the consuming application, the team that owns it, and where relevant the end customer or tenant. Those identifiers are logged with the token counts, and the cost is computed downstream from token counts times the rate for the model that served the call.

    Why the model identity matters for attribution. Once automatic routing is in play, the rate applied to a request depends on which model handled it, which varies per request. Attribution that assumes a single rate produces a number that drifts from the invoice as the routing distribution shifts. Recording the served model per request is what keeps computed cost reconcilable against the bill. Platforms that return this in routing metadata remove the need for separate instrumentation, which is covered in Model Routing for AI Applications: How to Balance Cost, Quality, Latency, and Reliability.

    The shared-GPU version of the problem. A dedicated GPU cluster serving several teams has one invoice and no inherent split. FOCUS 1.3, ratified in December 2025, added split cost allocation for shared resources specifically to close this gap. The practical mechanism is allocating cluster cost in proportion to measured consumption, typically GPU-seconds or token throughput per team, which requires per-team usage telemetry that the cluster does not produce by default.

    The Unit Economics Ladder

    Unit economics for AI works as a ladder, and most organisations stop one or two rungs below where the number becomes useful.

    Rung one: cost per million tokens. The normalisation that makes GPU price and throughput comparable across deployment options. For self-hosted inference it is monthly GPU cost divided by monthly token throughput at achieved utilisation. For API usage it is the posted rate. This is the engineering metric, and it is the right level for comparing a model on H100 against the same model on H200, or self-hosting against an API.

    Rung two: cost per inference or cost per task. Tokens per request multiplied by the rate, which converts an abstract token figure into something an engineer can reason about per feature. This is where agentic workloads reveal themselves: a task that consumes 40 model calls has a cost per task that bears no resemblance to its cost per request.

    Rung three: cost per unit of business work. The full, quality-adjusted cost of completing one meaningful unit: resolving a support request, processing a document, approving a transaction, generating a qualified lead. This is the rung where the number becomes a business metric rather than an infrastructure one.

    Why the third rung is the one that matters. A CEO, CFO, or product leader does not think in tokens or in chargeback. They think in unit margins, cost per booking, cost per call, cost per transaction, and cost per dollar of profit. A FinOps practice reporting token consumption to leadership is reporting in the wrong unit, and the recommendation from FinOps X 2026 was explicit: report in executive units rather than FinOps ones.

    The reframe worth carrying. The metric to chase is value per token rather than cost per token, and the only way to earn it is optimising across every layer rather than negotiating the rate.

    The Utilisation Problem

    For any organisation running its own GPU capacity, utilisation is the dominant cost lever and it is usually far worse than assumed.

    The number. Around 5 percent average enterprise GPU utilisation across 23,000 clusters. The FinOps Foundation’s guidance treats any utilisation below 100 percent as a waste signal, which is a deliberately aggressive framing, but the gap between 5 percent and any reasonable target is where most of the money sits.

    Why it happens. Capacity provisioned for peak that idles through off-peak. Development and experimentation clusters left running. Jobs sized for a model that was later quantised and now needs half the hardware. Reserved capacity bought during a period of high demand that has since shifted.

    Why it is invisible without split allocation. A single cluster invoice attributed to a platform team shows a cost, not a waste. FOCUS 1.3’s split allocation makes the waste visible per team, which is generally what makes it actionable: a team that can see it is paying for 20 GPU-hours and using two has a reason to act that an aggregate platform line item does not create.

    The structural fix for bursty workloads. Capacity that scales down during quiet periods converts the idle hours from a fixed cost into no cost at all. This is the economic argument for reserved baseline capacity combined with elastic burst rather than provisioning for peak, and it is covered from the infrastructure side in GMI Prime Inference.

    Showback Before Chargeback

    Both terms describe cost allocation to consuming teams, and the difference is whether money actually moves.

    Showback reports what each team, application, or business unit consumed, without debiting anyone’s budget. It is an awareness mechanism. Teams see their spend, compare it against peers, and optimise because the number is visible and attributed.

    Chargeback debits the cost from the consuming business unit’s operational budget. It creates real financial accountability and real incentive, and it also creates real disputes when the attribution is wrong.

    The sequencing rule. Showback first, for long enough that the attribution is trusted. Every source in the FinOps literature agrees on this, and the reason is practical: the first chargeback cycle with flawed attribution consumes more organisational energy in dispute resolution than the savings are worth, and it damages confidence in the FinOps function at exactly the point where it needs credibility.

    What has to be true before chargeback. Attribution covering effectively all spend rather than most of it, with the unallocated remainder small enough to absorb centrally. Reports that reconcile against the invoice. A defined process for disputing an allocation. And an agreed treatment for shared infrastructure that the consuming teams accept as fair before the first bill lands.

    Maturity by spend. The practical tiers observed in the field: at $50,000 to $250,000 per month, dedicated AI FinOps tooling with chargeback and weekly review becomes worth the overhead. At $250,000 to $1 million per month, unit economics, automation, and active commitment portfolio management become necessary. Below the first threshold, showback with monthly review is usually the right level of investment.

    Forecasting Under Token Billing

    Traditional cost forecasting models built around resource allocation do not generalise to token consumption, and the reason is that the drivers are different.

    What a cloud forecast models. Provisioned capacity, growth in that capacity, and commitment coverage. The quantities are relatively stable and change through deliberate provisioning decisions.

    What an AI forecast has to model. Request volume, tokens per request, the model mix and therefore the blended rate, cache hit rate, and the eight cost buckets that are not the token invoice. Each of these moves independently, and two of them move without anyone deciding anything: the model mix shifts when routing distribution changes, and tokens per request grow when prompts, contexts, or agent step counts grow.

    The forecast that works. Build it from drivers rather than from trend extrapolation. Forecast request volume from product plans, tokens per request from current measurement, and blended rate from the observed routing distribution. Then sanity-check against the trailing trend rather than deriving the forecast from it.

    The variance sources to watch. A shift in the routing distribution toward more expensive models raises spend with no change in volume. A prompt change that alters verbosity changes output token consumption before any quality metric moves. Context accumulation in agentic sessions grows input tokens per step as a session progresses. Each of these breaks a forecast built on a flat rate per request.

    The nine-bucket warning. The token invoice is one line. The others include GPU infrastructure, data storage and movement, vector database and retrieval infrastructure, observability and evaluation tooling, the AI tooling subscriptions accumulated across departments, engineering time, and the fine-tuning or training runs that are bursty and easy to omit from a steady-state forecast. A forecast built on the token bill alone is wrong by construction.

    FOCUS and Why Standardisation Matters Here

    The FinOps Open Cost and Usage Specification defines standard fields and terminology for technology billing data, and three recent revisions address AI specifically.

    Revision

    Ratified

    What it added for AI

    FOCUS 1.2

    29 May 2025

    Virtual-currency and token-lifecycle support, making per-token billing normalisable

    FOCUS 1.3

    4 December 2025

    Split cost allocation for shared resources, closing the shared-GPU-cluster gap

    FOCUS 1.4

    4 June 2026

    Invoice Detail and Billing Period datasets for reconciliation, expanded commitment data for cross-provider analysis

    Why this matters practically. Before token-lifecycle support existed in the specification, every provider’s token billing had its own shape and cross-provider comparison required custom normalisation per vendor. Standardised fields make a multi-provider AI cost report assemblable rather than a bespoke engineering project.

    The remaining gap. FOCUS standardises the billing export. It does not solve attribution, because the provider still does not know which of your applications made the call. Standardisation helps with reconciliation and cross-provider comparison; the allocation dimension still has to be added at request time by you.

    The Operating Model That Works

    The organisational pattern that succeeds mirrors cloud FinOps a decade ago, and the division of responsibility is the part worth copying.

    A small central function owns rates and negotiation, commitment portfolio, tooling, the allocation methodology, and reporting standards. This is where the expertise concentrates and where cross-team consistency is enforced.

    Engineering teams own their unit economics. Cost per task, per feature, per customer, for the workloads they build. The central function supplies the data and the method; the team owns the number and the decisions that move it.

    Finance owns budget allocation and the chargeback mechanism, and increasingly participates in the forecast rather than receiving it.

    The cross-functional requirement. Effective AI cost governance needs data science, MLOps, platform engineering, and finance in the same conversation, because the levers sit in different places. Model selection is a data science decision with a cost consequence. Utilisation is a platform decision. Prompt length is an application decision. None of these is visible to finance, and none of them is framed as a cost decision by the team making it unless someone connects the two.

    The Levers, Ranked

    The measured outcomes across published case studies point at a consistent ordering.

    Utilisation, for self-hosted capacity. Moving from single-digit utilisation toward a 60 to 85 percent target is the largest single change available, and it requires no negotiation and no quality tradeoff. Consolidating workloads, scaling down off-peak, and right-sizing after quantisation are the mechanics.

    Model right-sizing. Routing requests to the smallest model that meets the quality requirement. Analysis of 84 Bedrock deployments found cost per answer falling from $0.41 to $0.07 once routing, caching, and right-sizing were applied together, an 83 percent reduction.

    Prompt caching. Correctly configured caching produces substantial input-token reduction for workloads with stable prefixes, at no quality cost. The mechanics and the structural requirements are covered in LLM Inference Cost Optimization: Caching, Batching, and Routing.

    Prompt and context pruning. Reducing tokens per request directly reduces cost per request. Accumulated context in agentic sessions and unnecessarily verbose system prompts are the usual targets.

    Tool rationalisation. Departmental AI subscriptions accumulate without central visibility. Consolidating overlapping tools is a finance lever rather than an engineering one, and it frequently recovers more than an engineering optimisation would.

    Commitment management. Reserved and committed capacity for the predictable baseline, on-demand or burst for the variable portion. This is classic FinOps applied to GPU capacity, and it only works once the baseline is actually known, which requires the attribution work to be done first.

    What GMI Cloud Provides for the Attribution Layer

    Three platform properties map onto the FinOps requirements above.

    Per-request routing metadata. Every response identifies the model that served the request, which is the field that makes computed cost reconcilable against the invoice once automatic model selection is in play. Without it, cost attribution assumes a rate that drifts as the routing distribution shifts. The metadata also indicates whether a fallback was applied, which matters when the fallback model has a different rate, as described in Primary and Fallback Models: How Auto-Routing Improves Reliability in AI Applications.

    Reserved baseline with elastic burst. The structural answer to the utilisation problem for bursty workloads. Capacity that scales down during quiet hours converts idle GPU-hours into no cost rather than into a fixed line item, which is the difference between 5 percent utilisation on provisioned-for-peak capacity and a meaningful utilisation figure on a right-sized baseline.

    Per-minute billing with no minimum commitment. For the experimentation and evaluation workloads that are bursty by nature, hourly or per-minute billing means the idle gap between runs is not paid for. This matters for the fine-tuning and evaluation bucket specifically, which is one of the nine and one of the easiest to over-provision.

    Transparent per-GPU-hour pricing. Cost per million tokens computed from a known hourly rate and measured throughput is a number the engineering team can own. Blended or bundled pricing makes the same calculation an estimate. For comparing deployment options, the transparency is what makes the comparison possible, which is the subject of GPU Cloud Computing Cost Comparison and TCO.

    Conclusion

    AI cost management is not cloud cost management with a new resource type. It has a billing unit that finance has not handled before, two cost structures requiring opposite optimisations, and an attribution gap that providers do not close for you.

    The sequence that works starts with attribution, because nothing downstream is trustworthy without it. Then showback, for long enough that teams believe the numbers. Then chargeback, once the allocation is defensible. Unit economics climbs from cost per million tokens through cost per task to cost per unit of business work, and only the third rung produces a number that leadership can act on.

    Two findings anchor the effort. Enterprise GPU utilisation averaging around 5 percent means the largest available saving for self-hosted capacity requires no negotiation and no quality tradeoff. And the token invoice being one of nine cost buckets means a forecast built on it alone is wrong before it is finished.

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    Because cheaper inference invites more of it, and because the token bill is not where most of the cost sits. Per-inference prices have collapsed and Gartner projects a further 90 percent drop by 2030, while total spend climbs through longer contexts, more agentic steps per task, more features, and more users. Underneath that, the infrastructure layer carries the larger problem: Cast AI’s 2026 analysis across 23,000 clusters found enterprise GPU utilisation averaging around 5 percent, meaning 95 percent of paid GPU time produces nothing. Token-billed API usage cannot idle, but a GPU fleet can and mostly does.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started