The open versus closed decision used to be a capability decision. At the end of 2023, the best closed model scored roughly 88 percent on MMLU while the best open alternative managed around 70.5 percent, a 17.5-point gap that settled the argument for anything demanding. By early 2026 that gap is effectively zero on knowledge benchmarks and in the single digits on most reasoning tasks, with Qwen 3.5 scoring 88.4 on GPQA Diamond and outperforming every closed model except the most expensive frontier options.
September 30, 2026
.webp)
The open versus closed decision used to be a capability decision. At the end of 2023, the best closed model scored roughly 88 percent on MMLU while the best open alternative managed around 70.5 percent, a 17.5-point gap that settled the argument for anything demanding. By early 2026 that gap is effectively zero on knowledge benchmarks and in the single digits on most reasoning tasks, with Qwen 3.5 scoring 88.4 on GPQA Diamond and outperforming every closed model except the most expensive frontier options. The decision did not disappear when the capability gap closed. It moved: from “which is better” to “which set of constraints does your deployment actually have,” and the answer now turns on data residency, total cost of ownership at your specific volume, and whether your team can absorb the operational burden that owning the infrastructure creates.
Most models marketed as open source are open-weight. The Open Source Initiative published a formal Open Source AI Definition, and models that publish weights under permissive licences while withholding training data do not meet it. The distinction matters for legal review, not just terminology.
The realistic floor for a self-hosted production deployment is $125,000 to $190,000 annually once staff time, infrastructure, and operations are counted. Token price comparisons that ignore this line item consistently blow their projections.
GMI Cloud removes the operational burden from the open-weight side of the decision by providing managed access to 170-plus open and closed models through one API, plus dedicated GPU capacity for teams that need self-hosted deployment without datacentre operations.
The break-even against frontier APIs sits between 2 and 10 million tokens per day, depending on model size and input-output ratio. Below that range, self-hosting rarely justifies its fixed costs on economics alone.
Data residency, not cost, is the constraint that most often forces the decision. When regulation requires that data never leave a jurisdiction or a network perimeter, open-weight self-hosting is the only path, and the economics become a secondary consideration.
The consensus 2026 architecture is hybrid. Closed models where frontier reasoning and low operational overhead matter most, open-weight models for private deployment, fine-tuning, cost control, and data residency. This is not a compromise between two positions; it is the position most production teams converge on.
The phrase “open source LLM” is used loosely enough that it obscures the properties that matter for a production decision. Getting the categories right prevents legal surprises late in a procurement cycle.
Open source, formally. The Open Source Initiative’s Open Source AI Definition requires access to the training data, the training code, and the weights, under terms that permit use, study, modification, and sharing. Very few released models meet this standard.
Open-weight. The weights are published, typically under a permissive licence such as MIT or Apache 2.0, but the training data and often the training code are withheld. This covers the models most teams mean when they say open source: the Llama family, Qwen, DeepSeek, GLM, Mistral, and Kimi releases. The weights can be downloaded, run, fine-tuned, and deployed on infrastructure you control, which is what matters operationally.
Open-weight with restrictions. Some licences permit commercial use with conditions: usage thresholds above which separate licensing is required, prohibited use categories, attribution requirements, or restrictions on using outputs to train competing models. Meta’s Llama licence, Kimi K3’s Modified MIT, and several others carry terms that a standard open-source licence review would not anticipate.
Closed. Weights are not published. Access is through the vendor’s API or through cloud marketplace channels such as Amazon Bedrock, Google Vertex AI, or Microsoft Foundry. GPT, Claude, and Gemini flagships sit here.
Why the distinction matters in practice. Legal review treats “open source” and “open-weight with commercial restrictions” very differently. A team that assumed Apache 2.0 terms and discovers a usage threshold or a prohibited-use clause during a compliance review loses weeks. Read the actual licence for each model in the candidate pool before the architecture depends on it.
The economics changed because the capability argument stopped being decisive.
Period | Best closed | Best open | Gap |
End of 2023 | ~88% MMLU | ~70.5% MMLU | 17.5 points |
Early 2026 | Frontier | Effectively parity on knowledge | ~0 points |
On reasoning tasks the gap persists at single digits. On agentic coding and complex multi-step reasoning, the frontier closed models retain a measurable lead: GPT-5.3 Codex and the Claude Opus line continue to push the ceiling, and the open-weight field follows closely rather than leading.
Where open-weight models now lead or match: knowledge retrieval, instruction following, multilingual tasks, code generation on well-specified problems, and cost per unit of quality across the board.
Where closed frontier models retain an advantage: long-horizon agentic execution, complex multi-file reasoning, and the tail of hard problems where the difference between 85 and 95 percent task completion determines whether the feature works.
The evaluation landscape also shifted. The Hugging Face Open LLM Leaderboard was retired and archived in 2025, with the community moving toward human-preference evaluation such as Chatbot Arena and task-specific frameworks such as Stanford HELM. This matters for the decision because it means aggregate leaderboard position is a weaker signal than it was, and task-specific evaluation on your own workload is a stronger one.
The decision is rarely a single-factor call. Five dimensions each have a threshold above which they become decisive.
Dimension 1: Data residency and privacy.
This is the dimension that most often overrides everything else. When regulation requires that data never leave a jurisdiction, or that it never be processed by a third party, open-weight self-hosting is the only compliant path. Healthcare records under HIPAA, EU personal data under GDPR where the transfer mechanism is unavailable, financial records under sector-specific rules, and government or defence workloads all produce this constraint.
The nuance: closed models accessed through cloud marketplace channels can satisfy many data residency requirements. A model available through Amazon Bedrock, running in a specific AWS region under an existing BAA, brings the model inside a compliance framework the organisation already has. This path covers a substantial share of regulated workloads without self-hosting.
Where it does not work: when the requirement is that no third party processes the data at all, rather than that the processing occurs in a specific jurisdiction under contract. That requirement is only satisfied by running the weights on infrastructure you control.
Dimension 2: Volume and cost.
Token price comparison alone is misleading, but the gap at the extremes is real. A tuned self-hosted 8B model runs near $0.20 per million tokens marginal cost. Frontier flagship input pricing runs $5.00 per million for GPT-5.5 as of June 2026, with output billing between $25 and $180 per million depending on the model and mode.
The break-even is volume-dependent, and the published estimates cluster in a consistent range: 2 to 5 million tokens per day against frontier API pricing for a reserved cloud GPU deployment, and 5 to 10 million tokens per day for a 70B model against a premium API tier. Below roughly 2 million tokens per day, the fixed costs of self-hosting rarely pay back.
Dimension 3: Customisation requirements.
Fine-tuning on proprietary data is available on open weights and unavailable on closed frontier models. For most workloads this does not matter, because prompting and retrieval cover the customisation requirement. For a specific set, it is decisive: teams with substantial proprietary domain corpora where a fine-tuned model measurably outperforms prompting on the target task have a requirement that no amount of prompt engineering satisfies.
The practical caveat: most teams that need fine-tuning fine-tune a smaller model rather than a frontier-scale one, because the cost difference is two orders of magnitude and the quality difference on a narrow domain task is often small.
Dimension 4: Operational capacity.
This is the dimension most consistently underestimated. Self-hosting requires at minimum one MLOps engineer to manage deployment, monitoring, serving framework selection and tuning, quantisation decisions, uptime and redundancy, and model migration when a better model releases. There is no SLA unless the team builds one.
A team without GPU infrastructure experience treating self-hosting as a cost optimisation is taking on an operational risk that frequently costs more than the savings.
Dimension 5: Vendor dependency and lifecycle control.
Closed models carry a set of risks that are manageable but should be planned for rather than discovered: pricing changes and restructuring, rate limits and availability outside your control, model updates that alter behaviour you relied on, deprecation on the vendor’s schedule rather than yours, and limited transparency into model internals.
Open weights invert this. The model is never deprecated without your explicit decision, which allows building long-term products without the threat of forced migration. The open-weight ecosystem is also highly competitive, with better models releasing continuously, so a self-hosted architecture can swap in improvements on the team’s own schedule.
Token price is the tip of the iceberg. The cost drivers below the surface are where self-hosting projections fail.
The headline comparison that misleads. A self-hosted Llama 3-8B at roughly $0.20 per million tokens against a frontier flagship at $15 to $30 per million input tokens looks like a settled argument. It is not, because the $0.20 figure is marginal cost on already-provisioned hardware.
The fixed costs.
Cloud GPU rental for a production 70B deployment runs $6,000 to $17,000 or more per month. H100 80GB pricing stabilised in early 2026 at $1.49 to $3.90 per hour depending on provider, with A100 80GB dropping to $0.66 to $0.78 on spot markets.
Operations headcount is the line item teams most often omit. At least one MLOps engineer, in practice often a portion of several engineers’ time, to handle deployment, monitoring, incident response, framework upgrades, and migration.
The idle penalty applies to any workload with bursty traffic. A deployment provisioned for peak that runs at 20 to 40 percent average utilisation is paying a 60 to 80 percent idle penalty on the GPU line, which changes the effective cost per token substantially.
The realistic floor. A minimal production self-hosted deployment costs $125,000 to $190,000 annually once staff time, infrastructure, and operations are counted. Any comparison that shows self-hosting winning at low volume has almost certainly omitted part of this.
The quantisation tradeoff. Self-hosting permits quantisation, which reduces the hardware requirement meaningfully. Q4_K_M retains approximately 92 percent of original quality; AWQ retains approximately 95 percent. Whether that loss is acceptable is workload-specific, and it should be measured on the target task rather than assumed.
Where the managed middle ground sits. The framing above presents a binary: managed API with no operational burden and per-token pricing, or self-hosted with full control and full operational responsibility. A third option exists and is frequently the right answer: managed infrastructure running open-weight models, where the provider handles the serving stack, the hardware, and the availability, while the team retains model choice and data residency control.
GMI Prime Inference occupies this position: dedicated GPU capacity with pre-configured serving frameworks and per-model runtime tuning, which removes the MLOps burden while preserving the open-weight advantages of model choice, fine-tuned weight deployment, and single-tenant data isolation.
The calculation that supports the decision requires four inputs: monthly token volume split between input and output, the frontier API rate for the model you would otherwise use, the fully loaded self-hosting cost at your required capacity, and your achievable GPU utilisation.
The formula.
Self-hosted effective cost per million tokens equals monthly fixed cost divided by monthly token throughput. Monthly token throughput equals tokens per second at your batch size, times seconds per month, times utilisation rate.
A worked example. A 70B model on a single H200 at $2.60 per hour produces roughly 2,000 to 3,000 output tokens per second with continuous batching at moderate batch sizes. At 70 percent utilisation over 730 hours: approximately 4.6 to 6.9 billion output tokens per month at $1,898 in GPU cost.
That is $0.28 to $0.41 per million output tokens on compute alone. Adding a conservative $8,000 per month for a fraction of an MLOps engineer’s time brings the effective cost to roughly $1.44 to $2.15 per million output tokens at that volume.
Against a frontier API at $25 per million output tokens, self-hosting wins decisively at this volume. Against a mid-tier managed open-weight API at $0.60 per million output tokens, it does not.
This is the comparison most teams get wrong. The relevant alternative is frequently not a frontier closed model. It is the same open-weight model served by a managed provider. That comparison has a much higher break-even, because the managed provider is also running the model efficiently at scale and passing most of that efficiency through.
The frontier-versus-self-hosted comparison answers whether to use an open-weight model. The managed-versus-self-hosted comparison answers whether to run it yourself, and it is the second question that determines infrastructure strategy.
Most organisations in 2026 run a hybrid strategy, and the reasoning behind it is structural rather than a hedge.
Closed frontier models for the workloads where frontier reasoning capability is the requirement and the volume is low enough that per-token pricing is acceptable. Complex agentic execution, multi-file code reasoning, and the hard tail of problems where task completion rate determines whether the feature works.
Managed open-weight models for the majority of production traffic: instruction following, knowledge retrieval, summarisation, classification, routing, and code generation on well-specified problems. This is typically 60 to 80 percent of request volume, and open-weight models handle it at equivalent user-perceived quality for a fraction of the cost.
Self-hosted open-weight models for workloads with data residency requirements no managed option satisfies, for fine-tuned models on proprietary data, and for sustained high-volume workloads where the break-even favours owned capacity.
The migration pattern that works. Start with closed APIs for speed, ship the product, then migrate specific high-volume or privacy-sensitive workloads to open-weight models when the data or the economics demand it. A 500-person software company adding AI code review, document summarisation, and meeting notes to internal tools shipped its first feature in two weeks on closed APIs; the self-hosting path would have taken two months of infrastructure work before the first feature existed.
At 5,000 requests per day across all features, that deployment cost roughly $800 per month via APIs against $3,000 or more per month self-hosting on A100s. The decision was straightforward, and it would reverse at higher volume.
Why routing makes the hybrid practical. A hybrid architecture requires deciding which model handles which request, and doing that manually per call site does not scale. As covered in GMI Cloud’s guide to model routing, automatic routing based on detected task type sends each request to the appropriate model tier without the application needing to encode the decision. This is what converts the hybrid from an architectural aspiration into a working configuration.
Scenario A: Regulated healthcare, moderate volume.
A healthcare technology company processing clinical documentation. Volume around 3 million tokens per day. HIPAA applies to every request.
The residency dimension dominates. If the organisation has a BAA with AWS or Azure, a closed frontier model through Bedrock or Foundry brings the model inside an existing compliance framework with no infrastructure work, and the volume is below the self-hosting break-even. If the requirement is that no third party processes PHI at all, open-weight self-hosting on single-tenant infrastructure is the only path, and the economics become secondary.
Answer: closed model through a cloud marketplace channel if the BAA path is available, self-hosted open-weight on dedicated single-tenant capacity if it is not.
Scenario B: Consumer product, high volume, no regulated data.
A consumer application serving 50 million tokens per day of summarisation and conversational traffic. No regulated data. Latency matters for user experience.
Volume is well above the break-even, the task profile is squarely in the range where open-weight models match closed frontier quality, and there is no residency constraint forcing a particular path.
Answer: managed open-weight models for the bulk of traffic, with closed frontier reserved for the subset of requests that genuinely need it. Self-hosting becomes worth evaluating if a single model dominates the traffic and the volume grows further.
Scenario C: Enterprise internal tooling, low volume, no MLOps team.
A 500-person company adding AI features to internal tools. 5,000 requests per day. No regulated data. Engineering team has no GPU infrastructure experience.
Volume is far below the break-even, operational capacity is the binding constraint, and speed to first shipped feature matters more than per-token cost.
Answer: managed APIs, closed or open-weight depending on the task profile. Revisit when volume grows by an order of magnitude or a regulated data requirement appears.
The hybrid architecture requires access to both sides of the decision through a consistent interface, plus infrastructure for the workloads that need self-hosting.
Unified access to both categories. GMI Cloud’s model library provides 170-plus models, open-weight and closed, through a single OpenAI-compatible API. For a hybrid architecture this matters more than it appears: the alternative is maintaining separate provider integrations with different SDKs, different parameter handling, and different error semantics for each tier.
Managed open-weight inference. For the majority of production traffic that open-weight models handle well, managed inference removes the operational burden entirely while capturing the cost advantage. This is the tier that eliminates the $125,000 to $190,000 annual operational floor from the calculation for most workloads.
Dedicated capacity for self-hosted requirements. For data residency, fine-tuned weights, or sustained volume above the break-even, Prime Inference provides reserved single-tenant GPU capacity with pre-configured serving frameworks, per-model runtime tuning, and region-pinned endpoints. This covers the self-hosting requirements without the datacentre operations burden that pure self-hosting carries.
Measuring the break-even on your own workload. The published break-even ranges are directional. The number that matters is computed from your token volume, your input-output ratio, and your achievable utilisation. GMI Cloud’s on-demand infrastructure bills hourly with no minimum commitment, which makes it practical to measure actual throughput and cost per token on your workload before committing to a deployment model.
The open versus closed decision stopped being a capability decision when the benchmark gap closed to effectively zero on knowledge tasks and single digits on reasoning. What remains is a constraint-matching exercise across five dimensions, and two of them dominate in practice.
Data residency is the constraint that most often decides the question outright. When regulation requires that no third party process the data, open-weight self-hosting is the only compliant path and the economics become secondary. When the requirement is jurisdictional rather than absolute, closed models through cloud marketplace channels frequently satisfy it with no infrastructure work.
Volume is the constraint that decides the rest. Below roughly 2 million tokens per day, the $125,000 to $190,000 annual operational floor for self-hosting rarely pays back against managed alternatives. Above 5 to 10 million tokens per day against premium API pricing, it does.
The answer most production teams reach is neither pole. Closed frontier models for the workloads that need frontier capability, managed open-weight models for the 60 to 80 percent of traffic that does not, and self-hosted deployment for the specific workloads where residency or volume demands it. Routing makes that hybrid practical rather than an architectural burden.
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
