• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Engineering

    Fine-Tuning Infrastructure for LLMs: LoRA, QLoRA, and Full Fine-Tuning GPU Requirements

    A practical guide to fine-tuning LLMs with LoRA, QLoRA, and full fine-tuning. Learn how GPU memory requirements differ, what impacts VRAM usage, when each method makes sense, and how to choose the right infrastructure for efficient training and deployment.

    September 30, 2026

    The memory arithmetic for fine-tuning does not resemble the arithmetic for inference, and teams that size a training run using inference intuition get the answer wrong by an order of magnitude. Serving a 70B model at FP8 needs roughly 70 GB of weights plus KV cache. Full fine-tuning the same model needs 560 to 640 GB, because gradients and optimizer states dominate the budget rather than weights. Adam stores two FP32 running statistics per trainable parameter regardless of the training precision, which means 560 GB of momentum buffers alone before weights, gradients, or activations are counted. That single fact is what pushes full fine-tuning of serious models into multi-node cluster territory and what makes LoRA and QLoRA the pragmatic default for almost every production use case.

    • Optimizer state is the memory component most teams omit. Adam and AdamW store a first and second moment per trainable parameter at FP32, giving trainable_params × 4 bytes × 2. This is also why moving from LoRA to QLoRA does not reduce optimizer memory: both train the same small adapter parameter count.

    • GMI Cloud provides the GPU capacity for all three regimes, from a single H100 for QLoRA on a 70B model to multi-node clusters with 3.2 Tbps InfiniBand for full fine-tuning runs that require it.

    • At QLoRA with realistic batch and sequence settings, activations dominate, not weights. A Llama 3.1 8B QLoRA run at rank 16, batch 4, sequence 2048 uses 3.81 GB for weights and 16.00 GB for activations out of 23.87 GB total. Sizing on weights alone produces an out-of-memory error at the first training step.

    • The quality ladder is roughly 100 / 90 to 95 / 80 to 90. Full fine-tuning sets the reference, LoRA recovers 90 to 95 percent of it, and QLoRA lands at 80 to 90 percent. For most production domain adaptation, the QLoRA gap is not detectable in the final application.

    • The decision that precedes the infrastructure question is whether to fine-tune at all. Fine-tuning teaches style, format, and task behaviour. It does not reliably teach new facts. A team fine-tuning to inject knowledge that changes weekly has chosen the wrong tool and will pay for it twice.

    • Deployment has two paths with different operational profiles. Merging the adapter into base weights gives zero inference latency overhead and one artefact to serve. Keeping adapters separate allows many fine-tuned variants to share one base model in memory, which is decisive for multi-tenant deployments.

    The Four Memory Components

    Fine-tuning memory splits into four categories, and the ratio between them changes completely depending on which method you use.

    Weights. The base model parameters. At BF16, a 70B model is 140 GB. At 4-bit NF4 quantization, roughly 35 to 40 GB. This is the only component that inference also pays.

    Gradients. One gradient value per trainable parameter, typically at the training dtype. In full fine-tuning every parameter is trainable, so gradients match the weight footprint. In LoRA, only the adapter parameters are trainable, so gradients drop to a fraction of a percent of the model size.

    Optimizer states. This is the component teams consistently underestimate. Adam and AdamW store two FP32 running statistics per trainable parameter: a first moment tracking the mean of gradients, and a second moment tracking the mean of squared gradients. Both are stored at FP32 regardless of whether training runs in BF16 or FP16.

    The formula is trainable_params × 4 bytes × 2 moments. For a 70B model with every parameter trainable: 70 billion × 4 × 2 equals 560 GB, for momentum buffers alone.

    This also explains a property that surprises people: switching from LoRA to QLoRA does not change optimizer memory at all. Both methods train the same small set of adapter parameters, so the optimizer state is identical. QLoRA saves memory on the frozen base weights, not on the optimizer.

    Activations. Intermediate values retained during the forward pass for use in the backward pass. Activation memory scales with batch size and sequence length rather than with model size, which makes it the component that varies most between configurations of the same job.

    Gradient checkpointing trades compute for activation memory by recomputing activations during the backward pass instead of storing them. It typically reduces activation memory substantially at a cost of roughly 20 to 30 percent additional compute time, and it is enabled by default in most production fine-tuning configurations.

    The Three Methods and What They Actually Cost

    Full fine-tuning updates every parameter. Maximum quality, maximum memory, and the only method that can meaningfully change the model’s underlying capabilities rather than its behaviour.

    LoRA freezes the base model and trains small low-rank matrices injected into the attention and feed-forward layers. Trainable parameters drop to roughly 0.1 to 1 percent of the total. The base model still sits in memory at full precision, so weight memory is unchanged, but gradients and optimizer states collapse.

    QLoRA quantizes the frozen base model to 4-bit NF4 and trains LoRA adapters on top of it. NF4 is information-theoretically optimal for normally distributed weights, which is what most neural network layers produce. Two additional techniques make it practical: double quantization, which quantizes the quantization constants themselves, and paged optimizers, which offload optimizer states to CPU RAM in pages when GPU memory tightens.

    The comparison on Llama 3.1 8B:

    Method

    Total VRAM

    Weight memory

    Trainable parameters

    Full fine-tuning

    117.3 GB

    15.0 GB

    100%

    LoRA

    27.6 GB

    15.0 GB

    0.08%

    QLoRA 4-bit

    23.9 GB

    3.8 GB

    0.08%

    Two things stand out. Full fine-tuning an 8B model needs more VRAM than a single H100 provides, which is why even small-model full fine-tuning is a multi-GPU job. And the gap between LoRA and QLoRA on total VRAM is smaller than the gap on weight memory, because the components QLoRA does not reduce, activations and optimizer state, make up the majority of a LoRA run.

    The 70B picture:

    Method

    VRAM required

    Practical hardware

    Full fine-tuning

    560 to 640 GB

    8 GPUs minimum

    LoRA at FP16

    140 to 180 GB

    2 GPUs (192 GB combined)

    QLoRA 4-bit

    ~46 GB

    Single H100 80GB or H200

    The 70B QLoRA number is the one that changed the field. A model that requires 672 GB for full-precision full fine-tuning drops to roughly 46 GB, which fits comfortably on a single H100 with headroom for reasonable batch sizes and sequence lengths.

    The Activation Memory Trap

    The single most common sizing error in QLoRA runs comes from budgeting for weights and forgetting activations.

    A concrete breakdown from a Llama 3.1 8B QLoRA run at rank 16, 4-bit quantization, batch size 4, sequence length 2048:

    Component

    Memory

    Model weights

    3.81 GB

    LoRA adapters

    0.01 GB

    Gradients

    0.01 GB

    Optimizer states

    0.05 GB

    Activations

    16.00 GB

    CUDA overhead

    3.98 GB

    Total

    23.87 GB

    Activations consume 67 percent of the total. Weights consume 16 percent. A team sizing this job from the quantized weight footprint would provision an 8 GB GPU and hit an out-of-memory error at the first step.

    What drives activation memory.

    Sequence length is the strongest driver. Doubling from 512 to 1024 tokens increases memory roughly 1.5 to 2 times. Training on long documents or long conversation histories can push activation memory above everything else combined.

    Batch size increases activation memory proportionally, but does not increase weights, gradients, or optimizer states, which are fixed. This is why out-of-memory errors during fine-tuning are usually solved by reducing batch size rather than by changing method.

    Gradient accumulation is the standard workaround. Rather than processing a large batch in one step, process several small batches and accumulate gradients before applying the optimizer update. This achieves the convergence behaviour of a large effective batch size at the activation memory cost of a small one, at the cost of more steps to reach the same number of samples processed.

    Paged optimizers for the marginal case. Setting optim="paged_adamw_32bit" offloads optimizer states to CPU RAM in pages when GPU memory tightens. It adds CPU-to-GPU transfer overhead but can recover 2 to 6 GB on smaller models where optimizer state is a meaningful fraction of the total. On large models with LoRA, optimizer state is already negligible, so this setting has little effect.

    The Quality Tradeoff, Honestly

    Published figures for the quality ladder are reasonably consistent across sources, with one area of genuine disagreement worth flagging.

    LoRA recovers 90 to 95 percent of full fine-tuning quality on most tasks. The gap narrows as the LoRA rank increases, at the cost of more trainable parameters and correspondingly more gradient and optimizer memory. Rank 16 is a common production default; rank 32 or 64 narrows the gap further where quality matters more than memory.

    QLoRA typically lands at 80 to 90 percent of full fine-tuning performance in the comparisons that report a gap. The original QLoRA research reported near state-of-the-art results on several tasks, arguing that the quality impact is minimal because quantization applies primarily to the forward pass while gradients accumulate in higher precision.

    Both characterisations can be true depending on the task. The practical resolution is that the gap is task-dependent and should be measured rather than assumed. For domain adaptation where the goal is teaching a consistent output format or a domain vocabulary, QLoRA is usually indistinguishable in the final application. For tasks requiring precise numerical reasoning or where small errors compound, the gap is more likely to be measurable.

    The measurement that settles it for your case. Fine-tune with QLoRA, evaluate on a held-out set from your actual task distribution, and compare against the base model and against a LoRA run if the budget allows. If QLoRA closes most of the gap between base and target performance, the additional cost of LoRA at full precision is not justified.

    The Decision That Comes Before the Infrastructure

    The most expensive fine-tuning mistake is not a sizing error. It is fine-tuning when a different technique was the right answer.

    What fine-tuning reliably teaches. Output format and structure. Domain vocabulary and phrasing conventions. Task-specific behaviour patterns such as how to respond to a particular category of request. Tone and register. Implicit rules that are tedious to express in a prompt.

    What fine-tuning does not reliably teach. New factual knowledge, especially knowledge that changes. A model fine-tuned on a product catalogue produces confident answers about products that were discontinued after the training cut-off. Retrieval is the correct mechanism for facts, and fine-tuning is the correct mechanism for behaviour.

    The three-way decision.

    Prompting handles the case where the desired behaviour can be described in instructions and few-shot examples, and where the instruction overhead per request is acceptable. This covers more cases than teams expect, and it should be exhausted before fine-tuning is considered.

    Retrieval handles the case where the model needs access to information it does not have: current facts, proprietary documents, user-specific context. Retrieval-augmented generation keeps the knowledge outside the model where it can be updated without retraining.

    Fine-tuning handles the case where the behaviour is hard to specify in a prompt, where the prompt overhead is prohibitive at volume, or where consistency across many requests matters more than flexibility. A model that has learned the output format natively does not need 2,000 tokens of format instruction on every request, which is both a cost saving and a reliability improvement.

    The hybrid that works. Fine-tune for format and behaviour, retrieve for facts. A fine-tuned model that reliably produces the correct output structure, combined with retrieval that supplies current information, outperforms either technique alone on most production tasks.

    Data Requirements

    The question of how many examples are needed has a consistent practical answer that is lower than most teams expect.

    For format and style adaptation, 500 to 2,000 well-constructed examples frequently suffice. The model already knows how to write; it is learning which of its existing behaviours to apply.

    For task-specific behaviour, 1,000 to 10,000 examples depending on task complexity and the variety of inputs the model must handle.

    For domain adaptation on specialised vocabulary, larger datasets help, but the returns diminish. A model that has seen 50,000 examples of legal drafting is not markedly better than one that has seen 10,000 well-chosen ones.

    Quality dominates quantity. A thousand carefully constructed examples that cover the input distribution, including edge cases, outperform ten thousand scraped examples with inconsistent output quality. The model learns the patterns present in the data, including the mistakes.

    Safe starting hyperparameters for a LoRA or QLoRA run: 1 to 3 epochs, learning rate between 1e-4 and 2e-4, rank 16, and gradient checkpointing enabled. More epochs on a small dataset produces memorisation rather than generalisation, and the resulting model performs worse on inputs that differ from the training examples.

    Deployment: Merged Weights Versus Adapter Serving

    A completed fine-tune produces an adapter, and there are two ways to serve it with different operational consequences.

    Merged weights. The adapter is mathematically merged into the base model weights, producing a single model artefact. Inference runs exactly as it would on the base model with no additional computation, which means zero inference latency overhead.

    The cost is that each fine-tuned variant becomes a full model to store and load. Serving five fine-tuned variants of a 70B model means five 70 GB artefacts and five model instances if they run concurrently.

    This is the right choice for a single fine-tuned model serving all traffic, or for export to formats such as GGUF for local deployment.

    Adapter serving. The base model is loaded once, and adapters are applied at inference time. vLLM and Hugging Face TGI both support this pattern, including serving multiple adapters against a single base model instance.

    The cost is a small inference overhead from applying the adapter computation, and additional serving complexity. The benefit is decisive for multi-tenant deployments: one 70 GB base model in memory can serve dozens of customer-specific adapters, each a few hundred megabytes, rather than requiring a full model instance per customer.

    The decision rule. One fine-tuned model serving all traffic: merge. Many fine-tuned variants sharing a base model, particularly per-customer or per-domain adapters: serve adapters separately and accept the modest overhead.

    For teams deploying merged fine-tuned weights, GMI Prime Inference supports loading custom weights on dedicated GPU capacity with the serving stack pre-configured, which means the fine-tuned artefact deploys the same way a standard open-weight model does.

    Cost: What a Fine-Tuning Run Actually Costs

    The compute cost of fine-tuning is usually smaller than teams expect, and the surrounding costs are larger.

    A worked QLoRA example. Fine-tuning a 70B model with QLoRA on a single H100, on 5,000 examples at sequence length 2048, for 3 epochs. At typical throughput this runs in roughly 8 to 16 hours depending on hardware and configuration. On GMI Cloud H100 capacity, that is a single-digit to low double-digit dollar figure for the compute.

    A worked LoRA example. The same job at FP16 LoRA requires two GPUs with 192 GB combined for a 70B model, and runs for a similar duration. Roughly double the compute cost of the QLoRA run, still a modest absolute figure.

    A full fine-tuning example. The same model at full fine-tuning requires 8 GPUs minimum and substantially longer wall-clock time. This is where the cost becomes material, and it is the reason full fine-tuning is rarely the right choice for domain adaptation.

    Where the cost actually accumulates. Dataset construction and cleaning is typically the largest line item, measured in engineering and domain-expert time rather than GPU hours. Evaluation infrastructure to determine whether the fine-tune improved anything is the second. Iteration is the third: the first fine-tune is rarely the one that ships, and each iteration repeats the compute cost.

    Budgeting a fine-tuning project at the GPU cost of a single run understates it by a large multiple. Budgeting at ten runs plus the data and evaluation work is closer to reality.

    Infrastructure Selection by Method and Model Size

    The GPU requirement follows directly from the memory arithmetic above.

    QLoRA on models up to 70B: a single H100 or H200. The 70B QLoRA footprint of roughly 46 GB fits on an 80 GB H100 with headroom for batch size and sequence length. H200’s 141 GB provides substantially more headroom, which translates to larger batches and longer sequences without gradient accumulation.

    LoRA at FP16 on 70B: two GPUs minimum. The 140 to 180 GB requirement needs 192 GB of combined VRAM, which is two H100s or two H200s. Two H200s at 282 GB combined provide comfortable headroom.

    Full fine-tuning on 70B: 8 GPUs and inter-node bandwidth. At 560 to 640 GB, this is an 8-GPU H100 node at minimum. Full fine-tuning also makes inter-GPU communication a bottleneck rather than a detail, because gradients must be synchronised across all GPUs at every step.

    This is where the network fabric matters. As covered in GMI Cloud’s analysis of GPU cloud for LLM training, 3.2 Tbps InfiniBand keeps communication overhead in the 5 to 10 percent range, against 30 to 50 percent on standard Ethernet. For a full fine-tuning run measured in days, that difference determines whether the job completes in three days or five.

    The MoE caveat. Mixture-of-Experts architectures break the standard formulas. The full expert weight matrix must be present regardless of how few experts activate per token, and the expert routing adds communication patterns that dense models do not have. Llama 4’s Scout and Maverick variants, Qwen3-235B, and similar MoE models need architecture-specific sizing rather than the parameter-count formulas above.

    Running Fine-Tuning Workloads on GMI Cloud

    Three infrastructure properties matter for fine-tuning specifically, and they differ from the properties that matter for inference.

    Bare metal without hypervisor overhead. Training runs are long and compute-intensive, which means the 10 to 15 percent hypervisor overhead on virtualised instances translates directly into a proportionally longer run. On a three-day full fine-tune, that is most of a day.

    High-bandwidth interconnect for multi-GPU jobs. LoRA on two GPUs and full fine-tuning on eight both require gradient synchronisation between GPUs at every step. NVLink within a node and InfiniBand between nodes are what keep that synchronisation from dominating the run time.

    Hourly billing without commitment. Fine-tuning is bursty by nature: an intensive run, then a gap while the results are evaluated, then another run. Reserved capacity sized for the training peak sits idle between runs. GMI Cloud’s on-demand infrastructure bills hourly with no minimum, which matches the actual usage shape of a fine-tuning project.

    And then serving the result. A completed fine-tune needs somewhere to run. Prime Inference supports custom weight loading on dedicated capacity with the serving stack pre-configured, which means the path from a finished training run to a production endpoint does not require assembling a separate serving environment.

    Conclusion

    Fine-tuning memory is dominated by components that inference does not pay. Optimizer state at FP32, two moments per trainable parameter, is 560 GB for a 70B full fine-tune before anything else is counted, and it is the number that pushes full fine-tuning into multi-node territory for any serious model.

    LoRA and QLoRA collapse that number by reducing trainable parameters to a fraction of a percent, which is why a 70B model that needs 672 GB for full-precision full fine-tuning runs in roughly 46 GB with QLoRA on a single H100. The quality cost is 5 to 10 percentage points for LoRA and 10 to 20 for QLoRA against the full fine-tuning reference, and for most production domain adaptation that gap is not detectable in the application.

    Two errors account for most failed fine-tuning projects. Sizing on weight memory and forgetting activations, which at QLoRA can be 67 percent of the total. And fine-tuning to inject facts, which is what retrieval is for. Fine-tuning teaches behaviour, format, and style reliably. It teaches facts unreliably and expensively.

    Run fine-tuning and serving on GMI Cloud

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    Because three of the four memory components do not exist during inference. Inference pays for weights and KV cache. Training additionally pays for gradients, one value per trainable parameter, and optimizer states, which for Adam and AdamW are two FP32 running statistics per trainable parameter regardless of training precision. The formula is trainable_params × 4 bytes × 2 moments, which for a 70B model with every parameter trainable is 560 GB for momentum buffers alone. Adding weights, gradients, and activations brings the total to roughly 560 to 640 GB, against 70 GB to serve the same model at FP8.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started