Fine-Tuning VRAM Calculator
Pick a model, the method and the training settings to see the peak GPU memory split into weights, adapters, gradients, optimizer states, activations and logits, which GPUs it fits, and what to change to fit a smaller one.
14.8 GB estimated peak VRAM on one GPU
- Model weights
- 7.3 GB
- LoRA adapters
- 0.2 GB
- Gradients
- 0.2 GB
- Optimizer states
- 0.3 GB
- Activations
- 1.7 GB
- Logits and loss
- 3.4 GB
- CUDA context and buffers
- 1.8 GB
Trainable parameters: 41,943,040 (0.52% of the model).
Activations and logits: 2,634 KB per token, for 2,048 tokens per step.
Paged 8-bit states are the same size. When the GPU runs short, bitsandbytes moves them to system RAM instead of failing, which is slower.
MoE model: “all linear” adapts every expert, as PEFT does when experts are separate layers. Activations use the experts one token passes through.
Which GPUs fit
| Memory | GPUs | Result | Longest sequence that fits |
|---|---|---|---|
| 8 GB | RTX 4060 8GB | No | – |
| 12 GB | RTX 3060 12GB, Arc B580 12GB, RTX 4070 12GB, RTX 5070 12GB | No | 512 |
| 16 GB | RTX 4060 Ti 16GB, RTX 5060 Ti 16GB, RX 9070 XT 16GB, RTX 4070 Ti Super 16GB, RTX 4080 Super 16GB, RTX 5070 Ti 16GB, RTX 5080 16GB | Fits | 2,304 |
| 24 GB | RTX 3090, RTX 4090, RX 7900 XTX | Fits | 5,376 |
| 32 GB | RTX 5090 | Fits | 8,448 |
| 48 GB | L40S | Fits | 14,848 |
| 80 GB | A100 80GB, H100 SXM | Fits | 27,648 |
| 96 GB | RTX PRO 6000 Blackwell | Fits | 34,048 |
| 120 GB | DGX Spark (128 GB) | Fits | 43,520 |
| 141 GB | H200 | Fits | 51,968 |
| 180 GB | B200 | Fits | 67,584 |
“Fits” leaves at least 0.5 GB free; “Tight” fits with less, so a long sample or fragmentation can still run out. One GPU, no offloading. Macs are left out: training there runs on MLX, which this does not model.
Changes to try in order, each on top of the ones before, until the setup fits with 0.5 GB to spare:
VRAM for common setups
2,048 tokens, micro-batch 1, gradient checkpointing on, LoRA rank 16 on all linear layers, AdamW; full fine-tuning in mixed precision. Each figure opens the calculator with that setup.
| Model | QLoRA, Transformers | QLoRA, Unsloth | LoRA 16-bit | Full |
|---|---|---|---|---|
| Llama 3.2 3B | 9.2 GB | 4.0 GB | 11.5 GB | 52.8 GB |
| Qwen3 4B | 10.4 GB | 4.6 GB | 13.8 GB | 65.5 GB |
| Llama 3.1 8B | 14.8 GB | 7.8 GB | 21.3 GB | 125 GB |
| Qwen3 8B | 16.3 GB | 8.1 GB | 22.2 GB | 128 GB |
| Qwen3 14B | 21.8 GB | 12.3 GB | 35.4 GB | 227 GB |
| Gemma 3 27B | 33.1 GB | 19.1 GB | 63.4 GB | 419 GB |
| Qwen3 32B | 32.6 GB | 22.3 GB | 70.9 GB | 495 GB |
| Llama 3.3 70B | 55.6 GB | 42.8 GB | 143 GB | 1,060 GB |
Checked against published numbers
15 public measurements with their settings, next to this calculator’s estimate for the same setup. The median error is 1.6% and the largest 7.5%. The activation constants and runtime costs were fitted to the Unsloth-blog rows (both frameworks), so those match partly by construction; the three PEFT and Giles Thomas rows were not used for fitting.
| Framework | Model | Setup | Published | Estimate | Error | Source |
|---|---|---|---|---|---|---|
| PEFT | Llama 3.2 3B | LoRA r=32 on q and v, batch 4 × 768 tokens, no checkpointing Assumed: Peak taken at a full 4 × 768 batch | 20.8 GB | 22.0 GB | +5.9% | github.com 2026-07-15 |
| PEFT | Llama 3.2 3B | Full fine-tune in BF16, batch 4 × 768 tokens, no checkpointing Assumed: Peak taken at a full 4 × 768 batch | 36.2 GB | 37.7 GB | +4.2% | github.com 2026-07-15 |
| HF Trainer + DeepSpeed | Qwen1.5 0.5B | Full fine-tune, mixed precision, 2,048 tokens, no checkpointing | 13.9 GB | 12.9 GB | −7.5% | gilesthomas.com 2024-07 |
| HF + FA2 | Llama 3.1 8B | Longest sequence on a 12 GB card: 932 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 12 GB | 12.2 GB | +1.6% | unsloth.ai 2024-12-10 |
| HF + FA2 | Llama 3.1 8B | Longest sequence on a 24 GB card: 5,789 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 24 GB | 24.4 GB | +1.6% | unsloth.ai 2024-12-10 |
| HF + FA2 | Llama 3.1 8B | Longest sequence on an 80 GB card: 28,454 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 80 GB | 81.3 GB | +1.6% | unsloth.ai 2024-12-10 |
| HF + FA2 | Llama 3.3 70B | Longest sequence on an 80 GB card: 6,916 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 80 GB | 76.2 GB | −4.8% | unsloth.ai 2024-12-10 |
| HF + FA2 | Mistral 7B v0.2 | Longest sequence on an 8 GB card: 1,696 tokens | 8 GB | 8.31 GB | +3.8% | unsloth.ai 2024-04-09 |
| HF + FA2 | Mistral 7B v0.2 | Longest sequence on a 24 GB card: 14,099 tokens | 24 GB | 23.9 GB | −0.5% | unsloth.ai 2024-04-09 |
| HF + FA2 | Mistral 7B v0.2 | Longest sequence on an 80 GB card: 57,510 tokens | 80 GB | 78.4 GB | −2.0% | unsloth.ai 2024-04-09 |
| Unsloth | Llama 3.1 8B | Longest sequence on an 8 GB card: 2,972 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 8 GB | 8.14 GB | +1.7% | unsloth.ai 2024-12-10 |
| Unsloth | Llama 3.1 8B | Longest sequence on a 24 GB card: 78,475 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 24 GB | 24.0 GB | +0.1% | unsloth.ai 2024-12-10 |
| Unsloth | Llama 3.1 8B | Longest sequence on an 80 GB card: 342,733 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 80 GB | 79.7 GB | −0.4% | unsloth.ai 2024-12-10 |
| Unsloth | Llama 3.3 70B | Longest sequence on a 48 GB card: 12,106 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 48 GB | 47.8 GB | −0.4% | unsloth.ai 2024-12-10 |
| Unsloth | Llama 3.3 70B | Longest sequence on an 80 GB card: 89,389 tokens Assumed: Optimizer not stated; 8-bit AdamW assumed (Unsloth notebooks) | 80 GB | 80.1 GB | +0.1% | unsloth.ai 2024-12-10 |
“Published” is the card size for the longest-sequence rows and PyTorch’s peak for the others; the estimate is compared like for like (without the 0.5 GB CUDA context for PyTorch peaks).
What is not checked
We found no public measurement with its settings for LoRA on a 16-bit base with gradient checkpointing, Unsloth without gradient checkpointing, MoE models, or micro-batches above 4, so those estimates combine checked parts but are not checked themselves. Unsloth’s longest-sequence tables grow exactly linearly with the card size, so Unsloth partly extrapolated them. Full fine-tuning in mixed precision has one check, 8% low, because DeepSpeed keeps an extra FP16 copy of the weights.
Next to rule-of-thumb tables
Unsloth’s requirements page lists an “absolute minimum” per model size without the sequence length, rank or batch. It is guidance, not a measurement, so it is not in the error figures. The estimate column is Unsloth at 512 tokens, rank 16, batch 1.
| Model | QLoRA, Unsloth table | QLoRA, estimate | LoRA 16-bit, Unsloth table | LoRA 16-bit, estimate |
|---|---|---|---|---|
| Llama 3.1 8B | 6 GB | 7.5 GB | 22 GB | 16.9 GB |
| Qwen3 14B | 8.5 GB | 11.9 GB | 33 GB | 29.8 GB |
| Qwen3 32B | 26 GB | 21.8 GB | 76 GB | 64.4 GB |
| Llama 3.3 70B | 41 GB | 42.2 GB | 164 GB | 136 GB |
Unsloth’s own benchmark implies more than its table for 8B: at rank 32 it fits 2,972 tokens on an 8 GB card, and its figures grow in a straight line that puts a very short sample at about 7.4 GB. LLaMA-Factory’s table (marked “estimated”) gives a 7B model 120 GB for full fine-tuning in 32 bits, 60 GB in pure 16-bit, 16 GB for LoRA and 6 GB for 4-bit QLoRA; for Llama 3.1 8B at 2,048 tokens with Transformers this calculator gives 125 GB, 65.2 GB, 21.3 GB and 14.8 GB.
How it is calculated
- Weights. LoRA: 2 bytes per parameter. QLoRA: every linear layer except the output layer in bitsandbytes NF4, 4 bits plus an 8-bit scale per 64 weights and an FP32 scale per 256 blocks with double quantization (4.127 bits; 4.5 without); embeddings, output layer and norms stay BF16 in Unsloth and become FP32 in Transformers, because
prepare_model_for_kbit_trainingcasts them. - LoRA parameters. r × (din + dout) for each adapted matrix in each layer, from the config’s hidden size, head counts and feed-forward width (all experts in an MoE layer). Llama 3.1 8B at r = 16 on all linear layers: 41,943,040. Adapters and their gradients are FP32 (4 + 4 bytes); AdamW adds 8 bytes, 8-bit AdamW 2.
- Full fine-tuning. Mixed precision: 16 bytes per parameter with AdamW (BF16 weights 2, FP32 master copy 4, gradients 2, m and v 8, as counted in the ZeRO paper), 10 with 8-bit AdamW. Pure BF16: 8 (PyTorch’s AdamW keeps m and v in the weights’ dtype), 6 with 8-bit AdamW.
- Activations. A layer saves about 10h + 2q + 4kv + 8f bytes per token in BF16 with flash attention (h hidden size; q and kv the query and key/value widths; f the feed-forward width, for MoE the experts a token uses), plus a BF16 copy of each adapted layer’s input. Without checkpointing every layer keeps this. Transformers’ checkpointing keeps each layer’s input (2h bytes) plus 2.6 times one layer while it is recomputed; Unsloth’s offloaded checkpointing keeps 0.92 times one layer on the GPU and the inputs in system RAM.
- Logits. Transformers computes the loss on FP32 logits: about 14 bytes per vocabulary entry per token, 1.8 MB for Llama 3’s 128,256 entries. Unsloth’s fused cross-entropy (Apple’s Cut Cross Entropy) never builds them.
- Runtime. 0.6 GB (Transformers) or 1.2 GB (Unsloth) for the CUDA context, cuBLAS workspace and allocator. QLoRA adds a dequantized copy of the largest matrix, and Transformers QLoRA a BF16 copy of the FP32 output layer.
- Not modelled. Several GPUs (FSDP, DeepSpeed ZeRO), CPU offload of weights or optimizer, evaluation or generation during training, and DPO or GRPO (reference model, rollouts).
Sources: bitsandbytes and the QLoRA paper (NF4, double quantization), PEFT’s prepare_model_for_kbit_training, the ZeRO paper (16 bytes per parameter), Korthikanti et al. (activation memory), Unsloth on offloaded checkpointing and Cut Cross Entropy.
Sources read on .
To run the model after training, see the LLM VRAM calculator. How the site’s inference estimates compare with real runs: predicted vs measured.
How to use
- Pick a model from the list or load any Hugging Face model by its id.
- Choose full fine-tuning, LoRA on the 16-bit model, or QLoRA on a 4-bit base, and the framework: Transformers with PEFT (and TRL), or Unsloth.
- Set the LoRA rank and target modules, the sequence length, the micro-batch, gradient checkpointing and the optimizer.
- Read the total and the parts, check the GPU table, and pick a GPU under “Make it fit” to see which changes get the setup under its memory.
Frequently asked questions
How much VRAM do I need to fine-tune an 8B model with QLoRA?
About 7.8 GB with Unsloth and 14.8 GB with Transformers + PEFT, for Llama 3.1 8B at 2,048 tokens, micro-batch 1, rank 16 on all linear layers and gradient checkpointing on (9.1 GB and 30.3 GB at 8,192 tokens). Most of the gap is the loss: Transformers builds FP32 logits over the 128,256-entry vocabulary, about 1.8 MB per token, and casts the embedding and output layer to FP32, while Unsloth fuses the loss and keeps them in BF16. Unsloth measured 2,972 tokens on an 8 GB card at rank 32 where HF + FA2 ran out of memory; our estimates for those runs are within 2%.
How much more memory does LoRA or full fine-tuning take than QLoRA?
For Llama 3.1 8B at 2,048 tokens with Transformers: QLoRA 14.8 GB, LoRA on the 16-bit model 21.3 GB, and full fine-tuning 125 GB in mixed precision with AdamW (50.2 GB in pure BF16 with 8-bit AdamW). Full fine-tuning pays 16 bytes per parameter for the weights, gradients and optimizer states before any activations, so an 8B model needs several 80 GB GPUs or offloading.
How much does gradient checkpointing save?
For LoRA on Llama 3.1 8B at 4,096 tokens with Transformers, 53.0 GB without it and 26.5 GB with it. Without checkpointing every layer keeps its activations for the backward pass; with it, only each layer's input is kept and one layer is recomputed at a time, which costs about 20–30% more time per step. Unsloth's version also moves the layer inputs to system RAM.
Can I fine-tune a 70B model on one GPU?
With QLoRA, yes: about 42.8 GB for Llama 3.3 70B at 2,048 tokens with Unsloth, so a 48 GB card fits (Unsloth measured 12,106 tokens on 48 GB), and 55.6 GB with Transformers, which needs an 80 GB card. LoRA on the 16-bit model or full fine-tuning needs several GPUs with FSDP or DeepSpeed, which this calculator does not model.
How accurate is this?
Across 15 published measurements the median error is 1.6% and the largest 7.5%; the table on this page lists each with its settings and source. Most of them were also used to fit the activation constants, and some setups (LoRA on a 16-bit base with checkpointing, MoE models) have no public measurement to check against, as the page says.
Is GB here GB or GiB?
GiB (1024³ bytes), the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The published measurements are shown in the same unit.
More calculators
- Image & Video Model VRAM Calculator Peak VRAM and system RAM for FLUX.2, Wan 2.2, LTX-2 and Qwen-Image in ComfyUI, from exact file sizes, checked against public runs. Open →
- Jev Alternatives You Can Run Locally Open models that replace the API-only Jev on your own machine: Laya, zero-shot encoders and small LLMs, with sizes and VRAM. Open →
- Laya VRAM Requirements Real file sizes and run-time memory of the three Laya decision-encoder checkpoints, and three ways to run them. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open →
Updated