Why a 24 GB GPU doesn't give you 24 GB for an LLM: CUDA context, display, Windows WDDM and runtime buffers
An RTX 3090 sold as 24 GB reports 24,126 MiB to CUDA on Linux and had 23,863 MiB (23.3 GB) free when llama.cpp loaded, and 12 GB RTX 3060s on Windows had 11,238–11,255 MiB free; this site's calculator plans for the gap by charging each GPU 0.5 GB for the CUDA context and runtime, adding 10% of weights and cache for buffers and leaving 0.5 GB free, which leaves 20.9 GB for weights and KV cache on a 24 GB card.
What CUDA reports before a model loads
llama.cpp prints two numbers per GPU. Device 0: …, VRAM: N MiB is the card's total as CUDA reports it
(totalGlobalMem, ggml-cuda.cu:339); using device CUDA0 (…) - N MiB free
is what cudaMemGetInfo says is free when the model starts loading (ggml-cuda.cu:5056,
src/
| GPU | OS | Total (CUDA) | Free at load | Short of the size it is sold as |
|---|---|---|---|---|
| RTX 3090 2026-03-11 | Linux | 24,126 MiB | 23,863 MiB (23.30 GB) | 713 MiB |
| RTX 4090 2026-04-04 | Linux | 24,077 MiB | 23,580 MiB (23.03 GB) | 996 MiB |
| RTX 5080 2026-04-01 | Linux | 15,841 MiB | 15,557 MiB (15.19 GB) | 827 MiB |
| RTX 3090 Ti 2026-03-22; the report does not say what else was on the card | Linux | 24,109 MiB | 22,346 MiB (21.82 GB) | 2,230 MiB |
| RTX 3060 12GB 2026-03-09 | Windows | 12,287 MiB | 11,238 MiB (10.97 GB) | 1,050 MiB |
| RTX 3060 12GB 2026-04-12 | Windows | 12,287 MiB | 11,245 MiB (10.98 GB) | 1,043 MiB |
| RTX 3060 12GB 2026-05-06; second GPU of two | Windows | 12,287 MiB | 11,255 MiB (10.99 GB) | 1,033 MiB |
| RTX 5060 Ti 16GB 2026-05-06; first GPU of two | Windows | 16,310 MiB | 15,173 MiB (14.82 GB) | 1,211 MiB |
| RTX 4080 Super 2026-05-14; first GPU of two | Windows | 16,375 MiB | 15,061 MiB (14.71 GB) | 1,323 MiB |
On Linux CUDA already reports a total a few hundred MiB below the size on the box (24,126 MiB for a 3090, not 24,576), and in these logs 713–996 MiB in all was unavailable. On Windows the total matches the box (12,287 MiB for a 3060) but 1,033–1,323 MiB was taken before loading, also on cards that were the second GPU. None of the reports says whether a monitor was attached or what else ran, so no figure for the desktop alone is given here.
The CUDA context
NVIDIA does not publish a fixed size for a process's CUDA context. Its documentation says the default module loading is lazy, loading
kernels only when they are first used, and that eager loading (CUDA_MODULE_LOADING=EAGER) means a “higher startup time
and GPU memory footprint” (CUDA programming guide, environment variables). vLLM measures what
a process holds outside PyTorch's allocator as non_torch_memory (the GPU's used memory minus what PyTorch reserved,
mem_utils.py:153-160); its startup logs show
0.10 GiB on an RTX 3090 and 0.25 GiB on an RTX 4090.
Windows: WDDM and shared GPU memory
- Task Manager splits GPU memory into dedicated (the card's VRAM) and shared. Shared GPU memory “represents normal system memory that can be used by either the GPU or the CPU”, and “Windows has a policy whereby the GPU is only allowed to use half of physical memory at any given instant” (Microsoft DirectX blog, GPUs in the Task Manager).
-
When a CUDA program on Windows runs out of VRAM, NVIDIA's driver can fall back (community reports; NVIDIA's page could not be fetched) to that system memory instead of failing; the NVIDIA
Control Panel's CUDA - Sysmem Fallback Policy chooses between the two (setting name as documented by the community,
nvidia
Profile , which links NVIDIA's own answer 5490).Inspector #166 - Falling back is slow: one user with two RTX 3090s on Windows Server saw 95 GB of shared memory in use and about 0.2 tokens/s, and 3–4 tokens/s after turning GPU offload off (ollama #10479; setting the fallback policy to prefer no fallback did not help there). So on Windows a model that does not fit can load and run very slowly rather than fail.
llama.cpp's own buffers
After the weights and the KV cache, llama.cpp reserves a compute buffer on every device, sized for a worst-case batch of
min(n_ctx, n_ubatch) tokens (llama-context.cpp:627-628,
671-674) and logged per device (733-737).
The defaults are -b 2048 and -ub 512 (common/-ub is the setting that changes the reserved batch (arg.cpp:1673-1677). Read at
llama.cpp 19e28a2.
- Gemma 4 31B Q5_K_M, -c 128000 -b 512 -ub 256, RTX 5060 Ti + RTX 3060: 218.04 MiB and 261.25 MiB per GPU, plus 539.04 MiB on the host.
- Qwen3.6 27B Q4_K_M, --ctx-size 131072, RTX 2080 Ti + Tesla P100: 668.03 MiB and 523.54 MiB per GPU.
With --fit (on by default) llama-server sizes context and layers so that --fit-target, 1,024 MiB per
device by default (common.h), stays free after its own measurement; the
--fit preview replays that with this site's estimate.
vLLM: --gpu-memory-utilization
gpu_memory_utilization is “the fraction of GPU memory to be used for the model executor”, default 0.92, and it is
per instance (vllm/ceil(total × utilization) of the total memory and stops with “Free memory on device … is less than desired GPU
memory utilization” if less than that is free (v1/
| GPU | Total to CUDA | Requested at 0.92 | Room for everything else |
|---|---|---|---|
| RTX 3060 12GB | 11.63 GiB | 10.70 GB | 0.93 GB |
| RTX 3090 | 23.56 GiB | 21.68 GB | 1.88 GB |
| RTX 4090 | 23.55 GiB | 21.67 GB | 1.88 GB |
| RTX 5090 | 31.39 GiB | 28.88 GB | 2.51 GB |
Totals are the ones the vLLM memory calculator uses, read from startup logs where one exists (the RTX 3090's 23.56 GiB is vLLM #15877).
What this site assumes, and where it differs
- Each card counts at the size it is sold with (24 GB for a 3090), read as GiB.
- Every GPU pays 0.5 GB for the CUDA context and runtime; 10% of weights and KV cache is added for compute buffers and fragmentation; a setting counts as a fit only with 0.5 GB left. On a 24 GB card that is 20.9 GB of weights and cache, 23.5 GB in all.
- The desktop, other programs and Windows' share are not modelled.
That 23.5 GB limit is above the free memory in every log above, by 0.20 GB to 1.68 GB. What keeps the site's fits honest is that its totals run high: on the accuracy page whole-card totals came out 5.6%–8.1% above measured use (4 runs, one RTX 4090 report). If that holds, a setting right at the limit would really use:
| Log | Site limit | Would really use | Free in the log | Fits? |
|---|---|---|---|---|
| RTX 3090, Linux | 23.5 GB | 21.75 GB–22.25 GB | 23.30 GB | yes |
| RTX 4090, Linux | 23.5 GB | 21.75 GB–22.25 GB | 23.03 GB | yes |
| RTX 5080, Linux | 15.5 GB | 14.34 GB–14.67 GB | 15.19 GB | yes |
| RTX 3090 Ti, Linux | 23.5 GB | 21.75 GB–22.25 GB | 21.82 GB | no |
| RTX 3060 12GB, Windows | 11.5 GB | 10.64 GB–10.89 GB | 10.97 GB | yes |
| RTX 3060 12GB, Windows | 11.5 GB | 10.64 GB–10.89 GB | 10.98 GB | yes |
| RTX 3060 12GB, Windows | 11.5 GB | 10.64 GB–10.89 GB | 10.99 GB | yes |
| RTX 5060 Ti 16GB, Windows | 15.5 GB | 14.34 GB–14.67 GB | 14.82 GB | yes |
| RTX 4080 Super, Windows | 15.5 GB | 14.34 GB–14.67 GB | 14.71 GB | yes |
The RTX 3090 Ti log, 2.2 GB short of its size with nothing said about what else ran, is the exception. On Windows the margin left in these logs is only 0.03 GB–0.14 GB, and one calibration report is thin evidence, so for a setting the calculator marks close to the limit, close the browser and other GPU programs first, or pick the next smaller context. The calculator shows the total; on several cards each one pays the runtime share (splitting a model across GPUs), and a quantized cache frees context room (KV cache quantization).
Questions
How much VRAM is actually usable on a 24 GB GPU?
In public llama.cpp logs on Linux, CUDA reported 23,863 MiB free on an RTX 3090 and 23,580 MiB free on an RTX 4090 at load, 713–996 MiB short of 24 GiB. The model then also needs its compute buffers. This site budgets 20.9 GB of weights and KV cache for a 24 GB card.
Does Windows use more VRAM than Linux?
In the public logs found, yes: on Windows 1,033–1,323 MiB of each card was already unavailable when llama.cpp loaded, against 713–996 MiB on Linux, including cards that were the second GPU in the machine. The logs do not say what held it, so this is a pattern across a handful of reports, not a measured Windows cost.
What happens when a model spills into shared GPU memory on Windows?
Shared GPU memory is ordinary system RAM the GPU can use, at most half of physical memory at a time (Microsoft). Reading weights from it over PCIe is much slower than VRAM: one Ollama user with two RTX 3090s on Windows saw 95 GB of shared memory in use and about 0.2 tokens/s, against 3–4 tokens/s with the GPU turned off. The NVIDIA Control Panel has a "CUDA - Sysmem Fallback Policy" setting for this; in that report, setting it to prefer no fallback did not help.
What is the compute buffer in the llama.cpp log?
Scratch memory for the forward pass, reserved per device at startup for a worst-case batch of min
What does --gpu-memory-utilization mean in vLLM?
The share of the GPU's total memory this vLLM instance takes (default 0.92). vLLM computes total × utilization at startup and refuses to start if less than that is free, so on a RTX 3090 (23.56 GiB to CUDA) everything else on the card must fit in about 1.88 GB. Inside its share, whatever the weights, activations and CUDA graphs leave becomes KV cache.
Sources read on .