Why a 24 GB GPU doesn't give you 24 GB for an LLM: CUDA context, display, Windows WDDM and runtime buffers

An RTX 3090 sold as 24 GB reports 24,126 MiB to CUDA on Linux and had 23,863 MiB (23.3 GB) free when llama.cpp loaded, and 12 GB RTX 3060s on Windows had 11,238–11,255 MiB free; this site's calculator plans for the gap by charging each GPU 0.5 GB for the CUDA context and runtime, adding 10% of weights and cache for buffers and leaving 0.5 GB free, which leaves 20.9 GB for weights and KV cache on a 24 GB card.

What CUDA reports before a model loads

llama.cpp prints two numbers per GPU. Device 0: …, VRAM: N MiB is the card's total as CUDA reports it (totalGlobalMem, ggml-cuda.cu:339); using device CUDA0 (…) - N MiB free is what cudaMemGetInfo says is free when the model starts loading (ggml-cuda.cu:5056, src/llama.cpp:303-308). The difference is whatever the driver, the desktop and other programs already hold on that card. From public reports:

GPUOSTotal (CUDA)Free at loadShort of the size it is sold as
RTX 3090 2026-03-11 Linux 24,126 MiB 23,863 MiB (23.30 GB) 713 MiB
RTX 4090 2026-04-04 Linux 24,077 MiB 23,580 MiB (23.03 GB) 996 MiB
RTX 5080 2026-04-01 Linux 15,841 MiB 15,557 MiB (15.19 GB) 827 MiB
RTX 3090 Ti 2026-03-22; the report does not say what else was on the card Linux 24,109 MiB 22,346 MiB (21.82 GB) 2,230 MiB
RTX 3060 12GB 2026-03-09 Windows 12,287 MiB 11,238 MiB (10.97 GB) 1,050 MiB
RTX 3060 12GB 2026-04-12 Windows 12,287 MiB 11,245 MiB (10.98 GB) 1,043 MiB
RTX 3060 12GB 2026-05-06; second GPU of two Windows 12,287 MiB 11,255 MiB (10.99 GB) 1,033 MiB
RTX 5060 Ti 16GB 2026-05-06; first GPU of two Windows 16,310 MiB 15,173 MiB (14.82 GB) 1,211 MiB
RTX 4080 Super 2026-05-14; first GPU of two Windows 16,375 MiB 15,061 MiB (14.71 GB) 1,323 MiB

On Linux CUDA already reports a total a few hundred MiB below the size on the box (24,126 MiB for a 3090, not 24,576), and in these logs 713–996 MiB in all was unavailable. On Windows the total matches the box (12,287 MiB for a 3060) but 1,033–1,323 MiB was taken before loading, also on cards that were the second GPU. None of the reports says whether a monitor was attached or what else ran, so no figure for the desktop alone is given here.

The CUDA context

NVIDIA does not publish a fixed size for a process's CUDA context. Its documentation says the default module loading is lazy, loading kernels only when they are first used, and that eager loading (CUDA_MODULE_LOADING=EAGER) means a “higher startup time and GPU memory footprint” (CUDA programming guide, environment variables). vLLM measures what a process holds outside PyTorch's allocator as non_torch_memory (the GPU's used memory minus what PyTorch reserved, mem_utils.py:153-160); its startup logs show 0.10 GiB on an RTX 3090 and 0.25 GiB on an RTX 4090.

Windows: WDDM and shared GPU memory

llama.cpp's own buffers

After the weights and the KV cache, llama.cpp reserves a compute buffer on every device, sized for a worst-case batch of min(n_ctx, n_ubatch) tokens (llama-context.cpp:627-628, 671-674) and logged per device (733-737). The defaults are -b 2048 and -ub 512 (common/common.h:453-454); -ub is the setting that changes the reserved batch (arg.cpp:1673-1677). Read at llama.cpp 19e28a2.

With --fit (on by default) llama-server sizes context and layers so that --fit-target, 1,024 MiB per device by default (common.h), stays free after its own measurement; the --fit preview replays that with this site's estimate.

vLLM: --gpu-memory-utilization

gpu_memory_utilization is “the fraction of GPU memory to be used for the model executor”, default 0.92, and it is per instance (vllm/config/cache.py:103-111). At startup vLLM requests ceil(total × utilization) of the total memory and stops with “Free memory on device … is less than desired GPU memory utilization” if less than that is free (v1/worker/utils.py:539-558). Of the requested share, what the weights, the activation peak, non-torch memory and the CUDA graph estimate leave is the KV cache (gpu_worker.py:673-695). Read at ac7f3e1 (v0.30.0 logic).

GPUTotal to CUDARequested at 0.92Room for everything else
RTX 3060 12GB11.63 GiB10.70 GB0.93 GB
RTX 309023.56 GiB21.68 GB1.88 GB
RTX 409023.55 GiB21.67 GB1.88 GB
RTX 509031.39 GiB28.88 GB2.51 GB

Totals are the ones the vLLM memory calculator uses, read from startup logs where one exists (the RTX 3090's 23.56 GiB is vLLM #15877).

What this site assumes, and where it differs

That 23.5 GB limit is above the free memory in every log above, by 0.20 GB to 1.68 GB. What keeps the site's fits honest is that its totals run high: on the accuracy page whole-card totals came out 5.6%–8.1% above measured use (4 runs, one RTX 4090 report). If that holds, a setting right at the limit would really use:

LogSite limitWould really useFree in the logFits?
RTX 3090, Linux 23.5 GB 21.75 GB–22.25 GB 23.30 GB yes
RTX 4090, Linux 23.5 GB 21.75 GB–22.25 GB 23.03 GB yes
RTX 5080, Linux 15.5 GB 14.34 GB–14.67 GB 15.19 GB yes
RTX 3090 Ti, Linux 23.5 GB 21.75 GB–22.25 GB 21.82 GB no
RTX 3060 12GB, Windows 11.5 GB 10.64 GB–10.89 GB 10.97 GB yes
RTX 3060 12GB, Windows 11.5 GB 10.64 GB–10.89 GB 10.98 GB yes
RTX 3060 12GB, Windows 11.5 GB 10.64 GB–10.89 GB 10.99 GB yes
RTX 5060 Ti 16GB, Windows 15.5 GB 14.34 GB–14.67 GB 14.82 GB yes
RTX 4080 Super, Windows 15.5 GB 14.34 GB–14.67 GB 14.71 GB yes

The RTX 3090 Ti log, 2.2 GB short of its size with nothing said about what else ran, is the exception. On Windows the margin left in these logs is only 0.03 GB–0.14 GB, and one calibration report is thin evidence, so for a setting the calculator marks close to the limit, close the browser and other GPU programs first, or pick the next smaller context. The calculator shows the total; on several cards each one pays the runtime share (splitting a model across GPUs), and a quantized cache frees context room (KV cache quantization).

Questions

How much VRAM is actually usable on a 24 GB GPU?

In public llama.cpp logs on Linux, CUDA reported 23,863 MiB free on an RTX 3090 and 23,580 MiB free on an RTX 4090 at load, 713–996 MiB short of 24 GiB. The model then also needs its compute buffers. This site budgets 20.9 GB of weights and KV cache for a 24 GB card.

Does Windows use more VRAM than Linux?

In the public logs found, yes: on Windows 1,033–1,323 MiB of each card was already unavailable when llama.cpp loaded, against 713–996 MiB on Linux, including cards that were the second GPU in the machine. The logs do not say what held it, so this is a pattern across a handful of reports, not a measured Windows cost.

What happens when a model spills into shared GPU memory on Windows?

Shared GPU memory is ordinary system RAM the GPU can use, at most half of physical memory at a time (Microsoft). Reading weights from it over PCIe is much slower than VRAM: one Ollama user with two RTX 3090s on Windows saw 95 GB of shared memory in use and about 0.2 tokens/s, against 3–4 tokens/s with the GPU turned off. The NVIDIA Control Panel has a "CUDA - Sysmem Fallback Policy" setting for this; in that report, setting it to prefer no fallback did not help.

What is the compute buffer in the llama.cpp log?

Scratch memory for the forward pass, reserved per device at startup for a worst-case batch of min(context, --ubatch-size) tokens (default ubatch 512). The log prints it as "compute buffer size = N MiB" for each GPU and for the host. Public logs show 218–668 MiB per GPU for 27–31B models with 128K context.

What does --gpu-memory-utilization mean in vLLM?

The share of the GPU's total memory this vLLM instance takes (default 0.92). vLLM computes total × utilization at startup and refuses to start if less than that is free, so on a RTX 3090 (23.56 GiB to CUDA) everything else on the card must fit in about 1.88 GB. Inside its share, whatever the weights, activations and CUDA graphs leave becomes KV cache.

Sources read on .