How to split an LLM across two or more GPUs: VRAM per card in llama.cpp and vLLM
With llama.cpp's default layer split, each GPU holds the weights and KV cache of the layers it is given, in proportion to its free memory, plus its own runtime buffers: Llama 3.1 70B at Q4_K_M with 8K context needs about 23.7 GB + 23.2 GB on two RTX 3090s, and Gemma 4 31B about 11.1 GB + 11.3 GB; vLLM's --tensor-parallel-size N instead gives every GPU 1/N of the weights and of the KV heads.
llama.cpp: the split modes
Read at master 19e28a2 (2026-09-29). From the help text of --split-mode (common/
- none: one GPU only, the one
--main-gpunames; the others are dropped (src/llama.cpp:288-301 ). - layer (default): “split layers and KV across GPUs (pipelined)”.
- row: “split weight across GPUs by rows (parallelized)”; the weights go into a split buffer across the devices (src/
llama-model.cpp:1136-1162 ), and--main-gpuis the GPU “for intermediate results and KV” (arg.cpp:2854-2862). - tensor: “split weights and KV across GPUs (parallelized, EXPERIMENTAL)”, and only for some architectures.
What each GPU holds in layer mode
- Split proportions.
--tensor-splittakes one proportion per GPU, e.g.3,1(arg.cpp:2827-2853). Without it, each GPU's free memory at load time is used; either way the proportions are summed and normalised to cut points (llama-model.cpp:1553-1584). - Layers. With every layer on a GPU there are layers + 1 slots, the last one being the output layer, and slot i goes to the first GPU whose cut point is above i ÷ (layers + 1) (1586-1598). Two equal cards and Llama 3.1 70B's 80 layers: 41 and 39 layers, the output layer on the second.
- Embeddings. The input (token embedding) layer always stays on the CPU (1600-1602);
the output layer goes to the last GPU (1610-1611). A model with tied embeddings such as
Gemma 4 loads a duplicate of the embedding as its output matrix (src/
models/ ), so that matrix is in system RAM and on the last GPU.gemma4.cpp:49-53 - KV cache. Each layer's K and V are allocated on that layer's device, unless the cache is kept on the CPU
(src/
llama-kv-cache.cpp:214-220 ), so the cache splits in the same proportion as the layers. - Buffers. Every device gets its own compute buffer, logged per device
(src/
llama-context.cpp:733-737 ). When several GPUs hold every layer and the KV cache in layer mode, llama.cpp can turn on pipeline parallelism (if every GPU backend supports it), which its source notes “increases memory usage” (427-434).
Checked against llama.cpp logs
Public logs of layer splits print one model, KV and compute buffer per GPU. Below, the site's split of Gemma 4 31B's weights (the rule above, with the site's whole-model bits per weight) next to two of them. The Q8_0 log's CPU buffer is exactly the embedding at 8.5 bits per weight; K-quant files keep some tensors at other types, so the Q5_K_M one differs more.
| Log | Layers per GPU | CPU (embedding) | GPU model buffers, logged | Split here | Error |
|---|---|---|---|---|---|
| Gemma 4 31B Q8_0 on 2× RTX 3090, -ngl 999 -c 32768, no -ts 2026-04-03 | 31 + 29 | 1428.0 MiB logged, 1428.0 here | 15326 MiB + 15783 MiB | 15635 MiB + 16054 MiB | +2.0%, +1.7% |
| Gemma 4 31B Q5_K_M on an RTX 5060 Ti 16GB + RTX 3060 12GB, -sm layer -ts 0.6,0.43 -c 128000, q8_0/q4_0 cache 2026-05-06 | 36 + 24 | 1102.5 MiB logged, 952.6 here | 11798 MiB + 9021 MiB | 12111 MiB + 9027 MiB | +2.7%, +0.1% |
In the second log the KV cache splits 2681.25 : 1787.50 MiB, exactly
36 : 24, the layer counts the rule gives for -ts 0.6,0.43. A third public log,
Qwen3.6 27B Q4_K_M on an RTX 2080 Ti + Tesla P100 16GB, --split-mode layer --tensor-split 22,16 --ctx-size 131072, shows the same pattern: CUDA0 model buffer size = 8500.51 MiB; CUDA1 model buffer size = 6845.15 MiB; CUDA0 KV buffer size = 4608.00 MiB; CUDA1 KV buffer size = 3584.00 MiB; CUDA0 compute buffer size = 668.03 MiB; CUDA1 compute buffer size = 523.54 MiB.
Memory per card at Q4_K_M
Default --tensor-split, with each card's free memory taken as its size; one request, FP16 KV cache. Per card: its
layers' weights (plus the output layer on the last card) and its layers' share of the KV cache, with the site's 10% overhead and
the 0.5 GB every GPU needs for its runtime. “Fits” means no card is over its size and the group keeps 0.5 GB spare, as on the
multi-GPU pages. The token embedding (0.7 GB to 0.8 GB
for these models) stays in system RAM. GB means GiB.
Qwen3.8 27B (64 layers)
| GPUs | 8K context: per card | 32K context: per card |
|---|---|---|
| 2× RTX 3090 | 8.8 GB + 9.1 GB 33 / 31+out layers; fits | 9.7 GB + 9.9 GB 33 / 31+out layers; fits |
| RTX 3090 + RTX 3060 12GB | 11.6 GB + 6.3 GB 44 / 20+out layers; fits | 12.8 GB + 6.9 GB 44 / 20+out layers; fits |
| 2× RTX 4090 | 8.8 GB + 9.1 GB 33 / 31+out layers; fits | 9.7 GB + 9.9 GB 33 / 31+out layers; fits |
| 4× RTX 3090 | 4.8 GB + 4.5 GB + 4.5 GB + 5.1 GB 17 / 16 / 16 / 15+out layers; fits | 5.2 GB + 5.0 GB + 5.0 GB + 5.5 GB 17 / 16 / 16 / 15+out layers; fits |
| 2× RTX 5090 | 8.8 GB + 9.1 GB 33 / 31+out layers; fits | 9.7 GB + 9.9 GB 33 / 31+out layers; fits |
Gemma 4 31B (60 layers)
| GPUs | 8K context: per card | 32K context: per card |
|---|---|---|
| 2× RTX 3090 | 11.1 GB + 11.3 GB 31 / 29+out layers; fits | 12.2 GB + 12.3 GB 31 / 29+out layers; fits |
| RTX 3090 + RTX 3060 12GB | 14.5 GB + 7.9 GB 41 / 19+out layers; fits | 15.9 GB + 8.5 GB 41 / 19+out layers; fits |
| 2× RTX 4090 | 11.1 GB + 11.3 GB 31 / 29+out layers; fits | 12.2 GB + 12.3 GB 31 / 29+out layers; fits |
| 4× RTX 3090 | 6.0 GB + 5.6 GB + 5.6 GB + 6.2 GB 16 / 15 / 15 / 14+out layers; fits | 6.5 GB + 6.1 GB + 6.1 GB + 6.6 GB 16 / 15 / 15 / 14+out layers; fits |
| 2× RTX 5090 | 11.1 GB + 11.3 GB 31 / 29+out layers; fits | 12.2 GB + 12.3 GB 31 / 29+out layers; fits |
Llama 3.1 70B (80 layers)
| GPUs | 8K context: per card | 32K context: per card |
|---|---|---|
| 2× RTX 3090 | 23.7 GB + 23.2 GB 41 / 39+out layers; fits | 27.9 GB + 27.2 GB 41 / 39+out layers; a card is over its size |
| RTX 3090 + RTX 3060 12GB | 31.0 GB + 15.8 GB 54 / 26+out layers; a card is over its size | 36.6 GB + 18.5 GB 54 / 26+out layers; a card is over its size |
| 2× RTX 4090 | 23.7 GB + 23.2 GB 41 / 39+out layers; fits | 27.9 GB + 27.2 GB 41 / 39+out layers; a card is over its size |
| 4× RTX 3090 | 12.4 GB + 11.8 GB + 11.8 GB + 11.9 GB 21 / 20 / 20 / 19+out layers; fits | 14.5 GB + 13.9 GB + 13.9 GB + 13.8 GB 21 / 20 / 20 / 19+out layers; fits |
| 2× RTX 5090 | 23.7 GB + 23.2 GB 41 / 39+out layers; fits | 27.9 GB + 27.2 GB 41 / 39+out layers; fits |
On four RTX 3090s, Llama 3.1 70B at 32K needs about 14.5 GB + 13.9 GB + 13.9 GB + 13.8 GB. Group pages with every model:
2× RTX 3090, 2× RTX 4090, 4× RTX 3090 and 2× RTX 5090;
all setups are on the GPU index. To see what llama-server's automatic --fit would choose
on a given amount of free memory, use the --fit preview.
vLLM: tensor and pipeline parallelism
Read at main ac7f3e1, the code the site's vLLM calculator follows (v0.30.0 has the same logic).
-
--tensor-parallel-size N: every linear layer is cut into N equal parts, its output or input dimension divided by N (linear.py:494, 1681), so each GPU holds 1/N of the weights. The attention heads must divide by N (config/model.py:1502-1509 ), and each GPU caches max(1, KV heads ÷ N) KV heads, copying heads when there are fewer than GPUs (1608-1628). Because every GPU gets the same share, the smallest card sets the limit. -
--gpu-memory-utilization(default 0.92) is the fraction of each GPU's memory vLLM takes; whatever the weights and activations leave of it becomes KV cache (config/cache.py:103-111 ). -
--pipeline-parallel-size N: the layers are divided evenly between N stages, any remainder going to the stages before the last, and theVLLM_PP_LAYER_PARTITIONenvironment variable sets the layer counts by hand (distributed/utils.py:128-167 ). Each stage holds its layers' weights and KV cache, so with that variable a smaller card can be given fewer layers.
Per GPU with the vLLM memory calculator's estimate, AWQ 4-bit weights, --max-model-len 32768, tensor parallel over the group:
| Model | GPUs | Weights per GPU | KV cache memory per GPU | KV cache, tokens |
|---|---|---|---|---|
| Qwen3.8 27B | 2× RTX 3090, TP 2 | 6.9 GB | 13.3 GB | 435,888 |
| Qwen3.8 27B | 2× RTX 4090, TP 2 | 6.9 GB | 13.3 GB | 435,584 |
| Qwen3.8 27B | 4× RTX 3090, TP 4 | 3.4 GB | 16.7 GB | 1,096,992 |
| Qwen3.8 27B | 2× RTX 5090, TP 2 | 6.9 GB | 20.5 GB | 671,936 |
| Gemma 4 31B | 2× RTX 3090, TP 2 | 7.7 GB | 12.4 GB | 126,996 |
| Gemma 4 31B | 2× RTX 4090, TP 2 | 7.7 GB | 12.4 GB | 126,909 |
| Gemma 4 31B | 4× RTX 3090, TP 4 | 3.9 GB | 16.3 GB | 333,002 |
| Gemma 4 31B | 2× RTX 5090, TP 2 | 7.7 GB | 19.6 GB | 200,559 |
| Llama 3.1 70B | 2× RTX 3090, TP 2 | 17.5 GB | 2.7 GB | 17,824less than one 32K request |
| Llama 3.1 70B | 2× RTX 4090, TP 2 | 17.5 GB | 2.7 GB | 17,760less than one 32K request |
| Llama 3.1 70B | 4× RTX 3090, TP 4 | 8.7 GB | 11.4 GB | 150,048 |
| Llama 3.1 70B | 2× RTX 5090, TP 2 | 17.5 GB | 9.9 GB | 65,040 |
Questions
Can I run one model on an RTX 3090 and an RTX 3060 together?
Yes, with llama.cpp's default layer split. Without --tensor-split it divides the layers by each card's free memory, so the 24 GB card takes about two thirds of them: Gemma 4 31B at Q4_K_M with 32K context comes to about 15.9 GB + 8.5 GB (41 layers and 19 layers). vLLM's tensor parallelism gives both GPUs the same share of every weight matrix, so there the 12 GB card sets the limit.
Does the main GPU need more memory than the others?
Not in the default layer mode, where --main-gpu is not used for the model. What differs is the last GPU: it gets the output layer, which for Gemma 4 31B is a copy of the 262,144 × 5,376 embedding matrix, while the input embedding stays in system RAM. --main-gpu matters with --split-mode none (the only GPU used) and row (the GPU that holds the KV cache and intermediate results).
What is the difference between --split-mode layer and row?
layer, the default, gives each GPU whole layers with their KV cache and runs them one after another. row splits each weight matrix across the GPUs by rows so they work on every layer together, and keeps the KV cache and intermediate results on the main GPU. There is also an experimental tensor mode that splits both weights and KV cache.
How do I set --tensor-split?
Give one proportion per GPU, in the order llama.cpp lists them: -ts 3,1 puts three quarters of the layers on the first card. The proportions are normalised, so 24,12 and 2,1 are the same. Leaving it out splits by the free memory each GPU reports at load time.
How much memory does each GPU need with vLLM tensor parallelism?
Each GPU gets 1/N of the weights and of the KV heads (at least one head each, so heads are copied when there are more GPUs than KV heads), and each reserves its own gpu_memory_utilization share of its memory. With AWQ 4-bit weights, Llama 3.1 70B on four RTX 3090s keeps about 8.7 GB of weights on each card.
Sources read on .