How to split an LLM across two or more GPUs: VRAM per card in llama.cpp and vLLM

With llama.cpp's default layer split, each GPU holds the weights and KV cache of the layers it is given, in proportion to its free memory, plus its own runtime buffers: Llama 3.1 70B at Q4_K_M with 8K context needs about 23.7 GB + 23.2 GB on two RTX 3090s, and Gemma 4 31B about 11.1 GB + 11.3 GB; vLLM's --tensor-parallel-size N instead gives every GPU 1/N of the weights and of the KV heads.

llama.cpp: the split modes

Read at master 19e28a2 (2026-09-29). From the help text of --split-mode (common/arg.cpp:2803-2826):

What each GPU holds in layer mode

  1. Split proportions. --tensor-split takes one proportion per GPU, e.g. 3,1 (arg.cpp:2827-2853). Without it, each GPU's free memory at load time is used; either way the proportions are summed and normalised to cut points (llama-model.cpp:1553-1584).
  2. Layers. With every layer on a GPU there are layers + 1 slots, the last one being the output layer, and slot i goes to the first GPU whose cut point is above i ÷ (layers + 1) (1586-1598). Two equal cards and Llama 3.1 70B's 80 layers: 41 and 39 layers, the output layer on the second.
  3. Embeddings. The input (token embedding) layer always stays on the CPU (1600-1602); the output layer goes to the last GPU (1610-1611). A model with tied embeddings such as Gemma 4 loads a duplicate of the embedding as its output matrix (src/models/gemma4.cpp:49-53), so that matrix is in system RAM and on the last GPU.
  4. KV cache. Each layer's K and V are allocated on that layer's device, unless the cache is kept on the CPU (src/llama-kv-cache.cpp:214-220), so the cache splits in the same proportion as the layers.
  5. Buffers. Every device gets its own compute buffer, logged per device (src/llama-context.cpp:733-737). When several GPUs hold every layer and the KV cache in layer mode, llama.cpp can turn on pipeline parallelism (if every GPU backend supports it), which its source notes “increases memory usage” (427-434).

Checked against llama.cpp logs

Public logs of layer splits print one model, KV and compute buffer per GPU. Below, the site's split of Gemma 4 31B's weights (the rule above, with the site's whole-model bits per weight) next to two of them. The Q8_0 log's CPU buffer is exactly the embedding at 8.5 bits per weight; K-quant files keep some tensors at other types, so the Q5_K_M one differs more.

LogLayers per GPUCPU (embedding)GPU model buffers, loggedSplit hereError
Gemma 4 31B Q8_0 on 2× RTX 3090, -ngl 999 -c 32768, no -ts 2026-04-03 31 + 29 1428.0 MiB logged, 1428.0 here 15326 MiB + 15783 MiB 15635 MiB + 16054 MiB +2.0%, +1.7%
Gemma 4 31B Q5_K_M on an RTX 5060 Ti 16GB + RTX 3060 12GB, -sm layer -ts 0.6,0.43 -c 128000, q8_0/q4_0 cache 2026-05-06 36 + 24 1102.5 MiB logged, 952.6 here 11798 MiB + 9021 MiB 12111 MiB + 9027 MiB +2.7%, +0.1%

In the second log the KV cache splits 2681.25 : 1787.50 MiB, exactly 36 : 24, the layer counts the rule gives for -ts 0.6,0.43. A third public log, Qwen3.6 27B Q4_K_M on an RTX 2080 Ti + Tesla P100 16GB, --split-mode layer --tensor-split 22,16 --ctx-size 131072, shows the same pattern: CUDA0 model buffer size = 8500.51 MiB; CUDA1 model buffer size = 6845.15 MiB; CUDA0 KV buffer size = 4608.00 MiB; CUDA1 KV buffer size = 3584.00 MiB; CUDA0 compute buffer size = 668.03 MiB; CUDA1 compute buffer size = 523.54 MiB.

Memory per card at Q4_K_M

Default --tensor-split, with each card's free memory taken as its size; one request, FP16 KV cache. Per card: its layers' weights (plus the output layer on the last card) and its layers' share of the KV cache, with the site's 10% overhead and the 0.5 GB every GPU needs for its runtime. “Fits” means no card is over its size and the group keeps 0.5 GB spare, as on the multi-GPU pages. The token embedding (0.7 GB to 0.8 GB for these models) stays in system RAM. GB means GiB.

Qwen3.8 27B (64 layers)

GPUs8K context: per card32K context: per card
2× RTX 3090 8.8 GB + 9.1 GB 33 / 31+out layers; fits 9.7 GB + 9.9 GB 33 / 31+out layers; fits
RTX 3090 + RTX 3060 12GB 11.6 GB + 6.3 GB 44 / 20+out layers; fits 12.8 GB + 6.9 GB 44 / 20+out layers; fits
2× RTX 4090 8.8 GB + 9.1 GB 33 / 31+out layers; fits 9.7 GB + 9.9 GB 33 / 31+out layers; fits
4× RTX 3090 4.8 GB + 4.5 GB + 4.5 GB + 5.1 GB 17 / 16 / 16 / 15+out layers; fits 5.2 GB + 5.0 GB + 5.0 GB + 5.5 GB 17 / 16 / 16 / 15+out layers; fits
2× RTX 5090 8.8 GB + 9.1 GB 33 / 31+out layers; fits 9.7 GB + 9.9 GB 33 / 31+out layers; fits

Gemma 4 31B (60 layers)

GPUs8K context: per card32K context: per card
2× RTX 3090 11.1 GB + 11.3 GB 31 / 29+out layers; fits 12.2 GB + 12.3 GB 31 / 29+out layers; fits
RTX 3090 + RTX 3060 12GB 14.5 GB + 7.9 GB 41 / 19+out layers; fits 15.9 GB + 8.5 GB 41 / 19+out layers; fits
2× RTX 4090 11.1 GB + 11.3 GB 31 / 29+out layers; fits 12.2 GB + 12.3 GB 31 / 29+out layers; fits
4× RTX 3090 6.0 GB + 5.6 GB + 5.6 GB + 6.2 GB 16 / 15 / 15 / 14+out layers; fits 6.5 GB + 6.1 GB + 6.1 GB + 6.6 GB 16 / 15 / 15 / 14+out layers; fits
2× RTX 5090 11.1 GB + 11.3 GB 31 / 29+out layers; fits 12.2 GB + 12.3 GB 31 / 29+out layers; fits

Llama 3.1 70B (80 layers)

GPUs8K context: per card32K context: per card
2× RTX 3090 23.7 GB + 23.2 GB 41 / 39+out layers; fits 27.9 GB + 27.2 GB 41 / 39+out layers; a card is over its size
RTX 3090 + RTX 3060 12GB 31.0 GB + 15.8 GB 54 / 26+out layers; a card is over its size 36.6 GB + 18.5 GB 54 / 26+out layers; a card is over its size
2× RTX 4090 23.7 GB + 23.2 GB 41 / 39+out layers; fits 27.9 GB + 27.2 GB 41 / 39+out layers; a card is over its size
4× RTX 3090 12.4 GB + 11.8 GB + 11.8 GB + 11.9 GB 21 / 20 / 20 / 19+out layers; fits 14.5 GB + 13.9 GB + 13.9 GB + 13.8 GB 21 / 20 / 20 / 19+out layers; fits
2× RTX 5090 23.7 GB + 23.2 GB 41 / 39+out layers; fits 27.9 GB + 27.2 GB 41 / 39+out layers; fits

On four RTX 3090s, Llama 3.1 70B at 32K needs about 14.5 GB + 13.9 GB + 13.9 GB + 13.8 GB. Group pages with every model: 2× RTX 3090, 2× RTX 4090, 4× RTX 3090 and 2× RTX 5090; all setups are on the GPU index. To see what llama-server's automatic --fit would choose on a given amount of free memory, use the --fit preview.

vLLM: tensor and pipeline parallelism

Read at main ac7f3e1, the code the site's vLLM calculator follows (v0.30.0 has the same logic).

Per GPU with the vLLM memory calculator's estimate, AWQ 4-bit weights, --max-model-len 32768, tensor parallel over the group:

ModelGPUsWeights per GPUKV cache memory per GPUKV cache, tokens
Qwen3.8 27B 2× RTX 3090, TP 2 6.9 GB 13.3 GB 435,888
Qwen3.8 27B 2× RTX 4090, TP 2 6.9 GB 13.3 GB 435,584
Qwen3.8 27B 4× RTX 3090, TP 4 3.4 GB 16.7 GB 1,096,992
Qwen3.8 27B 2× RTX 5090, TP 2 6.9 GB 20.5 GB 671,936
Gemma 4 31B 2× RTX 3090, TP 2 7.7 GB 12.4 GB 126,996
Gemma 4 31B 2× RTX 4090, TP 2 7.7 GB 12.4 GB 126,909
Gemma 4 31B 4× RTX 3090, TP 4 3.9 GB 16.3 GB 333,002
Gemma 4 31B 2× RTX 5090, TP 2 7.7 GB 19.6 GB 200,559
Llama 3.1 70B 2× RTX 3090, TP 2 17.5 GB 2.7 GB 17,824less than one 32K request
Llama 3.1 70B 2× RTX 4090, TP 2 17.5 GB 2.7 GB 17,760less than one 32K request
Llama 3.1 70B 4× RTX 3090, TP 4 8.7 GB 11.4 GB 150,048
Llama 3.1 70B 2× RTX 5090, TP 2 17.5 GB 9.9 GB 65,040

Questions

Can I run one model on an RTX 3090 and an RTX 3060 together?

Yes, with llama.cpp's default layer split. Without --tensor-split it divides the layers by each card's free memory, so the 24 GB card takes about two thirds of them: Gemma 4 31B at Q4_K_M with 32K context comes to about 15.9 GB + 8.5 GB (41 layers and 19 layers). vLLM's tensor parallelism gives both GPUs the same share of every weight matrix, so there the 12 GB card sets the limit.

Does the main GPU need more memory than the others?

Not in the default layer mode, where --main-gpu is not used for the model. What differs is the last GPU: it gets the output layer, which for Gemma 4 31B is a copy of the 262,144 × 5,376 embedding matrix, while the input embedding stays in system RAM. --main-gpu matters with --split-mode none (the only GPU used) and row (the GPU that holds the KV cache and intermediate results).

What is the difference between --split-mode layer and row?

layer, the default, gives each GPU whole layers with their KV cache and runs them one after another. row splits each weight matrix across the GPUs by rows so they work on every layer together, and keeps the KV cache and intermediate results on the main GPU. There is also an experimental tensor mode that splits both weights and KV cache.

How do I set --tensor-split?

Give one proportion per GPU, in the order llama.cpp lists them: -ts 3,1 puts three quarters of the layers on the first card. The proportions are normalised, so 24,12 and 2,1 are the same. Leaving it out splits by the free memory each GPU reports at load time.

How much memory does each GPU need with vLLM tensor parallelism?

Each GPU gets 1/N of the weights and of the KV heads (at least one head each, so heads are copied when there are more GPUs than KV heads), and each reserves its own gpu_memory_utilization share of its memory. With AWQ 4-bit weights, Llama 3.1 70B on four RTX 3090s keeps about 8.7 GB of weights on each card.

Sources read on .