What LLMs can a CPU-only PC with 64 GB of dual-channel DDR4-3200 run?

A CPU-only PC with 64 GB of dual-channel DDR4-3200 gives a model 58 GB usable of memory (system RAM minus about 6 GB for the OS and apps) and 51.2 GB/s of bandwidth. Of the 61 open models tracked here, 38 fit at Q4_K_M with an 8,192-token context, and 3 more at a lower precision, each leaving at least 0.5 GB free on this PC.

With no GPU, llama.cpp (or Ollama and LM Studio on top of it) keeps the whole model in the 64 GB of DDR4-3200 and reads all of its active weights over the memory bus for every token. A model fits here when weights, cache and overhead stay within the RAM minus 6 GB held back for the operating system and the apps next to it; close the browser and that reserve shrinks. MoE offload does not apply, since everything already runs from RAM, but mixture-of-experts models still help most: only their active experts are read per token. The speeds use 50–80% of the rated 51.2 GB/s for dense models and 30–55% for MoE models, from public CPU-only llama.cpp runs. Reading the prompt is compute-bound on a CPU, often tens of tokens per second, so a long prompt can take minutes before the first word; a GPU of any size speeds that part up even when the weights stay in RAM.

Models that fit in 64 GB of DDR4-3200 (dual channel, CPU only)

The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-20b (14.8 GB, 5.9–11 tokens/s on this PC) is listed at MXFP4 only, as published.

Model Best precision Memory Longest context Tokens/s
Mistral Medium 3.5 128B 127.7B GGUF IQ3_XXS 57.5 GB 7K 0.5–0.7
Qwen3.5 122B-A10B 125.1B, 10B active GGUF Q2_K 54.4 GB 129K 3.5–6.4
Nemotron 3 Super 120B-A12B 123.6B, 12B active GGUF Q2_K 53.6 GB 256K (full) 3.0–5.5
AliceAI Foundation 80B-A3B 81.3B, 3B active GGUF Q4_K_M 51.1 GB 256K (full) 7.5–14
Qwen3-Coder-Next 79.7B, 3B active GGUF Q4_K_M 50.1 GB 256K (full) 7.5–14
Llama 3.1 70B 70.6B GGUF Q5_K_M 54.5 GB 16K 0.5–0.8
K2-Horizon MoVA 36B-A4B 37.4B, 4B active GGUF Q8_0 42.9 GB 78K 2.6–4.8
Ornith 1.5 35B-A3B 36.0B, 3B active GGUF Q8_0 39.8 GB 256K (full) 4.5–8.4
Qwen3.6 35B-A3B 36.0B, 3B active GGUF Q8_0 39.8 GB 256K (full) 4.5–8.4
Ornith 1.0 35B 35.1B, 3B active GGUF Q8_0 38.9 GB 256K (full) 4.5–8.4
LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active GGUF Q8_0 36.0 GB 64K (full) 3.3–6.1
Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active GGUF Q8_0 34.9 GB 256K (full) 4.1–7.4
Gemma 4 31B 31.3B GGUF Q8_0 36.5 GB 252K 0.7–1.2
GLM-4.7 Flash 31.2B, 3B active GGUF Q8_0 34.9 GB 198K (full) 4.2–7.7
Xing 4.0 29B-A4B 31.2B, 4B active GGUF Q8_0 34.9 GB 256K (full) 3.3–6.1
Qwen3-Coder 30B-A3B 30.5B, 3.3B active GGUF Q8_0 34.6 GB 230K 3.5–6.5
Qwen3 30B-A3B 30.5B, 3.3B active GGUF Q8_0 34.6 GB 40K (full) 3.5–6.5
Muse Glimmer 30B 29.8B GGUF Q8_0 33.1 GB 128K (full) 0.8–1.3
Granite 4.2 30B 29.3B GGUF Q8_0 34.6 GB 91K 0.8–1.2
Qwen3.6 27B 27.8B GGUF Q8_0 31.3 GB 256K (full) 0.9–1.4
Qwen3.8 27B 27.8B GGUF Q8_0 31.3 GB 256K (full) 0.9–1.4
Hemmingway-1 27B 27.3B FP16 / BF16 57.0 GB 14K 0.5–0.7
Gemma 4 26B-A4B 25.8B, 4B active FP16 / BF16 53.9 GB 176K 1.8–3.3
gpt-oss-20b 20.9B, 3.6B active As published (MXFP4) 14.8 GB 128K (full) 5.9–11
Gemma 4 12B 12.0B FP16 / BF16 25.7 GB 256K (full) 1.0–1.7
ZDTaichu 5.0 9B 9.8B FP16 / BF16 20.8 GB 128K (full) 1.3–2.1
Ornith 1.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 1.3–2.1
Qwen3.5 9B 9.7B FP16 / BF16 20.6 GB 256K (full) 1.3–2.1
Ornith 1.0 9B 9.4B FP16 / BF16 20.1 GB 256K (full) 1.3–2.1
MiMo V2.6 Distill Qwen 9B 9.4B FP16 / BF16 20.1 GB 256K (full) 1.3–2.1
Granite 4.2 8B 8.8B FP16 / BF16 19.9 GB 128K (full) 1.3–2.2
LFM2.5 8B-A1B 8.5B, 1.5B active FP16 / BF16 18.0 GB 128,000 (full) 4.9–9.0
Qwen3 8B 8.2B FP16 / BF16 18.5 GB 40K (full) 1.5–2.3
Llama 3.1 8B 8.0B FP16 / BF16 18.1 GB 128K (full) 1.5–2.4
Gemma 4 E4B 8.0B FP16 / BF16 17.1 GB 128K (full) 1.6–2.5
Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active FP16 / BF16 16.7 GB 128K (full) 5.7–11
Spark-X2.5 4B 4.1B FP16 / BF16 9.35 GB 1M (full) 3.0–4.7
Nemotron 3 Nano 4B 4.0B FP16 / BF16 8.78 GB 256K (full) 3.2–5.1
Granite 4.2 3B 3.7B FP16 / BF16 8.69 GB 128K (full) 3.2–5.1
MiniCPM5 2B 2.5B FP16 / BF16 6.02 GB 128K (full) 4.7–7.6
Limite 1B Violetto 1.0B FP16 / BF16 2.79 GB 128K (full) 11–18

What one GPU adds to this PC

With a graphics card, llama.cpp's --n-cpu-moe keeps the routed experts of the first layers in this PC's 64 GB of DDR4-3200 and everything else on the card, so a mixture-of-experts model can run far bigger than the card and far faster than the CPU alone. Each measured GGUF file below at its own 32K context (or its whole window): tokens per second on the CPU alone at MXFP4 where it fits, then the smallest --n-cpu-moe and tokens per second with an RTX 5060 Ti 16GB or an RTX 4090, the experts read at 51.2 GB/s. "More RAM" means the experts left on the CPU need more than the 58 GB this PC leaves for a model.

Model GGUF CPU only --n-cpu-moe, RTX 5060 Ti 16GB--n-cpu-moe, RTX 4090 Tokens/s, RTX 5060 Ti 16GBTokens/s, RTX 4090
gpt-oss-20b MXFP4 4.8–8.8 00 37–6478–138
Gemma 4 26B-A4B UD-Q4_K_M 4.5–8.2 30 29–5073–128
Qwen3 30B-A3B Q4_K_M 2.9–5.4 150 17–2854–93
GLM-4.7 Flash Q4_K_M 4.3–7.8 120 21–3665–114
Nemotron 3 Nano 30B-A3B Q4_K_M 6.6–12 210 21–3694–167
Ornith 1.5 35B-A3B Q4_K_M 6.1–11 130 30–5296–171
Qwen3.6 35B-A3B UD-Q4_K_M 6.1–11 130 27–4581–142
Qwen3-Coder-Next Q4_K_M 5.8–11 3526 16–2623–40
gpt-oss-120b MXFP4 Does not fit 2924 8–1310–17
Qwen3.5 122B-A10B Q4_K_M Does not fit 4236 6–98–13
Qwen3.8 Flash Next UD-Q4_K_XL Does not fit 42 (more RAM)37 —10–17

Too big for 64 GB of DDR4-3200 (dual channel, CPU only) at 8K context

How much memory each of these lacks at Q4_K_M (gpt-oss-120b at MXFP4) with 8K context, one request and 0.5 GB kept free, out of the 58 GB this PC leaves for a model; more RAM or a smaller quant is the way in.

Best models for 64 GB of DDR4-3200 (dual channel, CPU only) with room for long context

The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the PC takes at that precision and the writing speed with that context in the cache, from 51.2 GB/s of bandwidth.

ModelContextMemoryLeft overLongest contextTokens/s
AliceAI Foundation 80B-A3B 81.3B 32K 51.7 GB 6.29 GB 256K (full) 5.8–11
Qwen3-Coder-Next 79.7B 32K 50.7 GB 7.29 GB 256K (full) 5.8–11
Llama 3.1 70B 70.6B 32K 55.2 GB 2.77 GB 38K 0.5–0.8
K2-Horizon MoVA 36B-A4B 37.4B 32K 30.3 GB 27.7 GB 163K 1.7–3.2
Ornith 1.5 35B-A3B 36.0B 32K 23.5 GB 34.5 GB 256K (full) 6.1–11

llama-server commands for a CPU-only PC with 64 GB of dual-channel DDR4-3200

GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on this PC. -ngl 0 keeps every layer in system RAM, even in a CUDA or Vulkan build of llama.cpp.

Llama 3.1 70B, Q4_K_M with 38K tokens: 0.5–0.7 tokens/s on this PC

llama-server -hf bartowski/Meta-Llama-3.1-70B-Instruct-GGUF:Q4_K_M -c 38912 -ngl 0

Meta-Llama-3.1-70B-Instruct-Q4_K_M.gguf, 42.5 GB: 57.3 GB used and 726 MB free of this PC's 58 GB usable, 0.5–0.7 tokens/s once the 38K cache is full.

K2-Horizon MoVA 36B-A4B, Q4_K_M with 163K tokens: 0.4–0.8 tokens/s on this PC

llama-server -hf IFM/K2-Horizon-MoVA-36B-A4B-GGUF:Q4_K_M -c 166912 -ngl 0

K2-Horizon-MoVA-36B-A4B-Q4_K_M.gguf, 22.4 GB: 57.3 GB used and 689 MB free of this PC's 58 GB usable, 0.4–0.8 tokens/s once the 163K cache is full.

Qwen3-Coder-Next, Q4_K_M with 256K tokens: 1.9–3.4 tokens/s on this PC

llama-server -hf unsloth/Qwen3-Coder-Next-GGUF:Q4_K_M -c 262144 -ngl 0

Qwen3-Coder-Next-Q4_K_M.gguf, 48.5 GB: 56.8 GB used and 1.18 GB free of this PC's 58 GB usable, 1.9–3.4 tokens/s once the 256K cache is full.

How fast this PC writes as the context fills

Tokens per second at Q4_K_M for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the PC holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 51.2 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, CPU only, DDR5-5600 dual channel (64 GB), 89.6 GB/s, writes 74% faster on average.

Model8K32KLongestAt the longestQ8_0 at 8KBandwidth ceiling at 8K1,000-token reply at 32K
AliceAI Foundation 80B-A3B 7.5–14 5.8–11 256K 1.9–3.4 Does not fit 25 94–172 s
Qwen3-Coder-Next 7.5–14 5.8–11 256K 1.9–3.4 Does not fit 25 94–172 s
Llama 3.1 70B 0.6–0.9 0.5–0.8 38K 0.5–0.7 Does not fit 1 1305–2088 s
K2-Horizon MoVA 36B-A4B 3.8–7.0 1.7–3.2 163K 0.4–0.8 2.6–4.8 13 315–578 s
Ornith 1.5 35B-A3B 7.7–14 6.1–11 256K 2.1–3.9 4.5–8.4 26 89–163 s
Qwen3.6 35B-A3B 7.7–14 6.1–11 256K 2.1–3.9 4.5–8.4 26 89–163 s
Ornith 1.0 35B 7.7–14 6.1–11 256K 2.1–3.9 4.5–8.4 26 89–163 s
LLM-jp-4.1 32B-A3B Thinking 5.3–9.8 3.4–6.3 64K 2.3–4.3 3.3–6.1 18 159–292 s
Nemotron 3 Nano 30B-A3B 7.0–13 6.6–12 256K 4.1–7.5 4.1–7.4 24 83–152 s
Gemma 4 31B 1.2–2.0 1.1–1.8 256K 0.6–1.0 0.7–1.2 2 559–895 s

64 GB of DDR4-3200 (dual channel, CPU only) against the other CPU-only PCs with 58 GB usable

The same models fit on every CPU-only PC with 58 GB usable, so speed is what separates them. By bandwidth the CPU only, DDR4-3200 dual channel (64 GB) is the slowest of the 2 CPU-only PCs with 58 GB usable at 51.2 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the CPU only, DDR5-5600 dual channel (64 GB), at 89.6 GB/s, writes 74% faster than the CPU only, DDR4-3200 dual channel (64 GB). Each cell below is the other CPU-only PC's tokens per second minus this one's, midpoints of the estimated ranges.

Model CPU only, DDR4-3200 dual channel (64 GB) tokens/s CPU only, DDR5-5600 dual channel (64 GB)
AliceAI Foundation 80B-A3B 7.5–14 +7.9
Qwen3-Coder-Next 7.5–14 +7.9
Llama 3.1 70B 0.6–0.9 +0.5
K2-Horizon MoVA 36B-A4B 3.8–7.0 +4.0
Ornith 1.5 35B-A3B 7.7–14 +8.0
Qwen3.6 35B-A3B 7.7–14 +8.0
Ornith 1.0 35B 7.7–14 +8.0
LLM-jp-4.1 32B-A3B Thinking 5.3–9.8 +5.6
Nemotron 3 Nano 30B-A3B 7.0–13 +7.4
Gemma 4 31B 1.2–2.0 +1.2
GLM-4.7 Flash 6.7–12 +7.1
Xing 4.0 29B-A4B 5.4–10 +5.7

Near misses on 64 GB of DDR4-3200 (dual channel, CPU only)

Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.

ModelShort by at Q4_K_MFits instead
Nemotron 3 Super 120B-A12B 123.6B 19.4 GB GGUF IQ3_XXS with 32K, 53.0 GB
Qwen3.5 122B-A10B 125.1B 20.9 GB GGUF IQ3_XXS with 32K, 54.2 GB

How these numbers are worked out

Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.

Nearby GPUs and setups

How many of the tracked models each one holds at Q4_K_M with 8K context, against 38 here.

SetupHow it relatesMemoryBandwidthModels at Q4_K_M
CPU only, DDR5-5600 dual channel (64 GB) Same memory 58 GB usable 89.6 GB/s 38 (same)
CPU only, DDR5-5600 dual channel (32 GB) Next size down 26 GB usable 89.6 GB/s 35 (−3)
CPU only, DDR5-5600 dual channel (128 GB) Next size up 122 GB usable 89.6 GB/s 43 (+5)

Model numbers read from Hugging Face on .