What LLMs can a CPU-only PC with 128 GB of quad-channel DDR4-3200 run?
A CPU-only PC with 128 GB of quad-channel DDR4-3200 gives a model 122 GB usable of memory (system RAM minus about 6 GB for the OS and apps) and 102.4 GB/s of bandwidth. Of the 61 open models tracked here, 43 fit at Q4_K_M with an 8,192-token context, and 2 more at a lower precision, each leaving at least 0.5 GB free on this PC.
With no GPU, llama.cpp (or Ollama and LM Studio on top of it) keeps the whole model in the 128 GB of DDR4-3200 and reads all of its active weights over the memory bus for every token. A model fits here when weights, cache and overhead stay within the RAM minus 6 GB held back for the operating system and the apps next to it; close the browser and that reserve shrinks. MoE offload does not apply, since everything already runs from RAM, but mixture-of-experts models still help most: only their active experts are read per token. The speeds use 50–80% of the rated 102.4 GB/s for dense models and 30–55% for MoE models, from public CPU-only llama.cpp runs. Reading the prompt is compute-bound on a CPU, often tens of tokens per second, so a long prompt can take minutes before the first word; a GPU of any size speeds that part up even when the weights stay in RAM.
Models that fit in 128 GB of DDR4-3200 (quad channel, CPU only)
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models come first. gpt-oss-120b (67.7 GB, 9.5–18 tokens/s on this PC) and gpt-oss-20b (14.8 GB, 12–22 tokens/s on this PC) are listed at MXFP4 only, as published.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|---|---|---|---|
| MiniMax M2.7 228.7B, 10B active | GGUF Q3_K_M | 117 GB | 24K | 4.4–8.1 |
| Step 3.7 Flash 196B-A11B 201.4B, 11B active | GGUF Q3_K_M | 102 GB | 256K (full) | 5.2–9.5 |
| Qwen3.8 Flash Next 180.0B, 6B active | GGUF Q4_K_M | 112 GB | 256K (full) | 7.9–15 |
| Mistral Medium 3.5 128B 127.7B | GGUF Q6_K | 111 GB | 36K | 0.5–0.8 |
| Qwen3.5 122B-A10B 125.1B, 10B active | GGUF Q6_K | 106 GB | 256K (full) | 3.6–6.7 |
| Nemotron 3 Super 120B-A12B 123.6B, 12B active | GGUF Q6_K | 104 GB | 256K (full) | 3.1–5.7 |
| gpt-oss-120b 116.8B, 5.1B active | As published (MXFP4) | 67.7 GB | 128K (full) | 9.5–18 |
| AliceAI Foundation 80B-A3B 81.3B, 3B active | GGUF Q8_0 | 89.2 GB | 256K (full) | 8.9–16 |
| Qwen3-Coder-Next 79.7B, 3B active | GGUF Q8_0 | 87.4 GB | 256K (full) | 8.9–16 |
| Llama 3.1 70B 70.6B | GGUF Q8_0 | 80.0 GB | 128K (full) | 0.7–1.1 |
| K2-Horizon MoVA 36B-A4B 37.4B, 4B active | FP16 / BF16 | 78.9 GB | 214K | 3.2–5.8 |
| Ornith 1.5 35B-A3B 36.0B, 3B active | FP16 / BF16 | 74.3 GB | 256K (full) | 4.9–9.1 |
| Qwen3.6 35B-A3B 36.0B, 3B active | FP16 / BF16 | 74.3 GB | 256K (full) | 4.9–9.1 |
| Ornith 1.0 35B 35.1B, 3B active | FP16 / BF16 | 72.6 GB | 256K (full) | 4.9–9.1 |
| LLM-jp-4.1 32B-A3B Thinking 32.1B, 3.8B active | FP16 / BF16 | 66.9 GB | 64K (full) | 3.7–6.9 |
| Nemotron 3 Nano 30B-A3B 31.6B, 3.5B active | FP16 / BF16 | 65.3 GB | 256K (full) | 4.3–8.0 |
| Gemma 4 31B 31.3B | FP16 / BF16 | 66.6 GB | 256K (full) | 0.8–1.3 |
| GLM-4.7 Flash 31.2B, 3B active | FP16 / BF16 | 64.9 GB | 198K (full) | 4.7–8.7 |
| Xing 4.0 29B-A4B 31.2B, 4B active | FP16 / BF16 | 64.8 GB | 256K (full) | 3.6–6.7 |
| Qwen3-Coder 30B-A3B 30.5B, 3.3B active | FP16 / BF16 | 63.9 GB | 256K (full) | 4.1–7.6 |
| Qwen3 30B-A3B 30.5B, 3.3B active | FP16 / BF16 | 63.9 GB | 40K (full) | 4.1–7.6 |
| Muse Glimmer 30B 29.8B | FP16 / BF16 | 61.7 GB | 128K (full) | 0.9–1.4 |
| Granite 4.2 30B 29.3B | FP16 / BF16 | 62.7 GB | 128K (full) | 0.8–1.3 |
| Qwen3.6 27B 27.8B | FP16 / BF16 | 58.0 GB | 256K (full) | 0.9–1.5 |
| Qwen3.8 27B 27.8B | FP16 / BF16 | 58.0 GB | 256K (full) | 0.9–1.5 |
| Hemmingway-1 27B 27.3B | FP16 / BF16 | 57.0 GB | 256K (full) | 0.9–1.5 |
| Gemma 4 26B-A4B 25.8B, 4B active | FP16 / BF16 | 53.9 GB | 256K (full) | 3.6–6.6 |
| gpt-oss-20b 20.9B, 3.6B active | As published (MXFP4) | 14.8 GB | 128K (full) | 12–22 |
| Gemma 4 12B 12.0B | FP16 / BF16 | 25.7 GB | 256K (full) | 2.1–3.3 |
| ZDTaichu 5.0 9B 9.8B | FP16 / BF16 | 20.8 GB | 128K (full) | 2.6–4.1 |
| Ornith 1.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 2.6–4.2 |
| Qwen3.5 9B 9.7B | FP16 / BF16 | 20.6 GB | 256K (full) | 2.6–4.2 |
| Ornith 1.0 9B 9.4B | FP16 / BF16 | 20.1 GB | 256K (full) | 2.7–4.3 |
| MiMo V2.6 Distill Qwen 9B 9.4B | FP16 / BF16 | 20.1 GB | 256K (full) | 2.7–4.3 |
| Granite 4.2 8B 8.8B | FP16 / BF16 | 19.9 GB | 128K (full) | 2.7–4.3 |
| LFM2.5 8B-A1B 8.5B, 1.5B active | FP16 / BF16 | 18.0 GB | 128,000 (full) | 9.8–18 |
| Qwen3 8B 8.2B | FP16 / BF16 | 18.5 GB | 40K (full) | 2.9–4.6 |
| Llama 3.1 8B 8.0B | FP16 / BF16 | 18.1 GB | 128K (full) | 3.0–4.8 |
| Gemma 4 E4B 8.0B | FP16 / BF16 | 17.1 GB | 128K (full) | 3.2–5.1 |
| Ling 3.0 Tiny 7.9B-A1.3B 7.9B, 1.3B active | FP16 / BF16 | 16.7 GB | 128K (full) | 11–21 |
| Spark-X2.5 4B 4.1B | FP16 / BF16 | 9.35 GB | 1M (full) | 5.9–9.4 |
| Nemotron 3 Nano 4B 4.0B | FP16 / BF16 | 8.78 GB | 256K (full) | 6.3–10 |
| Granite 4.2 3B 3.7B | FP16 / BF16 | 8.69 GB | 128K (full) | 6.3–10 |
| MiniCPM5 2B 2.5B | FP16 / BF16 | 6.02 GB | 128K (full) | 9.4–15 |
| Limite 1B Violetto 1.0B | FP16 / BF16 | 2.79 GB | 128K (full) | 22–36 |
What one GPU adds to this PC
With a graphics card, llama.cpp's --n-cpu-moe keeps the routed experts of the first layers in this PC's
128 GB of DDR4-3200 and everything else on the card, so a mixture-of-experts model can run far bigger than the card and far faster than
the CPU alone. Each measured GGUF file below at its own 32K context (or its whole window): tokens per second on the CPU alone
at MXFP4 where it fits, then the smallest --n-cpu-moe and tokens per second with an
RTX 5060 Ti 16GB or an RTX 4090, the experts read at 102.4 GB/s.
"More RAM" means the experts left on the CPU need more than the 122 GB this PC leaves for a model.
| Model | GGUF | CPU only | --n-cpu-moe, RTX 5060 Ti 16GB | --n-cpu-moe, RTX 4090 | Tokens/s, RTX 5060 Ti 16GB | Tokens/s, RTX 4090 |
|---|---|---|---|---|---|---|
| gpt-oss-20b | MXFP4 | 9.5–17 | 0 | 0 | 37–64 | 78–138 |
| Gemma 4 26B-A4B | UD-Q4_K_M | 8.9–16 | 3 | 0 | 32–55 | 73–128 |
| Qwen3 30B-A3B | Q4_K_M | 5.8–11 | 15 | 0 | 21–35 | 54–93 |
| GLM-4.7 Flash | Q4_K_M | 8.5–16 | 12 | 0 | 26–44 | 65–114 |
| Nemotron 3 Nano 30B-A3B | Q4_K_M | 13–24 | 21 | 0 | 30–52 | 94–167 |
| Ornith 1.5 35B-A3B | Q4_K_M | 12–22 | 13 | 0 | 38–65 | 96–171 |
| Qwen3.6 35B-A3B | UD-Q4_K_M | 12–22 | 13 | 0 | 32–55 | 81–142 |
| Qwen3-Coder-Next | Q4_K_M | 12–21 | 35 | 26 | 24–40 | 37–64 |
| gpt-oss-120b | MXFP4 | 7.4–14 | 29 | 24 | 13–22 | 18–31 |
| Qwen3.5 122B-A10B | Q4_K_M | 4.5–8.2 | 42 | 36 | 9–15 | 13–22 |
| Qwen3.8 Flash Next | UD-Q4_K_XL | 6.9–13 | 42 | 37 | 11–19 | 17–29 |
Too big for 128 GB of DDR4-3200 (quad channel, CPU only) at 8K context
How much memory each of these lacks at Q4_K_M with 8K context, one request and 0.5 GB kept free, out of the 122 GB this PC leaves for a model; more RAM or a smaller quant is the way in.
- DeepSeek V4 Flash: 60.1 GB short
- DeepSeek V4 Flash 0731: 68.3 GB short
- MiMo V2.6 Flash: 72.0 GB short
- IQuest-Q1 320B-A15B: 79.6 GB short
- GLM-5.3 Flash: 78.3 GB short
- MiniMax M3: 145 GB short
- DeepSeek V3 / R1: 304 GB short
- DeepSeek V3.2: 304 GB short
- GLM-5.3: 347 GB short
- GLM-5.2: 347 GB short
- DeepSeek V4.1 Flash: 353 GB short
- Hy4 Preview 770B-A49B: 363 GB short
- MiMo V2.6 Pro: 514 GB short
- DeepSeek V4 Pro: 871 GB short
- Qwen3.8 2.4T-A95B: 1,396 GB short
- Kimi K3: 1,602 GB short
Best models for 128 GB of DDR4-3200 (quad channel, CPU only) with room for long context
The biggest models that still fit at Q4_K_M with 32,768 tokens of context (or their whole window, if shorter), the memory left over, the longest context the PC takes at that precision and the writing speed with that context in the cache, from 102.4 GB/s of bandwidth. gpt-oss-120b is at MXFP4, as published.
| Model | Context | Memory | Left over | Longest context | Tokens/s |
|---|---|---|---|---|---|
| Qwen3.8 Flash Next 180.0B | 32K | 113 GB | 9.11 GB | 256K (full) | 6.9–13 |
| Mistral Medium 3.5 128B 127.7B | 32K | 91.8 GB | 30.2 GB | 110K | 0.6–0.9 |
| Qwen3.5 122B-A10B 125.1B | 32K | 78.9 GB | 43.1 GB | 256K (full) | 4.5–8.2 |
| Nemotron 3 Super 120B-A12B 123.6B | 32K | 77.4 GB | 44.6 GB | 256K (full) | 4.1–7.5 |
| gpt-oss-120b 116.8B, MXFP4 | 32K | 68.6 GB | 53.4 GB | 128K (full) | 7.4–14 |
llama-server commands for a CPU-only PC with 128 GB of quad-channel DDR4-3200
GGUF repos checked 2026-09-29; -c is the longest context with 0.5 GB free on this PC. -ngl 0 keeps every layer in system RAM, even in a CUDA or Vulkan build of llama.cpp.
Mistral Medium 3.5 128B, Q4_K_M with 110K tokens: 0.4–0.7 tokens/s on this PC
llama-server -hf unsloth/Mistral-Medium-3.5-128B-GGUF:Q4_K_M -c 112640 -ngl 0 Q4_K_M/
Qwen3.5 122B-A10B, Q4_K_M with 256K tokens: 2.5–4.5 tokens/s on this PC
llama-server -hf unsloth/Qwen3.5-122B-A10B-GGUF:Q4_K_M -c 262144 -ngl 0 Q4_K_M/
How fast this PC writes as the context fills
Tokens per second at Q4_K_M (gpt-oss-120b at MXFP4) for the biggest models that fit with 32K tokens, with 8K and 32K tokens in the cache and at the longest context the PC holds, and at Q8_0 with 8K where that fits. Each token reads the active weights and the whole cache once, so the speed follows the 102.4 GB/s of bandwidth and falls as the cache grows. The ceiling is that bandwidth divided by the bytes read per token, which no runtime reaches; the reply time is for 1,000 new tokens with 32K (or the longest context) in the cache, prompt processing aside. With the same memory, CPU only, DDR5-5600 dual channel (128 GB), 89.6 GB/s, writes 12% slower on average.
| Model | 8K | 32K | Longest | At the longest | Q8_0 at 8K | Bandwidth ceiling at 8K | 1,000-token reply at 32K |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Flash Next | 7.9–15 | 6.9–13 | 256K | 3.0–5.6 | Does not fit | 27 | 79–146 s |
| Mistral Medium 3.5 128B | 0.6–1.0 | 0.6–0.9 | 110K | 0.4–0.7 | Does not fit | 1 | 1088–1741 s |
| Qwen3.5 122B-A10B | 4.9–9.0 | 4.5–8.2 | 256K | 2.5–4.5 | Does not fit | 16 | 122–225 s |
| Nemotron 3 Super 120B-A12B | 4.2–7.7 | 4.1–7.5 | 256K | 3.2–6.0 | Does not fit | 14 | 134–247 s |
| gpt-oss-120b | 9.5–18 | 7.4–14 | 128K | 4.0–7.3 | MXFP4 only | 32 | 73–134 s |
| AliceAI Foundation 80B-A3B | 15–28 | 12–21 | 256K | 3.7–6.8 | 8.9–16 | 51 | 47–87 s |
| Qwen3-Coder-Next | 15–28 | 12–21 | 256K | 3.7–6.8 | 8.9–16 | 51 | 47–87 s |
| Llama 3.1 70B | 1.1–1.8 | 1.0–1.5 | 128K | 0.6–1.0 | 0.7–1.1 | 2 | 653–1045 s |
| K2-Horizon MoVA 36B-A4B | 7.5–14 | 3.4–6.3 | 474K | 0.3–0.6 | 5.2–9.6 | 25 | 158–290 s |
| Ornith 1.5 35B-A3B | 15–28 | 12–22 | 256K | 4.2–7.8 | 9.0–17 | 52 | 45–82 s |
128 GB of DDR4-3200 (quad channel, CPU only) against the other CPU-only PCs with 122 GB usable
The same models fit on every CPU-only PC with 122 GB usable, so speed is what separates them. By bandwidth the CPU only, DDR4-3200 quad channel (128 GB) is the fastest of the 2 CPU-only PCs with 122 GB usable at 102.4 GB/s. Over the 12 biggest models that fit at Q4_K_M with 8K context, the CPU only, DDR5-5600 dual channel (128 GB), at 89.6 GB/s, writes 12% slower than the CPU only, DDR4-3200 quad channel (128 GB). Each cell below is the other CPU-only PC's tokens per second minus this one's, midpoints of the estimated ranges.
| Model | CPU only, DDR4-3200 quad channel (128 GB) tokens/s | CPU only, DDR5-5600 dual channel (128 GB) |
|---|---|---|
| Qwen3.8 Flash Next | 7.9–15 | −1.4 |
| Mistral Medium 3.5 128B | 0.6–1.0 | −0.1 |
| Qwen3.5 122B-A10B | 4.9–9.0 | −0.9 |
| Nemotron 3 Super 120B-A12B | 4.2–7.7 | −0.7 |
| gpt-oss-120b | 9.5–18 | −1.7 |
| AliceAI Foundation 80B-A3B | 15–28 | −2.6 |
| Qwen3-Coder-Next | 15–28 | −2.6 |
| Llama 3.1 70B | 1.1–1.8 | −0.2 |
| K2-Horizon MoVA 36B-A4B | 7.5–14 | −1.3 |
| Ornith 1.5 35B-A3B | 15–28 | −2.7 |
| Qwen3.6 35B-A3B | 15–28 | −2.7 |
| Ornith 1.0 35B | 15–28 | −2.7 |
Near misses on 128 GB of DDR4-3200 (quad channel, CPU only)
Models that miss at Q4_K_M with 32,768 tokens of context (or their whole window), or leave under 0.5 GB free there, but fit with a shorter context or 3-bit weights, or are short by at most a quarter of the memory and fit with MoE offload. Smallest shortfall first.
| Model | Short by at Q4_K_M | Fits instead |
|---|---|---|
| Step 3.7 Flash 196B-A11B 201.4B | 5.10 GB | GGUF Q3_K_M with 32K, 103 GB |
| MiniMax M2.7 228.7B | 28.8 GB | GGUF IQ3_XXS with 32K, 106 GB |
How these numbers are worked out
Memory is the weights at each precision plus the KV cache for 8,192 tokens in FP16 and runtime overhead (0.5 GB plus 10%), computed from each model's files on Hugging Face with the LLM VRAM Calculator. Speed is estimated from memory bandwidth and the parameters read per token with the LLM Speed Calculator. Both tools take any other model from Hugging Face.
Nearby GPUs and setups
How many of the tracked models each one holds at Q4_K_M with 8K context, against 43 here.
| Setup | How it relates | Memory | Bandwidth | Models at Q4_K_M |
|---|---|---|---|---|
| CPU only, DDR5-5600 dual channel (128 GB) | Same memory | 122 GB usable | 89.6 GB/s | 43 (same) |
| CPU only, DDR5-5600 dual channel (64 GB) | Next size down | 58 GB usable | 89.6 GB/s | 38 (−5) |
| CPU only, DDR4-3200 quad channel (256 GB) | Next size up | 250 GB usable | 102.4 GB/s | 50 (+7) |
Model numbers read from Hugging Face on .