Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD
Pick a Qwen3.8 27B file, the context length and KV cache type, and see the VRAM it takes, which cards fit, whether it runs on mainline llama.cpp or needs PrismML's fork, and the llama-server command.
– estimated VRAM
- Model file
- –
- KV cache
- –
- mmproj
- –
- Buffers + CUDA
- 1.00 GB
| GPU memory | 8 GB | 12 GB | 16 GB | 24 GB |
|---|---|---|---|---|
| Result | – | – | – | – |
“Tight”: fits with under 0.5 GB to spare. Buffers are fixed at 1 GB; the RTX 4070 run below used about 0.9 GB more than this estimate at 131K filled context, so leave headroom, and about 0.5–1 GB more if the GPU drives your screen.
Engine: –
llama-server command
What the three formats are
Ternary Bonsai 2 (PrismML, 2026-09-16) is Qwen3.8 27B retrained to ternary weights (−1, 0, +1) with one FP16 scale per group of 128 and a Hadamard rotation. Its PTQ1_0 (1.75 bits) and PQ2_0 (2.13 bits) types run only in PrismML's llama.cpp fork. GSQ-RCO (ISTA-DASLab, 2026-08-28) stores the weights in standard GGUF types, IQ2_XS to IQ3_S, so mainline llama.cpp loads it; each size also comes as an -mtp file with the multi-token-prediction layer. Unsloth UD files are the usual dynamic quants, shown for comparison; they include the MTP layer.
Every Qwen3.8 27B file compared
Size: bytes on Hugging Face. Bits: as published, and file size ÷ 27.8B parameters (the -mtp files carry an extra layer). Score: ByteShape's normalized score vs BF16. PPL: ISTA-DASLab's self-reported wikitext perplexity (BF16 7.05, Unsloth UD-IQ2_S 8.02). Speed: ByteShape decode tokens/s. Total: file + 32K q8_0 KV cache + 1 GB. A dash means not published.
| File | Size | Bits (published) | Bits (size ÷ params) | Engine | Score | PPL | RTX 4090 tok/s | RTX 5060 Ti tok/s | Total, 32K q8_0 |
|---|---|---|---|---|---|---|---|---|---|
| Bonsai 2 PTQ1_0 | 5.54 GB5,946,648,928 B | 1.75 | 1.71 | PrismML fork | 91.44% | — | 95.34 | 45.81 | 7.60 GB |
| Bonsai 2 PTQ1_0-mtp | 6.53 GB7,012,820,512 B | 1.75 | 2.02 | PrismML fork | — | — | — | — | 8.59 GB |
| Bonsai 2 PQ2_0 | 6.71 GB7,206,168,928 B | 2.13 | 2.08 | PrismML fork | 91.65% | — | 89.26 | 48.73 | 8.77 GB |
| GSQ-RCO IQ2_XS | 7.84 GB8,422,841,472 B | 2.50 | 2.43 | mainline | 86.47% | 7.69 | — | — | 9.91 GB |
| GSQ-RCO IQ2_XS-mtp | 8.17 GB8,771,311,680 B | 2.50 | 2.53 | mainline | — | — | — | — | 10.2 GB |
| GSQ-RCO IQ2_S | 8.62 GB9,259,510,912 B | 2.75 | 2.67 | mainline | 93.64% | — | — | — | 10.7 GB |
| GSQ-RCO IQ2_S-mtp | 8.95 GB9,607,981,120 B | 2.75 | 2.77 | mainline | — | — | — | — | 11.0 GB |
| GSQ-RCO IQ3_XXS | 9.40 GB10,094,357,632 B | 3.00 | 2.91 | mainline | 94.38% | 7.20 | 69.26 | 33.95 | 11.5 GB |
| GSQ-RCO IQ3_XXS-mtp | 9.73 GB10,442,827,840 B | 3.00 | 3.01 | mainline | — | — | — | — | 11.8 GB |
| GSQ-RCO IQ3_S | 11.0 GB11,771,546,784 B | 3.50 | 3.39 | mainline | 99.43% | 7.07 | 62.41 | 30.72 | 13.0 GB |
| GSQ-RCO IQ3_S-mtp | 11.3 GB12,120,016,960 B | 3.50 | 3.49 | mainline | — | — | — | — | 13.4 GB |
| Unsloth UD-IQ3_XXS | 10.2 GB10,934,860,704 B | — | 3.15 | mainline | 93.59% | — | 67.24 | 32.80 | 12.2 GB |
| Unsloth UD-IQ4_XS | 13.3 GB14,252,845,984 B | — | 4.10 | mainline | 99.20% | — | 55.62 | — | 15.3 GB |
| Unsloth UD-Q4_K_M | 15.3 GB16,464,440,224 B | — | 4.74 | mainline | 97.03% | — | 49.77 | — | 17.4 GB |
KL divergence vs BF16, measured independently
| File | Mean KLD | Same top-1 token | PPL |
|---|---|---|---|
| Bonsai 2 PQ2_0 | 0.34 | 77.8% | 8.35 |
| Unsloth UD-Q4_K_XL | 0.0087 | 96.9% | 6.34 |
Bonsai 2 is retrained rather than rounded from BF16, so a high KLD measures how far its answers differ from the original, not how good they are; ByteShape's task score above puts PQ2_0 at 91.65%. huggingface.co 2026-09-23
Which one should I pick
Totals: the whole file on the GPU, no mmproj, no MTP, 1 GB of buffers. Score: ByteShape's normalized score vs BF16.
- 8 GB
- Bonsai 2 PTQ1_0 is the only file that fits with room for context: 5.54 GB of weights, 7.10 GB in total at 32K with q4_0 KV (7.60 GB with q8_0). It needs the PrismML fork. An RTX 2060 Super 8GB ran it at 32K in 6.2–6.5 GB at 16.8 tokens/s. PQ2_0 is tight even at 8K (7.85 GB), and the smallest GSQ-RCO file, IQ2_XS, needs 8.99 GB at 8K.
- 12 GB
- On mainline llama.cpp, GSQ-RCO IQ3_XXS (9.40 GB, score 94.38%) takes 11.5 GB at 32K with q8_0 KV, or 11.5 GB at 64K with q4_0. Unsloth UD-IQ3_XXS is larger (10.2 GB) and scores 93.59%. GSQ-RCO IQ3_S needs 12.5 GB at 32K even with q4_0, over 12 GB. For long context, Bonsai 2 PTQ1_0 reaches 256K in 11.0 GB with q4_0 KV (an RTX 3060 12GB reported 11.7 GB at 262K) and PQ2_0 128K in 9.96 GB.
- 16 GB
- GSQ-RCO IQ3_S (11.0 GB, score 99.43%, the highest here) takes 14.0 GB at 32K with f16 KV, 14.1 GB at 64K with q8_0 and 14.2 GB at 128K with q4_0. Unsloth UD-IQ4_XS (13.3 GB, 99.20%) takes 15.3 GB at 32K with q8_0. For 256K, GSQ-RCO IQ3_XXS needs 14.9 GB with q4_0 KV; an RTX 5060 Ti 16GB ran it at 240K context at about 30 tokens/s.
- 24 GB
- Unsloth UD-Q4_K_M (15.3 GB, 97.03%) takes 20.3 GB at 64K with f16 KV, and UD-IQ4_XS 22.3 GB at 128K with f16. On an RTX 4090 ByteShape measured 49.77 tokens/s for UD-Q4_K_M, 55.62 for UD-IQ4_XS and 62.41 for GSQ-RCO IQ3_S, which also scores higher (99.43%).
Qwen3.8 Flash Next: GSQ-RCO shards
ISTA-DASLab's GSQ-RCO build of Qwen3.8 Flash Next (180B MoE, 2026-09-07) comes as two files per size: a weights shard and the same 26.8 GB n-gram shard (IQ4_NL). llama.cpp maps the n-gram shard and reads its rows on demand
| Quant | Weights shard | n-gram shard | Download | Ollama |
|---|---|---|---|---|
| Q2_0 | 35.0 GB | 26.8 GB | 61.9 GB | no (discussion #23) |
| IQ2_XS | 36.5 GB | 26.8 GB | 63.4 GB | not reported |
| IQ3_XXS | 43.8 GB | 26.8 GB | 70.6 GB | not reported |
| IQ3_S | 51.1 GB | 26.8 GB | 77.9 GB | not reported |
ISTA-DASLab/
Known issues
Bonsai 2 items are from PrismML's KNOWN_ISSUES.md (last checked there 2026-09-23) unless another source is linked.
- Bonsai 2 PTQ1_0 and PQ2_0 (GGML types 143 and 142) need PrismML's fork, PrismML-Eng/
llama.cpp (release prism-b10743, 2026-09-25). The pull request that would have added them to mainline llama.cpp, #29077, was closed unmerged on 2026-09-22; the feature request #29058 is still open. PrismML-Eng/llama.cpp · llama.cpp PR #29077 · issue #29058 - Stock llama.cpp, Ollama and LM Studio (GGUF) cannot load PTQ1_0 or PQ2_0. The F16 Bonsai file loads in stock llama.cpp but gives garbled output, because it relies on metadata only the PrismML build applies. KNOWN_ISSUES.md
- Quantizing another model to PTQ1_0 or PQ2_0 with llama-quantize gives garbage: without Prism's Hadamard rotation metadata the file loads without a warning and generates nonsense. KNOWN_ISSUES.md
- PQ2_0 on Vulkan silently runs on the CPU in the binary release PrismML listed on 2026-09-23, at under 2 tokens/s, while the log still says every layer is offloaded. Fixed in source (PR #238); PTQ1_0 runs on the GPU. SYCL builds cannot run either type yet. KNOWN_ISSUES.md
- PQ2_0 segfaults while loading on AVX-512 CPUs, including AMD Zen 4 and Zen 5, even with every layer on the GPU. Fixed in source (PR #245). KNOWN_ISSUES.md
- ROCm/HIP aborts on consumer RDNA2 GPUs such as gfx1030; PrismML suggests the Vulkan build. CUDA builds from source may need -DGGML_
CUDA_ for q8_0/q4_0 KV with flash attention. KNOWN_ISSUES.mdFA_ ALL_ QUANTS=ON - Which Bonsai type decodes faster depends on the GPU: PTQ1_0 on Ada and L4, PQ2_0 on Blackwell, Hopper, Ampere, Volta and Turing. PQ2_0 is also faster at prompt processing. KNOWN_ISSUES.md
- Bonsai 2 PQ2_0 differs from BF16 far more than a 4-bit quant: mean KLD 0.340 and 77.8% matching top-1 tokens, against 0.0087 and 96.9% for Unsloth UD-Q4_K_XL. It is retrained, so this measures difference, not task skill. discussion #54
- The GSQ-RCO -mtp files use llama.cpp's built-in MTP
(--spec-type draft-mtp), added by PR #22673 on 2026-05-16; older builds cannot use the layer. The PrismML fork refused MTP for Bonsai files until PR #205 (2026-09-21); the grafted PTQ1_0-mtp file is a community upload. llama.cpp PR #22673 - The Qwen3.8 Flash Next GSQ-RCO Q2_0 shard does not load in Ollama. discussion #23
What people measured
| GPU | File | Context | KV | Engine | VRAM | tok/s | Source |
|---|---|---|---|---|---|---|---|
| RTX 4070 12GB | Bonsai 2 PTQ1_0 | 0 filled | q4_0 | PrismML fork | 7,561 MiB | 51.93 | github.com 2026-09-20 |
| RTX 4070 12GB | Bonsai 2 PTQ1_0 | 16K filled | q4_0 | PrismML fork | 7,721 MiB | 44.54 | github.com 2026-09-20 |
| RTX 4070 12GB | Bonsai 2 PTQ1_0 | 64K filled | q4_0 | PrismML fork | 8,586 MiB | 30.69 | github.com 2026-09-20 |
| RTX 4070 12GB | Bonsai 2 PTQ1_0 | 131K filled | q4_0 | PrismML fork | 9,918 MiB | 21.77 | github.com 2026-09-20 |
| RTX 3060 12GB | Bonsai 2 PTQ1_0 | 64K | not stated | not stated | 7.3 GB | — | github.com 2026-09-19 |
| RTX 3060 12GB | Bonsai 2 PTQ1_0 | 262K | not stated | not stated | 11.7 GB | — | github.com 2026-09-19 |
| RTX 3060 12GB | Bonsai 2 PTQ1_0 | not stated | not stated | stock build | — | 26.32 | github.com 2026-09-19 |
| RTX 2060 Super 8GB | Bonsai 2 PTQ1_0 | 32K | not stated | not stated | 6.2–6.5 GB | 16.8 | huggingface.co 2026-09-22 |
| RTX 5060 Ti 16GB | GSQ-RCO IQ3_XXS | 240K | not stated | mainline llama.cpp | — | ~30 | huggingface.co 2026-09-01 |
| RTX 5080 16GB | GSQ-RCO IQ3_S-mtp | 262K | not stated | llama.cpp fork | ~15.36 GiB | ~60 at start | huggingface.co 2026-09-11 |
File sizes read from the Hugging Face API on : prism-ml/
How to use
- Pick the file. Bonsai 2 files need PrismML's llama.cpp fork; GSQ-RCO and Unsloth UD run on mainline llama.cpp.
- Set the context length and KV cache type. q8_0 halves the cache and q4_0 cuts it to about a quarter.
- Choose whether to load the vision projector (mmproj) and, for files with an MTP layer, whether to turn on MTP.
- Read the total and the GPU table, then copy the llama-server command.
Frequently asked questions
Can Qwen3.8 27B run on an 8 GB GPU?
Only as Ternary Bonsai 2 PTQ1_0: the file is 5.54 GB, and with a 32K q4_0 KV cache and 1 GB of buffers the total is 7.10 GB. An RTX 2060 Super 8GB reported 6.2–6.5 GB at 32K and 16.8 tokens/s. It needs PrismML's llama.cpp fork.
Does Ternary Bonsai 2 work in Ollama, LM Studio or mainline llama.cpp?
No. PTQ1_0 and PQ2_0 are new GGML types that only PrismML's fork (PrismML-Eng/
Is GSQ-RCO better than Unsloth UD at the same size?
On ByteShape's normalized score, GSQ-RCO IQ3_XXS (9.40 GB) scores 94.38% and Unsloth UD-IQ3_XXS (10.2 GB) 93.59%. ISTA-DASLab's own wikitext perplexity is 7.69 for GSQ-RCO IQ2_XS and 8.02 for Unsloth UD-IQ2_S, against 7.05 for BF16. GSQ-RCO IQ3_S (11.0 GB) scores 99.43%, above UD-IQ4_XS (13.3 GB) at 99.20%.
How good is Bonsai 2 compared with a 4-bit quant?
ByteShape scores PTQ1_0 at 91.44% and PQ2_0 at 91.65% of BF16, below GSQ-RCO IQ2_S (93.64%). An independent test measured a mean KL divergence of 0.340 for PQ2_0 against 0.0087 for Unsloth UD-Q4_K_XL. Bonsai 2 is retrained, so KLD shows how different its answers are, not how good.
What does an -mtp file cost, and is GB here GB or GiB?
The GSQ-RCO -mtp files are 0.32 GB larger (IQ3_XXS: 9.40 GB, -mtp 9.73 GB) and add a small draft KV cache when you turn MTP on with --spec-type draft-mtp. GB means GiB (bytes ÷ 1024³), the unit nvidia-smi and GPU memory sizes use, as everywhere on this site.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open → - MoE Offload Calculator
(--n-cpu-moe) The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes. Open → - Qwen-Image-2.1 VRAM Calculator Peak VRAM for every DiT, text encoder and VAE combination, with real measurements. Open →
Updated