Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD

Pick a Qwen3.8 27B file, the context length and KV cache type, and see the VRAM it takes, which cards fit, whether it runs on mainline llama.cpp or needs PrismML's fork, and the llama-server command.

– estimated VRAM

Model file
–
KV cache
–
mmproj
–
Buffers + CUDA
1.00 GB
GPU memory8 GB12 GB16 GB24 GB
Result––––

“Tight”: fits with under 0.5 GB to spare. Buffers are fixed at 1 GB; the RTX 4070 run below used about 0.9 GB more than this estimate at 131K filled context, so leave headroom, and about 0.5–1 GB more if the GPU drives your screen.

Engine: –

llama-server command

What the three formats are

Ternary Bonsai 2 (PrismML, 2026-09-16) is Qwen3.8 27B retrained to ternary weights (−1, 0, +1) with one FP16 scale per group of 128 and a Hadamard rotation. Its PTQ1_0 (1.75 bits) and PQ2_0 (2.13 bits) types run only in PrismML's llama.cpp fork. GSQ-RCO (ISTA-DASLab, 2026-08-28) stores the weights in standard GGUF types, IQ2_XS to IQ3_S, so mainline llama.cpp loads it; each size also comes as an -mtp file with the multi-token-prediction layer. Unsloth UD files are the usual dynamic quants, shown for comparison; they include the MTP layer.

Every Qwen3.8 27B file compared

Size: bytes on Hugging Face. Bits: as published, and file size ÷ 27.8B parameters (the -mtp files carry an extra layer). Score: ByteShape's normalized score vs BF16. PPL: ISTA-DASLab's self-reported wikitext perplexity (BF16 7.05, Unsloth UD-IQ2_S 8.02). Speed: ByteShape decode tokens/s. Total: file + 32K q8_0 KV cache + 1 GB. A dash means not published.

FileSizeBits (published)Bits (size ÷ params) EngineScorePPL RTX 4090 tok/sRTX 5060 Ti tok/sTotal, 32K q8_0
Bonsai 2 PTQ1_0 5.54 GB5,946,648,928 B 1.75 1.71 PrismML fork 91.44% — 95.34 45.81 7.60 GB
Bonsai 2 PTQ1_0-mtp 6.53 GB7,012,820,512 B 1.75 2.02 PrismML fork — — — — 8.59 GB
Bonsai 2 PQ2_0 6.71 GB7,206,168,928 B 2.13 2.08 PrismML fork 91.65% — 89.26 48.73 8.77 GB
GSQ-RCO IQ2_XS 7.84 GB8,422,841,472 B 2.50 2.43 mainline 86.47% 7.69 — — 9.91 GB
GSQ-RCO IQ2_XS-mtp 8.17 GB8,771,311,680 B 2.50 2.53 mainline — — — — 10.2 GB
GSQ-RCO IQ2_S 8.62 GB9,259,510,912 B 2.75 2.67 mainline 93.64% — — — 10.7 GB
GSQ-RCO IQ2_S-mtp 8.95 GB9,607,981,120 B 2.75 2.77 mainline — — — — 11.0 GB
GSQ-RCO IQ3_XXS 9.40 GB10,094,357,632 B 3.00 2.91 mainline 94.38% 7.20 69.26 33.95 11.5 GB
GSQ-RCO IQ3_XXS-mtp 9.73 GB10,442,827,840 B 3.00 3.01 mainline — — — — 11.8 GB
GSQ-RCO IQ3_S 11.0 GB11,771,546,784 B 3.50 3.39 mainline 99.43% 7.07 62.41 30.72 13.0 GB
GSQ-RCO IQ3_S-mtp 11.3 GB12,120,016,960 B 3.50 3.49 mainline — — — — 13.4 GB
Unsloth UD-IQ3_XXS 10.2 GB10,934,860,704 B — 3.15 mainline 93.59% — 67.24 32.80 12.2 GB
Unsloth UD-IQ4_XS 13.3 GB14,252,845,984 B — 4.10 mainline 99.20% — 55.62 — 15.3 GB
Unsloth UD-Q4_K_M 15.3 GB16,464,440,224 B — 4.74 mainline 97.03% — 49.77 — 17.4 GB

KL divergence vs BF16, measured independently

FileMean KLDSame top-1 tokenPPL
Bonsai 2 PQ2_00.3477.8%8.35
Unsloth UD-Q4_K_XL0.008796.9%6.34

Bonsai 2 is retrained rather than rounded from BF16, so a high KLD measures how far its answers differ from the original, not how good they are; ByteShape's task score above puts PQ2_0 at 91.65%. huggingface.co 2026-09-23

Which one should I pick

Totals: the whole file on the GPU, no mmproj, no MTP, 1 GB of buffers. Score: ByteShape's normalized score vs BF16.

8 GB
Bonsai 2 PTQ1_0 is the only file that fits with room for context: 5.54 GB of weights, 7.10 GB in total at 32K with q4_0 KV (7.60 GB with q8_0). It needs the PrismML fork. An RTX 2060 Super 8GB ran it at 32K in 6.2–6.5 GB at 16.8 tokens/s. PQ2_0 is tight even at 8K (7.85 GB), and the smallest GSQ-RCO file, IQ2_XS, needs 8.99 GB at 8K.
12 GB
On mainline llama.cpp, GSQ-RCO IQ3_XXS (9.40 GB, score 94.38%) takes 11.5 GB at 32K with q8_0 KV, or 11.5 GB at 64K with q4_0. Unsloth UD-IQ3_XXS is larger (10.2 GB) and scores 93.59%. GSQ-RCO IQ3_S needs 12.5 GB at 32K even with q4_0, over 12 GB. For long context, Bonsai 2 PTQ1_0 reaches 256K in 11.0 GB with q4_0 KV (an RTX 3060 12GB reported 11.7 GB at 262K) and PQ2_0 128K in 9.96 GB.
16 GB
GSQ-RCO IQ3_S (11.0 GB, score 99.43%, the highest here) takes 14.0 GB at 32K with f16 KV, 14.1 GB at 64K with q8_0 and 14.2 GB at 128K with q4_0. Unsloth UD-IQ4_XS (13.3 GB, 99.20%) takes 15.3 GB at 32K with q8_0. For 256K, GSQ-RCO IQ3_XXS needs 14.9 GB with q4_0 KV; an RTX 5060 Ti 16GB ran it at 240K context at about 30 tokens/s.
24 GB
Unsloth UD-Q4_K_M (15.3 GB, 97.03%) takes 20.3 GB at 64K with f16 KV, and UD-IQ4_XS 22.3 GB at 128K with f16. On an RTX 4090 ByteShape measured 49.77 tokens/s for UD-Q4_K_M, 55.62 for UD-IQ4_XS and 62.41 for GSQ-RCO IQ3_S, which also scores higher (99.43%).

Qwen3.8 Flash Next: GSQ-RCO shards

ISTA-DASLab's GSQ-RCO build of Qwen3.8 Flash Next (180B MoE, 2026-09-07) comes as two files per size: a weights shard and the same 26.8 GB n-gram shard (IQ4_NL). llama.cpp maps the n-gram shard and reads its rows on demand (--lazy-mode), so it sits on disk or in the page cache, not in VRAM. Even the smallest weights shard is larger than a 32 GB card, so part of the experts has to live in system RAM.

QuantWeights shardn-gram shardDownloadOllama
Q2_0 35.0 GB 26.8 GB 61.9 GB no (discussion #23)
IQ2_XS 36.5 GB 26.8 GB 63.4 GB not reported
IQ3_XXS 43.8 GB 26.8 GB 70.6 GB not reported
IQ3_S 51.1 GB 26.8 GB 77.9 GB not reported

ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF 2026-09-07 To size the split, use the MoE offload calculator: it finds the smallest --n-cpu-moe for your GPU.

Known issues

Bonsai 2 items are from PrismML's KNOWN_ISSUES.md (last checked there 2026-09-23) unless another source is linked.

  • Bonsai 2 PTQ1_0 and PQ2_0 (GGML types 143 and 142) need PrismML's fork, PrismML-Eng/llama.cpp (release prism-b10743, 2026-09-25). The pull request that would have added them to mainline llama.cpp, #29077, was closed unmerged on 2026-09-22; the feature request #29058 is still open. PrismML-Eng/llama.cpp · llama.cpp PR #29077 · issue #29058
  • Stock llama.cpp, Ollama and LM Studio (GGUF) cannot load PTQ1_0 or PQ2_0. The F16 Bonsai file loads in stock llama.cpp but gives garbled output, because it relies on metadata only the PrismML build applies. KNOWN_ISSUES.md
  • Quantizing another model to PTQ1_0 or PQ2_0 with llama-quantize gives garbage: without Prism's Hadamard rotation metadata the file loads without a warning and generates nonsense. KNOWN_ISSUES.md
  • PQ2_0 on Vulkan silently runs on the CPU in the binary release PrismML listed on 2026-09-23, at under 2 tokens/s, while the log still says every layer is offloaded. Fixed in source (PR #238); PTQ1_0 runs on the GPU. SYCL builds cannot run either type yet. KNOWN_ISSUES.md
  • PQ2_0 segfaults while loading on AVX-512 CPUs, including AMD Zen 4 and Zen 5, even with every layer on the GPU. Fixed in source (PR #245). KNOWN_ISSUES.md
  • ROCm/HIP aborts on consumer RDNA2 GPUs such as gfx1030; PrismML suggests the Vulkan build. CUDA builds from source may need -DGGML_CUDA_FA_ALL_QUANTS=ON for q8_0/q4_0 KV with flash attention. KNOWN_ISSUES.md
  • Which Bonsai type decodes faster depends on the GPU: PTQ1_0 on Ada and L4, PQ2_0 on Blackwell, Hopper, Ampere, Volta and Turing. PQ2_0 is also faster at prompt processing. KNOWN_ISSUES.md
  • Bonsai 2 PQ2_0 differs from BF16 far more than a 4-bit quant: mean KLD 0.340 and 77.8% matching top-1 tokens, against 0.0087 and 96.9% for Unsloth UD-Q4_K_XL. It is retrained, so this measures difference, not task skill. discussion #54
  • The GSQ-RCO -mtp files use llama.cpp's built-in MTP (--spec-type draft-mtp), added by PR #22673 on 2026-05-16; older builds cannot use the layer. The PrismML fork refused MTP for Bonsai files until PR #205 (2026-09-21); the grafted PTQ1_0-mtp file is a community upload. llama.cpp PR #22673
  • The Qwen3.8 Flash Next GSQ-RCO Q2_0 shard does not load in Ollama. discussion #23

What people measured

GPUFileContextKVEngineVRAMtok/sSource
RTX 4070 12GBBonsai 2 PTQ1_00 filledq4_0PrismML fork 7,561 MiB51.93 github.com 2026-09-20
RTX 4070 12GBBonsai 2 PTQ1_016K filledq4_0PrismML fork 7,721 MiB44.54 github.com 2026-09-20
RTX 4070 12GBBonsai 2 PTQ1_064K filledq4_0PrismML fork 8,586 MiB30.69 github.com 2026-09-20
RTX 4070 12GBBonsai 2 PTQ1_0131K filledq4_0PrismML fork 9,918 MiB21.77 github.com 2026-09-20
RTX 3060 12GBBonsai 2 PTQ1_064Knot statednot stated 7.3 GB— github.com 2026-09-19
RTX 3060 12GBBonsai 2 PTQ1_0262Knot statednot stated 11.7 GB— github.com 2026-09-19
RTX 3060 12GBBonsai 2 PTQ1_0not statednot statedstock build —26.32 github.com 2026-09-19
RTX 2060 Super 8GBBonsai 2 PTQ1_032Knot statednot stated 6.2–6.5 GB16.8 huggingface.co 2026-09-22
RTX 5060 Ti 16GBGSQ-RCO IQ3_XXS240Knot statedmainline llama.cpp —~30 huggingface.co 2026-09-01
RTX 5080 16GBGSQ-RCO IQ3_S-mtp262Knot statedllama.cpp fork ~15.36 GiB~60 at start huggingface.co 2026-09-11

File sizes read from the Hugging Face API on : prism-ml/Ternary-Bonsai-2-27B-gguf, sudoingX/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF, ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF, unsloth/Qwen3.8-27B-GGUF. Scores and speeds: ByteShape (2026-09-14). For other models and quant types, see the LLM VRAM calculator.

How to use

  1. Pick the file. Bonsai 2 files need PrismML's llama.cpp fork; GSQ-RCO and Unsloth UD run on mainline llama.cpp.
  2. Set the context length and KV cache type. q8_0 halves the cache and q4_0 cuts it to about a quarter.
  3. Choose whether to load the vision projector (mmproj) and, for files with an MTP layer, whether to turn on MTP.
  4. Read the total and the GPU table, then copy the llama-server command.

Frequently asked questions

Can Qwen3.8 27B run on an 8 GB GPU?

Only as Ternary Bonsai 2 PTQ1_0: the file is 5.54 GB, and with a 32K q4_0 KV cache and 1 GB of buffers the total is 7.10 GB. An RTX 2060 Super 8GB reported 6.2–6.5 GB at 32K and 16.8 tokens/s. It needs PrismML's llama.cpp fork.

Does Ternary Bonsai 2 work in Ollama, LM Studio or mainline llama.cpp?

No. PTQ1_0 and PQ2_0 are new GGML types that only PrismML's fork (PrismML-Eng/llama.cpp) implements. The pull request to add them to mainline, #29077, was closed unmerged on 2026-09-22, and Ollama and LM Studio (GGUF) cannot load the files.

Is GSQ-RCO better than Unsloth UD at the same size?

On ByteShape's normalized score, GSQ-RCO IQ3_XXS (9.40 GB) scores 94.38% and Unsloth UD-IQ3_XXS (10.2 GB) 93.59%. ISTA-DASLab's own wikitext perplexity is 7.69 for GSQ-RCO IQ2_XS and 8.02 for Unsloth UD-IQ2_S, against 7.05 for BF16. GSQ-RCO IQ3_S (11.0 GB) scores 99.43%, above UD-IQ4_XS (13.3 GB) at 99.20%.

How good is Bonsai 2 compared with a 4-bit quant?

ByteShape scores PTQ1_0 at 91.44% and PQ2_0 at 91.65% of BF16, below GSQ-RCO IQ2_S (93.64%). An independent test measured a mean KL divergence of 0.340 for PQ2_0 against 0.0087 for Unsloth UD-Q4_K_XL. Bonsai 2 is retrained, so KLD shows how different its answers are, not how good.

What does an -mtp file cost, and is GB here GB or GiB?

The GSQ-RCO -mtp files are 0.32 GB larger (IQ3_XXS: 9.40 GB, -mtp 9.73 GB) and add a small draft KV cache when you turn MTP on with --spec-type draft-mtp. GB means GiB (bytes ÷ 1024³), the unit nvidia-smi and GPU memory sizes use, as everywhere on this site.

More calculators

Updated