Ternary Bonsai 2 27B VRAM Requirements on 8, 12 and 16 GB GPUs

Ternary Bonsai 2 27B runs on an 8 GB GPU: its 5.54 GB PTQ1_0 file needs 7.10 GB of VRAM with 32K context and a q4_0 KV cache, a 12 GB card holds the full 256K context in 11.0 GB, and a 16 GB card runs the larger PQ2_0 file at 256K in 12.2 GB. It runs only in PrismML's llama.cpp fork, not in Ollama, LM Studio or mainline llama.cpp.

Which GPUs run Bonsai 2 27B, and at what context

Each cell is the longest context that fits the card with at least 0.5 GB to spare, and the VRAM it takes: the whole file on the GPU (-ngl 99), the KV cache of the 16 attention layers (the other 48 layers keep a fixed-size state), and 1 GB of compute buffers. Text only, no MTP.

File, KV cache8 GB12 GB16 GB24 GB
PTQ1_0, f16 8K · 7.04 GB64K · 10.5 GB128K · 14.5 GB256K · 22.5 GB
PTQ1_0, q8_0 16K · 7.07 GB128K · 10.8 GB256K · 15.0 GB256K · 15.0 GB
PTQ1_0, q4_0 32K · 7.10 GB256K · 11.0 GB256K · 11.0 GB256K · 11.0 GB
PQ2_0, f16 4K · 7.96 GB
tight: under 0.5 GB to spare
32K · 9.71 GB64K · 11.7 GB128K · 15.7 GB
PQ2_0, q8_0 8K · 7.98 GB
tight: under 0.5 GB to spare
64K · 9.84 GB128K · 12.0 GB256K · 16.2 GB
PQ2_0, q4_0 16K · 7.99 GB
tight: under 0.5 GB to spare
128K · 9.96 GB256K · 12.2 GB256K · 12.2 GB

With the vision projector loaded (the 0.59 GB Q8_0 file, for image input) the same cards hold: 8 GB, PTQ1_0 at 16K (7.41 GB); 12 GB, PTQ1_0 or PQ2_0 at 128K (9.37 GB and 10.5 GB); 16 GB, both at 256K (11.6 GB and 12.8 GB), all with a q4_0 KV cache.

Between "fits" and "does not fit": PTQ1_0 at 64K with q4_0 takes 7.66 GB, which loads on an 8 GB card only when nothing else uses it (no desktop on that GPU). A 24 GB card runs every file at the full 256K, even with an f16 cache for PTQ1_0 (22.5 GB).

Every Bonsai 2 27B file, in bytes

Bytes as the Hugging Face API reports them. "Decimal GB" divides by 1,000,000,000; the site's GB (like nvidia-smi) divides by 1024³, which is why the same file reads 5.95 GB in one place and 5.54 GB here.

FileBytesDecimal GBGB (GiB) Bits/weightRuns in
Ternary-Bonsai-2-27B-PTQ1_0.gguf
Smallest file, the one for 8 GB
5,946,648,928 5.95 GB 5.54 GB 1.75 PrismML llama.cpp fork
Ternary-Bonsai-2-27B-PQ2_0.gguf
Larger, slightly higher score
7,206,168,928 7.21 GB 6.71 GB 2.13 PrismML llama.cpp fork
Ternary-Bonsai-2-27B-PTQ1_0-mtp.gguf
PTQ1_0 with the MTP layer grafted back (community)
7,012,820,512 7.01 GB 6.53 GB 1.75 PrismML llama.cpp fork
Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf
Vision projector
629,246,976 0.63 GB 0.59 GB Q8_0 Standard GGUF, loaded by the fork with the model
Ternary-Bonsai-2-27B-mmproj-BF16.gguf
Vision projector
931,145,856 0.93 GB 0.87 GB BF16 Standard GGUF, loaded by the fork with the model
Ternary-Bonsai-2-27B-F16.gguf
Unquantized reference; stock llama.cpp loads it but the output is garbled
53,808,408,928 53.81 GB 50.1 GB 16 PrismML llama.cpp fork
model.safetensors
Apple Silicon pack; MLX apps and LM Studio (MLX) cannot load it yet
8,595,477,990 8.60 GB 8.01 GB 2.25 Bundled MLX runtime

Why you see 5.95 GB, 7.2 GB and 8.60 GB

5.95 GB is the PTQ1_0 GGUF, 5,946,648,928 bytes, in decimal gigabytes, which is how the model card's own table lists it (and PQ2_0 as 7.21 GB). In GiB, the unit GPU memory is sold and reported in, it is 5.54 GB. "Under 6 GB" and "5.54 GB" describe the same file.

7.2 GB is the PQ2_0 GGUF (7,206,168,928 bytes, 6.71 GB), the 2.13-bit file; the first community MTP grafts were built on it.

8.60 GB is a different download: the MLX 2-bit pack for Apple Silicon, one model.safetensors of 8,595,477,990 bytes (8.01 GB). Its README splits it into a 7.67 GB language model and a 0.92 GB vision tower, and explains the larger size: MLX stores a scale and a bias per group of 128 weights, so the same ternary weights cost 2.25 bits each instead of 1.75 in the GGUF PTQ1_0 packing. Articles that quote 8.60 GB for "Bonsai 2 27B" are describing the Mac pack, not the file a PC with an NVIDIA card downloads.

How to run it

PTQ1_0 and PQ2_0 are new GGML tensor types that only PrismML's llama.cpp fork implements (release prism-b10743, 2026-09-25). The pull request that would add them to mainline, #29077, was closed unmerged on 2026-09-22, so Ollama, LM Studio and stock llama.cpp builds cannot load these files. With the fork, the commands are the usual llama-server ones:

huggingface-cli download prism-ml/Ternary-Bonsai-2-27B-gguf Ternary-Bonsai-2-27B-PTQ1_0.gguf --local-dir .

8 GB card, 32K context:

llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 32768 -fa on -ctk q4_0 -ctv q4_0

12 GB card, full 256K context:

llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 262144 -fa on -ctk q4_0 -ctv q4_0

16 GB card, PQ2_0 at 256K:

llama-server -m Ternary-Bonsai-2-27B-PQ2_0.gguf -ngl 99 -c 262144 -fa on -ctk q4_0 -ctv q4_0

The community file with the MTP layer grafted back (sudoingX/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF, 6.53 GB) adds 0.99 GB and lets the fork run --spec-type draft-mtp. Its author measured 25.0 tokens/s without and 27.1 with one draft token on an RTX 3060 12GB at 128K.

On a Mac, llama.cpp from the fork runs the same GGUF files through Metal. The MLX pack is 8.01 GB of weights; with this site's method and an f16 cache that is about 11.0 GB at 32K and 17.0 GB at 128K. Its PACK-RUNTIME.md says ordinary MLX loaders do not apply the required Hadamard transforms, so use the runtime bundled in the repo, and that vision and MTP are not included, although the README counts a vision tower in the 8.60 GB. See What can my PC run for how much of a Mac's memory the GPU may use.

What people measured

Units as each source wrote them. Filled context means the prompt actually used that many tokens.

GPUContextKVVRAMTokens/sSource
RTX 4070 12GB0 filledq4_07,561 MiB51.93 github.com 2026-09-20
RTX 4070 12GB16K filledq4_07,721 MiB44.54 github.com 2026-09-20
RTX 4070 12GB64K filledq4_08,586 MiB30.69 github.com 2026-09-20
RTX 4070 12GB131K filledq4_09,918 MiB21.77 github.com 2026-09-20
RTX 3060 12GB64Knot stated7.3 GB— github.com 2026-09-19
RTX 3060 12GB262Knot stated11.7 GB— github.com 2026-09-19
RTX 3060 12GBnot statednot stated—26.32 github.com 2026-09-19
RTX 2060 Super 8GB32Knot stated6.2–6.5 GB16.8 huggingface.co 2026-09-22

The estimate is close for short prompts and low at long ones: at 131K filled context the RTX 4070 run used 9,918 MiB, against 8.79 GB here, so leave about 1 GB more headroom for very long prompts.

File sizes read from the Hugging Face API on : prism-ml/Ternary-Bonsai-2-27B-gguf (3,581,027 downloads, 2,263 likes, Apache 2.0), prism-ml/Ternary-Bonsai-2-27B-mlx-2bit and sudoingX/Ternary-Bonsai-2-27B-PTQ1_0-MTP-GGUF. The architecture is Qwen3.8 27B's. To compare Bonsai 2 with GSQ-RCO and Unsloth UD quants of the same model, or to try any context, KV type and projector setting, open the Qwen3.8 27B quant calculator.

Frequently asked questions

Can Ternary Bonsai 2 27B run on an 8 GB GPU?

Yes, as the PTQ1_0 file: 5.54 GB of weights, 7.10 GB in total with a 32K q4_0 KV cache and 1 GB of buffers. An RTX 2060 Super 8GB reported 6.2–6.5 GB at 32K and 16.8 tokens/s. PQ2_0 is only a tight fit on 8 GB: 7.85 GB at 8K, under 0.5 GB to spare.

Is Bonsai 2 27B 5.95 GB or 8.60 GB?

Both, for different files. 5.95 GB is the PTQ1_0 GGUF, 5,946,648,928 bytes, in decimal gigabytes (5.54 GiB). 8.60 GB is the MLX 2-bit pack for Apple Silicon, 8,595,477,990 bytes: MLX stores a scale and a bias per group, so each weight costs 2.25 bits instead of 1.75.

How much context fits on a 12 GB or 16 GB GPU?

On 12 GB, PTQ1_0 reaches the full 256K with a q4_0 cache (11.0 GB) and PQ2_0 128K (9.96 GB). On 16 GB, PQ2_0 reaches 256K with q4_0 (12.2 GB) and PTQ1_0 256K even with q8_0 (15.0 GB). Only 16 of the 64 layers keep a KV cache, so long context is cheap.

Does Bonsai 2 work in Ollama or LM Studio?

Not yet. PTQ1_0 and PQ2_0 are implemented only in PrismML's llama.cpp fork; the pull request to add them to mainline, #29077, was closed unmerged on 2026-09-22. Ollama and LM Studio build on mainline llama.cpp and cannot load the files.

Is GB here GB or GiB?

GiB (bytes ÷ 1024³), the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The file table also lists decimal GB, to match the numbers other sites quote.

More calculators

Updated