Ternary Bonsai 2 27B VRAM Requirements on 8, 12 and 16 GB GPUs
Ternary Bonsai 2 27B runs on an 8 GB GPU: its 5.54 GB PTQ1_0 file needs 7.10 GB of VRAM with 32K context and a q4_0 KV cache, a 12 GB card holds the full 256K context in 11.0 GB, and a 16 GB card runs the larger PQ2_0 file at 256K in 12.2 GB. It runs only in PrismML's llama.cpp fork, not in Ollama, LM Studio or mainline llama.cpp.
Which GPUs run Bonsai 2 27B, and at what context
Each cell is the longest context that fits the card with at least 0.5 GB to spare, and the VRAM it takes: the whole file on the GPU
| File, KV cache | 8 GB | 12 GB | 16 GB | 24 GB |
|---|---|---|---|---|
| PTQ1_0, f16 | 8K · 7.04 GB | 64K · 10.5 GB | 128K · 14.5 GB | 256K · 22.5 GB |
| PTQ1_0, q8_0 | 16K · 7.07 GB | 128K · 10.8 GB | 256K · 15.0 GB | 256K · 15.0 GB |
| PTQ1_0, q4_0 | 32K · 7.10 GB | 256K · 11.0 GB | 256K · 11.0 GB | 256K · 11.0 GB |
| PQ2_0, f16 | 4K · 7.96 GB tight: under 0.5 GB to spare | 32K · 9.71 GB | 64K · 11.7 GB | 128K · 15.7 GB |
| PQ2_0, q8_0 | 8K · 7.98 GB tight: under 0.5 GB to spare | 64K · 9.84 GB | 128K · 12.0 GB | 256K · 16.2 GB |
| PQ2_0, q4_0 | 16K · 7.99 GB tight: under 0.5 GB to spare | 128K · 9.96 GB | 256K · 12.2 GB | 256K · 12.2 GB |
With the vision projector loaded (the 0.59 GB Q8_0 file, for image input) the same cards hold: 8 GB, PTQ1_0 at 16K (7.41 GB); 12 GB, PTQ1_0 or PQ2_0 at 128K (9.37 GB and 10.5 GB); 16 GB, both at 256K (11.6 GB and 12.8 GB), all with a q4_0 KV cache.
Between "fits" and "does not fit": PTQ1_0 at 64K with q4_0 takes 7.66 GB, which loads on an 8 GB card only when nothing else uses it (no desktop on that GPU). A 24 GB card runs every file at the full 256K, even with an f16 cache for PTQ1_0 (22.5 GB).
Every Bonsai 2 27B file, in bytes
Bytes as the Hugging Face API reports them. "Decimal GB" divides by 1,000,000,000; the site's GB (like nvidia-smi) divides by 1024³, which is why the same file reads 5.95 GB in one place and 5.54 GB here.
| File | Bytes | Decimal GB | GB (GiB) | Bits/weight | Runs in |
|---|---|---|---|---|---|
Ternary-Bonsai-2-27B-PTQ1_0.ggufSmallest file, the one for 8 GB | 5,946,648,928 | 5.95 GB | 5.54 GB | 1.75 | PrismML llama.cpp fork |
Ternary-Bonsai-2-27B-PQ2_0.ggufLarger, slightly higher score | 7,206,168,928 | 7.21 GB | 6.71 GB | 2.13 | PrismML llama.cpp fork |
Ternary-Bonsai-2-27B-PTQ1_0-mtp.ggufPTQ1_0 with the MTP layer grafted back (community) | 7,012,820,512 | 7.01 GB | 6.53 GB | 1.75 | PrismML llama.cpp fork |
Ternary-Bonsai-2-27B-mmproj-Q8_0.ggufVision projector | 629,246,976 | 0.63 GB | 0.59 GB | Q8_0 | Standard GGUF, loaded by the fork with the model |
Ternary-Bonsai-2-27B-mmproj-BF16.ggufVision projector | 931,145,856 | 0.93 GB | 0.87 GB | BF16 | Standard GGUF, loaded by the fork with the model |
Ternary-Bonsai-2-27B-F16.ggufUnquantized reference; stock llama.cpp loads it but the output is garbled | 53,808,408,928 | 53.81 GB | 50.1 GB | 16 | PrismML llama.cpp fork |
model.safetensorsApple Silicon pack; MLX apps and LM Studio (MLX) cannot load it yet | 8,595,477,990 | 8.60 GB | 8.01 GB | 2.25 | Bundled MLX runtime |
Why you see 5.95 GB, 7.2 GB and 8.60 GB
5.95 GB is the PTQ1_0 GGUF, 5,946,648,928 bytes, in decimal gigabytes, which is how the model card's own table lists it (and PQ2_0 as 7.21 GB). In GiB, the unit GPU memory is sold and reported in, it is 5.54 GB. "Under 6 GB" and "5.54 GB" describe the same file.
7.2 GB is the PQ2_0 GGUF (7,206,168,928 bytes, 6.71 GB), the 2.13-bit file; the first community MTP grafts were built on it.
8.60 GB is a different download: the MLX 2-bit pack for Apple Silicon, one model.safetensors of 8,595,477,990 bytes (8.01 GB). Its README splits it into a 7.67 GB language model and a 0.92 GB vision tower, and explains the larger size: MLX stores a scale and a bias per group of 128 weights, so the same ternary weights cost 2.25 bits each instead of 1.75 in the GGUF PTQ1_0 packing. Articles that quote 8.60 GB for "Bonsai 2 27B" are describing the Mac pack, not the file a PC with an NVIDIA card downloads.
How to run it
PTQ1_0 and PQ2_0 are new GGML tensor types that only PrismML's llama.cpp fork implements (release prism-b10743, 2026-09-25). The pull request that would add them to mainline, #29077, was closed unmerged on 2026-09-22, so Ollama, LM Studio and stock llama.cpp builds cannot load these files. With the fork, the commands are the usual llama-server ones:
huggingface-cli download prism-ml/Ternary-Bonsai-2-27B-gguf Ternary-Bonsai-2-27B-PTQ1_0.gguf --local-dir . 8 GB card, 32K context:
llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 32768 -fa on -ctk q4_0 -ctv q4_0 12 GB card, full 256K context:
llama-server -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -c 262144 -fa on -ctk q4_0 -ctv q4_0 16 GB card, PQ2_0 at 256K:
llama-server -m Ternary-Bonsai-2-27B-PQ2_0.gguf -ngl 99 -c 262144 -fa on -ctk q4_0 -ctv q4_0 The community file with the MTP layer grafted back (sudoingX/--spec-type draft-mtp. Its author measured 25.0 tokens/s without and 27.1 with one draft token on an RTX 3060 12GB at 128K.
On a Mac, llama.cpp from the fork runs the same GGUF files through Metal. The MLX pack is 8.01 GB of weights; with this site's method and an f16 cache that is about 11.0 GB at 32K and 17.0 GB at 128K. Its PACK-RUNTIME.md says ordinary MLX loaders do not apply the required Hadamard transforms, so use the runtime bundled in the repo, and that vision and MTP are not included, although the README counts a vision tower in the 8.60 GB. See What can my PC run for how much of a Mac's memory the GPU may use.
What people measured
Units as each source wrote them. Filled context means the prompt actually used that many tokens.
| GPU | Context | KV | VRAM | Tokens/s | Source |
|---|---|---|---|---|---|
| RTX 4070 12GB | 0 filled | q4_0 | 7,561 MiB | 51.93 | github.com 2026-09-20 |
| RTX 4070 12GB | 16K filled | q4_0 | 7,721 MiB | 44.54 | github.com 2026-09-20 |
| RTX 4070 12GB | 64K filled | q4_0 | 8,586 MiB | 30.69 | github.com 2026-09-20 |
| RTX 4070 12GB | 131K filled | q4_0 | 9,918 MiB | 21.77 | github.com 2026-09-20 |
| RTX 3060 12GB | 64K | not stated | 7.3 GB | — | github.com 2026-09-19 |
| RTX 3060 12GB | 262K | not stated | 11.7 GB | — | github.com 2026-09-19 |
| RTX 3060 12GB | not stated | not stated | — | 26.32 | github.com 2026-09-19 |
| RTX 2060 Super 8GB | 32K | not stated | 6.2–6.5 GB | 16.8 | huggingface.co 2026-09-22 |
The estimate is close for short prompts and low at long ones: at 131K filled context the RTX 4070 run used 9,918 MiB, against 8.79 GB here, so leave about 1 GB more headroom for very long prompts.
File sizes read from the Hugging Face API on : prism-ml/
Frequently asked questions
Can Ternary Bonsai 2 27B run on an 8 GB GPU?
Yes, as the PTQ1_0 file: 5.54 GB of weights, 7.10 GB in total with a 32K q4_0 KV cache and 1 GB of buffers. An RTX 2060 Super 8GB reported 6.2–6.5 GB at 32K and 16.8 tokens/s. PQ2_0 is only a tight fit on 8 GB: 7.85 GB at 8K, under 0.5 GB to spare.
Is Bonsai 2 27B 5.95 GB or 8.60 GB?
Both, for different files. 5.95 GB is the PTQ1_0 GGUF, 5,946,648,928 bytes, in decimal gigabytes (5.54 GiB). 8.60 GB is the MLX 2-bit pack for Apple Silicon, 8,595,477,990 bytes: MLX stores a scale and a bias per group, so each weight costs 2.25 bits instead of 1.75.
How much context fits on a 12 GB or 16 GB GPU?
On 12 GB, PTQ1_0 reaches the full 256K with a q4_0 cache (11.0 GB) and PQ2_0 128K (9.96 GB). On 16 GB, PQ2_0 reaches 256K with q4_0 (12.2 GB) and PTQ1_0 256K even with q8_0 (15.0 GB). Only 16 of the 64 layers keep a KV cache, so long context is cheap.
Does Bonsai 2 work in Ollama or LM Studio?
Not yet. PTQ1_0 and PQ2_0 are implemented only in PrismML's llama.cpp fork; the pull request to add them to mainline, #29077, was closed unmerged on 2026-09-22. Ollama and LM Studio build on mainline llama.cpp and cannot load the files.
Is GB here GB or GiB?
GiB (bytes ÷ 1024³), the unit nvidia-smi and GPU memory sizes use, as everywhere on this site. The file table also lists decimal GB, to match the numbers other sites quote.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- Image & Video Model VRAM Calculator Peak VRAM and system RAM for FLUX.2, Wan 2.2, LTX-2 and Qwen-Image in ComfyUI, from exact file sizes, checked against public runs. Open →
- Jev Alternatives You Can Run Locally Open models that replace the API-only Jev on your own machine: Laya, zero-shot encoders and small LLMs, with sizes and VRAM. Open →
- Laya VRAM Requirements Real file sizes and run-time memory of the three Laya decision-encoder checkpoints, and three ways to run them. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
Updated