Laya VRAM Requirements: Every Checkpoint, File and Runtime
Laya is a ModernBERT-based decision encoder, not a chat LLM: it takes a text and typed questions and returns calibrated probabilities in one forward pass. Pick the checkpoint, the file and the input size to see the peak memory and which GPUs fit.
– estimated peak memory
Weights: – the size of the file; the rest is activations for one forward pass plus 0.5 GB for the runtime.
| GPU memory | 4 GB | 6 GB | 8 GB | 12 GB | 16 GB | 24 GB |
|---|---|---|---|---|---|---|
| Result | – | – | – | – | – | – |
With no GPU the same amount comes out of system RAM: –.
llama.cpp only: the Python process that runs the decision head also loads the original checkpoint, about – of system RAM on top.
"Tight" means it fits only if the runtime uses a fused attention kernel, which never stores the full attention-score matrix.
Can a 16 GB or 24 GB GPU, or a CPU, run Laya?
Yes, all three. The largest published file is 850 MB, and with the default input length the peak is 0.76 GB–1.40 GB for every file, so a 16 GB or 24 GB card uses less than a tenth of its memory, and a 4 GB card is enough. With no GPU the model runs from system RAM, and 3 GB free is enough even for the llama.cpp setup, which loads the model twice. What changes with the hardware is speed, not whether it runs: tens of milliseconds per request on a GPU, a few hundred on a CPU (see the measurements below).
Every Laya file: size and memory
Each checkpoint at its default input length (512 tokens for English, 1,024 for typed-decisions and multilingual), one row per pass. Size is the published file in MB as Hugging Face lists it; memory is in GB (GiB).
| Checkpoint | Runtime | Format | File | Size | Peak memory | 16 GB | 24 GB | CPU only (RAM) |
|---|---|---|---|---|---|---|---|---|
| Laya (English) | PyTorch (laya package) | F16 | model.safetensors | 843 MB | 1.30 GB–1.30 GB | Fits | Fits | 1.30 GB |
| Laya (English) | ggmlc laya binary | F16 | laya_english_f16.gguf | 846 MB | 1.31 GB–1.33 GB | Fits | Fits | 1.33 GB |
| Laya (English) | ggmlc laya binary | Q8_0 | laya_english_q8_0.gguf | 452 MB | 0.95 GB–0.96 GB | Fits | Fits | 0.96 GB |
| Laya (English) | ggmlc laya binary | UD_Q4_K_M | laya_english_ud_q4_k_m.gguf | 420 MB | 0.92 GB–0.93 GB | Fits | Fits | 0.93 GB |
| Laya (English) | llama.cpp + head | F16 | laya-F16.gguf | 791 MB plus decision head 106 MB | 1.26 GB–1.28 GB | Fits | Fits | 2.66 GB |
| Laya (English) | llama.cpp + head | Q8_0 | laya-Q8_0.gguf | 421 MB plus decision head 106 MB | 0.92 GB–0.93 GB | Fits | Fits | 2.32 GB |
| Laya (English) | llama.cpp + head | Q6_K | laya-Q6_K.gguf | 344 MB plus decision head 106 MB | 0.85 GB–0.86 GB | Fits | Fits | 2.24 GB |
| Laya (English) | llama.cpp + head | Q4_K_M | laya-Q4_K_M.gguf | 272 MB plus decision head 106 MB | 0.78 GB–0.79 GB | Fits | Fits | 2.18 GB |
| Laya typed-decisions | PyTorch (laya package) | F16 | model.safetensors | 843 MB | 1.31 GB–1.34 GB | Fits | Fits | 1.34 GB |
| Laya typed-decisions | ggmlc laya binary | F16 | laya_typed_decisions_f16.gguf | 850 MB | 1.34 GB–1.40 GB | Fits | Fits | 1.40 GB |
| Laya typed-decisions | ggmlc laya binary | Q8_0 | laya_typed_decisions_q8_0.gguf | 455 MB | 0.97 GB–1.04 GB | Fits | Fits | 1.04 GB |
| Laya typed-decisions | ggmlc laya binary | UD_Q4_K_M | laya_typed_decisions_ud_q4_k_m.gguf | 424 MB | 0.94 GB–1.01 GB | Fits | Fits | 1.01 GB |
| Laya typed-decisions | llama.cpp + head | F16 | laya-typed-decisions-F16.gguf | 791 MB plus decision head 106 MB | 1.29 GB–1.35 GB | Fits | Fits | 2.73 GB |
| Laya typed-decisions | llama.cpp + head | Q8_0 | laya-typed-decisions-Q8_0.gguf | 421 MB plus decision head 106 MB | 0.94 GB–1.00 GB | Fits | Fits | 2.39 GB |
| Laya typed-decisions | llama.cpp + head | Q6_K | laya-typed-decisions-Q6_K.gguf | 344 MB plus decision head 106 MB | 0.87 GB–0.93 GB | Fits | Fits | 2.32 GB |
| Laya typed-decisions | llama.cpp + head | Q4_K_M | laya-typed-decisions-Q4_K_M.gguf | 272 MB plus decision head 106 MB | 0.80 GB–0.87 GB | Fits | Fits | 2.25 GB |
| Laya multilingual | PyTorch (laya package) | F16 | model.safetensors | 644 MB | 1.11 GB–1.14 GB | Fits | Fits | 1.14 GB |
| Laya multilingual | ggmlc laya binary | F16 | laya_multilingual_f16.gguf | 663 MB | 1.15 GB–1.19 GB | Fits | Fits | 1.19 GB |
| Laya multilingual | ggmlc laya binary | Q8_0 | laya_multilingual_q8_0.gguf | 362 MB | 0.86 GB–0.91 GB | Fits | Fits | 0.91 GB |
| Laya multilingual | ggmlc laya binary | UD_Q4_K_M | laya_multilingual_ud_q4_k_m.gguf | 524 MB | 1.02 GB–1.06 GB | Fits | Fits | 1.06 GB |
| Laya multilingual | llama.cpp + head | F16 | laya-multilingual-F16.gguf | 629 MB plus decision head 60 MB | 1.11 GB–1.16 GB | Fits | Fits | 2.32 GB |
| Laya multilingual | llama.cpp + head | Q8_0 | laya-multilingual-Q8_0.gguf | 341 MB plus decision head 60 MB | 0.85 GB–0.89 GB | Fits | Fits | 2.05 GB |
| Laya multilingual | llama.cpp + head | Q6_K | laya-multilingual-Q6_K.gguf | 271 MB plus decision head 60 MB | 0.78 GB–0.83 GB | Fits | Fits | 1.98 GB |
| Laya multilingual | llama.cpp + head | Q4_K_M | laya-multilingual-Q4_K_M.gguf | 249 MB plus decision head 60 MB | 0.76 GB–0.81 GB | Fits | Fits | 1.96 GB |
Longer inputs and bigger batches
Memory grows with rows × tokens, and the attention scores with the square of the length. Even the heaviest case here, laya-multilingual reading one 8,192-token document, stays under 5 GB.
| Workload | Peak memory | 4 GB | 6 GB | 8 GB | 16 GB | 24 GB |
|---|---|---|---|---|---|---|
| Laya English, PyTorch, 32 rows × 512 tokens | 1.68 GB–1.93 GB | Fits | Fits | Fits | Fits | Fits |
| Laya English, PyTorch, 64 rows × 512 tokens | 2.08 GB–2.58 GB | Fits | Fits | Fits | Fits | Fits |
| multilingual, PyTorch, 1 document × 8,192 tokens | 1.21 GB–2.71 GB | Fits | Fits | Fits | Fits | Fits |
| multilingual, ggmlc F16, 1 document × 8,192 tokens | 1.34 GB–4.34 GB | Tight | Fits | Fits | Fits | Fits |
Three ways to run Laya locally
1. PyTorch: the laya Python package
The reference way, and the only one that runs all three checkpoints and routes each request to the right one. It downloads the safetensors files on first use and runs on CUDA, Apple GPUs or the CPU. PyPI, GitHub.
pip install laya from laya import Router
router = Router() # Router(preload=True) loads all three checkpoints up front
result = router.predict(
"We were billed twice for March. Please refund the duplicate.",
{"churn_risk": {"type": "noul", "instructions": "Does the user threaten to leave?"}},
)
print(result["answers"]["churn_risk"]["noul"]) # probability of yes 2. ggmlc: the standalone laya binary
A single binary from the ggmlc releases runs the mys/laya-*-GGUF files with no Python: CUDA or Metal when present, else the CPU. It has a CLI, an HTTP server with a web page and a JSON-RPC daemon. These GGUF files are ggmlc's own format: llama.cpp cannot load them.
huggingface-cli download mys/laya-GGUF laya_english_f16.gguf --local-dir . laya decide laya_english_f16.gguf --preset email --device auto 3. llama.cpp: encoder in llama-server, head in Python
The fr0stbit3/laya-*-gguf files hold only the ModernBERT encoder, which llama-server serves as per-token embeddings. The decision head is the sibling -head.safetensors file and runs in Python with the laya package, which still downloads the original weights once. The model card calls these files lightly tested; prefer F16 or Q8_0.
llama-server -m laya-F16.gguf --embeddings --pooling none -c 2048 -ub 2048 -b 2048 --port 8080 Ollama
Laya is not in the official Ollama library. A community upload, s1gnature/laya, is an embedding model: ollama embed returns the encoder's vectors, not Laya's answers, because the decision head is not part of it. ollama run does not chat with it: Laya does not generate text.
Laya is not LLaVA
The names are close, the models are not. LLaVA is a vision-language chat model (a 7B–34B LLM with an image encoder) that writes text about pictures; its default Ollama downloads are 4.7–20 GB. Laya is a 0.3–0.4B text encoder with a decision head: you give it a text or JSON "state" and typed questions (choice, score, yes/no), and it returns calibrated probabilities in one forward pass. It never generates tokens, reads no images and has no KV cache, which is why none of the LLM calculators on this site apply to it.
Published speed measurements
| Hardware | Runtime | Workload | Time | Source |
|---|---|---|---|---|
| Tesla T4 (16 GB) | PyTorch (laya package) | Laya English, 1 question / 10 questions in one call | 39.5 ms / 158.6 ms | convaiinnovations/ |
| Tesla T4 (16 GB) | PyTorch (laya package) | Laya multilingual, 1 question / 10 questions in one call | 32.8 ms / 72.3 ms | convaiinnovations/ |
| CPU | PyTorch, Router(preload=True) | one request, checkpoints already loaded | 193–464 ms | convaiinnovations/ |
| RTX 4050 Laptop (6 GB) | ggmlc laya, F16, --cuda-graph | one yes/no question / the 7-question email preset in one B=7 pass | 25 ms / 143 ms | mys/laya-GGUF |
| Apple GPU | PyTorch (laya package) | Laya multilingual, one 4,000-token document, max_len=8192 | about 1.7 s | convaiinnovations/ |
File sizes read from the Hugging Face API on ; layer counts from each checkpoint's encoder/config.json. The weights are exact; activation memory is an estimate from the architecture (one encoder layer's tensors at a time, in BF16 for PyTorch and F32 for ggml) plus a fixed 0.5 GB for the runtime. Repositories: convaiinnovations/
For chat LLMs, use the LLM VRAM calculator.
Frequently asked questions
How much VRAM does Laya need?
Under 1.5 GB. The largest file is 850 MB (typed-decisions, ggmlc F16), and one forward pass at the default input length peaks at 0.76–1.40 GB including 0.5 GB for the runtime. A batch of 64 rows × 512 tokens still needs only about 2.1–2.6 GB.
Can Laya run without a GPU?
Yes. The laya Python package and the ggmlc laya binary both run on the CPU, using the same 1–1.5 GB as the table, from system RAM. The official model card reports 193–464 ms per request on a CPU with the checkpoints preloaded.
Is Laya the same as LLaVA?
No. LLaVA is a 7B–34B vision-language chat model; Laya is a 0.3–0.4B text classification and decision encoder that neither generates text nor reads images.
Can Ollama run Laya?
Laya is not in the official Ollama library. The community upload s1gnature/laya only returns the encoder's vectors through ollama embed; it has no decision head, so it does not give Laya's answers. Use the laya Python package or the ggmlc laya binary to run the whole model.
What is the difference between the three Laya checkpoints?
Laya English and typed-decisions are ModernBERT-large (421 million parameters, 28 layers); multilingual is mmBERT-base (322 million, 22 layers), covers 100+ languages and reads documents of up to 8,192 tokens.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- Image & Video Model VRAM Calculator Peak VRAM and system RAM for FLUX.2, Wan 2.2, LTX-2 and Qwen-Image in ComfyUI, from exact file sizes, checked against public runs. Open →
- Jev Alternatives You Can Run Locally Open models that replace the API-only Jev on your own machine: Laya, zero-shot encoders and small LLMs, with sizes and VRAM. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- What Can My PC Run? Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open →
Updated