Laya VRAM Requirements: Every Checkpoint, File and Runtime

Laya is a ModernBERT-based decision encoder, not a chat LLM: it takes a text and typed questions and returns calibrated probabilities in one forward pass. Pick the checkpoint, the file and the input size to see the peak memory and which GPUs fit.

– estimated peak memory

Weights: – the size of the file; the rest is activations for one forward pass plus 0.5 GB for the runtime.

GPU memory4 GB6 GB8 GB12 GB16 GB24 GB
Result––––––

With no GPU the same amount comes out of system RAM: –.

"Tight" means it fits only if the runtime uses a fused attention kernel, which never stores the full attention-score matrix.

Can a 16 GB or 24 GB GPU, or a CPU, run Laya?

Yes, all three. The largest published file is 850 MB, and with the default input length the peak is 0.76 GB–1.40 GB for every file, so a 16 GB or 24 GB card uses less than a tenth of its memory, and a 4 GB card is enough. With no GPU the model runs from system RAM, and 3 GB free is enough even for the llama.cpp setup, which loads the model twice. What changes with the hardware is speed, not whether it runs: tens of milliseconds per request on a GPU, a few hundred on a CPU (see the measurements below).

Every Laya file: size and memory

Each checkpoint at its default input length (512 tokens for English, 1,024 for typed-decisions and multilingual), one row per pass. Size is the published file in MB as Hugging Face lists it; memory is in GB (GiB).

CheckpointRuntimeFormat FileSizePeak memory 16 GB24 GBCPU only (RAM)
Laya (English) PyTorch (laya package) F16 model.safetensors 843 MB 1.30 GB–1.30 GB Fits Fits 1.30 GB
Laya (English) ggmlc laya binary F16 laya_english_f16.gguf 846 MB 1.31 GB–1.33 GB Fits Fits 1.33 GB
Laya (English) ggmlc laya binary Q8_0 laya_english_q8_0.gguf 452 MB 0.95 GB–0.96 GB Fits Fits 0.96 GB
Laya (English) ggmlc laya binary UD_Q4_K_M laya_english_ud_q4_k_m.gguf 420 MB 0.92 GB–0.93 GB Fits Fits 0.93 GB
Laya (English) llama.cpp + head F16 laya-F16.gguf 791 MB
plus decision head 106 MB
1.26 GB–1.28 GB Fits Fits 2.66 GB
Laya (English) llama.cpp + head Q8_0 laya-Q8_0.gguf 421 MB
plus decision head 106 MB
0.92 GB–0.93 GB Fits Fits 2.32 GB
Laya (English) llama.cpp + head Q6_K laya-Q6_K.gguf 344 MB
plus decision head 106 MB
0.85 GB–0.86 GB Fits Fits 2.24 GB
Laya (English) llama.cpp + head Q4_K_M laya-Q4_K_M.gguf 272 MB
plus decision head 106 MB
0.78 GB–0.79 GB Fits Fits 2.18 GB
Laya typed-decisions PyTorch (laya package) F16 model.safetensors 843 MB 1.31 GB–1.34 GB Fits Fits 1.34 GB
Laya typed-decisions ggmlc laya binary F16 laya_typed_decisions_f16.gguf 850 MB 1.34 GB–1.40 GB Fits Fits 1.40 GB
Laya typed-decisions ggmlc laya binary Q8_0 laya_typed_decisions_q8_0.gguf 455 MB 0.97 GB–1.04 GB Fits Fits 1.04 GB
Laya typed-decisions ggmlc laya binary UD_Q4_K_M laya_typed_decisions_ud_q4_k_m.gguf 424 MB 0.94 GB–1.01 GB Fits Fits 1.01 GB
Laya typed-decisions llama.cpp + head F16 laya-typed-decisions-F16.gguf 791 MB
plus decision head 106 MB
1.29 GB–1.35 GB Fits Fits 2.73 GB
Laya typed-decisions llama.cpp + head Q8_0 laya-typed-decisions-Q8_0.gguf 421 MB
plus decision head 106 MB
0.94 GB–1.00 GB Fits Fits 2.39 GB
Laya typed-decisions llama.cpp + head Q6_K laya-typed-decisions-Q6_K.gguf 344 MB
plus decision head 106 MB
0.87 GB–0.93 GB Fits Fits 2.32 GB
Laya typed-decisions llama.cpp + head Q4_K_M laya-typed-decisions-Q4_K_M.gguf 272 MB
plus decision head 106 MB
0.80 GB–0.87 GB Fits Fits 2.25 GB
Laya multilingual PyTorch (laya package) F16 model.safetensors 644 MB 1.11 GB–1.14 GB Fits Fits 1.14 GB
Laya multilingual ggmlc laya binary F16 laya_multilingual_f16.gguf 663 MB 1.15 GB–1.19 GB Fits Fits 1.19 GB
Laya multilingual ggmlc laya binary Q8_0 laya_multilingual_q8_0.gguf 362 MB 0.86 GB–0.91 GB Fits Fits 0.91 GB
Laya multilingual ggmlc laya binary UD_Q4_K_M laya_multilingual_ud_q4_k_m.gguf 524 MB 1.02 GB–1.06 GB Fits Fits 1.06 GB
Laya multilingual llama.cpp + head F16 laya-multilingual-F16.gguf 629 MB
plus decision head 60 MB
1.11 GB–1.16 GB Fits Fits 2.32 GB
Laya multilingual llama.cpp + head Q8_0 laya-multilingual-Q8_0.gguf 341 MB
plus decision head 60 MB
0.85 GB–0.89 GB Fits Fits 2.05 GB
Laya multilingual llama.cpp + head Q6_K laya-multilingual-Q6_K.gguf 271 MB
plus decision head 60 MB
0.78 GB–0.83 GB Fits Fits 1.98 GB
Laya multilingual llama.cpp + head Q4_K_M laya-multilingual-Q4_K_M.gguf 249 MB
plus decision head 60 MB
0.76 GB–0.81 GB Fits Fits 1.96 GB

Longer inputs and bigger batches

Memory grows with rows × tokens, and the attention scores with the square of the length. Even the heaviest case here, laya-multilingual reading one 8,192-token document, stays under 5 GB.

WorkloadPeak memory4 GB6 GB8 GB16 GB24 GB
Laya English, PyTorch, 32 rows × 512 tokens1.68 GB–1.93 GBFitsFitsFitsFitsFits
Laya English, PyTorch, 64 rows × 512 tokens2.08 GB–2.58 GBFitsFitsFitsFitsFits
multilingual, PyTorch, 1 document × 8,192 tokens1.21 GB–2.71 GBFitsFitsFitsFitsFits
multilingual, ggmlc F16, 1 document × 8,192 tokens1.34 GB–4.34 GBTightFitsFitsFitsFits

Three ways to run Laya locally

1. PyTorch: the laya Python package

The reference way, and the only one that runs all three checkpoints and routes each request to the right one. It downloads the safetensors files on first use and runs on CUDA, Apple GPUs or the CPU. PyPI, GitHub.

pip install laya
from laya import Router

router = Router()  # Router(preload=True) loads all three checkpoints up front
result = router.predict(
    "We were billed twice for March. Please refund the duplicate.",
    {"churn_risk": {"type": "noul", "instructions": "Does the user threaten to leave?"}},
)
print(result["answers"]["churn_risk"]["noul"])  # probability of yes

2. ggmlc: the standalone laya binary

A single binary from the ggmlc releases runs the mys/laya-*-GGUF files with no Python: CUDA or Metal when present, else the CPU. It has a CLI, an HTTP server with a web page and a JSON-RPC daemon. These GGUF files are ggmlc's own format: llama.cpp cannot load them.

huggingface-cli download mys/laya-GGUF laya_english_f16.gguf --local-dir .
laya decide laya_english_f16.gguf --preset email --device auto

3. llama.cpp: encoder in llama-server, head in Python

The fr0stbit3/laya-*-gguf files hold only the ModernBERT encoder, which llama-server serves as per-token embeddings. The decision head is the sibling -head.safetensors file and runs in Python with the laya package, which still downloads the original weights once. The model card calls these files lightly tested; prefer F16 or Q8_0.

llama-server -m laya-F16.gguf --embeddings --pooling none -c 2048 -ub 2048 -b 2048 --port 8080

Ollama

Laya is not in the official Ollama library. A community upload, s1gnature/laya, is an embedding model: ollama embed returns the encoder's vectors, not Laya's answers, because the decision head is not part of it. ollama run does not chat with it: Laya does not generate text.

Laya is not LLaVA

The names are close, the models are not. LLaVA is a vision-language chat model (a 7B–34B LLM with an image encoder) that writes text about pictures; its default Ollama downloads are 4.7–20 GB. Laya is a 0.3–0.4B text encoder with a decision head: you give it a text or JSON "state" and typed questions (choice, score, yes/no), and it returns calibrated probabilities in one forward pass. It never generates tokens, reads no images and has no KV cache, which is why none of the LLM calculators on this site apply to it.

Published speed measurements

HardwareRuntimeWorkloadTimeSource
Tesla T4 (16 GB)PyTorch (laya package)Laya English, 1 question / 10 questions in one call39.5 ms / 158.6 ms convaiinnovations/laya
Tesla T4 (16 GB)PyTorch (laya package)Laya multilingual, 1 question / 10 questions in one call32.8 ms / 72.3 ms convaiinnovations/laya
CPUPyTorch, Router(preload=True)one request, checkpoints already loaded193–464 ms convaiinnovations/laya
RTX 4050 Laptop (6 GB)ggmlc laya, F16, --cuda-graphone yes/no question / the 7-question email preset in one B=7 pass25 ms / 143 ms mys/laya-GGUF
Apple GPUPyTorch (laya package)Laya multilingual, one 4,000-token document, max_len=8192about 1.7 s convaiinnovations/laya

File sizes read from the Hugging Face API on ; layer counts from each checkpoint's encoder/config.json. The weights are exact; activation memory is an estimate from the architecture (one encoder layer's tensors at a time, in BF16 for PyTorch and F32 for ggml) plus a fixed 0.5 GB for the runtime. Repositories: convaiinnovations/laya, mys/laya-GGUF, fr0stbit3/laya-gguf, convaiinnovations/laya-typed-decisions, mys/laya-typed-decisions-GGUF, fr0stbit3/laya-typed-decisions-gguf, convaiinnovations/laya-multilingual, mys/laya-multilingual-GGUF, fr0stbit3/laya-multilingual-gguf

For chat LLMs, use the LLM VRAM calculator.

Frequently asked questions

How much VRAM does Laya need?

Under 1.5 GB. The largest file is 850 MB (typed-decisions, ggmlc F16), and one forward pass at the default input length peaks at 0.76–1.40 GB including 0.5 GB for the runtime. A batch of 64 rows × 512 tokens still needs only about 2.1–2.6 GB.

Can Laya run without a GPU?

Yes. The laya Python package and the ggmlc laya binary both run on the CPU, using the same 1–1.5 GB as the table, from system RAM. The official model card reports 193–464 ms per request on a CPU with the checkpoints preloaded.

Is Laya the same as LLaVA?

No. LLaVA is a 7B–34B vision-language chat model; Laya is a 0.3–0.4B text classification and decision encoder that neither generates text nor reads images.

Can Ollama run Laya?

Laya is not in the official Ollama library. The community upload s1gnature/laya only returns the encoder's vectors through ollama embed; it has no decision head, so it does not give Laya's answers. Use the laya Python package or the ggmlc laya binary to run the whole model.

What is the difference between the three Laya checkpoints?

Laya English and typed-decisions are ModernBERT-large (421 million parameters, 28 layers); multilingual is mmBERT-base (322 million, 22 layers), covers 100+ languages and reads documents of up to 8,192 tokens.

More calculators

Updated