Changelog
ModelVRAM had 10 changes on September 30, 2026, the latest day below. This page lists what changed for people using the site, newest first: new calculators, model pages, data and fixes. Subscribe with the RSS feed.
- Accuracy page: prompt speed and image/video checks: The accuracy page now also lists all 32 llama-bench runs behind the time-to-first-token estimate, with the error per hardware family, and the ComfyUI load logs behind the image and video file sizes.
- Vendor-reported benchmark scores next to VRAM: GPQA, MMLU-Pro and LiveCodeBench scores that makers published in their own model cards, linked to the source: on model pages, as a GPQA column in the per-size lists, and as a y-axis option on the chart.
- Bespoke Nimble 9B model page: VRAM per quant and context for Bespoke Nimble 9B, with its GGUF files and Ollama tags.
- Article: the llama-server --fit context regression: Why llama.cpp builds b10999–b11200 picked contexts such as 828,160 tokens, how --fit sizes context, slots and layers now, and the flags to set.
- GPU index page: Every GPU, Mac, CPU-only PC and multi-GPU setup on one page, with how many models each runs.
- Shows cloud GPU options when a result doesn’t fit: When a model does not fit a local card, the calculators name the memory it needs and the cloud GPU class that holds it, with plain, unsponsored links.
- Load models by their Ollama name: Type an Ollama library name such as qwen3:30b-a3b into the LLM VRAM calculator to load that model.
- Apple LensVLM 9B model page: VRAM for Apple’s LensVLM-9B per quant and context, its GGUF files and mmproj, and what its visual text compression saves in KV cache.
- Latency and throughput estimates for Laya and Jev alternatives: How long one decision takes and how many per second a GPU or CPU handles, on the Laya and Jev alternatives pages.
- Ternary Bonsai 2 27B VRAM page: Real file sizes of the PTQ1_0, PQ2_0 and MLX files and the longest context on 8, 12 and 16 GB GPUs.
- llama-server --fit preview: The LLM VRAM calculator shows what llama-server’s automatic --fit would pick for context, slots and GPU layers, and the explicit flags instead.
- Chart: every model by the VRAM it needs: A scatter plot of all tracked models by VRAM at Q4_K_M against parameters, with 8–80 GB card lines.
- 2026 Local LLM VRAM Report: The largest open model per memory tier and the 32K KV cache ranking, with CSV downloads under CC BY 4.0.
- Image and video model VRAM calculator: Peak VRAM and system RAM for FLUX.2, Wan 2.2, LTX-2 and Qwen-Image 2.1 in ComfyUI, checked against public ComfyUI runs.
- Jev alternatives you can run locally: Open models that do the API-only Jev’s job on your own GPU or CPU, with file sizes and VRAM.
- Laya VRAM requirements: Real file sizes of Laya’s checkpoints in PyTorch, ggmlc and GGUF, and whether a 16 GB or 24 GB GPU or a CPU runs it.
- Predicted vs measured: The estimates next to 20 public llama.cpp and vLLM measurements, with the error of each.
- KV cache counted as llama.cpp allocates it: Gemma 4’s global layers and sliding-window caches now match llama.cpp’s logs to the byte.
- Open dataset and best LLMs per VRAM size: Every model’s VRAM estimates as CSV and JSON under CC BY 4.0, and the models that fit 8 to 96 GB.
- What can my PC run?: Detects your GPU in the browser and lists the local LLMs it runs, with the best quantization and speed.
- vLLM KV cache and concurrency calculator: The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, checked against startup logs.
- Fine-tuning VRAM calculator: GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against 15 published runs.
- MoE offload calculator: The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes.
- Qwen-Image 2.1 VRAM calculator: Peak VRAM for every DiT, text encoder and VAE combination of Qwen-Image 2.1, with real measurements.
- LLM VRAM and speed calculators: The first two tools: how much GPU memory a model needs from its Hugging Face files, and how fast it writes on your hardware.
Updated .