# ModelVRAM > How much GPU memory a model needs, how fast it runs on your hardware and what fits your GPU, from the model files on Hugging Face. No install, no sign-up. Estimates are worked out from each model's config.json and file sizes on Hugging Face (architecture-level KV cache: MLA, sliding windows, hybrid linear layers), and compared with public measurements on the accuracy page. The site is static: every page and JSON file is generated at build time; the calculators run in the browser and only contact Hugging Face. Data version 1.2026-09-29b, model data checked 2026-09-29. English pages, most also in Chinese under /zh/. ## Calculators - [Fine-Tuning VRAM Calculator](https://modelvram.com/fine-tuning-vram-calculator/): GPU memory for full fine-tuning, LoRA and QLoRA of any Hugging Face model in Transformers or Unsloth, split into parts and checked against 15 published runs. - [vLLM KV Cache & Concurrency Calculator](https://modelvram.com/vllm-memory-calculator/): How many tokens of KV cache vLLM allocates and how many requests fit, for any model and GPU, tensor parallel and FP8 cache. Checked against real vLLM logs. - [What Can My PC Run?](https://modelvram.com/what-can-i-run/): Detects your GPU in the browser, then lists the local LLMs that fit its VRAM: best quantization, longest context, tokens/s and MoE offload. Nothing is uploaded. - [MiniMax H3 VRAM Calculator (ComfyUI)](https://modelvram.com/minimax-h3-vram-calculator/): How much VRAM and RAM MiniMax H3 needs in ComfyUI: exact DiT, Qwen3-VL-32B encoder and VAE file sizes, measured runs and what streams from RAM on 8–96 GB GPUs. - [MoE Offload Calculator (--n-cpu-moe)](https://modelvram.com/moe-offload-calculator/): Find the --n-cpu-moe value that fits a MoE GGUF on your GPU, the system RAM the offloaded experts take and the speed, from the real tensor sizes of each file. - [Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD](https://modelvram.com/new-gguf-quants/): Qwen3.8 27B in Ternary Bonsai 2 (PTQ1_0, PQ2_0), GSQ-RCO and Unsloth UD GGUF: exact sizes, VRAM with KV cache, quality, speed and which fit 8–24 GB GPUs. - [Qwen-Image-2.1 VRAM Calculator](https://modelvram.com/qwen-image-2-1-vram-calculator/): How much VRAM Qwen-Image-2.1 needs for every mix of DiT, Qwen3-VL text encoder and VAE: exact file sizes, peak memory measured on real GPUs and which cards fit. - [Speculative Decoding VRAM Calculator (MTP & draft models)](https://modelvram.com/speculative-decoding-vram-calculator/): Extra VRAM of speculative decoding in llama.cpp: MTP heads, DFlash drafters and draft models for Qwen3.8 27B and Gemma 4 31B, from exact GGUF tensor bytes. - [LLM Speed Calculator](https://modelvram.com/llm-speed-calculator/): Estimate how many tokens per second an LLM writes on your GPU or Mac from memory bandwidth, active parameters and context length. Any model on Hugging Face. - [LLM VRAM Calculator](https://modelvram.com/llm-vram-calculator/): Estimate how much GPU memory an LLM needs to run: weights, KV cache and overhead for any quantization and context length. Load any model from Hugging Face. ## Reference pages - [Model pages](https://modelvram.com/llm-vram-calculator/qwen3.6-35b-a3b/): one per tracked model (55), e.g. Qwen3.6 35B-A3B (MoE): VRAM per precision and context, longest context per GPU, speed, llama-server command - [GPU pages](https://modelvram.com/llm-vram-calculator/gpu/rtx-4060-8gb/): one per GPU or multi-GPU setup (73): which models fit, best precision, longest context, tokens/s - [Can I run it?](https://modelvram.com/can-i-run/): yes/no answers for popular model and GPU pairs at Q4_K_M with 32K context - [Best LLMs for your VRAM](https://modelvram.com/best-llm-for-vram/): the models that fit 8 to 96 GB - [Predicted vs measured](https://modelvram.com/accuracy/): public llama.cpp and vLLM measurements next to the site's predictions - [KV cache explained](https://modelvram.com/llm-vram-calculator/kv-cache/): how much memory long context needs - [GGUF quantization explained](https://modelvram.com/llm-vram-calculator/gguf-quantization/): what Q4_K_M, IQ4_XS and the other names mean ## Data and API - [Open dataset](https://modelvram.com/data/): CSV (https://modelvram.com/data/modelvram-vram.csv) and JSON (https://modelvram.com/data/modelvram-vram.json), one row per model - [/api/v1/models.json](https://modelvram.com/api/v1/models.json): every model with parameters, KV bytes per token (f16, q8_0) and weights per precision - [/api/v1/models/.json](https://modelvram.com/api/v1/models/qwen3.6-35b-a3b.json): totals per precision at 4K-128K and the longest context, checked GGUF repos, smallest setup, a "Can I run" verdict per GPU - [/api/v1/gpus.json](https://modelvram.com/api/v1/gpus.json): every GPU and multi-GPU setup - [/api/v1/gpus/.json](https://modelvram.com/api/v1/gpus/rtx-4090.json): the models one setup holds and the ones it does not - [Agent guide](https://modelvram.com/agent.md): the calculators' URL parameters, to open a calculator with a given setup ## License and citation Data under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Cite as "Data: ModelVRAM (modelvram.com)", with a link to https://modelvram.com/data/ or to the model or GPU page you quote, and the data version (1.2026-09-29b). Figures are estimates for one request with 10% + 0.5 GB overhead, not measured peaks. ## Optional - [Full text for LLMs](https://modelvram.com/llms-full.txt): this file, the agent guide and a table of every model - [VRAM badges](https://modelvram.com/badge/): README badges per model - [Calculation code](https://github.com/159753a52/llm-vram-calculator): formulas and tests - [Submit a measurement](https://github.com/159753a52/llm-vram-calculator/issues/new?template=measurement.yml): GitHub issue form for real memory readings