# Using ModelVRAM from an agent

Two ways to get numbers: read the static JSON API (no key, CORS open), or open a calculator with URL parameters and read the page. Both use the same formulas. Data version 1.2026-09-29b, license CC BY 4.0.

## JSON API

```sh
curl -s https://modelvram.com/api/v1/models.json
curl -s https://modelvram.com/api/v1/models/qwen3.6-35b-a3b.json
curl -s https://modelvram.com/api/v1/gpus/rtx-4090.json
```

Model ids are the model page slugs (`qwen3.6-35b-a3b`), GPU ids the GPU page slugs (`rtx-4090`, `2x-rtx-3090`); the index files list them all. Only the 55 tracked models are in the API; for any other Hugging Face model, use the VRAM calculator with `model=<org>/<repo>`.

## LLM VRAM calculator: https://modelvram.com/llm-vram-calculator/

| Parameter | Values |
|---|---|
| `model` | a tracked model id (`Qwen/Qwen3.6-35B-A3B`), any Hugging Face repo id or URL (loaded from config.json and the file list), or `custom` |
| `precision` | `native` (As published), `fp32` (FP32), `bf16` (FP16 / BF16), `fp8` (FP8 / INT8), `int4` (INT4 (AWQ / GPTQ)), `q8_0` (GGUF Q8_0), `q6_k` (GGUF Q6_K), `q5_k_m` (GGUF Q5_K_M), `q4_k_m` (GGUF Q4_K_M), `iq4_xs` (GGUF IQ4_XS), `q3_k_m` (GGUF Q3_K_M), `iq3_xxs` (GGUF IQ3_XXS), `q2_k` (GGUF Q2_K) |
| `context` | tokens per request, e.g. `32768` |
| `requests` | requests at the same time (parallel sequences), e.g. `1` |
| `kv` | `fp16` (FP16 / BF16), `fp8` (FP8), `q8_0` (Q8_0), `q4_0` (Q4_0) |
| `overhead` | extra overhead in percent, default `10` |
| `params`, `layers`, `kvHeads`, `headDim`, `mla`, `sliding`, `window` | with `model=custom`: parameters in billions, layers, KV heads, head dimension, MLA cache size (0 = none), sliding-window layers, window in tokens |

Example: https://modelvram.com/llm-vram-calculator/?model=Qwen%2FQwen3.6-35B-A3B&precision=q4_k_m&context=32768&requests=1&kv=fp16 (Qwen3.6 35B-A3B (MoE) at Q4_K_M with 32K tokens: 23.5 GB).

## LLM speed calculator: https://modelvram.com/llm-speed-calculator/

| Parameter | Values |
|---|---|
| `model` | as above (tracked id or Hugging Face repo) |
| `gpu` | index into this list: `0` RTX 4060 8GB, `1` RTX 3060 12GB, `2` Arc B580 12GB, `3` RTX 4070 12GB, `4` RTX 5070 12GB, `5` RTX 4060 Ti 16GB, `6` RTX 5060 Ti 16GB, `7` RX 9070 XT 16GB, `8` RTX 4070 Ti Super 16GB, `9` RTX 4080 Super 16GB, `10` RTX 5070 Ti 16GB, `11` RTX 5080 16GB, `12` RTX 3090, `13` RTX 4090, `14` RX 7900 XTX, `15` RTX 5090, `16` M4 Pro Mac (64 GB), `17` M4 Max Mac (128 GB), `18` M3 Ultra Mac Studio (512 GB), `19` Ryzen AI Max+ 395 (128 GB), `20` DGX Spark (128 GB), `21` L40S, `22` RTX PRO 6000 Blackwell, `23` A100 80GB, `24` H100 SXM, `25` H200, `26` B200, `27` M1 Mac (8 GB), `28` M1 Mac (16 GB), `29` M2 or M3 Mac (8 GB), `30` M2 or M3 Mac (16 GB), `31` M2 or M3 Mac (24 GB), `32` M4 Mac (16 GB), `33` M4 Mac (24 GB), `34` M4 Mac (32 GB), `35` M5 Mac (16 GB), `36` M5 Mac (24 GB), `37` M5 Mac (32 GB), `38` M1 Pro or M2 Pro Mac (16 GB), `39` M1 Pro or M2 Pro Mac (32 GB), `40` M3 Pro Mac (18 GB), `41` M3 Pro Mac (36 GB), `42` M4 Pro Mac (24 GB), `43` M4 Pro Mac (48 GB), `44` M5 Pro Mac (24 GB), `45` M5 Pro Mac (48 GB), `46` M5 Pro Mac (64 GB), `47` M1 Max or M2 Max Mac (32 GB), `48` M1–M3 Max Mac (64 GB), `49` M2 Max Mac (96 GB), `50` M3 Max Mac (36 GB), `51` M3 Max Mac (48 GB), `52` M3 Max Mac (96 GB), `53` M3 Max Mac (128 GB), `54` M4 Max Mac (36 GB), `55` M4 Max Mac (48 GB), `56` M4 Max Mac (64 GB), `57` M5 Max Mac (36 GB), `58` M5 Max Mac (48 GB), `59` M5 Max Mac (64 GB), `60` M5 Max Mac (128 GB), `61` CPU only, DDR4-3200 dual channel (16 GB), `62` CPU only, DDR4-3200 dual channel (32 GB), `63` CPU only, DDR4-3200 dual channel (64 GB), `64` CPU only, DDR5-5600 dual channel (32 GB), `65` CPU only, DDR5-5600 dual channel (64 GB), `66` CPU only, DDR5-5600 dual channel (128 GB), `67` CPU only, DDR5-6000 dual channel (32 GB), `68` CPU only, DDR5-6000 dual channel (64 GB), `69` CPU only, DDR5-6000 dual channel (96 GB), `70` CPU only, DDR4-3200 quad channel (64 GB), `71` CPU only, DDR4-3200 quad channel (128 GB), `72` CPU only, DDR4-3200 quad channel (256 GB) |
| `gpus` | cards in tensor parallel: `1`, `2`, `4` or `8` |
| `precision`, `context`, `kv` | as in the VRAM calculator |
| `active` | active parameters in billions, for a MoE model loaded from Hugging Face |

## MoE offload calculator (--n-cpu-moe): https://modelvram.com/moe-offload-calculator/

| Parameter | Values |
|---|---|
| `file` | one of the measured GGUF files below (`<repo>/<path>`) |
| `gpu` | GPU name exactly as listed on the page, e.g. `RTX 4090` |
| `ram` | index: `0` DDR4-3200, 2 channels (51 GB/s), `1` DDR5-5600, 2 channels (90 GB/s), `2` DDR5-6400, 2 channels (102 GB/s), `3` DDR5-6400, 4 channels (205 GB/s), `4` DDR5-4800, 8 channels (307 GB/s) |
| `ctx` | context in tokens |
| `kv` | cache bits: `16` (f16), `8.5` (q8_0), `4.5` (q4_0) |

- `unsloth/Qwen3.6-35B-A3B-GGUF/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf`
- `unsloth/Qwen3.6-35B-A3B-GGUF/Qwen3.6-35B-A3B-UD-IQ4_XS.gguf`
- `unsloth/Qwen3.6-35B-A3B-GGUF/Qwen3.6-35B-A3B-Q8_0.gguf`
- `ornith-ai/Ornith-1.5-35B-A3B-GGUF/Ornith-1.5-35B-Q4_K_M.gguf`
- `ornith-ai/Ornith-1.5-35B-A3B-GGUF/Ornith-1.5-35B-Q8_0.gguf`
- `unsloth/Qwen3-30B-A3B-GGUF/Qwen3-30B-A3B-Q4_K_M.gguf`
- `unsloth/Qwen3-30B-A3B-GGUF/Qwen3-30B-A3B-Q8_0.gguf`
- `unsloth/gemma-4-26B-A4B-it-GGUF/gemma-4-26B-A4B-it-UD-Q4_K_M.gguf`
- `unsloth/GLM-4.7-Flash-GGUF/GLM-4.7-Flash-Q4_K_M.gguf`
- `unsloth/GLM-4.7-Flash-GGUF/GLM-4.7-Flash-UD-Q4_K_XL.gguf`
- `unsloth/Nemotron-3-Nano-30B-A3B-GGUF/Nemotron-3-Nano-30B-A3B-Q4_K_M.gguf`
- `ggml-org/gpt-oss-20b-GGUF/gpt-oss-20b-MXFP4.gguf`
- `ggml-org/gpt-oss-120b-GGUF/gpt-oss-120b-MXFP4.gguf`
- `unsloth/Qwen3-Coder-Next-GGUF/Qwen3-Coder-Next-Q4_K_M.gguf`
- `unsloth/Qwen3-Coder-Next-GGUF/Qwen3-Coder-Next-IQ4_XS.gguf`
- `unsloth/Qwen3-Coder-Next-GGUF/Qwen3-Coder-Next-UD-Q4_K_XL.gguf`
- `unsloth/Qwen3.5-122B-A10B-GGUF/Q4_K_M/Qwen3.5-122B-A10B-Q4_K_M-0000x-of-00003.gguf`
- `unsloth/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-0000x-of-00004.gguf`
- `unsloth/Qwen3.8-Flash-Next-GGUF/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-0000x-of-00003.gguf`
- `unsloth/Qwen3.8-Flash-Next-GGUF/UD-IQ3_XXS/Qwen3.8-Flash-Next-UD-IQ3_XXS-0000x-of-00003.gguf`
- `unsloth/Qwen3.8-Flash-Next-GGUF/UD-Q2_K_XL/Qwen3.8-Flash-Next-UD-Q2_K_XL-0000x-of-00003.gguf`
- `unsloth/Qwen3.8-Flash-Next-GGUF/UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-0000x-of-00003.gguf`
- `unsloth/Qwen3.8-Flash-Next-GGUF/UD-Q5_K_XL/Qwen3.8-Flash-Next-UD-Q5_K_XL-0000x-of-00006.gguf`
- `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-0000x-of-00002.gguf`

## Other calculators

These read the option values of their form fields; open the page, pick a setup and copy the address, which the page keeps up to date.

- https://modelvram.com/speculative-decoding-vram-calculator/: `main`, `draft`, `ctx`, `kv`, `dkv`, `n`
- https://modelvram.com/new-gguf-quants/: `file`, `ctx`, `kv`, `mmproj`, `mtp`
- https://modelvram.com/qwen-image-2-1-vram-calculator/: `dit`, `enc`, `vae`, `mode`, `work`
- https://modelvram.com/minimax-h3-vram-calculator/: `dit`, `enc`, `vae`, `cn`, `mode`, `res`, `len`

## Pages by URL

- Model: `https://modelvram.com/llm-vram-calculator/<model id>/`
- GPU or setup: `https://modelvram.com/llm-vram-calculator/gpu/<gpu id>/` (`rtx-4060-8gb` RTX 4060 8GB, `rtx-3060-12gb` RTX 3060 12GB, `arc-b580-12gb` Arc B580 12GB, ...)
- Can I run: `https://modelvram.com/can-i-run/<model id>-on-<gpu id>/`, for the pairs listed on https://modelvram.com/can-i-run/

## Citing

"Data: ModelVRAM (modelvram.com)", https://modelvram.com/data/, data version 1.2026-09-29b. The numbers are estimates for the stated setting; https://modelvram.com/accuracy/ shows how they compare with measurements.
