What can my PC run?
The page reads your GPU from the browser (WebGL and WebGPU), asks you to confirm it and the memory, then lists every tracked model that fits, with the same rule as the GPU pages: 0.5 GB left free.
Your hardware
Reading what your browser reports about its GPU…
Detection runs in your browser: nothing about your hardware is sent anywhere, and the page makes no request to do it.
What your browser reported
- WebGL renderer
–- WebGPU adapter
–
Browsers report at most 8 GB, so pick what your PC has. Used for the MoE offload list.
A Mac’s GPU may use about 75% of unified memory by default; sudo sysctl iogpu.wired_limit_mb can raise it.
Models that fit
The most precise weights that still fit with 8K tokens of context and one request, the memory that takes, the longest context at that precision and the writing speed. Bigger models first. Each precision opens that setting in the VRAM calculator.
| Model | Best precision | Memory | Longest context | Tokens/s |
|---|
None of the tracked models fits, even at 2-bit precision.
MoE models that fit with experts in system RAM
| Model | GGUF | --n-cpu-moe | On the GPU | In RAM | Tokens/s |
|---|
Near misses
Too big at Q4_K_M with 32K context, but close. Shortfall first, then what fits instead.
| Model | Short by at Q4_K_M | Fits instead |
|---|
More for this hardware
How the detection works, and its limits
The page asks your browser for two things: the renderer name from WebGL’s WEBGL_debug_renderer_info extension (for example “ANGLE (NVIDIA, NVIDIA GeForce RTX 4090 …)”) and WebGPU’s adapter.info (vendor, architecture and sometimes the model). A script finds the card or chip in that name, maps it to the GPU list this site uses and asks you to confirm. It all happens on your machine; nothing is sent.
Browsers tell a page only so much. None reports the size of the VRAM or unified memory. Safari calls every Mac “Apple GPU”, Firefox gives a coarse name ending in “or similar” for privacy, laptops with two GPUs often run the browser on the integrated one, and navigator.deviceMemory stops at 8 GB. So the result is a guess for you to check. Once it is right, every figure follows the GPU pages’ rule: the most precise setting that fits with 8K context and 0.5 GB left free, and speed as a range from memory bandwidth.
How to check your GPU and its VRAM
| System | Command or place | What it shows |
|---|---|---|
| Windows | Task Manager > Performance > GPU | Model and “Dedicated GPU memory” |
| Windows, Linux (NVIDIA) | nvidia-smi --query-gpu=name,memory.total --format=csv | Model and total VRAM |
| macOS | system_profiler SPDisplaysDataType | Chip and GPU core count |
| macOS | sysctl hw.memsize | Unified memory, in bytes |
| Linux (AMD) | rocm-smi --showmeminfo vram | Total and used VRAM |
| Linux | lspci | grep -iE "vga|3d" | Which cards are installed (not their memory) |
What fits on common VRAM sizes
One common card with its own page per size, at Q4_K_M (MXFP4 for gpt-oss) with 8K context and 0.5 GB left free.
| VRAM | Example card | Fit at Q4_K_M | Biggest model | Tokens/s |
|---|---|---|---|---|
| 8 GB | RTX 4060 8GB | 13 of 55 | ZDTaichu 5.0 9B | 23–32 |
| 12 GB | RTX 3060 12GB | 14 of 55 | Gemma 4 12B | 24–34 |
| 16 GB | RTX 5060 Ti 16GB | 15 of 55 | gpt-oss-20b | 48–83 |
| 24 GB | RTX 4090 | 28 of 55 | Ornith 1.5 35B-A3B | 124–226 |
| 32 GB | RTX 5090 | 29 of 55 | K2-Horizon MoVA 36B-A4B | 111–200 |
How close the estimates come to real runs: predicted vs measured. Or ask the question the other way round on the “Can I run” pages.
How to use
- Open the page: it guesses your GPU from what the browser reports. Nothing is sent anywhere.
- Check the guess. Pick the right card or Mac if it is wrong, and the memory size if the card comes in more than one.
- Set your system RAM (browsers report at most 8 GB). It decides which MoE models run with experts in RAM.
- Read the models that fit, the near misses and the offload plans, then open any of them in the calculators.
Frequently asked questions
Which LLM can my PC run?
It depends mostly on your GPU’s memory (VRAM), or on a Mac its unified memory. At Q4_K_M with an 8K context and 0.5 GB left free, the biggest of the 55 tracked models that fit are: 8 GB, up to ZDTaichu 5.0 9B; 12 GB, up to Gemma 4 12B; 16 GB, up to gpt-oss-20b; 24 GB, up to Ornith 1.5 35B-A3B; 32 GB, up to K2-Horizon MoVA 36B-A4B. MoE models can go further with some experts in system RAM, at a lower speed. This page detects your GPU and lists every model for it.
How do I check my GPU VRAM on Windows, Mac or Linux?
Windows: Task Manager > Performance > GPU shows “Dedicated GPU memory”; with an NVIDIA card, nvidia-smi --query-gpu=name,memory.total --format=csv prints the name and size. Mac: system_profiler SPDisplaysDataType names the chip, and sysctl hw.memsize gives the unified memory in bytes (Apple menu > About This Mac shows both). Linux: nvidia-smi for NVIDIA, rocm-smi --showmeminfo vram for AMD, and lspci | grep -iE "vga|3d" to see which card is installed.
Why did the page guess wrong or ask me to pick?
Browsers hide some of it. Safari reports every Mac as “Apple GPU”, Firefox gives a coarse name followed by “or similar”, no browser reports the memory size, and navigator.deviceMemory stops at 8 GB. Laptops with two GPUs often run the browser on the integrated one. So the result is a guess to confirm, never a reading of your hardware.
Is any data sent anywhere?
No. The GPU name comes from the WebGL and WebGPU APIs in your browser, and all the fitting runs in the page’s own JavaScript. Nothing about your hardware is uploaded, stored or used to track you.
How accurate are these estimates?
Memory comes from each model’s real config and weight sizes, plus a 10% overhead and 0.5 GB of runtime; the accuracy page compares them with public measurements. Speeds are ranges from memory bandwidth, not benchmarks, and laptop GPUs run slower than their desktop names suggest.
More calculators
- Fine-Tuning VRAM Calculator GPU memory for full fine-tuning, LoRA and QLoRA in Transformers or Unsloth, checked against published runs. Open →
- vLLM KV Cache & Concurrency Calculator The KV cache pool vLLM allocates, in tokens, and how many requests fit at once, with the vllm serve command. Open →
- MiniMax H3 VRAM Calculator
(ComfyUI) VRAM and system RAM for MiniMax H3 video in ComfyUI: pruned, INT8, NVFP4 and GGUF files on 8–96 GB GPUs. Open → - MoE Offload Calculator
(--n-cpu-moe) The smallest llama.cpp --n-cpu-moe that fits your GPU, from real GGUF tensor sizes. Open → - Qwen3.8 27B GGUF Quants: Bonsai 2 vs GSQ-RCO vs UD Ternary Bonsai 2, GSQ-RCO and Unsloth UD files of Qwen3.8 27B: VRAM, quality, speed and the engine each needs. Open →
- Qwen-Image-2.1 VRAM Calculator Peak VRAM for every DiT, text encoder and VAE combination, with real measurements. Open →
Updated