Can I run this LLM on my GPU?
One answer per model and card: whether it fits at Q4_K_M with 32K tokens of context and one request, and if not, the shorter context, smaller quant, --n-cpu-moe setting or number of cards that makes it run. 148 pairs have their own page; the rest fit with room to spare at BF16 and link to the model's page.
"Yes" means Q4_K_M (MXFP4 for gpt-oss) fits at 32K with at least 0.5 GB free; "up to" names the most precise setting that does. The figures use the same estimate as the VRAM calculator: weights, FP16 KV cache and 0.5 GB plus 10% overhead. Card not in the table? Detect your GPU in the browser and see every model it runs.