IQuest-Q1 320B-A15B VRAM requirements
IQuest-Q1 320B-A15B has 320.3B parameters, of which about 15B are used per token; all 256 experts still have to be in memory. With an 8K-token context and one request it needs about 201 GB of GPU memory at Q4_K_M, 331 GB at FP8 and 659 GB at FP16/BF16. The published weights take 597 GB (BF16). It does not fit on a 24 GB card at Q4_K_M; the smallest setup here that holds it with an 8K context is the M3 Ultra Mac Studio (512 GB, 384 GB usable).
201 GB at Q4_K_M, 8K context, one request
- FP8
- 331 GB
- FP16 / BF16
- 659 GB
- Published weights
- 597 GB
- Smallest setup, Q4_K_M
- M3 Ultra Mac Studio (512 GB, 384 GB usable)
Worked example: IQuest-Q1 320B-A15B with 32K tokens
Inputs: IQuest-Q1 320B-A15B · Q4_K_M weights · 32,768 tokens of context · one request · FP16 KV cache
- Weights (320.3B at Q4_K_M)
- 180 GB
- KV cache (100 KB per token)
- 4.23 GB
- Buffers and runtime (0.5 GB + 10%)
- 19.0 GB
- Total
- 204 GB
The smallest setup here that holds it is the M3 Ultra Mac Studio (512 GB, 384 GB usable), with about 180 GB to spare. Change the inputs in the calculator
What makes IQuest-Q1 320B-A15B's memory use different
Its 63 sliding-window layers keep only the last 4,096 tokens (llama.cpp gives them 4,608 cells: the window plus a 512-token batch, rounded up to 256): at its 512K limit they hold 1.11 GB of cache instead of the 126 GB they would need with full attention.
As a mixture-of-experts model it reads about 8.45 GB of its 180 GB Q4_K_M weights per generated token (5%), so it writes like a much smaller model while needing memory for all of them.
How much VRAM does IQuest-Q1 320B-A15B need?
IQuest-Q1 320B-A15B needs 201 GB at Q4_K_M, 351 GB at Q8_0 and 659 GB at FP16 with 8,192 tokens of context and one request. Each total below is the weights plus the FP16 KV cache and the runtime overhead (0.5 GB plus 10%), and each row opens the calculator with that setting. What the GGUF names mean.
| Precision | Weights | Total | Smallest setup |
|---|---|---|---|
| As published (BF16) | 597 GB | 659 GB | 8× H200 (1,128 GB) |
| FP16 / BF16 | 597 GB | 659 GB | 8× H200 (1,128 GB) |
| FP8 / INT8 | 298 GB | 331 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| INT4 (AWQ / GPTQ) | 158 GB | 177 GB | B200 (180 GB) |
| GGUF Q8_0 | 317 GB | 351 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GGUF Q6_K | 245 GB | 272 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GGUF Q5_K_M | 211 GB | 235 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GGUF Q4_K_M | 180 GB | 201 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GGUF IQ4_XS | 162 GB | 181 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GGUF Q3_K_M | 146 GB | 163 GB | B200 (180 GB) |
| GGUF IQ3_XXS | 123 GB | 138 GB | H200 (141 GB) |
| GGUF Q2_K | 125 GB | 140 GB | H200 (141 GB) |
Fine-tuning IQuest-Q1 320B-A15B? IQuest-Q1 320B-A15B VRAM for LoRA, QLoRA and full training.
IQuest-Q1 320B-A15B at 8K, 32K, 128K and 512K (full) tokens of context
25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens. Each extra token of context adds 100 KB of FP16 cache per request once the sliding windows are full. The last column is the smallest setup here, one card or a group, that holds it at Q4_K_M. How the KV cache works.
| Context | KV cache, FP16 | Total at Q4_K_M | Total at Q8_0 | Smallest setup, Q4_K_M |
|---|---|---|---|---|
| 8K tokens | 1.89 GB | 201 GB | 351 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| 32K tokens | 4.23 GB | 204 GB | 354 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| 128K tokens | 13.6 GB | 214 GB | 364 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| 512K tokens | 51.1 GB | 255 GB | 405 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
How fast IQuest-Q1 320B-A15B writes
Tokens per second for one request with 8,192 tokens of context, estimated from memory bandwidth and the parameters read per token. A dash means it does not fit on one card. Try other settings in the speed calculator.
| Hardware | Bandwidth | Q4_K_M | FP8 |
|---|---|---|---|
| RTX 3060 12GB | 360 GB/s | — | — |
| RTX 4090 | 1,008 GB/s | — | — |
| RTX 5090 | 1,792 GB/s | — | — |
| M4 Max Mac (128 GB) | 546 GB/s | — | — |
| M3 Ultra Mac Studio (512 GB) | 819 GB/s | 21–36 | 14–24 |
| H100 SXM | 3,350 GB/s | — | — |
| H200 | 4,800 GB/s | — | — |
Longest context on a multi-GPU group
No single card in this table holds it at Q4_K_M, Q8_0 or FP8, so these rows use GPU groups: how many tokens of context IQuest-Q1 320B-A15B fits on each, in one tensor-parallel group with one request, an FP16 KV cache and 0.5 GB left free. "Full" means the model's whole context window fits.
| Group | Memory | Q4_K_M | Q8_0 | FP8 |
|---|---|---|---|---|
| 2× RTX 3060 12GB | 24 GB | No | No | No |
| 2× RTX 3090 | 48 GB | No | No | No |
| 2× RTX 4090 | 48 GB | No | No | No |
| 2× RTX 5090 | 64 GB | No | No | No |
| 4× RTX 3090 | 96 GB | No | No | No |
| 8× H100 SXM | 640 GB | 512K (full) | 512K (full) | 512K (full) |
| 8× H200 | 1,128 GB | 512K (full) | 512K (full) | 512K (full) |
Model details
- Parameters
- 320.3B (320,318,615,552)
- Experts
- 256 routed experts, all loaded
- Active per token
- 15B
- Layers
- 25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens
- Attention cache
- 8 KV heads × 128
- Context length
- 524,288 tokens
- Published weights
- 597 GB (BF16)
- On Hugging Face
- IQuestLab/IQuest-Q1
Why this estimate looks this way
At Q4_K_M and an 8K-token context, IQuest-Q1 320B-A15B uses 180 GB for weights, 1.89 GB for its FP16 KV cache and 18.7 GB for estimated runtime overhead, totaling 201 GB. The overhead is a 0.5 GB base plus 10% of weights and cache.
This is a mixture-of-experts model: all 256 experts contribute to the 320.3B parameters held in memory, even though only about 15B parameters run per token. Using only active parameters would understate VRAM.
Its attention layout matters for long context: 25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens. The FP16 cache grows by about 100 KB per additional token per request once the sliding windows are full.
The architecture and published checkpoint size come from the model's config.json and weight files. These are estimates rather than measured peak memory; inference engines can reserve extra buffers or preallocate the full configured cache.
IQuest-Q1 320B-A15B next to similar-sized models
The three models here closest to it in parameter count, at Q4_K_M with 8,192 tokens of context.
| Model | Parameters | Total | KV per token | Smallest setup |
|---|---|---|---|---|
| IQuest-Q1 320B-A15B | 320.3B, 15B active | 201 GB | 100 KB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GLM-5.3 Flash | 321.3B, 18B active | 200 GB | 12 KB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| MiMo V2.6 Flash | 310.8B, 15B active | 193 GB | 23 KB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| DeepSeek V4 Flash 0731 | 304.2B, 13B active | 190 GB | 86 KB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
Badge for your model card
Paste it into a Hugging Face or GitHub README; it shows the Q4_K_M total above and links to this page. Q8_0 and FP16 badges.
Numbers read from the model files on Hugging Face on . Compare it with other models in the reproducible model and GPU report.