IQuest-Q1 320B-A15B VRAM requirements

IQuest-Q1 320B-A15B has 320.3B parameters, of which about 15B are used per token; all 256 experts still have to be in memory. With an 8K-token context and one request it needs about 201 GB of GPU memory at Q4_K_M, 331 GB at FP8 and 659 GB at FP16/BF16. The published weights take 597 GB (BF16). It does not fit on a 24 GB card at Q4_K_M; the smallest setup here that holds it with an 8K context is the M3 Ultra Mac Studio (512 GB, 384 GB usable).

201 GB at Q4_K_M, 8K context, one request

FP8
331 GB
FP16 / BF16
659 GB
Published weights
597 GB
Smallest setup, Q4_K_M
M3 Ultra Mac Studio (512 GB, 384 GB usable)
Open IQuest-Q1 320B-A15B in the calculator

Worked example: IQuest-Q1 320B-A15B with 32K tokens

Inputs: IQuest-Q1 320B-A15B · Q4_K_M weights · 32,768 tokens of context · one request · FP16 KV cache

Weights (320.3B at Q4_K_M)
180 GB
KV cache (100 KB per token)
4.23 GB
Buffers and runtime (0.5 GB + 10%)
19.0 GB
Total
204 GB

The smallest setup here that holds it is the M3 Ultra Mac Studio (512 GB, 384 GB usable), with about 180 GB to spare. Change the inputs in the calculator

What makes IQuest-Q1 320B-A15B's memory use different

Its 63 sliding-window layers keep only the last 4,096 tokens (llama.cpp gives them 4,608 cells: the window plus a 512-token batch, rounded up to 256): at its 512K limit they hold 1.11 GB of cache instead of the 126 GB they would need with full attention.

As a mixture-of-experts model it reads about 8.45 GB of its 180 GB Q4_K_M weights per generated token (5%), so it writes like a much smaller model while needing memory for all of them.

How much VRAM does IQuest-Q1 320B-A15B need?

IQuest-Q1 320B-A15B needs 201 GB at Q4_K_M, 351 GB at Q8_0 and 659 GB at FP16 with 8,192 tokens of context and one request. Each total below is the weights plus the FP16 KV cache and the runtime overhead (0.5 GB plus 10%), and each row opens the calculator with that setting. What the GGUF names mean.

Fine-tuning IQuest-Q1 320B-A15B? IQuest-Q1 320B-A15B VRAM for LoRA, QLoRA and full training.

IQuest-Q1 320B-A15B at 8K, 32K, 128K and 512K (full) tokens of context

25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens. Each extra token of context adds 100 KB of FP16 cache per request once the sliding windows are full. The last column is the smallest setup here, one card or a group, that holds it at Q4_K_M. How the KV cache works.

ContextKV cache, FP16Total at Q4_K_MTotal at Q8_0Smallest setup, Q4_K_M
8K tokens 1.89 GB 201 GB 351 GB M3 Ultra Mac Studio (512 GB, 384 GB usable)
32K tokens 4.23 GB 204 GB 354 GB M3 Ultra Mac Studio (512 GB, 384 GB usable)
128K tokens 13.6 GB 214 GB 364 GB M3 Ultra Mac Studio (512 GB, 384 GB usable)
512K tokens 51.1 GB 255 GB 405 GB M3 Ultra Mac Studio (512 GB, 384 GB usable)

How fast IQuest-Q1 320B-A15B writes

Tokens per second for one request with 8,192 tokens of context, estimated from memory bandwidth and the parameters read per token. A dash means it does not fit on one card. Try other settings in the speed calculator.

HardwareBandwidthQ4_K_MFP8
RTX 3060 12GB 360 GB/s ——
RTX 4090 1,008 GB/s ——
RTX 5090 1,792 GB/s ——
M4 Max Mac (128 GB) 546 GB/s ——
M3 Ultra Mac Studio (512 GB) 819 GB/s 21–3614–24
H100 SXM 3,350 GB/s ——
H200 4,800 GB/s ——

Longest context on a multi-GPU group

No single card in this table holds it at Q4_K_M, Q8_0 or FP8, so these rows use GPU groups: how many tokens of context IQuest-Q1 320B-A15B fits on each, in one tensor-parallel group with one request, an FP16 KV cache and 0.5 GB left free. "Full" means the model's whole context window fits.

GroupMemoryQ4_K_MQ8_0FP8
2× RTX 3060 12GB 24 GB NoNoNo
2× RTX 3090 48 GB NoNoNo
2× RTX 4090 48 GB NoNoNo
2× RTX 5090 64 GB NoNoNo
4× RTX 3090 96 GB NoNoNo
8× H100 SXM 640 GB 512K (full)512K (full)512K (full)
8× H200 1,128 GB 512K (full)512K (full)512K (full)

Model details

Parameters
320.3B (320,318,615,552)
Experts
256 routed experts, all loaded
Active per token
15B
Layers
25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens
Attention cache
8 KV heads × 128
Context length
524,288 tokens
Published weights
597 GB (BF16)
On Hugging Face
IQuestLab/IQuest-Q1

Why this estimate looks this way

At Q4_K_M and an 8K-token context, IQuest-Q1 320B-A15B uses 180 GB for weights, 1.89 GB for its FP16 KV cache and 18.7 GB for estimated runtime overhead, totaling 201 GB. The overhead is a 0.5 GB base plus 10% of weights and cache.

This is a mixture-of-experts model: all 256 experts contribute to the 320.3B parameters held in memory, even though only about 15B parameters run per token. Using only active parameters would understate VRAM.

Its attention layout matters for long context: 25 of its 88 layers use full attention and 63 keep a sliding window of 4,096 tokens. The FP16 cache grows by about 100 KB per additional token per request once the sliding windows are full.

The architecture and published checkpoint size come from the model's config.json and weight files. These are estimates rather than measured peak memory; inference engines can reserve extra buffers or preallocate the full configured cache.

IQuest-Q1 320B-A15B next to similar-sized models

The three models here closest to it in parameter count, at Q4_K_M with 8,192 tokens of context.

ModelParametersTotalKV per tokenSmallest setup
IQuest-Q1 320B-A15B 320.3B, 15B active 201 GB 100 KB M3 Ultra Mac Studio (512 GB, 384 GB usable)
GLM-5.3 Flash 321.3B, 18B active 200 GB 12 KB M3 Ultra Mac Studio (512 GB, 384 GB usable)
MiMo V2.6 Flash 310.8B, 15B active 193 GB 23 KB M3 Ultra Mac Studio (512 GB, 384 GB usable)
DeepSeek V4 Flash 0731 304.2B, 13B active 190 GB 86 KB M3 Ultra Mac Studio (512 GB, 384 GB usable)

Badge for your model card

Paste it into a Hugging Face or GitHub README; it shows the Q4_K_M total above and links to this page. Q8_0 and FP16 badges.

IQuest-Q1 320B-A15B VRAM: 201 GB at Q4_K_M, 8K context

Numbers read from the model files on Hugging Face on . Compare it with other models in the reproducible model and GPU report.