Hy4 Preview 770B-A49B VRAM requirements
Hy4 Preview 770B-A49B has 780.0B parameters (the 770B in the name leaves out the MTP layer, which the safetensors count includes), of which about 49B are used per token; all 256 experts still have to be in memory. With an 8K-token context and one request it needs about 485 GB of GPU memory at Q4_K_M, 800 GB at FP8 and 1,599 GB at FP16/BF16. The published weights take 1,453 GB (BF16). It does not fit on a 24 GB card at Q4_K_M; the smallest setup here that holds it with an 8K context is 8× H100 SXM (640 GB).
485 GB at Q4_K_M, 8K context, one request
- FP8
- 800 GB
- FP16 / BF16
- 1,599 GB
- Published weights
- 1,453 GB
- Smallest setup, Q4_K_M
- 8× H100 SXM (640 GB)
Worked example: Hy4 Preview 770B-A49B with 32K tokens
Inputs: Hy4 Preview 770B-A49B · Q4_K_M weights · 32,768 tokens of context · one request · FP16 KV cache
- Weights (780.0B at Q4_K_M)
- 439 GB
- KV cache (90 KB per token)
- 2.82 GB
- Buffers and runtime (0.5 GB + 10%)
- 44.7 GB
- Total
- 487 GB
The smallest setup here that holds it is 8× H100 SXM (640 GB): split across the 8 cards, with the 0.5 GB runtime base on each of them, it takes 491 GB and leaves about 149 GB to spare. Change the inputs in the calculator
What makes Hy4 Preview 770B-A49B's memory use different
It uses multi-head latent attention: each layer caches a latent of 576 values per token instead of 4,096 for 8 full key and value heads, 7.1× less.
Its sparse-attention indexer adds 3 KB per token on top of the main cache.
As a mixture-of-experts model it reads about 27.6 GB of its 439 GB Q4_K_M weights per generated token (6%), so it writes like a much smaller model while needing memory for all of them.
How much VRAM does Hy4 Preview 770B-A49B need?
Hy4 Preview 770B-A49B needs 485 GB at Q4_K_M, 850 GB at Q8_0 and 1,599 GB at FP16 with 8,192 tokens of context and one request. Each total below is the weights plus the FP16 KV cache and the runtime overhead (0.5 GB plus 10%), and each row opens the calculator with that setting. What the GGUF names mean.
| Precision | Weights | Total | Smallest setup |
|---|---|---|---|
| As published (BF16) | 1,453 GB | 1,599 GB | None listed |
| FP16 / BF16 | 1,453 GB | 1,599 GB | None listed |
| FP8 / INT8 | 726 GB | 800 GB | 8× H200 (1,128 GB) |
| INT4 (AWQ / GPTQ) | 386 GB | 426 GB | 8× H100 SXM (640 GB) |
| GGUF Q8_0 | 772 GB | 850 GB | 8× H200 (1,128 GB) |
| GGUF Q6_K | 596 GB | 656 GB | 8× H200 (1,128 GB) |
| GGUF Q5_K_M | 515 GB | 568 GB | 8× H100 SXM (640 GB) |
| GGUF Q4_K_M | 439 GB | 485 GB | 8× H100 SXM (640 GB) |
| GGUF IQ4_XS | 395 GB | 436 GB | 8× H100 SXM (640 GB) |
| GGUF Q3_K_M | 355 GB | 392 GB | 8× H100 SXM (640 GB) |
| GGUF IQ3_XXS | 300 GB | 331 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
| GGUF Q2_K | 304 GB | 336 GB | M3 Ultra Mac Studio (512 GB, 384 GB usable) |
Fine-tuning Hy4 Preview 770B-A49B? Hy4 Preview 770B-A49B VRAM for LoRA, QLoRA and full training.
Hy4 Preview 770B-A49B at 8K, 32K, 128K and 1M (full) tokens of context
All 78 layers use multi-head latent attention. Each extra token of context adds 90 KB of FP16 cache per request. The last column is the smallest setup here, one card or a group, that holds it at Q4_K_M. How the KV cache works.
| Context | KV cache, FP16 | Total at Q4_K_M | Total at Q8_0 | Smallest setup, Q4_K_M |
|---|---|---|---|---|
| 8K tokens | 723 MB | 485 GB | 850 GB | 8× H100 SXM (640 GB) |
| 32K tokens | 2.82 GB | 487 GB | 853 GB | 8× H100 SXM (640 GB) |
| 128K tokens | 11.3 GB | 496 GB | 862 GB | 8× H100 SXM (640 GB) |
| 1M tokens | 90.4 GB | 583 GB | 949 GB | 8× H100 SXM (640 GB) |
How fast Hy4 Preview 770B-A49B writes on a multi-GPU group
Tokens per second for one request with 8,192 tokens of context, estimated from memory bandwidth and the parameters read per token. No single card or machine holds it at Q4_K_M or FP8 (the precisions below), so these rows are the GPU groups here that do, in tensor parallel with their combined bandwidth. A dash means it does not fit at that precision. Try other settings in the speed calculator.
| Hardware | Bandwidth | Q4_K_M | FP8 |
|---|---|---|---|
| 8× H100 SXM | 26,800 GB/s | 137–280 | — |
| 8× H200 | 38,400 GB/s | 163–347 | 128–257 |
Longest context on a multi-GPU group
No single card or machine holds it at Q4_K_M, Q8_0 or FP8 (the precisions below), so these rows use GPU groups: how many tokens of context Hy4 Preview 770B-A49B fits on each, in one tensor-parallel group with one request, an FP16 KV cache and 0.5 GB left free. "Full" means the model's whole context window fits.
| Group | Memory | Q4_K_M | Q8_0 | FP8 |
|---|---|---|---|---|
| 2× RTX 3060 12GB | 24 GB | No | No | No |
| 2× RTX 3090 | 48 GB | No | No | No |
| 2× RTX 4090 | 48 GB | No | No | No |
| 2× RTX 5090 | 64 GB | No | No | No |
| 4× RTX 3090 | 96 GB | No | No | No |
| 8× H100 SXM | 640 GB | 1M (full) | No | No |
| 8× H200 | 1,128 GB | 1M (full) | 1M (full) | 1M (full) |
Model details
- Parameters
- 780.0B (779,960,992,733); the 770B in the name leaves out the MTP layer, which the safetensors count includes
- Experts
- 256 routed experts, all loaded
- Active per token
- 49B
- Layers
- All 78 layers use multi-head latent attention
- Attention cache
- 576 values per token (compressed latent)
- Context length
- 1,048,576 tokens
- Published weights
- 1,453 GB (BF16)
- On Hugging Face
- tencent/Hy4-preview
Why this estimate looks this way
At Q4_K_M and an 8K-token context, Hy4 Preview 770B-A49B uses 439 GB for weights, 723 MB for its FP16 KV cache and 44.5 GB for estimated runtime overhead, totaling 485 GB. The overhead is a 0.5 GB base plus 10% of weights and cache.
This is a mixture-of-experts model: all 256 experts contribute to the 780.0B parameters held in memory, even though only about 49B parameters run per token. Using only active parameters would understate VRAM.
Its attention layout matters for long context: all 78 layers use multi-head latent attention. The FP16 cache grows by about 90 KB per additional token per request.
The architecture and published checkpoint size come from the model's config.json and weight files. These are estimates rather than measured peak memory; inference engines can reserve extra buffers or preallocate the full configured cache.
Hy4 Preview 770B-A49B next to similar-sized models
The three models here closest to it in parameter count, at Q4_K_M with 8,192 tokens of context.
| Model | Parameters | Total | KV per token | Smallest setup |
|---|---|---|---|---|
| Hy4 Preview 770B-A49B | 780.0B, 49B active | 485 GB | 90 KB | 8× H100 SXM (640 GB) |
| DeepSeek V4.1 Flash | 763.2B, 16B active | 474 GB | 80 KB | 8× H100 SXM (640 GB) |
| GLM-5.3 | 753.3B, 40B active | 468 GB | 90 KB | 8× H100 SXM (640 GB) |
| GLM-5.2 | 753.3B, 40B active | 468 GB | 90 KB | 8× H100 SXM (640 GB) |
Badge for your model card
Paste it into a Hugging Face or GitHub README; it shows the Q4_K_M total above and links to this page. Q8_0 and FP16 badges.
Numbers read from the model files on Hugging Face on . Compare it with other models in the reproducible model and GPU report.