DeepSeek V4 Pro 1.6T
DeepSeek V4 Pro 1.6T needs roughly 1091.4 GB VRAM at Q4_K_M quantization (3584.1 GB at FP16). 0 GPUs we track can run it fully in VRAM at 8k context.
0 GPUs run this natively · 0 with CPU offload
DeepSeek V4 Pro 1.6T is a Mixture of Experts (MoE) model with 1600B total parameters but only 49B active per token developed by DeepSeek. April 2026 1.6T parameter MoE with hybrid CSA/HCA attention and 1M token context.
To run DeepSeek V4 Pro 1.6T locally: Q4_K_M needs roughly 974.4GB of weights, about 1,091.4GB total at a short context; no single machine this site tracks fits it, datacenter-only. The Flash variant (284B/13B active) is far more practical at roughly 173.0GB of weights. As a MoE model, inference speed depends on active parameters (49B) rather than total size.
MMLU-Pro 87.5%, GPQA 90.1%, new frontier benchmark. Requires 27% of V3's inference FLOPs at 1M context.
VRAM at each quantization
Calculated at 8k context. Since KV cache scales linearly with context, longer sessions need more VRAM than shown here.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 6400.0 GB | 0.10 GB | 7168.1 GB |
| BF16 | 3200.0 GB | 0.10 GB | 3584.1 GB |
| FP16 | 3200.0 GB | 0.10 GB | 3584.1 GB |
| Q8_0 | 1700.8 GB | 0.10 GB | 1905.0 GB |
| Q6_K | 1313.6 GB | 0.10 GB | 1471.3 GB |
| Q5_K_M | 1139.2 GB | 0.10 GB | 1276.0 GB |
| Q4_K_Mrec | 974.4 GB | 0.10 GB | 1091.4 GB |
| Q3_K_M | 769.6 GB | 0.10 GB | 862.1 GB |
| Q2_K | 609.6 GB | 0.10 GB | 682.9 GB |
| NVFP4cuda | 800.0 GB | 0.10 GB | 896.1 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run DeepSeek V4 Pro 1.6T natively (0)
No single GPU in our list fits this model at Q4 with 8k context. Browse all GPUs or try multi-GPU setups.
Notes
1.6T MoE with hybrid CSA/HCA attention and 1M token context. Requires 27% of V3.2's inference FLOPs and 10% of its KV cache at 1M context; kvHeads/headDim approximates the width of one compressed KV entry. Superseded by the DeepSeek V4 Pro 0813 build, which keeps this architecture unchanged and re-post-trains it for agentic work.
Compare DeepSeek V4 Pro 1.6T with other models
Frequently asked questions
- What are the VRAM requirements for DeepSeek V4 Pro 1.6T?
- DeepSeek V4 Pro 1.6T requires approximately 1091.4 GB of VRAM at Q4_K_M quantization, 1905.0 GB at Q8, and 3584.1 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does DeepSeek V4 Pro 1.6T have?
- DeepSeek V4 Pro 1.6T has 1600 billion total parameters, but only 49 billion are active per token thanks to its Mixture of Experts (MoE) architecture. This makes inference significantly faster than the total parameter count suggests.
- How capable is DeepSeek V4 Pro 1.6T?
- DeepSeek V4 Pro 1.6T achieves an MMLU-Pro score of 87.5, placing it among the most capable open-weight models available, competitive with frontier systems on general knowledge and reasoning.
- Can DeepSeek V4 Pro 1.6T run on a 16 GB GPU?
- No. At Q4_K_M, DeepSeek V4 Pro 1.6T needs 1091.4 GB of VRAM, more than 16 GB. You will need a multi-GPU server.
- Can DeepSeek V4 Pro 1.6T run on a 24 GB GPU?
- No. Even at Q4_K_M, DeepSeek V4 Pro 1.6T needs 1091.4 GB. Consider a multi-GPU server with 1092 GB+ of combined VRAM.
- What is the smallest quantization for DeepSeek V4 Pro 1.6T that fits in 24 GB of VRAM?
- DeepSeek V4 Pro 1.6T cannot fit in 24 GB of VRAM at any standard quantization level. The minimum needed is 682.9 GB at Q2_K.
- What GPU do I need to run DeepSeek V4 Pro 1.6T locally?
- You need a multi-GPU server. At Q4_K_M, DeepSeek V4 Pro 1.6T needs 1091.4 GB VRAM, more than any single consumer GPU. That's roughly 14x 80 GB datacenter GPUs (H100, A100, or similar) pooled together.