DeepSeek R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B needs roughly 50.8 GB VRAM at Q4_K_M quantization (159.8 GB at FP16). 40 GPUs we track can run it fully in VRAM at 8k context.
40 GPUs run this natively · 35 with CPU offload
DeepSeek R1 Distill Llama 70B is a 70B parameter dense model developed by DeepSeek. 70B distillation of DeepSeek-R1's reasoning capabilities into Llama-3.3 architecture.
To run DeepSeek R1 Distill Llama 70B locally: Q4_K_M ~35-40GB — same requirements as Llama-3.3-70B. Best way to get R1-style reasoning locally.
MMLU-Pro 70.0%, GPQA 65.2%, Math 94.5% — inherits R1's reasoning strength at practical size.
VRAM at each quantization
DeepSeek R1 Distill Llama 70B natively supports a longer context window, but the table below is capped at 8k for comparability — KV cache grows linearly with context length.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 280.0 GB | 2.68 GB | 316.6 GB |
| BF16 | 140.0 GB | 2.68 GB | 159.8 GB |
| FP16 | 140.0 GB | 2.68 GB | 159.8 GB |
| Q8_0 | 74.4 GB | 2.68 GB | 86.3 GB |
| Q6_K | 57.5 GB | 2.68 GB | 67.4 GB |
| Q5_K_M | 49.8 GB | 2.68 GB | 58.8 GB |
| Q4_K_Mrec | 42.6 GB | 2.68 GB | 50.8 GB |
| Q3_K_M | 33.7 GB | 2.68 GB | 40.7 GB |
| Q2_K | 26.7 GB | 2.68 GB | 32.9 GB |
| NVFP4cuda | 35.0 GB | 2.68 GB | 42.2 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run DeepSeek R1 Distill Llama 70B natively (40)
- NVIDIA H100 80GBNVFP4 · 57.8 t/s
- NVIDIA A100 80GBNVFP4 · 35.2 t/s
- NVIDIA A100 40GBQ2_K · 34.4 t/s
- NVIDIA L40SNVFP4 · 14.9 t/s
- NVIDIA RTX A6000NVFP4 · 13.2 t/s
- NVIDIA RTX 6000 AdaNVFP4 · 16.6 t/s
- NVIDIA RTX Pro 6000NVFP4 · 23.2 t/s
- NVIDIA DGX Spark (128GB)NVFP4 · 4.7 t/s
- AMD Radeon PRO W7900Q3_K_M · 15.4 t/s
- AMD Instinct MI300XBF16 · 24.1 t/s
- AMD Strix Halo (128GB)Q8_0 · 2.2 t/s
- AMD Strix Halo (96GB)Q8_0 · 2.2 t/s
- AMD Strix Halo (64GB)Q4_K_M · 3.7 t/s
- Apple M5 Max (128GB)Q8_0 · 6.4 t/s
- Apple M5 Max (64GB)Q4_K_M · 10.8 t/s
- Apple M5 Max (48GB)Q2_K · 16.7 t/s
- Apple M5 Pro (48GB)Q2_K · 8.4 t/s
- Apple M4 Ultra (384GB)FP32 · 3.1 t/s
- Apple M4 Ultra (192GB)BF16 · 6.1 t/s
- Apple M4 Max (128GB)Q8_0 · 5.7 t/s
- Apple M4 Max (96GB)Q8_0 · 5.7 t/s
- Apple M4 Max (64GB)Q4_K_M · 9.6 t/s
- Apple M4 Max (48GB)Q2_K · 14.9 t/s
- Apple M4 Pro (48GB)Q2_K · 7.4 t/s
- Apple M3 Ultra (512GB)FP32 · 2.3 t/s
- Apple M3 Ultra (256GB)BF16 · 4.6 t/s
- Apple M3 Ultra (96GB)Q8_0 · 8.5 t/s
- Apple M3 Max (128GB)Q8_0 · 4.2 t/s
- Apple M3 Max (96GB)Q8_0 · 4.2 t/s
- Apple M3 Max (64GB)Q4_K_M · 7.1 t/s
- Apple M3 Max (48GB)Q2_K · 10.9 t/s
- Apple M2 Ultra (384GB)FP32 · 2.3 t/s
- Apple M2 Ultra (192GB)BF16 · 4.5 t/s
- Apple M2 Max (96GB)Q8_0 · 4.2 t/s
- Apple M2 Max (64GB)Q4_K_M · 7.1 t/s
- Apple M1 Ultra (128GB)Q8_0 · 8.3 t/s
- Apple M1 Ultra (64GB)Q4_K_M · 14.1 t/s
- Apple M1 Max (64GB)Q4_K_M · 7.1 t/s
- Intel Data Center GPU Max 1550Q8_0 · 27.6 t/s
- Intel Data Center GPU Max 1100Q3_K_M · 22 t/s
Plus 35 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5090NVFP4 · 3.1 t/s
- NVIDIA RTX 5080Q3_K_M · 1.1 t/s
- NVIDIA RTX 5070 TiQ3_K_M · 1.1 t/s
- NVIDIA RTX 5070Q2_K · 1.3 t/s
- NVIDIA RTX 5060 Ti 16GBQ3_K_M · 1.1 t/s
- NVIDIA RTX 5060Q2_K · 1.1 t/s
- NVIDIA RTX 5050Q2_K · 1.1 t/s
- NVIDIA RTX 4090NVFP4 · 1.6 t/s
- NVIDIA RTX 4080Q3_K_M · 1.1 t/s
- NVIDIA RTX 4070 TiQ2_K · 1.3 t/s
- NVIDIA RTX 4070Q2_K · 1.3 t/s
- NVIDIA RTX 4060 Ti 16GBQ3_K_M · 1.1 t/s
- NVIDIA RTX 4060Q2_K · 1.1 t/s
- NVIDIA RTX 3090NVFP4 · 1.6 t/s
- NVIDIA RTX 3090 TiNVFP4 · 1.6 t/s
- NVIDIA RTX 3080 10GBQ2_K · 1.2 t/s
- NVIDIA RTX 3060 12GBQ2_K · 1.3 t/s
- NVIDIA RTX 4000 AdaNVFP4 · 1.2 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 1.5 t/s
- NVIDIA RTX 5000 AdaNVFP4 · 2.7 t/s
- AMD Radeon RX 7900 XTXQ3_K_M · 1.7 t/s
- AMD Radeon RX 7900 XTQ3_K_M · 1.4 t/s
- AMD Radeon RX 7900 GREQ3_K_M · 1.1 t/s
- AMD Radeon RX 6800 XTQ3_K_M · 1.1 t/s
- AMD Radeon PRO W7800Q4_K_M · 1.5 t/s
- AMD Radeon AI Pro 9700 32GBQ4_K_M · 1.5 t/s
- Intel Arc B580 12GBQ2_K · 1.3 t/s
- Intel Arc B570 10GBQ2_K · 1.2 t/s
- Intel Arc Pro B70 24GBQ3_K_M · 1.6 t/s
- Intel Arc Pro B60 24GBQ3_K_M · 1.6 t/s
- Intel Arc A770 16GBQ3_K_M · 1.1 t/s
- Intel Arc A770 8GBQ2_K · 1.1 t/s
- Intel Arc A750 8GBQ2_K · 1.1 t/s
- Intel Arc A580 8GBQ2_K · 1.1 t/s
- Intel Arc Pro A60 12GBQ2_K · 1.3 t/s
Notes
Reasoning model — outputs long chains-of-thought before answering.
Compare DeepSeek R1 Distill Llama 70B with other models
How to run DeepSeek R1 Distill Llama 70B locally
Q4_K_M needs 50.8 GB — needs a workstation or datacenter GPU (48–80 GB).
Ollama
ollama run deepseek-r1:70bllama.cpp
./llama-cli -m deepseek-r1-distill-llama-70b.Q4_K_M.gguf -c 8192 -ngl 99LM Studio: Search for 'DeepSeek R1 Distill Llama 70B' in LM Studio. Requires 40+ GB of VRAM at Q4 -- plan for dual GPUs or a workstation card.
Why this quantization? This 70B dense model distills DeepSeek R1's reasoning into the Llama architecture. Q4_K_M at roughly 40 GB for weights is the minimum viable quantization for fitting on prosumer hardware. The model's extraordinary MATH score (94.5) and strong GPQA (65.2) are well-preserved at Q4. Grouped-query attention with 8 KV heads keeps the KV cache manageable alongside the heavy weight footprint.
Who is DeepSeek R1 Distill Llama 70B for?
Users with 48 GB+ VRAM setups who want the absolute best open-weight reasoning model that runs locally without a server rack. This distill outperforms the 32B variant on most benchmarks, making it the choice when you have the hardware to support it and need maximum reasoning capability.
Best for
- Olympiad-level mathematics and formal proofs (MATH: 94.5)
- Advanced scientific reasoning and research (GPQA: 65.2)
- Complex code generation with multi-step reasoning chains
- Situations where the 32B distill's accuracy isn't quite sufficient
Not ideal for
- Users with a single 24 GB GPU -- the 32B distill is your model
- Latency-sensitive applications -- 70B parameters means slower inference than the 32B variant
- General-purpose chat where reasoning overhead is unnecessary
Continue reading
Frequently asked questions
- What are the VRAM requirements for DeepSeek R1 Distill Llama 70B?
- DeepSeek R1 Distill Llama 70B requires approximately 50.8 GB of VRAM at Q4_K_M quantization, 86.3 GB at Q8, and 159.8 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does DeepSeek R1 Distill Llama 70B have?
- DeepSeek R1 Distill Llama 70B has 70 billion parameters.
- Is DeepSeek R1 Distill Llama 70B good at reasoning and math?
- Yes. With a MATH score of 94.5 and MMLU-Pro of 70, DeepSeek R1 Distill Llama 70B handles complex multi-step reasoning, analytical tasks, and problem-solving well.
- Can DeepSeek R1 Distill Llama 70B run on a 16 GB GPU?
- No. At Q4_K_M, DeepSeek R1 Distill Llama 70B needs 50.8 GB of VRAM — more than 16 GB. You will need a multi-GPU server.
- Can DeepSeek R1 Distill Llama 70B run on a 24 GB GPU?
- No. Even at Q4_K_M, DeepSeek R1 Distill Llama 70B needs 50.8 GB. Consider a multi-GPU server with 80 GB+ total VRAM.
- What is the smallest quantization for DeepSeek R1 Distill Llama 70B that fits in 24 GB of VRAM?
- DeepSeek R1 Distill Llama 70B cannot fit in 24 GB of VRAM at any standard quantization level. The minimum needed is 32.9 GB at Q2_K.
- What GPU do I need to run DeepSeek R1 Distill Llama 70B locally?
- You need a multi-GPU server. At Q4_K_M, DeepSeek R1 Distill Llama 70B needs 50.8 GB VRAM, more than any single consumer GPU. Consider 2–4× H100 or A100 GPUs.