Mixtral 8x22B Instruct v0.1
Mixtral 8x22B Instruct v0.1 needs roughly 98.3 GB VRAM at Q4_K_M quantization (317.9 GB at FP16). 22 GPUs we track can run it fully in VRAM at 8k context.
22 GPUs run this natively · 6 with CPU offload
Mixtral 8x22B Instruct v0.1 is a Mixture of Experts (MoE) model with 141B total parameters but only 39B active per token developed by Mistral AI. April 2024 large MoE — 141B total parameters with ~39B active. 64K context.
To run Mixtral 8x22B Instruct v0.1 locally: Q4_K_M ~80-90GB — requires 80GB GPU or dual 48GB. Multi-GPU or Mac Studio territory. As a MoE model, inference speed depends on active parameters (39B) rather than total size.
Significant upgrade over 8x7B in reasoning, coding, and knowledge. MMLU-Pro 40.0%, HumanEval 76.2%.
VRAM at each quantization
Mixtral 8x22B Instruct v0.1 natively supports a longer context window, but the table below is capped at 8k for comparability — KV cache grows linearly with context length.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 564.0 GB | 1.88 GB | 633.8 GB |
| BF16 | 282.0 GB | 1.88 GB | 317.9 GB |
| FP16 | 282.0 GB | 1.88 GB | 317.9 GB |
| Q8_0 | 149.9 GB | 1.88 GB | 170.0 GB |
| Q6_K | 115.8 GB | 1.88 GB | 131.8 GB |
| Q5_K_M | 100.4 GB | 1.88 GB | 114.5 GB |
| Q4_K_Mrec | 85.9 GB | 1.88 GB | 98.3 GB |
| Q3_K_M | 67.8 GB | 1.88 GB | 78.1 GB |
| Q2_K | 53.7 GB | 1.88 GB | 62.3 GB |
| NVFP4cuda | 70.5 GB | 1.88 GB | 81.1 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run Mixtral 8x22B Instruct v0.1 natively (22)
- NVIDIA H100 80GBQ2_K · 42.4 t/s
- NVIDIA A100 80GBQ2_K · 25.8 t/s
- NVIDIA RTX Pro 6000NVFP4 · 13.1 t/s
- NVIDIA DGX Spark (128GB)NVFP4 · 2.7 t/s
- AMD Instinct MI300XQ8_0 · 24.6 t/s
- AMD Strix Halo (128GB)Q5_K_M · 1.8 t/s
- AMD Strix Halo (96GB)Q3_K_M · 2.6 t/s
- Apple M5 Max (128GB)Q5_K_M · 5.2 t/s
- Apple M4 Ultra (384GB)BF16 · 3.3 t/s
- Apple M4 Ultra (192GB)Q8_0 · 6.2 t/s
- Apple M4 Max (128GB)Q5_K_M · 4.6 t/s
- Apple M4 Max (96GB)Q3_K_M · 6.8 t/s
- Apple M3 Ultra (512GB)BF16 · 2.5 t/s
- Apple M3 Ultra (256GB)Q8_0 · 4.7 t/s
- Apple M3 Ultra (96GB)Q3_K_M · 10.2 t/s
- Apple M3 Max (128GB)Q5_K_M · 3.4 t/s
- Apple M3 Max (96GB)Q3_K_M · 5 t/s
- Apple M2 Ultra (384GB)BF16 · 2.4 t/s
- Apple M2 Ultra (192GB)Q8_0 · 4.6 t/s
- Apple M2 Max (96GB)Q3_K_M · 5 t/s
- Apple M1 Ultra (128GB)Q5_K_M · 6.8 t/s
- Intel Data Center GPU Max 1550Q5_K_M · 22.5 t/s
Plus 6 GPUs that run it with CPU offload (slower)
- NVIDIA A100 40GBQ2_K · 1.5 t/s
- NVIDIA L40SQ2_K · 2.2 t/s
- NVIDIA RTX A6000Q2_K · 2.2 t/s
- NVIDIA RTX 6000 AdaQ2_K · 2.3 t/s
- AMD Radeon PRO W7900Q2_K · 2.2 t/s
- Intel Data Center GPU Max 1100Q2_K · 2.4 t/s
Notes
MoE: 141B total / 39B active — needs a lot of VRAM but runs fast when it fits.
Compare Mixtral 8x22B Instruct v0.1 with other models
Continue reading
Frequently asked questions
- What are the VRAM requirements for Mixtral 8x22B Instruct v0.1?
- Mixtral 8x22B Instruct v0.1 requires approximately 98.3 GB of VRAM at Q4_K_M quantization, 170.0 GB at Q8, and 317.9 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Mixtral 8x22B Instruct v0.1 have?
- Mixtral 8x22B Instruct v0.1 has 141 billion total parameters, but only 39 billion are active per token thanks to its Mixture of Experts (MoE) architecture. This makes inference significantly faster than the total parameter count suggests.
- How capable is Mixtral 8x22B Instruct v0.1?
- Mixtral 8x22B Instruct v0.1 has an MMLU-Pro score of 40, making it well-suited for lightweight tasks, prototyping, and resource-constrained environments.
- Can Mixtral 8x22B Instruct v0.1 run on a 16 GB GPU?
- No. At Q4_K_M, Mixtral 8x22B Instruct v0.1 needs 98.3 GB of VRAM — more than 16 GB. You will need a multi-GPU server.
- Can Mixtral 8x22B Instruct v0.1 run on a 24 GB GPU?
- No. Even at Q4_K_M, Mixtral 8x22B Instruct v0.1 needs 98.3 GB. Consider a multi-GPU server with 80 GB+ total VRAM.
- What is the smallest quantization for Mixtral 8x22B Instruct v0.1 that fits in 24 GB of VRAM?
- Mixtral 8x22B Instruct v0.1 cannot fit in 24 GB of VRAM at any standard quantization level. The minimum needed is 62.3 GB at Q2_K.
- What GPU do I need to run Mixtral 8x22B Instruct v0.1 locally?
- You need a multi-GPU server. At Q4_K_M, Mixtral 8x22B Instruct v0.1 needs 98.3 GB VRAM, more than any single consumer GPU. Consider 2–4× H100 or A100 GPUs.