Mistral Small 22B
Mistral Small 22B needs roughly 17.3 GB VRAM at Q4_K_M quantization (51.8 GB at FP16). 84 GPUs we track can run it fully in VRAM at 8k context.
84 GPUs run this natively · 21 with CPU offload
Mistral Small 22B is a 22.2B parameter dense model developed by Mistral AI. September 2024 22B model with strong general capabilities.
To run Mistral Small 22B locally: Q4_K_M needs ~17.3GB total once KV cache and overhead are counted, past a 16GB card's ~15.2GB effective capacity -- a 24GB GPU is the real minimum, with comfortable headroom to spare.
MMLU-Pro 49.2%, HumanEval 81.1%, solid mid-range performer.
VRAM at each quantization
Numbers here are computed at 8k context. Because KV cache grows linearly with context length, expect higher totals at longer sequence lengths.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 88.8 GB | 1.88 GB | 101.6 GB |
| BF16 | 44.4 GB | 1.88 GB | 51.8 GB |
| FP16 | 44.4 GB | 1.88 GB | 51.8 GB |
| Q8_0 | 23.6 GB | 1.88 GB | 28.5 GB |
| Q6_K | 18.2 GB | 1.88 GB | 22.5 GB |
| Q5_K_M | 15.8 GB | 1.88 GB | 19.8 GB |
| Q4_K_Mrec | 13.5 GB | 1.88 GB | 17.3 GB |
| Q3_K_M | 10.7 GB | 1.88 GB | 14.1 GB |
| Q2_K | 8.5 GB | 1.88 GB | 11.6 GB |
| NVFP4cuda | 11.1 GB | 1.88 GB | 14.5 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Mistral Small 22B natively (84)
- NVIDIA RTX 5090NVFP4 · 89.7 t/s
- NVIDIA RTX 5080NVFP4 · 48.1 t/s
- NVIDIA RTX 5070 TiNVFP4 · 44.9 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 22.4 t/s
- NVIDIA RTX 4090Q6_K · 32.6 t/s
Show 79 more
- NVIDIA RTX 4080Q3_K_M · 37.1 t/s
- NVIDIA RTX 4070 Ti SUPERQ3_K_M · 34.8 t/s
- NVIDIA RTX 4060 Ti 16GBQ3_K_M · 14.9 t/s
- NVIDIA RTX 3090Q6_K · 30.3 t/s
- NVIDIA RTX 3090 TiQ6_K · 32.6 t/s
- NVIDIA B300 288GBBF16 · 112.4 t/s
- NVIDIA B200 180GBBF16 · 112.4 t/s
- NVIDIA H200 141GBBF16 · 67.4 t/s
- NVIDIA H100 80GBBF16 · 47.1 t/s
- NVIDIA A100 80GBBF16 · 28.6 t/s
- NVIDIA A100 40GBQ8_0 · 39.7 t/s
- NVIDIA L40SQ8_0 · 22 t/s
- NVIDIA RTX A6000Q8_0 · 19.6 t/s
- NVIDIA RTX 4000 AdaQ4_K_M · 13.5 t/s
- NVIDIA RTX 4500 AdaQ6_K · 14 t/s
- NVIDIA RTX 5000 AdaQ8_0 · 14.7 t/s
- NVIDIA RTX 6000 AdaQ8_0 · 24.5 t/s
- NVIDIA RTX Pro 6000BF16 · 18.9 t/s
- NVIDIA DGX Spark (128GB)BF16 · 3.8 t/s
- AMD Radeon RX 7900 XTXQ6_K · 31 t/s
- AMD Radeon RX 7900 XTQ4_K_M · 33.8 t/s
- AMD Radeon RX 7900 GREQ3_K_M · 29.8 t/s
- AMD Radeon RX 6800 XTQ3_K_M · 26.5 t/s
- AMD Radeon PRO W7800Q8_0 · 14.7 t/s
- AMD Radeon PRO W7900Q8_0 · 22 t/s
- AMD Instinct MI300XBF16 · 74.4 t/s
- AMD Radeon AI PRO R9700 32GBQ8_0 · 16.3 t/s
- AMD Strix Halo (128GB)BF16 · 3.6 t/s
- AMD Strix Halo (96GB)BF16 · 3.6 t/s
- AMD Strix Halo (64GB)BF16 · 3.6 t/s
- AMD Strix Halo (32GB)Q6_K · 8.3 t/s
- Apple M5 Ultra (512GB)BF16 · 20.7 t/s
- Apple M5 Ultra (256GB)BF16 · 20.7 t/s
- Apple M5 Ultra (96GB)BF16 · 20.7 t/s
- Apple M5 Max (128GB)BF16 · 10.6 t/s
- Apple M5 Max (64GB)BF16 · 10.6 t/s
- Apple M5 Max (48GB)Q8_0 · 19.3 t/s
- Apple M5 Max (36GB)Q6_K · 18.3 t/s
- Apple M5 Pro (64GB)BF16 · 5.3 t/s
- Apple M5 Pro (48GB)Q8_0 · 9.6 t/s
- Apple M5 Pro (24GB)Q3_K_M · 19.6 t/s
- Apple M5 (32GB)Q6_K · 6.1 t/s
- Apple M6 (32GB)Q6_K · 6.8 t/s
- Apple M4 Max (128GB)BF16 · 9.4 t/s
- Apple M4 Max (64GB)BF16 · 9.4 t/s
- Apple M4 Max (48GB)Q8_0 · 17.1 t/s
- Apple M4 Max (36GB)Q6_K · 16.3 t/s
- Apple M4 Pro (48GB)Q8_0 · 8.6 t/s
- Apple M4 Pro (24GB)Q3_K_M · 17.4 t/s
- Apple M4 (32GB)Q6_K · 4.8 t/s
- Apple M3 Ultra (512GB)BF16 · 14.2 t/s
- Apple M3 Ultra (256GB)BF16 · 14.2 t/s
- Apple M3 Ultra (96GB)BF16 · 14.2 t/s
- Apple M3 Max (128GB)BF16 · 6.9 t/s
- Apple M3 Max (96GB)BF16 · 5.2 t/s
- Apple M3 Max (64GB)BF16 · 6.9 t/s
- Apple M3 Max (48GB)Q8_0 · 12.6 t/s
- Apple M3 Max (36GB)Q6_K · 11.9 t/s
- Apple M3 Pro (36GB)Q6_K · 6 t/s
- Apple M3 (24GB)Q3_K_M · 6.4 t/s
- Apple M2 Ultra (192GB)BF16 · 13.8 t/s
- Apple M2 Ultra (64GB)BF16 · 13.8 t/s
- Apple M2 Max (96GB)BF16 · 6.9 t/s
- Apple M2 Max (64GB)BF16 · 6.9 t/s
- Apple M2 Max (32GB)Q6_K · 15.9 t/s
- Apple M2 Pro (32GB)Q6_K · 8 t/s
- Apple M2 (24GB)Q3_K_M · 6.4 t/s
- Apple M1 Ultra (128GB)BF16 · 13.8 t/s
- Apple M1 Ultra (64GB)BF16 · 13.8 t/s
- Apple M1 Max (64GB)BF16 · 6.9 t/s
- Apple M1 Max (32GB)Q6_K · 15.9 t/s
- Apple M1 Pro (32GB)Q6_K · 8 t/s
- Intel Arc Pro B70 32GBQ8_0 · 15.5 t/s
- Intel Arc Pro B60 24GBQ6_K · 12.3 t/s
- Intel Arc Pro B50 16GBQ3_K_M · 11.6 t/s
- Intel Arc A770 16GBQ3_K_M · 29 t/s
- Intel Data Center GPU Max 1550BF16 · 46 t/s
- Intel Data Center GPU Max 1100Q8_0 · 31.4 t/s
- Intel Arc 140V (32GB)Q6_K · 4.4 t/s
Plus 21 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5070NVFP4 · 11.2 t/s
- NVIDIA RTX 5060 Ti 8GBNVFP4 · 4.1 t/s
- NVIDIA RTX 5060NVFP4 · 4.1 t/s
- NVIDIA RTX 5050NVFP4 · 4 t/s
- NVIDIA RTX 4070 TiQ8_0 · 1.7 t/s
- NVIDIA RTX 4070 SUPERQ8_0 · 1.7 t/s
- NVIDIA RTX 4070Q8_0 · 1.7 t/s
- NVIDIA RTX 4060Q8_0 · 1.3 t/s
- NVIDIA RTX 3080 10GBQ8_0 · 1.5 t/s
- NVIDIA RTX 3060 12GBQ8_0 · 1.6 t/s
- Intel Arc B580 12GBQ8_0 · 1.7 t/s
- Intel Arc B570 10GBQ8_0 · 1.5 t/s
- Intel Arc A770 8GBQ8_0 · 1.4 t/s
- Intel Arc A750 8GBQ8_0 · 1.4 t/s
- Intel Arc A580 8GBQ8_0 · 1.4 t/s
- Intel Arc A380 6GBQ8_0 · 1.2 t/s
- Intel Arc A310 4GBQ8_0 · 1.1 t/s
- Intel Arc Pro A60 12GBQ8_0 · 1.6 t/s
- Intel Arc Pro A50 6GBQ8_0 · 1.2 t/s
- Intel Arc Pro A40 6GBQ8_0 · 1.2 t/s
- CPU only (system RAM)Q6_K · 2 t/s
Continue reading
Frequently asked questions
- What are the VRAM requirements for Mistral Small 22B?
- Mistral Small 22B requires approximately 17.3 GB of VRAM at Q4_K_M quantization, 28.5 GB at Q8, and 51.8 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Mistral Small 22B have?
- Mistral Small 22B has 22.2 billion parameters.
- How capable is Mistral Small 22B?
- Mistral Small 22B has an MMLU-Pro score of 49.2, making it well-suited for lightweight tasks, prototyping, and resource-constrained environments.
- Can Mistral Small 22B run on a 16 GB GPU?
- No. At Q4_K_M, Mistral Small 22B needs 17.3 GB of VRAM, more than 16 GB. You will need a 24 GB GPU like the RTX 4090 or RTX 3090.
- Can Mistral Small 22B run on a 24 GB GPU?
- Yes. Mistral Small 22B fits in a 24 GB GPU at Q4_K_M, requiring 17.3 GB VRAM. GPUs with 24 GB include the RTX 4090, RTX 3090, and RTX 3090 Ti.
- What is the smallest quantization for Mistral Small 22B that fits in 24 GB of VRAM?
- At NVFP4, Mistral Small 22B needs 14.5 GB, the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Mistral Small 22B locally?
- A 24 GB GPU is the minimum. At Q4_K_M, Mistral Small 22B needs 17.3 GB VRAM. Good options: RTX 4090 (24 GB), RTX 3090 (24 GB).