Gemma 2 9B Instruct
Gemma 2 9B Instruct needs roughly 9.4 GB VRAM at Q4_K_M quantization (23.8 GB at FP16). 111 GPUs we track can run it fully in VRAM at 8k context.
111 GPUs run this natively · 5 with CPU offload
Gemma 2 9B Instruct is a 9.2B parameter dense model developed by Google. June 2024 9B model with knowledge distillation from 27B teacher, best performance for its size class.
To run Gemma 2 9B Instruct locally: Q5_K_M ~6-7GB and runs on 8GB GPUs. Excellent quality-per-VRAM ratio.
MMLU-Pro 32.0%, competitive with models 2-3× larger. Trained 50× beyond compute-optimal.
VRAM at each quantization
The table below assumes 8k context. KV cache size scales linearly with how much context you use.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 36.8 GB | 2.82 GB | 44.4 GB |
| BF16 | 18.4 GB | 2.82 GB | 23.8 GB |
| FP16 | 18.4 GB | 2.82 GB | 23.8 GB |
| Q8_0 | 9.8 GB | 2.82 GB | 14.1 GB |
| Q6_K | 7.5 GB | 2.82 GB | 11.6 GB |
| Q5_K_Mrec | 6.5 GB | 2.82 GB | 10.5 GB |
| Q4_K_M | 5.6 GB | 2.82 GB | 9.4 GB |
| Q3_K_M | 4.4 GB | 2.82 GB | 8.1 GB |
| Q2_K | 3.5 GB | 2.82 GB | 7.1 GB |
| NVFP4cuda | 4.6 GB | 2.82 GB | 8.3 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Gemma 2 9B Instruct natively (111)
- NVIDIA RTX 5090BF16 · 54.9 t/s
- NVIDIA RTX 5080NVFP4 · 84.1 t/s
- NVIDIA RTX 5070 TiNVFP4 · 78.5 t/s
- NVIDIA RTX 5070NVFP4 · 58.9 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 39.3 t/s
Show 106 more
- NVIDIA RTX 5060 Ti 8GBQ2_K · 46 t/s
- NVIDIA RTX 5060Q2_K · 46 t/s
- NVIDIA RTX 5050Q2_K · 32.9 t/s
- NVIDIA RTX 4090Q8_0 · 52 t/s
- NVIDIA RTX 4080Q8_0 · 37 t/s
- NVIDIA RTX 4070 Ti SUPERQ8_0 · 34.7 t/s
- NVIDIA RTX 4070 TiQ5_K_M · 35 t/s
- NVIDIA RTX 4070 SUPERQ5_K_M · 35 t/s
- NVIDIA RTX 4070Q5_K_M · 35 t/s
- NVIDIA RTX 4060 Ti 16GBQ8_0 · 14.9 t/s
- NVIDIA RTX 4060Q2_K · 28 t/s
- NVIDIA RTX 3090Q8_0 · 48.3 t/s
- NVIDIA RTX 3090 TiQ8_0 · 52 t/s
- NVIDIA RTX 3080 10GBQ4_K_M · 58.7 t/s
- NVIDIA RTX 3060 12GBQ5_K_M · 25 t/s
- NVIDIA B300 288GBBF16 · 245.1 t/s
- NVIDIA B200 180GBBF16 · 245.1 t/s
- NVIDIA H200 141GBBF16 · 147 t/s
- NVIDIA H100 80GBBF16 · 102.6 t/s
- NVIDIA A100 80GBBF16 · 62.5 t/s
- NVIDIA A100 40GBBF16 · 47.6 t/s
- NVIDIA L40SBF16 · 26.5 t/s
- NVIDIA RTX A6000BF16 · 23.5 t/s
- NVIDIA RTX 4000 AdaQ8_0 · 16.5 t/s
- NVIDIA RTX 4500 AdaQ8_0 · 22.3 t/s
- NVIDIA RTX 5000 AdaBF16 · 17.6 t/s
- NVIDIA RTX 6000 AdaBF16 · 29.4 t/s
- NVIDIA RTX Pro 6000BF16 · 41.2 t/s
- NVIDIA DGX Spark (128GB)BF16 · 8.4 t/s
- AMD Radeon RX 7900 XTXQ8_0 · 49.5 t/s
- AMD Radeon RX 7900 XTQ8_0 · 41.3 t/s
- AMD Radeon RX 7900 GREQ8_0 · 29.7 t/s
- AMD Radeon RX 6800 XTQ8_0 · 26.4 t/s
- AMD Radeon PRO W7800BF16 · 17.6 t/s
- AMD Radeon PRO W7900BF16 · 26.5 t/s
- AMD Instinct MI300XBF16 · 162.4 t/s
- AMD Radeon AI PRO R9700 32GBBF16 · 19.6 t/s
- AMD Strix Halo (128GB)BF16 · 7.8 t/s
- AMD Strix Halo (96GB)BF16 · 7.8 t/s
- AMD Strix Halo (64GB)BF16 · 7.8 t/s
- AMD Strix Halo (32GB)BF16 · 7.8 t/s
- Apple M5 Ultra (512GB)BF16 · 45.2 t/s
- Apple M5 Ultra (256GB)BF16 · 45.2 t/s
- Apple M5 Ultra (96GB)BF16 · 45.2 t/s
- Apple M5 Max (128GB)BF16 · 23.1 t/s
- Apple M5 Max (64GB)BF16 · 23.1 t/s
- Apple M5 Max (48GB)BF16 · 23.1 t/s
- Apple M5 Max (36GB)BF16 · 17.3 t/s
- Apple M5 Pro (64GB)BF16 · 11.6 t/s
- Apple M5 Pro (48GB)BF16 · 11.6 t/s
- Apple M5 Pro (24GB)Q8_0 · 19.5 t/s
- Apple M5 (32GB)BF16 · 5.8 t/s
- Apple M5 (16GB)Q2_K · 19.4 t/s
- Apple M6 (32GB)BF16 · 6.4 t/s
- Apple M6 (16GB)Q2_K · 21.5 t/s
- Apple M4 Max (128GB)BF16 · 20.6 t/s
- Apple M4 Max (64GB)BF16 · 20.6 t/s
- Apple M4 Max (48GB)BF16 · 20.6 t/s
- Apple M4 Max (36GB)BF16 · 15.5 t/s
- Apple M4 Pro (48GB)BF16 · 10.3 t/s
- Apple M4 Pro (24GB)Q8_0 · 17.3 t/s
- Apple M4 (32GB)BF16 · 4.5 t/s
- Apple M4 (16GB)Q2_K · 15.2 t/s
- Apple M3 Ultra (512GB)BF16 · 30.9 t/s
- Apple M3 Ultra (256GB)BF16 · 30.9 t/s
- Apple M3 Ultra (96GB)BF16 · 30.9 t/s
- Apple M3 Max (128GB)BF16 · 15.1 t/s
- Apple M3 Max (96GB)BF16 · 11.3 t/s
- Apple M3 Max (64GB)BF16 · 15.1 t/s
- Apple M3 Max (48GB)BF16 · 15.1 t/s
- Apple M3 Max (36GB)BF16 · 11.3 t/s
- Apple M3 Pro (36GB)BF16 · 5.7 t/s
- Apple M3 Pro (18GB)Q4_K_M · 14.2 t/s
- Apple M3 (24GB)Q8_0 · 6.4 t/s
- Apple M3 (16GB)Q2_K · 12.7 t/s
- Apple M2 Ultra (192GB)BF16 · 30.2 t/s
- Apple M2 Ultra (64GB)BF16 · 30.2 t/s
- Apple M2 Max (96GB)BF16 · 15.1 t/s
- Apple M2 Max (64GB)BF16 · 15.1 t/s
- Apple M2 Max (32GB)BF16 · 15.1 t/s
- Apple M2 Pro (32GB)BF16 · 7.5 t/s
- Apple M2 Pro (16GB)Q2_K · 25.3 t/s
- Apple M2 (24GB)Q8_0 · 6.4 t/s
- Apple M2 (16GB)Q2_K · 12.7 t/s
- Apple M1 Ultra (128GB)BF16 · 30.2 t/s
- Apple M1 Ultra (64GB)BF16 · 30.2 t/s
- Apple M1 Max (64GB)BF16 · 15.1 t/s
- Apple M1 Max (32GB)BF16 · 15.1 t/s
- Apple M1 Pro (32GB)BF16 · 7.5 t/s
- Apple M1 Pro (16GB)Q2_K · 25.3 t/s
- Apple M1 (16GB)Q2_K · 8.6 t/s
- Intel Arc B580 12GBQ5_K_M · 31.6 t/s
- Intel Arc B570 10GBQ4_K_M · 29.3 t/s
- Intel Arc Pro B70 32GBBF16 · 18.6 t/s
- Intel Arc Pro B60 24GBQ8_0 · 19.6 t/s
- Intel Arc Pro B50 16GBQ8_0 · 11.6 t/s
- Intel Arc A770 16GBQ8_0 · 28.9 t/s
- Intel Arc A770 8GBQ2_K · 52.6 t/s
- Intel Arc A750 8GBQ2_K · 52.6 t/s
- Intel Arc A580 8GBQ2_K · 52.6 t/s
- Intel Arc Pro A60 12GBQ5_K_M · 26.6 t/s
- Intel Data Center GPU Max 1550BF16 · 100.4 t/s
- Intel Data Center GPU Max 1100BF16 · 37.6 t/s
- Intel Arc 140V (32GB)BF16 · 4.2 t/s
- Intel Arc 140V (16GB)Q2_K · 14.1 t/s
- Intel Arc 130V (16GB)Q2_K · 14.1 t/s
Plus 5 GPUs that run it with CPU offload (slower)
- Intel Arc A380 6GBBF16 · 1.5 t/s
- Intel Arc A310 4GBBF16 · 1.3 t/s
- Intel Arc Pro A50 6GBBF16 · 1.5 t/s
- Intel Arc Pro A40 6GBBF16 · 1.5 t/s
- CPU only (system RAM)BF16 · 1.9 t/s
Compare Gemma 2 9B Instruct with other models
Frequently asked questions
- What are the VRAM requirements for Gemma 2 9B Instruct?
- Gemma 2 9B Instruct requires approximately 9.4 GB of VRAM at Q4_K_M quantization, 14.1 GB at Q8, and 23.8 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 2 9B Instruct have?
- Gemma 2 9B Instruct has 9.2 billion parameters.
- How capable is Gemma 2 9B Instruct?
- Gemma 2 9B Instruct has an MMLU-Pro score of 32, making it well-suited for lightweight tasks, prototyping, and resource-constrained environments.
- Can Gemma 2 9B Instruct run on a 16 GB GPU?
- Yes. Gemma 2 9B Instruct needs 9.4 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 5070 Ti.
- What is the smallest quantization for Gemma 2 9B Instruct that fits in 24 GB of VRAM?
- At BF16, Gemma 2 9B Instruct needs 23.8 GB, the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 2 9B Instruct locally?
- A 16 GB GPU is enough. At Q4_K_M, Gemma 2 9B Instruct needs 9.4 GB VRAM. Good options: RTX 4080 (16 GB), RTX 5070 Ti (16 GB).