Gemma 4 31B
Gemma 4 31B needs roughly 24.8 GB VRAM at Q4_K_M quantization (73.0 GB at FP16). 63 GPUs we track can run it fully in VRAM at 8k context.
63 GPUs run this natively · 27 with CPU offload
Gemma 4 31B is a 31B parameter dense model developed by Google. April 2026 dense model optimized for workstations and servers. 256K context with multimodal support.
To run Gemma 4 31B locally: Q4_K_M ~18-20GB — fits on 24GB GPUs. Good choice for M4 Max or RTX 4090 owners.
31B dense architecture with vision capabilities — Google's workstation-focused offering.
VRAM at each quantization
Numbers here are computed at 8k context. Because KV cache grows linearly with context length, expect higher totals at longer sequence lengths.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 124.0 GB | 3.22 GB | 142.5 GB |
| BF16 | 62.0 GB | 3.22 GB | 73.0 GB |
| FP16 | 62.0 GB | 3.22 GB | 73.0 GB |
| Q8_0 | 33.0 GB | 3.22 GB | 40.5 GB |
| Q6_K | 25.4 GB | 3.22 GB | 32.1 GB |
| Q5_K_M | 22.1 GB | 3.22 GB | 28.3 GB |
| Q4_K_Mrec | 18.9 GB | 3.22 GB | 24.8 GB |
| Q3_K_M | 14.9 GB | 3.22 GB | 20.3 GB |
| Q2_K | 11.8 GB | 3.22 GB | 16.8 GB |
| NVFP4cuda | 15.5 GB | 3.22 GB | 21.0 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Gemma 4 31B natively (63)
- NVIDIA RTX 5090NVFP4 · 62.2 t/s
- NVIDIA RTX 4090NVFP4 · 35 t/s
- NVIDIA RTX 3090NVFP4 · 32.5 t/s
- NVIDIA RTX 3090 TiNVFP4 · 35 t/s
- NVIDIA H100 80GBBF16 · 33.4 t/s
- NVIDIA A100 80GBBF16 · 20.3 t/s
- NVIDIA A100 40GBNVFP4 · 54 t/s
- NVIDIA L40SNVFP4 · 30 t/s
- NVIDIA RTX A6000NVFP4 · 26.7 t/s
- NVIDIA RTX 4000 AdaQ2_K · 13.8 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 15 t/s
- NVIDIA RTX 5000 AdaNVFP4 · 20 t/s
- NVIDIA RTX 6000 AdaNVFP4 · 33.3 t/s
- NVIDIA RTX Pro 6000BF16 · 13.4 t/s
- NVIDIA DGX Spark (128GB)BF16 · 2.7 t/s
- AMD Radeon RX 7900 XTXQ3_K_M · 34.4 t/s
- AMD Radeon RX 7900 XTQ2_K · 34.6 t/s
- AMD Radeon PRO W7800Q5_K_M · 14.8 t/s
- AMD Radeon PRO W7900Q8_0 · 15.5 t/s
- AMD Instinct MI300XFP32 · 27.1 t/s
- AMD Radeon AI Pro 9700 32GBQ5_K_M · 16.4 t/s
- AMD Strix Halo (128GB)BF16 · 2.6 t/s
- AMD Strix Halo (96GB)BF16 · 2.6 t/s
- AMD Strix Halo (64GB)Q8_0 · 4.6 t/s
- Apple M5 Max (128GB)BF16 · 7.5 t/s
- Apple M5 Max (64GB)Q8_0 · 13.6 t/s
- Apple M5 Max (48GB)Q6_K · 17.1 t/s
- Apple M5 Pro (48GB)Q6_K · 8.6 t/s
- Apple M5 Pro (36GB)Q4_K_M · 11.1 t/s
- Apple M5 (32GB)Q3_K_M · 6.8 t/s
- Apple M4 Ultra (384GB)FP32 · 6.9 t/s
- Apple M4 Ultra (192GB)FP32 · 6.9 t/s
- Apple M4 Max (128GB)BF16 · 6.7 t/s
- Apple M4 Max (96GB)BF16 · 6.7 t/s
- Apple M4 Max (64GB)Q8_0 · 12.1 t/s
- Apple M4 Max (48GB)Q6_K · 15.2 t/s
- Apple M4 Pro (48GB)Q6_K · 7.6 t/s
- Apple M4 (32GB)Q3_K_M · 5.3 t/s
- Apple M3 Ultra (512GB)FP32 · 5.2 t/s
- Apple M3 Ultra (256GB)FP32 · 5.2 t/s
- Apple M3 Ultra (96GB)BF16 · 10 t/s
- Apple M3 Max (128GB)BF16 · 4.9 t/s
- Apple M3 Max (96GB)BF16 · 4.9 t/s
- Apple M3 Max (64GB)Q8_0 · 8.8 t/s
- Apple M3 Max (48GB)Q6_K · 11.2 t/s
- Apple M3 Max (36GB)Q4_K_M · 14.5 t/s
- Apple M3 Pro (36GB)Q4_K_M · 5.4 t/s
- Apple M2 Ultra (384GB)FP32 · 5 t/s
- Apple M2 Ultra (192GB)FP32 · 5 t/s
- Apple M2 Max (96GB)BF16 · 4.9 t/s
- Apple M2 Max (64GB)Q8_0 · 8.8 t/s
- Apple M2 Max (32GB)Q3_K_M · 17.6 t/s
- Apple M2 Pro (32GB)Q3_K_M · 8.8 t/s
- Apple M1 Ultra (128GB)BF16 · 9.8 t/s
- Apple M1 Ultra (64GB)Q8_0 · 17.7 t/s
- Apple M1 Max (64GB)Q8_0 · 8.8 t/s
- Apple M1 Max (32GB)Q3_K_M · 17.6 t/s
- Apple M1 Pro (32GB)Q3_K_M · 8.8 t/s
- Intel Arc Pro B70 24GBQ3_K_M · 16.3 t/s
- Intel Arc Pro B60 24GBQ3_K_M · 13.6 t/s
- Intel Data Center GPU Max 1550BF16 · 32.6 t/s
- Intel Data Center GPU Max 1100Q8_0 · 22.1 t/s
- Intel Arc 140V (32GB)Q3_K_M · 4.9 t/s
Plus 27 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5080NVFP4 · 6.1 t/s
- NVIDIA RTX 5070 TiNVFP4 · 6 t/s
- NVIDIA RTX 5070NVFP4 · 3.1 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 5.2 t/s
- NVIDIA RTX 5060NVFP4 · 2.1 t/s
- NVIDIA RTX 5050NVFP4 · 2.1 t/s
- NVIDIA RTX 4080NVFP4 · 5.8 t/s
- NVIDIA RTX 4070 TiNVFP4 · 3.1 t/s
- NVIDIA RTX 4070NVFP4 · 3.1 t/s
- NVIDIA RTX 4060 Ti 16GBNVFP4 · 4.5 t/s
- NVIDIA RTX 4060NVFP4 · 2 t/s
- NVIDIA RTX 3080 10GBNVFP4 · 2.6 t/s
- NVIDIA RTX 3060 12GBNVFP4 · 2.9 t/s
- AMD Radeon RX 7900 GREQ8_0 · 1.1 t/s
- AMD Radeon RX 6800 XTQ8_0 · 1.1 t/s
- Intel Arc B580 12GBQ6_K · 1.4 t/s
- Intel Arc B570 10GBQ6_K · 1.2 t/s
- Intel Arc A770 16GBQ8_0 · 1.1 t/s
- Intel Arc A770 8GBQ6_K · 1.2 t/s
- Intel Arc A750 8GBQ6_K · 1.2 t/s
- Intel Arc A580 8GBQ6_K · 1.2 t/s
- Intel Arc A380 6GBQ5_K_M · 1.2 t/s
- Intel Arc A310 4GBQ5_K_M · 1.1 t/s
- Intel Arc Pro A60 12GBQ6_K · 1.4 t/s
- Intel Arc Pro A50 6GBQ5_K_M · 1.2 t/s
- Intel Arc Pro A40 6GBQ5_K_M · 1.2 t/s
- CPU only (system RAM)Q4_K_M · 1.8 t/s
Notes
Dense model optimized for workstations/servers.
Compare Gemma 4 31B with other models
Continue reading
Frequently asked questions
- What are the VRAM requirements for Gemma 4 31B?
- Gemma 4 31B requires approximately 24.8 GB of VRAM at Q4_K_M quantization, 40.5 GB at Q8, and 73.0 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 4 31B have?
- Gemma 4 31B has 31 billion parameters.
- How capable is Gemma 4 31B?
- Gemma 4 31B achieves an MMLU-Pro score of 85.2, placing it among the most capable open-weight models available — competitive with frontier systems on general knowledge and reasoning.
- Can Gemma 4 31B run on a 16 GB GPU?
- No. At Q4_K_M, Gemma 4 31B needs 24.8 GB of VRAM — more than 16 GB. You will need a 48 GB GPU like the RTX 6000 Ada or a dual-GPU setup.
- Can Gemma 4 31B run on a 24 GB GPU?
- No. Even at Q4_K_M, Gemma 4 31B needs 24.8 GB. Consider a 48 GB card like the RTX 6000 Ada or a dual RTX 4090 setup.
- What is the smallest quantization for Gemma 4 31B that fits in 24 GB of VRAM?
- At NVFP4, Gemma 4 31B needs 21.0 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 4 31B locally?
- You need a 48 GB GPU or a dual-GPU setup. At Q4_K_M, Gemma 4 31B needs 24.8 GB VRAM. Options: RTX 6000 Ada (48 GB), A6000 (48 GB), or 2× RTX 4090.