Gemma 4 31B
Gemma 4 31B needs roughly 22.6 GB VRAM at Q4_K_M quantization (70.5 GB at FP16). 84 GPUs we track can run it fully in VRAM at 8k context.
84 GPUs run this natively · 21 with CPU offload
Gemma 4 31B is a 30.7B parameter dense model developed by Google. April 2026 dense model optimized for workstations and servers. 256K context with multimodal support.
To run Gemma 4 31B locally: Q4_K_M ~18-20GB and fits on 24GB GPUs. Good choice for M4 Max or RTX 4090 owners.
31B dense architecture with vision capabilities, Google's workstation-focused offering.
VRAM at each quantization
Numbers here are computed at 8k context. This model's hybrid attention stack means KV cache grows much more slowly than context length, unlike a conventional full-attention model.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 122.8 GB | 1.51 GB | 139.2 GB |
| BF16 | 61.4 GB | 1.51 GB | 70.5 GB |
| FP16 | 61.4 GB | 1.51 GB | 70.5 GB |
| Q8_0 | 32.6 GB | 1.51 GB | 38.2 GB |
| Q6_K | 25.2 GB | 1.51 GB | 29.9 GB |
| Q5_K_M | 21.9 GB | 1.51 GB | 26.2 GB |
| Q4_K_Mrec | 18.7 GB | 1.51 GB | 22.6 GB |
| Q3_K_M | 14.8 GB | 1.51 GB | 18.2 GB |
| Q2_K | 11.7 GB | 1.51 GB | 14.8 GB |
| NVFP4cuda | 15.3 GB | 1.51 GB | 18.9 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Gemma 4 31B natively (84)
- NVIDIA RTX 5090NVFP4 · 69.1 t/s
- NVIDIA RTX 5080Q2_K · 47.2 t/s
- NVIDIA RTX 5070 TiQ2_K · 44.1 t/s
- NVIDIA RTX 5060 Ti 16GBQ2_K · 22 t/s
- NVIDIA RTX 4090Q4_K_M · 32.4 t/s
Show 79 more
- NVIDIA RTX 4080Q2_K · 35.3 t/s
- NVIDIA RTX 4070 Ti SUPERQ2_K · 33.1 t/s
- NVIDIA RTX 4060 Ti 16GBQ2_K · 14.2 t/s
- NVIDIA RTX 3090Q4_K_M · 30.1 t/s
- NVIDIA RTX 3090 TiQ4_K_M · 32.4 t/s
- NVIDIA B300 288GBBF16 · 82.7 t/s
- NVIDIA B200 180GBBF16 · 82.7 t/s
- NVIDIA H200 141GBBF16 · 49.6 t/s
- NVIDIA H100 80GBBF16 · 34.6 t/s
- NVIDIA A100 80GBBF16 · 21.1 t/s
- NVIDIA A100 40GBQ6_K · 37.8 t/s
- NVIDIA L40SQ8_0 · 16.4 t/s
- NVIDIA RTX A6000Q8_0 · 14.6 t/s
- NVIDIA RTX 4000 AdaQ3_K_M · 12.8 t/s
- NVIDIA RTX 4500 AdaQ4_K_M · 13.9 t/s
- NVIDIA RTX 5000 AdaQ6_K · 14 t/s
- NVIDIA RTX 6000 AdaQ8_0 · 18.3 t/s
- NVIDIA RTX Pro 6000BF16 · 13.9 t/s
- NVIDIA DGX Spark (128GB)BF16 · 2.8 t/s
- AMD Radeon RX 7900 XTXQ4_K_M · 30.9 t/s
- AMD Radeon RX 7900 XTQ3_K_M · 31.9 t/s
- AMD Radeon RX 7900 GREQ2_K · 28.3 t/s
- AMD Radeon RX 6800 XTQ2_K · 25.2 t/s
- AMD Radeon PRO W7800Q6_K · 14 t/s
- AMD Radeon PRO W7900Q8_0 · 16.4 t/s
- AMD Instinct MI300XBF16 · 54.8 t/s
- AMD Radeon AI PRO R9700 32GBQ6_K · 15.6 t/s
- AMD Strix Halo (128GB)BF16 · 2.6 t/s
- AMD Strix Halo (96GB)BF16 · 2.6 t/s
- AMD Strix Halo (64GB)Q8_0 · 4.9 t/s
- AMD Strix Halo (32GB)Q4_K_M · 8.2 t/s
- Apple M5 Ultra (512GB)BF16 · 15.3 t/s
- Apple M5 Ultra (256GB)BF16 · 15.3 t/s
- Apple M5 Ultra (96GB)BF16 · 15.3 t/s
- Apple M5 Max (128GB)BF16 · 7.8 t/s
- Apple M5 Max (64GB)Q8_0 · 14.4 t/s
- Apple M5 Max (48GB)Q8_0 · 14.4 t/s
- Apple M5 Max (36GB)Q5_K_M · 15.7 t/s
- Apple M5 Pro (64GB)Q8_0 · 7.2 t/s
- Apple M5 Pro (48GB)Q8_0 · 7.2 t/s
- Apple M5 Pro (24GB)Q2_K · 18.6 t/s
- Apple M5 (32GB)Q4_K_M · 6.1 t/s
- Apple M6 (32GB)Q4_K_M · 6.7 t/s
- Apple M4 Max (128GB)BF16 · 6.9 t/s
- Apple M4 Max (64GB)Q8_0 · 12.8 t/s
- Apple M4 Max (48GB)Q8_0 · 12.8 t/s
- Apple M4 Max (36GB)Q5_K_M · 14 t/s
- Apple M4 Pro (48GB)Q8_0 · 6.4 t/s
- Apple M4 Pro (24GB)Q2_K · 16.5 t/s
- Apple M4 (32GB)Q4_K_M · 4.8 t/s
- Apple M3 Ultra (512GB)BF16 · 10.4 t/s
- Apple M3 Ultra (256GB)BF16 · 10.4 t/s
- Apple M3 Ultra (96GB)BF16 · 10.4 t/s
- Apple M3 Max (128GB)BF16 · 5.1 t/s
- Apple M3 Max (96GB)BF16 · 3.8 t/s
- Apple M3 Max (64GB)Q8_0 · 9.4 t/s
- Apple M3 Max (48GB)Q8_0 · 9.4 t/s
- Apple M3 Max (36GB)Q5_K_M · 10.3 t/s
- Apple M3 Pro (36GB)Q5_K_M · 5.1 t/s
- Apple M3 (24GB)Q2_K · 6.1 t/s
- Apple M2 Ultra (192GB)BF16 · 10.2 t/s
- Apple M2 Ultra (64GB)Q8_0 · 18.7 t/s
- Apple M2 Max (96GB)BF16 · 5.1 t/s
- Apple M2 Max (64GB)Q8_0 · 9.4 t/s
- Apple M2 Max (32GB)Q4_K_M · 15.8 t/s
- Apple M2 Pro (32GB)Q4_K_M · 7.9 t/s
- Apple M2 (24GB)Q2_K · 6.1 t/s
- Apple M1 Ultra (128GB)BF16 · 10.2 t/s
- Apple M1 Ultra (64GB)Q8_0 · 18.7 t/s
- Apple M1 Max (64GB)Q8_0 · 9.4 t/s
- Apple M1 Max (32GB)Q4_K_M · 15.8 t/s
- Apple M1 Pro (32GB)Q4_K_M · 7.9 t/s
- Intel Arc Pro B70 32GBQ6_K · 14.8 t/s
- Intel Arc Pro B60 24GBQ4_K_M · 12.2 t/s
- Intel Arc Pro B50 16GBQ2_K · 11 t/s
- Intel Arc A770 16GBQ2_K · 27.6 t/s
- Intel Data Center GPU Max 1550BF16 · 33.8 t/s
- Intel Data Center GPU Max 1100Q8_0 · 23.4 t/s
- Intel Arc 140V (32GB)Q4_K_M · 4.4 t/s
Plus 21 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5070NVFP4 · 4.1 t/s
- NVIDIA RTX 5060 Ti 8GBNVFP4 · 2.5 t/s
- NVIDIA RTX 5060NVFP4 · 2.5 t/s
- NVIDIA RTX 5050NVFP4 · 2.5 t/s
- NVIDIA RTX 4070 TiQ6_K · 1.5 t/s
- NVIDIA RTX 4070 SUPERQ6_K · 1.5 t/s
- NVIDIA RTX 4070Q6_K · 1.5 t/s
- NVIDIA RTX 4060Q6_K · 1.2 t/s
- NVIDIA RTX 3080 10GBQ6_K · 1.4 t/s
- NVIDIA RTX 3060 12GBQ6_K · 1.5 t/s
- Intel Arc B580 12GBQ6_K · 1.5 t/s
- Intel Arc B570 10GBQ6_K · 1.4 t/s
- Intel Arc A770 8GBQ6_K · 1.3 t/s
- Intel Arc A750 8GBQ6_K · 1.3 t/s
- Intel Arc A580 8GBQ6_K · 1.3 t/s
- Intel Arc A380 6GBQ6_K · 1.1 t/s
- Intel Arc A310 4GBQ5_K_M · 1.2 t/s
- Intel Arc Pro A60 12GBQ6_K · 1.5 t/s
- Intel Arc Pro A50 6GBQ6_K · 1.1 t/s
- Intel Arc Pro A40 6GBQ6_K · 1.1 t/s
- CPU only (system RAM)Q5_K_M · 1.7 t/s
Notes
Dense, not MoE: all 30.7B parameters run on every token. Like Gemma 3 and the sibling Gemma 4 26B MoE, its 60 attention layers repeat a 5-local/1-global sliding-window pattern (1,024-token window), so only 10 of 60 layers scale their KV cache with the full context; the other 50 stay capped regardless of how long the conversation runs. The 10 full-attention layers also use a narrower KV width (4 KV heads x 512 head_dim) than the 50 sliding-window layers (16 x 256). Native context is 256k tokens (262,144).
Compare Gemma 4 31B with other models
Continue reading
Frequently asked questions
- What are the VRAM requirements for Gemma 4 31B?
- Gemma 4 31B requires approximately 22.6 GB of VRAM at Q4_K_M quantization, 38.2 GB at Q8, and 70.5 GB at FP16. These numbers assume 8k context window; its hybrid attention stack caches far fewer than all layers, so VRAM grows much slower than linearly with context.
- How many parameters does Gemma 4 31B have?
- Gemma 4 31B has 30.7 billion parameters.
- Is Gemma 4 31B good for coding?
- Yes. Gemma 4 31B scores 80.0 on LiveCodeBench, demonstrating strong code generation and completion capabilities.
- Can Gemma 4 31B run on a 16 GB GPU?
- No. At Q4_K_M, Gemma 4 31B needs 22.6 GB of VRAM, more than 16 GB. You will need a 24 GB GPU like the RTX 4090 or RTX 3090.
- Can Gemma 4 31B run on a 24 GB GPU?
- Yes. Gemma 4 31B fits in a 24 GB GPU at Q4_K_M, requiring 22.6 GB VRAM. GPUs with 24 GB include the RTX 4090, RTX 3090, and RTX 3090 Ti.
- What is the smallest quantization for Gemma 4 31B that fits in 24 GB of VRAM?
- At NVFP4, Gemma 4 31B needs 18.9 GB, the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 4 31B locally?
- A 24 GB GPU is the minimum. At Q4_K_M, Gemma 4 31B needs 22.6 GB VRAM. Good options: RTX 4090 (24 GB), RTX 3090 (24 GB).