Gemma 4 31B

Gemma 4 31B needs roughly 22.6 GB VRAM at Q4_K_M quantization (70.5 GB at FP16). 84 GPUs we track can run it fully in VRAM at 8k context.

84 GPUs run this natively · 21 with CPU offload

Google30.7B params256k contextApache 2.0Commercial use ok

Gemma 4 31B is a 30.7B parameter dense model developed by Google. April 2026 dense model optimized for workstations and servers. 256K context with multimodal support.

To run Gemma 4 31B locally: Q4_K_M ~18-20GB and fits on 24GB GPUs. Good choice for M4 Max or RTX 4090 owners.

31B dense architecture with vision capabilities, Google's workstation-focused offering.

VRAM at each quantization

Numbers here are computed at 8k context. This model's hybrid attention stack means KV cache grows much more slowly than context length, unlike a conventional full-attention model.

QuantWeightsKV cacheTotal
FP32122.8 GB1.51 GB139.2 GB
BF1661.4 GB1.51 GB70.5 GB
FP1661.4 GB1.51 GB70.5 GB
Q8_032.6 GB1.51 GB38.2 GB
Q6_K25.2 GB1.51 GB29.9 GB
Q5_K_M21.9 GB1.51 GB26.2 GB
Q4_K_Mrec18.7 GB1.51 GB22.6 GB
Q3_K_M14.8 GB1.51 GB18.2 GB
Q2_K11.7 GB1.51 GB14.8 GB
NVFP4cuda15.3 GB1.51 GB18.9 GB

KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.

Benchmarks

GPUs that run Gemma 4 31B natively (84)

Show 79 more
Plus 21 GPUs that run it with CPU offload (slower)

Notes

Dense, not MoE: all 30.7B parameters run on every token. Like Gemma 3 and the sibling Gemma 4 26B MoE, its 60 attention layers repeat a 5-local/1-global sliding-window pattern (1,024-token window), so only 10 of 60 layers scale their KV cache with the full context; the other 50 stay capped regardless of how long the conversation runs. The 10 full-attention layers also use a narrower KV width (4 KV heads x 512 head_dim) than the 50 sliding-window layers (16 x 256). Native context is 256k tokens (262,144).

Hugging Face ↗Ollama ↗Released 2026-04-02

Compare Gemma 4 31B with other models

Frequently asked questions

What are the VRAM requirements for Gemma 4 31B?
Gemma 4 31B requires approximately 22.6 GB of VRAM at Q4_K_M quantization, 38.2 GB at Q8, and 70.5 GB at FP16. These numbers assume 8k context window; its hybrid attention stack caches far fewer than all layers, so VRAM grows much slower than linearly with context.
How many parameters does Gemma 4 31B have?
Gemma 4 31B has 30.7 billion parameters.
Is Gemma 4 31B good for coding?
Yes. Gemma 4 31B scores 80.0 on LiveCodeBench, demonstrating strong code generation and completion capabilities.
Can Gemma 4 31B run on a 16 GB GPU?
No. At Q4_K_M, Gemma 4 31B needs 22.6 GB of VRAM, more than 16 GB. You will need a 24 GB GPU like the RTX 4090 or RTX 3090.
Can Gemma 4 31B run on a 24 GB GPU?
Yes. Gemma 4 31B fits in a 24 GB GPU at Q4_K_M, requiring 22.6 GB VRAM. GPUs with 24 GB include the RTX 4090, RTX 3090, and RTX 3090 Ti.
What is the smallest quantization for Gemma 4 31B that fits in 24 GB of VRAM?
At NVFP4, Gemma 4 31B needs 18.9 GB, the highest-quality quantization that fits in 24 GB of VRAM.
What GPU do I need to run Gemma 4 31B locally?
A 24 GB GPU is the minimum. At Q4_K_M, Gemma 4 31B needs 22.6 GB VRAM. Good options: RTX 4090 (24 GB), RTX 3090 (24 GB).