CanItRun Logocanitrun.

Gemma 4 31B

Gemma 4 31B needs roughly 24.8 GB VRAM at Q4_K_M quantization (73.0 GB at FP16). 63 GPUs we track can run it fully in VRAM at 8k context.

63 GPUs run this natively · 27 with CPU offload

Google31B params250k contextApache 2.0Commercial use ok

Gemma 4 31B is a 31B parameter dense model developed by Google. April 2026 dense model optimized for workstations and servers. 256K context with multimodal support.

To run Gemma 4 31B locally: Q4_K_M ~18-20GB — fits on 24GB GPUs. Good choice for M4 Max or RTX 4090 owners.

31B dense architecture with vision capabilities — Google's workstation-focused offering.

VRAM at each quantization

Numbers here are computed at 8k context. Because KV cache grows linearly with context length, expect higher totals at longer sequence lengths.

QuantWeightsKV cacheTotal
FP32124.0 GB3.22 GB142.5 GB
BF1662.0 GB3.22 GB73.0 GB
FP1662.0 GB3.22 GB73.0 GB
Q8_033.0 GB3.22 GB40.5 GB
Q6_K25.4 GB3.22 GB32.1 GB
Q5_K_M22.1 GB3.22 GB28.3 GB
Q4_K_Mrec18.9 GB3.22 GB24.8 GB
Q3_K_M14.9 GB3.22 GB20.3 GB
Q2_K11.8 GB3.22 GB16.8 GB
NVFP4cuda15.5 GB3.22 GB21.0 GB

KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.

Benchmarks

GPUs that run Gemma 4 31B natively (63)

Plus 27 GPUs that run it with CPU offload (slower)

Notes

Dense model optimized for workstations/servers.

Hugging Face ↗Ollama ↗Released 2026-04-02

Compare Gemma 4 31B with other models

Frequently asked questions

What are the VRAM requirements for Gemma 4 31B?
Gemma 4 31B requires approximately 24.8 GB of VRAM at Q4_K_M quantization, 40.5 GB at Q8, and 73.0 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
How many parameters does Gemma 4 31B have?
Gemma 4 31B has 31 billion parameters.
How capable is Gemma 4 31B?
Gemma 4 31B achieves an MMLU-Pro score of 85.2, placing it among the most capable open-weight models available — competitive with frontier systems on general knowledge and reasoning.
Can Gemma 4 31B run on a 16 GB GPU?
No. At Q4_K_M, Gemma 4 31B needs 24.8 GB of VRAM — more than 16 GB. You will need a 48 GB GPU like the RTX 6000 Ada or a dual-GPU setup.
Can Gemma 4 31B run on a 24 GB GPU?
No. Even at Q4_K_M, Gemma 4 31B needs 24.8 GB. Consider a 48 GB card like the RTX 6000 Ada or a dual RTX 4090 setup.
What is the smallest quantization for Gemma 4 31B that fits in 24 GB of VRAM?
At NVFP4, Gemma 4 31B needs 21.0 GB — the highest-quality quantization that fits in 24 GB of VRAM.
What GPU do I need to run Gemma 4 31B locally?
You need a 48 GB GPU or a dual-GPU setup. At Q4_K_M, Gemma 4 31B needs 24.8 GB VRAM. Options: RTX 6000 Ada (48 GB), A6000 (48 GB), or 2× RTX 4090.