CanItRun Logocanitrun.

Gemma 4 12B (Unified)

Gemma 4 12B (Unified) needs roughly 11.2 GB VRAM at Q4_K_M quantization (30.5 GB at FP16). 84 GPUs we track can run it fully in VRAM at 8k context.

84 GPUs run this natively · 11 with CPU offload

Google12B params256k contextApache 2.0Commercial use ok

Gemma 4 12B (Unified) is a 12B parameter dense model developed by Google. June 2026 release using an encoder-free 'Unified' architecture — image, audio, and video are projected directly into the transformer instead of routing through separate encoders. 256K native context, Apache 2.0.

To run Gemma 4 12B (Unified) locally: Q4_K_M needs ~6.8GB for weights (~11GB total at 8K context) — comfortable on any 16GB GPU with headroom to spare, unlike the 26B/31B siblings which need aggressive quantization to fit the same tier.

MMLU-Pro 77.2 at 12B lands within reach of the 26B and 31B siblings — the best quality-per-GB in the Gemma 4 lineup.

VRAM at each quantization

Gemma 4 12B (Unified) natively supports a longer context window, but the table below is capped at 8k for comparability — KV cache grows linearly with context length.

QuantWeightsKV cacheTotal
FP3248.0 GB3.22 GB57.4 GB
BF1624.0 GB3.22 GB30.5 GB
FP1624.0 GB3.22 GB30.5 GB
Q8_012.0 GB3.22 GB17.1 GB
Q6_K9.8 GB3.22 GB14.6 GB
Q5_K_M7.7 GB3.22 GB12.3 GB
Q4_K_Mrec6.8 GB3.22 GB11.2 GB
Q3_K_M5.2 GB3.22 GB9.4 GB
Q2_K4.0 GB3.22 GB8.0 GB
NVFP4cuda6.0 GB3.22 GB10.3 GB

KV cache is calculated at 8k context (FP16). Note that NVFP4 only runs on CUDA GPUs. Turn on TurboQuant in the calculator above for lower KV cache estimates.

Benchmarks

GPUs that run Gemma 4 12B (Unified) natively (84)

Plus 11 GPUs that run it with CPU offload (slower)

Notes

Encoder-free 'unified' architecture projects image/audio/video directly into the transformer instead of using separate encoders. Close behind the 26B and 31B siblings on most benchmarks at under half the size.

Hugging Face ↗Ollama ↗Released 2026-06-03

Frequently asked questions

What are the VRAM requirements for Gemma 4 12B (Unified)?
Gemma 4 12B (Unified) requires approximately 11.2 GB of VRAM at Q4_K_M quantization, 17.0 GB at Q8, and 30.5 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
How many parameters does Gemma 4 12B (Unified) have?
Gemma 4 12B (Unified) has 12 billion parameters.
How capable is Gemma 4 12B (Unified)?
Gemma 4 12B (Unified) achieves an MMLU-Pro score of 77.2, placing it among the most capable open-weight models available — competitive with frontier systems on general knowledge and reasoning.
Can Gemma 4 12B (Unified) run on a 16 GB GPU?
Yes. Gemma 4 12B (Unified) needs 11.2 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 4070 Ti Super.
What is the smallest quantization for Gemma 4 12B (Unified) that fits in 24 GB of VRAM?
At NVFP4, Gemma 4 12B (Unified) needs 10.3 GB — the highest-quality quantization that fits in 24 GB of VRAM.
What GPU do I need to run Gemma 4 12B (Unified) locally?
A 16 GB GPU is enough. At Q4_K_M, Gemma 4 12B (Unified) needs 11.2 GB VRAM. Good options: RTX 4080 (16 GB), RTX 4070 Ti Super (16 GB).