Gemma 4 12B (Unified)
Gemma 4 12B (Unified) needs roughly 11.2 GB VRAM at Q4_K_M quantization (30.5 GB at FP16). 84 GPUs we track can run it fully in VRAM at 8k context.
84 GPUs run this natively · 11 with CPU offload
Gemma 4 12B (Unified) is a 12B parameter dense model developed by Google. June 2026 release using an encoder-free 'Unified' architecture — image, audio, and video are projected directly into the transformer instead of routing through separate encoders. 256K native context, Apache 2.0.
To run Gemma 4 12B (Unified) locally: Q4_K_M needs ~6.8GB for weights (~11GB total at 8K context) — comfortable on any 16GB GPU with headroom to spare, unlike the 26B/31B siblings which need aggressive quantization to fit the same tier.
MMLU-Pro 77.2 at 12B lands within reach of the 26B and 31B siblings — the best quality-per-GB in the Gemma 4 lineup.
VRAM at each quantization
Gemma 4 12B (Unified) natively supports a longer context window, but the table below is capped at 8k for comparability — KV cache grows linearly with context length.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 48.0 GB | 3.22 GB | 57.4 GB |
| BF16 | 24.0 GB | 3.22 GB | 30.5 GB |
| FP16 | 24.0 GB | 3.22 GB | 30.5 GB |
| Q8_0 | 12.0 GB | 3.22 GB | 17.1 GB |
| Q6_K | 9.8 GB | 3.22 GB | 14.6 GB |
| Q5_K_M | 7.7 GB | 3.22 GB | 12.3 GB |
| Q4_K_Mrec | 6.8 GB | 3.22 GB | 11.2 GB |
| Q3_K_M | 5.2 GB | 3.22 GB | 9.4 GB |
| Q2_K | 4.0 GB | 3.22 GB | 8.0 GB |
| NVFP4cuda | 6.0 GB | 3.22 GB | 10.3 GB |
KV cache is calculated at 8k context (FP16). Note that NVFP4 only runs on CUDA GPUs. Turn on TurboQuant in the calculator above for lower KV cache estimates.
Benchmarks
GPUs that run Gemma 4 12B (Unified) natively (84)
- NVIDIA RTX 5090NVFP4 · 298.7 t/s
- NVIDIA RTX 5080NVFP4 · 160 t/s
- NVIDIA RTX 5070 TiNVFP4 · 149.3 t/s
- NVIDIA RTX 5070NVFP4 · 112 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 74.7 t/s
- NVIDIA RTX 4090NVFP4 · 168 t/s
- NVIDIA RTX 4080NVFP4 · 119.5 t/s
- NVIDIA RTX 4070 TiNVFP4 · 84 t/s
- NVIDIA RTX 4070NVFP4 · 84 t/s
- NVIDIA RTX 4060 Ti 16GBNVFP4 · 48 t/s
- NVIDIA RTX 3090NVFP4 · 156 t/s
- NVIDIA RTX 3090 TiNVFP4 · 168 t/s
- NVIDIA RTX 3080 10GBQ3_K_M · 147.3 t/s
- NVIDIA RTX 3060 12GBNVFP4 · 60 t/s
- NVIDIA H100 80GBFP32 · 69.8 t/s
- NVIDIA A100 80GBFP32 · 42.5 t/s
- NVIDIA A100 40GBBF16 · 64.8 t/s
- NVIDIA L40SBF16 · 36 t/s
- NVIDIA RTX A6000BF16 · 32 t/s
- NVIDIA RTX 4000 AdaNVFP4 · 53.3 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 72 t/s
- NVIDIA RTX 5000 AdaNVFP4 · 96 t/s
- NVIDIA RTX 6000 AdaBF16 · 40 t/s
- NVIDIA RTX Pro 6000FP32 · 28 t/s
- NVIDIA DGX Spark (128GB)FP32 · 5.7 t/s
- AMD Radeon RX 7900 XTXQ8_0 · 80 t/s
- AMD Radeon RX 7900 XTQ8_0 · 66.7 t/s
- AMD Radeon RX 7900 GREQ6_K · 58.5 t/s
- AMD Radeon RX 6800 XTQ6_K · 52 t/s
- AMD Radeon PRO W7800Q8_0 · 48 t/s
- AMD Radeon PRO W7900BF16 · 36 t/s
- AMD Instinct MI300XFP32 · 110.4 t/s
- AMD Radeon AI Pro 9700 32GBQ8_0 · 53.3 t/s
- AMD Strix Halo (128GB)FP32 · 5.3 t/s
- AMD Strix Halo (96GB)FP32 · 5.3 t/s
- AMD Strix Halo (64GB)BF16 · 10.7 t/s
- Apple M5 Max (128GB)FP32 · 12.8 t/s
- Apple M5 Max (64GB)BF16 · 25.6 t/s
- Apple M5 Max (48GB)BF16 · 25.6 t/s
- Apple M5 Pro (48GB)BF16 · 12.8 t/s
- Apple M5 Pro (36GB)Q8_0 · 25.6 t/s
- Apple M5 Pro (24GB)Q6_K · 31.2 t/s
- Apple M5 (32GB)Q8_0 · 12.8 t/s
- Apple M4 Ultra (384GB)FP32 · 22.8 t/s
- Apple M4 Ultra (192GB)FP32 · 22.8 t/s
- Apple M4 Max (128GB)FP32 · 11.4 t/s
- Apple M4 Max (96GB)FP32 · 11.4 t/s
- Apple M4 Max (64GB)BF16 · 22.8 t/s
- Apple M4 Max (48GB)BF16 · 22.8 t/s
- Apple M4 Pro (48GB)BF16 · 11.4 t/s
- Apple M4 Pro (24GB)Q6_K · 27.7 t/s
- Apple M4 (32GB)Q8_0 · 10 t/s
- Apple M3 Ultra (512GB)FP32 · 17.1 t/s
- Apple M3 Ultra (256GB)FP32 · 17.1 t/s
- Apple M3 Ultra (96GB)FP32 · 17.1 t/s
- Apple M3 Max (128GB)FP32 · 8.3 t/s
- Apple M3 Max (96GB)FP32 · 8.3 t/s
- Apple M3 Max (64GB)BF16 · 16.7 t/s
- Apple M3 Max (48GB)BF16 · 16.7 t/s
- Apple M3 Max (36GB)Q8_0 · 33.3 t/s
- Apple M3 Pro (36GB)Q8_0 · 12.5 t/s
- Apple M3 Pro (18GB)Q3_K_M · 29.1 t/s
- Apple M3 (24GB)Q6_K · 10.2 t/s
- Apple M2 Ultra (384GB)FP32 · 16.7 t/s
- Apple M2 Ultra (192GB)FP32 · 16.7 t/s
- Apple M2 Max (96GB)FP32 · 8.3 t/s
- Apple M2 Max (64GB)BF16 · 16.7 t/s
- Apple M2 Max (32GB)Q8_0 · 33.3 t/s
- Apple M2 Pro (32GB)Q8_0 · 16.7 t/s
- Apple M2 (24GB)Q6_K · 10.2 t/s
- Apple M1 Ultra (128GB)FP32 · 16.7 t/s
- Apple M1 Ultra (64GB)BF16 · 33.3 t/s
- Apple M1 Max (64GB)BF16 · 16.7 t/s
- Apple M1 Max (32GB)Q8_0 · 33.3 t/s
- Apple M1 Pro (32GB)Q8_0 · 16.7 t/s
- Intel Arc B580 12GBQ4_K_M · 67.5 t/s
- Intel Arc B570 10GBQ3_K_M · 73.6 t/s
- Intel Arc Pro B70 24GBQ8_0 · 38 t/s
- Intel Arc Pro B60 24GBQ8_0 · 31.7 t/s
- Intel Arc A770 16GBQ6_K · 56.9 t/s
- Intel Arc Pro A60 12GBQ4_K_M · 56.8 t/s
- Intel Data Center GPU Max 1550FP32 · 68.3 t/s
- Intel Data Center GPU Max 1100BF16 · 51.2 t/s
- Intel Arc 140V (32GB)Q8_0 · 11.4 t/s
Plus 11 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5060BF16 · 4.7 t/s
- NVIDIA RTX 5050BF16 · 3.3 t/s
- NVIDIA RTX 4060BF16 · 2.8 t/s
- Intel Arc A770 8GBBF16 · 5.3 t/s
- Intel Arc A750 8GBBF16 · 5.3 t/s
- Intel Arc A580 8GBBF16 · 5.3 t/s
- Intel Arc A380 6GBBF16 · 1.9 t/s
- Intel Arc A310 4GBQ8_0 · 2.6 t/s
- Intel Arc Pro A50 6GBBF16 · 2 t/s
- Intel Arc Pro A40 6GBBF16 · 2 t/s
- CPU only (system RAM)Q8_0 · 0.9 t/s
Notes
Encoder-free 'unified' architecture projects image/audio/video directly into the transformer instead of using separate encoders. Close behind the 26B and 31B siblings on most benchmarks at under half the size.
Continue reading
Frequently asked questions
- What are the VRAM requirements for Gemma 4 12B (Unified)?
- Gemma 4 12B (Unified) requires approximately 11.2 GB of VRAM at Q4_K_M quantization, 17.0 GB at Q8, and 30.5 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 4 12B (Unified) have?
- Gemma 4 12B (Unified) has 12 billion parameters.
- How capable is Gemma 4 12B (Unified)?
- Gemma 4 12B (Unified) achieves an MMLU-Pro score of 77.2, placing it among the most capable open-weight models available — competitive with frontier systems on general knowledge and reasoning.
- Can Gemma 4 12B (Unified) run on a 16 GB GPU?
- Yes. Gemma 4 12B (Unified) needs 11.2 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 4070 Ti Super.
- What is the smallest quantization for Gemma 4 12B (Unified) that fits in 24 GB of VRAM?
- At NVFP4, Gemma 4 12B (Unified) needs 10.3 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 4 12B (Unified) locally?
- A 16 GB GPU is enough. At Q4_K_M, Gemma 4 12B (Unified) needs 11.2 GB VRAM. Good options: RTX 4080 (16 GB), RTX 4070 Ti Super (16 GB).