Gemma 3 12B Instruct
Gemma 3 12B Instruct needs roughly 9.5 GB VRAM at Q4_K_M quantization (28.5 GB at FP16). 99 GPUs we track can run it fully in VRAM at 8k context.
99 GPUs run this natively · 5 with CPU offload
Gemma 3 12B Instruct is a 12.2B parameter dense model developed by Google. Mid-size Gemma 3 with multimodal capabilities and 128K context.
To run Gemma 3 12B Instruct locally: Q5_K_M ~8-9GB — fits on 12GB GPUs comfortably.
12B sweet spot with vision support — balances quality and accessibility.
VRAM at each quantization
Numbers here are computed at 8k context. Because KV cache grows linearly with context length, expect higher totals at longer sequence lengths.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 48.8 GB | 1.06 GB | 55.8 GB |
| BF16 | 24.4 GB | 1.06 GB | 28.5 GB |
| FP16 | 24.4 GB | 1.06 GB | 28.5 GB |
| Q8_0 | 13.0 GB | 1.06 GB | 15.7 GB |
| Q6_K | 10.0 GB | 1.06 GB | 12.4 GB |
| Q5_K_Mrec | 8.7 GB | 1.06 GB | 10.9 GB |
| Q4_K_M | 7.4 GB | 1.06 GB | 9.5 GB |
| Q3_K_M | 5.9 GB | 1.06 GB | 7.8 GB |
| Q2_K | 4.7 GB | 1.06 GB | 6.4 GB |
| NVFP4cuda | 6.1 GB | 1.06 GB | 8.0 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run Gemma 3 12B Instruct natively (99)
- NVIDIA RTX 5090BF16 · 45.8 t/s
- NVIDIA RTX 5080NVFP4 · 87.2 t/s
- NVIDIA RTX 5070 TiNVFP4 · 81.4 t/s
- NVIDIA RTX 5070NVFP4 · 61 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 40.7 t/s
- NVIDIA RTX 5060Q2_K · 51 t/s
- NVIDIA RTX 5050Q2_K · 36.5 t/s
- NVIDIA RTX 4090NVFP4 · 91.5 t/s
- NVIDIA RTX 4080NVFP4 · 65.1 t/s
- NVIDIA RTX 4070 TiNVFP4 · 45.8 t/s
- NVIDIA RTX 4070NVFP4 · 45.8 t/s
- NVIDIA RTX 4060 Ti 16GBNVFP4 · 26.2 t/s
- NVIDIA RTX 4060Q2_K · 31 t/s
- NVIDIA RTX 3090NVFP4 · 85 t/s
- NVIDIA RTX 3090 TiNVFP4 · 91.5 t/s
- NVIDIA RTX 3080 10GBNVFP4 · 69 t/s
- NVIDIA RTX 3060 12GBNVFP4 · 32.7 t/s
- NVIDIA H100 80GBFP32 · 43.7 t/s
- NVIDIA A100 80GBFP32 · 26.6 t/s
- NVIDIA A100 40GBBF16 · 39.7 t/s
- NVIDIA L40SBF16 · 22.1 t/s
- NVIDIA RTX A6000BF16 · 19.6 t/s
- NVIDIA RTX 4000 AdaNVFP4 · 29.1 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 39.2 t/s
- NVIDIA RTX 5000 AdaBF16 · 14.7 t/s
- NVIDIA RTX 6000 AdaBF16 · 24.5 t/s
- NVIDIA RTX Pro 6000FP32 · 17.5 t/s
- NVIDIA DGX Spark (128GB)FP32 · 3.6 t/s
- AMD Radeon RX 7900 XTXQ8_0 · 44.5 t/s
- AMD Radeon RX 7900 XTQ8_0 · 37.1 t/s
- AMD Radeon RX 7900 GREQ6_K · 33.8 t/s
- AMD Radeon RX 6800 XTQ6_K · 30.1 t/s
- AMD Radeon PRO W7800BF16 · 14.7 t/s
- AMD Radeon PRO W7900BF16 · 22.1 t/s
- AMD Instinct MI300XFP32 · 69.1 t/s
- AMD Radeon AI Pro 9700 32GBBF16 · 16.3 t/s
- AMD Strix Halo (128GB)FP32 · 3.3 t/s
- AMD Strix Halo (96GB)FP32 · 3.3 t/s
- AMD Strix Halo (64GB)FP32 · 3.3 t/s
- Apple M5 Max (128GB)FP32 · 9.9 t/s
- Apple M5 Max (64GB)FP32 · 9.9 t/s
- Apple M5 Max (48GB)BF16 · 19.3 t/s
- Apple M5 Pro (48GB)BF16 · 9.6 t/s
- Apple M5 Pro (36GB)Q8_0 · 17.5 t/s
- Apple M5 Pro (24GB)Q8_0 · 17.5 t/s
- Apple M5 (32GB)Q8_0 · 8.7 t/s
- Apple M5 (16GB)Q3_K_M · 17.7 t/s
- Apple M4 Ultra (384GB)FP32 · 17.5 t/s
- Apple M4 Ultra (192GB)FP32 · 17.5 t/s
- Apple M4 Max (128GB)FP32 · 8.8 t/s
- Apple M4 Max (96GB)FP32 · 8.8 t/s
- Apple M4 Max (64GB)FP32 · 8.8 t/s
- Apple M4 Max (48GB)BF16 · 17.2 t/s
- Apple M4 Pro (48GB)BF16 · 8.6 t/s
- Apple M4 Pro (24GB)Q8_0 · 15.6 t/s
- Apple M4 (32GB)Q8_0 · 6.8 t/s
- Apple M4 (16GB)Q3_K_M · 13.9 t/s
- Apple M3 Ultra (512GB)FP32 · 13.1 t/s
- Apple M3 Ultra (256GB)FP32 · 13.1 t/s
- Apple M3 Ultra (96GB)FP32 · 13.1 t/s
- Apple M3 Max (128GB)FP32 · 6.4 t/s
- Apple M3 Max (96GB)FP32 · 6.4 t/s
- Apple M3 Max (64GB)FP32 · 6.4 t/s
- Apple M3 Max (48GB)BF16 · 12.6 t/s
- Apple M3 Max (36GB)Q8_0 · 22.8 t/s
- Apple M3 Pro (36GB)Q8_0 · 8.6 t/s
- Apple M3 Pro (18GB)Q4_K_M · 14.1 t/s
- Apple M3 (24GB)Q8_0 · 5.7 t/s
- Apple M3 (16GB)Q3_K_M · 11.6 t/s
- Apple M2 Ultra (384GB)FP32 · 12.8 t/s
- Apple M2 Ultra (192GB)FP32 · 12.8 t/s
- Apple M2 Max (96GB)FP32 · 6.4 t/s
- Apple M2 Max (64GB)FP32 · 6.4 t/s
- Apple M2 Max (32GB)Q8_0 · 22.8 t/s
- Apple M2 Pro (32GB)Q8_0 · 11.4 t/s
- Apple M2 Pro (16GB)Q3_K_M · 23.1 t/s
- Apple M2 (24GB)Q8_0 · 5.7 t/s
- Apple M2 (16GB)Q3_K_M · 11.6 t/s
- Apple M1 Ultra (128GB)FP32 · 12.8 t/s
- Apple M1 Ultra (64GB)FP32 · 12.8 t/s
- Apple M1 Max (64GB)FP32 · 6.4 t/s
- Apple M1 Max (32GB)Q8_0 · 22.8 t/s
- Apple M1 Pro (32GB)Q8_0 · 11.4 t/s
- Apple M1 Pro (16GB)Q3_K_M · 23.1 t/s
- Apple M1 (16GB)Q3_K_M · 7.9 t/s
- Intel Arc B580 12GBQ5_K_M · 30.4 t/s
- Intel Arc B570 10GBQ3_K_M · 35.7 t/s
- Intel Arc Pro B70 24GBQ8_0 · 21.1 t/s
- Intel Arc Pro B60 24GBQ8_0 · 17.6 t/s
- Intel Arc A770 16GBQ6_K · 32.9 t/s
- Intel Arc A770 8GBQ2_K · 58.3 t/s
- Intel Arc A750 8GBQ2_K · 58.3 t/s
- Intel Arc A580 8GBQ2_K · 58.3 t/s
- Intel Arc Pro A60 12GBQ5_K_M · 25.6 t/s
- Intel Data Center GPU Max 1550FP32 · 42.7 t/s
- Intel Data Center GPU Max 1100BF16 · 31.4 t/s
- Intel Arc 140V (32GB)Q8_0 · 6.3 t/s
- Intel Arc 140V (16GB)Q3_K_M · 12.9 t/s
- Intel Arc 130V (16GB)Q3_K_M · 12.9 t/s
Plus 5 GPUs that run it with CPU offload (slower)
- Intel Arc A380 6GBBF16 · 1.2 t/s
- Intel Arc A310 4GBBF16 · 1.1 t/s
- Intel Arc Pro A50 6GBBF16 · 1.2 t/s
- Intel Arc Pro A40 6GBBF16 · 1.2 t/s
- CPU only (system RAM)Q8_0 · 2.9 t/s
Compare Gemma 3 12B Instruct with other models
How to run Gemma 3 12B Instruct locally
Q5_K_M needs 10.9 GB — fits a single high-end consumer GPU (24 GB).
Ollama
ollama run gemma3:12bllama.cpp
./llama-cli -m gemma-3-12b-it.Q5_K_M.gguf -c 8192 -ngl 99LM Studio: Search for 'Gemma 3 12B' in LM Studio. The Q5_K_M variant runs well on 12-16 GB GPUs. Supports both text and image inputs.
Why this quantization? At 12.2B parameters, Q5_K_M uses roughly 8 GB of VRAM for weights, fitting comfortably on a 12-16 GB GPU with room for the KV cache. The 5:1 local/global attention pattern keeps cache overhead manageable. Q5 preserves the model's solid MMLU-Pro score (60.6) better than Q4 would, and the additional VRAM cost over Q4 is only about 1 GB.
Who is Gemma 3 12B Instruct for?
Users with mid-range GPUs (12-16 GB) who want multimodal capabilities without stepping up to a 24 GB card. A good middle ground for developers who need both text and vision understanding on consumer hardware.
Best for
- Image-to-text tasks like describing photos, reading charts, or parsing screenshots
- General chat and writing assistance on mid-range hardware
- Multilingual text generation and translation
- Building multimodal applications with a manageable model footprint
Not ideal for
- Heavy reasoning or math tasks -- Phi-4 14B significantly outperforms at a similar size
- Code-specialized tasks where Qwen 2.5 Coder is purpose-built
- Production deployments requiring frontier-level accuracy
Continue reading
Frequently asked questions
- What are the VRAM requirements for Gemma 3 12B Instruct?
- Gemma 3 12B Instruct requires approximately 9.5 GB of VRAM at Q4_K_M quantization, 15.7 GB at Q8, and 28.5 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 3 12B Instruct have?
- Gemma 3 12B Instruct has 12.2 billion parameters.
- How capable is Gemma 3 12B Instruct?
- With an MMLU-Pro score of 60.6, Gemma 3 12B Instruct delivers solid general-purpose performance suitable for most everyday tasks and professional use.
- Can Gemma 3 12B Instruct run on a 16 GB GPU?
- Yes. Gemma 3 12B Instruct needs 9.5 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 4070 Ti Super.
- What is the smallest quantization for Gemma 3 12B Instruct that fits in 24 GB of VRAM?
- At NVFP4, Gemma 3 12B Instruct needs 8.0 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 3 12B Instruct locally?
- A 16 GB GPU is enough. At Q4_K_M, Gemma 3 12B Instruct needs 9.5 GB VRAM. Good options: RTX 4080 (16 GB), RTX 4070 Ti Super (16 GB).