Gemma 3 27B Instruct
Gemma 3 27B Instruct needs roughly 20.1 GB VRAM at Q4_K_M quantization (62.2 GB at FP16). 75 GPUs we track can run it fully in VRAM at 8k context.
75 GPUs run this natively · 19 with CPU offload
Gemma 3 27B Instruct is a 27B parameter dense model developed by Google. March 2025 multimodal model with native vision via SigLIP 400M encoder. 128K context with 5:1 local/global attention interleaving.
To run Gemma 3 27B Instruct locally: Q4_K_M needs ~32-33GB with context — 24GB GPU can run it but 32GB+ recommended for vision tasks. KV-cache optimization gives 60% memory reduction.
MMLU-Pro 67.5%, LiveCodeBench 29.7%, MMMU (vision) 64.9% — Chatbot Arena Elo 1338 ranks it best open non-thinking model.
VRAM at each quantization
Numbers here are computed at 8k context. Because KV cache grows linearly with context length, expect higher totals at longer sequence lengths.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 108.0 GB | 1.54 GB | 122.7 GB |
| BF16 | 54.0 GB | 1.54 GB | 62.2 GB |
| FP16 | 54.0 GB | 1.54 GB | 62.2 GB |
| Q8_0 | 28.7 GB | 1.54 GB | 33.9 GB |
| Q6_K | 22.2 GB | 1.54 GB | 26.6 GB |
| Q5_K_M | 19.2 GB | 1.54 GB | 23.3 GB |
| Q4_K_Mrec | 16.4 GB | 1.54 GB | 20.1 GB |
| Q3_K_M | 13.0 GB | 1.54 GB | 16.3 GB |
| Q2_K | 10.3 GB | 1.54 GB | 13.3 GB |
| NVFP4cuda | 13.5 GB | 1.54 GB | 16.9 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run Gemma 3 27B Instruct natively (75)
- NVIDIA RTX 5090NVFP4 · 77.4 t/s
- NVIDIA RTX 5080Q2_K · 52.7 t/s
- NVIDIA RTX 5070 TiQ2_K · 49.2 t/s
- NVIDIA RTX 5060 Ti 16GBQ2_K · 24.6 t/s
- NVIDIA RTX 4090NVFP4 · 43.6 t/s
- NVIDIA RTX 4080Q2_K · 39.4 t/s
- NVIDIA RTX 4060 Ti 16GBQ2_K · 15.8 t/s
- NVIDIA RTX 3090NVFP4 · 40.4 t/s
- NVIDIA RTX 3090 TiNVFP4 · 43.6 t/s
- NVIDIA H100 80GBBF16 · 39.2 t/s
- NVIDIA A100 80GBBF16 · 23.9 t/s
- NVIDIA A100 40GBNVFP4 · 67.2 t/s
- NVIDIA L40SNVFP4 · 37.3 t/s
- NVIDIA RTX A6000NVFP4 · 33.2 t/s
- NVIDIA RTX 4000 AdaNVFP4 · 13.8 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 18.7 t/s
- NVIDIA RTX 5000 AdaNVFP4 · 24.9 t/s
- NVIDIA RTX 6000 AdaNVFP4 · 41.5 t/s
- NVIDIA RTX Pro 6000BF16 · 15.7 t/s
- NVIDIA DGX Spark (128GB)BF16 · 3.2 t/s
- AMD Radeon RX 7900 XTXQ4_K_M · 34.7 t/s
- AMD Radeon RX 7900 XTQ3_K_M · 35.8 t/s
- AMD Radeon RX 7900 GREQ2_K · 31.6 t/s
- AMD Radeon RX 6800 XTQ2_K · 28.1 t/s
- AMD Radeon PRO W7800Q6_K · 15.8 t/s
- AMD Radeon PRO W7900Q8_0 · 18.6 t/s
- AMD Instinct MI300XFP32 · 31.4 t/s
- AMD Radeon AI Pro 9700 32GBQ6_K · 17.5 t/s
- AMD Strix Halo (128GB)BF16 · 3 t/s
- AMD Strix Halo (96GB)BF16 · 3 t/s
- AMD Strix Halo (64GB)Q8_0 · 5.5 t/s
- Apple M5 Max (128GB)BF16 · 8.8 t/s
- Apple M5 Max (64GB)Q8_0 · 16.2 t/s
- Apple M5 Max (48GB)Q8_0 · 16.2 t/s
- Apple M5 Pro (48GB)Q8_0 · 8.1 t/s
- Apple M5 Pro (36GB)Q6_K · 10.4 t/s
- Apple M5 Pro (24GB)Q2_K · 20.8 t/s
- Apple M5 (32GB)Q5_K_M · 5.9 t/s
- Apple M4 Ultra (384GB)FP32 · 8 t/s
- Apple M4 Ultra (192GB)FP32 · 8 t/s
- Apple M4 Max (128GB)BF16 · 7.9 t/s
- Apple M4 Max (96GB)BF16 · 7.9 t/s
- Apple M4 Max (64GB)Q8_0 · 14.4 t/s
- Apple M4 Max (48GB)Q8_0 · 14.4 t/s
- Apple M4 Pro (48GB)Q8_0 · 7.2 t/s
- Apple M4 Pro (24GB)Q2_K · 18.5 t/s
- Apple M4 (32GB)Q5_K_M · 4.6 t/s
- Apple M3 Ultra (512GB)FP32 · 6 t/s
- Apple M3 Ultra (256GB)FP32 · 6 t/s
- Apple M3 Ultra (96GB)BF16 · 11.8 t/s
- Apple M3 Max (128GB)BF16 · 5.8 t/s
- Apple M3 Max (96GB)BF16 · 5.8 t/s
- Apple M3 Max (64GB)Q8_0 · 10.6 t/s
- Apple M3 Max (48GB)Q8_0 · 10.6 t/s
- Apple M3 Max (36GB)Q6_K · 13.5 t/s
- Apple M3 Pro (36GB)Q6_K · 5.1 t/s
- Apple M3 (24GB)Q2_K · 6.8 t/s
- Apple M2 Ultra (384GB)FP32 · 5.8 t/s
- Apple M2 Ultra (192GB)FP32 · 5.8 t/s
- Apple M2 Max (96GB)BF16 · 5.8 t/s
- Apple M2 Max (64GB)Q8_0 · 10.6 t/s
- Apple M2 Max (32GB)Q5_K_M · 15.4 t/s
- Apple M2 Pro (32GB)Q5_K_M · 7.7 t/s
- Apple M2 (24GB)Q2_K · 6.8 t/s
- Apple M1 Ultra (128GB)BF16 · 11.5 t/s
- Apple M1 Ultra (64GB)Q8_0 · 21.2 t/s
- Apple M1 Max (64GB)Q8_0 · 10.6 t/s
- Apple M1 Max (32GB)Q5_K_M · 15.4 t/s
- Apple M1 Pro (32GB)Q5_K_M · 7.7 t/s
- Intel Arc Pro B70 24GBQ4_K_M · 16.5 t/s
- Intel Arc Pro B60 24GBQ4_K_M · 13.7 t/s
- Intel Arc A770 16GBQ2_K · 30.8 t/s
- Intel Data Center GPU Max 1550BF16 · 38.3 t/s
- Intel Data Center GPU Max 1100Q8_0 · 26.4 t/s
- Intel Arc 140V (32GB)Q5_K_M · 4.3 t/s
Plus 19 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5070NVFP4 · 5.8 t/s
- NVIDIA RTX 5060NVFP4 · 3.1 t/s
- NVIDIA RTX 5050NVFP4 · 3 t/s
- NVIDIA RTX 4070 TiNVFP4 · 5.5 t/s
- NVIDIA RTX 4070NVFP4 · 5.5 t/s
- NVIDIA RTX 4060NVFP4 · 2.9 t/s
- NVIDIA RTX 3080 10GBNVFP4 · 4.1 t/s
- NVIDIA RTX 3060 12GBNVFP4 · 5.1 t/s
- Intel Arc B580 12GBQ8_0 · 1.3 t/s
- Intel Arc B570 10GBQ8_0 · 1.2 t/s
- Intel Arc A770 8GBQ6_K · 1.5 t/s
- Intel Arc A750 8GBQ6_K · 1.5 t/s
- Intel Arc A580 8GBQ6_K · 1.5 t/s
- Intel Arc A380 6GBQ6_K · 1.3 t/s
- Intel Arc A310 4GBQ6_K · 1.2 t/s
- Intel Arc Pro A60 12GBQ8_0 · 1.3 t/s
- Intel Arc Pro A50 6GBQ6_K · 1.3 t/s
- Intel Arc Pro A40 6GBQ6_K · 1.3 t/s
- CPU only (system RAM)Q6_K · 1.7 t/s
Notes
5:1 local/global attention interleaving.
Compare Gemma 3 27B Instruct with other models
How to run Gemma 3 27B Instruct locally
Q4_K_M needs 20.1 GB — fits a single high-end consumer GPU (24 GB).
Ollama
ollama run gemma3:27bllama.cpp
./llama-cli -m gemma-3-27b-it.Q4_K_M.gguf -c 8192 -ngl 99LM Studio: Search for 'Gemma 3 27B' in LM Studio. The Q4_K_M variant fits on a single 24 GB GPU. This model supports both text and image inputs.
Why this quantization? At 27B dense parameters, Q4_K_M keeps the model within reach of a single 24 GB GPU (roughly 17 GB for weights plus room for KV cache). Gemma 3 27B uses a 5:1 local/global attention interleaving pattern that keeps the KV cache smaller than a standard full-attention model at the same size. This architectural efficiency means Q4 preserves strong performance (MMLU-Pro: 67.5) while fitting on mainstream hardware.
Who is Gemma 3 27B Instruct for?
Users with a single 24 GB GPU who want a strong multimodal model that handles both text and images. Ideal for developers exploring vision-language tasks locally, or anyone who wants Google's latest open model running on their workstation.
Best for
- Multimodal tasks combining text and image understanding
- High-quality general chat and instruction following on a single 24 GB GPU
- Content analysis that includes visual elements (charts, screenshots, documents)
- Multilingual text generation leveraging Google's training data
- Fine-tuning for domain-specific vision-language applications
Not ideal for
- Users who only need text -- comparably sized text-only models like Qwen 2.5 32B may perform better on pure-text benchmarks
- Workloads requiring very long context -- despite 128K support, practical context is limited by VRAM
- Users who need Apache 2.0 licensing -- Gemma uses its own license
Continue reading
Frequently asked questions
- What are the VRAM requirements for Gemma 3 27B Instruct?
- Gemma 3 27B Instruct requires approximately 20.1 GB of VRAM at Q4_K_M quantization, 33.9 GB at Q8, and 62.2 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 3 27B Instruct have?
- Gemma 3 27B Instruct has 27 billion parameters.
- How capable is Gemma 3 27B Instruct?
- With an MMLU-Pro score of 67.5, Gemma 3 27B Instruct delivers solid general-purpose performance suitable for most everyday tasks and professional use.
- Can Gemma 3 27B Instruct run on a 16 GB GPU?
- No. At Q4_K_M, Gemma 3 27B Instruct needs 20.1 GB of VRAM — more than 16 GB. You will need a 24 GB GPU like the RTX 4090 or RTX 3090.
- Can Gemma 3 27B Instruct run on a 24 GB GPU?
- Yes. Gemma 3 27B Instruct fits in a 24 GB GPU at Q4_K_M, requiring 20.1 GB VRAM. GPUs with 24 GB include the RTX 4090, RTX 3090, and RTX 3090 Ti.
- What is the smallest quantization for Gemma 3 27B Instruct that fits in 24 GB of VRAM?
- At NVFP4, Gemma 3 27B Instruct needs 16.8 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 3 27B Instruct locally?
- A 24 GB GPU is the minimum. At Q4_K_M, Gemma 3 27B Instruct needs 20.1 GB VRAM. Good options: RTX 4090 (24 GB), RTX 3090 (24 GB).