CanItRun Logocanitrun.

Gemma 2 27B Instruct

Gemma 2 27B Instruct needs roughly 22.0 GB VRAM at Q4_K_M quantization (64.4 GB at FP16). 75 GPUs we track can run it fully in VRAM at 8k context.

75 GPUs run this natively · 19 with CPU offload

Google27.2B params8k contextGemmaCommercial use ok

Gemma 2 27B Instruct is a 27.2B parameter dense model developed by Google. June 2024 release with 8K context — short context but efficient architecture.

To run Gemma 2 27B Instruct locally: Q4_K_M ~16-18GB — fits on 24GB GPU with room to spare. Good mid-range option.

MMLU-Pro 38.0%, strong for its size. The 8K context keeps KV cache tiny even at full context.

VRAM at each quantization

Figures below assume 8k context; KV cache grows linearly as context length increases.

QuantWeightsKV cacheTotal
FP32108.8 GB3.09 GB125.3 GB
BF1654.4 GB3.09 GB64.4 GB
FP1654.4 GB3.09 GB64.4 GB
Q8_028.9 GB3.09 GB35.8 GB
Q6_K22.3 GB3.09 GB28.5 GB
Q5_K_M19.4 GB3.09 GB25.1 GB
Q4_K_Mrec16.6 GB3.09 GB22.0 GB
Q3_K_M13.1 GB3.09 GB18.1 GB
Q2_K10.4 GB3.09 GB15.1 GB
NVFP4cuda13.6 GB3.09 GB18.7 GB

KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.

Benchmarks

GPUs that run Gemma 2 27B Instruct natively (75)

Plus 19 GPUs that run it with CPU offload (slower)

Notes

Short 8k context — KV cache is tiny even at full context.

Hugging Face ↗Ollama ↗Released 2024-06-27

Compare Gemma 2 27B Instruct with other models

Frequently asked questions

What are the VRAM requirements for Gemma 2 27B Instruct?
Gemma 2 27B Instruct requires approximately 22.0 GB of VRAM at Q4_K_M quantization, 35.8 GB at Q8, and 64.4 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
How many parameters does Gemma 2 27B Instruct have?
Gemma 2 27B Instruct has 27.2 billion parameters.
How capable is Gemma 2 27B Instruct?
Gemma 2 27B Instruct has an MMLU-Pro score of 38, making it well-suited for lightweight tasks, prototyping, and resource-constrained environments.
Can Gemma 2 27B Instruct run on a 16 GB GPU?
No. At Q4_K_M, Gemma 2 27B Instruct needs 22.0 GB of VRAM — more than 16 GB. You will need a 24 GB GPU like the RTX 4090 or RTX 3090.
Can Gemma 2 27B Instruct run on a 24 GB GPU?
Yes. Gemma 2 27B Instruct fits in a 24 GB GPU at Q4_K_M, requiring 22.0 GB VRAM. GPUs with 24 GB include the RTX 4090, RTX 3090, and RTX 3090 Ti.
What is the smallest quantization for Gemma 2 27B Instruct that fits in 24 GB of VRAM?
At NVFP4, Gemma 2 27B Instruct needs 18.7 GB — the highest-quality quantization that fits in 24 GB of VRAM.
What GPU do I need to run Gemma 2 27B Instruct locally?
A 24 GB GPU is the minimum. At Q4_K_M, Gemma 2 27B Instruct needs 22.0 GB VRAM. Good options: RTX 4090 (24 GB), RTX 3090 (24 GB).