CanItRun Logocanitrun.

Gemma 4 E2B

Gemma 4 E2B needs roughly 1.8 GB VRAM at Q4_K_M quantization (4.9 GB at FP16). 103 GPUs we track can run it fully in VRAM at 8k context.

103 GPUs run this natively · 1 with CPU offload

Google2B params125k contextApache 2.0Commercial use ok

Gemma 4 E2B is a 2B parameter dense model developed by Google. April 2026, the smaller sibling in Google's elastic Gemma 4 pair, using the same MatFormer/per-layer-embedding approach as Gemma 3n's E2B to shrink runtime memory below what the raw 2B parameter count would suggest. Multimodal (text, vision, audio), Apache 2.0.

To run Gemma 4 E2B locally: Q8_0 needs roughly 2GB — runs on entry-level 8GB GPUs, higher-end phones, and other edge hardware with headroom to spare.

MMLU-Pro 60.0 is high for a 2B-class model, reflecting the same efficiency techniques that made Gemma 3n's E2B punch above its size.

VRAM at each quantization

Gemma 4 E2B natively supports a longer context window, but the table below is capped at 8k for comparability — KV cache grows linearly with context length.

QuantWeightsKV cacheTotal
FP328.0 GB0.40 GB9.4 GB
BF164.0 GB0.40 GB4.9 GB
FP164.0 GB0.40 GB4.9 GB
Q8_0rec2.1 GB0.40 GB2.8 GB
Q6_K1.6 GB0.40 GB2.3 GB
Q5_K_M1.4 GB0.40 GB2.0 GB
Q4_K_M1.2 GB0.40 GB1.8 GB
Q3_K_M1.0 GB0.40 GB1.5 GB
Q2_K0.8 GB0.40 GB1.3 GB
NVFP4cuda1.0 GB0.40 GB1.6 GB

KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.

Benchmarks

GPUs that run Gemma 4 E2B natively (103)

Plus 1 GPUs that run it with CPU offload (slower)
Hugging Face ↗Ollama ↗Released 2026-04-02

Frequently asked questions

What are the VRAM requirements for Gemma 4 E2B?
Gemma 4 E2B requires approximately 1.8 GB of VRAM at Q4_K_M quantization, 2.8 GB at Q8, and 4.9 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
How many parameters does Gemma 4 E2B have?
Gemma 4 E2B has 2 billion parameters.
How capable is Gemma 4 E2B?
With an MMLU-Pro score of 60, Gemma 4 E2B delivers solid general-purpose performance suitable for most everyday tasks and professional use.
Can Gemma 4 E2B run on a 16 GB GPU?
Yes. Gemma 4 E2B needs 1.8 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 4070 Ti Super.
What is the smallest quantization for Gemma 4 E2B that fits in 24 GB of VRAM?
At FP32, Gemma 4 E2B needs 9.4 GB — the highest-quality quantization that fits in 24 GB of VRAM.
What GPU do I need to run Gemma 4 E2B locally?
A 16 GB GPU is enough. At Q4_K_M, Gemma 4 E2B needs 1.8 GB VRAM. Good options: RTX 4080 (16 GB), RTX 4070 Ti Super (16 GB).