Gemma 4 E2B
Gemma 4 E2B needs roughly 1.8 GB VRAM at Q4_K_M quantization (4.9 GB at FP16). 103 GPUs we track can run it fully in VRAM at 8k context.
103 GPUs run this natively · 1 with CPU offload
Gemma 4 E2B is a 2B parameter dense model developed by Google. April 2026, the smaller sibling in Google's elastic Gemma 4 pair, using the same MatFormer/per-layer-embedding approach as Gemma 3n's E2B to shrink runtime memory below what the raw 2B parameter count would suggest. Multimodal (text, vision, audio), Apache 2.0.
To run Gemma 4 E2B locally: Q8_0 needs roughly 2GB — runs on entry-level 8GB GPUs, higher-end phones, and other edge hardware with headroom to spare.
MMLU-Pro 60.0 is high for a 2B-class model, reflecting the same efficiency techniques that made Gemma 3n's E2B punch above its size.
VRAM at each quantization
Gemma 4 E2B natively supports a longer context window, but the table below is capped at 8k for comparability — KV cache grows linearly with context length.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 8.0 GB | 0.40 GB | 9.4 GB |
| BF16 | 4.0 GB | 0.40 GB | 4.9 GB |
| FP16 | 4.0 GB | 0.40 GB | 4.9 GB |
| Q8_0rec | 2.1 GB | 0.40 GB | 2.8 GB |
| Q6_K | 1.6 GB | 0.40 GB | 2.3 GB |
| Q5_K_M | 1.4 GB | 0.40 GB | 2.0 GB |
| Q4_K_M | 1.2 GB | 0.40 GB | 1.8 GB |
| Q3_K_M | 1.0 GB | 0.40 GB | 1.5 GB |
| Q2_K | 0.8 GB | 0.40 GB | 1.3 GB |
| NVFP4cuda | 1.0 GB | 0.40 GB | 1.6 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Gemma 4 E2B natively (103)
- NVIDIA RTX 5090FP32 · 138.6 t/s
- NVIDIA RTX 5080FP32 · 74.3 t/s
- NVIDIA RTX 5070 TiFP32 · 69.3 t/s
- NVIDIA RTX 5070FP32 · 52 t/s
- NVIDIA RTX 5060 Ti 16GBFP32 · 34.7 t/s
- NVIDIA RTX 5060BF16 · 66.1 t/s
- NVIDIA RTX 5050BF16 · 47.2 t/s
- NVIDIA RTX 4090FP32 · 78 t/s
- NVIDIA RTX 4080FP32 · 55.5 t/s
- NVIDIA RTX 4070 TiFP32 · 39 t/s
- NVIDIA RTX 4070FP32 · 39 t/s
- NVIDIA RTX 4060 Ti 16GBFP32 · 22.3 t/s
- NVIDIA RTX 4060BF16 · 40.2 t/s
- NVIDIA RTX 3090FP32 · 72.4 t/s
- NVIDIA RTX 3090 TiFP32 · 78 t/s
- NVIDIA RTX 3080 10GBFP32 · 58.8 t/s
- NVIDIA RTX 3060 12GBFP32 · 27.8 t/s
- NVIDIA H100 80GBFP32 · 259.1 t/s
- NVIDIA A100 80GBFP32 · 157.7 t/s
- NVIDIA A100 40GBFP32 · 120.3 t/s
- NVIDIA L40SFP32 · 66.8 t/s
- NVIDIA RTX A6000FP32 · 59.4 t/s
- NVIDIA RTX 4000 AdaFP32 · 24.8 t/s
- NVIDIA RTX 4500 AdaFP32 · 33.4 t/s
- NVIDIA RTX 5000 AdaFP32 · 44.6 t/s
- NVIDIA RTX 6000 AdaFP32 · 74.3 t/s
- NVIDIA RTX Pro 6000FP32 · 104 t/s
- NVIDIA DGX Spark (128GB)FP32 · 21.1 t/s
- AMD Radeon RX 7900 XTXFP32 · 74.3 t/s
- AMD Radeon RX 7900 XTFP32 · 61.9 t/s
- AMD Radeon RX 7900 GREFP32 · 44.6 t/s
- AMD Radeon RX 6800 XTFP32 · 39.6 t/s
- AMD Radeon PRO W7800FP32 · 44.6 t/s
- AMD Radeon PRO W7900FP32 · 66.8 t/s
- AMD Instinct MI300XFP32 · 410 t/s
- AMD Radeon AI Pro 9700 32GBFP32 · 49.5 t/s
- AMD Strix Halo (128GB)FP32 · 19.8 t/s
- AMD Strix Halo (96GB)FP32 · 19.8 t/s
- AMD Strix Halo (64GB)FP32 · 19.8 t/s
- Apple M5 Max (128GB)FP32 · 58.5 t/s
- Apple M5 Max (64GB)FP32 · 58.5 t/s
- Apple M5 Max (48GB)FP32 · 58.5 t/s
- Apple M5 Pro (48GB)FP32 · 29.2 t/s
- Apple M5 Pro (36GB)FP32 · 29.2 t/s
- Apple M5 Pro (24GB)FP32 · 29.2 t/s
- Apple M5 (32GB)FP32 · 14.6 t/s
- Apple M5 (16GB)BF16 · 27.8 t/s
- Apple M4 Ultra (384GB)FP32 · 104 t/s
- Apple M4 Ultra (192GB)FP32 · 104 t/s
- Apple M4 Max (128GB)FP32 · 52 t/s
- Apple M4 Max (96GB)FP32 · 52 t/s
- Apple M4 Max (64GB)FP32 · 52 t/s
- Apple M4 Max (48GB)FP32 · 52 t/s
- Apple M4 Pro (48GB)FP32 · 26 t/s
- Apple M4 Pro (24GB)FP32 · 26 t/s
- Apple M4 (32GB)FP32 · 11.4 t/s
- Apple M4 (16GB)BF16 · 21.8 t/s
- Apple M3 Ultra (512GB)FP32 · 78 t/s
- Apple M3 Ultra (256GB)FP32 · 78 t/s
- Apple M3 Ultra (96GB)FP32 · 78 t/s
- Apple M3 Max (128GB)FP32 · 38.1 t/s
- Apple M3 Max (96GB)FP32 · 38.1 t/s
- Apple M3 Max (64GB)FP32 · 38.1 t/s
- Apple M3 Max (48GB)FP32 · 38.1 t/s
- Apple M3 Max (36GB)FP32 · 38.1 t/s
- Apple M3 Pro (36GB)FP32 · 14.3 t/s
- Apple M3 Pro (18GB)FP32 · 14.3 t/s
- Apple M3 (24GB)FP32 · 9.5 t/s
- Apple M3 (16GB)BF16 · 18.2 t/s
- Apple M2 Ultra (384GB)FP32 · 76.2 t/s
- Apple M2 Ultra (192GB)FP32 · 76.2 t/s
- Apple M2 Max (96GB)FP32 · 38.1 t/s
- Apple M2 Max (64GB)FP32 · 38.1 t/s
- Apple M2 Max (32GB)FP32 · 38.1 t/s
- Apple M2 Pro (32GB)FP32 · 19 t/s
- Apple M2 Pro (16GB)BF16 · 36.3 t/s
- Apple M2 (24GB)FP32 · 9.5 t/s
- Apple M2 (16GB)BF16 · 18.2 t/s
- Apple M1 Ultra (128GB)FP32 · 76.2 t/s
- Apple M1 Ultra (64GB)FP32 · 76.2 t/s
- Apple M1 Max (64GB)FP32 · 38.1 t/s
- Apple M1 Max (32GB)FP32 · 38.1 t/s
- Apple M1 Pro (32GB)FP32 · 19 t/s
- Apple M1 Pro (16GB)BF16 · 36.3 t/s
- Apple M1 (16GB)BF16 · 12.4 t/s
- Intel Arc B580 12GBFP32 · 35.3 t/s
- Intel Arc B570 10GBFP32 · 29.4 t/s
- Intel Arc Pro B70 24GBFP32 · 35.3 t/s
- Intel Arc Pro B60 24GBFP32 · 29.4 t/s
- Intel Arc A770 16GBFP32 · 43.3 t/s
- Intel Arc A770 8GBBF16 · 75.6 t/s
- Intel Arc A750 8GBBF16 · 75.6 t/s
- Intel Arc A580 8GBBF16 · 75.6 t/s
- Intel Arc A380 6GBBF16 · 27.5 t/s
- Intel Arc A310 4GBQ8_0 · 31.9 t/s
- Intel Arc Pro A60 12GBFP32 · 29.7 t/s
- Intel Arc Pro A50 6GBBF16 · 28.3 t/s
- Intel Arc Pro A40 6GBBF16 · 28.3 t/s
- Intel Data Center GPU Max 1550FP32 · 253.4 t/s
- Intel Data Center GPU Max 1100FP32 · 95.1 t/s
- Intel Arc 140V (32GB)FP32 · 10.6 t/s
- Intel Arc 140V (16GB)BF16 · 20.2 t/s
- Intel Arc 130V (16GB)BF16 · 20.2 t/s
Plus 1 GPUs that run it with CPU offload (slower)
- CPU only (system RAM)FP32 · 4.8 t/s
Frequently asked questions
- What are the VRAM requirements for Gemma 4 E2B?
- Gemma 4 E2B requires approximately 1.8 GB of VRAM at Q4_K_M quantization, 2.8 GB at Q8, and 4.9 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 4 E2B have?
- Gemma 4 E2B has 2 billion parameters.
- How capable is Gemma 4 E2B?
- With an MMLU-Pro score of 60, Gemma 4 E2B delivers solid general-purpose performance suitable for most everyday tasks and professional use.
- Can Gemma 4 E2B run on a 16 GB GPU?
- Yes. Gemma 4 E2B needs 1.8 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 4070 Ti Super.
- What is the smallest quantization for Gemma 4 E2B that fits in 24 GB of VRAM?
- At FP32, Gemma 4 E2B needs 9.4 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 4 E2B locally?
- A 16 GB GPU is enough. At Q4_K_M, Gemma 4 E2B needs 1.8 GB VRAM. Good options: RTX 4080 (16 GB), RTX 4070 Ti Super (16 GB).