Gemma 3 1B Instruct
Gemma 3 1B Instruct needs roughly 1.1 GB VRAM at Q4_K_M quantization (2.6 GB at FP16). 115 GPUs we track can run it fully in VRAM at 8k context.
115 GPUs run this natively · 1 with CPU offload
Gemma 3 1B Instruct is a 1B parameter dense model developed by Google. March 2025 release, the smallest model in the Gemma 3 family. Unlike the 4B/12B/27B siblings, the 1B checkpoint is text-only (no SigLIP vision encoder) and caps out at 32K context rather than 128K.
To run Gemma 3 1B Instruct locally: Q8_0 needs roughly 1GB and runs on any GPU, most phones, and even Raspberry Pi-class hardware without issue.
MMLU-Pro 14.7 is modest, as expected at 1B scale, but the tradeoff buys very low latency for simple classification, autocomplete, and on-device assistant tasks.
VRAM at each quantization
Calculated at 8k context. Since KV cache scales linearly with context, longer sessions need more VRAM than shown here.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 4.0 GB | 0.33 GB | 4.8 GB |
| BF16 | 2.0 GB | 0.33 GB | 2.6 GB |
| FP16 | 2.0 GB | 0.33 GB | 2.6 GB |
| Q8_0rec | 1.1 GB | 0.33 GB | 1.6 GB |
| Q6_K | 0.8 GB | 0.33 GB | 1.3 GB |
| Q5_K_M | 0.7 GB | 0.33 GB | 1.2 GB |
| Q4_K_M | 0.6 GB | 0.33 GB | 1.1 GB |
| Q3_K_M | 0.5 GB | 0.33 GB | 0.9 GB |
| Q2_K | 0.4 GB | 0.33 GB | 0.8 GB |
| NVFP4cuda | 0.5 GB | 0.33 GB | 0.9 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run Gemma 3 1B Instruct natively (115)
- NVIDIA RTX 5090BF16 · 500.5 t/s
- NVIDIA RTX 5080BF16 · 268.1 t/s
- NVIDIA RTX 5070 TiBF16 · 250.3 t/s
- NVIDIA RTX 5070BF16 · 187.7 t/s
- NVIDIA RTX 5060 Ti 16GBBF16 · 125.1 t/s
Show 110 more
- NVIDIA RTX 5060 Ti 8GBBF16 · 125.1 t/s
- NVIDIA RTX 5060BF16 · 125.1 t/s
- NVIDIA RTX 5050BF16 · 89.4 t/s
- NVIDIA RTX 4090BF16 · 281.5 t/s
- NVIDIA RTX 4080BF16 · 200.3 t/s
- NVIDIA RTX 4070 Ti SUPERBF16 · 187.7 t/s
- NVIDIA RTX 4070 TiBF16 · 140.8 t/s
- NVIDIA RTX 4070 SUPERBF16 · 140.8 t/s
- NVIDIA RTX 4070BF16 · 140.8 t/s
- NVIDIA RTX 4060 Ti 16GBBF16 · 80.4 t/s
- NVIDIA RTX 4060BF16 · 76 t/s
- NVIDIA RTX 3090BF16 · 261.4 t/s
- NVIDIA RTX 3090 TiBF16 · 281.5 t/s
- NVIDIA RTX 3080 10GBBF16 · 212.3 t/s
- NVIDIA RTX 3060 12GBBF16 · 100.6 t/s
- NVIDIA B300 288GBBF16 · 2234.5 t/s
- NVIDIA B200 180GBBF16 · 2234.5 t/s
- NVIDIA H200 141GBBF16 · 1340.7 t/s
- NVIDIA H100 80GBBF16 · 935.7 t/s
- NVIDIA A100 80GBBF16 · 569.5 t/s
- NVIDIA A100 40GBBF16 · 434.3 t/s
- NVIDIA L40SBF16 · 241.3 t/s
- NVIDIA RTX A6000BF16 · 214.5 t/s
- NVIDIA RTX 4000 AdaBF16 · 89.4 t/s
- NVIDIA RTX 4500 AdaBF16 · 120.7 t/s
- NVIDIA RTX 5000 AdaBF16 · 160.9 t/s
- NVIDIA RTX 6000 AdaBF16 · 268.1 t/s
- NVIDIA RTX Pro 6000BF16 · 375.4 t/s
- NVIDIA DGX Spark (128GB)BF16 · 76.3 t/s
- AMD Radeon RX 7900 XTXBF16 · 268.1 t/s
- AMD Radeon RX 7900 XTBF16 · 223.4 t/s
- AMD Radeon RX 7900 GREBF16 · 160.9 t/s
- AMD Radeon RX 6800 XTBF16 · 143 t/s
- AMD Radeon PRO W7800BF16 · 160.9 t/s
- AMD Radeon PRO W7900BF16 · 241.3 t/s
- AMD Instinct MI300XBF16 · 1480.3 t/s
- AMD Radeon AI PRO R9700 32GBBF16 · 178.8 t/s
- AMD Strix Halo (128GB)BF16 · 71.5 t/s
- AMD Strix Halo (96GB)BF16 · 71.5 t/s
- AMD Strix Halo (64GB)BF16 · 71.5 t/s
- AMD Strix Halo (32GB)BF16 · 71.5 t/s
- Apple M5 Ultra (512GB)BF16 · 412.5 t/s
- Apple M5 Ultra (256GB)BF16 · 412.5 t/s
- Apple M5 Ultra (96GB)BF16 · 412.5 t/s
- Apple M5 Max (128GB)BF16 · 211.1 t/s
- Apple M5 Max (64GB)BF16 · 211.1 t/s
- Apple M5 Max (48GB)BF16 · 211.1 t/s
- Apple M5 Max (36GB)BF16 · 158.1 t/s
- Apple M5 Pro (64GB)BF16 · 105.5 t/s
- Apple M5 Pro (48GB)BF16 · 105.5 t/s
- Apple M5 Pro (24GB)BF16 · 105.5 t/s
- Apple M5 (32GB)BF16 · 52.6 t/s
- Apple M5 (16GB)BF16 · 52.6 t/s
- Apple M6 (32GB)BF16 · 58.4 t/s
- Apple M6 (16GB)BF16 · 58.4 t/s
- Apple M4 Max (128GB)BF16 · 187.7 t/s
- Apple M4 Max (64GB)BF16 · 187.7 t/s
- Apple M4 Max (48GB)BF16 · 187.7 t/s
- Apple M4 Max (36GB)BF16 · 140.9 t/s
- Apple M4 Pro (48GB)BF16 · 93.8 t/s
- Apple M4 Pro (24GB)BF16 · 93.8 t/s
- Apple M4 (32GB)BF16 · 41.3 t/s
- Apple M4 (16GB)BF16 · 41.3 t/s
- Apple M3 Ultra (512GB)BF16 · 281.5 t/s
- Apple M3 Ultra (256GB)BF16 · 281.5 t/s
- Apple M3 Ultra (96GB)BF16 · 281.5 t/s
- Apple M3 Max (128GB)BF16 · 137.5 t/s
- Apple M3 Max (96GB)BF16 · 103.1 t/s
- Apple M3 Max (64GB)BF16 · 137.5 t/s
- Apple M3 Max (48GB)BF16 · 137.5 t/s
- Apple M3 Max (36GB)BF16 · 103.1 t/s
- Apple M3 Pro (36GB)BF16 · 51.6 t/s
- Apple M3 Pro (18GB)BF16 · 51.6 t/s
- Apple M3 (24GB)BF16 · 34.4 t/s
- Apple M3 (16GB)BF16 · 34.4 t/s
- Apple M2 Ultra (192GB)BF16 · 275 t/s
- Apple M2 Ultra (64GB)BF16 · 275 t/s
- Apple M2 Max (96GB)BF16 · 137.5 t/s
- Apple M2 Max (64GB)BF16 · 137.5 t/s
- Apple M2 Max (32GB)BF16 · 137.5 t/s
- Apple M2 Pro (32GB)BF16 · 68.8 t/s
- Apple M2 Pro (16GB)BF16 · 68.8 t/s
- Apple M2 (24GB)BF16 · 34.4 t/s
- Apple M2 (16GB)BF16 · 34.4 t/s
- Apple M1 Ultra (128GB)BF16 · 275 t/s
- Apple M1 Ultra (64GB)BF16 · 275 t/s
- Apple M1 Max (64GB)BF16 · 137.5 t/s
- Apple M1 Max (32GB)BF16 · 137.5 t/s
- Apple M1 Pro (32GB)BF16 · 68.8 t/s
- Apple M1 Pro (16GB)BF16 · 68.8 t/s
- Apple M1 (16GB)BF16 · 23.4 t/s
- Intel Arc B580 12GBBF16 · 127.4 t/s
- Intel Arc B570 10GBBF16 · 106.1 t/s
- Intel Arc Pro B70 32GBBF16 · 169.8 t/s
- Intel Arc Pro B60 24GBBF16 · 106.1 t/s
- Intel Arc Pro B50 16GBBF16 · 62.6 t/s
- Intel Arc A770 16GBBF16 · 156.4 t/s
- Intel Arc A770 8GBBF16 · 143 t/s
- Intel Arc A750 8GBBF16 · 143 t/s
- Intel Arc A580 8GBBF16 · 143 t/s
- Intel Arc A380 6GBBF16 · 52 t/s
- Intel Arc A310 4GBBF16 · 34.6 t/s
- Intel Arc Pro A60 12GBBF16 · 107.3 t/s
- Intel Arc Pro A50 6GBBF16 · 53.6 t/s
- Intel Arc Pro A40 6GBBF16 · 53.6 t/s
- Intel Data Center GPU Max 1550BF16 · 915 t/s
- Intel Data Center GPU Max 1100BF16 · 343.3 t/s
- Intel Arc 140V (32GB)BF16 · 38.3 t/s
- Intel Arc 140V (16GB)BF16 · 38.3 t/s
- Intel Arc 130V (16GB)BF16 · 38.3 t/s
Plus 1 GPUs that run it with CPU offload (slower)
- CPU only (system RAM)BF16 · 17.2 t/s
Notes
Text-only; larger siblings support vision.
Compare Gemma 3 1B Instruct with other models
Frequently asked questions
- What are the VRAM requirements for Gemma 3 1B Instruct?
- Gemma 3 1B Instruct requires approximately 1.1 GB of VRAM at Q4_K_M quantization, 1.6 GB at Q8, and 2.6 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Gemma 3 1B Instruct have?
- Gemma 3 1B Instruct has 1 billion parameters.
- How capable is Gemma 3 1B Instruct?
- Gemma 3 1B Instruct has an MMLU-Pro score of 14.7, making it well-suited for lightweight tasks, prototyping, and resource-constrained environments.
- Can Gemma 3 1B Instruct run on a 16 GB GPU?
- Yes. Gemma 3 1B Instruct needs 1.1 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 5070 Ti.
- What is the smallest quantization for Gemma 3 1B Instruct that fits in 24 GB of VRAM?
- At BF16, Gemma 3 1B Instruct needs 2.6 GB, the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Gemma 3 1B Instruct locally?
- A 16 GB GPU is enough. At Q4_K_M, Gemma 3 1B Instruct needs 1.1 GB VRAM. Good options: RTX 4080 (16 GB), RTX 5070 Ti (16 GB).