Qwen3 14B
Qwen3 14B needs roughly 11.6 GB VRAM at Q4_K_M quantization (34.7 GB at FP16). 104 GPUs we track can run it fully in VRAM at 8k context.
104 GPUs run this natively · 12 with CPU offload
Qwen3 14B is a 14.8B parameter dense model developed by Alibaba. April 2025 dense model from the Qwen3 line, released under Apache 2.0 with a unified thinking/non-thinking mode toggled via the chat template or an enable_thinking flag.
To run Qwen3 14B locally: Q5_K_M needs roughly 10-11GB and fits on a 12GB GPU with some headroom, or comfortably on 16GB. Thinking mode generates longer outputs, so budget extra KV cache for reasoning-heavy prompts.
MMLU-Pro 61.03 puts it well ahead of Qwen2.5-14B on general reasoning, reflecting Qwen3's larger 36-trillion-token pretraining corpus.
VRAM at each quantization
Qwen3 14B natively supports a longer context window, but the table below is capped at 8k for comparability; KV cache grows linearly with context length.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 59.2 GB | 1.34 GB | 67.8 GB |
| BF16 | 29.6 GB | 1.34 GB | 34.7 GB |
| FP16 | 29.6 GB | 1.34 GB | 34.7 GB |
| Q8_0 | 15.7 GB | 1.34 GB | 19.1 GB |
| Q6_K | 12.2 GB | 1.34 GB | 15.1 GB |
| Q5_K_Mrec | 10.5 GB | 1.34 GB | 13.3 GB |
| Q4_K_M | 9.0 GB | 1.34 GB | 11.6 GB |
| Q3_K_M | 7.1 GB | 1.34 GB | 9.5 GB |
| Q2_K | 5.6 GB | 1.34 GB | 7.8 GB |
| NVFP4cuda | 7.4 GB | 1.34 GB | 9.8 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Qwen3 14B natively (104)
- NVIDIA RTX 5090NVFP4 · 133.2 t/s
- NVIDIA RTX 5080NVFP4 · 71.4 t/s
- NVIDIA RTX 5070 TiNVFP4 · 66.6 t/s
- NVIDIA RTX 5070NVFP4 · 50 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 33.3 t/s
Show 99 more
- NVIDIA RTX 4090Q8_0 · 38.4 t/s
- NVIDIA RTX 4080Q6_K · 34.5 t/s
- NVIDIA RTX 4070 Ti SUPERQ6_K · 32.4 t/s
- NVIDIA RTX 4070 TiQ3_K_M · 38.7 t/s
- NVIDIA RTX 4070 SUPERQ3_K_M · 38.7 t/s
- NVIDIA RTX 4070Q3_K_M · 38.7 t/s
- NVIDIA RTX 4060 Ti 16GBQ6_K · 13.9 t/s
- NVIDIA RTX 3090Q8_0 · 35.6 t/s
- NVIDIA RTX 3090 TiQ8_0 · 38.4 t/s
- NVIDIA RTX 3080 10GBQ3_K_M · 58.4 t/s
- NVIDIA RTX 3060 12GBQ3_K_M · 27.7 t/s
- NVIDIA B300 288GBBF16 · 168.1 t/s
- NVIDIA B200 180GBBF16 · 168.1 t/s
- NVIDIA H200 141GBBF16 · 100.8 t/s
- NVIDIA H100 80GBBF16 · 70.4 t/s
- NVIDIA A100 80GBBF16 · 42.8 t/s
- NVIDIA A100 40GBBF16 · 32.7 t/s
- NVIDIA L40SBF16 · 18.1 t/s
- NVIDIA RTX A6000BF16 · 16.1 t/s
- NVIDIA RTX 4000 AdaQ6_K · 15.4 t/s
- NVIDIA RTX 4500 AdaQ8_0 · 16.4 t/s
- NVIDIA RTX 5000 AdaQ8_0 · 21.9 t/s
- NVIDIA RTX 6000 AdaBF16 · 20.2 t/s
- NVIDIA RTX Pro 6000BF16 · 28.2 t/s
- NVIDIA DGX Spark (128GB)BF16 · 5.7 t/s
- AMD Radeon RX 7900 XTXQ8_0 · 36.5 t/s
- AMD Radeon RX 7900 XTQ6_K · 38.5 t/s
- AMD Radeon RX 7900 GREQ6_K · 27.7 t/s
- AMD Radeon RX 6800 XTQ6_K · 24.7 t/s
- AMD Radeon PRO W7800Q8_0 · 21.9 t/s
- AMD Radeon PRO W7900BF16 · 18.1 t/s
- AMD Instinct MI300XBF16 · 111.3 t/s
- AMD Radeon AI PRO R9700 32GBQ8_0 · 24.4 t/s
- AMD Strix Halo (128GB)BF16 · 5.4 t/s
- AMD Strix Halo (96GB)BF16 · 5.4 t/s
- AMD Strix Halo (64GB)BF16 · 5.4 t/s
- AMD Strix Halo (32GB)Q8_0 · 9.7 t/s
- Apple M5 Ultra (512GB)BF16 · 31 t/s
- Apple M5 Ultra (256GB)BF16 · 31 t/s
- Apple M5 Ultra (96GB)BF16 · 31 t/s
- Apple M5 Max (128GB)BF16 · 15.9 t/s
- Apple M5 Max (64GB)BF16 · 15.9 t/s
- Apple M5 Max (48GB)BF16 · 15.9 t/s
- Apple M5 Max (36GB)Q8_0 · 21.6 t/s
- Apple M5 Pro (64GB)BF16 · 7.9 t/s
- Apple M5 Pro (48GB)BF16 · 7.9 t/s
- Apple M5 Pro (24GB)Q6_K · 18.2 t/s
- Apple M5 (32GB)Q8_0 · 7.2 t/s
- Apple M5 (16GB)Q2_K · 17.5 t/s
- Apple M6 (32GB)Q8_0 · 8 t/s
- Apple M6 (16GB)Q2_K · 19.5 t/s
- Apple M4 Max (128GB)BF16 · 14.1 t/s
- Apple M4 Max (64GB)BF16 · 14.1 t/s
- Apple M4 Max (48GB)BF16 · 14.1 t/s
- Apple M4 Max (36GB)Q8_0 · 19.2 t/s
- Apple M4 Pro (48GB)BF16 · 7.1 t/s
- Apple M4 Pro (24GB)Q6_K · 16.2 t/s
- Apple M4 (32GB)Q8_0 · 5.6 t/s
- Apple M4 (16GB)Q2_K · 13.8 t/s
- Apple M3 Ultra (512GB)BF16 · 21.2 t/s
- Apple M3 Ultra (256GB)BF16 · 21.2 t/s
- Apple M3 Ultra (96GB)BF16 · 21.2 t/s
- Apple M3 Max (128GB)BF16 · 10.3 t/s
- Apple M3 Max (96GB)BF16 · 7.8 t/s
- Apple M3 Max (64GB)BF16 · 10.3 t/s
- Apple M3 Max (48GB)BF16 · 10.3 t/s
- Apple M3 Max (36GB)Q8_0 · 14.1 t/s
- Apple M3 Pro (36GB)Q8_0 · 7 t/s
- Apple M3 Pro (18GB)Q3_K_M · 14.2 t/s
- Apple M3 (24GB)Q6_K · 5.9 t/s
- Apple M3 (16GB)Q2_K · 11.5 t/s
- Apple M2 Ultra (192GB)BF16 · 20.7 t/s
- Apple M2 Ultra (64GB)BF16 · 20.7 t/s
- Apple M2 Max (96GB)BF16 · 10.3 t/s
- Apple M2 Max (64GB)BF16 · 10.3 t/s
- Apple M2 Max (32GB)Q8_0 · 18.7 t/s
- Apple M2 Pro (32GB)Q8_0 · 9.4 t/s
- Apple M2 Pro (16GB)Q2_K · 22.9 t/s
- Apple M2 (24GB)Q6_K · 5.9 t/s
- Apple M2 (16GB)Q2_K · 11.5 t/s
- Apple M1 Ultra (128GB)BF16 · 20.7 t/s
- Apple M1 Ultra (64GB)BF16 · 20.7 t/s
- Apple M1 Max (64GB)BF16 · 10.3 t/s
- Apple M1 Max (32GB)Q8_0 · 18.7 t/s
- Apple M1 Pro (32GB)Q8_0 · 9.4 t/s
- Apple M1 Pro (16GB)Q2_K · 22.9 t/s
- Apple M1 (16GB)Q2_K · 7.8 t/s
- Intel Arc B580 12GBQ3_K_M · 35 t/s
- Intel Arc B570 10GBQ3_K_M · 29.2 t/s
- Intel Arc Pro B70 32GBQ8_0 · 23.1 t/s
- Intel Arc Pro B60 24GBQ8_0 · 14.5 t/s
- Intel Arc Pro B50 16GBQ6_K · 10.8 t/s
- Intel Arc A770 16GBQ6_K · 27 t/s
- Intel Arc Pro A60 12GBQ3_K_M · 29.5 t/s
- Intel Data Center GPU Max 1550BF16 · 68.8 t/s
- Intel Data Center GPU Max 1100BF16 · 25.8 t/s
- Intel Arc 140V (32GB)Q8_0 · 5.2 t/s
- Intel Arc 140V (16GB)Q2_K · 12.8 t/s
- Intel Arc 130V (16GB)Q2_K · 12.8 t/s
Plus 12 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5060 Ti 8GBNVFP4 · 13.9 t/s
- NVIDIA RTX 5060NVFP4 · 13.9 t/s
- NVIDIA RTX 5050NVFP4 · 12.2 t/s
- NVIDIA RTX 4060Q8_0 · 2.4 t/s
- Intel Arc A770 8GBQ8_0 · 2.5 t/s
- Intel Arc A750 8GBQ8_0 · 2.5 t/s
- Intel Arc A580 8GBQ8_0 · 2.5 t/s
- Intel Arc A380 6GBQ8_0 · 2 t/s
- Intel Arc A310 4GBQ8_0 · 1.7 t/s
- Intel Arc Pro A50 6GBQ8_0 · 2 t/s
- Intel Arc Pro A40 6GBQ8_0 · 2 t/s
- CPU only (system RAM)Q8_0 · 2.3 t/s
Notes
Supports thinking and non-thinking modes.
Continue reading
Frequently asked questions
- What are the VRAM requirements for Qwen3 14B?
- Qwen3 14B requires approximately 11.6 GB of VRAM at Q4_K_M quantization, 19.1 GB at Q8, and 34.7 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Qwen3 14B have?
- Qwen3 14B has 14.8 billion parameters.
- How capable is Qwen3 14B?
- With an MMLU-Pro score of 61.03, Qwen3 14B delivers solid general-purpose performance suitable for most everyday tasks and professional use.
- Can Qwen3 14B run on a 16 GB GPU?
- Yes. Qwen3 14B needs 11.6 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 5070 Ti.
- What is the smallest quantization for Qwen3 14B that fits in 24 GB of VRAM?
- At NVFP4, Qwen3 14B needs 9.8 GB, the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Qwen3 14B locally?
- A 16 GB GPU is enough. At Q4_K_M, Qwen3 14B needs 11.6 GB VRAM. Good options: RTX 4080 (16 GB), RTX 5070 Ti (16 GB).