GPT-OSS 20B
GPT-OSS 20B needs roughly 14.8 GB VRAM at Q4_K_M quantization (47.5 GB at FP16). 84 GPUs we track can run it fully in VRAM at 8k context.
84 GPUs run this natively · 11 with CPU offload
GPT-OSS 20B is a Mixture of Experts (MoE) model with 21B total parameters but only 4B active per token developed by OpenAI. August 2025 21B MoE with 4B active — matches o3-mini on key benchmarks.
To run GPT-OSS 20B locally: Q5_K_M ~14-16GB — fits on 16GB GPUs. Best reasoning model for 16GB hardware. As a MoE model, inference speed depends on active parameters (4B) rather than total size.
GPQA 71.5% at 21B scale — exceptional reasoning efficiency.
VRAM at each quantization
Figures below assume 8k context; KV cache grows linearly as context length increases.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 84.0 GB | 0.40 GB | 94.5 GB |
| BF16 | 42.0 GB | 0.40 GB | 47.5 GB |
| FP16 | 42.0 GB | 0.40 GB | 47.5 GB |
| Q8_0 | 22.3 GB | 0.40 GB | 25.4 GB |
| Q6_K | 17.2 GB | 0.40 GB | 19.8 GB |
| Q5_K_Mrec | 14.9 GB | 0.40 GB | 17.2 GB |
| Q4_K_M | 12.8 GB | 0.40 GB | 14.8 GB |
| Q3_K_M | 10.1 GB | 0.40 GB | 11.8 GB |
| Q2_K | 8.0 GB | 0.40 GB | 9.4 GB |
| NVFP4cuda | 10.5 GB | 0.40 GB | 12.2 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run GPT-OSS 20B natively (84)
- NVIDIA RTX 5090NVFP4 · 164.8 t/s
- NVIDIA RTX 5080NVFP4 · 88.3 t/s
- NVIDIA RTX 5070 TiNVFP4 · 82.4 t/s
- NVIDIA RTX 5070Q2_K · 79.7 t/s
- NVIDIA RTX 5060 Ti 16GBNVFP4 · 41.2 t/s
- NVIDIA RTX 4090NVFP4 · 92.7 t/s
- NVIDIA RTX 4080NVFP4 · 65.9 t/s
- NVIDIA RTX 4070 TiQ2_K · 59.8 t/s
- NVIDIA RTX 4070Q2_K · 59.8 t/s
- NVIDIA RTX 4060 Ti 16GBNVFP4 · 26.5 t/s
- NVIDIA RTX 3090NVFP4 · 86.1 t/s
- NVIDIA RTX 3090 TiNVFP4 · 92.7 t/s
- NVIDIA RTX 3080 10GBQ2_K · 90.1 t/s
- NVIDIA RTX 3060 12GBQ2_K · 42.7 t/s
- NVIDIA H100 80GBBF16 · 80.4 t/s
- NVIDIA A100 80GBBF16 · 49 t/s
- NVIDIA A100 40GBNVFP4 · 143 t/s
- NVIDIA L40SNVFP4 · 79.4 t/s
- NVIDIA RTX A6000NVFP4 · 70.6 t/s
- NVIDIA RTX 4000 AdaNVFP4 · 29.4 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 39.7 t/s
- NVIDIA RTX 5000 AdaNVFP4 · 53 t/s
- NVIDIA RTX 6000 AdaNVFP4 · 88.3 t/s
- NVIDIA RTX Pro 6000BF16 · 32.3 t/s
- NVIDIA DGX Spark (128GB)FP32 · 3.3 t/s
- AMD Radeon RX 7900 XTXQ6_K · 55 t/s
- AMD Radeon RX 7900 XTQ5_K_M · 52.5 t/s
- AMD Radeon RX 7900 GREQ4_K_M · 43.9 t/s
- AMD Radeon RX 6800 XTQ4_K_M · 39 t/s
- AMD Radeon PRO W7800Q8_0 · 25.7 t/s
- AMD Radeon PRO W7900Q8_0 · 38.5 t/s
- AMD Instinct MI300XFP32 · 64.1 t/s
- AMD Radeon AI Pro 9700 32GBQ8_0 · 28.5 t/s
- AMD Strix Halo (128GB)FP32 · 3.1 t/s
- AMD Strix Halo (96GB)BF16 · 6.1 t/s
- AMD Strix Halo (64GB)BF16 · 6.1 t/s
- Apple M5 Max (128GB)FP32 · 9.1 t/s
- Apple M5 Max (64GB)BF16 · 18.1 t/s
- Apple M5 Max (48GB)Q8_0 · 33.7 t/s
- Apple M5 Pro (48GB)Q8_0 · 16.8 t/s
- Apple M5 Pro (36GB)Q8_0 · 16.8 t/s
- Apple M5 Pro (24GB)Q4_K_M · 28.8 t/s
- Apple M5 (32GB)Q6_K · 10.8 t/s
- Apple M4 Ultra (384GB)FP32 · 16.3 t/s
- Apple M4 Ultra (192GB)FP32 · 16.3 t/s
- Apple M4 Max (128GB)FP32 · 8.1 t/s
- Apple M4 Max (96GB)BF16 · 16.1 t/s
- Apple M4 Max (64GB)BF16 · 16.1 t/s
- Apple M4 Max (48GB)Q8_0 · 30 t/s
- Apple M4 Pro (48GB)Q8_0 · 15 t/s
- Apple M4 Pro (24GB)Q4_K_M · 25.6 t/s
- Apple M4 (32GB)Q6_K · 8.5 t/s
- Apple M3 Ultra (512GB)FP32 · 12.2 t/s
- Apple M3 Ultra (256GB)FP32 · 12.2 t/s
- Apple M3 Ultra (96GB)BF16 · 24.2 t/s
- Apple M3 Max (128GB)FP32 · 6 t/s
- Apple M3 Max (96GB)BF16 · 11.8 t/s
- Apple M3 Max (64GB)BF16 · 11.8 t/s
- Apple M3 Max (48GB)Q8_0 · 22 t/s
- Apple M3 Max (36GB)Q8_0 · 22 t/s
- Apple M3 Pro (36GB)Q8_0 · 8.2 t/s
- Apple M3 Pro (18GB)Q2_K · 21.9 t/s
- Apple M3 (24GB)Q4_K_M · 9.4 t/s
- Apple M2 Ultra (384GB)FP32 · 11.9 t/s
- Apple M2 Ultra (192GB)FP32 · 11.9 t/s
- Apple M2 Max (96GB)BF16 · 11.8 t/s
- Apple M2 Max (64GB)BF16 · 11.8 t/s
- Apple M2 Max (32GB)Q6_K · 28.2 t/s
- Apple M2 Pro (32GB)Q6_K · 14.1 t/s
- Apple M2 (24GB)Q4_K_M · 9.4 t/s
- Apple M1 Ultra (128GB)FP32 · 11.9 t/s
- Apple M1 Ultra (64GB)BF16 · 23.6 t/s
- Apple M1 Max (64GB)BF16 · 11.8 t/s
- Apple M1 Max (32GB)Q6_K · 28.2 t/s
- Apple M1 Pro (32GB)Q6_K · 14.1 t/s
- Intel Arc B580 12GBQ2_K · 54.1 t/s
- Intel Arc B570 10GBQ2_K · 45.1 t/s
- Intel Arc Pro B70 24GBQ6_K · 26.1 t/s
- Intel Arc Pro B60 24GBQ6_K · 21.8 t/s
- Intel Arc A770 16GBQ4_K_M · 42.7 t/s
- Intel Arc Pro A60 12GBQ2_K · 45.5 t/s
- Intel Data Center GPU Max 1550FP32 · 39.6 t/s
- Intel Data Center GPU Max 1100Q8_0 · 54.8 t/s
- Intel Arc 140V (32GB)Q6_K · 7.8 t/s
Plus 11 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5060NVFP4 · 9.7 t/s
- NVIDIA RTX 5050NVFP4 · 9.1 t/s
- NVIDIA RTX 4060NVFP4 · 8.8 t/s
- Intel Arc A770 8GBQ8_0 · 2.5 t/s
- Intel Arc A750 8GBQ8_0 · 2.5 t/s
- Intel Arc A580 8GBQ8_0 · 2.5 t/s
- Intel Arc A380 6GBQ8_0 · 2.1 t/s
- Intel Arc A310 4GBQ8_0 · 1.9 t/s
- Intel Arc Pro A50 6GBQ8_0 · 2.1 t/s
- Intel Arc Pro A40 6GBQ8_0 · 2.1 t/s
- CPU only (system RAM)Q8_0 · 2.7 t/s
Notes
Smaller sibling of GPT-OSS 120B. Matches o3-mini on key benchmarks; runs on 16 GB of VRAM.
Frequently asked questions
- What are the VRAM requirements for GPT-OSS 20B?
- GPT-OSS 20B requires approximately 14.8 GB of VRAM at Q4_K_M quantization, 25.5 GB at Q8, and 47.5 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does GPT-OSS 20B have?
- GPT-OSS 20B has 21 billion total parameters, but only 4 billion are active per token thanks to its Mixture of Experts (MoE) architecture. This makes inference significantly faster than the total parameter count suggests.
- How capable is GPT-OSS 20B?
- With an MMLU-Pro score of 67.86, GPT-OSS 20B delivers solid general-purpose performance suitable for most everyday tasks and professional use.
- Can GPT-OSS 20B run on a 16 GB GPU?
- Yes. GPT-OSS 20B needs 14.8 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 4070 Ti Super.
- What is the smallest quantization for GPT-OSS 20B that fits in 24 GB of VRAM?
- At NVFP4, GPT-OSS 20B needs 12.2 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run GPT-OSS 20B locally?
- A 16 GB GPU is enough. At Q4_K_M, GPT-OSS 20B needs 14.8 GB VRAM. Good options: RTX 4080 (16 GB), RTX 4070 Ti Super (16 GB).