GPT-OSS 120B
GPT-OSS 120B needs roughly 80.5 GB VRAM at Q4_K_M quantization (262.8 GB at FP16). 29 GPUs we track can run it fully in VRAM at 8k context.
29 GPUs run this natively · 10 with CPU offload
GPT-OSS 120B is a Mixture of Experts (MoE) model with 117B total parameters but only 5B active per token developed by OpenAI. August 2025 117B MoE with 5B active. Alternating sliding+full attention. Apache 2.0.
To run GPT-OSS 120B locally: Q4_K_M ~70-80GB — fits on 80GB GPU or Mac Studio. Best open reasoning model at this size. As a MoE model, inference speed depends on active parameters (5B) rather than total size.
GPQA 80.1% — near-parity with o4-mini. Fits on single 80GB GPU at Q4.
VRAM at each quantization
Calculated at 8k context. Since KV cache scales linearly with context, longer sessions need more VRAM than shown here.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 468.0 GB | 0.60 GB | 524.8 GB |
| BF16 | 234.0 GB | 0.60 GB | 262.8 GB |
| FP16 | 234.0 GB | 0.60 GB | 262.8 GB |
| Q8_0 | 124.4 GB | 0.60 GB | 140.0 GB |
| Q6_K | 96.1 GB | 0.60 GB | 108.3 GB |
| Q5_K_M | 83.3 GB | 0.60 GB | 94.0 GB |
| Q4_K_Mrec | 71.3 GB | 0.60 GB | 80.5 GB |
| Q3_K_M | 56.3 GB | 0.60 GB | 63.7 GB |
| Q2_K | 44.6 GB | 0.60 GB | 50.6 GB |
| NVFP4cuda | 58.5 GB | 0.60 GB | 66.2 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run GPT-OSS 120B natively (29)
- NVIDIA H100 80GBNVFP4 · 243.6 t/s
- NVIDIA A100 80GBNVFP4 · 148.3 t/s
- NVIDIA RTX Pro 6000NVFP4 · 97.7 t/s
- NVIDIA DGX Spark (128GB)NVFP4 · 19.9 t/s
- AMD Instinct MI300XQ8_0 · 188 t/s
- AMD Strix Halo (128GB)Q6_K · 11.6 t/s
- AMD Strix Halo (96GB)Q4_K_M · 15.5 t/s
- AMD Strix Halo (64GB)Q2_K · 23.9 t/s
- Apple M5 Max (128GB)Q6_K · 34.4 t/s
- Apple M5 Max (64GB)Q2_K · 70.6 t/s
- Apple M4 Ultra (384GB)BF16 · 25.7 t/s
- Apple M4 Ultra (192GB)Q8_0 · 47.7 t/s
- Apple M4 Max (128GB)Q6_K · 30.6 t/s
- Apple M4 Max (96GB)Q4_K_M · 40.6 t/s
- Apple M4 Max (64GB)Q2_K · 62.8 t/s
- Apple M3 Ultra (512GB)BF16 · 19.3 t/s
- Apple M3 Ultra (256GB)Q8_0 · 35.8 t/s
- Apple M3 Ultra (96GB)Q4_K_M · 60.9 t/s
- Apple M3 Max (128GB)Q6_K · 22.4 t/s
- Apple M3 Max (96GB)Q4_K_M · 29.8 t/s
- Apple M3 Max (64GB)Q2_K · 46 t/s
- Apple M2 Ultra (384GB)BF16 · 18.9 t/s
- Apple M2 Ultra (192GB)Q8_0 · 34.9 t/s
- Apple M2 Max (96GB)Q4_K_M · 29.8 t/s
- Apple M2 Max (64GB)Q2_K · 46 t/s
- Apple M1 Ultra (128GB)Q6_K · 44.8 t/s
- Apple M1 Ultra (64GB)Q2_K · 92 t/s
- Apple M1 Max (64GB)Q2_K · 46 t/s
- Intel Data Center GPU Max 1550Q6_K · 149 t/s
Plus 10 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5090Q2_K · 10.5 t/s
- NVIDIA A100 40GBQ3_K_M · 8.3 t/s
- NVIDIA L40SNVFP4 · 10.6 t/s
- NVIDIA RTX A6000NVFP4 · 10.5 t/s
- NVIDIA RTX 5000 AdaQ2_K · 9.7 t/s
- NVIDIA RTX 6000 AdaNVFP4 · 10.8 t/s
- AMD Radeon PRO W7800Q2_K · 9.7 t/s
- AMD Radeon PRO W7900Q3_K_M · 12.4 t/s
- AMD Radeon AI Pro 9700 32GBQ2_K · 9.8 t/s
- Intel Data Center GPU Max 1100Q3_K_M · 13 t/s
Notes
Alternating sliding+full attention MoE. Near-parity with o4-mini; fits on a single 80 GB GPU at q4.
Frequently asked questions
- What are the VRAM requirements for GPT-OSS 120B?
- GPT-OSS 120B requires approximately 80.5 GB of VRAM at Q4_K_M quantization, 140.0 GB at Q8, and 262.8 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does GPT-OSS 120B have?
- GPT-OSS 120B has 117 billion total parameters, but only 5 billion are active per token thanks to its Mixture of Experts (MoE) architecture. This makes inference significantly faster than the total parameter count suggests.
- How capable is GPT-OSS 120B?
- GPT-OSS 120B achieves an MMLU-Pro score of 80.7, placing it among the most capable open-weight models available — competitive with frontier systems on general knowledge and reasoning.
- Can GPT-OSS 120B run on a 16 GB GPU?
- No. At Q4_K_M, GPT-OSS 120B needs 80.5 GB of VRAM — more than 16 GB. You will need a multi-GPU server.
- Can GPT-OSS 120B run on a 24 GB GPU?
- No. Even at Q4_K_M, GPT-OSS 120B needs 80.5 GB. Consider a multi-GPU server with 80 GB+ total VRAM.
- What is the smallest quantization for GPT-OSS 120B that fits in 24 GB of VRAM?
- GPT-OSS 120B cannot fit in 24 GB of VRAM at any standard quantization level. The minimum needed is 50.6 GB at Q2_K.
- What GPU do I need to run GPT-OSS 120B locally?
- You need a multi-GPU server. At Q4_K_M, GPT-OSS 120B needs 80.5 GB VRAM, more than any single consumer GPU. Consider 2–4× H100 or A100 GPUs.