Qwen3 30B-A3B (MoE)
Qwen3 30B-A3B (MoE) needs roughly 21.4 GB VRAM at Q4_K_M quantization (68.1 GB at FP16). 75 GPUs we track can run it fully in VRAM at 8k context.
75 GPUs run this natively · 19 with CPU offload
Qwen3 30B-A3B (MoE) is a Mixture of Experts (MoE) model with 30B total parameters but only 3B active per token developed by Alibaba. Ultra-efficient MoE with 30B total parameters but only 3B active per token.
To run Qwen3 30B-A3B (MoE) locally: Q4_K_M needs ~18-20GB — runs on 24GB GPUs with excellent speed due to low active parameter count. As a MoE model, inference speed depends on active parameters (3B) rather than total size.
MoE architecture delivers 30B-class quality at 3B inference cost — exceptional tokens/sec when it fits.
VRAM at each quantization
Numbers here are computed at 8k context. Because KV cache grows linearly with context length, expect higher totals at longer sequence lengths.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 120.0 GB | 0.81 GB | 135.3 GB |
| BF16 | 60.0 GB | 0.81 GB | 68.1 GB |
| FP16 | 60.0 GB | 0.81 GB | 68.1 GB |
| Q8_0 | 31.9 GB | 0.81 GB | 36.6 GB |
| Q6_K | 24.6 GB | 0.81 GB | 28.5 GB |
| Q5_K_M | 21.4 GB | 0.81 GB | 24.8 GB |
| Q4_K_Mrec | 18.3 GB | 0.81 GB | 21.4 GB |
| Q3_K_M | 14.4 GB | 0.81 GB | 17.1 GB |
| Q2_K | 11.4 GB | 0.81 GB | 13.7 GB |
| NVFP4cuda | 15.0 GB | 0.81 GB | 17.7 GB |
Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.
Benchmarks
GPUs that run Qwen3 30B-A3B (MoE) natively (75)
- NVIDIA RTX 5090NVFP4 · 200.6 t/s
- NVIDIA RTX 5080Q2_K · 135.2 t/s
- NVIDIA RTX 5070 TiQ2_K · 126.2 t/s
- NVIDIA RTX 5060 Ti 16GBQ2_K · 63.1 t/s
- NVIDIA RTX 4090NVFP4 · 112.9 t/s
- NVIDIA RTX 4080Q2_K · 101 t/s
- NVIDIA RTX 4060 Ti 16GBQ2_K · 40.6 t/s
- NVIDIA RTX 3090NVFP4 · 104.8 t/s
- NVIDIA RTX 3090 TiNVFP4 · 112.9 t/s
- NVIDIA H100 80GBBF16 · 104.7 t/s
- NVIDIA A100 80GBBF16 · 63.7 t/s
- NVIDIA A100 40GBNVFP4 · 174.1 t/s
- NVIDIA L40SNVFP4 · 96.7 t/s
- NVIDIA RTX A6000NVFP4 · 86 t/s
- NVIDIA RTX 4000 AdaNVFP4 · 35.8 t/s
- NVIDIA RTX 4500 AdaNVFP4 · 48.4 t/s
- NVIDIA RTX 5000 AdaNVFP4 · 64.5 t/s
- NVIDIA RTX 6000 AdaNVFP4 · 107.5 t/s
- NVIDIA RTX Pro 6000BF16 · 42 t/s
- NVIDIA DGX Spark (128GB)BF16 · 8.5 t/s
- AMD Radeon RX 7900 XTXQ4_K_M · 90.5 t/s
- AMD Radeon RX 7900 XTQ3_K_M · 92.6 t/s
- AMD Radeon RX 7900 GREQ2_K · 81.1 t/s
- AMD Radeon RX 6800 XTQ2_K · 72.1 t/s
- AMD Radeon PRO W7800Q6_K · 41.5 t/s
- AMD Radeon PRO W7900Q8_0 · 49.1 t/s
- AMD Instinct MI300XFP32 · 84.4 t/s
- AMD Radeon AI Pro 9700 32GBQ6_K · 46.1 t/s
- AMD Strix Halo (128GB)BF16 · 8 t/s
- AMD Strix Halo (96GB)BF16 · 8 t/s
- AMD Strix Halo (64GB)Q8_0 · 14.6 t/s
- Apple M5 Max (128GB)BF16 · 23.6 t/s
- Apple M5 Max (64GB)Q8_0 · 43 t/s
- Apple M5 Max (48GB)Q8_0 · 43 t/s
- Apple M5 Pro (48GB)Q8_0 · 21.5 t/s
- Apple M5 Pro (36GB)Q5_K_M · 31 t/s
- Apple M5 Pro (24GB)Q2_K · 53.2 t/s
- Apple M5 (32GB)Q4_K_M · 17.8 t/s
- Apple M4 Ultra (384GB)FP32 · 21.4 t/s
- Apple M4 Ultra (192GB)FP32 · 21.4 t/s
- Apple M4 Max (128GB)BF16 · 21 t/s
- Apple M4 Max (96GB)BF16 · 21 t/s
- Apple M4 Max (64GB)Q8_0 · 38.2 t/s
- Apple M4 Max (48GB)Q8_0 · 38.2 t/s
- Apple M4 Pro (48GB)Q8_0 · 19.1 t/s
- Apple M4 Pro (24GB)Q2_K · 47.3 t/s
- Apple M4 (32GB)Q4_K_M · 13.9 t/s
- Apple M3 Ultra (512GB)FP32 · 16.1 t/s
- Apple M3 Ultra (256GB)FP32 · 16.1 t/s
- Apple M3 Ultra (96GB)BF16 · 31.5 t/s
- Apple M3 Max (128GB)BF16 · 15.4 t/s
- Apple M3 Max (96GB)BF16 · 15.4 t/s
- Apple M3 Max (64GB)Q8_0 · 28 t/s
- Apple M3 Max (48GB)Q8_0 · 28 t/s
- Apple M3 Max (36GB)Q5_K_M · 40.4 t/s
- Apple M3 Pro (36GB)Q5_K_M · 15.1 t/s
- Apple M3 (24GB)Q2_K · 17.3 t/s
- Apple M2 Ultra (384GB)FP32 · 15.7 t/s
- Apple M2 Ultra (192GB)FP32 · 15.7 t/s
- Apple M2 Max (96GB)BF16 · 15.4 t/s
- Apple M2 Max (64GB)Q8_0 · 28 t/s
- Apple M2 Max (32GB)Q4_K_M · 46.4 t/s
- Apple M2 Pro (32GB)Q4_K_M · 23.2 t/s
- Apple M2 (24GB)Q2_K · 17.3 t/s
- Apple M1 Ultra (128GB)BF16 · 30.8 t/s
- Apple M1 Ultra (64GB)Q8_0 · 56 t/s
- Apple M1 Max (64GB)Q8_0 · 28 t/s
- Apple M1 Max (32GB)Q4_K_M · 46.4 t/s
- Apple M1 Pro (32GB)Q4_K_M · 23.2 t/s
- Intel Arc Pro B70 24GBQ4_K_M · 43 t/s
- Intel Arc Pro B60 24GBQ4_K_M · 35.8 t/s
- Intel Arc A770 16GBQ2_K · 78.9 t/s
- Intel Data Center GPU Max 1550BF16 · 102.3 t/s
- Intel Data Center GPU Max 1100Q8_0 · 69.9 t/s
- Intel Arc 140V (32GB)Q4_K_M · 12.9 t/s
Plus 19 GPUs that run it with CPU offload (slower)
- NVIDIA RTX 5070NVFP4 · 13.5 t/s
- NVIDIA RTX 5060NVFP4 · 7.7 t/s
- NVIDIA RTX 5050NVFP4 · 7.5 t/s
- NVIDIA RTX 4070 TiNVFP4 · 12.9 t/s
- NVIDIA RTX 4070NVFP4 · 12.9 t/s
- NVIDIA RTX 4060NVFP4 · 7.3 t/s
- NVIDIA RTX 3080 10GBNVFP4 · 10 t/s
- NVIDIA RTX 3060 12GBNVFP4 · 12.1 t/s
- Intel Arc B580 12GBQ8_0 · 3.2 t/s
- Intel Arc B570 10GBQ6_K · 4.2 t/s
- Intel Arc A770 8GBQ6_K · 3.8 t/s
- Intel Arc A750 8GBQ6_K · 3.8 t/s
- Intel Arc A580 8GBQ6_K · 3.8 t/s
- Intel Arc A380 6GBQ6_K · 3.4 t/s
- Intel Arc A310 4GBQ6_K · 3.1 t/s
- Intel Arc Pro A60 12GBQ8_0 · 3.2 t/s
- Intel Arc Pro A50 6GBQ6_K · 3.4 t/s
- Intel Arc Pro A40 6GBQ6_K · 3.4 t/s
- CPU only (system RAM)Q5_K_M · 5 t/s
Notes
30B total, only 3B active per token — fast inference when it fits.
Compare Qwen3 30B-A3B (MoE) with other models
Continue reading
Frequently asked questions
- What are the VRAM requirements for Qwen3 30B-A3B (MoE)?
- Qwen3 30B-A3B (MoE) requires approximately 21.4 GB of VRAM at Q4_K_M quantization, 36.6 GB at Q8, and 68.1 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Qwen3 30B-A3B (MoE) have?
- Qwen3 30B-A3B (MoE) has 30 billion total parameters, but only 3 billion are active per token thanks to its Mixture of Experts (MoE) architecture. This makes inference significantly faster than the total parameter count suggests.
- How capable is Qwen3 30B-A3B (MoE)?
- With an MMLU-Pro score of 61.49, Qwen3 30B-A3B (MoE) delivers solid general-purpose performance suitable for most everyday tasks and professional use.
- Can Qwen3 30B-A3B (MoE) run on a 16 GB GPU?
- No. At Q4_K_M, Qwen3 30B-A3B (MoE) needs 21.4 GB of VRAM — more than 16 GB. You will need a 24 GB GPU like the RTX 4090 or RTX 3090.
- Can Qwen3 30B-A3B (MoE) run on a 24 GB GPU?
- Yes. Qwen3 30B-A3B (MoE) fits in a 24 GB GPU at Q4_K_M, requiring 21.4 GB VRAM. GPUs with 24 GB include the RTX 4090, RTX 3090, and RTX 3090 Ti.
- What is the smallest quantization for Qwen3 30B-A3B (MoE) that fits in 24 GB of VRAM?
- At NVFP4, Qwen3 30B-A3B (MoE) needs 17.7 GB — the highest-quality quantization that fits in 24 GB of VRAM.
- What GPU do I need to run Qwen3 30B-A3B (MoE) locally?
- A 24 GB GPU is the minimum. At Q4_K_M, Qwen3 30B-A3B (MoE) needs 21.4 GB VRAM. Good options: RTX 4090 (24 GB), RTX 3090 (24 GB).