NVIDIA RTX 4080
The NVIDIA RTX 4080 has 16 GB VRAM and 717 GB/s memory bandwidth. It can run 45 of our 97 tracked models natively in VRAM at 8k context.
With 16 GB GDDR6X, the NVIDIA RTX 4080 is a consumer-tier GPU that can run 45 models natively. This site's calculator puts Qwen3 8B at its recommended Q5_K_M quant at 67.5 tok/s (7.73 GB, fits with room for context) and Qwen3 14B at its recommended Q5_K_M at 39.2 tok/s (13.31 GB, fits but with only about 1.9 GB of headroom left in the roughly 15.2 GB this card actually has usable). GPT-OSS 20B, a model OpenAI explicitly sized to fit a 16GB card, fits at its recommended Q4_K_M (14.55 GB, 62 tok/s), with barely 0.6 GB of that usable budget left over. Qwen 3.6 27B draws the real ceiling: its recommended Q4_K_M build (19.02 GB) overshoots this card's usable VRAM and drops to a CPU-offloaded 9.6 tok/s, but the next quant down, Q3_K_M (15.15 GB), still fits natively at 34.5 tok/s, decoding within a rounding error of GPT-OSS 20B despite having more total parameters. Qwen3 32B doesn't fit natively at any quantization on this site's standard ladder: even Q3_K_M needs 19.17 GB against roughly 15.2 GB of usable VRAM, landing at 9.1 tok/s offloaded rather than a native fit, so 27B-class is closer to this card's real dense-model ceiling than the 32B some marketing copy implies. GPT-OSS 120B doesn't fit at all, even with full system-RAM offload: its 62.6 GB MXFP4 weight floor alone is well past this card's roughly 41 GB usable VRAM+RAM budget. Measured against the RTX 4090 at the identical Q5_K_M quant, Qwen3 8B decodes at 94.9 tok/s there versus 67.5 tok/s here, a 40.6% gap that tracks the two cards' bandwidth difference (1,008 vs 717 GB/s, 40.6%) almost exactly, since decode is bandwidth-bound. Against the RTX 4070 Ti SUPER at the same Q5_K_M quant, the gap is much smaller, 67.5 vs 63.3 tok/s, a 6.6% edge closely matching the two cards' bandwidth difference (6.7%): the whole practical difference between paying $1,199 for this card and $799 for that one.
The NVIDIA RTX 4080 launched November 16, 2022 at a $1,199 MSRP, the second-tier Ada Lovelace consumer GPU beneath the RTX 4090. It pairs 16GB of GDDR6X on a 256-bit bus (roughly 717 GB/s bandwidth) with a cut-down AD103-300 die: 9,728 CUDA cores and 304 4th-gen Tensor Cores from 45.9 billion transistors on a 379mm² die. Fourteen months later, the RTX 4070 Ti SUPER launched at $799, exactly $400 less, on the identical AD103 die with the same 16GB capacity and 256-bit bus, giving up only 1,280 CUDA cores (8,448 vs 9,728) and 45 GB/s of bandwidth (672 vs 717 GB/s) to get there; NVIDIA replaced this card outright in January 2024 with the RTX 4080 SUPER, a near-full AD103 configuration (10,240 CUDA cores) at $999, $200 under this card's original price.
NVIDIA RTX 4080: Reviewers were broadly critical of this card's $1,199 launch price, a real jump from the $699 the RTX 3080 and RTX 2080 each launched at in the two generations before it (widely covered in launch-week press). The more durable story became clear fourteen months later: the RTX 4070 Ti SUPER launched January 24, 2024 at $799, on the identical AD103 die with this card's same 16GB capacity and same 256-bit bus, giving up only 1,280 CUDA cores (8,448 vs 9,728) and 45 GB/s of bandwidth (672 vs 717 GB/s, a 6.7% gap) to get there for $400 less. NVIDIA replaced this card outright the same month with the RTX 4080 SUPER, a near-full AD103 configuration (10,240 CUDA cores) at $999, $200 under this card's original price.
This site's calculator puts Qwen3 8B at its recommended Q5_K_M quant at 67.5 tok/s (7.73 GB, fits with room for context) and Qwen3 14B at its recommended Q5_K_M at 39.2 tok/s (13.31 GB, fits but with only about 1.9 GB of headroom left in the roughly 15.2 GB this card actually has usable). GPT-OSS 20B, a model OpenAI explicitly sized to fit a 16GB card, fits at its recommended Q4_K_M (14.55 GB, 62 tok/s), with barely 0.6 GB of that usable budget left over. Qwen 3.6 27B draws the real ceiling: its recommended Q4_K_M build (19.02 GB) overshoots this card's usable VRAM and drops to a CPU-offloaded 9.6 tok/s, but the next quant down, Q3_K_M (15.15 GB), still fits natively at 34.5 tok/s, decoding within a rounding error of GPT-OSS 20B despite having more total parameters. Qwen3 32B doesn't fit natively at any quantization on this site's standard ladder: even Q3_K_M needs 19.17 GB against roughly 15.2 GB of usable VRAM, landing at 9.1 tok/s offloaded rather than a native fit, so 27B-class is closer to this card's real dense-model ceiling than the 32B some marketing copy implies. GPT-OSS 120B doesn't fit at all, even with full system-RAM offload: its 62.6 GB MXFP4 weight floor alone is well past this card's roughly 41 GB usable VRAM+RAM budget. Measured against the RTX 4090 at the identical Q5_K_M quant, Qwen3 8B decodes at 94.9 tok/s there versus 67.5 tok/s here, a 40.6% gap that tracks the two cards' bandwidth difference (1,008 vs 717 GB/s, 40.6%) almost exactly, since decode is bandwidth-bound. Against the RTX 4070 Ti SUPER at the same Q5_K_M quant, the gap is much smaller, 67.5 vs 63.3 tok/s, a 6.6% edge closely matching the two cards' bandwidth difference (6.7%): the whole practical difference between paying $1,199 for this card and $799 for that one.
Full CUDA support on Ada Lovelace, the same mature llama.cpp, Ollama, vLLM, and TensorRT-LLM coverage every Ada card in this site's lineup gets, well ahead of Blackwell's rockier early sm_120 rollout. This card's 4th-gen Tensor Cores add hardware FP8 support, accelerated natively by TensorRT-LLM and vLLM on Ada and Hopper, but not NVFP4: that 4-bit format is Blackwell-only in this site's calculator, so Q2_K remains this card's smallest quant on the standard ladder. On value, the RTX 4070 Ti SUPER is the sharper comparison than the raw specs suggest: identical 16GB capacity on the identical AD103 die, giving up just 6.7% of this card's bandwidth (672 vs 717 GB/s) for $400 less at launch, and since local decode speed is bandwidth-bound, that 6.7% gap (67.5 vs 63.3 tok/s on Qwen3 8B at Q5_K_M above) is close to the entire practical difference between the two cards for LLM inference specifically. The RTX 4060 Ti 16GB undercuts both further at the same 16GB ceiling, but on a much narrower 288 GB/s bus: this site's calculator puts the same Qwen3 8B build at 27.1 tok/s there, about 40% of this card's speed, tracking the 288-vs-717 GB/s ratio (40.2%) closely. The same models fit either card; they just decode far more slowly on the cheaper one.
| Vendor | NVIDIA |
| Architecture | Ada Lovelace |
| VRAM | 16 GB |
| Memory type | GDDR6X |
| Memory bandwidth | 717 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2022 |
| Models (native) | 45 / 97 |
| Models (offload) | 12 / 97 |
AD104 to AD103 bought 33.3% more bandwidth; staying on AD103 bought 6.7% more
The RTX 4070 Ti, the RTX 4070 Ti SUPER, and this page's RTX 4080 form a real die lineage, not just three similarly-priced cards: the first step swaps dies entirely, the second one doesn't.
The RTX 4070 Ti's 504 GB/s comes off a fully-cut AD104 die at 12GB. Moving to the RTX 4070 Ti SUPER swaps onto the larger AD103 die this page's RTX 4080 also uses, gaining 4GB of VRAM (12GB to 16GB) and 33.3% more bandwidth (672 GB/s). This site's calculator shows that jump directly: Qwen3 8B at its recommended Q5_K_M decodes at 47.5 tok/s on the RTX 4070 Ti and 63.3 tok/s on the RTX 4070 Ti SUPER. The next step, from that Ti SUPER to this card, changes nothing about VRAM (16GB either way) and stays on the same AD103 die, just less cut down: 717 GB/s is only 6.7% more than 672 GB/s, and the same Qwen3 8B build reaches just 67.5 tok/s here, a real but much smaller gain for $400 more at launch ($1,199 vs $799). The die swap bought both capacity and most of the bandwidth gain across all three cards; staying on the fuller die bought a little more speed and nothing else.
Popular models for this GPU
Models this GPU runs natively in VRAM (45)
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q2_K · ~117.2 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/AQ2_K · ~117.2 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q2_K · ~109.8 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q2_K · ~35.3 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q2_K · ~101 t/s
Show 40 more
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q2_K · ~120.7 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/AQ3_K_M · ~34.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q2_K · ~34.7 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q2_K · ~39.4 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q3_K_M · ~34.5 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2Q3_K_M · ~34.5 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~52.9 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/AQ3_K_M · ~34.5 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q3_K_M · ~72 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q3_K_M · ~36.2 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q3_K_M · ~37.1 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q4_K_M · ~62 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q6_K · ~34.5 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q5_K_M · ~38.6 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q6_K · ~36.3 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q6_K · ~41 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q6_K · ~42.1 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q6_K · ~35.6 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q8_0 · ~37 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5Q8_0 · ~47.4 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/AQ8_0 · ~47.4 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3Q8_0 · ~48.7 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0Q8_0 · ~48.7 t/s
- Qwen3 8B8B · MMLU-Pro 56.7Q8_0 · ~48 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3Q8_0 · ~54.5 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0Q8_0 · ~53.1 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~54.8 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~51.7 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~43.1 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~53.7 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~63.5 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~71.7 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~76.7 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~105.9 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~93 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~144.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~169.6 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~200.3 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~423.4 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~441.5 t/s
Models that fit with CPU offload (12)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~1.6 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q3_K_M · ~1.1 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q3_K_M · ~1.1 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q3_K_M · ~1.1 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q5_K_M · ~1.4 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q5_K_M · ~1.2 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q6_K · ~1.5 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q6_K · ~1.6 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q8_0 · ~1.1 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q8_0 · ~1.1 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q8_0 · ~1.1 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q8_0 · ~1.1 t/s
Too large for this GPU (40)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
Compare NVIDIA RTX 4080 with other GPUs
- NVIDIA RTX 4080vsNVIDIA RTX 4090-8 GB VRAM
- NVIDIA RTX 4080vsNVIDIA RTX 5090-16 GB VRAM
- NVIDIA RTX 4080vsNVIDIA RTX 3080 10GB+6 GB VRAM
- NVIDIA RTX 4080vsAMD Radeon RX 6800 XT16 GB each
- NVIDIA RTX 4080vsApple M4 Pro (24GB)-8 GB VRAM
- NVIDIA RTX 4080vsNVIDIA RTX 4070 Ti+4 GB VRAM
- NVIDIA RTX 4080vsNVIDIA RTX 4060 Ti 16GB16 GB each
Continue reading
Frequently asked questions
- How much VRAM does the NVIDIA RTX 4080 have?
- The NVIDIA RTX 4080 has 16 GB of GDDR6X with 717 GB/s memory bandwidth.
- What is the NVIDIA RTX 4080 best for?
- With 16 GB of VRAM, the NVIDIA RTX 4080 handles smaller models (7B–14B) at Q4–Q5 quantization, ideal for entry-level local LLM experimentation and lightweight inference.
- What LLMs can the NVIDIA RTX 4080 run locally?
- The NVIDIA RTX 4080 can run 45 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q3_K_M, Ornith 1.5 35B-A3B (MoE) at Q2_K, Ornith 1.5 9B at Q8_0.
- Can the NVIDIA RTX 4080 run Gemma 4 31B?
- Yes. The NVIDIA RTX 4080 runs Gemma 4 31B natively in VRAM at Q2_K quantization, achieving approximately 35.3 tokens per second.
- Can the NVIDIA RTX 4080 run Qwen 3.6 27B?
- Yes. The NVIDIA RTX 4080 runs Qwen 3.6 27B natively in VRAM at Q3_K_M quantization, achieving approximately 34.5 tokens per second.
- Can the NVIDIA RTX 4080 run Qwen3 8B?
- Yes. The NVIDIA RTX 4080 runs Qwen3 8B natively in VRAM at Q8_0 quantization, achieving approximately 48 tokens per second.