NVIDIA RTX 5080

The NVIDIA RTX 5080 has 16 GB VRAM and 960 GB/s memory bandwidth. It can run 46 of our 99 tracked models natively in VRAM at 8k context.

With 16 GB GDDR7, the NVIDIA RTX 5080 is a consumer-tier GPU that can run 46 models natively. This site's calculator puts Qwen3 14B at 60.3 tok/s at Q4_K_M (11.6 GB, fits natively) and Llama 3.1 8B at 104.9 tok/s at Q4_K_M (6.66 GB), both comfortably inside the 16GB budget with headroom for context. The 16GB ceiling isn't a hard wall at 14B, either: dense models up to 22-31B (Mistral Small 22B, Qwen 3.6 27B, Gemma 3 27B, Gemma 4 31B) and MoE models up to 35B total parameters (Qwen3.5 35B-A3B) still fit natively at more aggressive Q2_K-Q3_K_M quantization, decoding anywhere from 46 to 157 tok/s depending on how many parameters actually activate per token. Above that range, CPU offload becomes the only option.

The NVIDIA RTX 5080 is the second-tier Blackwell GPU, built on the full GB203 die with 16GB GDDR7 on a 256-bit bus at 960 GB/s. Its 10,752 CUDA cores and 336 5th-gen Tensor Cores make it a strong 1440p–4K gaming card. For local LLM inference, 7B–14B dense models fit comfortably at Q4–Q8 with room for context; dense models up to roughly 27B and MoE models up to roughly 35B total parameters still fit natively at more aggressive Q2–Q3 quantization, with CPU offload only needed above that.

NVIDIA RTX 5080: Announced at CES on January 6, 2025 and released January 30, 2025 alongside the flagship RTX 5090, the RTX 5080 uses the full GB203-400-A1 die (84 SMs, 10,752 CUDA cores, 336 5th-gen Tensor Cores) with 16GB of GDDR7 on a 256-bit bus, at $999 MSRP for the Founders Edition (confirmed on NVIDIA's own newsroom announcement). Total board power is 360W. GamersNexus's launch review measured the generational gaming uplift over the RTX 4080 Super at just 7-20% depending on the game, called the value proposition "hard to get excited" about when the FPS gap over the same-price 4080 Super sometimes came to just 10 frames, and flagged the repeated 16GB VRAM ceiling, unchanged for a second straight generation, as a real limitation for creative workloads like Premiere and Blender rendering, not just a gaming nitpick.

This site's calculator puts Qwen3 14B at 60.3 tok/s at Q4_K_M (11.6 GB, fits natively) and Llama 3.1 8B at 104.9 tok/s at Q4_K_M (6.66 GB), both comfortably inside the 16GB budget with headroom for context. The 16GB ceiling isn't a hard wall at 14B, either: dense models up to 22-31B (Mistral Small 22B, Qwen 3.6 27B, Gemma 3 27B, Gemma 4 31B) and MoE models up to 35B total parameters (Qwen3.5 35B-A3B) still fit natively at more aggressive Q2_K-Q3_K_M quantization, decoding anywhere from 46 to 157 tok/s depending on how many parameters actually activate per token. Above that range, CPU offload becomes the only option.

Full CUDA support out of the box, but Blackwell's compute capability 12.0 (sm_120) was genuinely new hardware at launch and needs CUDA 12.8+ and an open-kernel driver from the 570 branch; an older llama.cpp, Ollama, or PyTorch build may not recognize this GPU's architecture at all. NVIDIA's own engineering blog measured that gap closing directly on this card: upgrading LM Studio to its CUDA-12.8 llama.cpp backend produced a real ~27% speedup decoding DeepSeek-R1-Distill-Llama-8B at Q4_K_M, benchmarked on an RTX 5080. The bigger long-run constraint is capacity, not compute: this card and the RTX 5070 Ti share the exact same 16GB ceiling and fit exactly the same models, so it's VRAM, not raw compute and not the price gap between them, that actually decides what runs.

VendorNVIDIA
ArchitectureBlackwell
VRAM16 GB
Memory typeGDDR7
Memory bandwidth960 GB/s
Compute backendCUDA
TierConsumer
Released2025
Models (native)46 / 99
Models (offload)12 / 99
Software: Full llama.cpp and Ollama support out of the box. CUDA 12.x recommended; driver ≥ 525 required.

The 5080 shares its VRAM ceiling with the card below it, and trails the one above it by a lot

Plotting every desktop Blackwell GPU this site tracks by VRAM and bandwidth together shows exactly where the RTX 5080 sits: pinned to the same 16GB as the cheaper card just below it, and a full tier behind the 5090 above it on both axes:

0950190001734VRAM (GB)Bandwidth (GB/s)NVIDIA RTX 5050NVIDIA RTX 5060NVIDIA RTX 5060 Ti 16GBNVIDIA RTX 5070NVIDIA RTX 5070 TiNVIDIA RTX 5080NVIDIA RTX 5090
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

The RTX 5080 (this page) and RTX 5070 Ti share the exact same 16GB VRAM ceiling; the only real difference on this chart is bandwidth, where the 5080's 960 GB/s beats the 5070 Ti's 896 GB/s by 7.1%. That's a far smaller gap than the one above it: the RTX 5090 has 2x this card's VRAM (32GB vs 16GB) and 86.7% more bandwidth (1,792 vs 960 GB/s), putting the 5080 much closer, on both axes, to the 5070 Ti below it than to the 5090 above it. A third card, the RTX 5060 Ti 16GB, shares this same 16GB ceiling too, but at less than half the bandwidth (448 GB/s), proof that VRAM capacity alone says nothing about how fast a model that fits will actually decode.

960 GB/s is a real generational jump: the 16GB ceiling isn't

The RTX 5080 and its direct predecessor, the RTX 4080, ship the exact same 16GB capacity on the same 256-bit bus; only the memory technology changed:

NVIDIA RTX 4080
717 GB/s
NVIDIA RTX 5080 (this page)
960 GB/s

GDDR6X to GDDR7 is a real jump: 717 GB/s to 960 GB/s, up 33.9%, on the identical 256-bit bus and identical 16GB capacity. Since decode is bandwidth-bound, that shows up almost exactly in this site's calculator at a fixed quant: Qwen3 14B at Q4_K_M (11.6 GB) decodes at 45.0 tok/s on the RTX 4080 and 60.3 tok/s on the RTX 5080 (this page), a 34.0% gain, within a rounding error of the bandwidth increase itself. VRAM capacity is the one number that didn't move: buying the two-generations-newer card gets a faster memory subsystem and Blackwell's new NVFP4 quant format (which the RTX 4080's Ada Lovelace architecture can't run at all), but the same 16GB ceiling reviewers already called disappointing at this card's own launch.

Same 16GB ceiling, same models: a real but modest speed edge

The RTX 5080 and RTX 5070 Ti fit exactly the same models, since both cap out at the same 16GB. The real question is how much faster the 5080 actually decodes for it; this site's calculator answers that directly, on a 14B-class model both cards fit at the same widely-used Q4_K_M quant:

NVIDIA RTX 5070 Ti
56.2 tok/s
NVIDIA RTX 5080 (this page)
60.3 tok/s

At Q4_K_M, Qwen3 14B decodes at 56.2 tok/s on the RTX 5070 Ti and 60.3 tok/s on the RTX 5080 (this page), a 7.3% gain, tracking the two cards' 7.1% bandwidth difference (896 vs 960 GB/s) almost exactly. That's the whole story: since both cards hit the identical 16GB ceiling, nothing about VRAM capacity changes between them, and the roughly $250 MSRP gap between the two ($999 vs $749 at launch, per NVIDIA's own pricing, not a figure this site's calculator tracks) buys a real but modest ~7% speedup, not access to any model the cheaper card can't already run.

Popular models for this GPU

Models this GPU runs natively in VRAM (46)

Show 41 more

Models that fit with CPU offload (12)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (41)

Frequently asked questions

How much VRAM does the NVIDIA RTX 5080 have?
The NVIDIA RTX 5080 has 16 GB of GDDR7 with 960 GB/s memory bandwidth.
What is the NVIDIA RTX 5080 best for?
With 16 GB of VRAM, the NVIDIA RTX 5080 handles smaller models (7B–14B) at Q4–Q5 quantization, ideal for entry-level local LLM experimentation and lightweight inference.
What LLMs can the NVIDIA RTX 5080 run locally?
The NVIDIA RTX 5080 can run 46 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q3_K_M, Ornith 1.5 35B-A3B (MoE) at Q2_K, Ornith 1.5 9B at NVFP4.
Can the NVIDIA RTX 5080 run Gemma 4 31B?
Yes. The NVIDIA RTX 5080 runs Gemma 4 31B natively in VRAM at Q2_K quantization, achieving approximately 47.2 tokens per second.
Can the NVIDIA RTX 5080 run Qwen 3.6 27B?
Yes. The NVIDIA RTX 5080 runs Qwen 3.6 27B natively in VRAM at Q3_K_M quantization, achieving approximately 46.1 tokens per second.
Can the NVIDIA RTX 5080 run Qwen3 8B?
Yes. The NVIDIA RTX 5080 runs Qwen3 8B natively in VRAM at NVFP4 quantization, achieving approximately 119.8 tokens per second.