NVIDIA RTX 5090

The NVIDIA RTX 5090 has 32 GB VRAM and 1792 GB/s memory bandwidth. It can run 53 of our 97 tracked models natively in VRAM at 8k context.

With 32 GB GDDR7, the NVIDIA RTX 5090 is a consumer-tier GPU that can run 53 models natively. This site's calculator puts Qwen3 32B at its recommended NVFP4 quant at 19.87 GB (65.7 tok/s), the largest dense model that fits this card natively with real headroom at 8k context, and Llama 3.1 8B at Q4_K_M at 195.9 tok/s, 77.8% faster than the same model's 110.2 tok/s on an RTX 4090, tracking the bandwidth gain almost exactly. 70B-class dense models (Llama 3.3 70B, Qwen 2.5 72B, Llama 3.1 70B) don't fit natively at any quantization at 8k context: even Llama 3.3 70B's smallest real build (NVFP4, 42.21 GB) and its more common Q4_K_M build (50.75 GB) both exceed this card's roughly 30 GB of usable VRAM, so it lands in the offload bucket at 1.6-3.1 tok/s rather than running natively. GPT-OSS 120B doesn't fit at all, even with full system-RAM offload: its 62.6 GB MXFP4 weight floor alone is close to this card's entire usable VRAM+RAM budget. Community benchmarks (hardware-corner.net) report comparable native-fit throughput in practice: Qwen3 8B generating at 185.91 tok/s and prefill above 10,000 tok/s at 4K context, though real llama.cpp batching means these aren't directly the same measurement as this site's steady-state estimate.

The NVIDIA RTX 5090 is the flagship Blackwell consumer GPU, launched January 30, 2025 at a $1,999 MSRP. It pairs 32GB of GDDR7 on a 512-bit bus (1,792 GB/s bandwidth, a 78% jump over the RTX 4090's 1,008 GB/s) with a cut-down GB202-300 die: 21,760 CUDA cores and 680 5th-gen Tensor Cores from the same 92.2-billion-transistor GB202 that powers the 96GB RTX Pro 6000 workstation card at its full, uncut 24,576-core configuration. Native NVFP4 support accelerates FP4 inference for newer quantized LLMs that older Ada/Ampere cards can't run at all.

NVIDIA RTX 5090: This is the highest-bandwidth desktop GPU this site tracks (1,792 GB/s, 77.8% more than the RTX 4090's 1,008 GB/s) from GDDR7 replacing GDDR6X on a similar-width bus rather than a wider one (Wikipedia's GeForce RTX 50 series spec table, cross-checked against NVIDIA's own product page). The bigger story eighteen months after launch is the price: an AI-datacenter-driven GDDR7 supply crunch pushed real street prices to $4,300-$4,600 by mid-2026 (videocardprices.com tracking, August 2026), more than double the $1,999 launch MSRP, a genuinely strange irony for an AI-capable consumer card, since the same AI boom that makes it desirable for local inference is also what's diverting the memory supply that would otherwise go into building more of them.

This site's calculator puts Qwen3 32B at its recommended NVFP4 quant at 19.87 GB (65.7 tok/s), the largest dense model that fits this card natively with real headroom at 8k context, and Llama 3.1 8B at Q4_K_M at 195.9 tok/s, 77.8% faster than the same model's 110.2 tok/s on an RTX 4090, tracking the bandwidth gain almost exactly. 70B-class dense models (Llama 3.3 70B, Qwen 2.5 72B, Llama 3.1 70B) don't fit natively at any quantization at 8k context: even Llama 3.3 70B's smallest real build (NVFP4, 42.21 GB) and its more common Q4_K_M build (50.75 GB) both exceed this card's roughly 30 GB of usable VRAM, so it lands in the offload bucket at 1.6-3.1 tok/s rather than running natively. GPT-OSS 120B doesn't fit at all, even with full system-RAM offload: its 62.6 GB MXFP4 weight floor alone is close to this card's entire usable VRAM+RAM budget. Community benchmarks (hardware-corner.net) report comparable native-fit throughput in practice: Qwen3 8B generating at 185.91 tok/s and prefill above 10,000 tok/s at 4K context, though real llama.cpp batching means these aren't directly the same measurement as this site's steady-state estimate.

Full CUDA support, but Blackwell's sm_120 compute capability was new enough at launch that official PyTorch wheels lagged for a long stretch: a GitHub issue asking PyTorch to add stable sm_120 support (pytorch/pytorch#159207) was still open in July 2025, six months after this card shipped, forcing early PyTorch-based tooling to build from source or use nightlies. llama.cpp-based tools (Ollama, LM Studio) mostly avoided that gap since llama.cpp compiles its own CUDA kernels. NVIDIA's own engineering blog reports up to 35% higher throughput from CUDA graph enablement and up to 15% from flash-attention kernels after upgrading to the CUDA 12.8 runtime Blackwell requires (a specific 27% LM Studio speedup was measured on the related RTX 5080, not this card), the software stack has matured substantially since, but a very old llama.cpp, Ollama, or PyTorch build still won't recognize this GPU's architecture. Native NVFP4 support is the other software-visible difference from Ada/Ampere: it's the only quantization format in this site's ladder that requires Blackwell specifically.

VendorNVIDIA
ArchitectureBlackwell
VRAM32 GB
Memory typeGDDR7
Memory bandwidth1792 GB/s
Compute backendCUDA
TierConsumer
Released2025
Models (native)53 / 97
Models (offload)9 / 97
Software: Full llama.cpp and Ollama support out of the box. CUDA 12.x recommended; driver ≥ 525 required.

The ceiling of the desktop Blackwell stack: exactly 4x the entry card on both axes

Plotting every desktop Blackwell GPU this site tracks by VRAM and bandwidth together shows the RTX 5090 alone in the top-right corner, and how unevenly the steps below it climb to reach it:

0950190001734VRAM (GB)Bandwidth (GB/s)NVIDIA RTX 5050NVIDIA RTX 5060NVIDIA RTX 5060 Ti 16GBNVIDIA RTX 5070NVIDIA RTX 5070 TiNVIDIA RTX 5080NVIDIA RTX 5090
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

This card (this page) sits at 32GB and 1,792 GB/s, precisely 4x the RTX 5060's 8GB and 448 GB/s on both axes at once, the only exact 4x match anywhere in the stack. The steps between them aren't even: the RTX 5060 Ti 16GB matches the plain 5060's 448 GB/s exactly (a VRAM-only upgrade), the RTX 5070 Ti's 896 GB/s is exactly double the 5060 Ti's on the same 16GB, and the RTX 5080 barely moves bandwidth past that (960 GB/s, +7.1%) at the same 16GB again. The jump from the RTX 5080 to this card is the biggest single step anywhere in the lineup: VRAM doubles (16GB to 32GB) while bandwidth nearly doubles again on top of that (960 to 1,792 GB/s, +86.7%), which is what actually separates this card from being 'one more Blackwell SKU' and makes it the only one in the desktop stack that runs 32B-class dense models with real headroom.

A flat generation, then the biggest bandwidth jump in three flagships

RTX 3090 to RTX 4090 was a quiet generation for local LLM use; this card's step up from the RTX 4090 was anything but. Tracking VRAM and bandwidth across NVIDIA's last three 24GB+ flagships:

NVIDIA RTX 3090
936 GB/s
NVIDIA RTX 4090
1008 GB/s
NVIDIA RTX 5090 (this page)
1792 GB/s

RTX 3090 to RTX 4090 barely moved either number: VRAM held flat at 24GB, and bandwidth grew just 7.7% (936 to 1,008 GB/s). RTX 4090 to this card flips that pattern hard: bandwidth jumps 77.8% (1,008 to 1,792 GB/s), the largest single-generation bandwidth gain across these three flagships, while VRAM grows a comparatively modest 33.3% (24GB to 32GB). Since decode is bandwidth-bound, this site's calculator measures that chain directly: Llama 3.1 8B at Q4_K_M runs 102.3 tok/s on the RTX 3090, 110.2 tok/s on the RTX 4090 (a 7.7% gain, tracking that generation's bandwidth bump almost exactly), and 195.9 tok/s on this card (a 77.8% gain over the 4090, again matching bandwidth almost exactly). The bigger practical change is what that bandwidth doesn't fix: 70B-class models still need CPU offload on all three cards at this site's 8k-context benchmark; this generation's jump makes what already fit faster, not more of what fits.

Same die, three tiers apart: what 64GB more VRAM buys on identical silicon

The RTX Pro 6000 Blackwell workstation card uses the same GB202 die as this page's RTX 5090: the full, uncut 24,576-CUDA-core configuration paired with 96GB of ECC GDDR7, instead of this card's cut-down 21,760 cores and 32GB. Counting how many of this site's 97 tracked models each GPU actually fits natively in VRAM at 8k context shows how much of that gap is capacity, not compute:

NVIDIA RTX 4090
52 / 97
NVIDIA RTX 5090 (this page)
53 / 97
NVIDIA RTX Pro 6000
68 / 97

One generation of Blackwell improvement over the RTX 4090 (more bandwidth, newer Tensor Cores, native NVFP4) moves the native-fit count by exactly one model (52 to 53 of 97), because VRAM capacity barely grew (24GB to 32GB) relative to what 70B-class models actually need. The RTX Pro 6000's same GB202 die with 96GB instead fits 68 of 97, 15 more than this card, including models this page's RTX 5090 can't run natively at all: Llama 3.3 70B (needs CPU offload here at 50.75 GB, Q4_K_M) and GPT-OSS 120B (doesn't fit at any quantization here, even with full system-RAM offload) both fit the RTX Pro 6000 natively: the 70B model at Q4_K_M and GPT-OSS 120B at NVFP4 (70.46 GB, 99.2 tok/s). Bandwidth decides how fast a model runs once it fits; on this exact silicon, VRAM capacity is what decides whether it fits at all.

Cloud GPU Rental

Don't want to buy a NVIDIA RTX 5090? RunPod is a cloud GPU rental service: rent one by the hour instead, no contract, no upfront hardware cost.

Pay by the hour · no contract · pods start in about a minute.

Rent a NVIDIA RTX 5090 on RunPod ↗ (+$5 signup credit)

Affiliate link: CanItRun may earn a commission. Doesn't affect the fit calculation above.

Popular models for this GPU

Models this GPU runs natively in VRAM (53)

Show 48 more

Models that fit with CPU offload (9)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (35)

Compare NVIDIA RTX 5090 with other GPUs

Frequently asked questions

How much VRAM does the NVIDIA RTX 5090 have?
The NVIDIA RTX 5090 has 32 GB of GDDR7 with 1792 GB/s memory bandwidth.
What is the NVIDIA RTX 5090 best for?
With 32 GB of VRAM, the NVIDIA RTX 5090 is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
What LLMs can the NVIDIA RTX 5090 run locally?
The NVIDIA RTX 5090 can run 53 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at NVFP4, Ornith 1.5 35B-A3B (MoE) at NVFP4, Ornith 1.5 9B at BF16.
Can the NVIDIA RTX 5090 run Gemma 4 31B?
Yes. The NVIDIA RTX 5090 runs Gemma 4 31B natively in VRAM at NVFP4 quantization, achieving approximately 69.1 tokens per second.
Can the NVIDIA RTX 5090 run Qwen 3.6 27B?
Yes. The NVIDIA RTX 5090 runs Qwen 3.6 27B natively in VRAM at NVFP4 quantization, achieving approximately 83 tokens per second.
Can the NVIDIA RTX 5090 run Qwen3 8B?
Yes. The NVIDIA RTX 5090 runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 67.7 tokens per second.
Can I rent the NVIDIA RTX 5090 instead of buying it?
Yes: RunPod and similar cloud GPU providers let you rent NVIDIA RTX 5090 instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.