NVIDIA RTX 5060
The NVIDIA RTX 5060 has 8 GB VRAM and 448 GB/s memory bandwidth. It can run 27 of our 99 tracked models natively in VRAM at 8k context.
With 8 GB GDDR7, the NVIDIA RTX 5060 is a consumer-tier GPU that can run 27 models natively. This site's calculator puts Qwen 2.5 7B at 43.4 tok/s at Q6_K (7.5 GB, fits natively) and Llama 3.1 8B at 49 tok/s at Q4_K_M (6.7 GB); Q6_K doesn't quite fit for the 8B model (8.56 GB, tips into offload), so Q4_K_M is the practical ceiling for 8B-class weights. Qwen2.5 14B doesn't fit natively even at Q3_K_M (9.7 GB total, spills into system RAM on this 8GB card); this is really a 7-8B-class GPU, not a 14B one. The 5th-gen Tensor Cores also run NVIDIA's native NVFP4 format, a 4-bit quantization scheme newer than the legacy GGUF K-quants, in hardware, and it's the fastest fit for both models this card actually handles: Qwen 2.5 7B reaches 68.2 tok/s at NVFP4 (4.78 GB) versus 43.4 tok/s at Q6_K, and Llama 3.1 8B reaches 57.4 tok/s at NVFP4 (5.68 GB) versus 49 tok/s at Q4_K_M. Those two NVFP4 figures are worth comparing against the RTX 5060 Ti 16GB, which shares this card's exact 448 GB/s bandwidth: this site's calculator returns the identical 68.2 tok/s for Qwen 2.5 7B and 57.4 tok/s for Llama 3.1 8B on both cards, bit-for-bit, since decode speed is bandwidth-bound and the two cards' memory subsystems are otherwise the same; the Ti's 16GB buys headroom for larger models, not a faster ceiling for the ones that already fit here.
The NVIDIA RTX 5060 is the entry-level Blackwell GPU with 8GB GDDR7 on a 128-bit bus (448 GB/s) and 3,840 CUDA cores. It is strictly a 1080p gaming card; for LLM inference, only small models like Gemma 4 E4B or Phi-4-mini fit comfortably in VRAM.
NVIDIA RTX 5060: Launched May 19, 2025 on the Blackwell GB206-250 die at a $299 MSRP, though real street prices ran $330-360 during launch week: $330 at Newegg, $360 in Taiwan. NVIDIA didn't seed review samples to press at all this time: GamersNexus bought its own two cards to run a delayed, self-funded review it titled "Forbidden Review," writing it would likely keep buying every future NVIDIA GPU it tests rather than depend on manufacturer samples, and calling the pattern "anti-consumer." The card itself: 3,840 CUDA cores, a 145W TDP, and PCIe 5.0 x8, half the lanes of a full x16 slot, a cut this card shares with the RTX 4060 before it, and one that costs little in practice on a Gen 4 or Gen 5 board.
This site's calculator puts Qwen 2.5 7B at 43.4 tok/s at Q6_K (7.5 GB, fits natively) and Llama 3.1 8B at 49 tok/s at Q4_K_M (6.7 GB); Q6_K doesn't quite fit for the 8B model (8.56 GB, tips into offload), so Q4_K_M is the practical ceiling for 8B-class weights. Qwen2.5 14B doesn't fit natively even at Q3_K_M (9.7 GB total, spills into system RAM on this 8GB card); this is really a 7-8B-class GPU, not a 14B one. The 5th-gen Tensor Cores also run NVIDIA's native NVFP4 format, a 4-bit quantization scheme newer than the legacy GGUF K-quants, in hardware, and it's the fastest fit for both models this card actually handles: Qwen 2.5 7B reaches 68.2 tok/s at NVFP4 (4.78 GB) versus 43.4 tok/s at Q6_K, and Llama 3.1 8B reaches 57.4 tok/s at NVFP4 (5.68 GB) versus 49 tok/s at Q4_K_M. Those two NVFP4 figures are worth comparing against the RTX 5060 Ti 16GB, which shares this card's exact 448 GB/s bandwidth: this site's calculator returns the identical 68.2 tok/s for Qwen 2.5 7B and 57.4 tok/s for Llama 3.1 8B on both cards, bit-for-bit, since decode speed is bandwidth-bound and the two cards' memory subsystems are otherwise the same; the Ti's 16GB buys headroom for larger models, not a faster ceiling for the ones that already fit here.
Compute capability 12.0 (sm_120) was genuinely new hardware at launch: stable PyTorch releases needed months to add sm_120 kernels (a GitHub issue tracking official support, pytorch/pytorch#159207, stayed open for months after this card's May 2025 launch), and NVIDIA's own engineering blog cites a ~27% LM Studio speedup after upgrading to the CUDA 12.8 runtime Blackwell requires. That gap has closed since launch, but a very old llama.cpp, Ollama, or PyTorch build may still not recognize this GPU's architecture. VRAM capacity is the bigger long-run constraint: this is the second straight generation at 8GB (the RTX 4060 was 8GB too), even as bandwidth jumped 64.7%; reviewers were near-unanimous that capacity, not compute, is what will limit this card first. It shares that same 448 GB/s ceiling with the RTX 5060 Ti 16GB one step up in the lineup, so the Ti's $429 MSRP buys double the VRAM on an identical memory subsystem rather than a faster one.
| Vendor | NVIDIA |
| Architecture | Blackwell |
| VRAM | 8 GB |
| Memory type | GDDR7 |
| Memory bandwidth | 448 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2025 |
| Models (native) | 27 / 99 |
| Models (offload) | 30 / 99 |
The floor of the current Blackwell stack
Plotting every desktop Blackwell GPU this site tracks by VRAM and bandwidth together shows where the RTX 5060 actually sits, and how little separates it from the next card up:
The RTX 5060 and RTX 5060 Ti 16GB share the exact same 448 GB/s; the Ti upgrade buys double the VRAM (16GB vs 8GB) and more CUDA cores, not a faster memory subsystem. Above that, VRAM and bandwidth climb together all the way to the RTX 5090's 32GB at 1,792 GB/s, four times this card's bandwidth. Only the RTX 5050 has less bandwidth than the RTX 5060 in the tracked Blackwell desktop lineup, and nothing in it has less VRAM.
Bandwidth jumped 65% generation over generation: VRAM didn't move at all
The RTX 4060 to RTX 5060 upgrade is unusual: one of these two numbers changed a lot, and the other didn't change even a little.
The RTX 4060 and RTX 5060 both ship 8GB: Blackwell's entry tier is the second straight generation stuck at that capacity, a point nearly every launch review raised. Bandwidth is the real generational story: 272 GB/s to 448 GB/s is a 64.7% jump, GDDR6 to faster GDDR7 on the same 128-bit bus. For LLM inference that means a model that fit on the RTX 4060 decodes meaningfully faster on the RTX 5060, but a model that didn't fit still doesn't.
The 8GB→16GB cliff, on the identical 448 GB/s
The RTX 5060 Ti 16GB isn't a faster RTX 5060; it's the same GB206 memory subsystem at the exact same 448 GB/s, just with twice the VRAM. That makes this the cleanest capacity-only comparison on the whole site: 28 of the 96 models this site tracks with a standard quant ladder (29.2%) reach a fits verdict on the Ti's 16GB at their best quant but only reach an offload verdict on this card's 8GB, at that identical quant. Here's a representative slice spanning dense 14B-27B models and 20B-35B MoE models:
None of these ten models fit this card's 8GB at any quant: Phi-4 spills earliest at just 9.34 GB (NVFP4), and Qwen 3.6 27B needs 15.15 GB (Q3_K_M), nearly double this card's usable VRAM. The throughput cost of that spill varies by architecture: dense models fall hardest, from 17.9 tok/s offloaded to 34.9 tok/s fully in VRAM for Phi-4, and from 3.8 to 21.5 tok/s for Qwen 3.6 27B, while MoE models, which only read their active experts, still take a real hit but from a higher floor, Qwen 3.5 35B-A3B going from 12.9 tok/s offloaded to 73.2 tok/s fully in VRAM. What makes this pair unusual is what doesn't change: for any model that already fits both cards' VRAM, decode speed is identical, not just close; this site's calculator returns the same 68.2 tok/s for Qwen 2.5 7B at NVFP4 and the same 57.4 tok/s for Llama 3.1 8B at NVFP4 on both cards, because bandwidth, not capacity, sets decode speed, and these two cards share the exact same bandwidth. The Ti's extra $130 (per NVIDIA's $299 vs $429 MSRPs) buys a wider set of models that run at all, never a faster version of the ones that already do.
Popular models for this GPU
Models this GPU runs natively in VRAM (27)
- Bonsai 27B27B · MMLU-Pro ~81.51-bit (Q1_0) · ~67.1 t/s
- Bonsai 2 27B27B · MMLU-Pro N/ATernary (Q2_0) · ~45.2 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q2_K · ~43.6 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q2_K · ~48.6 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q2_K · ~51 t/s
Show 22 more
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q2_K · ~46 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5NVFP4 · ~61.1 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ANVFP4 · ~61.1 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3NVFP4 · ~57.4 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0NVFP4 · ~57.4 t/s
- Qwen3 8B8B · MMLU-Pro 56.7NVFP4 · ~55.9 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3NVFP4 · ~68.2 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0NVFP4 · ~62 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6NVFP4 · ~116.3 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4NVFP4 · ~96.9 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4NVFP4 · ~56.9 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3NVFP4 · ~97.9 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0NVFP4 · ~114.7 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~44.8 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~48 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~66.1 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~58.1 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~90 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~106 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~125.1 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~264.6 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~275.9 t/s
Models that fit with CPU offload (30)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~1.1 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~1.1 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~1.1 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7NVFP4 · ~1.5 t/s
- Command-R 35B35B · MMLU-Pro 33.0NVFP4 · ~1.2 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3NVFP4 · ~8 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2NVFP4 · ~2 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/ANVFP4 · ~8 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0NVFP4 · ~2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5NVFP4 · ~2.3 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0NVFP4 · ~2.2 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3NVFP4 · ~2.2 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0NVFP4 · ~2.2 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3NVFP4 · ~8 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2NVFP4 · ~2.5 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5NVFP4 · ~7.7 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6NVFP4 · ~9.2 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/ANVFP4 · ~3.5 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0NVFP4 · ~2.6 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5NVFP4 · ~3.1 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2NVFP4 · ~3.5 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2NVFP4 · ~3.5 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/ANVFP4 · ~3.5 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6NVFP4 · ~8 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8NVFP4 · ~3.9 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2NVFP4 · ~4.1 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9NVFP4 · ~9.9 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0NVFP4 · ~13.9 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7NVFP4 · ~12.4 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~1.2 t/s
Too large for this GPU (42)
- Qwen 2.5 72B Instruct
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
- DeepSeek V4.1 Flash 552B
Continue reading
Frequently asked questions
- How much VRAM does the NVIDIA RTX 5060 have?
- The NVIDIA RTX 5060 has 8 GB of GDDR7 with 448 GB/s memory bandwidth.
- What is the NVIDIA RTX 5060 best for?
- With 8 GB of VRAM, the NVIDIA RTX 5060 is best for running compact models (1B–8B) at low quantization, suitable for edge inference, prototyping, and lightweight tasks.
- What LLMs can the NVIDIA RTX 5060 run locally?
- The NVIDIA RTX 5060 can run 27 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Ornith 1.5 9B at NVFP4, Qwen 3.5 9B at NVFP4, Bonsai 27B at 1-bit (Q1_0).
- Can the NVIDIA RTX 5060 run Gemma 4 31B?
- The NVIDIA RTX 5060 can run Gemma 4 31B with CPU offload at NVFP4 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 5060 run Qwen 3.6 27B?
- The NVIDIA RTX 5060 can run Qwen 3.6 27B with CPU offload at NVFP4 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 5060 run Qwen3 8B?
- Yes. The NVIDIA RTX 5060 runs Qwen3 8B natively in VRAM at NVFP4 quantization, achieving approximately 55.9 tokens per second.