NVIDIA RTX 5050
The NVIDIA RTX 5050 has 8 GB VRAM and 320 GB/s memory bandwidth. It can run 27 of our 99 tracked models natively in VRAM at 8k context.
With 8 GB GDDR6, the NVIDIA RTX 5050 is a consumer-tier GPU that can run 27 models natively. This card's roughly 7.6 GB of usable VRAM (this site reserves 5% of the spec-sheet 8GB) covers the small end of the ladder with real headroom, not just a bare fit: this site's calculator puts Llama 3.2 3B at 58.3 tok/s at its recommended Q6_K (3.99 GB), Qwen 2.5 3B at 73.1 tok/s at Q6_K (3.19 GB), and Gemma 3 4B at 54.9 tok/s at Q6_K (4.24 GB). The 7-8B class is the real boundary, and it's closer than a blanket "possible at aggressive quant" suggests: Llama 3.1 8B's recommended Q5_K_M (7.58 GB) and Qwen 2.5 7B's recommended Q6_K (7.51 GB) both clear the ~7.6 GB ceiling (Llama 3.1 8B by only about 20 MB), decoding at 30.7 and 31.0 tok/s. Qwen3 8B's recommended Q5_K_M (7.73 GB) tips about 130 MB over that same line into CPU offload, dropping to 30.1 tok/s despite needing barely more memory than the models that do fit natively. Overall, 27 of this site's 99 tracked models fit natively in this card's VRAM at 8k context, exactly as many as the otherwise-identical-capacity RTX 5060, because VRAM capacity, not bandwidth, decides the fits/offload line; the RTX 5060's extra bandwidth just makes those same 26 models faster (Llama 3.2 3B: 58.3 tok/s here vs 81.6 tok/s there, a 40% gap tracking the 448-vs-320 GB/s bandwidth ratio almost exactly). Independent llama.cpp/Ollama benchmarks specific to this card are sparse (it's a budget gaming SKU first, not one reviewers typically bench for LLM workloads), so treat the tok/s figures above as this site's own bandwidth-modeled estimates, not community-verified numbers.
The NVIDIA RTX 5050 is the most affordable Blackwell desktop GPU at $249 MSRP, with 2,560 CUDA cores and 8GB GDDR6 on a 128-bit bus (320 GB/s). Unlike the rest of the 50-series it still uses GDDR6. Suitable only for very small LLMs (3B–4B params) and entry-level 1080p gaming.
NVIDIA RTX 5050: Released July 1, 2025 at a $249 MSRP, the RTX 5050 is the smallest desktop Blackwell die this site tracks: a cut-down GB207-300 chip with 20 SMs, 2,560 CUDA cores, 80 5th-gen Tensor Cores, and 20 RT cores at a 130W TDP (NVIDIA's own spec page and Wikipedia's GeForce RTX 50-series table agree on every figure here). It's the only desktop 50-series card that still ships with GDDR6 instead of GDDR7: 8GB on a 128-bit bus at a 20 Gbps pin rate for 320 GB/s, versus the 28 Gbps GDDR7 every other desktop Blackwell card in this lineup uses. That's a cost decision, not a limitation of the die itself: NVIDIA's laptop GeForce RTX 5050, built on the same 2,560-core silicon, ships with 8GB of GDDR7 instead (VideoCardz's launch coverage). The result: this card ties the RTX 5060 for the smallest VRAM capacity in the desktop stack (8GB), but doesn't tie its bandwidth: 320 GB/s is 28.6% below the RTX 5060's 448 GB/s despite both sharing the exact same 128-bit bus, purely because one runs GDDR6 and the other GDDR7.
This card's roughly 7.6 GB of usable VRAM (this site reserves 5% of the spec-sheet 8GB) covers the small end of the ladder with real headroom, not just a bare fit: this site's calculator puts Llama 3.2 3B at 58.3 tok/s at its recommended Q6_K (3.99 GB), Qwen 2.5 3B at 73.1 tok/s at Q6_K (3.19 GB), and Gemma 3 4B at 54.9 tok/s at Q6_K (4.24 GB). The 7-8B class is the real boundary, and it's closer than a blanket "possible at aggressive quant" suggests: Llama 3.1 8B's recommended Q5_K_M (7.58 GB) and Qwen 2.5 7B's recommended Q6_K (7.51 GB) both clear the ~7.6 GB ceiling (Llama 3.1 8B by only about 20 MB), decoding at 30.7 and 31.0 tok/s. Qwen3 8B's recommended Q5_K_M (7.73 GB) tips about 130 MB over that same line into CPU offload, dropping to 30.1 tok/s despite needing barely more memory than the models that do fit natively. Overall, 27 of this site's 99 tracked models fit natively in this card's VRAM at 8k context, exactly as many as the otherwise-identical-capacity RTX 5060, because VRAM capacity, not bandwidth, decides the fits/offload line; the RTX 5060's extra bandwidth just makes those same 26 models faster (Llama 3.2 3B: 58.3 tok/s here vs 81.6 tok/s there, a 40% gap tracking the 448-vs-320 GB/s bandwidth ratio almost exactly). Independent llama.cpp/Ollama benchmarks specific to this card are sparse (it's a budget gaming SKU first, not one reviewers typically bench for LLM workloads), so treat the tok/s figures above as this site's own bandwidth-modeled estimates, not community-verified numbers.
Full CUDA support and the same sm_120 Blackwell compute capability as every other desktop 50-series card, so it inherits the identical early-driver rough edges: NVIDIA's own engineering blog measured a ~27% LM Studio/llama.cpp speedup on the sibling RTX 5080 just from upgrading to the CUDA 12.8 runtime Blackwell requires, and a very old llama.cpp, Ollama, or PyTorch build may still not recognize this GPU's architecture. The 5th-gen Tensor Cores support NVFP4 in hardware, but at 8GB there's little capacity left to exploit it; the format buys speed on models that already fit, not headroom for larger ones. At $249, this card undercuts the previous-generation RTX 4060's $299 MSRP by trading specs rather than simply beating them: The FPS Review's review (8/10) found the RTX 4060 packs 20% more CUDA cores (3,072 vs 2,560) while this card offers 18% more memory bandwidth (320 vs 272 GB/s), at a 130W TDP versus the 4060's 115W. Since LLM decode is bandwidth-bound rather than compute-bound, that trade favors this card's tok/s over its Ada-generation predecessor at the identical 8GB ceiling; the same review wished for 12GB of VRAM "to really be a game changer" at this price, and that 8GB ceiling is the real ecosystem gotcha for local LLM work: Ollama, LM Studio, and llama.cpp all install and run without issue, but the practical ceiling is a 7B-class model, not a compute or driver limitation.
| Vendor | NVIDIA |
| Architecture | Blackwell |
| VRAM | 8 GB |
| Memory type | GDDR6 |
| Memory bandwidth | 320 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2025 |
| Models (native) | 27 / 99 |
| Models (offload) | 30 / 99 |
The floor of the desktop Blackwell stack, on both axes at once
Plotting every desktop Blackwell GPU this site tracks by VRAM and bandwidth together shows this card anchoring the bottom-left corner; not close to the floor, but sitting on it on both axes simultaneously:
This card (this page) sits at 8GB and 320 GB/s, tied for the smallest VRAM capacity in the stack with the RTX 5060, but alone at the bottom on bandwidth: even the RTX 5060 gets 40% more (320 to 448 GB/s) from the exact same 128-bit bus, just by using 28 Gbps GDDR7 instead of this card's 20 Gbps GDDR6. The RTX 5090 at the opposite end has exactly 4x this card's VRAM (32GB vs 8GB) but 5.6x its bandwidth (1,792 vs 320 GB/s), a wider gap on bandwidth than on capacity, precisely because every other card in this lineup moved to GDDR7 and this one didn't. That combination, smallest capacity and the only card still on the older memory type, is what makes this the practical floor for local LLM use on desktop Blackwell: not a card to size a model around, but the one every other card in the lineup is a step up from.
The one clean GDDR6-vs-GDDR7 comparison in the lineup
The RTX 5050 and RTX 5060 share an identical 8GB capacity and an identical 128-bit bus, the only pair in the desktop Blackwell stack that isolates what one memory generation alone is worth:
320 GB/s vs 448 GB/s: the RTX 5060's 28 Gbps GDDR7 delivers 40% more bandwidth than this card's 20 Gbps GDDR6, on the same bus width and the same 8GB capacity. This site's calculator shows that gap directly in tok/s, not just spec-sheet bandwidth: Llama 3.2 3B at its recommended Q6_K quant decodes at 58.3 tok/s on this card versus 81.6 tok/s on the RTX 5060, a 40% gap that tracks the bandwidth ratio almost exactly, since decode is bandwidth-bound. What doesn't change is which models fit at all: both cards hit the identical fits/offload line, 27 of this site's 99 tracked models natively in VRAM at 8k context, because VRAM capacity, not bandwidth, decides that boundary. The RTX 5060's extra bandwidth makes the same 26 models faster; it doesn't add any new ones.
Popular models for this GPU
Models this GPU runs natively in VRAM (27)
- Bonsai 27B27B · MMLU-Pro ~81.51-bit (Q1_0) · ~48 t/s
- Bonsai 2 27B27B · MMLU-Pro N/ATernary (Q2_0) · ~32.3 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q2_K · ~31.2 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q2_K · ~34.7 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q2_K · ~36.5 t/s
Show 22 more
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q2_K · ~32.9 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5NVFP4 · ~43.6 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ANVFP4 · ~43.6 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3NVFP4 · ~41 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0NVFP4 · ~41 t/s
- Qwen3 8B8B · MMLU-Pro 56.7NVFP4 · ~39.9 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3NVFP4 · ~48.7 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0NVFP4 · ~44.3 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6NVFP4 · ~83.1 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4NVFP4 · ~69.2 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4NVFP4 · ~40.6 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3NVFP4 · ~69.9 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0NVFP4 · ~81.9 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~32 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~34.3 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~47.2 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~41.5 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~64.3 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~75.7 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~89.4 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~189 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~197.1 t/s
Models that fit with CPU offload (30)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~1.1 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~1.1 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~1.1 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7NVFP4 · ~1.5 t/s
- Command-R 35B35B · MMLU-Pro 33.0NVFP4 · ~1.2 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3NVFP4 · ~7.8 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2NVFP4 · ~1.9 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/ANVFP4 · ~7.8 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0NVFP4 · ~2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5NVFP4 · ~2.3 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0NVFP4 · ~2.1 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3NVFP4 · ~2.1 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0NVFP4 · ~2.1 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3NVFP4 · ~7.8 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2NVFP4 · ~2.5 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5NVFP4 · ~7.5 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6NVFP4 · ~8.9 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/ANVFP4 · ~3.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0NVFP4 · ~2.5 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5NVFP4 · ~3 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2NVFP4 · ~3.4 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2NVFP4 · ~3.4 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/ANVFP4 · ~3.4 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6NVFP4 · ~7.7 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8NVFP4 · ~3.8 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2NVFP4 · ~4 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9NVFP4 · ~9.4 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0NVFP4 · ~12.2 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7NVFP4 · ~11 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~1.2 t/s
Too large for this GPU (42)
- Qwen 2.5 72B Instruct
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
- DeepSeek V4.1 Flash 552B
Compare NVIDIA RTX 5050 with other GPUs
Frequently asked questions
- How much VRAM does the NVIDIA RTX 5050 have?
- The NVIDIA RTX 5050 has 8 GB of GDDR6 with 320 GB/s memory bandwidth.
- What is the NVIDIA RTX 5050 best for?
- With 8 GB of VRAM, the NVIDIA RTX 5050 is best for running compact models (1B–8B) at low quantization, suitable for edge inference, prototyping, and lightweight tasks.
- What LLMs can the NVIDIA RTX 5050 run locally?
- The NVIDIA RTX 5050 can run 27 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Ornith 1.5 9B at NVFP4, Qwen 3.5 9B at NVFP4, Bonsai 27B at 1-bit (Q1_0).
- Can the NVIDIA RTX 5050 run Gemma 4 31B?
- The NVIDIA RTX 5050 can run Gemma 4 31B with CPU offload at NVFP4 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 5050 run Qwen 3.6 27B?
- The NVIDIA RTX 5050 can run Qwen 3.6 27B with CPU offload at NVFP4 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 5050 run Qwen3 8B?
- Yes. The NVIDIA RTX 5050 runs Qwen3 8B natively in VRAM at NVFP4 quantization, achieving approximately 39.9 tokens per second.