NVIDIA RTX 5070
The NVIDIA RTX 5070 has 12 GB VRAM and 672 GB/s memory bandwidth. It can run 31 of our 99 tracked models natively in VRAM at 8k context.
With 12 GB GDDR7, the NVIDIA RTX 5070 is a consumer-tier GPU that can run 31 models natively. This site's calculator puts Qwen 2.5 7B at 65.1 tok/s at Q6_K (7.51 GB, fits natively) and Llama 3.1 8B at 73.5 tok/s at Q4_K_M (6.66 GB), both comfortable inside the 12GB ceiling with room for an 8K context window. Qwen2.5 14B is the boundary case this card's 12GB actually draws: Q3_K_M fits at 9.72 GB (50.3 tok/s), but the next step up, Q4_K_M, needs 11.83 GB and tips into an offload verdict, dropping to 41.4 tok/s despite needing well under a gigabyte more. The 5th-gen Tensor Cores also run NVIDIA's native NVFP4 format in hardware; this site's calculator estimates 102.3 tok/s for Qwen 2.5 7B at NVFP4 (4.78 GB), and NVFP4 is actually the one quant that lets Qwen2.5 14B fit natively at 48.7 tok/s (10.04 GB), beating Q4_K_M's offload result. Treat those NVFP4 numbers as a preview rather than the default plan: llama.cpp's Blackwell-native NVFP4 kernels only started landing in April 2026, and a llama.cpp GitHub discussion that May was still hashing out real correctness edge cases in the dequantization path. GGUF's Q3-Q6 quants remain the better-tested choice for most llama.cpp- or Ollama-based setups today.
The NVIDIA RTX 5070 launched March 5, 2025 on the GB205 die, with 12GB of GDDR7 on a 192-bit bus at 672 GB/s and 6,144 CUDA cores. It handles 7B-8B models comfortably at Q4-Q6 with room for context; 14B-class models sit right at the 12GB edge, fitting at some quant levels and tipping into CPU offload at others. That 12GB capacity is a real step down from the RTX 5070 Ti and RTX 5080 immediately above it, both of which jump to 16GB.
NVIDIA RTX 5070: Announced at CES on January 6, 2025 and launched March 5, 2025 at a $549 MSRP, the first GeForce card built on the GB205 die. NVIDIA's own spec page lists 6,144 CUDA cores, 192 5th-gen Tensor Cores, 12GB of GDDR7 on a 192-bit bus, and a 250W TGP (650W minimum PSU). The 28 Gbps GDDR7 modules push 672 GB/s of bandwidth, a real 33.3% jump over the RTX 4070 and RTX 4070 Ti's shared 504 GB/s, on the exact same 192-bit bus and identical 12GB capacity. NVIDIA's CES pitch was "RTX 4090 performance for $549" (Jensen Huang's keynote framing, leaning on DLSS 4 frame generation); independent 4K benchmarks without frame generation put it well behind the 4090.
This site's calculator puts Qwen 2.5 7B at 65.1 tok/s at Q6_K (7.51 GB, fits natively) and Llama 3.1 8B at 73.5 tok/s at Q4_K_M (6.66 GB), both comfortable inside the 12GB ceiling with room for an 8K context window. Qwen2.5 14B is the boundary case this card's 12GB actually draws: Q3_K_M fits at 9.72 GB (50.3 tok/s), but the next step up, Q4_K_M, needs 11.83 GB and tips into an offload verdict, dropping to 41.4 tok/s despite needing well under a gigabyte more. The 5th-gen Tensor Cores also run NVIDIA's native NVFP4 format in hardware; this site's calculator estimates 102.3 tok/s for Qwen 2.5 7B at NVFP4 (4.78 GB), and NVFP4 is actually the one quant that lets Qwen2.5 14B fit natively at 48.7 tok/s (10.04 GB), beating Q4_K_M's offload result. Treat those NVFP4 numbers as a preview rather than the default plan: llama.cpp's Blackwell-native NVFP4 kernels only started landing in April 2026, and a llama.cpp GitHub discussion that May was still hashing out real correctness edge cases in the dequantization path. GGUF's Q3-Q6 quants remain the better-tested choice for most llama.cpp- or Ollama-based setups today.
Like every desktop Blackwell card, this GPU reports compute capability 12.0 (sm_120), genuinely new hardware at launch, needing the CUDA 12.8 runtime before frameworks recognized it. NVIDIA's own engineering blog measured a ~27% LM Studio speedup after that runtime upgrade (benchmarked on the sibling RTX 5080 running DeepSeek-R1-Distill-Llama-8B, but the same sm_120 target and CUDA 12.8 requirement apply to this card too). GamersNexus's launch review, titled "NVIDIA is Selling Lies," was blunt about the tradeoff behind that CES "4090 performance" pitch: 12GB next to the RTX 4090's 24GB, calling it "getting absolutely clobbered for VRAM." That capacity, not compute, is what limits this card first for local LLM work; it's also why the RTX 5070 Ti and RTX 5080 sitting one step above it both jump straight to 16GB rather than a marginal bump.
| Vendor | NVIDIA |
| Architecture | Blackwell |
| VRAM | 12 GB |
| Memory type | GDDR7 |
| Memory bandwidth | 672 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2025 |
| Models (native) | 31 / 99 |
| Models (offload) | 27 / 99 |
The only step where moving up the stack costs you VRAM
Plotting every desktop Blackwell GPU this site tracks by VRAM and bandwidth together shows a real notch right at this card; the one point in the lineup where the tier below it actually holds more memory:
VRAM across the tracked Blackwell desktop stack runs 8, 8, 16, 12, 16, 16, 32 GB (RTX 5050 through RTX 5090), climbing everywhere except right here: the RTX 5060 Ti 16GB, one tier below this card, already ships 16GB, and both cards immediately above it, the RTX 5070 Ti and RTX 5080, jump straight back to that same 16GB (same totals either way; both are 16GB Blackwell cards). This card's own step is 12GB, the only dip in an otherwise-ascending capacity ladder. Bandwidth doesn't have the same anomaly: 320, 448, 448, 672, 896, 960, 1,792 GB/s climbs at every step but one tie (the RTX 5060 and RTX 5060 Ti 16GB share 448 GB/s); this card's 672 GB/s is a real, unshared step up, but bandwidth can't substitute for the capacity the tiers on both sides of it already have.
GDDR6X to GDDR7 is the real jump: 12GB doesn't move at all
The RTX 4070 and RTX 4070 Ti shared the exact same 504 GB/s bandwidth on Ada Lovelace, despite the Ti costing more. Tracking bandwidth from those two Ada-generation cards through to the RTX 5070:
The RTX 4070 and RTX 4070 Ti, this card's two direct Ada-generation predecessors, both shipped at 504 GB/s; Ada's 12GB tier never split by bandwidth, only by core count and clocks. The RTX 5070 (this page) breaks that plateau with GDDR7: 672 GB/s is a 33.3% jump over both, even though VRAM capacity holds flat at 12GB across all three generations of card. Since decode is bandwidth-bound, a model that already fit a 4070 or 4070 Ti decodes about a third faster on this card at the identical 12GB ceiling, but a model that didn't fit either Ada card still doesn't fit this one.
The 12GB→16GB cliff: what the extra 4GB actually buys
18 of the 99 models this site tracks land in a specific gap: each one reaches a fits verdict at its best quant on the RTX 5070 Ti's 16GB, but the identical build only reaches an offload verdict on this card's 12GB. That's a deliberately fixed comparison, the same exact quant on both cards, to isolate what the extra 4GB alone changes, so a couple of these tok/s figures read lower than the "CPU offload" list further down this page, which picks each card's own best-available quant independently and can land on a different one (often NVFP4). Here's how far each one spills past the line, and what that offload costs in tokens/sec:
That's 17.5% of the models this site tracks (17 of 99). All seventeen spill into a 12.12-15.19 GB band, comfortably under the 16GB ceiling, but every one of them tips this card's 12GB into an offload verdict, where weights and KV cache partially spill into system RAM. The throughput cost is steep: Qwen 3.5 35B-A3B goes from 146.4 tok/s fully in VRAM on the 16GB tier to 30.8 tok/s offloaded here, Nemotron 3 Nano 30B from 137.2 to 39.5, and Qwen3 30B-A3B from 126.2 to 44.5; all three are mixture-of-experts models, whose usual advantage (reading only the active parameters) gets undercut once part of that read has to cross into system RAM. Dense models take the same hit at smaller absolute numbers: Qwen 3.6 27B drops from 43.1 to 9.0 tok/s. None of these seventeen models fit natively on this card at any quant; the extra 4GB the 16GB tier adds isn't a speed upgrade so much as the difference between a model running at all and one that doesn't.
Popular models for this GPU
Models this GPU runs natively in VRAM (31)
- Bonsai 27B27B · MMLU-Pro ~81.5Ternary (Q2_0) · ~56.5 t/s
- Bonsai 2 27B27B · MMLU-Pro N/ATernary (Q2_0) · ~67.9 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q2_K · ~83.9 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0NVFP4 · ~50 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7NVFP4 · ~48.7 t/s
Show 26 more
- Phi-4 14B Instruct14B · MMLU-Pro 70.4NVFP4 · ~52.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6NVFP4 · ~58.7 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6NVFP4 · ~61 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2NVFP4 · ~47.4 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0NVFP4 · ~58.9 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5NVFP4 · ~91.6 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ANVFP4 · ~91.6 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3NVFP4 · ~86.1 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0NVFP4 · ~86.1 t/s
- Qwen3 8B8B · MMLU-Pro 56.7NVFP4 · ~83.9 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3NVFP4 · ~102.3 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0NVFP4 · ~93 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~51.4 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~48.5 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4NVFP4 · ~85.3 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~50.4 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~59.5 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~67.2 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~71.9 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~99.2 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~87.2 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~135 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~158.9 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~187.7 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~396.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~413.8 t/s
Models that fit with CPU offload (27)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~1.3 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~1.3 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~1.3 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~1.3 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7NVFP4 · ~2 t/s
- Command-R 35B35B · MMLU-Pro 33.0NVFP4 · ~1.4 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3NVFP4 · ~12.4 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2NVFP4 · ~2.8 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/ANVFP4 · ~12.4 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0NVFP4 · ~3 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5NVFP4 · ~3.6 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0NVFP4 · ~3.3 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3NVFP4 · ~3.3 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0NVFP4 · ~3.3 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3NVFP4 · ~13.3 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2NVFP4 · ~4.1 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5NVFP4 · ~13.5 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6NVFP4 · ~17.3 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/ANVFP4 · ~7.5 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0NVFP4 · ~4.2 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5NVFP4 · ~5.8 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2NVFP4 · ~7.6 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2NVFP4 · ~7.6 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/ANVFP4 · ~7.6 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8NVFP4 · ~9.6 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2NVFP4 · ~11.2 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9NVFP4 · ~49.3 t/s
Too large for this GPU (41)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
- DeepSeek V4.1 Flash 552B
Continue reading
Frequently asked questions
- How much VRAM does the NVIDIA RTX 5070 have?
- The NVIDIA RTX 5070 has 12 GB of GDDR7 with 672 GB/s memory bandwidth.
- What is the NVIDIA RTX 5070 best for?
- With 12 GB of VRAM, the NVIDIA RTX 5070 comfortably handles 7B–8B models and can stretch to some 13B–14B models at aggressive quantization, a solid entry point into local LLM inference.
- What LLMs can the NVIDIA RTX 5070 run locally?
- The NVIDIA RTX 5070 can run 31 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Ornith 1.5 9B at NVFP4, Gemma 4 26B (MoE) at Q2_K, Qwen 3.5 9B at NVFP4.
- Can the NVIDIA RTX 5070 run Gemma 4 31B?
- The NVIDIA RTX 5070 can run Gemma 4 31B with CPU offload at NVFP4 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 5070 run Qwen 3.6 27B?
- The NVIDIA RTX 5070 can run Qwen 3.6 27B with CPU offload at NVFP4 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 5070 run Qwen3 8B?
- Yes. The NVIDIA RTX 5070 runs Qwen3 8B natively in VRAM at NVFP4 quantization, achieving approximately 83.9 tokens per second.