NVIDIA RTX 4070 SUPER
The NVIDIA RTX 4070 SUPER has 12 GB VRAM and 504 GB/s memory bandwidth. It can run 29 of our 87 tracked models natively in VRAM at 8k context.
With 12 GB GDDR6X, the NVIDIA RTX 4070 SUPER is a consumer-tier GPU that can run 29 models natively. It handles 13B-class models comfortably.
The NVIDIA RTX 4070 SUPER launched January 17, 2024 at a $599 MSRP, slotting between the base RTX 4070 and RTX 4070 Ti on the same AD104 die. Its 7,168 CUDA cores are 22% more than the base 4070's 5,888, but VRAM stays at 12GB GDDR6X and bandwidth stays at 504 GB/s, identical to both neighboring cards. Since local LLM decode speed is bandwidth-bound rather than compute-bound, this site's calculator returns the same VRAM fits and the same tokens per second as the base RTX 4070 for every tracked model: the extra cores show up in gaming frame rates and prompt processing, not in decode throughput.
NVIDIA RTX 4070 SUPER: NVIDIA's January 2024 Super refresh gave the base RTX 4070 more compute without touching its memory: this card keeps the exact same 12GB GDDR6X capacity and 192-bit, 504 GB/s memory subsystem as both the RTX 4070 and RTX 4070 Ti, while raising CUDA core count to 7,168 on an AD104-350 die (35.8 billion transistors), 22% more than the base 4070's 5,888 and closer to the 4070 Ti's 7,680 (NVIDIA's own GeForce RTX 40 SUPER announcement; Wikipedia's GeForce RTX 40 series spec table cross-checks every figure here). Launched January 17, 2024 at $599, the same launch MSRP the original RTX 4070 carried. GamersNexus's launch review measured a real 15% gaming uplift over the base RTX 4070 and landed about 10% behind the RTX 4070 Ti.
Because local LLM decode speed on this site's calculator is set entirely by VRAM capacity and memory bandwidth, and this card shares both numbers exactly with the base RTX 4070 and the RTX 4070 Ti (12GB, 504 GB/s), it returns bit-for-bit identical results to both cards for every one of the 86 models this site tracks: Llama 3.1 8B decodes at 34.2 tok/s at Q8_0 (10.73 GB) on all three, and Qwen3 14B fits natively at 38.7 tok/s at Q3_K_M (9.48 GB) on all three too. GPT-OSS 20B, a model OpenAI explicitly sized for a 16GB card, needs CPU offload on all three 12GB cards alike (25.23 GB, 3.6 tok/s). The extra 1,280 CUDA cores this card has over the base RTX 4070 show up in gaming frame rates and prompt-processing throughput, which this site doesn't model, not in the decode tok/s figures above.
Full CUDA support on Ada Lovelace, the same mature llama.cpp, Ollama, vLLM, and TensorRT-LLM coverage every Ada card gets. Since this card, the base RTX 4070, and the RTX 4070 Ti all share the identical 12GB/504 GB/s memory subsystem, none of them can run NVFP4 (Blackwell-only in this site's calculator) and all three cap out at the same Q2_K floor on the standard quant ladder. The practical buying question this card actually answers isn't local LLM throughput, where it ties the cheaper base RTX 4070 exactly: it's whether the extra compute is worth the price gap for gaming or other compute-bound work. For local inference specifically, the 12GB ceiling is the real constraint on all three cards in this family; the RTX 4070 Ti SUPER, which does move the memory subsystem to 16GB and 672 GB/s, is the meaningful upgrade for anyone whose models don't fit in 12GB.
| Vendor | NVIDIA |
| Architecture | Ada Lovelace |
| VRAM | 12 GB |
| Memory type | GDDR6X |
| Memory bandwidth | 504 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2024 |
| Models (native) | 29 / 87 |
| Models (offload) | 24 / 87 |
Three different cards, one number: this site can't tell them apart
The base RTX 4070, this card, and the RTX 4070 Ti differ in price by up to $200 and in CUDA core count by nearly 1,800 cores, but all three share the identical 12GB GDDR6X capacity and 504 GB/s bandwidth. Running one real model at its best-fitting quant on all three shows exactly what that shared memory subsystem means for local LLM decode speed:
Llama 3.1 8B at Q8_0 (10.73 GB) decodes at 34.2 tok/s on all three cards: not close, identical, because this site's decode model is driven entirely by VRAM capacity and memory bandwidth, and all three share the exact same 12GB and 504 GB/s. The RTX 4070 Ti's extra 1,792 CUDA cores over the base RTX 4070, and this card's extra 1,280, change gaming frame rates and compute-bound work like prompt processing, neither of which this site models. For local inference specifically, the only way to actually move these numbers in the RTX 4070 family is the RTX 4070 Ti SUPER, which is also the only one of the four that changed its memory subsystem at all.
Popular models for this GPU
Models this GPU runs natively in VRAM (29)
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~37.2 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q2_K · ~63 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q3_K_M · ~38.7 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q3_K_M · ~37.7 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q4_K_M · ~33.2 t/s
Show 24 more
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q5_K_M · ~32.7 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q5_K_M · ~33.6 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q3_K_M · ~36.4 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q5_K_M · ~35 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5Q8_0 · ~33.3 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3Q8_0 · ~34.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0Q8_0 · ~34.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7Q8_0 · ~33.7 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3Q8_0 · ~38.3 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0Q8_0 · ~37.3 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~38.5 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~36.4 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4Q8_0 · ~45.1 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~37.8 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~44.6 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~50.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~53.9 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~39 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~39 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~52.5 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~62.7 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~75.7 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~156 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~184.5 t/s
Models that fit with CPU offload (24)
These use system RAM for layers that don't fit in VRAM — expect much slower inference.
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~1.3 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~1.3 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~1.3 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~1.3 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q4_K_M · ~1.4 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q4_K_M · ~1.2 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q6_K · ~4.7 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q6_K · ~1.2 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q6_K · ~1.3 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q6_K · ~1.4 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q6_K · ~1.4 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q6_K · ~1.4 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q6_K · ~1.4 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q6_K · ~4.8 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q6_K · ~1.5 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~3.2 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q8_0 · ~3.5 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro —Q8_0 · ~1.3 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~1.2 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~1.3 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~1.3 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q8_0 · ~1.5 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q8_0 · ~1.7 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~3.6 t/s
Too large for this GPU (34)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
Compare NVIDIA RTX 4070 SUPER with other GPUs
Frequently asked questions
- How much VRAM does the NVIDIA RTX 4070 SUPER have?
- The NVIDIA RTX 4070 SUPER has 12 GB of GDDR6X with 504 GB/s memory bandwidth.
- What is the NVIDIA RTX 4070 SUPER best for?
- With 12 GB of VRAM, the NVIDIA RTX 4070 SUPER is best for running compact models (1B–8B) at low quantization, suitable for edge inference, prototyping, and lightweight tasks.
- What LLMs can the NVIDIA RTX 4070 SUPER run locally?
- The NVIDIA RTX 4070 SUPER can run 29 of the 87 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at Q8_0, Llama 3.2 3B Instruct at BF16, Llama 3.2 1B Instruct at FP32.
- Can the NVIDIA RTX 4070 SUPER run Llama 3.3 70B Instruct?
- The NVIDIA RTX 4070 SUPER can run Llama 3.3 70B Instruct with CPU offload at Q2_K quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 4070 SUPER run Qwen 3.6 27B?
- The NVIDIA RTX 4070 SUPER can run Qwen 3.6 27B with CPU offload at Q8_0 quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 4070 SUPER run Llama 3.1 8B Instruct?
- Yes. The NVIDIA RTX 4070 SUPER runs Llama 3.1 8B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 34.2 tokens per second.