NVIDIA RTX 4070 Ti SUPER
The NVIDIA RTX 4070 Ti SUPER has 16 GB VRAM and 672 GB/s memory bandwidth. It can run 41 of our 87 tracked models natively in VRAM at 8k context.
With 16 GB GDDR6X, the NVIDIA RTX 4070 Ti SUPER is a consumer-tier GPU that can run 41 models natively. It handles smaller models (7B–14B) at Q4–Q5 quantization.
The NVIDIA RTX 4070 Ti SUPER launched January 24, 2024 at a $799 MSRP, moving off the RTX 4070 Ti's AD104 die onto the larger AD103 die shared with the RTX 4080. That switch grows VRAM from 12GB to 16GB and widens the memory bus from 192-bit to 256-bit (672 GB/s, up from 504 GB/s), while CUDA cores climb from 7,680 to 8,448. It is the only card in the RTX 4070 family whose memory subsystem actually changed: the RTX 4070 SUPER kept the same 192-bit, 504 GB/s memory as the RTX 4070 and RTX 4070 Ti on either side of it.
NVIDIA RTX 4070 Ti SUPER: NVIDIA's January 2024 Super refresh gave the RTX 4070 Ti a real memory upgrade, not just a clock bump: this card moves off the RTX 4070 Ti's AD104 die onto the larger AD103-275 die (45.9 billion transistors, up from AD104's 35.8 billion) that also powers the RTX 4080, gaining 4GB of VRAM and a wider 256-bit memory bus in the process (NVIDIA's own GeForce RTX 40 SUPER announcement; Wikipedia's GeForce RTX 40 series spec table cross-checks every figure here). Launched January 24, 2024 at $799, it landed $200 under the RTX 4080's $999 launch price for the exact same 16GB capacity, trading only 45 GB/s of bandwidth (672 vs 717 GB/s) and 1,280 fewer CUDA cores (8,448 vs 9,728) to get there. GamersNexus's launch review measured a real but modest 10-11% gaming uplift over the plain RTX 4070 Ti at 4K, a gap that narrows further at lower resolutions.
This site's calculator puts Llama 3.1 8B at 45.6 tok/s at Q8_0 (10.73 GB, fits with room to spare) and GPT-OSS 20B, a model OpenAI explicitly sized for a 16GB card, at 58.1 tok/s at Q4_K_M (14.55 GB). Both figures land within 10% of the RTX 4080's numbers for the same models (48.7 and 62 tok/s), tracking the two cards' 717-vs-672 GB/s bandwidth gap almost exactly, since decode is bandwidth-bound. The jump over this card's own predecessor is much larger: the plain RTX 4070 Ti needs CPU offload for GPT-OSS 20B (25.23 GB spills past its 12GB ceiling, 3.6 tok/s) and for Qwen 3.6 27B (32.75 GB, 1.3 tok/s), while this card runs both natively, GPT-OSS 20B at 58.1 tok/s and Qwen 3.6 27B at 32.3 tok/s (Q3_K_M, 15.15 GB). Measured against the Blackwell RTX 5070 Ti at the identical Q3_K_M quant, Qwen 3.6 27B decodes at 43.1 tok/s there versus 32.3 tok/s here, a 33.4% gap that matches the two cards' 896-vs-672 GB/s bandwidth difference (33.3%) almost exactly, since both fit the same models at the same quant and the only real difference is bandwidth.
Full CUDA support on Ada Lovelace, a mature target for llama.cpp, Ollama, vLLM, and TensorRT-LLM alike, unlike Blackwell's rockier early sm_120 rollout. This card's 4th-gen Tensor Cores add hardware FP8 support, accelerated natively by TensorRT-LLM and vLLM on Ada and Hopper, but not NVFP4: that format is Blackwell-only in this site's calculator, so Q2_K remains the smallest quant on this card's standard ladder, unlike the RTX 5070 Ti at the same 16GB. On value, this is the only card in the RTX 4070 family whose bandwidth actually improved over the base RTX 4070 and RTX 4070 Ti; the RTX 4070 SUPER released the same week kept the identical 504 GB/s memory subsystem as both, gaining only compute. For anyone comparing this card against its direct Blackwell successor, the RTX 5070 Ti, the honest baseline is this card's 672 GB/s, not the original RTX 4070 Ti's 504 GB/s: measured from here, the 5070 Ti's bandwidth gain is a smaller 33.3%, not the 77.8% a comparison against the original RTX 4070 Ti would suggest.
| Vendor | NVIDIA |
| Architecture | Ada Lovelace |
| VRAM | 16 GB |
| Memory type | GDDR6X |
| Memory bandwidth | 672 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2024 |
| Models (native) | 41 / 87 |
| Models (offload) | 12 / 87 |
The one 4070-family bandwidth jump, bookended by two flat ones
This card is the only Ada-generation 4070-family GPU whose memory subsystem actually changed. Tracking bandwidth from the original RTX 4070 Ti through this card to its Blackwell successor:
The RTX 4070 Ti to this card is a real jump: 504 to 672 GB/s, up 33.3%, alongside 4GB of extra VRAM (12GB to 16GB) from moving off the AD104 die onto AD103. That capacity jump matters as much as the bandwidth: Qwen 3.6 27B doesn't fit the RTX 4070 Ti's 12GB at any quantization and needs CPU offload (1.3 tok/s), while it fits this card natively at Q3_K_M (15.15 GB, 32.3 tok/s). The RTX 4070 SUPER released the same week took none of that: it kept the identical 504 GB/s and 12GB as the RTX 4070 Ti, gaining only CUDA cores. The step from here to the RTX 5070 Ti is a second bandwidth jump of almost the same size, 672 to 896 GB/s, up 33.3%, this time from GDDR6X to GDDR7 rather than a wider bus, with VRAM capacity unchanged at 16GB. Since decode is bandwidth-bound, this site's calculator measures that second jump almost exactly on a model both cards fit at the identical Q3_K_M quant: Qwen 3.6 27B decodes at 32.3 tok/s here versus 43.1 tok/s on the RTX 5070 Ti, a 33.4% gap matching the bandwidth ratio within a rounding error.
Popular models for this GPU
Models this GPU runs natively in VRAM (41)
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q2_K · ~109.8 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q2_K · ~102.9 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q2_K · ~33.1 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q2_K · ~94.6 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q2_K · ~113.2 t/s
Show 36 more
- Muse Glimmer 30B27.8B · MMLU-Pro —Q3_K_M · ~32.2 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q2_K · ~32.5 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q2_K · ~36.9 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q3_K_M · ~32.3 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~49.6 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q3_K_M · ~67.5 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q3_K_M · ~33.9 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q3_K_M · ~34.8 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q4_K_M · ~58.1 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q6_K · ~32.4 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q5_K_M · ~36.2 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q6_K · ~34 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q6_K · ~38.5 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q6_K · ~39.4 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q6_K · ~33.4 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q8_0 · ~34.7 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5Q8_0 · ~44.4 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3Q8_0 · ~45.6 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0Q8_0 · ~45.6 t/s
- Qwen3 8B8B · MMLU-Pro 56.7Q8_0 · ~45 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3Q8_0 · ~51.1 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0Q8_0 · ~49.7 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~51.4 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~48.5 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~40.4 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~50.4 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~59.5 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~34.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~38.7 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~52 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~51.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~70.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~83.5 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~100.9 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~207.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~246 t/s
Models that fit with CPU offload (12)
These use system RAM for layers that don't fit in VRAM — expect much slower inference.
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~1.6 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q3_K_M · ~1.1 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q3_K_M · ~1.1 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q3_K_M · ~1.1 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q5_K_M · ~1.4 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q5_K_M · ~1.2 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q6_K · ~1.5 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q6_K · ~1.6 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q8_0 · ~1.1 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q8_0 · ~1.1 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q8_0 · ~1.1 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q8_0 · ~1.1 t/s
Too large for this GPU (34)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
Compare NVIDIA RTX 4070 Ti SUPER with other GPUs
Frequently asked questions
- How much VRAM does the NVIDIA RTX 4070 Ti SUPER have?
- The NVIDIA RTX 4070 Ti SUPER has 16 GB of GDDR6X with 672 GB/s memory bandwidth.
- What is the NVIDIA RTX 4070 Ti SUPER best for?
- With 16 GB of VRAM, the NVIDIA RTX 4070 Ti SUPER handles smaller models (7B–14B) at Q4–Q5 quantization — ideal for entry-level local LLM experimentation and lightweight inference.
- What LLMs can the NVIDIA RTX 4070 Ti SUPER run locally?
- The NVIDIA RTX 4070 Ti SUPER can run 41 of the 87 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at Q8_0, Llama 3.2 3B Instruct at BF16, Llama 3.2 1B Instruct at FP32.
- Can the NVIDIA RTX 4070 Ti SUPER run Llama 3.3 70B Instruct?
- The NVIDIA RTX 4070 Ti SUPER can run Llama 3.3 70B Instruct with CPU offload at Q3_K_M quantization, but inference will be slower than native VRAM execution.
- Can the NVIDIA RTX 4070 Ti SUPER run Qwen 3.6 27B?
- Yes. The NVIDIA RTX 4070 Ti SUPER runs Qwen 3.6 27B natively in VRAM at Q3_K_M quantization, achieving approximately 32.3 tokens per second.
- Can the NVIDIA RTX 4070 Ti SUPER run Llama 3.1 8B Instruct?
- Yes. The NVIDIA RTX 4070 Ti SUPER runs Llama 3.1 8B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 45.6 tokens per second.