NVIDIA RTX 4070 Ti SUPER

The NVIDIA RTX 4070 Ti SUPER has 16 GB VRAM and 672 GB/s memory bandwidth. It can run 41 of our 87 tracked models natively in VRAM at 8k context.

With 16 GB GDDR6X, the NVIDIA RTX 4070 Ti SUPER is a consumer-tier GPU that can run 41 models natively. It handles smaller models (7B–14B) at Q4–Q5 quantization.

The NVIDIA RTX 4070 Ti SUPER launched January 24, 2024 at a $799 MSRP, moving off the RTX 4070 Ti's AD104 die onto the larger AD103 die shared with the RTX 4080. That switch grows VRAM from 12GB to 16GB and widens the memory bus from 192-bit to 256-bit (672 GB/s, up from 504 GB/s), while CUDA cores climb from 7,680 to 8,448. It is the only card in the RTX 4070 family whose memory subsystem actually changed: the RTX 4070 SUPER kept the same 192-bit, 504 GB/s memory as the RTX 4070 and RTX 4070 Ti on either side of it.

NVIDIA RTX 4070 Ti SUPER: NVIDIA's January 2024 Super refresh gave the RTX 4070 Ti a real memory upgrade, not just a clock bump: this card moves off the RTX 4070 Ti's AD104 die onto the larger AD103-275 die (45.9 billion transistors, up from AD104's 35.8 billion) that also powers the RTX 4080, gaining 4GB of VRAM and a wider 256-bit memory bus in the process (NVIDIA's own GeForce RTX 40 SUPER announcement; Wikipedia's GeForce RTX 40 series spec table cross-checks every figure here). Launched January 24, 2024 at $799, it landed $200 under the RTX 4080's $999 launch price for the exact same 16GB capacity, trading only 45 GB/s of bandwidth (672 vs 717 GB/s) and 1,280 fewer CUDA cores (8,448 vs 9,728) to get there. GamersNexus's launch review measured a real but modest 10-11% gaming uplift over the plain RTX 4070 Ti at 4K, a gap that narrows further at lower resolutions.

This site's calculator puts Llama 3.1 8B at 45.6 tok/s at Q8_0 (10.73 GB, fits with room to spare) and GPT-OSS 20B, a model OpenAI explicitly sized for a 16GB card, at 58.1 tok/s at Q4_K_M (14.55 GB). Both figures land within 10% of the RTX 4080's numbers for the same models (48.7 and 62 tok/s), tracking the two cards' 717-vs-672 GB/s bandwidth gap almost exactly, since decode is bandwidth-bound. The jump over this card's own predecessor is much larger: the plain RTX 4070 Ti needs CPU offload for GPT-OSS 20B (25.23 GB spills past its 12GB ceiling, 3.6 tok/s) and for Qwen 3.6 27B (32.75 GB, 1.3 tok/s), while this card runs both natively, GPT-OSS 20B at 58.1 tok/s and Qwen 3.6 27B at 32.3 tok/s (Q3_K_M, 15.15 GB). Measured against the Blackwell RTX 5070 Ti at the identical Q3_K_M quant, Qwen 3.6 27B decodes at 43.1 tok/s there versus 32.3 tok/s here, a 33.4% gap that matches the two cards' 896-vs-672 GB/s bandwidth difference (33.3%) almost exactly, since both fit the same models at the same quant and the only real difference is bandwidth.

Full CUDA support on Ada Lovelace, a mature target for llama.cpp, Ollama, vLLM, and TensorRT-LLM alike, unlike Blackwell's rockier early sm_120 rollout. This card's 4th-gen Tensor Cores add hardware FP8 support, accelerated natively by TensorRT-LLM and vLLM on Ada and Hopper, but not NVFP4: that format is Blackwell-only in this site's calculator, so Q2_K remains the smallest quant on this card's standard ladder, unlike the RTX 5070 Ti at the same 16GB. On value, this is the only card in the RTX 4070 family whose bandwidth actually improved over the base RTX 4070 and RTX 4070 Ti; the RTX 4070 SUPER released the same week kept the identical 504 GB/s memory subsystem as both, gaining only compute. For anyone comparing this card against its direct Blackwell successor, the RTX 5070 Ti, the honest baseline is this card's 672 GB/s, not the original RTX 4070 Ti's 504 GB/s: measured from here, the 5070 Ti's bandwidth gain is a smaller 33.3%, not the 77.8% a comparison against the original RTX 4070 Ti would suggest.

VendorNVIDIA
ArchitectureAda Lovelace
VRAM16 GB
Memory typeGDDR6X
Memory bandwidth672 GB/s
Compute backendCUDA
TierConsumer
Released2024
Models (native)41 / 87
Models (offload)12 / 87
Software: Full llama.cpp and Ollama support out of the box. CUDA 12.x recommended; driver ≥ 525 required.

The one 4070-family bandwidth jump, bookended by two flat ones

This card is the only Ada-generation 4070-family GPU whose memory subsystem actually changed. Tracking bandwidth from the original RTX 4070 Ti through this card to its Blackwell successor:

NVIDIA RTX 4070 Ti
504 GB/s
NVIDIA RTX 4070 Ti SUPER — this page
672 GB/s
NVIDIA RTX 5070 Ti
896 GB/s

The RTX 4070 Ti to this card is a real jump: 504 to 672 GB/s, up 33.3%, alongside 4GB of extra VRAM (12GB to 16GB) from moving off the AD104 die onto AD103. That capacity jump matters as much as the bandwidth: Qwen 3.6 27B doesn't fit the RTX 4070 Ti's 12GB at any quantization and needs CPU offload (1.3 tok/s), while it fits this card natively at Q3_K_M (15.15 GB, 32.3 tok/s). The RTX 4070 SUPER released the same week took none of that: it kept the identical 504 GB/s and 12GB as the RTX 4070 Ti, gaining only CUDA cores. The step from here to the RTX 5070 Ti is a second bandwidth jump of almost the same size, 672 to 896 GB/s, up 33.3%, this time from GDDR6X to GDDR7 rather than a wider bus, with VRAM capacity unchanged at 16GB. Since decode is bandwidth-bound, this site's calculator measures that second jump almost exactly on a model both cards fit at the identical Q3_K_M quant: Qwen 3.6 27B decodes at 32.3 tok/s here versus 43.1 tok/s on the RTX 5070 Ti, a 33.4% gap matching the bandwidth ratio within a rounding error.

Popular models for this GPU

Models this GPU runs natively in VRAM (41)

Show 36 more

Models that fit with CPU offload (12)

These use system RAM for layers that don't fit in VRAM — expect much slower inference.

Too large for this GPU (34)

Compare NVIDIA RTX 4070 Ti SUPER with other GPUs

Frequently asked questions

How much VRAM does the NVIDIA RTX 4070 Ti SUPER have?
The NVIDIA RTX 4070 Ti SUPER has 16 GB of GDDR6X with 672 GB/s memory bandwidth.
What is the NVIDIA RTX 4070 Ti SUPER best for?
With 16 GB of VRAM, the NVIDIA RTX 4070 Ti SUPER handles smaller models (7B–14B) at Q4–Q5 quantization — ideal for entry-level local LLM experimentation and lightweight inference.
What LLMs can the NVIDIA RTX 4070 Ti SUPER run locally?
The NVIDIA RTX 4070 Ti SUPER can run 41 of the 87 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at Q8_0, Llama 3.2 3B Instruct at BF16, Llama 3.2 1B Instruct at FP32.
Can the NVIDIA RTX 4070 Ti SUPER run Llama 3.3 70B Instruct?
The NVIDIA RTX 4070 Ti SUPER can run Llama 3.3 70B Instruct with CPU offload at Q3_K_M quantization, but inference will be slower than native VRAM execution.
Can the NVIDIA RTX 4070 Ti SUPER run Qwen 3.6 27B?
Yes. The NVIDIA RTX 4070 Ti SUPER runs Qwen 3.6 27B natively in VRAM at Q3_K_M quantization, achieving approximately 32.3 tokens per second.
Can the NVIDIA RTX 4070 Ti SUPER run Llama 3.1 8B Instruct?
Yes. The NVIDIA RTX 4070 Ti SUPER runs Llama 3.1 8B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 45.6 tokens per second.