NVIDIA RTX 4000 Ada

The NVIDIA RTX 4000 Ada has 20 GB VRAM and 320 GB/s memory bandwidth. It can run 51 of our 94 tracked models natively in VRAM at 8k context.

With 20 GB GDDR6, the NVIDIA RTX 4000 Ada is a workstation-tier GPU that can run 51 models natively. It handles smaller models (7B–14B) at Q4–Q5 quantization.

The NVIDIA RTX 4000 Ada is the entry-level Ada Lovelace workstation GPU, built on a 160-bit memory bus with 20GB ECC GDDR6 at 320 GB/s. With 6,144 CUDA cores and a single-slot form factor, it is the most compact professional GPU in the Ada lineup. Its 20GB VRAM comfortably handles 13B models at Q8_0 and 14B–20B models at Q4_K_M, exceeding what typical consumer 16GB cards can hold, while fitting into thermally constrained workstations and SFF builds.

NVIDIA RTX 4000 Ada: 2023 Ada Lovelace workstation entry point: 20GB ECC GDDR6 on a 160-bit bus at 320 GB/s, single-slot form factor. The most compact professional GPU in the Ada lineup.

13B models fit at Q8_0; 14B-20B models fit at Q4_K_M, more headroom than a typical consumer 16GB card. ~15-22 t/s for 7B Q4.

Full CUDA with ECC memory. Single-slot cooling makes it the pick for thermally constrained SFF workstations where an RTX 4500/5000 Ada won't physically fit.

VendorNVIDIA
ArchitectureAda Lovelace
VRAM20 GB
Memory typeGDDR6
Memory bandwidth320 GB/s
Compute backendCUDA
TierWorkstation
Released2023
Models (native)51 / 94
Models (offload)6 / 94
Software: Full llama.cpp and Ollama support out of the box. CUDA 12.x recommended; driver ≥ 525 required.

Popular models for this GPU

Models this GPU runs natively in VRAM (51)

Show 46 more

Models that fit with CPU offload (6)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (37)

Frequently asked questions

How much VRAM does the NVIDIA RTX 4000 Ada have?
The NVIDIA RTX 4000 Ada has 20 GB of GDDR6 with 320 GB/s memory bandwidth.
What is the NVIDIA RTX 4000 Ada best for?
With 20 GB of VRAM, the NVIDIA RTX 4000 Ada handles smaller models (7B–14B) at Q4–Q5 quantization, ideal for entry-level local LLM experimentation and lightweight inference.
What LLMs can the NVIDIA RTX 4000 Ada run locally?
The NVIDIA RTX 4000 Ada can run 51 of the 94 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q3_K_M, Muse Glimmer 30B at Q3_K_M, Ornith 1.5 9B at Q8_0.
Can the NVIDIA RTX 4000 Ada run Gemma 4 31B?
Yes. The NVIDIA RTX 4000 Ada runs Gemma 4 31B natively in VRAM at Q3_K_M quantization, achieving approximately 12.8 tokens per second.
Can the NVIDIA RTX 4000 Ada run Qwen 3.6 27B?
Yes. The NVIDIA RTX 4000 Ada runs Qwen 3.6 27B natively in VRAM at Q3_K_M quantization, achieving approximately 15.4 tokens per second.
Can the NVIDIA RTX 4000 Ada run Qwen3 8B?
Yes. The NVIDIA RTX 4000 Ada runs Qwen3 8B natively in VRAM at Q8_0 quantization, achieving approximately 21.4 tokens per second.