NVIDIA RTX 6000 Ada

The NVIDIA RTX 6000 Ada has 48 GB VRAM and 960 GB/s memory bandwidth. It can run 57 of our 97 tracked models natively in VRAM at 8k context.

With 48 GB GDDR6, the NVIDIA RTX 6000 Ada is a workstation-tier GPU that can run 57 models natively. 70B at Q3_K_M native (Q4_K_M's ~51GB total is just past this card's effective capacity), with ~25% higher tokens/sec than A6000 at the same quant. ~30-45 t/s for 7B, ~17 t/s for 70B at Q3_K_M.

The NVIDIA RTX 6000 Ada Generation is the Ada Lovelace successor to the RTX A6000, upgrading memory bandwidth from 768 to 960 GB/s while keeping the same 48GB GDDR6 VRAM and workstation form factor. The jump in bandwidth meaningfully improves inference tokens-per-second on larger models. Like its predecessor, it supports NVLink, ECC memory, and fits in standard workstations, making it the top-tier single-GPU option for on-prem LLM workloads that need professional reliability.

NVIDIA RTX 6000 Ada: October 2022 Ada workstation with 48GB GDDR6 at 960 GB/s, A6000 successor.

70B at Q3_K_M native (Q4_K_M's ~51GB total is just past this card's effective capacity), with ~25% higher tokens/sec than A6000 at the same quant. ~30-45 t/s for 7B, ~17 t/s for 70B at Q3_K_M.

Full CUDA with ECC. NVLink support. Top single-GPU option for on-prem LLM needing professional reliability.

VendorNVIDIA
ArchitectureAda Lovelace
VRAM48 GB
Memory typeGDDR6
Memory bandwidth960 GB/s
Compute backendCUDA
TierWorkstation
Released2022
Models (native)57 / 97
Models (offload)8 / 97
Software: Full llama.cpp and Ollama support out of the box. CUDA 12.x recommended; driver ≥ 525 required.

Cloud GPU Rental

Don't want to buy a NVIDIA RTX 6000 Ada? RunPod is a cloud GPU rental service: rent one by the hour instead, no contract, no upfront hardware cost.

Pay by the hour · no contract · pods start in about a minute.

Rent a NVIDIA RTX 6000 Ada on RunPod ↗ (+$5 signup credit)

Affiliate link: CanItRun may earn a commission. Doesn't affect the fit calculation above.

Popular models for this GPU

Models this GPU runs natively in VRAM (57)

Show 52 more

Models that fit with CPU offload (8)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (32)

Compare NVIDIA RTX 6000 Ada with other GPUs

Frequently asked questions

How much VRAM does the NVIDIA RTX 6000 Ada have?
The NVIDIA RTX 6000 Ada has 48 GB of GDDR6 with 960 GB/s memory bandwidth.
What is the NVIDIA RTX 6000 Ada best for?
With 48 GB of VRAM, the NVIDIA RTX 6000 Ada is ideal for running 70B-class models at Q3-class quantization and large MoE models, a workstation sweet spot for local inference.
What LLMs can the NVIDIA RTX 6000 Ada run locally?
The NVIDIA RTX 6000 Ada can run 57 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q8_0, Ornith 1.5 35B-A3B (MoE) at Q8_0, Ornith 1.5 9B at BF16.
Can the NVIDIA RTX 6000 Ada run Gemma 4 31B?
Yes. The NVIDIA RTX 6000 Ada runs Gemma 4 31B natively in VRAM at Q8_0 quantization, achieving approximately 18.3 tokens per second.
Can the NVIDIA RTX 6000 Ada run Qwen 3.6 27B?
Yes. The NVIDIA RTX 6000 Ada runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 21.3 tokens per second.
Can the NVIDIA RTX 6000 Ada run Qwen3 8B?
Yes. The NVIDIA RTX 6000 Ada runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 36.3 tokens per second.
Can I rent the NVIDIA RTX 6000 Ada instead of buying it?
Yes: RunPod and similar cloud GPU providers let you rent NVIDIA RTX 6000 Ada instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.