NVIDIA RTX 3090
The NVIDIA RTX 3090 has 24 GB VRAM and 936 GB/s memory bandwidth. It can run 52 of our 97 tracked models natively in VRAM at 8k context.
With 24 GB GDDR6X, the NVIDIA RTX 3090 is a consumer-tier GPU that can run 52 models natively. This site's calculator puts Qwen3 8B at 100.1 tok/s at Q4_K_M (6.81 GB, fits with room to spare) and Qwen3 14B at its recommended Q8_0 quant at 35.6 tok/s (19.12 GB), both comfortably inside the 24GB budget with headroom for context. Qwen 3.8 27B fits too: its recommended Q4_K_M build leaves real headroom (19.02 GB, 35.8 tok/s), and even the higher-quality Q5_K_M still fits, just with well under a gigabyte of this card's ~22.8 GB usable budget left over (22.13 GB, 30.8 tok/s). The real ceiling sits around 32-34B, the same pattern the RTX 4090 shows one generation later: Qwen3 32B and Yi 1.5 34B both need Q3_K_M to fit natively (19.17 GB at 35.5 tok/s, and 20.79 GB at 32.8 tok/s) at 8k context, while the more common Q4_K_M build tips both into CPU offload (23.88 GB at 28.5 tok/s, and 25.72 GB at 22.7 tok/s), a modest falloff rather than a cliff. MoE models do better: Qwen3 30B-A3B fits at Q4_K_M (21.36 GB, 88.2 tok/s) and Qwen3.5 35B-A3B fits at Q3_K_M (19.04 GB, 122.2 tok/s), since only about 3B of their parameters activate per token. GPT-OSS 20B, a model OpenAI explicitly sized for a 16GB card, fits here with real headroom at Q6_K (19.54 GB, 60.5 tok/s). 70B-class dense models (Llama 3.3 70B, Qwen2.5 72B) don't fit a single card at any quantization at 8k context: even the smallest real builds, Q2_K at 32.88-33.73 GB, need CPU offload and crawl at 3.0-3.3 tok/s. This is exactly where a second RTX 3090 changes the picture: this site's calculator pools two cards at roughly 43.2 GB of usable VRAM, and Llama 3.3 70B fits there natively at Q3_K_M (40.72 GB, 15.1 tok/s), a roughly 9x jump over the single-card offload result of 1.7 tok/s at that same quant. The more commonly distributed Q4_K_M build doesn't quite clear that pooled 43.2 GB ceiling either (50.75 GB total), so it still spills a few GB into system RAM even across two cards, landing at 6.1 tok/s rather than running fully native; Q3_K_M, not Q4_K_M, is the honest answer for what two RTX 3090s actually run natively at 70B.
The NVIDIA RTX 3090 launched September 24, 2020 at a $1,499 MSRP, the flagship Ampere consumer GPU and the first GeForce card to put 24GB of VRAM within reach of a home workstation. It pairs 24GB of GDDR6X on a 384-bit bus (936 GB/s bandwidth) with a cut-down GA102 die: 10,496 CUDA cores from 82 of the die's 84 streaming multiprocessors, one bin below the fully unlocked 10,752-core GA102 configuration NVIDIA reserved for the RTX 3090 Ti and the RTX A6000 workstation card. Its most durable local-inference feature has nothing to do with raw specs: the RTX 3090 and RTX 3090 Ti are the last GeForce cards NVIDIA shipped with a working NVLink bridge, letting two cards pool VRAM into roughly 48GB for models like Llama 3.3 70B that don't fit on any single 24GB card, a trick no consumer GPU since (RTX 4090, RTX 5090) can replicate because NVIDIA dropped the connector.
NVIDIA RTX 3090: Launched September 24, 2020 at a $1,499 MSRP, the RTX 3090 was the flagship of NVIDIA's Ampere generation and the first GeForce card to carry 24GB of VRAM (NVIDIA's own RTX 3090/3090 Ti product page confirms every core spec here, cross-checked against Wikipedia's GeForce RTX 30 series spec table). Its GA102 die is cut down to 10,496 of the silicon's 10,752 CUDA cores (82 of 84 streaming multiprocessors), one bin below the fully unlocked GA102 configuration NVIDIA reserved for the RTX 3090 Ti and the professional RTX A6000. Five years after launch it still commands real money on the used market: community-tracked listings in 2026 commonly cite roughly $700 to $1,050 for a working card, evidence of sustained demand from local-AI hobbyists rather than the price a five-year-old flagship GPU would normally fetch on its own. The RTX 3090 and RTX 3090 Ti are also the last GeForce cards NVIDIA shipped with a working NVLink bridge: NVIDIA's own spec page lists NVLink support for both, and a bridge (commonly cited around 112.5 GB/s of bidirectional bandwidth, though NVIDIA's own product page doesn't publish that figure directly) lets two cards pool memory into one address space. The RTX 4090 and RTX 5090 that followed dropped the connector entirely, which is the real reason a used RTX 3090 pair, not a newer single card, is still the cheapest way into a 48GB pool for local inference.
This site's calculator puts Qwen3 8B at 100.1 tok/s at Q4_K_M (6.81 GB, fits with room to spare) and Qwen3 14B at its recommended Q8_0 quant at 35.6 tok/s (19.12 GB), both comfortably inside the 24GB budget with headroom for context. Qwen 3.8 27B fits too: its recommended Q4_K_M build leaves real headroom (19.02 GB, 35.8 tok/s), and even the higher-quality Q5_K_M still fits, just with well under a gigabyte of this card's ~22.8 GB usable budget left over (22.13 GB, 30.8 tok/s). The real ceiling sits around 32-34B, the same pattern the RTX 4090 shows one generation later: Qwen3 32B and Yi 1.5 34B both need Q3_K_M to fit natively (19.17 GB at 35.5 tok/s, and 20.79 GB at 32.8 tok/s) at 8k context, while the more common Q4_K_M build tips both into CPU offload (23.88 GB at 28.5 tok/s, and 25.72 GB at 22.7 tok/s), a modest falloff rather than a cliff. MoE models do better: Qwen3 30B-A3B fits at Q4_K_M (21.36 GB, 88.2 tok/s) and Qwen3.5 35B-A3B fits at Q3_K_M (19.04 GB, 122.2 tok/s), since only about 3B of their parameters activate per token. GPT-OSS 20B, a model OpenAI explicitly sized for a 16GB card, fits here with real headroom at Q6_K (19.54 GB, 60.5 tok/s). 70B-class dense models (Llama 3.3 70B, Qwen2.5 72B) don't fit a single card at any quantization at 8k context: even the smallest real builds, Q2_K at 32.88-33.73 GB, need CPU offload and crawl at 3.0-3.3 tok/s. This is exactly where a second RTX 3090 changes the picture: this site's calculator pools two cards at roughly 43.2 GB of usable VRAM, and Llama 3.3 70B fits there natively at Q3_K_M (40.72 GB, 15.1 tok/s), a roughly 9x jump over the single-card offload result of 1.7 tok/s at that same quant. The more commonly distributed Q4_K_M build doesn't quite clear that pooled 43.2 GB ceiling either (50.75 GB total), so it still spills a few GB into system RAM even across two cards, landing at 6.1 tok/s rather than running fully native; Q3_K_M, not Q4_K_M, is the honest answer for what two RTX 3090s actually run natively at 70B.
Full CUDA support under compute capability 8.6, the same mature Ampere target every RTX 30-series card shares: llama.cpp, Ollama, and every major inference tool has had five years to optimize for it, with none of Blackwell's early sm_120 driver caveats. Ampere predates Ada's hardware FP8 support and Blackwell's NVFP4, so GGUF's K-quants (down to Q2_K) are the only quantization path this site's calculator applies here; there's no newer low-bit format waiting to be unlocked the way there is on RTX 40- and 50-series cards. On the NVLink question specifically: llama.cpp's default multi-GPU behavior is pipeline/layer-split, where only one GPU computes at a time and the data crossing between cards each token is a small activation vector, not the whole model, so this site's calculator applies the identical multi-GPU bandwidth model whether or not an NVLink bridge is physically installed, matching this site's own multi-GPU guide's position that the bridge buys little decode speed for that common case. The bridge's real value is pooling capacity to roughly 48GB, not raw tokens per second, and even that takes real setup work: user reports on NVIDIA's own developer forums describe needing an SLI-certified motherboard, matched PCIe lane widths, and current drivers before peer-to-peer transfers reliably engage, friction that a bridge-free PCIe layer-split pair, which needs none of that, avoids entirely.
| Vendor | NVIDIA |
| Architecture | Ampere |
| VRAM | 24 GB |
| Memory type | GDDR6X |
| Memory bandwidth | 936 GB/s |
| Compute backend | CUDA |
| Tier | Consumer |
| Released | 2020 |
| Models (native) | 52 / 97 |
| Models (offload) | 7 / 97 |
The card three generations of 24GB+ flagships get measured from
This card started the chain that led to the RTX 4090 and RTX 5090, NVIDIA's last three 24GB+ consumer flagships. Tracking VRAM and bandwidth across all three from here:
This card's step to the RTX 4090 was the smallest bandwidth move of the three: just 7.7% (936 to 1,008 GB/s) at an unchanged 24GB, nearly a placeholder generation on the number that decides decode speed, even though the die, process node, and Tensor Core generation all changed underneath. The RTX 5090 two generations later moved much further past this card: 91.5% more bandwidth (936 to 1,792 GB/s) and 33.3% more VRAM (24GB to 32GB), the largest gap between any two cards in this three-flagship chain. This site's calculator tracks that chain directly on Qwen3 8B at Q4_K_M: 100.1 tok/s on this card, 107.8 tok/s on the RTX 4090 (a 7.7% gain matching the bandwidth increase almost exactly), and 191.6 tok/s on the RTX 5090. What didn't change in either step: none of these three cards run a 70B-class model natively at this site's 8k-context benchmark on a single card; that only became possible for this card by adding a second one over NVLink, not by waiting for a newer single GPU.
What a second RTX 3090 actually buys: 48GB, not double the speed
NVIDIA dropped NVLink after this generation, so a pair of RTX 3090s bridged together (or just split by layer across two PCIe slots, which this site's calculator treats the same way) is still the cheapest way to reach a 48GB pool. Comparing what fits on one RTX 3090 against what fits on two, at each model's best matching quant once pooled:
Both 70B-class dense models here reach a native "fits" verdict on two pooled RTX 3090s at Q3_K_M (40.72-41.79 GB against roughly 43.2 GB of usable pooled VRAM) that a single card can't reach at any quantization: on one card the identical Q3_K_M build has to spill into system RAM, decoding at just 1.6-1.7 tok/s, versus 14.7-15.1 tok/s once both cards hold it natively, roughly a 9x speedup. That speedup comes from removing the CPU-offload penalty, not from NVLink itself: this site's calculator applies the same layer-split bandwidth model whether or not a physical bridge is installed, since llama.cpp's default pipeline split only moves one small activation vector between cards per token, not the full weight set. The pooled 43.2 GB ceiling has its own limit too: even across two cards, the more commonly distributed Q4_K_M build of these same models (50.75-52.12 GB) still spills a few GB into system RAM rather than running fully native, and GPT-OSS 120B doesn't fit at all, even pooled and fully offloaded, since its 70.46 GB weight floor alone exceeds two RTX 3090s' combined ~69.2 GB of VRAM+RAM budget.
Cloud GPU Rental
Don't want to buy a NVIDIA RTX 3090? RunPod is a cloud GPU rental service: rent one by the hour instead, no contract, no upfront hardware cost.
Pay by the hour · no contract · pods start in about a minute.
Rent a NVIDIA RTX 3090 on RunPod ↗ (+$5 signup credit)Affiliate link: CanItRun may earn a commission. Doesn't affect the fit calculation above.
Popular models for this GPU
Models this GPU runs natively in VRAM (52)
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q2_K · ~34.9 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q3_K_M · ~122.2 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q3_K_M · ~32.1 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/AQ3_K_M · ~122.2 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q3_K_M · ~32.8 t/s
Show 47 more
- Qwen3 32B32.8B · MMLU-Pro 65.5Q3_K_M · ~35.5 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q3_K_M · ~34.2 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q3_K_M · ~34.2 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q3_K_M · ~34.2 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q4_K_M · ~93.2 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q4_K_M · ~30.1 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q4_K_M · ~88.2 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q4_K_M · ~99.1 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/AQ5_K_M · ~30.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q4_K_M · ~31 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q4_K_M · ~33.8 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q5_K_M · ~30.8 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2Q5_K_M · ~30.8 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~69.1 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/AQ5_K_M · ~30.8 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q5_K_M · ~64.7 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q5_K_M · ~33 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q6_K · ~30.3 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q6_K · ~60.5 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q8_0 · ~35.6 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q8_0 · ~35.3 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q8_0 · ~37.5 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q8_0 · ~42.5 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q8_0 · ~43.4 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q8_0 · ~38.1 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q8_0 · ~48.3 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~33.3 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ABF16 · ~33.3 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~35.6 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~35.6 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~35.4 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~38.8 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~39.1 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~71.5 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~67.6 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~56.2 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~70.1 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~82.9 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~93.6 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~100.2 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~138.2 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~121.4 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~188.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~221.4 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~261.4 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~552.8 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~576.4 t/s
Models that fit with CPU offload (7)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q2_K · ~3.1 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q2_K · ~3.1 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q3_K_M · ~1.6 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q3_K_M · ~1.7 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q3_K_M · ~1.7 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q3_K_M · ~1.7 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q6_K · ~1.4 t/s
Too large for this GPU (38)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.6 355B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
Compare NVIDIA RTX 3090 with other GPUs
Continue reading
Frequently asked questions
- How much VRAM does the NVIDIA RTX 3090 have?
- The NVIDIA RTX 3090 has 24 GB of GDDR6X with 936 GB/s memory bandwidth.
- What is the NVIDIA RTX 3090 best for?
- With 24 GB of VRAM, the NVIDIA RTX 3090 is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the NVIDIA RTX 3090 run locally?
- The NVIDIA RTX 3090 can run 52 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q5_K_M, Ornith 1.5 35B-A3B (MoE) at Q3_K_M, Ornith 1.5 9B at BF16.
- Can the NVIDIA RTX 3090 run Gemma 4 31B?
- Yes. The NVIDIA RTX 3090 runs Gemma 4 31B natively in VRAM at Q4_K_M quantization, achieving approximately 30.1 tokens per second.
- Can the NVIDIA RTX 3090 run Qwen 3.6 27B?
- Yes. The NVIDIA RTX 3090 runs Qwen 3.6 27B natively in VRAM at Q5_K_M quantization, achieving approximately 30.8 tokens per second.
- Can the NVIDIA RTX 3090 run Qwen3 8B?
- Yes. The NVIDIA RTX 3090 runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 35.4 tokens per second.
- Can I rent the NVIDIA RTX 3090 instead of buying it?
- Yes: RunPod and similar cloud GPU providers let you rent NVIDIA RTX 3090 instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.