NVIDIA A100 40GB
The NVIDIA A100 40GB has 40 GB VRAM and 1555 GB/s memory bandwidth. It can run 57 of our 97 tracked models natively in VRAM at 8k context.
With 40 GB HBM2, the NVIDIA A100 40GB is a datacenter-tier GPU that can run 57 models natively. This site's calculator only fits Llama 3.3 70B at the most aggressive Q2_K (32.88 GB, 34.4 tok/s) in 40GB; every higher-quality quant either needs CPU offload or doesn't fit outright. The card is far more comfortable in the 8B-14B range: Qwen 2.5 14B fits natively even at full BF16 (34.73 GB, 32.6 tok/s), and Llama 3.1 8B is small enough to run at full, uncompressed FP32 precision (37.04 GB, 30.6 tok/s), a precision tier most models this site tracks never reach on any single GPU. GPT-OSS 120B (MoE) doesn't fit at any quantization; its smallest build alone needs 62.60 GB of weights. Across the 97 tracked models, 57 fit fully in VRAM at 8k context, 8 fewer than the 80GB card manages with double the capacity.
The NVIDIA A100 40GB is the PCIe variant of NVIDIA's Ampere data center GPU, the predecessor to the H100. Its 40GB of HBM2 memory and 1,555 GB/s bandwidth make it well-suited for running 7B–34B models at high throughput in cloud or on-prem inference clusters. This is the GPU behind many AWS p4d and Azure NDv4 instances and remains widely deployed in production LLM serving.
NVIDIA A100 40GB: The original A100 configuration, unveiled at GTC in May 2020 as Ampere's debut datacenter GPU, before the 80GB HBM2e variant followed that November. This 40GB card uses the same 826mm² GA100 die (TSMC 7nm, 54.2 billion transistors, 6,912 CUDA cores) as the 80GB card, but with older HBM2 instead of HBM2e, capping bandwidth at 1,555 GB/s, 24% below the 80GB card's 2,039 GB/s despite identical silicon. Ships as either a 400W SXM4 module or a 250W PCIe card, and like the 80GB variant supports Multi-Instance GPU (MIG) partitioning into up to seven isolated slices.
This site's calculator only fits Llama 3.3 70B at the most aggressive Q2_K (32.88 GB, 34.4 tok/s) in 40GB; every higher-quality quant either needs CPU offload or doesn't fit outright. The card is far more comfortable in the 8B-14B range: Qwen 2.5 14B fits natively even at full BF16 (34.73 GB, 32.6 tok/s), and Llama 3.1 8B is small enough to run at full, uncompressed FP32 precision (37.04 GB, 30.6 tok/s), a precision tier most models this site tracks never reach on any single GPU. GPT-OSS 120B (MoE) doesn't fit at any quantization; its smallest build alone needs 62.60 GB of weights. Across the 97 tracked models, 57 fit fully in VRAM at 8k context, 8 fewer than the 80GB card manages with double the capacity.
Full CUDA support under compute capability 8.0 (sm_80), identical software stack to the 80GB card: llama.cpp, vLLM, and TensorRT-LLM all treat the two capacities the same way, just with a lower ceiling. Common on AWS p4d and Azure NDv4 instances, where several 40GB cards get pooled over NVLink (600 GB/s per GPU on SXM4) for models too large for one card. The cheapest way into genuine datacenter-class CUDA hardware on the secondhand market now that 80GB+ cards dominate new deployments, though the smaller ceiling and MIG partitioning make it better suited to serving many small models at once than any single large one.
| Vendor | NVIDIA |
| Architecture | Ampere |
| VRAM | 40 GB |
| Memory type | HBM2 |
| Memory bandwidth | 1555 GB/s |
| Compute backend | CUDA |
| Tier | Datacenter |
| Released | 2020 |
| Models (native) | 57 / 97 |
| Models (offload) | 7 / 97 |
Less capacity than same-tier workstation cards, far more bandwidth
VRAM and bandwidth don't move together across NVIDIA's stack; a card can trail on one and lead on the other. Plotting this card against a same-tier Ampere workstation card and its own 80GB sibling:
At 40GB, this card trails the RTX A6000's 48GB by 8GB, but at 1,555 GB/s it has more than double the RTX A6000's bandwidth (768 GB/s, +102%). The gap holds against other workstation-tier 48GB cards too: the L40S's 864 GB/s is still 80% below this card's 1,555 GB/s. That shows up directly in decode speed: this site's calculator puts Llama 3.1 8B at Q4_K_M at 170.0 tok/s on this card, versus 84.0 tok/s on the RTX A6000 and 94.5 tok/s on the L40S, nearly twice the throughput on a card with less memory to work with. HBM beats GDDR6 on bandwidth even at a capacity disadvantage, and the 80GB version of this same card pushes that same HBM advantage to 2,039 GB/s with double the capacity too.
Same die, different memory: HBM2 to HBM2e is a real jump, not just capacity
The 40GB and 80GB A100 aren't different chips; both use the same 826mm² GA100 die. The only thing that changes between them is the memory: HBM2 on this card, newer HBM2e on the 80GB card. Tracking bandwidth from this card through the next NVIDIA generation:
This card's 1,555 GB/s HBM2 jumps 31% to 2,039 GB/s just by switching to HBM2e on the 80GB card, same GA100 die, same architecture, purely a memory upgrade. The next generation goes further still: the H100's HBM3 adds another 64% on top of that, to 3,350 GB/s, on a new GH100 die. Since decode is bandwidth-bound, that compounds directly into tokens per second: this site's calculator measures Llama 3.1 8B at Q4_K_M at 170.0 tok/s on this card, 222.9 tok/s on the 80GB card, and 366.2 tok/s on the H100, each step tracking its bandwidth gain almost exactly.
Popular models for this GPU
Models this GPU runs natively in VRAM (57)
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~33.6 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~34.4 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~34.4 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~34.4 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q4_K_M · ~37.1 t/s
Show 52 more
- Command-R 35B35B · MMLU-Pro 33.0Q4_K_M · ~31.5 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q6_K · ~120.6 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q6_K · ~32.7 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/AQ6_K · ~120.6 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q6_K · ~33.4 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q6_K · ~35.8 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q6_K · ~35.1 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q6_K · ~35.1 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q6_K · ~35.1 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q6_K · ~116.9 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q6_K · ~37.8 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~88.4 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q8_0 · ~94.6 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/AQ8_0 · ~34 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~31.6 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~33.4 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~34.6 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2Q8_0 · ~34.6 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~114.7 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/AQ8_0 · ~34.6 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q8_0 · ~73 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q8_0 · ~37.6 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q8_0 · ~39.7 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~78 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~32.7 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~32.6 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~34.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~39.3 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~39.7 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~37.1 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~47.6 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~55.3 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ABF16 · ~55.3 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~59.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~59.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~58.7 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~64.5 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~64.9 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~118.9 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~112.2 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~93.4 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~116.5 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~137.7 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~155.5 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~166.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~229.6 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~201.7 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~312.5 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~367.8 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~434.3 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~918.3 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~957.6 t/s
Models that fit with CPU offload (7)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0Q2_K · ~1.5 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro N/AQ2_K · ~1.7 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q2_K · ~9.4 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q2_K · ~7.7 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q3_K_M · ~2.5 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q3_K_M · ~4.1 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q3_K_M · ~4.1 t/s
Too large for this GPU (33)
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.6 355B
- GLM-4.7 358B
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
Compare NVIDIA A100 40GB with other GPUs
Frequently asked questions
- How much VRAM does the NVIDIA A100 40GB have?
- The NVIDIA A100 40GB has 40 GB of HBM2 with 1555 GB/s memory bandwidth.
- What is the NVIDIA A100 40GB best for?
- With 40 GB of VRAM, the NVIDIA A100 40GB is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the NVIDIA A100 40GB run locally?
- The NVIDIA A100 40GB can run 57 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q8_0, Ornith 1.5 35B-A3B (MoE) at Q6_K, Ornith 1.5 9B at BF16.
- Can the NVIDIA A100 40GB run Gemma 4 31B?
- Yes. The NVIDIA A100 40GB runs Gemma 4 31B natively in VRAM at Q6_K quantization, achieving approximately 37.8 tokens per second.
- Can the NVIDIA A100 40GB run Qwen 3.6 27B?
- Yes. The NVIDIA A100 40GB runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 34.6 tokens per second.
- Can the NVIDIA A100 40GB run Qwen3 8B?
- Yes. The NVIDIA A100 40GB runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 58.7 tokens per second.