NVIDIA B200 180GB
The NVIDIA B200 180GB has 180 GB VRAM and 8000 GB/s memory bandwidth. It can run 69 of our 85 tracked models natively in VRAM at 8k context.
With 180 GB HBM3e, the NVIDIA B200 180GB is a datacenter-tier GPU that can run 69 models natively. It runs 70B-class dense models and large MoE models entirely in VRAM.
NVIDIA B200 180GB: NVIDIA's first Blackwell datacenter GPU, unveiled at GTC in March 2024 and shipping in HGX/DGX B200 and GB200 NVL72 systems through 2024-2025, with 180GB of HBM3e at 8 TB/s — more than double the H100's 80GB capacity and 3,350 GB/s bandwidth. Its 2nd-gen Transformer Engine adds native FP4 (NVFP4) alongside FP8, and 5th-gen NVLink scales to 1.8 TB/s per GPU, double the H100/H200's 900 GB/s.
This site's calculator runs Llama 3.3 70B at full BF16 (159.81 GB, 36.4 tok/s) and GPT-OSS 120B (MoE) at NVFP4 (70.46 GB, 590.5 tok/s) natively — no quantization compromise on the 70B-class model. GLM 4.5 (355B, MoE) fits entirely in VRAM at Q2_K (154.94 GB, 118.9 tok/s), something the H200 141GB can't do at any quantization without CPU offload. Across the 85 tracked models, 69 fit fully in VRAM at 8k context, versus the H200's 66 and the H100's 59.
Needs CUDA Toolkit 12.8 or newer — the first Toolkit release with Blackwell (sm_100) support, so an older CUDA, driver, or framework build won't recognize this GPU's architecture. vLLM and TensorRT-LLM both shipped Blackwell/NVFP4 kernels within that 12.8 window. Cloud-only for almost every buyer, rentable by the hour on RunPod and other GPU clouds.
| Vendor | NVIDIA |
| Architecture | Blackwell |
| VRAM | 180 GB |
| Memory type | HBM3e |
| Memory bandwidth | 8000 GB/s |
| Compute backend | CUDA |
| Tier | Datacenter |
| Released | 2024 |
| Models (native) | 69 / 85 |
| Models (offload) | 3 / 85 |
Bandwidth-led this generation, capacity-led the one before it
Plotting NVIDIA's last four datacenter flagships by VRAM and bandwidth together shows the B200 sitting at a real inflection point — the generation where bandwidth, not capacity, did most of the work:
H200 to B200 (this page) is a bandwidth-led jump: 4,800 to 8,000 GB/s is a 67% gain, while capacity grows a smaller 28% (141GB to 180GB). That's the reverse of the step before it — H100 to H200 grew capacity 76% (80GB to 141GB) on a smaller 43% bandwidth gain (3,350 to 4,800 GB/s) — and the reverse of the step after it, where B300 adds 60% more capacity (180GB to 288GB) at the exact same 8,000 GB/s. Since decode speed is bandwidth-bound, the B200's 67% bandwidth gain over the H200 shows up directly in tokens/sec on any model that fits both; its smaller capacity gain is what decides which additional models fit at all.
Cloud GPU Rental
Don't want to buy a NVIDIA B200 180GB? RunPod is a cloud GPU rental service — rent one by the hour instead, no contract, no upfront hardware cost.
Pay by the hour · no contract · pods start in about a minute.
Rent a NVIDIA B200 180GB on RunPod ↗ (+$5 signup credit)Affiliate link — CanItRun may earn a commission. Doesn't affect the fit calculation above.
Popular models for this GPU
Models this GPU runs natively in VRAM (69)
- GLM-4.7 358B358B · MMLU-Pro 84.3Q2_K · ~118.9 t/s
- GLM-4.5 355B355B · MMLU-Pro 84.6Q2_K · ~118.9 t/s
- GLM-4.6 355B355B · MMLU-Pro 83.2Q2_K · ~118.9 t/s
- DeepSeek V4 Flash 284B284B · MMLU-Pro 86.3NVFP4 · ~239.2 t/s
- DeepSeek V4 Flash 0731 284B284B · MMLU-Pro —UD-IQ3_XXS · ~326.2 t/s
Show 64 more
- Qwen3 235B-A22B (MoE)235B · MMLU-Pro 84.4NVFP4 · ~136 t/s
- MiniMax M2.5 229B229B · MMLU-Pro 84.8NVFP4 · ~277.4 t/s
- MiniMax M2.7 229B229B · MMLU-Pro 86.0NVFP4 · ~277.4 t/s
- Step 3.7 Flash198B · MMLU-Pro —NVFP4 · ~262.1 t/s
- Step 3.5 Flash196.81B · MMLU-Pro 84.4NVFP4 · ~262.1 t/s
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0NVFP4 · ~77.8 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro —NVFP4 · ~77.7 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7NVFP4 · ~276.4 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7NVFP4 · ~250.7 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7NVFP4 · ~590.5 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3NVFP4 · ~167.6 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4NVFP4 · ~241.4 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9NVFP4 · ~241.4 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1BF16 · ~35.5 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9BF16 · ~36.4 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0BF16 · ~36.4 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4BF16 · ~36.4 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7BF16 · ~59.7 t/s
- Command-R 35B35B · MMLU-Pro 33.0FP32 · ~34.5 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3FP32 · ~129.5 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2FP32 · ~36.6 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0FP32 · ~37.2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5FP32 · ~39.2 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0FP32 · ~39.3 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3FP32 · ~39.3 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0FP32 · ~39.3 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3FP32 · ~128.6 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2FP32 · ~40.9 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5FP32 · ~127.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0FP32 · ~46.5 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5FP32 · ~47.5 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2FP32 · ~47.9 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~590.2 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6FP32 · ~101.9 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8FP32 · ~53.4 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2FP32 · ~57.3 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9FP32 · ~107.9 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0FP32 · ~85.9 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7FP32 · ~86.1 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4FP32 · ~90.7 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6FP32 · ~103.7 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~104.3 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2FP32 · ~101.5 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~131.3 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~143.4 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~157.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~157.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~156.6 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~168.4 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~172.9 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~315.1 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~305.8 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~282.3 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~319.5 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~378.5 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~409.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~461.3 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~618.9 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~618.3 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~834 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~994.6 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~1201.7 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~2475.4 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~2928.7 t/s
Models that fit with CPU offload (3)
These use system RAM for layers that don't fit in VRAM — expect much slower inference.
Too large for this GPU (13)
Frequently asked questions
- How much VRAM does the NVIDIA B200 180GB have?
- The NVIDIA B200 180GB has 180 GB of HBM3e with 8000 GB/s memory bandwidth.
- What is the NVIDIA B200 180GB best for?
- With 180 GB of VRAM, the NVIDIA B200 180GB is a server-class GPU that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
- What LLMs can the NVIDIA B200 180GB run locally?
- The NVIDIA B200 180GB can run 69 of the 85 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at BF16, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the NVIDIA B200 180GB run Llama 3.3 70B Instruct?
- Yes. The NVIDIA B200 180GB runs Llama 3.3 70B Instruct natively in VRAM at BF16 quantization, achieving approximately 36.4 tokens per second.
- Can the NVIDIA B200 180GB run Qwen 3.6 27B?
- Yes. The NVIDIA B200 180GB runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 47.9 tokens per second.
- Can the NVIDIA B200 180GB run Llama 3.1 8B Instruct?
- Yes. The NVIDIA B200 180GB runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 157.2 tokens per second.
- Can I rent the NVIDIA B200 180GB instead of buying it?
- Yes — RunPod and similar cloud GPU providers let you rent NVIDIA B200 180GB instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.