CanItRun Logocanitrun.

NVIDIA B200 180GB

The NVIDIA B200 180GB has 180 GB VRAM and 8000 GB/s memory bandwidth. It can run 69 of our 85 tracked models natively in VRAM at 8k context.

With 180 GB HBM3e, the NVIDIA B200 180GB is a datacenter-tier GPU that can run 69 models natively. It runs 70B-class dense models and large MoE models entirely in VRAM.

NVIDIA B200 180GB: NVIDIA's first Blackwell datacenter GPU, unveiled at GTC in March 2024 and shipping in HGX/DGX B200 and GB200 NVL72 systems through 2024-2025, with 180GB of HBM3e at 8 TB/s — more than double the H100's 80GB capacity and 3,350 GB/s bandwidth. Its 2nd-gen Transformer Engine adds native FP4 (NVFP4) alongside FP8, and 5th-gen NVLink scales to 1.8 TB/s per GPU, double the H100/H200's 900 GB/s.

This site's calculator runs Llama 3.3 70B at full BF16 (159.81 GB, 36.4 tok/s) and GPT-OSS 120B (MoE) at NVFP4 (70.46 GB, 590.5 tok/s) natively — no quantization compromise on the 70B-class model. GLM 4.5 (355B, MoE) fits entirely in VRAM at Q2_K (154.94 GB, 118.9 tok/s), something the H200 141GB can't do at any quantization without CPU offload. Across the 85 tracked models, 69 fit fully in VRAM at 8k context, versus the H200's 66 and the H100's 59.

Needs CUDA Toolkit 12.8 or newer — the first Toolkit release with Blackwell (sm_100) support, so an older CUDA, driver, or framework build won't recognize this GPU's architecture. vLLM and TensorRT-LLM both shipped Blackwell/NVFP4 kernels within that 12.8 window. Cloud-only for almost every buyer, rentable by the hour on RunPod and other GPU clouds.

VendorNVIDIA
ArchitectureBlackwell
VRAM180 GB
Memory typeHBM3e
Memory bandwidth8000 GB/s
Compute backendCUDA
TierDatacenter
Released2024
Models (native)69 / 85
Models (offload)3 / 85
Software: Needs CUDA Toolkit 12.8 or newer — the first Toolkit release with Blackwell (sm_100) support; older CUDA builds won't recognize this GPU's architecture. vLLM and TensorRT-LLM are the recommended engines for NVFP4 batched inference.

Bandwidth-led this generation, capacity-led the one before it

Plotting NVIDIA's last four datacenter flagships by VRAM and bandwidth together shows the B200 sitting at a real inflection point — the generation where bandwidth, not capacity, did most of the work:

0425085000150300VRAM (GB)Bandwidth (GB/s)NVIDIA H100 80GBNVIDIA H200 141GBNVIDIA B200 180GBNVIDIA B300 288GB
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

H200 to B200 (this page) is a bandwidth-led jump: 4,800 to 8,000 GB/s is a 67% gain, while capacity grows a smaller 28% (141GB to 180GB). That's the reverse of the step before it — H100 to H200 grew capacity 76% (80GB to 141GB) on a smaller 43% bandwidth gain (3,350 to 4,800 GB/s) — and the reverse of the step after it, where B300 adds 60% more capacity (180GB to 288GB) at the exact same 8,000 GB/s. Since decode speed is bandwidth-bound, the B200's 67% bandwidth gain over the H200 shows up directly in tokens/sec on any model that fits both; its smaller capacity gain is what decides which additional models fit at all.

Cloud GPU Rental

Don't want to buy a NVIDIA B200 180GB? RunPod is a cloud GPU rental service — rent one by the hour instead, no contract, no upfront hardware cost.

Pay by the hour · no contract · pods start in about a minute.

Rent a NVIDIA B200 180GB on RunPod ↗ (+$5 signup credit)

Affiliate link — CanItRun may earn a commission. Doesn't affect the fit calculation above.

Popular models for this GPU

Models this GPU runs natively in VRAM (69)

Show 64 more

Models that fit with CPU offload (3)

These use system RAM for layers that don't fit in VRAM — expect much slower inference.

Too large for this GPU (13)

Frequently asked questions

How much VRAM does the NVIDIA B200 180GB have?
The NVIDIA B200 180GB has 180 GB of HBM3e with 8000 GB/s memory bandwidth.
What is the NVIDIA B200 180GB best for?
With 180 GB of VRAM, the NVIDIA B200 180GB is a server-class GPU that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
What LLMs can the NVIDIA B200 180GB run locally?
The NVIDIA B200 180GB can run 69 of the 85 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at BF16, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
Can the NVIDIA B200 180GB run Llama 3.3 70B Instruct?
Yes. The NVIDIA B200 180GB runs Llama 3.3 70B Instruct natively in VRAM at BF16 quantization, achieving approximately 36.4 tokens per second.
Can the NVIDIA B200 180GB run Qwen 3.6 27B?
Yes. The NVIDIA B200 180GB runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 47.9 tokens per second.
Can the NVIDIA B200 180GB run Llama 3.1 8B Instruct?
Yes. The NVIDIA B200 180GB runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 157.2 tokens per second.
Can I rent the NVIDIA B200 180GB instead of buying it?
Yes — RunPod and similar cloud GPU providers let you rent NVIDIA B200 180GB instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.