CanItRun Logocanitrun.

NVIDIA B300 288GB

The NVIDIA B300 288GB has 288 GB VRAM and 8000 GB/s memory bandwidth. It can run 74 of our 85 tracked models natively in VRAM at 8k context.

With 288 GB HBM3e, the NVIDIA B300 288GB is a datacenter-tier GPU that can run 74 models natively. It handles the largest open-weight models, including 405B-class frontier releases, at some quantization.

NVIDIA B300 288GB: Blackwell Ultra, NVIDIA's highest-capacity single-GPU datacenter accelerator, announced at GTC in March 2025 and shipping in HGX B300 NVL16 and GB300 NVL72 systems since the second half of 2025. 288GB of HBM3e — 60% more than the B200's 180GB, from 12-Hi memory stacks in place of the B200's 8-Hi stacks — at the same 8 TB/s per-GPU bandwidth and 1.8 TB/s NVLink 5. TDP rises to 1,400W, up from the B200's roughly 1,000W.

This site's calculator runs GPT-OSS 120B (MoE) at full BF16 (262.42 GB, 151.6 tok/s) and GLM 4.5 (355B, MoE) at NVFP4 (202.26 GB, 92.2 tok/s) natively — both need a lower quant, or don't fit at all, on the B200. Llama 3.1 405B, the largest dense model this site tracks, fits at NVFP4 (231.54 GB, 25.2 tok/s) with room to spare. DeepSeek V3 (671B, MoE) gets close but still needs CPU offload even at Q2_K (286.9 GB, 109.5 tok/s) — the 288GB ceiling covers everything this site tracks except the very largest frontier MoE releases. 74 of the 85 tracked models fit fully in VRAM at 8k context, more than any other GPU here.

Same CUDA Toolkit 12.8+ requirement as the B200 — both are sm_100 Blackwell parts, so a CUDA/driver stack that already supports the B200 supports the B300 unchanged. vLLM and TensorRT-LLM both support NVFP4 batched inference on Blackwell Ultra. Cloud-only for nearly every buyer, and the newest addition to on-demand GPU cloud pricing pages like RunPod's.

VendorNVIDIA
ArchitectureBlackwell
VRAM288 GB
Memory typeHBM3e
Memory bandwidth8000 GB/s
Compute backendCUDA
TierDatacenter
Released2025
Models (native)74 / 85
Models (offload)2 / 85
Software: Needs CUDA Toolkit 12.8 or newer — the first Toolkit release with Blackwell (sm_100) support. vLLM and TensorRT-LLM both have day-one Blackwell Ultra kernels for FP8/NVFP4 batched inference.

The B300 grew capacity 60% over the B200 without adding a byte of bandwidth

Plotting NVIDIA's last four datacenter flagships by VRAM and bandwidth together shows how unevenly the two numbers move — this generation's jump is capacity-only:

0425085000150300VRAM (GB)Bandwidth (GB/s)NVIDIA H100 80GBNVIDIA H200 141GBNVIDIA B200 180GBNVIDIA B300 288GB
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

B300 (this page) and B200 share the exact same 8,000 GB/s bandwidth — the entire generational upgrade is 108GB of extra HBM3e capacity, a 60% jump from the B200's 180GB to 288GB, achieved by stacking memory 12-Hi instead of 8-Hi rather than widening the memory bus. That's the opposite of the two steps before it: H100 to H200 was capacity-led (80GB to 141GB, +76%) on a smaller 43% bandwidth gain (3,350 to 4,800 GB/s), and H200 to B200 flipped to bandwidth-led (+67% to 8,000 GB/s) on a smaller 28% capacity gain. Since decode speed is bandwidth-bound, a model that already fit on a B200 decodes at the identical tok/s on a B300 — the B300's advantage is entirely about which larger models fit at all, not how fast the ones that already fit run.

Cloud GPU Rental

Don't want to buy a NVIDIA B300 288GB? RunPod is a cloud GPU rental service — rent one by the hour instead, no contract, no upfront hardware cost.

Pay by the hour · no contract · pods start in about a minute.

Rent a NVIDIA B300 288GB on RunPod ↗ (+$5 signup credit)

Affiliate link — CanItRun may earn a commission. Doesn't affect the fit calculation above.

Popular models for this GPU

Models this GPU runs natively in VRAM (74)

Show 69 more

Models that fit with CPU offload (2)

These use system RAM for layers that don't fit in VRAM — expect much slower inference.

Too large for this GPU (9)

Frequently asked questions

How much VRAM does the NVIDIA B300 288GB have?
The NVIDIA B300 288GB has 288 GB of HBM3e with 8000 GB/s memory bandwidth.
What is the NVIDIA B300 288GB best for?
With 288 GB of VRAM, the NVIDIA B300 288GB is a server-class GPU designed for running the largest open-weight models (70B–405B) at high quantization with ample context.
What LLMs can the NVIDIA B300 288GB run locally?
The NVIDIA B300 288GB can run 74 of the 85 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at BF16, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
Can the NVIDIA B300 288GB run Llama 3.3 70B Instruct?
Yes. The NVIDIA B300 288GB runs Llama 3.3 70B Instruct natively in VRAM at BF16 quantization, achieving approximately 36.4 tokens per second.
Can the NVIDIA B300 288GB run Qwen 3.6 27B?
Yes. The NVIDIA B300 288GB runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 47.9 tokens per second.
Can the NVIDIA B300 288GB run Llama 3.1 8B Instruct?
Yes. The NVIDIA B300 288GB runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 157.2 tokens per second.
Can I rent the NVIDIA B300 288GB instead of buying it?
Yes — RunPod and similar cloud GPU providers let you rent NVIDIA B300 288GB instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.