NVIDIA B300 288GB
The NVIDIA B300 288GB has 288 GB VRAM and 8000 GB/s memory bandwidth. It can run 74 of our 85 tracked models natively in VRAM at 8k context.
With 288 GB HBM3e, the NVIDIA B300 288GB is a datacenter-tier GPU that can run 74 models natively. It handles the largest open-weight models, including 405B-class frontier releases, at some quantization.
NVIDIA B300 288GB: Blackwell Ultra, NVIDIA's highest-capacity single-GPU datacenter accelerator, announced at GTC in March 2025 and shipping in HGX B300 NVL16 and GB300 NVL72 systems since the second half of 2025. 288GB of HBM3e — 60% more than the B200's 180GB, from 12-Hi memory stacks in place of the B200's 8-Hi stacks — at the same 8 TB/s per-GPU bandwidth and 1.8 TB/s NVLink 5. TDP rises to 1,400W, up from the B200's roughly 1,000W.
This site's calculator runs GPT-OSS 120B (MoE) at full BF16 (262.42 GB, 151.6 tok/s) and GLM 4.5 (355B, MoE) at NVFP4 (202.26 GB, 92.2 tok/s) natively — both need a lower quant, or don't fit at all, on the B200. Llama 3.1 405B, the largest dense model this site tracks, fits at NVFP4 (231.54 GB, 25.2 tok/s) with room to spare. DeepSeek V3 (671B, MoE) gets close but still needs CPU offload even at Q2_K (286.9 GB, 109.5 tok/s) — the 288GB ceiling covers everything this site tracks except the very largest frontier MoE releases. 74 of the 85 tracked models fit fully in VRAM at 8k context, more than any other GPU here.
Same CUDA Toolkit 12.8+ requirement as the B200 — both are sm_100 Blackwell parts, so a CUDA/driver stack that already supports the B200 supports the B300 unchanged. vLLM and TensorRT-LLM both support NVFP4 batched inference on Blackwell Ultra. Cloud-only for nearly every buyer, and the newest addition to on-demand GPU cloud pricing pages like RunPod's.
| Vendor | NVIDIA |
| Architecture | Blackwell |
| VRAM | 288 GB |
| Memory type | HBM3e |
| Memory bandwidth | 8000 GB/s |
| Compute backend | CUDA |
| Tier | Datacenter |
| Released | 2025 |
| Models (native) | 74 / 85 |
| Models (offload) | 2 / 85 |
The B300 grew capacity 60% over the B200 without adding a byte of bandwidth
Plotting NVIDIA's last four datacenter flagships by VRAM and bandwidth together shows how unevenly the two numbers move — this generation's jump is capacity-only:
B300 (this page) and B200 share the exact same 8,000 GB/s bandwidth — the entire generational upgrade is 108GB of extra HBM3e capacity, a 60% jump from the B200's 180GB to 288GB, achieved by stacking memory 12-Hi instead of 8-Hi rather than widening the memory bus. That's the opposite of the two steps before it: H100 to H200 was capacity-led (80GB to 141GB, +76%) on a smaller 43% bandwidth gain (3,350 to 4,800 GB/s), and H200 to B200 flipped to bandwidth-led (+67% to 8,000 GB/s) on a smaller 28% capacity gain. Since decode speed is bandwidth-bound, a model that already fit on a B200 decodes at the identical tok/s on a B300 — the B300's advantage is entirely about which larger models fit at all, not how fast the ones that already fit run.
Cloud GPU Rental
Don't want to buy a NVIDIA B300 288GB? RunPod is a cloud GPU rental service — rent one by the hour instead, no contract, no upfront hardware cost.
Pay by the hour · no contract · pods start in about a minute.
Rent a NVIDIA B300 288GB on RunPod ↗ (+$5 signup credit)Affiliate link — CanItRun may earn a commission. Doesn't affect the fit calculation above.
Popular models for this GPU
Models this GPU runs natively in VRAM (74)
- Nemotron 3 Ultra 550B-A55B550B · MMLU-Pro 86.8Q2_K · ~73.5 t/s
- MiniMax M1 456B456B · MMLU-Pro 81.1NVFP4 · ~65.5 t/s
- MiniMax M3428B · MMLU-Pro —NVFP4 · ~132.2 t/s
- Llama 3.1 405B Instruct405B · MMLU-Pro 73.3NVFP4 · ~25.2 t/s
- Llama 4 Maverick 400B400B · MMLU-Pro 80.5NVFP4 · ~160.7 t/s
Show 69 more
- GLM-4.7 358B358B · MMLU-Pro 84.3NVFP4 · ~92.2 t/s
- GLM-4.5 355B355B · MMLU-Pro 84.6NVFP4 · ~92.2 t/s
- GLM-4.6 355B355B · MMLU-Pro 83.2NVFP4 · ~92.2 t/s
- DeepSeek V4 Flash 284B284B · MMLU-Pro 86.3NVFP4 · ~239.2 t/s
- DeepSeek V4 Flash 0731 284B284B · MMLU-Pro —UD-Q8_K_XL · ~209.9 t/s
- Qwen3 235B-A22B (MoE)235B · MMLU-Pro 84.4NVFP4 · ~136 t/s
- MiniMax M2.5 229B229B · MMLU-Pro 84.8NVFP4 · ~277.4 t/s
- MiniMax M2.7 229B229B · MMLU-Pro 86.0NVFP4 · ~277.4 t/s
- Step 3.7 Flash198B · MMLU-Pro —NVFP4 · ~262.1 t/s
- Step 3.5 Flash196.81B · MMLU-Pro 84.4NVFP4 · ~262.1 t/s
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0NVFP4 · ~77.8 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro —NVFP4 · ~77.7 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7NVFP4 · ~276.4 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7BF16 · ~64.4 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7BF16 · ~151.6 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3BF16 · ~44.8 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4BF16 · ~63.8 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9BF16 · ~63.8 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1BF16 · ~35.5 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9BF16 · ~36.4 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0BF16 · ~36.4 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4BF16 · ~36.4 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7FP32 · ~30 t/s
- Command-R 35B35B · MMLU-Pro 33.0FP32 · ~34.5 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3FP32 · ~129.5 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2FP32 · ~36.6 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0FP32 · ~37.2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5FP32 · ~39.2 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0FP32 · ~39.3 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3FP32 · ~39.3 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0FP32 · ~39.3 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3FP32 · ~128.6 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2FP32 · ~40.9 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5FP32 · ~127.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0FP32 · ~46.5 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5FP32 · ~47.5 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2FP32 · ~47.9 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~590.2 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6FP32 · ~101.9 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8FP32 · ~53.4 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2FP32 · ~57.3 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9FP32 · ~107.9 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0FP32 · ~85.9 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7FP32 · ~86.1 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4FP32 · ~90.7 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6FP32 · ~103.7 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~104.3 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2FP32 · ~101.5 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~131.3 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~143.4 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~157.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~157.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~156.6 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~168.4 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~172.9 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~315.1 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~305.8 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~282.3 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~319.5 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~378.5 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~409.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~461.3 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~618.9 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~618.3 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~834 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~994.6 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~1201.7 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~2475.4 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~2928.7 t/s
Models that fit with CPU offload (2)
These use system RAM for layers that don't fit in VRAM — expect much slower inference.
Too large for this GPU (9)
Frequently asked questions
- How much VRAM does the NVIDIA B300 288GB have?
- The NVIDIA B300 288GB has 288 GB of HBM3e with 8000 GB/s memory bandwidth.
- What is the NVIDIA B300 288GB best for?
- With 288 GB of VRAM, the NVIDIA B300 288GB is a server-class GPU designed for running the largest open-weight models (70B–405B) at high quantization with ample context.
- What LLMs can the NVIDIA B300 288GB run locally?
- The NVIDIA B300 288GB can run 74 of the 85 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at BF16, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the NVIDIA B300 288GB run Llama 3.3 70B Instruct?
- Yes. The NVIDIA B300 288GB runs Llama 3.3 70B Instruct natively in VRAM at BF16 quantization, achieving approximately 36.4 tokens per second.
- Can the NVIDIA B300 288GB run Qwen 3.6 27B?
- Yes. The NVIDIA B300 288GB runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 47.9 tokens per second.
- Can the NVIDIA B300 288GB run Llama 3.1 8B Instruct?
- Yes. The NVIDIA B300 288GB runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 157.2 tokens per second.
- Can I rent the NVIDIA B300 288GB instead of buying it?
- Yes — RunPod and similar cloud GPU providers let you rent NVIDIA B300 288GB instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.