NVIDIA H100 80GB

The NVIDIA H100 80GB has 80 GB VRAM and 3350 GB/s memory bandwidth. It can run 65 of our 97 tracked models natively in VRAM at 8k context.

With 80 GB HBM3, the NVIDIA H100 80GB is a datacenter-tier GPU that can run 65 models natively. Runs 70B at Q4 native with headroom to spare in 80GB, though the largest frontier-class models need a bigger card or a multi-GPU pool. ~50-80 t/s for 7B, ~20-30 t/s for 70B Q4.

The NVIDIA H100 80GB is the Hopper-generation datacenter GPU with 80GB HBM3 at an industry-leading 3,350 GB/s bandwidth. With FP8 Transformer Engine and 4th-gen NVLink, it delivers the fastest single-GPU LLM inference available: running 70B models at Q8 and 405B models across multi-GPU NVLink clusters. The de facto standard for production LLM serving.

NVIDIA H100 80GB: March 2022 Hopper architecture with 80GB HBM3 at 3350 GB/s, datacenter flagship.

Runs 70B at Q4 native with headroom to spare in 80GB, though the largest frontier-class models need a bigger card or a multi-GPU pool. ~50-80 t/s for 7B, ~20-30 t/s for 70B Q4.

Best-in-class throughput with vLLM and TensorRT-LLM. Excellent multi-GPU NVLink scaling. Cloud-only for most users.

VendorNVIDIA
ArchitectureHopper
VRAM80 GB
Memory typeHBM3
Memory bandwidth3350 GB/s
Compute backendCUDA
TierDatacenter
Released2022
Models (native)65 / 97
Models (offload)6 / 97
Software: Best-in-class inference throughput. vLLM and TensorRT-LLM are recommended; excellent multi-GPU NVLink scaling.

Cloud GPU Rental

Don't want to buy a NVIDIA H100 80GB? RunPod is a cloud GPU rental service: rent one by the hour instead, no contract, no upfront hardware cost.

Pay by the hour · no contract · pods start in about a minute.

Rent a NVIDIA H100 80GB on RunPod ↗ (+$5 signup credit)

Affiliate link: CanItRun may earn a commission. Doesn't affect the fit calculation above.

Popular models for this GPU

Models this GPU runs natively in VRAM (65)

Show 60 more

Models that fit with CPU offload (6)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (26)

Compare NVIDIA H100 80GB with other GPUs

Frequently asked questions

How much VRAM does the NVIDIA H100 80GB have?
The NVIDIA H100 80GB has 80 GB of HBM3 with 3350 GB/s memory bandwidth.
What is the NVIDIA H100 80GB best for?
With 80 GB of VRAM, the NVIDIA H100 80GB is a server-class GPU that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
What LLMs can the NVIDIA H100 80GB run locally?
The NVIDIA H100 80GB can run 65 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at BF16, Ornith 1.5 35B-A3B (MoE) at Q8_0, Qwen 3.5 122B-A10B (MoE) at Q3_K_M.
Can the NVIDIA H100 80GB run Gemma 4 31B?
Yes. The NVIDIA H100 80GB runs Gemma 4 31B natively in VRAM at BF16 quantization, achieving approximately 34.6 tokens per second.
Can the NVIDIA H100 80GB run Qwen 3.6 27B?
Yes. The NVIDIA H100 80GB runs Qwen 3.6 27B natively in VRAM at BF16 quantization, achieving approximately 39.9 tokens per second.
Can the NVIDIA H100 80GB run Qwen3 8B?
Yes. The NVIDIA H100 80GB runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 126.5 tokens per second.
Can I rent the NVIDIA H100 80GB instead of buying it?
Yes: RunPod and similar cloud GPU providers let you rent NVIDIA H100 80GB instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.