CanItRun Logocanitrun.

NVIDIA H200 141GB

The NVIDIA H200 141GB has 141 GB VRAM and 4800 GB/s memory bandwidth. It can run 66 of our 85 tracked models natively in VRAM at 8k context.

With 141 GB HBM3e, the NVIDIA H200 141GB is a datacenter-tier GPU that can run 66 models natively. It runs 70B-class dense models and large MoE models entirely in VRAM.

NVIDIA H200 141GB: Announced November 13, 2023 at SC23 and shipping since Q2 2024, the H200 keeps the H100's GH100 Hopper die but swaps in 141GB of HBM3e at 4.8 TB/s — nearly double the H100's 80GB capacity and 1.4x its 3,350 GB/s bandwidth (NVIDIA's own comparison figures). Ships as an SXM module (up to 700W) or a PCIe-based NVL card (up to 600W).

This site's calculator puts Llama 3.3 70B at Q8_0 (86.35 GB, 40.5 tok/s) and GPT-OSS 120B (MoE) at Q6_K (107.93 GB, 218.7 tok/s) — both fit natively in the 141GB pool. GLM 4.5 (355B, MoE) is the real ceiling: even Q2_K needs CPU offload at 154.94 GB, decoding at 14.5 tok/s. Across the 85 models this site tracks, the H200 runs 66 of them fully in VRAM at 8k context — 7 more than the H100 80GB's 59, purely from the extra memory headroom.

Full CUDA support from day one — same GH100 die and sm_90 compute capability as the H100, so no toolkit or driver upgrade is needed to recognize it, though server vendors typically ship a BIOS/firmware update for the higher-capacity HBM3e stacks. vLLM and TensorRT-LLM both treat it as a drop-in H100 replacement with more headroom for KV cache and larger batches. Cloud-only for nearly everyone — available on AWS, Azure, GCP, OCI, CoreWeave, Lambda, and other GPU clouds by the hour.

VendorNVIDIA
ArchitectureHopper
VRAM141 GB
Memory typeHBM3e
Memory bandwidth4800 GB/s
Compute backendCUDA
TierDatacenter
Released2024
Models (native)66 / 85
Models (offload)3 / 85
Software: Typically cloud-accessed. vLLM and TensorRT-LLM give best batched-inference performance.

Memory capacity and bandwidth don't climb together, generation to generation

Plotting NVIDIA's last four datacenter flagships by VRAM and bandwidth together shows how unevenly the two numbers actually move — there's no single "the generational upgrade" pattern:

0425085000150300VRAM (GB)Bandwidth (GB/s)NVIDIA H100 80GBNVIDIA H200 141GBNVIDIA B200 180GBNVIDIA B300 288GB
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

H100 to H200 (this page) is a capacity-led jump — 80GB to 141GB, up 76%, on a smaller 43% bandwidth gain (3,350 to 4,800 GB/s). The next step, H200 to B200, flips that: capacity rises a smaller 28% but bandwidth jumps 67% to 8,000 GB/s. B200 to B300 flips again — 60% more capacity (180GB to 288GB) at the exact same 8,000 GB/s bandwidth. Each step trades differently between how much a GPU can hold and how fast it can read it, so "newer generation" alone doesn't tell you whether a model that didn't fit the last one now will, or whether one that already fit will just run faster.

Cloud GPU Rental

Don't want to buy a NVIDIA H200 141GB? RunPod is a cloud GPU rental service — rent one by the hour instead, no contract, no upfront hardware cost.

Pay by the hour · no contract · pods start in about a minute.

Rent a NVIDIA H200 141GB on RunPod ↗ (+$5 signup credit)

Affiliate link — CanItRun may earn a commission. Doesn't affect the fit calculation above.

Popular models for this GPU

Models this GPU runs natively in VRAM (66)

Show 61 more

Models that fit with CPU offload (3)

These use system RAM for layers that don't fit in VRAM — expect much slower inference.

Too large for this GPU (16)

Frequently asked questions

How much VRAM does the NVIDIA H200 141GB have?
The NVIDIA H200 141GB has 141 GB of HBM3e with 4800 GB/s memory bandwidth.
What is the NVIDIA H200 141GB best for?
With 141 GB of VRAM, the NVIDIA H200 141GB is a server-class GPU that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
What LLMs can the NVIDIA H200 141GB run locally?
The NVIDIA H200 141GB can run 66 of the 85 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q8_0, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
Can the NVIDIA H200 141GB run Llama 3.3 70B Instruct?
Yes. The NVIDIA H200 141GB runs Llama 3.3 70B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 40.5 tokens per second.
Can the NVIDIA H200 141GB run Qwen 3.6 27B?
Yes. The NVIDIA H200 141GB runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 28.7 tokens per second.
Can the NVIDIA H200 141GB run Llama 3.1 8B Instruct?
Yes. The NVIDIA H200 141GB runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 94.3 tokens per second.
Can I rent the NVIDIA H200 141GB instead of buying it?
Yes — RunPod and similar cloud GPU providers let you rent NVIDIA H200 141GB instances by the hour, with no long-term contract. This is often cheaper than buying if you only need it occasionally, and lets you try the GPU before committing to a purchase.