AMD Strix Halo (96GB)

The AMD Strix Halo (96GB) has 96 GB VRAM and 256 GB/s memory bandwidth. It can run 68 of our 97 tracked models natively in VRAM at 8k context.

With 96 GB LPDDR5X, the AMD Strix Halo (96GB) is a laptop-tier GPU that can run 68 models natively. This site's calculator puts Llama 3.1 8B at its recommended Q5_K_M (7.58 GB) at 24.6 tok/s and Qwen3 8B at its recommended Q5_K_M (7.73 GB) at 24.1 tok/s, identical to the 64GB and 128GB tiers since all three share the same 256 GB/s bandwidth. Independent Vulkan/RADV community benchmarks report closer to 45 tok/s for a similar 7B Q4 build; this site's decode-efficiency constant was calibrated against a discrete RTX 4090, not this unified-memory architecture, so treat the figures here as conservative. 96GB is the tier where large MoE and dense models genuinely open up: this site's own top-of-board picks include Mistral Medium 3.5 128B, a 128B dense model, fitting at Q3_K_M (72.26 GB, 2.6 tok/s), GPT-OSS 120B fitting at Q4 (80.15 GB, 15.6 tok/s, the smallest tier where it fits at all), and Qwen 3.5 122B-A10B, a 122B MoE release, at Q4_K_M (83.44 GB, 8.1 tok/s). More conventional 70B-class dense models fit comfortably too: Llama 3.3 70B reaches Q4_K_M (50.75 GB, 3.7 tok/s), and Qwen 2.5 72B lands close behind (52.12 GB, 3.6 tok/s). 68 of the 97 models this site tracks fit natively at this capacity, six more than the 64GB tier, four fewer than 128GB. None of those tok/s figures capture this platform's most-discussed real-world weakness, though: prompt processing. Strix Halo's iGPU is bandwidth-bound rather than compute-bound, so long prompts are genuinely slow to chew through before the first output token appears. Independent benchmarking (datahardware.ai) measured GPT-OSS 120B prefilling at only about 340 tok/s on this chip, roughly a fifth of the ~1,700 tok/s NVIDIA's compute-bound DGX Spark reaches on the identical model, and real document text prefills 24-33% slower still than the synthetic prompts most benchmarks use. The practical effect: a 12,000-token prompt needs roughly 35 seconds of processing before generation even starts, versus about 7 seconds on DGX Spark. Decode itself keeps sliding well past this page's 8k-context estimate too: one independent long-context benchmark found generation speed dropping by roughly two-thirds once the KV cache filled to around 76k tokens.

AMD Strix Halo (96GB): AMD unveiled Strix Halo (retail name: Ryzen AI Max 300 series) at CES on January 6, 2025. This 96GB tier is a genuinely distinct piece of hardware, not just a BIOS setting on a 128GB unit: AMD's Ryzen AI Max+ 395 supports a Variable Graphics Memory split that can dedicate up to 96GB of a 128GB system to the GPU, but vendors like GMKtec (EVO-X2) and X+ (RIVAL) also sell 96GB as its own soldered memory configuration, four LPDDR5X-8000 channels populated with 24GB modules instead of 32GB or 16GB ones. Either way the chip behind it is the full Ryzen AI Max+ 395 (16 Zen 5 cores, Radeon 8060S iGPU, 40 RDNA 3.5 compute units) and the same 256 GB/s theoretical bandwidth as every other tier this site tracks, quad-channel LPDDR5X-8000 over a 256-bit bus, unaffected by how much of the pool is set aside for the GPU. Independent community bandwidth testing (a Level1Techs forum benchmark thread) measures around 215 GB/s actually achieved in practice, about 84% of that ceiling. This tier's pricing has moved with the rest of the lineup: GMKtec's EVO-X2 96GB has listed around $2,349, part of the same 2026 LPDDR5X/DRAM shortage pushing unified-memory AI-PC pricing up broadly (compute-market.com cites the identical shortage behind NVIDIA's own DGX Spark listing around $4,699 in mid-2026), not a price this specific configuration set on its own.

This site's calculator puts Llama 3.1 8B at its recommended Q5_K_M (7.58 GB) at 24.6 tok/s and Qwen3 8B at its recommended Q5_K_M (7.73 GB) at 24.1 tok/s, identical to the 64GB and 128GB tiers since all three share the same 256 GB/s bandwidth. Independent Vulkan/RADV community benchmarks report closer to 45 tok/s for a similar 7B Q4 build; this site's decode-efficiency constant was calibrated against a discrete RTX 4090, not this unified-memory architecture, so treat the figures here as conservative. 96GB is the tier where large MoE and dense models genuinely open up: this site's own top-of-board picks include Mistral Medium 3.5 128B, a 128B dense model, fitting at Q3_K_M (72.26 GB, 2.6 tok/s), GPT-OSS 120B fitting at Q4 (80.15 GB, 15.6 tok/s, the smallest tier where it fits at all), and Qwen 3.5 122B-A10B, a 122B MoE release, at Q4_K_M (83.44 GB, 8.1 tok/s). More conventional 70B-class dense models fit comfortably too: Llama 3.3 70B reaches Q4_K_M (50.75 GB, 3.7 tok/s), and Qwen 2.5 72B lands close behind (52.12 GB, 3.6 tok/s). 68 of the 97 models this site tracks fit natively at this capacity, six more than the 64GB tier, four fewer than 128GB. None of those tok/s figures capture this platform's most-discussed real-world weakness, though: prompt processing. Strix Halo's iGPU is bandwidth-bound rather than compute-bound, so long prompts are genuinely slow to chew through before the first output token appears. Independent benchmarking (datahardware.ai) measured GPT-OSS 120B prefilling at only about 340 tok/s on this chip, roughly a fifth of the ~1,700 tok/s NVIDIA's compute-bound DGX Spark reaches on the identical model, and real document text prefills 24-33% slower still than the synthetic prompts most benchmarks use. The practical effect: a 12,000-token prompt needs roughly 35 seconds of processing before generation even starts, versus about 7 seconds on DGX Spark. Decode itself keeps sliding well past this page's 8k-context estimate too: one independent long-context benchmark found generation speed dropping by roughly two-thirds once the KV cache filled to around 76k tokens.

Vulkan (RADV or AMDVLK) via llama.cpp works everywhere; ROCm 6.4+ targets this chip directly as gfx1151 on Linux. Independent backend testing (soothill.io, kyuz0's toolboxes benchmarks) found no single fastest backend: ROCm usually wins prompt processing, Vulkan usually wins token generation. If your specific machine actually ships 128GB total with a BIOS-configurable split rather than 96GB soldered, AMD's Adrenalin driver on Windows exposes the Variable Graphics Memory setting directly; on Linux the equivalent is a small UMA carve-out plus the amdgpu.gttsize kernel parameter.

VendorAMD
ArchitectureRDNA 3.5
CPU cores16-core Zen 5 (Ryzen AI Max+ 395), Radeon 8060S iGPU (40 CUs)
VRAM96 GB (unified)
Memory typeLPDDR5X
Memory bandwidth256 GB/s
Compute backendVULKAN
TierLaptop
Released2025
Models (native)68 / 97
Models (offload)0 / 97
Software: Vulkan (RADV or AMDVLK) via llama.cpp works cross-platform. ROCm 6.4+ on Linux targets this chip as gfx1151; independent benchmarks show no single fastest backend, ROCm usually wins prompt processing, Vulkan usually wins token generation.

Strix Halo's bandwidth doesn't change with capacity, but its peers' does

All four Strix Halo memory tiers this site tracks, 32GB through 128GB, share the exact same LPDDR5X-8000 memory subsystem. Plotted against two other unified-memory systems at similar capacities, the pattern is a flat line where a discrete GPU's would slope upward with price:

0425850070140VRAM (GB)Bandwidth (GB/s)AMD Strix Halo (32GB)AMD Strix Halo (64GB)AMD Strix Halo (96GB)AMD Strix Halo (128GB)NVIDIA DGX Spark (128GB)Apple M3 Ultra (96GB)
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

Every Strix Halo point sits at exactly 256 GB/s regardless of capacity, a 4x range in GB with zero change in bandwidth. Apple's M3 Ultra shares this page's exact 96GB capacity but reaches 819 GB/s, roughly 3.2x Strix Halo's bandwidth, at the identical GB figure. NVIDIA's DGX Spark, at a larger 128GB capacity, only edges 17 GB/s ahead of this page's 256 GB/s (273 GB/s, about 6.6% more) despite the extra 32GB. Since LLM decode is bandwidth-bound, buying a bigger Strix Halo unit buys headroom for larger models, not a faster ceiling for the ones that already fit on a smaller one: this site's calculator returns the identical 24.6 tok/s for Llama 3.1 8B on every Strix Halo tier from 32GB to 128GB.

Popular models for this GPU

Models this GPU runs natively in VRAM (68)

Show 63 more

Too large for this GPU (29)

Compare AMD Strix Halo (96GB) with other GPUs

Frequently asked questions

How much VRAM does the AMD Strix Halo (96GB) have?
The AMD Strix Halo (96GB) has 96 GB of LPDDR5X with 256 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
What is the AMD Strix Halo (96GB) best for?
With 96 GB of unified memory, the AMD Strix Halo (96GB) is a high-capacity laptop platform that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
What LLMs can the AMD Strix Halo (96GB) run locally?
The AMD Strix Halo (96GB) can run 68 of the 97 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen3.8-Flash-Next at Q2_K, Qwen 3.8 27B at BF16, Ornith 1.5 35B-A3B (MoE) at BF16.
Can the AMD Strix Halo (96GB) run Gemma 4 31B?
Yes. The AMD Strix Halo (96GB) runs Gemma 4 31B natively in VRAM at BF16 quantization, achieving approximately 2.6 tokens per second.
Can the AMD Strix Halo (96GB) run Qwen 3.6 27B?
Yes. The AMD Strix Halo (96GB) runs Qwen 3.6 27B natively in VRAM at BF16 quantization, achieving approximately 3.1 tokens per second.
Can the AMD Strix Halo (96GB) run Qwen3 8B?
Yes. The AMD Strix Halo (96GB) runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 9.7 tokens per second.