AMD Radeon AI PRO R9700 32GB

The AMD Radeon AI PRO R9700 32GB has 32 GB VRAM and 640 GB/s memory bandwidth. It can run 54 of our 99 tracked models natively in VRAM at 8k context.

With 32 GB GDDR6, the AMD Radeon AI PRO R9700 32GB is a workstation-tier GPU that can run 54 models natively. At 8k context, dense ~32-34B models (Qwen3 32B, Qwen2.5 32B, and Yi 1.5 34B) all fit natively at Q5_K_M (27.7-29.7 GB) at 15.7-16.8 tok/s, and reach 18.1-19.5 tok/s once stepped down to Q4_K_M. Qwen 3.6 27B fits even more comfortably, natively up to Q6_K (25.4 GB) at 18.3 tok/s. 8B-class dense models fit at full BF16 with room to spare (24.4 tok/s for Llama 3.1 8B, 26.5 tok/s for Qwen2.5 7B) and jump to 70-81.6 tok/s once quantized to Q4_K_M. 70B dense models don't fit natively at all: Llama 3.3 70B needs CPU offload even at Q4_K_M, dropping to roughly 1.5 tok/s. Sparse MoE models do far better on the same 32GB: Qwen3.5-35B-A3B fits natively at Q5_K_M (28.1 GB) and reaches 57.1 tok/s, since decode only has to read its ~3B active parameters per token instead of the full 35B.

The AMD Radeon AI PRO R9700 is a workstation-class AI accelerator built on RDNA 4 (Navi 48, the same die as the RX 9070 XT) with 32GB of GDDR6 at 640 GB/s across a 256-bit bus and 64MB of Infinity Cache. Its 64 compute units pack 4,096 stream processors and 128 second-generation AI Accelerators, which AMD rates at 47.8 TFLOPS of FP32 vector throughput and up to 191 TFLOPS of FP16 matrix throughput. Launched at a $1,299 MSRP, it targets single-GPU local inference and fine-tuning where ECC-validated datacenter hardware isn't required.

AMD Radeon AI PRO R9700 32GB: Officially the "Radeon AI PRO R9700": many retailers shorten it to "AI Pro 9700," but AMD's own model number keeps the "R." AMD announced it at Computex in May 2025, shipped it to workstation OEMs that July, and didn't bring it to general retail until October 27, 2025, a five-month gap between announcement and being buyable off the shelf. The 640 GB/s bandwidth (32GB of GDDR6 on a 256-bit bus) is the number every tok/s estimate on this page is built from; see how it stacks up against AMD's own W7800/W7900 below.

At 8k context, dense ~32-34B models (Qwen3 32B, Qwen2.5 32B, and Yi 1.5 34B) all fit natively at Q5_K_M (27.7-29.7 GB) at 15.7-16.8 tok/s, and reach 18.1-19.5 tok/s once stepped down to Q4_K_M. Qwen 3.6 27B fits even more comfortably, natively up to Q6_K (25.4 GB) at 18.3 tok/s. 8B-class dense models fit at full BF16 with room to spare (24.4 tok/s for Llama 3.1 8B, 26.5 tok/s for Qwen2.5 7B) and jump to 70-81.6 tok/s once quantized to Q4_K_M. 70B dense models don't fit natively at all: Llama 3.3 70B needs CPU offload even at Q4_K_M, dropping to roughly 1.5 tok/s. Sparse MoE models do far better on the same 32GB: Qwen3.5-35B-A3B fits natively at Q5_K_M (28.1 GB) and reaches 57.1 tok/s, since decode only has to read its ~3B active parameters per token instead of the full 35B.

This is AMD's newest architecture in the lineup, so its ROCm support is the least battle-tested of any card on this site; see the Software note in the spec table below for the exact version and OS requirements. On Linux, Phoronix's ROCm 7.0 testing reported the R9700 beating the older Radeon PRO W7900 by roughly 47% on vLLM throughput, and a dual-card setup edging out a single RTX 6000 Ada, reviewer-reported figures, not verified by this site. Public llama.cpp benchmarks on GitHub report roughly 127-163 tok/s decode on Qwen3.5-35B-A3B at short context via Vulkan/ROCm, not directly comparable to the 8k-context estimate above, but in the same range.

VendorAMD
ArchitectureRDNA 4
VRAM32 GB
Memory typeGDDR6
Memory bandwidth640 GB/s
Compute backendROCM
TierWorkstation
Released2025
Models (native)54 / 99
Models (offload)9 / 99
Software: Requires ROCm 6.4.1 or newer (7.x recommended); this RDNA 4 chip (gfx1201) has no precompiled kernels in older ROCm builds. ROCm 7.0.x is validated only on Ubuntu 24.04.3, Ubuntu 22.04.5, or RHEL 9.6; other distros and Windows should use the Vulkan backend in llama.cpp instead, since AMD's Windows HIP builds have also shipped without gfx1201 kernels.

Where the AI Pro 9700 lands in AMD's own workstation stack

AMD sells three single-GPU workstation cards spanning 32-48GB, two architecture generations apart. Here's how the AI Pro 9700's bandwidth compares to its RDNA 3 stablemates:

AMD Radeon PRO W7800
576 GB/s
AMD Radeon AI PRO R9700 32GB (this page)
640 GB/s
AMD Radeon PRO W7900
864 GB/s

At 640 GB/s, the AI Pro 9700 has about 11% more bandwidth than the same-capacity 32GB Radeon PRO W7800 (576 GB/s), while launching at roughly half the W7800's $2,499 MSRP, since RDNA 4's newer memory controllers outrun RDNA 3's at a lower price. It still trails the 48GB Radeon PRO W7900's 864 GB/s by about 35%, and that extra 16GB of VRAM is what actually matters most: the W7900 fits Llama 3.3 70B natively at Q3_K_M (40.7 GB, 15.4 tok/s), where the 32GB 9700 needs CPU offload for any 70B model. Since decode speed is bandwidth-bound, the AI Pro 9700 reads as a faster, cheaper W7800 rather than a budget W7900; the bandwidth bump helps every model that already fits in 32GB, but it can't substitute for the 16GB of extra capacity that unlocks 70B-class models.

Popular models for this GPU

Models this GPU runs natively in VRAM (54)

Show 49 more

Models that fit with CPU offload (9)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (36)

Compare AMD Radeon AI PRO R9700 32GB with other GPUs

Frequently asked questions

How much VRAM does the AMD Radeon AI PRO R9700 32GB have?
The AMD Radeon AI PRO R9700 32GB has 32 GB of GDDR6 with 640 GB/s memory bandwidth.
What is the AMD Radeon AI PRO R9700 32GB best for?
With 32 GB of VRAM, the AMD Radeon AI PRO R9700 32GB is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
What LLMs can the AMD Radeon AI PRO R9700 32GB run locally?
The AMD Radeon AI PRO R9700 32GB can run 54 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q6_K, Ornith 1.5 35B-A3B (MoE) at Q5_K_M, Ornith 1.5 9B at BF16.
Can the AMD Radeon AI PRO R9700 32GB run Gemma 4 31B?
Yes. The AMD Radeon AI PRO R9700 32GB runs Gemma 4 31B natively in VRAM at Q6_K quantization, achieving approximately 15.6 tokens per second.
Can the AMD Radeon AI PRO R9700 32GB run Qwen 3.6 27B?
Yes. The AMD Radeon AI PRO R9700 32GB runs Qwen 3.6 27B natively in VRAM at Q6_K quantization, achieving approximately 18.3 tokens per second.
Can the AMD Radeon AI PRO R9700 32GB run Qwen3 8B?
Yes. The AMD Radeon AI PRO R9700 32GB runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 24.2 tokens per second.