AMD Radeon AI PRO R9700 32GB
The AMD Radeon AI PRO R9700 32GB has 32 GB VRAM and 640 GB/s memory bandwidth. It can run 54 of our 99 tracked models natively in VRAM at 8k context.
With 32 GB GDDR6, the AMD Radeon AI PRO R9700 32GB is a workstation-tier GPU that can run 54 models natively. At 8k context, dense ~32-34B models (Qwen3 32B, Qwen2.5 32B, and Yi 1.5 34B) all fit natively at Q5_K_M (27.7-29.7 GB) at 15.7-16.8 tok/s, and reach 18.1-19.5 tok/s once stepped down to Q4_K_M. Qwen 3.6 27B fits even more comfortably, natively up to Q6_K (25.4 GB) at 18.3 tok/s. 8B-class dense models fit at full BF16 with room to spare (24.4 tok/s for Llama 3.1 8B, 26.5 tok/s for Qwen2.5 7B) and jump to 70-81.6 tok/s once quantized to Q4_K_M. 70B dense models don't fit natively at all: Llama 3.3 70B needs CPU offload even at Q4_K_M, dropping to roughly 1.5 tok/s. Sparse MoE models do far better on the same 32GB: Qwen3.5-35B-A3B fits natively at Q5_K_M (28.1 GB) and reaches 57.1 tok/s, since decode only has to read its ~3B active parameters per token instead of the full 35B.
The AMD Radeon AI PRO R9700 is a workstation-class AI accelerator built on RDNA 4 (Navi 48, the same die as the RX 9070 XT) with 32GB of GDDR6 at 640 GB/s across a 256-bit bus and 64MB of Infinity Cache. Its 64 compute units pack 4,096 stream processors and 128 second-generation AI Accelerators, which AMD rates at 47.8 TFLOPS of FP32 vector throughput and up to 191 TFLOPS of FP16 matrix throughput. Launched at a $1,299 MSRP, it targets single-GPU local inference and fine-tuning where ECC-validated datacenter hardware isn't required.
AMD Radeon AI PRO R9700 32GB: Officially the "Radeon AI PRO R9700": many retailers shorten it to "AI Pro 9700," but AMD's own model number keeps the "R." AMD announced it at Computex in May 2025, shipped it to workstation OEMs that July, and didn't bring it to general retail until October 27, 2025, a five-month gap between announcement and being buyable off the shelf. The 640 GB/s bandwidth (32GB of GDDR6 on a 256-bit bus) is the number every tok/s estimate on this page is built from; see how it stacks up against AMD's own W7800/W7900 below.
At 8k context, dense ~32-34B models (Qwen3 32B, Qwen2.5 32B, and Yi 1.5 34B) all fit natively at Q5_K_M (27.7-29.7 GB) at 15.7-16.8 tok/s, and reach 18.1-19.5 tok/s once stepped down to Q4_K_M. Qwen 3.6 27B fits even more comfortably, natively up to Q6_K (25.4 GB) at 18.3 tok/s. 8B-class dense models fit at full BF16 with room to spare (24.4 tok/s for Llama 3.1 8B, 26.5 tok/s for Qwen2.5 7B) and jump to 70-81.6 tok/s once quantized to Q4_K_M. 70B dense models don't fit natively at all: Llama 3.3 70B needs CPU offload even at Q4_K_M, dropping to roughly 1.5 tok/s. Sparse MoE models do far better on the same 32GB: Qwen3.5-35B-A3B fits natively at Q5_K_M (28.1 GB) and reaches 57.1 tok/s, since decode only has to read its ~3B active parameters per token instead of the full 35B.
This is AMD's newest architecture in the lineup, so its ROCm support is the least battle-tested of any card on this site; see the Software note in the spec table below for the exact version and OS requirements. On Linux, Phoronix's ROCm 7.0 testing reported the R9700 beating the older Radeon PRO W7900 by roughly 47% on vLLM throughput, and a dual-card setup edging out a single RTX 6000 Ada, reviewer-reported figures, not verified by this site. Public llama.cpp benchmarks on GitHub report roughly 127-163 tok/s decode on Qwen3.5-35B-A3B at short context via Vulkan/ROCm, not directly comparable to the 8k-context estimate above, but in the same range.
| Vendor | AMD |
| Architecture | RDNA 4 |
| VRAM | 32 GB |
| Memory type | GDDR6 |
| Memory bandwidth | 640 GB/s |
| Compute backend | ROCM |
| Tier | Workstation |
| Released | 2025 |
| Models (native) | 54 / 99 |
| Models (offload) | 9 / 99 |
Where the AI Pro 9700 lands in AMD's own workstation stack
AMD sells three single-GPU workstation cards spanning 32-48GB, two architecture generations apart. Here's how the AI Pro 9700's bandwidth compares to its RDNA 3 stablemates:
At 640 GB/s, the AI Pro 9700 has about 11% more bandwidth than the same-capacity 32GB Radeon PRO W7800 (576 GB/s), while launching at roughly half the W7800's $2,499 MSRP, since RDNA 4's newer memory controllers outrun RDNA 3's at a lower price. It still trails the 48GB Radeon PRO W7900's 864 GB/s by about 35%, and that extra 16GB of VRAM is what actually matters most: the W7900 fits Llama 3.3 70B natively at Q3_K_M (40.7 GB, 15.4 tok/s), where the 32GB 9700 needs CPU offload for any 70B model. Since decode speed is bandwidth-bound, the AI Pro 9700 reads as a faster, cheaper W7800 rather than a budget W7900; the bandwidth bump helps every model that already fits in 32GB, but it can't substitute for the 16GB of extra capacity that unlocks 70B-class models.
Popular models for this GPU
Models this GPU runs natively in VRAM (54)
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q3_K_M · ~19.1 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q2_K · ~17.3 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q5_K_M · ~57.1 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q5_K_M · ~15.4 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/AQ5_K_M · ~57.1 t/s
Show 49 more
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q5_K_M · ~15.7 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q5_K_M · ~16.8 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q5_K_M · ~16.5 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q5_K_M · ~16.5 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q5_K_M · ~16.5 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q6_K · ~48.1 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q6_K · ~15.6 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q6_K · ~46.1 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q6_K · ~50.4 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/AQ6_K · ~18.1 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q6_K · ~16.4 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q6_K · ~17.5 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q6_K · ~18.3 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2Q6_K · ~18.3 t/s
- Bonsai 27B27B · MMLU-Pro ~81.5Ternary (Q2_0) · ~53.8 t/s
- Bonsai 2 27B27B · MMLU-Pro N/ATernary (Q2_0) · ~64.6 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/AQ6_K · ~18.3 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q6_K · ~38.6 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q8_0 · ~15.5 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q8_0 · ~16.3 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~32.1 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q8_0 · ~24.4 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q8_0 · ~24.1 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q8_0 · ~25.6 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~16.2 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~16.3 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q8_0 · ~26 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~19.6 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~22.8 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ABF16 · ~22.8 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~24.4 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~24.4 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~24.2 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~26.5 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~26.7 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~48.9 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~46.2 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~38.4 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~48 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~56.7 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~64 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~68.5 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~94.5 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~83 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~128.6 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~151.4 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~178.8 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~378 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~394.1 t/s
Models that fit with CPU offload (9)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q2_K · ~5 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q2_K · ~4.1 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q2_K · ~2.9 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q2_K · ~4.7 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q2_K · ~4.7 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q4_K_M · ~1.4 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q4_K_M · ~1.5 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q4_K_M · ~1.5 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q4_K_M · ~1.5 t/s
Too large for this GPU (36)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.6 355B
- GLM-4.7 358B
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- Qwen3.8-Flash-Next
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
- DeepSeek V4.1 Flash 552B
Compare AMD Radeon AI PRO R9700 32GB with other GPUs
Frequently asked questions
- How much VRAM does the AMD Radeon AI PRO R9700 32GB have?
- The AMD Radeon AI PRO R9700 32GB has 32 GB of GDDR6 with 640 GB/s memory bandwidth.
- What is the AMD Radeon AI PRO R9700 32GB best for?
- With 32 GB of VRAM, the AMD Radeon AI PRO R9700 32GB is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the AMD Radeon AI PRO R9700 32GB run locally?
- The AMD Radeon AI PRO R9700 32GB can run 54 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q6_K, Ornith 1.5 35B-A3B (MoE) at Q5_K_M, Ornith 1.5 9B at BF16.
- Can the AMD Radeon AI PRO R9700 32GB run Gemma 4 31B?
- Yes. The AMD Radeon AI PRO R9700 32GB runs Gemma 4 31B natively in VRAM at Q6_K quantization, achieving approximately 15.6 tokens per second.
- Can the AMD Radeon AI PRO R9700 32GB run Qwen 3.6 27B?
- Yes. The AMD Radeon AI PRO R9700 32GB runs Qwen 3.6 27B natively in VRAM at Q6_K quantization, achieving approximately 18.3 tokens per second.
- Can the AMD Radeon AI PRO R9700 32GB run Qwen3 8B?
- Yes. The AMD Radeon AI PRO R9700 32GB runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 24.2 tokens per second.