Intel Arc Pro B70 32GB
The Intel Arc Pro B70 32GB has 32 GB VRAM and 608 GB/s memory bandwidth. It can run 53 of our 94 tracked models natively in VRAM at 8k context.
With 32 GB GDDR6, the Intel Arc Pro B70 32GB is a workstation-tier GPU that can run 53 models natively. It comfortably runs 7B–32B models at Q4; 70B-class models typically need CPU offload.
The Intel Arc Pro B70 launched March 25, 2026 at $949 as Intel's first "Big Battlemage" card, the full-die BMG-G31 workstation GPU Intel had been previewing since CES 2025. Where the original CES announcement pointed at a 24GB card, the shipping product doubled down: 32GB of GDDR6 on a 256-bit bus (608 GB/s), 32 Xe cores, 256 XMX engines rated at 367 INT8 TOPS, and a 230W board power target. That 32GB framebuffer runs 7-8B models at full BF16 precision, 14B models at near-lossless Q8, and 27-31B dense models at Q6, with the Vulkan backend providing usable LLM inference speeds on both Linux and Windows.
Intel Arc Pro B70 32GB: sits above Intel's existing Arc Pro B60 in the same Battlemage Pro lineup, which shipped first, in 2025, at 24GB of GDDR6 and 380 GB/s. This card followed as the first "Big Battlemage" part on the larger, full BMG-G31 die: 8GB more capacity and 60% more bandwidth than the B60 (608 vs 380 GB/s), launched March 25, 2026 at $949. 32 Xe cores and 256 XMX engines are rated at 367 INT8 TOPS, in a 230W board power envelope on PCIe 5.0.
53 of this site's 87 tracked models fit natively in VRAM at 8k context. Llama 3.1 8B and Qwen 2.5 7B are the only two that run at full BF16 precision (19.12 GB at 23.1 tok/s, 17.55 GB at 25.2 tok/s); Qwen3 14B is the largest to reach near-lossless Q8 (19.12 GB, 23.1 tok/s). Qwen 3.6 27B is the biggest dense model with real headroom to spare, fitting Q6 at 25.43 GB and 17.4 tok/s, well under this card's 30.4 GB usable ceiling.
Phoronix's Linux review found it working well on the fully open-source Mesa/Xe driver stack, with llama.cpp's SYCL backend and an April 2026 preview OpenVINO backend both usable alongside the default Vulkan path. Independent community benchmarks put Qwen 3.6 27B at roughly 14 to 22 tok/s depending on quantization, bracketing this site's own 17.4 to 23.3 tok/s Q6/Q4 estimates well. ISV-certified for professional workloads on Linux and Windows.
| Vendor | Intel |
| Architecture | Xe2-HPG (Battlemage) |
| VRAM | 32 GB |
| Memory type | GDDR6 |
| Memory bandwidth | 608 GB/s |
| Compute backend | VULKAN |
| Tier | Workstation |
| Released | 2026 |
| Models (native) | 53 / 94 |
| Models (offload) | 9 / 94 |
Three vendors land on the exact same 32GB ceiling, at three very different prices
Plotting this card against its closest capacity peers and its nearest price-tier neighbor shows something the spec sheet alone doesn't: 32GB isn't a niche Intel number here, it's a ceiling three different vendors converge on from three very different starting points.
This card (this page), the $1,299 Radeon AI PRO R9700, and the $1,999 RTX 5090 all sit at exactly 32GB, and because VRAM capacity, not bandwidth, decides what fits, all three run the identical set of 53 of this site's 87 tracked models natively at 8k context, from Llama 3.1 8B at full BF16 precision to Gemma 4 31B squeezed into Q6 with just 0.48GB to spare. Price and speed are what actually separate them: the RTX 5090 costs more than double this card's $949 for 2.9x the bandwidth (1,792 vs 608 GB/s), while the Radeon AI PRO R9700 costs $350 more (27% over this card's price) for a 608-to-640 GB/s gap worth a real but modest 5.3% faster decode on every model both cards fit. The nearest card by price, the RTX 5080 at $999, only reaches 16GB: 8 fewer models fit its VRAM (45 of 87) than fit this card's 32GB, even though its 960 GB/s bandwidth is 58% higher. Qwen 3.6 27B needs Q3 quantization to fit the RTX 5080's 16GB; this card's 32GB holds the same model at the meaningfully higher-quality Q6.
Popular models for this GPU
Models this GPU runs natively in VRAM (53)
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q3_K_M · ~18.2 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q2_K · ~16.4 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q5_K_M · ~54.2 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q5_K_M · ~14.6 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/AQ5_K_M · ~54.2 t/s
Show 48 more
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q5_K_M · ~14.9 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q5_K_M · ~16 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q5_K_M · ~15.6 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q5_K_M · ~15.6 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q5_K_M · ~15.6 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q6_K · ~45.7 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q6_K · ~14.8 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q6_K · ~43.8 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q6_K · ~47.8 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/AQ6_K · ~17.2 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q6_K · ~15.5 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q6_K · ~16.7 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q6_K · ~17.4 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2Q6_K · ~17.4 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~44.9 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/AQ6_K · ~17.4 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q6_K · ~36.7 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q8_0 · ~14.7 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q8_0 · ~15.5 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~30.5 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q8_0 · ~23.1 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q8_0 · ~22.9 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q8_0 · ~24.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~15.4 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~15.5 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q8_0 · ~24.7 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~18.6 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~21.6 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ABF16 · ~21.6 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~23.1 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~23.1 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~23 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~25.2 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~25.4 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~23.9 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~23.2 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~21.5 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~24.3 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~28.8 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~31.1 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~35.1 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~47 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~47 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~63.4 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~75.6 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~91.3 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~188.1 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~222.6 t/s
Models that fit with CPU offload (9)
These use system RAM for layers that don't fit in VRAM, so expect much slower inference.
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q2_K · ~4.1 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q2_K · ~4 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q2_K · ~2.9 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q2_K · ~4.6 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q2_K · ~4.6 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q4_K_M · ~1.4 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q4_K_M · ~1.5 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q4_K_M · ~1.5 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q4_K_M · ~1.5 t/s
Too large for this GPU (32)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.6 355B
- GLM-4.7 358B
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- DeepSeek V4 Flash 0731 284B
- Qwen3.8 2.4T-A95B
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
Compare Intel Arc Pro B70 32GB with other GPUs
Frequently asked questions
- How much VRAM does the Intel Arc Pro B70 32GB have?
- The Intel Arc Pro B70 32GB has 32 GB of GDDR6 with 608 GB/s memory bandwidth.
- What is the Intel Arc Pro B70 32GB best for?
- With 32 GB of VRAM, the Intel Arc Pro B70 32GB is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the Intel Arc Pro B70 32GB run locally?
- The Intel Arc Pro B70 32GB can run 53 of the 94 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q6_K, Muse Glimmer 30B at Q6_K, Ornith 1.5 9B at BF16.
- Can the Intel Arc Pro B70 32GB run Gemma 4 31B?
- Yes. The Intel Arc Pro B70 32GB runs Gemma 4 31B natively in VRAM at Q6_K quantization, achieving approximately 14.8 tokens per second.
- Can the Intel Arc Pro B70 32GB run Qwen 3.6 27B?
- Yes. The Intel Arc Pro B70 32GB runs Qwen 3.6 27B natively in VRAM at Q6_K quantization, achieving approximately 17.4 tokens per second.
- Can the Intel Arc Pro B70 32GB run Qwen3 8B?
- Yes. The Intel Arc Pro B70 32GB runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 23 tokens per second.