Intel Arc Pro B70 32GB

The Intel Arc Pro B70 32GB has 32 GB VRAM and 608 GB/s memory bandwidth. It can run 53 of our 94 tracked models natively in VRAM at 8k context.

With 32 GB GDDR6, the Intel Arc Pro B70 32GB is a workstation-tier GPU that can run 53 models natively. It comfortably runs 7B–32B models at Q4; 70B-class models typically need CPU offload.

The Intel Arc Pro B70 launched March 25, 2026 at $949 as Intel's first "Big Battlemage" card, the full-die BMG-G31 workstation GPU Intel had been previewing since CES 2025. Where the original CES announcement pointed at a 24GB card, the shipping product doubled down: 32GB of GDDR6 on a 256-bit bus (608 GB/s), 32 Xe cores, 256 XMX engines rated at 367 INT8 TOPS, and a 230W board power target. That 32GB framebuffer runs 7-8B models at full BF16 precision, 14B models at near-lossless Q8, and 27-31B dense models at Q6, with the Vulkan backend providing usable LLM inference speeds on both Linux and Windows.

Intel Arc Pro B70 32GB: sits above Intel's existing Arc Pro B60 in the same Battlemage Pro lineup, which shipped first, in 2025, at 24GB of GDDR6 and 380 GB/s. This card followed as the first "Big Battlemage" part on the larger, full BMG-G31 die: 8GB more capacity and 60% more bandwidth than the B60 (608 vs 380 GB/s), launched March 25, 2026 at $949. 32 Xe cores and 256 XMX engines are rated at 367 INT8 TOPS, in a 230W board power envelope on PCIe 5.0.

53 of this site's 87 tracked models fit natively in VRAM at 8k context. Llama 3.1 8B and Qwen 2.5 7B are the only two that run at full BF16 precision (19.12 GB at 23.1 tok/s, 17.55 GB at 25.2 tok/s); Qwen3 14B is the largest to reach near-lossless Q8 (19.12 GB, 23.1 tok/s). Qwen 3.6 27B is the biggest dense model with real headroom to spare, fitting Q6 at 25.43 GB and 17.4 tok/s, well under this card's 30.4 GB usable ceiling.

Phoronix's Linux review found it working well on the fully open-source Mesa/Xe driver stack, with llama.cpp's SYCL backend and an April 2026 preview OpenVINO backend both usable alongside the default Vulkan path. Independent community benchmarks put Qwen 3.6 27B at roughly 14 to 22 tok/s depending on quantization, bracketing this site's own 17.4 to 23.3 tok/s Q6/Q4 estimates well. ISV-certified for professional workloads on Linux and Windows.

VendorIntel
ArchitectureXe2-HPG (Battlemage)
VRAM32 GB
Memory typeGDDR6
Memory bandwidth608 GB/s
Compute backendVULKAN
TierWorkstation
Released2026
Models (native)53 / 94
Models (offload)9 / 94
Software: Vulkan backend works in llama.cpp; SYCL backend available with oneAPI toolkit. Primarily a workstation/professional card.

Three vendors land on the exact same 32GB ceiling, at three very different prices

Plotting this card against its closest capacity peers and its nearest price-tier neighbor shows something the spec sheet alone doesn't: 32GB isn't a niche Intel number here, it's a ceiling three different vendors converge on from three very different starting points.

0950190001734VRAM (GB)Bandwidth (GB/s)NVIDIA RTX 5080Intel Arc Pro B70 32GBAMD Radeon AI PRO R9700 32GBNVIDIA RTX 4090NVIDIA RTX 5090
VRAM and memory bandwidth, from each card's real spec sheet. A card further right holds bigger models; a card further up decodes them faster once they fit.

This card (this page), the $1,299 Radeon AI PRO R9700, and the $1,999 RTX 5090 all sit at exactly 32GB, and because VRAM capacity, not bandwidth, decides what fits, all three run the identical set of 53 of this site's 87 tracked models natively at 8k context, from Llama 3.1 8B at full BF16 precision to Gemma 4 31B squeezed into Q6 with just 0.48GB to spare. Price and speed are what actually separate them: the RTX 5090 costs more than double this card's $949 for 2.9x the bandwidth (1,792 vs 608 GB/s), while the Radeon AI PRO R9700 costs $350 more (27% over this card's price) for a 608-to-640 GB/s gap worth a real but modest 5.3% faster decode on every model both cards fit. The nearest card by price, the RTX 5080 at $999, only reaches 16GB: 8 fewer models fit its VRAM (45 of 87) than fit this card's 32GB, even though its 960 GB/s bandwidth is 58% higher. Qwen 3.6 27B needs Q3 quantization to fit the RTX 5080's 16GB; this card's 32GB holds the same model at the meaningfully higher-quality Q6.

Popular models for this GPU

Models this GPU runs natively in VRAM (53)

Show 48 more

Models that fit with CPU offload (9)

These use system RAM for layers that don't fit in VRAM, so expect much slower inference.

Too large for this GPU (32)

Compare Intel Arc Pro B70 32GB with other GPUs

Frequently asked questions

How much VRAM does the Intel Arc Pro B70 32GB have?
The Intel Arc Pro B70 32GB has 32 GB of GDDR6 with 608 GB/s memory bandwidth.
What is the Intel Arc Pro B70 32GB best for?
With 32 GB of VRAM, the Intel Arc Pro B70 32GB is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
What LLMs can the Intel Arc Pro B70 32GB run locally?
The Intel Arc Pro B70 32GB can run 53 of the 94 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q6_K, Muse Glimmer 30B at Q6_K, Ornith 1.5 9B at BF16.
Can the Intel Arc Pro B70 32GB run Gemma 4 31B?
Yes. The Intel Arc Pro B70 32GB runs Gemma 4 31B natively in VRAM at Q6_K quantization, achieving approximately 14.8 tokens per second.
Can the Intel Arc Pro B70 32GB run Qwen 3.6 27B?
Yes. The Intel Arc Pro B70 32GB runs Qwen 3.6 27B natively in VRAM at Q6_K quantization, achieving approximately 17.4 tokens per second.
Can the Intel Arc Pro B70 32GB run Qwen3 8B?
Yes. The Intel Arc Pro B70 32GB runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 23 tokens per second.