Apple M4 Pro
The Apple M4 Pro ships in 24–48 GB unified-memory configurations at 273 GB/s. Across those configurations it runs 51 of our 84 tracked models natively in VRAM at 8k context.
More memory means more of our tracked models fit natively — see which configuration you need below.
| Configuration | Bandwidth | CPU cores | Native models | + Offload |
|---|---|---|---|---|
| 48 GB | 273 GB/s | 14 (10P + 4E) | 51 / 84 | 0 |
| 24 GB | 273 GB/s | 12 (8P + 4E) | 39 / 84 | 0 |
How much of the Apple M4 Pro's memory is actually usable?
macOS and background apps need a slice of the pool before a model gets to use it — this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:
M4 Pro's bandwidth recovery, from its own vantage point
M3 Pro cut memory bandwidth 25% below what M1 Pro and M2 Pro had already shipped — a real regression, not a rumor (see the M3 Pro page's own bandwidth note for that story). M4 Pro is the generation that reversed it. Tracking the Pro tier's bandwidth across five generations:
M1 Pro and M2 Pro both shipped at 200 GB/s, M3 Pro cut that to 150 GB/s, and M4 Pro (this page) jumped 82% past M3 Pro to 273 GB/s — not just a recovery, but the largest single-generation bandwidth increase anywhere in the Pro tier's history, pushing 36.5% past the old M1 Pro/M2 Pro ceiling. M5 Pro added another 12.5% on top of that to reach 307 GB/s. Since decode is bandwidth-bound, a used M4 Pro already out-decodes an M1 Pro or M2 Pro of the same capacity by more than a third, on top of running cooler and more efficiently.
Apple M4 Pro (48GB)
With 48 GB LPDDR5X at 273 GB/s, this configuration runs 51 models natively. It handles 70B-class models at Q4 quantization.
Apple M4 Pro (48GB): the upgraded M4 Pro configuration, available in the MacBook Pro and Mac mini, standard on the 14-core CPU (10P+4E)/20-core GPU die, at 273 GB/s — a 75% jump over the M3 Pro generation's 150 GB/s and the largest single-generation Pro-tier bandwidth increase Apple has shipped (see this page's bandwidth note for the full multi-generation picture).
With 40 GB of real headroom, this calculator's highest-fitting quant for a dense 70B model here is an aggressive Q2_K (32.88 GB, 7.4 tok/s) — technically native, but a real quality trade-off; the 10-15 tok/s sometimes quoted for "70B on M4 Pro" describes faster, higher-bandwidth configurations, not this one. MoE models fare far better at this size: Qwen 3.5 35B-A3B reaches near-full Q6_K (32.37 GB, 26.1 tok/s), and Command-R 35B fits at Q5_K_M (39.94 GB, 6.1 tok/s). 51 of the 84 tracked models fit natively.
Full llama.cpp K-quants support and MLX optimization. Nothing this configuration fits approaches the default ~36 GB Metal working-set limit (75% of 48 GB) — the largest fit here is under 40 GB — so the manual wired_limit override isn't necessary at this capacity.
Models the 48 GB configuration runs natively (51)
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~7.3 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~7.4 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~7.4 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~7.4 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q5_K_M · ~6.9 t/s
Show 46 more
- Command-R 35B35B · MMLU-Pro 33.0Q5_K_M · ~6.1 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q6_K · ~26.1 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q6_K · ~7.1 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q6_K · ~7.2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q6_K · ~7.7 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q6_K · ~7.6 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4Q6_K · ~7.6 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q6_K · ~7.6 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q8_0 · ~19.7 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2Q6_K · ~7.6 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~19.1 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~6.8 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~7.2 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~7.5 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~24.8 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q8_0 · ~15.8 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q8_0 · ~8.1 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q8_0 · ~8.6 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~16.9 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~7.1 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~7 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~7.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~8.5 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~8.6 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~8 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~10.3 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~12 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~6.6 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~6.6 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~6.6 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~7.1 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~7.3 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~13.2 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~12.8 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~11.9 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~13.4 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~15.9 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~17.2 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~19.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~26 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~26 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~35 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~41.8 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~50.5 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~104 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~123 t/s
Apple M4 Pro (24GB)
With 24 GB LPDDR5X at 273 GB/s, this configuration runs 39 models natively. It comfortably runs 7B–32B models at Q4; 70B-class models typically need CPU offload.
Apple M4 Pro (24GB): the base M4 Pro configuration, available in the MacBook Pro and Mac mini, standard on a smaller 12-core CPU (8P+4E)/16-core GPU die than the 48GB build's 14-core CPU/20-core GPU — though both share the identical 273 GB/s memory bandwidth; Apple doesn't cut bandwidth for the smaller-die configuration on this chip.
With 16 GB of real headroom, this configuration is where mixture-of-experts models pull dramatically ahead of same-size dense ones: Qwen 3.5 35B-A3B — a 35B-parameter MoE model — fits at Q2_K (15.12 GB) at 54.9 tok/s, essentially real-time, while the dense Qwen3 32B needs the same Q2_K precision at a near-identical footprint (15.50 GB) but manages only 15.8 tok/s, a 3.5x gap from MoE decode reading far fewer bytes per token. Llama 3.1 8B and Qwen2.5 7B both clear Q8_0, near-full precision, at 22.8 and 25.5 tok/s. 39 of the 84 tracked models fit natively.
MLX and llama.cpp's Metal backend are both mature here. Nothing this configuration fits comes close to the default ~18 GB Metal working-set limit (75% of 24 GB) — the largest fit is around 15.5 GB — so the manual wired_limit override doesn't apply at this capacity.
Models the 24 GB configuration runs natively (39)
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q2_K · ~54.9 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q2_K · ~15.8 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q2_K · ~51.4 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q2_K · ~47.3 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q2_K · ~16.2 t/s
Show 34 more
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q2_K · ~18.5 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q3_K_M · ~16.1 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~24.8 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q3_K_M · ~33.8 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q3_K_M · ~16.9 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q3_K_M · ~17.4 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q4_K_M · ~29.1 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q6_K · ~16.2 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q6_K · ~16 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q6_K · ~17 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q6_K · ~19.2 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q8_0 · ~15.6 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q6_K · ~16.7 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0Q8_0 · ~17.3 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5Q8_0 · ~22.2 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3Q8_0 · ~22.8 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0Q8_0 · ~22.8 t/s
- Qwen3 8B8B · MMLU-Pro 56.7Q8_0 · ~22.5 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3Q8_0 · ~25.5 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0Q8_0 · ~24.9 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~25.7 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~24.2 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~20.2 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~25.2 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~15.9 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~17.2 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~19.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~26 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~26 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~35 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~41.8 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~50.5 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~104 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~123 t/s
Too large for any Apple M4 Pro configuration (33)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Scout 109B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.5 Air 106B
- GLM-4.6 355B
- GLM-4.6V 106B
- GLM-4.7 358B
- Qwen 3.5 122B-A10B (MoE)
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Nemotron 3 Super 120B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
Compare Apple M4 Pro with other GPUs
Continue reading
Frequently asked questions
- How much memory does the Apple M4 Pro have?
- The Apple M4 Pro ships in 2 unified-memory configurations: 48 GB and 24 GB, all at 273 GB/s.
- Should I get the 24 GB or 48 GB Apple M4 Pro?
- Both run everything that fits natively in 24 GB. The extra memory in the 48 GB configuration additionally fits Qwen 2.5 72B Instruct, Llama 3.3 70B Instruct, DeepSeek R1 Distill Llama 70B, and 9 more models natively in VRAM — worth the upgrade if you plan to run any of those.
- How much VRAM does the Apple M4 Pro (48GB) have?
- The Apple M4 Pro (48GB) has 48 GB of LPDDR5X with 273 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M4 Pro (48GB) best for?
- With 48 GB of VRAM, the Apple M4 Pro (48GB) is ideal for running 70B-class models at Q4 quantization and large MoE models — a workstation sweet spot for local inference.
- What LLMs can the Apple M4 Pro (48GB) run locally?
- The Apple M4 Pro (48GB) can run 51 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q2_K, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M4 Pro (48GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M4 Pro (48GB) runs Llama 3.3 70B Instruct natively in VRAM at Q2_K quantization, achieving approximately 7.4 tokens per second.
Show 8 more questions
- Can the Apple M4 Pro (48GB) run Qwen 3.6 27B?
- Yes. The Apple M4 Pro (48GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 7.5 tokens per second.
- Can the Apple M4 Pro (48GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M4 Pro (48GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 6.6 tokens per second.
- How much VRAM does the Apple M4 Pro (24GB) have?
- The Apple M4 Pro (24GB) has 24 GB of LPDDR5X with 273 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M4 Pro (24GB) best for?
- With 24 GB of VRAM, the Apple M4 Pro (24GB) is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the Apple M4 Pro (24GB) run locally?
- The Apple M4 Pro (24GB) can run 39 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at Q8_0, Llama 3.2 3B Instruct at FP32, Llama 3.2 1B Instruct at FP32.
- Can the Apple M4 Pro (24GB) run Llama 3.3 70B Instruct?
- The Apple M4 Pro (24GB) does not have enough VRAM to run Llama 3.3 70B Instruct. You would need more VRAM or a lower quantization level.
- Can the Apple M4 Pro (24GB) run Qwen 3.6 27B?
- Yes. The Apple M4 Pro (24GB) runs Qwen 3.6 27B natively in VRAM at Q3_K_M quantization, achieving approximately 16.1 tokens per second.
- Can the Apple M4 Pro (24GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M4 Pro (24GB) runs Llama 3.1 8B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 22.8 tokens per second.