Apple M4 Max
The Apple M4 Max ships in 36–128 GB unified-memory configurations at 410–546 GB/s. Across those configurations it runs 64 of our 84 tracked models natively in VRAM at 8k context.
More memory means more of our tracked models fit natively — see which configuration you need below.
| Configuration | Bandwidth | CPU cores | Native models | + Offload |
|---|---|---|---|---|
| 128 GB | 546 GB/s | 16 (12P + 4E) | 64 / 84 | 0 |
| 64 GB | 546 GB/s | 16 (12P + 4E) | 56 / 84 | 0 |
| 48 GB | 546 GB/s | 16 (12P + 4E) | 51 / 84 | 0 |
| 36 GB | 410 GB/s | 14 (10P + 4E) | 47 / 84 | 0 |
How much of the Apple M4 Max's memory is actually usable?
macOS and background apps need a slice of the pool before a model gets to use it — this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:
M4 Max's cheapest configuration outruns M3 Max's fastest one
Like the M3 Max before it, M4 Max ships on two physical dies: a cut-down 14-core CPU/32-core GPU die that only comes with 36GB of memory, and a full 16-core CPU/40-core GPU die at 48GB, 64GB, and 128GB. Comparing the binned base tier against each generation's full-die flagship shows just how far the floor moved in one generation:
M4 Max's binned 36GB configuration (this page) reaches 410 GB/s — 2.5% faster than M3 Max's full, non-binned 400 GB/s flagship die, even though this is the cheapest way into M4 Max. The full-die M4 Max climbs further still, to 546 GB/s, 36.5% ahead of the equivalent M3 Max tier. Since LLM decode is bandwidth-bound, that means the entry-level M4 Max MacBook Pro out-decodes last generation's top-spec M3 Max on every model both fit — a real generational leap, not just a bigger number on the base config's spec sheet.
Apple M4 Max (128GB)
With 128 GB LPDDR5X at 546 GB/s, this configuration runs 64 models natively. It runs 70B-class dense models and large MoE models entirely in VRAM.
Apple M4 Max (128GB): the maximum-memory M4 Max configuration, launched October 2024 in the MacBook Pro and, from March 2025, the Mac Studio, on the full, un-binned 16-core CPU (12P+4E)/40-core GPU die at 546 GB/s — the fastest memory bandwidth Apple has ever shipped in a laptop, and a real generational jump from the M3 Max's 400 GB/s ceiling.
With 120 GB of real headroom, this is the first M4 Max configuration to fit genuine frontier-scale MoE models: Qwen3 235B-A22B fits at Q2_K (102.05 GB, 14.8 tok/s) and GPT-OSS 120B reaches Q6_K (107.93 GB, 30.6 tok/s) — both about 37% faster than the identically-sized M3 Max 128GB manages on the same fits (10.8 and 22.4 tok/s there), tracking the 546-vs-400 GB/s bandwidth gap almost exactly. Dense models do well too: Llama 3.3 70B and Qwen2.5 72B both reach Q8_0 (86.35 GB and 88.73 GB), near-full precision, at 5.7 and 5.5 tok/s. 64 of the 84 tracked models fit natively.
MLX and llama.cpp's Metal backend are both fully optimized here, and Apple Silicon's efficiency advantage over a discrete-GPU workstation is largest at this end of the lineup — this much memory bandwidth in a laptop chassis has no real equivalent on the discrete-GPU side.
Models the 128 GB configuration runs natively (64)
- Qwen3 235B-A22B (MoE)235B · MMLU-Pro 84.4Q2_K · ~14.8 t/s
- MiniMax M2.5 229B229B · MMLU-Pro 84.8Q2_K · ~29.6 t/s
- MiniMax M2.7 229B229B · MMLU-Pro 86.0Q2_K · ~29.6 t/s
- Step 3.7 Flash198B · MMLU-Pro —Q3_K_M · ~22.8 t/s
- Step 3.5 Flash196.81B · MMLU-Pro 84.4Q3_K_M · ~22.8 t/s
Show 59 more
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0Q5_K_M · ~4.6 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro —Q5_K_M · ~4.6 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q6_K · ~14.8 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q6_K · ~13 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7Q6_K · ~30.6 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q6_K · ~8.9 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q6_K · ~12.7 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q6_K · ~12.7 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q8_0 · ~5.5 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q8_0 · ~5.7 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q8_0 · ~5.7 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q8_0 · ~5.7 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7BF16 · ~5 t/s
- Command-R 35B35B · MMLU-Pro 33.0BF16 · ~5.4 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3BF16 · ~21.7 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2BF16 · ~6.1 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0BF16 · ~6.2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5BF16 · ~6.5 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0BF16 · ~6.5 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4BF16 · ~6.5 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0BF16 · ~6.5 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3BF16 · ~21.4 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2BF16 · ~6.7 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5BF16 · ~21 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0BF16 · ~7.6 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5BF16 · ~7.9 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2BF16 · ~8 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~49.6 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6FP32 · ~8.6 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8FP32 · ~4.5 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2FP32 · ~4.8 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9FP32 · ~9.1 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0FP32 · ~7.2 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7FP32 · ~7.2 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4FP32 · ~7.6 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6FP32 · ~8.7 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~8.8 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2FP32 · ~8.5 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~11 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~12 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~13.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~13.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~13.2 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~14.1 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~14.5 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~26.5 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~25.7 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~23.7 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~26.8 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~31.8 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~34.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~38.7 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~52 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~51.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~70.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~83.5 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~100.9 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~207.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~246 t/s
Apple M4 Max (64GB)
With 64 GB LPDDR5X at 546 GB/s, this configuration runs 56 models natively. It handles 70B-class models at Q4 quantization.
Apple M4 Max (64GB): the mid-tier M4 Max configuration, available in the MacBook Pro and Mac Studio on the same full 16-core CPU/40-core GPU die as the 48GB and 128GB builds, all sharing the generation's 546 GB/s ceiling — only the 36GB base configuration ships on a slower, cut-down die (see this page's bandwidth note).
With 56 GB of real headroom, this is the first M4 Max tier where 70B-class dense models clear a real quantization instead of the aggressive Q2_K floor: Llama 3.3 70B reaches Q4_K_M (50.75 GB, 9.6 tok/s) and Qwen2.5 72B reaches Q4_K_M (52.12 GB, 9.4 tok/s). Command-R 35B goes further, hitting near-full-precision Q8_0 (53.70 GB, 9.1 tok/s). MoE models are faster still at this size: Qwen 3.5 35B-A3B fits at Q8_0 (41.86 GB, 40.5 tok/s). 56 of the 84 tracked models fit natively.
MLX and llama.cpp's Metal backend are both mature here. Nothing this configuration fits comes close to the default ~48 GB Metal working-set limit (75% of 64 GB) — the largest fit here is under 54 GB — so the manual wired_limit override other Apple Silicon pages need doesn't bind at this capacity.
Models the 64 GB configuration runs natively (56)
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q2_K · ~29.4 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q2_K · ~27.3 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q2_K · ~18 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q2_K · ~26 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q2_K · ~26 t/s
Show 51 more
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q4_K_M · ~9.4 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q4_K_M · ~9.6 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q4_K_M · ~9.6 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q4_K_M · ~9.6 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q6_K · ~12 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q8_0 · ~9.1 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q8_0 · ~40.5 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q8_0 · ~11.1 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q8_0 · ~11.3 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q8_0 · ~12.1 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q8_0 · ~11.9 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4Q8_0 · ~11.9 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q8_0 · ~11.9 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q8_0 · ~39.5 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2Q8_0 · ~12.1 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~38.2 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~13.6 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~14.4 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~14.9 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~49.6 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q8_0 · ~31.6 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8BF16 · ~8.9 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2BF16 · ~9.4 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9BF16 · ~18 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~14.1 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~14.1 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~14.9 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~17 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~8.8 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~16 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~11 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~12 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~13.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~13.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~13.2 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~14.1 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~14.5 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~26.5 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~25.7 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~23.7 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~26.8 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~31.8 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~34.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~38.7 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~52 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~51.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~70.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~83.5 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~100.9 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~207.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~246 t/s
Apple M4 Max (48GB)
With 48 GB LPDDR5X at 546 GB/s, this configuration runs 51 models natively. It handles 70B-class models at Q4 quantization.
Apple M4 Max (48GB): the entry point into M4 Max's full, un-binned 16-core CPU/40-core GPU die, available in the MacBook Pro and Mac Studio — the same 546 GB/s as the 64GB and 128GB builds, a real step up from the cut-down 36GB base configuration's 410 GB/s.
With 40 GB of real headroom — identical to the M4 Pro 48GB configuration — this tier fits the same 51 of 84 tracked models, but at exactly double the speed on every shared fit: Llama 3.3 70B reaches the same Q2_K (32.88 GB) at 14.9 tok/s here versus 7.4 tok/s on the M4 Pro 48GB, and Qwen2.5 72B reaches 14.5 tok/s versus 7.3 tok/s — both track the 546-vs-273 GB/s bandwidth ratio almost exactly. MoE models do particularly well: Qwen 3.5 35B-A3B fits at near-full Q6_K (32.37 GB, 52.1 tok/s).
MLX and llama.cpp's Metal backend are both fully supported. Choosing this over the identically-priced-per-GB M4 Pro 48GB buys roughly 2x decode speed on every model both fit — the Max tier's advantage here is bandwidth, not capacity.
Models the 48 GB configuration runs natively (51)
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q2_K · ~14.5 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q2_K · ~14.9 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q2_K · ~14.9 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q2_K · ~14.9 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q5_K_M · ~13.8 t/s
Show 46 more
- Command-R 35B35B · MMLU-Pro 33.0Q5_K_M · ~12.2 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q6_K · ~52.1 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q6_K · ~14.1 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q6_K · ~14.4 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q6_K · ~15.5 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q6_K · ~15.2 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4Q6_K · ~15.2 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q6_K · ~15.2 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q8_0 · ~39.5 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2Q6_K · ~15.2 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~38.2 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~13.6 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~14.4 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~14.9 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~49.6 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q8_0 · ~31.6 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q8_0 · ~16.3 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q8_0 · ~17.1 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~33.7 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~14.1 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~14.1 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~14.9 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~17 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~17.2 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~16 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~20.6 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~23.9 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~13.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~13.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~13.2 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~14.1 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~14.5 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~26.5 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~25.7 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~23.7 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~26.8 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~31.8 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~34.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~38.7 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~52 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~51.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~70.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~83.5 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~100.9 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~207.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~246 t/s
Apple M4 Max (36GB)
With 36 GB LPDDR5X at 410 GB/s, this configuration runs 47 models natively. It comfortably runs 7B–32B models at Q4; 70B-class models typically need CPU offload.
Apple M4 Max (36GB): the standard base configuration of Apple's M4 Max chip, launched October 2024 in the MacBook Pro and, from March 2025, the base Mac Studio, on a cut-down 14-core CPU (10P+4E)/32-core GPU die at 410 GB/s — 136 GB/s less than the 546 GB/s the 48GB, 64GB, and 128GB configurations get on the full die. Apple offers no way to pair 36GB with the faster die, the same binned-base-tier pattern the M3 Max shipped a generation earlier (see this page's bandwidth note).
With 28 GB of real headroom, this configuration fits the same 47 of 84 tracked models as the identically-sized M3 Max 36GB and M3 Pro 36GB — capacity decides what fits, not bandwidth — but runs about 37% faster than the M3 Max 36GB on every shared fit, since 410 GB/s is 37% more than that chip's 300 GB/s. Qwen3 32B reaches Q5_K_M (27.66 GB) at 13.3 tok/s here versus 9.7 tok/s on the M3 Max 36GB, and Qwen 3.5 35B-A3B reaches Q4_K_M (24.06 GB) at 52.4 tok/s versus 38.4 tok/s — both track the 410-vs-300 GB/s ratio closely. Even on this cut-down die, 410 GB/s outright beats the previous generation's full, non-binned M3 Max flagship die (400 GB/s).
MLX and llama.cpp's Metal backend are both mature here. The default macOS Metal working-set limit (~75% of 36 GB, 27 GB) sits just under this page's largest fit — Qwen3 32B's 27.66 GB Q5_K_M build — so raising it with `sudo sysctl iogpu.wired_limit_mb=28672` (leaving 8 GB for macOS) covers it with room to spare, the same fix and figure this site notes for the identically-sized M3 Max and M3 Pro 36GB configurations.
Models the 36 GB configuration runs natively (47)
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q3_K_M · ~15.1 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q2_K · ~13.6 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q4_K_M · ~52.4 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q4_K_M · ~14 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q4_K_M · ~14.3 t/s
Show 42 more
- Qwen3 32B32.8B · MMLU-Pro 65.5Q5_K_M · ~13.3 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q4_K_M · ~14.9 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4Q4_K_M · ~14.9 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q4_K_M · ~14.9 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q5_K_M · ~43.4 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2Q4_K_M · ~14.8 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q5_K_M · ~41.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q5_K_M · ~14.6 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q6_K · ~13.8 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q6_K · ~14.4 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~37.2 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q6_K · ~30.4 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q6_K · ~15.6 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q6_K · ~16.3 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q8_0 · ~25.3 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q8_0 · ~19.2 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q8_0 · ~19 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q8_0 · ~20.2 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q8_0 · ~22.9 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q8_0 · ~23.4 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q8_0 · ~20.5 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~15.5 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~18 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~19.2 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~19.2 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~19.1 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~20.9 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~21.1 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~19.9 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~19.3 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~17.8 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~20.2 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~23.9 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~25.8 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~29.1 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~39 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~39 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~52.6 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~62.7 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~75.8 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~156.1 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~184.7 t/s
Too large for any Apple M4 Max configuration (20)
Compare Apple M4 Max with other GPUs
- Apple M4 Max (128GB)vsApple M3 Max (128GB)128 GB each
- Apple M4 Max (128GB)vsApple M1 Ultra (128GB)128 GB each
- Apple M4 Max (128GB)vsNVIDIA RTX 5090+96 GB VRAM
- Apple M4 Max (128GB)vsApple M3 Ultra (96GB)+32 GB VRAM
- Apple M4 Max (128GB)vsAMD Strix Halo (128GB)128 GB each
- Apple M4 Max (64GB)vsApple M3 Max (64GB)64 GB each
- Apple M4 Max (64GB)vsAMD Strix Halo (64GB)64 GB each
- Apple M4 Max (48GB)vsApple M3 Max (48GB)48 GB each
- Apple M4 Max (36GB)vsApple M3 Max (36GB)36 GB each
Continue reading
Frequently asked questions
- How much memory does the Apple M4 Max have?
- The Apple M4 Max ships in 4 unified-memory configurations: 128 GB, 64 GB, 48 GB, 36 GB, at 410–546 GB/s depending on configuration.
- Should I get the 36 GB or 128 GB Apple M4 Max?
- Both run everything that fits natively in 36 GB. The extra memory in the 128 GB configuration additionally fits Qwen3 235B-A22B (MoE), MiniMax M2.5 229B, MiniMax M2.7 229B, and 14 more models natively in VRAM — worth the upgrade if you plan to run any of those.
- How much VRAM does the Apple M4 Max (128GB) have?
- The Apple M4 Max (128GB) has 128 GB of LPDDR5X with 546 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M4 Max (128GB) best for?
- With 128 GB of unified memory, the Apple M4 Max (128GB) is a high-capacity laptop platform that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
- What LLMs can the Apple M4 Max (128GB) run locally?
- The Apple M4 Max (128GB) can run 64 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q8_0, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M4 Max (128GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M4 Max (128GB) runs Llama 3.3 70B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 5.7 tokens per second.
Show 20 more questions
- Can the Apple M4 Max (128GB) run Qwen 3.6 27B?
- Yes. The Apple M4 Max (128GB) runs Qwen 3.6 27B natively in VRAM at BF16 quantization, achieving approximately 8 tokens per second.
- Can the Apple M4 Max (128GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M4 Max (128GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 13.2 tokens per second.
- How much VRAM does the Apple M4 Max (64GB) have?
- The Apple M4 Max (64GB) has 64 GB of LPDDR5X with 546 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M4 Max (64GB) best for?
- With 64 GB of VRAM, the Apple M4 Max (64GB) is ideal for running 70B-class models at Q4 quantization and large MoE models — a workstation sweet spot for local inference.
- What LLMs can the Apple M4 Max (64GB) run locally?
- The Apple M4 Max (64GB) can run 56 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q4_K_M, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M4 Max (64GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M4 Max (64GB) runs Llama 3.3 70B Instruct natively in VRAM at Q4_K_M quantization, achieving approximately 9.6 tokens per second.
- Can the Apple M4 Max (64GB) run Qwen 3.6 27B?
- Yes. The Apple M4 Max (64GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 14.9 tokens per second.
- Can the Apple M4 Max (64GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M4 Max (64GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 13.2 tokens per second.
- How much VRAM does the Apple M4 Max (48GB) have?
- The Apple M4 Max (48GB) has 48 GB of LPDDR5X with 546 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M4 Max (48GB) best for?
- With 48 GB of VRAM, the Apple M4 Max (48GB) is ideal for running 70B-class models at Q4 quantization and large MoE models — a workstation sweet spot for local inference.
- What LLMs can the Apple M4 Max (48GB) run locally?
- The Apple M4 Max (48GB) can run 51 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q2_K, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M4 Max (48GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M4 Max (48GB) runs Llama 3.3 70B Instruct natively in VRAM at Q2_K quantization, achieving approximately 14.9 tokens per second.
- Can the Apple M4 Max (48GB) run Qwen 3.6 27B?
- Yes. The Apple M4 Max (48GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 14.9 tokens per second.
- Can the Apple M4 Max (48GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M4 Max (48GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 13.2 tokens per second.
- How much VRAM does the Apple M4 Max (36GB) have?
- The Apple M4 Max (36GB) has 36 GB of LPDDR5X with 410 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M4 Max (36GB) best for?
- With 36 GB of VRAM, the Apple M4 Max (36GB) is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the Apple M4 Max (36GB) run locally?
- The Apple M4 Max (36GB) can run 47 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at BF16, Llama 3.2 3B Instruct at FP32, Llama 3.2 1B Instruct at FP32.
- Can the Apple M4 Max (36GB) run Llama 3.3 70B Instruct?
- The Apple M4 Max (36GB) does not have enough VRAM to run Llama 3.3 70B Instruct. You would need more VRAM or a lower quantization level.
- Can the Apple M4 Max (36GB) run Qwen 3.6 27B?
- Yes. The Apple M4 Max (36GB) runs Qwen 3.6 27B natively in VRAM at Q6_K quantization, achieving approximately 14.4 tokens per second.
- Can the Apple M4 Max (36GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M4 Max (36GB) runs Llama 3.1 8B Instruct natively in VRAM at BF16 quantization, achieving approximately 19.2 tokens per second.