Apple M1 Max
The Apple M1 Max ships in 32–64 GB unified-memory configurations at 400 GB/s. Across those configurations it runs 56 of our 84 tracked models natively in VRAM at 8k context.
More memory means more of our tracked models fit natively — see which configuration you need below.
| Configuration | Bandwidth | CPU cores | Native models | + Offload |
|---|---|---|---|---|
| 64 GB | 400 GB/s | 10 (8P + 2E) | 56 / 84 | 0 |
| 32 GB | 400 GB/s | 10 (8P + 2E) | 46 / 84 | 0 |
How much of the Apple M1 Max's memory is actually usable?
macOS and background apps need a slice of the pool before a model gets to use it — this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:
Apple M1 Max (64GB)
With 64 GB LPDDR5 at 400 GB/s, this configuration runs 56 models natively. It handles 70B-class models at Q4 quantization.
Apple M1 Max (64GB): the memory-maxed configuration of Apple's first Max-tier chip, launched October 2021 as an $800 build-to-order upgrade over the 32GB base on either the 14-inch or 16-inch MacBook Pro — the only way in the M1 generation's laptop lineup to pair more than 32 GB with 400 GB/s bandwidth.
With 56 GB of real headroom after this site's standard 8 GB OS reservation, this is the first M1-family laptop configuration that clears 70B-class dense models: this site's calculator fits Llama 3.3 70B (50.8 GB at Q4_K_M, ~7.1 tok/s) and Qwen2.5 72B (52.1 GB at Q4_K_M, ~6.9 tok/s) natively. That's a different capability tier from the 32GB sibling, which tops out around 32-35B models — the extra 32 GB doesn't just add headroom, it unlocks a whole model class. 56 of the 84 tracked models fit natively, 10 more than the 32GB config's 46.
Mature Metal and MLX support, same as the 32GB sibling. macOS's Metal working-set limit (roughly 75% of unified memory, which llama.cpp and MLX both read on startup) sits around 48 GB here — comfortably above what any of this site's fitting 70B-class quants need, unlike the tighter headroom on 16-32 GB Apple Silicon configurations.
Models the 64 GB configuration runs natively (56)
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q2_K · ~21.6 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q2_K · ~20 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q2_K · ~13.2 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q2_K · ~19.1 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q2_K · ~19.1 t/s
Show 51 more
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q4_K_M · ~6.9 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q4_K_M · ~7.1 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q4_K_M · ~7.1 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q4_K_M · ~7.1 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q6_K · ~8.8 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q8_0 · ~6.7 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q8_0 · ~29.6 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q8_0 · ~8.1 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q8_0 · ~8.3 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q8_0 · ~8.8 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q8_0 · ~8.7 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4Q8_0 · ~8.7 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q8_0 · ~8.7 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q8_0 · ~28.9 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2Q8_0 · ~8.8 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~28 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~10 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~10.6 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~10.9 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~36.3 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q8_0 · ~23.1 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8BF16 · ~6.5 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2BF16 · ~6.9 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9BF16 · ~13.2 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~10.3 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~10.3 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~10.9 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~12.4 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~6.4 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~11.8 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~8.1 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~8.8 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~9.7 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~9.7 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~9.6 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~10.4 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~10.6 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~19.4 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~18.8 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~17.4 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~19.7 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~23.3 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~25.2 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~28.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~38.1 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~38 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~51.3 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~61.2 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~74 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~152.3 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~180.2 t/s
Apple M1 Max (32GB)
With 32 GB LPDDR5 at 400 GB/s, this configuration runs 46 models natively. It comfortably runs 7B–32B models at Q4; 70B-class models typically need CPU offload.
Apple M1 Max (32GB): the base memory configuration of Apple's first Max-tier chip, launched October 2021 alongside the M1 Pro on the same 14-inch and 16-inch MacBook Pro chassis. It doubles the M1 Pro's memory bandwidth to 400 GB/s from an identical 10-core CPU (8P+2E) — Apple never sold an M1 Max below 32 GB, since pairing that much bandwidth with only 16 GB of capacity wouldn't have made sense.
With 24 GB of real headroom after the standard 8 GB reservation, this site's calculator fits Qwen3 32B natively at Q4_K_M (23.9 GB, ~15 tok/s) and Mixtral 8x7B at Q2_K (~18.3 tok/s) — both out of reach on 16 GB Apple Silicon configurations. Mid-size models get real precision headroom too: Llama 3.1 8B and Qwen2.5 7B fit at full BF16 rather than being squeezed into Q5/Q6, at roughly 19-20 tok/s — about double the identical-capacity M1 Pro 32GB's speed, since 400 GB/s is twice the M1 Pro's 200 GB/s. 46 of the 84 tracked models fit natively.
Mature Metal and MLX support. Because macOS's Metal working-set reservation stays roughly fixed regardless of installed memory, it eats a much smaller share of this 32 GB pool than it does on a 16 GB Apple Silicon config, leaving comfortable room to raise the ceiling further if a workload needs it.
Models the 32 GB configuration runs natively (46)
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q2_K · ~18.3 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q3_K_M · ~64.3 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q3_K_M · ~16.9 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q3_K_M · ~17.2 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q4_K_M · ~15 t/s
Show 41 more
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q3_K_M · ~18 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4Q3_K_M · ~18 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q3_K_M · ~18 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q4_K_M · ~49 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2Q3_K_M · ~17.6 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q4_K_M · ~46.4 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q4_K_M · ~16.3 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q5_K_M · ~15.4 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q5_K_M · ~16.2 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~36.3 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6Q5_K_M · ~34.1 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8Q6_K · ~15.2 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2Q6_K · ~15.9 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9Q6_K · ~31.8 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0Q8_0 · ~18.7 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7Q8_0 · ~18.6 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4Q8_0 · ~19.7 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6Q8_0 · ~22.4 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6Q8_0 · ~22.8 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2Q8_0 · ~20 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~15.1 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~17.5 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~18.7 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~18.7 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~18.6 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~20.4 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~20.5 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~19.4 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~18.8 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~17.4 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~19.7 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~23.3 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~25.2 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~28.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~38.1 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~38 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~51.3 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~61.2 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~74 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~152.3 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~180.2 t/s
Too large for any Apple M1 Max configuration (28)
- Mixtral 8x22B Instruct v0.1
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Maverick 400B
- Qwen3 235B-A22B (MoE)
- MiniMax M1 456B
- GPT-OSS 120B
- GLM-4.5 355B
- GLM-4.6 355B
- GLM-4.7 358B
- MiniMax M2.5 229B
- GLM-5 744B
- MiniMax M2.7 229B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- Mistral Medium 3.5 128B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- Step 3.5 Flash
- Step 3.7 Flash
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
Frequently asked questions
- How much memory does the Apple M1 Max have?
- The Apple M1 Max ships in 2 unified-memory configurations: 64 GB and 32 GB, all at 400 GB/s.
- Should I get the 32 GB or 64 GB Apple M1 Max?
- Both run everything that fits natively in 32 GB. The extra memory in the 64 GB configuration additionally fits Qwen 3.5 122B-A10B (MoE), Nemotron 3 Super 120B, Llama 4 Scout 109B, and 7 more models natively in VRAM — worth the upgrade if you plan to run any of those.
- How much VRAM does the Apple M1 Max (64GB) have?
- The Apple M1 Max (64GB) has 64 GB of LPDDR5 with 400 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M1 Max (64GB) best for?
- With 64 GB of VRAM, the Apple M1 Max (64GB) is ideal for running 70B-class models at Q4 quantization and large MoE models — a workstation sweet spot for local inference.
- What LLMs can the Apple M1 Max (64GB) run locally?
- The Apple M1 Max (64GB) can run 56 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q4_K_M, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M1 Max (64GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M1 Max (64GB) runs Llama 3.3 70B Instruct natively in VRAM at Q4_K_M quantization, achieving approximately 7.1 tokens per second.
Show 8 more questions
- Can the Apple M1 Max (64GB) run Qwen 3.6 27B?
- Yes. The Apple M1 Max (64GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 10.9 tokens per second.
- Can the Apple M1 Max (64GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M1 Max (64GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 9.7 tokens per second.
- How much VRAM does the Apple M1 Max (32GB) have?
- The Apple M1 Max (32GB) has 32 GB of LPDDR5 with 400 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M1 Max (32GB) best for?
- With 32 GB of VRAM, the Apple M1 Max (32GB) is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
- What LLMs can the Apple M1 Max (32GB) run locally?
- The Apple M1 Max (32GB) can run 46 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at BF16, Llama 3.2 3B Instruct at FP32, Llama 3.2 1B Instruct at FP32.
- Can the Apple M1 Max (32GB) run Llama 3.3 70B Instruct?
- The Apple M1 Max (32GB) does not have enough VRAM to run Llama 3.3 70B Instruct. You would need more VRAM or a lower quantization level.
- Can the Apple M1 Max (32GB) run Qwen 3.6 27B?
- Yes. The Apple M1 Max (32GB) runs Qwen 3.6 27B natively in VRAM at Q5_K_M quantization, achieving approximately 16.2 tokens per second.
- Can the Apple M1 Max (32GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M1 Max (32GB) runs Llama 3.1 8B Instruct natively in VRAM at BF16 quantization, achieving approximately 18.7 tokens per second.