CanItRun Logocanitrun.

Apple M1 Max

The Apple M1 Max ships in 32–64 GB unified-memory configurations at 400 GB/s. Across those configurations it runs 56 of our 84 tracked models natively in VRAM at 8k context.

More memory means more of our tracked models fit natively — see which configuration you need below.

Jump to:64 GB32 GB
ConfigurationBandwidthCPU coresNative models+ Offload
64 GB400 GB/s10 (8P + 2E)56 / 840
32 GB400 GB/s10 (8P + 2E)46 / 840
Vendor: Apple
Memory type: LPDDR5
Compute backend: METAL
Software: MLX gives the best performance on Apple Silicon; llama.cpp Metal backend is a solid alternative. Both are well-supported by Ollama.

How much of the Apple M1 Max's memory is actually usable?

macOS and background apps need a slice of the pool before a model gets to use it — this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:

macOS + background apps (8 GB reserved)usable for model weights + KV cacheVertical lines mark roughly where a 7B/14B/32B/... dense model at Q4_K_M lands.

Apple M1 Max (64GB)

With 64 GB LPDDR5 at 400 GB/s, this configuration runs 56 models natively. It handles 70B-class models at Q4 quantization.

Apple M1 Max (64GB): the memory-maxed configuration of Apple's first Max-tier chip, launched October 2021 as an $800 build-to-order upgrade over the 32GB base on either the 14-inch or 16-inch MacBook Pro — the only way in the M1 generation's laptop lineup to pair more than 32 GB with 400 GB/s bandwidth.

With 56 GB of real headroom after this site's standard 8 GB OS reservation, this is the first M1-family laptop configuration that clears 70B-class dense models: this site's calculator fits Llama 3.3 70B (50.8 GB at Q4_K_M, ~7.1 tok/s) and Qwen2.5 72B (52.1 GB at Q4_K_M, ~6.9 tok/s) natively. That's a different capability tier from the 32GB sibling, which tops out around 32-35B models — the extra 32 GB doesn't just add headroom, it unlocks a whole model class. 56 of the 84 tracked models fit natively, 10 more than the 32GB config's 46.

Mature Metal and MLX support, same as the 32GB sibling. macOS's Metal working-set limit (roughly 75% of unified memory, which llama.cpp and MLX both read on startup) sits around 48 GB here — comfortably above what any of this site's fitting 70B-class quants need, unlike the tighter headroom on 16-32 GB Apple Silicon configurations.

Models the 64 GB configuration runs natively (56)

Show 51 more

Apple M1 Max (32GB)

With 32 GB LPDDR5 at 400 GB/s, this configuration runs 46 models natively. It comfortably runs 7B–32B models at Q4; 70B-class models typically need CPU offload.

Apple M1 Max (32GB): the base memory configuration of Apple's first Max-tier chip, launched October 2021 alongside the M1 Pro on the same 14-inch and 16-inch MacBook Pro chassis. It doubles the M1 Pro's memory bandwidth to 400 GB/s from an identical 10-core CPU (8P+2E) — Apple never sold an M1 Max below 32 GB, since pairing that much bandwidth with only 16 GB of capacity wouldn't have made sense.

With 24 GB of real headroom after the standard 8 GB reservation, this site's calculator fits Qwen3 32B natively at Q4_K_M (23.9 GB, ~15 tok/s) and Mixtral 8x7B at Q2_K (~18.3 tok/s) — both out of reach on 16 GB Apple Silicon configurations. Mid-size models get real precision headroom too: Llama 3.1 8B and Qwen2.5 7B fit at full BF16 rather than being squeezed into Q5/Q6, at roughly 19-20 tok/s — about double the identical-capacity M1 Pro 32GB's speed, since 400 GB/s is twice the M1 Pro's 200 GB/s. 46 of the 84 tracked models fit natively.

Mature Metal and MLX support. Because macOS's Metal working-set reservation stays roughly fixed regardless of installed memory, it eats a much smaller share of this 32 GB pool than it does on a 16 GB Apple Silicon config, leaving comfortable room to raise the ceiling further if a workload needs it.

Models the 32 GB configuration runs natively (46)

Show 41 more

Too large for any Apple M1 Max configuration (28)

Frequently asked questions

How much memory does the Apple M1 Max have?
The Apple M1 Max ships in 2 unified-memory configurations: 64 GB and 32 GB, all at 400 GB/s.
Should I get the 32 GB or 64 GB Apple M1 Max?
Both run everything that fits natively in 32 GB. The extra memory in the 64 GB configuration additionally fits Qwen 3.5 122B-A10B (MoE), Nemotron 3 Super 120B, Llama 4 Scout 109B, and 7 more models natively in VRAM — worth the upgrade if you plan to run any of those.
How much VRAM does the Apple M1 Max (64GB) have?
The Apple M1 Max (64GB) has 64 GB of LPDDR5 with 400 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
What is the Apple M1 Max (64GB) best for?
With 64 GB of VRAM, the Apple M1 Max (64GB) is ideal for running 70B-class models at Q4 quantization and large MoE models — a workstation sweet spot for local inference.
What LLMs can the Apple M1 Max (64GB) run locally?
The Apple M1 Max (64GB) can run 56 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q4_K_M, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
Can the Apple M1 Max (64GB) run Llama 3.3 70B Instruct?
Yes. The Apple M1 Max (64GB) runs Llama 3.3 70B Instruct natively in VRAM at Q4_K_M quantization, achieving approximately 7.1 tokens per second.
Show 8 more questions
Can the Apple M1 Max (64GB) run Qwen 3.6 27B?
Yes. The Apple M1 Max (64GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 10.9 tokens per second.
Can the Apple M1 Max (64GB) run Llama 3.1 8B Instruct?
Yes. The Apple M1 Max (64GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 9.7 tokens per second.
How much VRAM does the Apple M1 Max (32GB) have?
The Apple M1 Max (32GB) has 32 GB of LPDDR5 with 400 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
What is the Apple M1 Max (32GB) best for?
With 32 GB of VRAM, the Apple M1 Max (32GB) is well-suited for running 7B–32B models at Q4 with room for context, making it a great all-rounder for local LLM inference.
What LLMs can the Apple M1 Max (32GB) run locally?
The Apple M1 Max (32GB) can run 46 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.1 8B Instruct at BF16, Llama 3.2 3B Instruct at FP32, Llama 3.2 1B Instruct at FP32.
Can the Apple M1 Max (32GB) run Llama 3.3 70B Instruct?
The Apple M1 Max (32GB) does not have enough VRAM to run Llama 3.3 70B Instruct. You would need more VRAM or a lower quantization level.
Can the Apple M1 Max (32GB) run Qwen 3.6 27B?
Yes. The Apple M1 Max (32GB) runs Qwen 3.6 27B natively in VRAM at Q5_K_M quantization, achieving approximately 16.2 tokens per second.
Can the Apple M1 Max (32GB) run Llama 3.1 8B Instruct?
Yes. The Apple M1 Max (32GB) runs Llama 3.1 8B Instruct natively in VRAM at BF16 quantization, achieving approximately 18.7 tokens per second.