CanItRun Logocanitrun.

Apple M1 Ultra

The Apple M1 Ultra ships in 64–128 GB unified-memory configurations at 800 GB/s. Across those configurations it runs 64 of our 84 tracked models natively in VRAM at 8k context.

More memory means more of our tracked models fit natively — see which configuration you need below.

Jump to:128 GB64 GB
ConfigurationBandwidthCPU coresNative models+ Offload
128 GB800 GB/s20 (16P + 4E)64 / 840
64 GB800 GB/s20 (16P + 4E)56 / 840
Vendor: Apple
Memory type: LPDDR5
Compute backend: METAL
Software: MLX gives the best performance on Apple Silicon; llama.cpp Metal backend is a solid alternative. Both are well-supported by Ollama.

How much of the Apple M1 Ultra's memory is actually usable?

macOS and background apps need a slice of the pool before a model gets to use it — this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:

macOS + background apps (8 GB reserved)usable for model weights + KV cacheVertical lines mark roughly where a 7B/14B/32B/... dense model at Q4_K_M lands.

Apple M1 Ultra (128GB)

With 128 GB LPDDR5 at 800 GB/s, this configuration runs 64 models natively. It runs 70B-class dense models and large MoE models entirely in VRAM.

Apple M1 Ultra (128GB): the maximum-memory build of the M1 generation, and the only M1-era Mac that can hold a 235B-class MoE model's weights at all. It shares the 64GB Ultra's 20-core CPU, 800 GB/s LPDDR5 bandwidth, and UltraFusion dual-die design (two M1 Max dies fused with a 2.5 TB/s silicon interposer) — Apple doubled the memory pool without changing anything else, as a build-to-order upgrade on the Mac Studio that launched at $3,999 in March 2022.

120 GB of real headroom (after this site's standard 8 GB unified-memory reservation) unlocks a tier no smaller M1-family configuration reaches: this site's calculator fits Qwen3 235B-A22B, a 235-billion-parameter mixture-of-experts model, at Q2_K quantization and roughly 21.7 tok/s — the 64GB Ultra can't hold its weights at any quantization. On models that already fit both configurations, though, more memory buys quality rather than speed: Llama 3.3 70B runs at Q4_K_M (~14.1 tok/s) on the 64GB Ultra but at the higher-precision Q8_0 (~8.3 tok/s) here, since this site always recommends the best quantization that fits — extra headroom lets it pick a heavier format, which reads more bytes per token at the identical 800 GB/s bandwidth and so runs slower, not faster.

Same MLX and llama.cpp Metal support as the 64GB Ultra. Mixture-of-experts models like Qwen3 235B-A22B benefit less from Apple Silicon's bandwidth than a dense model of the same total size — llama.cpp and MLX still keep every expert's weights resident in memory even though only a fraction activate per token, so the capacity story here is about fitting the weights at all, not about MoE inference being unusually fast on Metal.

Models the 128 GB configuration runs natively (64)

Show 59 more

Apple M1 Ultra (64GB)

With 64 GB LPDDR5 at 800 GB/s, this configuration runs 56 models natively. It handles 70B-class models at Q4 quantization.

Apple M1 Ultra (64GB): the base memory configuration of Apple's first Ultra-tier chip — two M1 Max dies fused with Apple's UltraFusion interposer for 2.5 TB/s of chip-to-chip bandwidth, launched March 2022 exclusively in the Mac Studio desktop starting at $3,999. It doubles the M1 Max's 400 GB/s bandwidth to 800 GB/s while sharing the M1 Max 64GB config's memory ceiling.

Same 56 GB of real headroom as the M1 Max 64GB config — capacity is identical, so the same 56 of the 84 models tracked on this site fit natively, including Llama 3.3 70B and Qwen2.5 72B at Q4_K_M. What changes is speed: this site's calculator projects roughly 14 tok/s for both 70B-class models here, almost exactly double the M1 Max 64GB's ~7 tok/s, because doubling bandwidth doubles decode throughput on a workload this bandwidth-bound, without changing what fits.

Full MLX and llama.cpp Metal support. MLX in particular was built with Apple Silicon's unified-memory bandwidth in mind and tends to extract a larger share of the chip's rated bandwidth than llama.cpp's more general Metal backend — worth trying both if decode speed matters more than llama.cpp's broader format support.

Models the 64 GB configuration runs natively (56)

Show 51 more

Too large for any Apple M1 Ultra configuration (20)

Compare Apple M1 Ultra with other GPUs

Frequently asked questions

How much memory does the Apple M1 Ultra have?
The Apple M1 Ultra ships in 2 unified-memory configurations: 128 GB and 64 GB, all at 800 GB/s.
Should I get the 64 GB or 128 GB Apple M1 Ultra?
Both run everything that fits natively in 64 GB. The extra memory in the 128 GB configuration additionally fits Qwen3 235B-A22B (MoE), MiniMax M2.5 229B, MiniMax M2.7 229B, and 5 more models natively in VRAM — worth the upgrade if you plan to run any of those.
How much VRAM does the Apple M1 Ultra (128GB) have?
The Apple M1 Ultra (128GB) has 128 GB of LPDDR5 with 800 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
What is the Apple M1 Ultra (128GB) best for?
With 128 GB of unified memory, the Apple M1 Ultra (128GB) is a high-capacity workstation platform that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
What LLMs can the Apple M1 Ultra (128GB) run locally?
The Apple M1 Ultra (128GB) can run 64 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q8_0, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
Can the Apple M1 Ultra (128GB) run Llama 3.3 70B Instruct?
Yes. The Apple M1 Ultra (128GB) runs Llama 3.3 70B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 8.3 tokens per second.
Show 8 more questions
Can the Apple M1 Ultra (128GB) run Qwen 3.6 27B?
Yes. The Apple M1 Ultra (128GB) runs Qwen 3.6 27B natively in VRAM at BF16 quantization, achieving approximately 11.7 tokens per second.
Can the Apple M1 Ultra (128GB) run Llama 3.1 8B Instruct?
Yes. The Apple M1 Ultra (128GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 19.4 tokens per second.
How much VRAM does the Apple M1 Ultra (64GB) have?
The Apple M1 Ultra (64GB) has 64 GB of LPDDR5 with 800 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
What is the Apple M1 Ultra (64GB) best for?
With 64 GB of VRAM, the Apple M1 Ultra (64GB) is ideal for running 70B-class models at Q4 quantization and large MoE models — a workstation sweet spot for local inference.
What LLMs can the Apple M1 Ultra (64GB) run locally?
The Apple M1 Ultra (64GB) can run 56 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q4_K_M, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
Can the Apple M1 Ultra (64GB) run Llama 3.3 70B Instruct?
Yes. The Apple M1 Ultra (64GB) runs Llama 3.3 70B Instruct natively in VRAM at Q4_K_M quantization, achieving approximately 14.1 tokens per second.
Can the Apple M1 Ultra (64GB) run Qwen 3.6 27B?
Yes. The Apple M1 Ultra (64GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 21.9 tokens per second.
Can the Apple M1 Ultra (64GB) run Llama 3.1 8B Instruct?
Yes. The Apple M1 Ultra (64GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 19.4 tokens per second.