Apple M3 Ultra
The Apple M3 Ultra ships in 96–512 GB unified-memory configurations at 819 GB/s. Across those configurations it runs 82 of our 84 tracked models natively in VRAM at 8k context.
More memory means more of our tracked models fit natively — see which configuration you need below.
| Configuration | Bandwidth | CPU cores | Native models | + Offload |
|---|---|---|---|---|
| 512 GB | 819 GB/s | 32 (24P + 8E) | 82 / 84 | 0 |
| 256 GB | 819 GB/s | 28 (20P + 8E) | 73 / 84 | 0 |
| 96 GB | 819 GB/s | 28 (20P + 8E) | 61 / 84 | 0 |
How much of the Apple M3 Ultra's memory is actually usable?
macOS and background apps need a slice of the pool before a model gets to use it — this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:
M3 Ultra is really two M3 Max dies, and it shows in the bandwidth
Apple builds M3 Ultra by fusing two complete M3 Max dies together with UltraFusion, a 2.5 TB/s silicon interposer first used on the M1 generation. The bandwidth math bears that out directly against the M3 Max lineup it's built from:
819 GB/s beats even the fastest full-die M3 Max (400 GB/s) by 2.05x, and the binned 96GB M3 Max (300 GB/s) by 2.73x — because UltraFusion is, functionally, twice the memory controllers rather than a new design. That's not just a spec-sheet number: at the identical 88 GB of real headroom the 96GB configurations of both chips share, this site's calculator projects GPT-OSS 120B decoding at 61.5 tok/s on this Ultra configuration versus 22.5 tok/s on the 96GB Max — real-time speed on a 117B-parameter MoE model, on a chip that's still a laptop-class die times two, not a discrete accelerator.
Apple M3 Ultra (512GB)
With 512 GB LPDDR5X at 819 GB/s, this configuration runs 82 models natively. It handles the largest open-weight models, including 405B-class frontier releases, at some quantization.
Apple M3 Ultra (512GB): the maximum-memory Mac Studio configuration Apple introduced in March 2025, and the only M3 Ultra memory tier that requires the 32-core CPU (24P+8E)/80-core GPU upgrade die — Apple never offered 512GB on the base 28-core/60-core chip. M3 Ultra is two complete M3 Max dies fused with UltraFusion, a 2.5 TB/s silicon interposer (184 billion transistors total) — the same packaging trick Apple has used for every Ultra chip since the M1 generation — which is why its 819 GB/s bandwidth tracks roughly double a single M3 Max die rather than a new memory architecture.
With 504 GB of real headroom after this site's standard 8 GB reservation, this configuration fits 82 of the 84 models tracked on this site natively in VRAM. DeepSeek V3 and DeepSeek R1 (671B parameters) both fit at Q4_K_M (458.25 GB, 8.7 tok/s), and Kimi K2.6 — a full 1 trillion parameters — fits at Q2_K (429.73 GB, 15.1 tok/s), one of the very few ways to hold a trillion-parameter model's weights on a single consumer-purchasable machine at all. GLM 4.5 and GLM 4.6 (355B) run at Q8_0 (426.11 GB, 5.6 tok/s) here, full-precision headroom the smaller 256GB tier can't offer them.
MLX and llama.cpp's Metal backend both run frontier MoE models here, but neither parallelizes expert routing across chips the way a multi-GPU CUDA rack would — everything is served off one chip's 819 GB/s memory bus. That's plenty of raw bandwidth for these token counts, but it means the extra headroom over the 256GB tier is almost entirely about which models fit at all, not about squeezing more tokens/sec out of ones that already fit elsewhere.
Models the 512 GB configuration runs natively (82)
- MiMo V2.5 Pro1020B · MMLU-Pro 68.5Q2_K · ~11.5 t/s
- Kimi K2.61000B · MMLU-Pro 87.2Q2_K · ~15.1 t/s
- Kimi K2.51000B · MMLU-Pro 87.1Q2_K · ~15.7 t/s
- Inkling975B · MMLU-Pro —Q2_K · ~12.1 t/s
- GLM-5.1 754B754B · MMLU-Pro 86.5Q3_K_M · ~8.1 t/s
Show 77 more
- GLM-5.2 753B753B · MMLU-Pro 80.6Q3_K_M · ~8.8 t/s
- GLM-5 744B744B · MMLU-Pro 85.7Q3_K_M · ~8.8 t/s
- DeepSeek V3 671B671B · MMLU-Pro 75.9Q4_K_M · ~8.7 t/s
- DeepSeek R1 671B671B · MMLU-Pro 85.0Q4_K_M · ~8.7 t/s
- Nemotron 3 Ultra 550B-A55B550B · MMLU-Pro 86.8Q5_K_M · ~5 t/s
- MiniMax M1 456B456B · MMLU-Pro 81.1Q6_K · ~5.1 t/s
- MiniMax M3428B · MMLU-Pro —Q6_K · ~10.2 t/s
- Llama 3.1 405B Instruct405B · MMLU-Pro 73.3Q8_0 · ~1.5 t/s
- Llama 4 Maverick 400B400B · MMLU-Pro 80.5Q8_0 · ~10.2 t/s
- GLM-4.7 358B358B · MMLU-Pro 84.3Q8_0 · ~5.6 t/s
- GLM-4.5 355B355B · MMLU-Pro 84.6Q8_0 · ~5.6 t/s
- GLM-4.6 355B355B · MMLU-Pro —Q8_0 · ~5.6 t/s
- DeepSeek V4 Flash 284B284B · MMLU-Pro 86.3Q8_0 · ~14 t/s
- Qwen3 235B-A22B (MoE)235B · MMLU-Pro 84.4Q8_0 · ~8.2 t/s
- MiniMax M2.5 229B229B · MMLU-Pro 84.8Q8_0 · ~17.5 t/s
- MiniMax M2.7 229B229B · MMLU-Pro 86.0Q8_0 · ~17.5 t/s
- Step 3.7 Flash198B · MMLU-Pro —BF16 · ~8.8 t/s
- Step 3.5 Flash196.81B · MMLU-Pro 84.4BF16 · ~8.8 t/s
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0BF16 · ~2.5 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro —BF16 · ~2.5 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7BF16 · ~9.5 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7BF16 · ~8.1 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7BF16 · ~19.1 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3FP32 · ~2.9 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4FP32 · ~4.1 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9FP32 · ~4.1 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1FP32 · ~2.3 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9FP32 · ~2.3 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0FP32 · ~2.3 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4FP32 · ~2.3 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7FP32 · ~3.8 t/s
- Command-R 35B35B · MMLU-Pro 33.0FP32 · ~4.3 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3FP32 · ~16.3 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2FP32 · ~4.6 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0FP32 · ~4.7 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5FP32 · ~4.9 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0FP32 · ~5 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4FP32 · ~5 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0FP32 · ~5 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3FP32 · ~16.2 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2FP32 · ~5.2 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5FP32 · ~16.1 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0FP32 · ~5.9 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5FP32 · ~6 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2FP32 · ~6 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~74.4 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6FP32 · ~12.8 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8FP32 · ~6.7 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2FP32 · ~7.2 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9FP32 · ~13.6 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0FP32 · ~10.8 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7FP32 · ~10.8 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4FP32 · ~11.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6FP32 · ~13.1 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~13.1 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2FP32 · ~12.8 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~16.5 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~18.1 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~19.8 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~19.8 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~19.7 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~21.2 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~21.8 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~39.7 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~38.5 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~35.6 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~40.3 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~47.7 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~51.6 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~58.1 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~78 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~77.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~105.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~125.3 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~151.4 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~311.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~369 t/s
Apple M3 Ultra (256GB)
With 256 GB LPDDR5X at 819 GB/s, this configuration runs 73 models natively. It handles the largest open-weight models, including 405B-class frontier releases, at some quantization.
Apple M3 Ultra (256GB): the mid-memory Mac Studio configuration Apple introduced alongside the 96GB base and 512GB flagship in March 2025 — available as a BTO memory upgrade on either the base 28-core CPU (20P+8E)/60-core GPU die or the 32-core/80-core upgrade, at the identical 819 GB/s bandwidth either way, since M3 Ultra's UltraFusion-fused dual-die design doesn't bin bandwidth by CPU/GPU core count the way M3 Max bins by memory size.
With 248 GB of real headroom, this is the smallest Apple Silicon configuration this site tracks that fits a 350B+-parameter frontier MoE model at all: GLM 4.5 and GLM 4.6 (355B) run at Q4_K_M (245.60 GB, 9.6 tok/s). DeepSeek V3 and Kimi K2.6 still don't fit here — both need the 512GB tier. The MoE-vs-dense gap is stark at frontier scale: Llama 4 Maverick (400B, MoE) fits at Q3_K_M (220.00 GB) at 20.9 tok/s, while Llama 3.1 405B (405B, dense) needs the same Q3_K_M precision at almost the same footprint (222.92 GB) but manages only 3.3 tok/s — a 6x speed gap between two similarly-sized models, because MoE decode only reads its active parameters off the memory bus.
Same MLX and llama.cpp Metal maturity as the rest of the M3 Ultra lineup. At this file scale — GLM 4.5's Q4_K_M build alone is a quarter-terabyte download — a fast SSD and network connection matter as much as the chip; this is the one tier where local storage speed genuinely gates how quickly a new model gets from download to first token.
Models the 256 GB configuration runs natively (73)
- Nemotron 3 Ultra 550B-A55B550B · MMLU-Pro 86.8Q2_K · ~9.3 t/s
- MiniMax M1 456B456B · MMLU-Pro 81.1Q2_K · ~10.7 t/s
- MiniMax M3428B · MMLU-Pro —Q3_K_M · ~17.3 t/s
- Llama 3.1 405B Instruct405B · MMLU-Pro 73.3Q3_K_M · ~3.3 t/s
- Llama 4 Maverick 400B400B · MMLU-Pro 80.5Q3_K_M · ~20.9 t/s
Show 68 more
- GLM-4.7 358B358B · MMLU-Pro 84.3Q4_K_M · ~9.6 t/s
- GLM-4.5 355B355B · MMLU-Pro 84.6Q4_K_M · ~9.6 t/s
- GLM-4.6 355B355B · MMLU-Pro —Q4_K_M · ~9.6 t/s
- DeepSeek V4 Flash 284B284B · MMLU-Pro 86.3Q5_K_M · ~20.8 t/s
- Qwen3 235B-A22B (MoE)235B · MMLU-Pro 84.4Q6_K · ~10.6 t/s
- MiniMax M2.5 229B229B · MMLU-Pro 84.8Q6_K · ~22.3 t/s
- MiniMax M2.7 229B229B · MMLU-Pro 86.0Q6_K · ~22.3 t/s
- Step 3.7 Flash198B · MMLU-Pro —Q8_0 · ~16.2 t/s
- Step 3.5 Flash196.81B · MMLU-Pro 84.4Q8_0 · ~16.2 t/s
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0Q8_0 · ~4.7 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro —Q8_0 · ~4.7 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q8_0 · ~17.4 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q8_0 · ~15.1 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7Q8_0 · ~35.7 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3BF16 · ~5.6 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4BF16 · ~8 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9BF16 · ~8 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1BF16 · ~4.5 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9BF16 · ~4.6 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0BF16 · ~4.6 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4BF16 · ~4.6 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7FP32 · ~3.8 t/s
- Command-R 35B35B · MMLU-Pro 33.0FP32 · ~4.3 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3FP32 · ~16.3 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2FP32 · ~4.6 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0FP32 · ~4.7 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5FP32 · ~4.9 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0FP32 · ~5 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4FP32 · ~5 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0FP32 · ~5 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3FP32 · ~16.2 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2FP32 · ~5.2 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5FP32 · ~16.1 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0FP32 · ~5.9 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5FP32 · ~6 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2FP32 · ~6 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~74.4 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6FP32 · ~12.8 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8FP32 · ~6.7 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2FP32 · ~7.2 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9FP32 · ~13.6 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0FP32 · ~10.8 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7FP32 · ~10.8 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4FP32 · ~11.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6FP32 · ~13.1 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~13.1 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2FP32 · ~12.8 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~16.5 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~18.1 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~19.8 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~19.8 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~19.7 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~21.2 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~21.8 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~39.7 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~38.5 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~35.6 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~40.3 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~47.7 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~51.6 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~58.1 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~78 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~77.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~105.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~125.3 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~151.4 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~311.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~369 t/s
Apple M3 Ultra (96GB)
With 96 GB LPDDR5X at 819 GB/s, this configuration runs 61 models natively. It runs 70B-class dense models and large MoE models entirely in VRAM.
Apple M3 Ultra (96GB): the base Mac Studio configuration Apple introduced in March 2025, standard on the 28-core CPU (20P+8E)/60-core GPU die. Because M3 Ultra is built from two complete M3 Max dies fused with UltraFusion, this configuration shares its exact 88 GB of real headroom with the 96GB M3 Max — the difference is that the Max's 96GB build pairs with a binned, bandwidth-reduced die (300 GB/s), while the Ultra reaches 819 GB/s regardless of memory size: 2.73x faster at an identical usable capacity.
Because usable capacity matches the 96GB M3 Max exactly, this configuration fits the same 61 of 84 tracked models natively — what changes is speed, not the model list. This site's calculator projects Llama 3.3 70B at Q8_0 (86.35 GB) running 8.5 tok/s here versus 3.1 tok/s on the 96GB M3 Max, and Qwen 3.6 27B at BF16 (61.08 GB) at 12.0 tok/s versus 4.4 tok/s — both track the 2.73x bandwidth ratio almost exactly. MoE models benefit even more in absolute terms: GPT-OSS 120B reaches Q4_K_M (80.15 GB) at 61.5 tok/s here, genuinely real-time on a 117B-parameter model, versus 22.5 tok/s on the bandwidth-limited Max.
MLX and llama.cpp's Metal backend are both mature on this chip. Because usable capacity matches the 96GB M3 Max, the same macOS Metal gotcha applies: the default working-set limit (~75% of unified memory, about 72 GB here) sits below the 86.35 GB the Q8_0 Llama 3.3 70B fit above needs, so raising it with `sudo sysctl iogpu.wired_limit_mb=90112` (leaving 8 GB for macOS) is necessary before loading anything close to that ceiling — the exact same fix and figure as the 96GB M3 Max.
Models the 96 GB configuration runs natively (61)
- Step 3.7 Flash198B · MMLU-Pro —Q2_K · ~42.3 t/s
- Step 3.5 Flash196.81B · MMLU-Pro 84.4Q2_K · ~42.3 t/s
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0Q3_K_M · ~10.2 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro —Q3_K_M · ~10.2 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q4_K_M · ~29.2 t/s
Show 56 more
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q4_K_M · ~26.1 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7Q4_K_M · ~61.5 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q4_K_M · ~17.6 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q5_K_M · ~21.8 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q5_K_M · ~21.8 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q6_K · ~10.6 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q8_0 · ~8.5 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q8_0 · ~8.5 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q8_0 · ~8.5 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q8_0 · ~14 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q8_0 · ~13.7 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3BF16 · ~32.5 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2BF16 · ~9.1 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0BF16 · ~9.3 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5BF16 · ~9.8 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0BF16 · ~9.8 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 50.4BF16 · ~9.8 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0BF16 · ~9.8 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3BF16 · ~32.1 t/s
- Gemma 4 31B31B · MMLU-Pro 85.2BF16 · ~10 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5BF16 · ~31.5 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0BF16 · ~11.4 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5BF16 · ~11.8 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2BF16 · ~12 t/s
- Bonsai 27B27B · MMLU-Pro 81.5Ternary (Q2_0) · ~74.4 t/s
- Gemma 4 26B (MoE)26B · MMLU-Pro 82.6BF16 · ~25.5 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8BF16 · ~13.3 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2BF16 · ~14.2 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9BF16 · ~27.1 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0FP32 · ~10.8 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7FP32 · ~10.8 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4FP32 · ~11.4 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6FP32 · ~13.1 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6FP32 · ~13.1 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2FP32 · ~12.8 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0FP32 · ~16.5 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5FP32 · ~18.1 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3FP32 · ~19.8 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0FP32 · ~19.8 t/s
- Qwen3 8B8B · MMLU-Pro 56.7FP32 · ~19.7 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3FP32 · ~21.2 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0FP32 · ~21.8 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6FP32 · ~39.7 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4FP32 · ~38.5 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4FP32 · ~35.6 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3FP32 · ~40.3 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0FP32 · ~47.7 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4FP32 · ~51.6 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8FP32 · ~58.1 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0FP32 · ~78 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0FP32 · ~77.9 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8FP32 · ~105.1 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5FP32 · ~125.3 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7FP32 · ~151.4 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0FP32 · ~311.9 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0FP32 · ~369 t/s
Too large for any Apple M3 Ultra configuration (2)
Compare Apple M3 Ultra with other GPUs
- Apple M3 Ultra (512GB)vsApple M2 Ultra (192GB)+320 GB VRAM
- Apple M3 Ultra (96GB)vsNVIDIA RTX 5090+64 GB VRAM
- Apple M3 Ultra (96GB)vsNVIDIA RTX Pro 600096 GB each
- Apple M3 Ultra (96GB)vsNVIDIA DGX Spark (128GB)-32 GB VRAM
- Apple M3 Ultra (96GB)vsApple M4 Max (128GB)-32 GB VRAM
- Apple M3 Ultra (96GB)vsAMD Strix Halo (96GB)96 GB each
Continue reading
Frequently asked questions
- How much memory does the Apple M3 Ultra have?
- The Apple M3 Ultra ships in 3 unified-memory configurations: 512 GB, 256 GB, 96 GB, all at 819 GB/s.
- Should I get the 96 GB or 512 GB Apple M3 Ultra?
- Both run everything that fits natively in 96 GB. The extra memory in the 512 GB configuration additionally fits MiMo V2.5 Pro, Kimi K2.6, Kimi K2.5, and 18 more models natively in VRAM — worth the upgrade if you plan to run any of those.
- How much VRAM does the Apple M3 Ultra (512GB) have?
- The Apple M3 Ultra (512GB) has 512 GB of LPDDR5X with 819 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M3 Ultra (512GB) best for?
- With 512 GB of unified memory, the Apple M3 Ultra (512GB) is a high-capacity workstation platform capable of running the largest open-weight models (70B–405B) at high quantization with ample context.
- What LLMs can the Apple M3 Ultra (512GB) run locally?
- The Apple M3 Ultra (512GB) can run 82 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at FP32, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M3 Ultra (512GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M3 Ultra (512GB) runs Llama 3.3 70B Instruct natively in VRAM at FP32 quantization, achieving approximately 2.3 tokens per second.
Show 14 more questions
- Can the Apple M3 Ultra (512GB) run Qwen 3.6 27B?
- Yes. The Apple M3 Ultra (512GB) runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 6 tokens per second.
- Can the Apple M3 Ultra (512GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M3 Ultra (512GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 19.8 tokens per second.
- How much VRAM does the Apple M3 Ultra (256GB) have?
- The Apple M3 Ultra (256GB) has 256 GB of LPDDR5X with 819 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M3 Ultra (256GB) best for?
- With 256 GB of unified memory, the Apple M3 Ultra (256GB) is a high-capacity workstation platform capable of running the largest open-weight models (70B–405B) at high quantization with ample context.
- What LLMs can the Apple M3 Ultra (256GB) run locally?
- The Apple M3 Ultra (256GB) can run 73 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at BF16, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M3 Ultra (256GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M3 Ultra (256GB) runs Llama 3.3 70B Instruct natively in VRAM at BF16 quantization, achieving approximately 4.6 tokens per second.
- Can the Apple M3 Ultra (256GB) run Qwen 3.6 27B?
- Yes. The Apple M3 Ultra (256GB) runs Qwen 3.6 27B natively in VRAM at FP32 quantization, achieving approximately 6 tokens per second.
- Can the Apple M3 Ultra (256GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M3 Ultra (256GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 19.8 tokens per second.
- How much VRAM does the Apple M3 Ultra (96GB) have?
- The Apple M3 Ultra (96GB) has 96 GB of LPDDR5X with 819 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M3 Ultra (96GB) best for?
- With 96 GB of unified memory, the Apple M3 Ultra (96GB) is a high-capacity workstation platform that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
- What LLMs can the Apple M3 Ultra (96GB) run locally?
- The Apple M3 Ultra (96GB) can run 61 of the 84 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Llama 3.3 70B Instruct at Q8_0, Llama 3.1 8B Instruct at FP32, Llama 3.2 3B Instruct at FP32.
- Can the Apple M3 Ultra (96GB) run Llama 3.3 70B Instruct?
- Yes. The Apple M3 Ultra (96GB) runs Llama 3.3 70B Instruct natively in VRAM at Q8_0 quantization, achieving approximately 8.5 tokens per second.
- Can the Apple M3 Ultra (96GB) run Qwen 3.6 27B?
- Yes. The Apple M3 Ultra (96GB) runs Qwen 3.6 27B natively in VRAM at BF16 quantization, achieving approximately 12 tokens per second.
- Can the Apple M3 Ultra (96GB) run Llama 3.1 8B Instruct?
- Yes. The Apple M3 Ultra (96GB) runs Llama 3.1 8B Instruct natively in VRAM at FP32 quantization, achieving approximately 19.8 tokens per second.