Apple M1 Ultra
The Apple M1 Ultra ships in 64–128 GB unified-memory configurations at 800 GB/s. Across those configurations it runs 73 of our 99 tracked models natively in VRAM at 8k context.
More memory means more of our tracked models fit natively; see which configuration you need below.
| Configuration | Bandwidth | CPU cores | Native models | + Offload |
|---|---|---|---|---|
| 128 GB | 800 GB/s | 20 (16P + 4E) | 73 / 99 | 0 |
| 64 GB | 800 GB/s | 20 (16P + 4E) | 63 / 99 | 0 |
How much of the Apple M1 Ultra's memory is actually usable?
macOS and background apps need a slice of the pool before a model gets to use it: this site reserves 8GB on every unified-memory GPU, the same baseline used everywhere else on this site. What's left is real headroom for a model's weights and KV cache:
Apple M1 Ultra (128GB)
With 128 GB LPDDR5 at 800 GB/s, this configuration runs 73 models natively. 120 GB of real headroom (after this site's standard 8 GB unified-memory reservation) unlocks a tier no smaller M1-family configuration reaches: this site's calculator fits Qwen3 235B-A22B, a 235-billion-parameter mixture-of-experts model, at Q2_K quantization and roughly 21.7 tok/s; the 64GB Ultra can't hold its weights at any quantization. On models that already fit both configurations, though, more memory buys quality rather than speed: Llama 3.3 70B runs at Q4_K_M (~14.1 tok/s) on the 64GB Ultra but at the higher-precision Q8_0 (~8.3 tok/s) here, since this site always recommends the best quantization that fits; extra headroom lets it pick a heavier format, which reads more bytes per token at the identical 800 GB/s bandwidth and so runs slower, not faster.
Apple M1 Ultra (128GB): the maximum-memory build of the M1 generation, and the only M1-era Mac that can hold a 235B-class MoE model's weights at all. It shares the 64GB Ultra's 20-core CPU, 800 GB/s LPDDR5 bandwidth, and UltraFusion dual-die design (two M1 Max dies fused with a 2.5 TB/s silicon interposer); Apple doubled the memory pool without changing anything else, as a build-to-order upgrade on the Mac Studio that launched at $3,999 in March 2022.
120 GB of real headroom (after this site's standard 8 GB unified-memory reservation) unlocks a tier no smaller M1-family configuration reaches: this site's calculator fits Qwen3 235B-A22B, a 235-billion-parameter mixture-of-experts model, at Q2_K quantization and roughly 21.7 tok/s; the 64GB Ultra can't hold its weights at any quantization. On models that already fit both configurations, though, more memory buys quality rather than speed: Llama 3.3 70B runs at Q4_K_M (~14.1 tok/s) on the 64GB Ultra but at the higher-precision Q8_0 (~8.3 tok/s) here, since this site always recommends the best quantization that fits; extra headroom lets it pick a heavier format, which reads more bytes per token at the identical 800 GB/s bandwidth and so runs slower, not faster.
Same MLX and llama.cpp Metal support as the 64GB Ultra. Mixture-of-experts models like Qwen3 235B-A22B benefit less from Apple Silicon's bandwidth than a dense model of the same total size: llama.cpp and MLX still keep every expert's weights resident in memory even though only a fraction activate per token, so the capacity story here is about fitting the weights at all, not about MoE inference being unusually fast on Metal.
Models the 128 GB configuration runs natively (73)
- DeepSeek V4 Flash 0731 284B284B · MMLU-Pro N/AUD-IQ3_XXS · ~40.1 t/s
- Qwen3 235B-A22B (MoE)235B · MMLU-Pro 84.4Q2_K · ~21.7 t/s
- MiniMax M2.5 229B229B · MMLU-Pro 84.8Q2_K · ~43.3 t/s
- MiniMax M2.7 229B229B · MMLU-Pro 86.0Q2_K · ~43.3 t/s
- Step 3.7 Flash198B · MMLU-Pro N/AQ3_K_M · ~33.4 t/s
Show 68 more
- Step 3.5 Flash196.81B · MMLU-Pro 84.4Q3_K_M · ~33.4 t/s
- Qwen3.8-Flash-Next180B · MMLU-Pro N/AQ3_K_M · ~65.2 t/s
- Mixtral 8x22B Instruct v0.1141B · MMLU-Pro 40.0Q5_K_M · ~6.8 t/s
- Mistral Medium 3.5 128B128B · MMLU-Pro N/AQ5_K_M · ~6.8 t/s
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q6_K · ~23.2 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q6_K · ~19.1 t/s
- GPT-OSS 120B117B · MMLU-Pro 80.7Q6_K · ~44.9 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q6_K · ~13 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q6_K · ~18.6 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q6_K · ~18.6 t/s
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q8_0 · ~8.1 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q8_0 · ~8.3 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q8_0 · ~8.3 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q8_0 · ~8.3 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7BF16 · ~7.4 t/s
- Command-R 35B35B · MMLU-Pro 33.0BF16 · ~7.9 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3BF16 · ~31.7 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2BF16 · ~8.9 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/ABF16 · ~31.7 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0BF16 · ~9 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5BF16 · ~9.6 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0BF16 · ~9.5 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3BF16 · ~9.5 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0BF16 · ~9.5 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3BF16 · ~31.3 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2BF16 · ~10.2 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5BF16 · ~30.8 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6BF16 · ~31.9 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/ABF16 · ~11.5 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0BF16 · ~11.1 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5BF16 · ~11.5 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2BF16 · ~11.7 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2BF16 · ~11.7 t/s
- Bonsai 27B27B · MMLU-Pro ~81.5Ternary (Q2_0) · ~82.7 t/s
- Bonsai 2 27B27B · MMLU-Pro N/ATernary (Q2_0) · ~99.4 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/ABF16 · ~11.7 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6BF16 · ~24.9 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8BF16 · ~13 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2BF16 · ~13.8 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9BF16 · ~26.4 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~20.7 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~20.6 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~21.8 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~24.9 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~25.1 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~23.5 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~30.2 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~35 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ABF16 · ~35 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~37.5 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~37.5 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~37.2 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~40.8 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~41.1 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~75.3 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~71.1 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~59.1 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~73.8 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~87.2 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~98.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~105.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~145.4 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~127.7 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~197.8 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~232.9 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~275 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~581.5 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~606.3 t/s
Apple M1 Ultra (64GB)
With 64 GB LPDDR5 at 800 GB/s, this configuration runs 63 models natively. Same 56 GB of real headroom as the M1 Max 64GB config; capacity is identical, so the same 63 of the 99 models tracked on this site fit natively, including Llama 3.3 70B and Qwen2.5 72B at Q4_K_M. What changes is speed: this site's calculator projects roughly 14 tok/s for both 70B-class models here, almost exactly double the M1 Max 64GB's ~7 tok/s, because doubling bandwidth doubles decode throughput on a workload this bandwidth-bound, without changing what fits.
Apple M1 Ultra (64GB): the base memory configuration of Apple's first Ultra-tier chip, two M1 Max dies fused with Apple's UltraFusion interposer for 2.5 TB/s of chip-to-chip bandwidth, launched March 2022 exclusively in the Mac Studio desktop starting at $3,999. It doubles the M1 Max's 400 GB/s bandwidth to 800 GB/s while sharing the M1 Max 64GB config's memory ceiling.
Same 56 GB of real headroom as the M1 Max 64GB config; capacity is identical, so the same 63 of the 99 models tracked on this site fit natively, including Llama 3.3 70B and Qwen2.5 72B at Q4_K_M. What changes is speed: this site's calculator projects roughly 14 tok/s for both 70B-class models here, almost exactly double the M1 Max 64GB's ~7 tok/s, because doubling bandwidth doubles decode throughput on a workload this bandwidth-bound, without changing what fits.
Full MLX and llama.cpp Metal support. MLX in particular was built with Apple Silicon's unified-memory bandwidth in mind and tends to extract a larger share of the chip's rated bandwidth than llama.cpp's more general Metal backend; worth trying both if decode speed matters more than llama.cpp's broader format support.
Models the 64 GB configuration runs natively (63)
- Qwen 3.5 122B-A10B (MoE)122B · MMLU-Pro 86.7Q2_K · ~49.6 t/s
- Nemotron 3 Super 120B120B · MMLU-Pro 83.7Q2_K · ~40.1 t/s
- Llama 4 Scout 109B109B · MMLU-Pro 74.3Q2_K · ~26.4 t/s
- GLM-4.5 Air 106B106B · MMLU-Pro 81.4Q2_K · ~38.1 t/s
- GLM-4.6V 106B106B · MMLU-Pro 79.9Q2_K · ~38.1 t/s
Show 58 more
- Qwen 2.5 72B Instruct72B · MMLU-Pro 71.1Q4_K_M · ~13.8 t/s
- Llama 3.3 70B Instruct70B · MMLU-Pro 68.9Q4_K_M · ~14.1 t/s
- DeepSeek R1 Distill Llama 70B70B · MMLU-Pro 70.0Q4_K_M · ~14.1 t/s
- Llama 3.1 70B Instruct70B · MMLU-Pro 66.4Q4_K_M · ~14.1 t/s
- Mixtral 8x7B Instruct v0.146.7B · MMLU-Pro 29.7Q6_K · ~17.6 t/s
- Command-R 35B35B · MMLU-Pro 33.0Q8_0 · ~13.3 t/s
- Qwen 3.5 35B-A3B (MoE)35B · MMLU-Pro 85.3Q8_0 · ~59.3 t/s
- Qwen 3.6 35B35B · MMLU-Pro 85.2Q8_0 · ~16.3 t/s
- Ornith 1.5 35B-A3B (MoE)35B · MMLU-Pro N/AQ8_0 · ~59.3 t/s
- Yi 1.5 34B Chat34.4B · MMLU-Pro 37.0Q8_0 · ~16.6 t/s
- Qwen3 32B32.8B · MMLU-Pro 65.5Q8_0 · ~17.7 t/s
- Qwen 2.5 32B Instruct32.5B · MMLU-Pro 69.0Q8_0 · ~17.4 t/s
- Qwen 2.5 Coder 32B Instruct32.5B · MMLU-Pro 62.3Q8_0 · ~17.4 t/s
- DeepSeek R1 Distill Qwen 32B32.5B · MMLU-Pro 65.0Q8_0 · ~17.4 t/s
- Nemotron 3 Nano 30B32B · MMLU-Pro 78.3Q8_0 · ~57.8 t/s
- Gemma 4 31B30.7B · MMLU-Pro 85.2Q8_0 · ~18.7 t/s
- Qwen3 30B-A3B (MoE)30B · MMLU-Pro 61.5Q8_0 · ~56 t/s
- Nemotron 3.5 Lightning 30B-A3B30B · MMLU-Pro 81.6Q8_0 · ~59.9 t/s
- Muse Glimmer 30B27.8B · MMLU-Pro N/AQ8_0 · ~21.5 t/s
- Gemma 2 27B Instruct27.2B · MMLU-Pro 38.0Q8_0 · ~20 t/s
- Gemma 3 27B Instruct27B · MMLU-Pro 67.5Q8_0 · ~21.2 t/s
- Qwen 3.6 27B27B · MMLU-Pro 86.2Q8_0 · ~21.9 t/s
- UI-Mate 27B27B · MMLU-Pro ~86.2Q8_0 · ~21.9 t/s
- Bonsai 27B27B · MMLU-Pro ~81.5Ternary (Q2_0) · ~82.7 t/s
- Bonsai 2 27B27B · MMLU-Pro N/ATernary (Q2_0) · ~99.4 t/s
- Qwen 3.8 27B27B · MMLU-Pro N/AQ8_0 · ~21.9 t/s
- Gemma 4 26B (MoE)25.2B · MMLU-Pro 82.6Q8_0 · ~46.2 t/s
- Mistral Small 3.1 24B Instruct24B · MMLU-Pro 66.8BF16 · ~13 t/s
- Mistral Small 22B22.2B · MMLU-Pro 49.2BF16 · ~13.8 t/s
- GPT-OSS 20B21B · MMLU-Pro 67.9BF16 · ~26.4 t/s
- Qwen3 14B14.8B · MMLU-Pro 61.0BF16 · ~20.7 t/s
- Qwen 2.5 14B Instruct14.7B · MMLU-Pro 63.7BF16 · ~20.6 t/s
- Phi-4 14B Instruct14B · MMLU-Pro 70.4BF16 · ~21.8 t/s
- Mistral Nemo 12B Instruct12.2B · MMLU-Pro 35.6BF16 · ~24.9 t/s
- Gemma 3 12B Instruct12.2B · MMLU-Pro 60.6BF16 · ~25.1 t/s
- Gemma 4 12B (Unified)12B · MMLU-Pro 77.2BF16 · ~23.5 t/s
- Gemma 2 9B Instruct9.2B · MMLU-Pro 32.0BF16 · ~30.2 t/s
- Qwen 3.5 9B9B · MMLU-Pro 82.5BF16 · ~35 t/s
- Ornith 1.5 9B9B · MMLU-Pro N/ABF16 · ~35 t/s
- Llama 3.1 8B Instruct8B · MMLU-Pro 48.3BF16 · ~37.5 t/s
- DeepSeek R1 Distill Llama 8B8B · MMLU-Pro 41.0BF16 · ~37.5 t/s
- Qwen3 8B8B · MMLU-Pro 56.7BF16 · ~37.2 t/s
- Qwen 2.5 7B Instruct7.6B · MMLU-Pro 56.3BF16 · ~40.8 t/s
- Mistral 7B Instruct v0.37.25B · MMLU-Pro 30.0BF16 · ~41.1 t/s
- Gemma 3 4B Instruct4B · MMLU-Pro 43.6BF16 · ~75.3 t/s
- Gemma 4 E4B4B · MMLU-Pro 69.4BF16 · ~71.1 t/s
- Phi-3.5 Mini Instruct3.8B · MMLU-Pro 47.4BF16 · ~59.1 t/s
- Phi-4-mini Instruct3.8B · MMLU-Pro 67.3BF16 · ~73.8 t/s
- Llama 3.2 3B Instruct3.2B · MMLU-Pro 24.0BF16 · ~87.2 t/s
- Qwen 2.5 3B Instruct3.1B · MMLU-Pro 32.4BF16 · ~98.4 t/s
- Gemma 2 2B Instruct2.6B · MMLU-Pro 17.8BF16 · ~105.4 t/s
- Gemma 4 E2B2B · MMLU-Pro 60.0BF16 · ~145.4 t/s
- SmolLM2 1.7B Instruct1.7B · MMLU-Pro 19.0BF16 · ~127.7 t/s
- Qwen 2.5 1.5B Instruct1.5B · MMLU-Pro 16.8BF16 · ~197.8 t/s
- Llama 3.2 1B Instruct1.24B · MMLU-Pro 12.5BF16 · ~232.9 t/s
- Gemma 3 1B Instruct1B · MMLU-Pro 14.7BF16 · ~275 t/s
- Qwen 2.5 0.5B Instruct0.5B · MMLU-Pro 10.0BF16 · ~581.5 t/s
- SmolLM2 360M Instruct0.36B · MMLU-Pro 8.0BF16 · ~606.3 t/s
Too large for any Apple M1 Ultra configuration (26)
- Llama 3.1 405B Instruct
- DeepSeek V3 671B
- DeepSeek R1 671B
- Llama 4 Maverick 400B
- MiniMax M1 456B
- GLM-4.5 355B
- GLM-4.6 355B
- GLM-4.7 358B
- GLM-5 744B
- Kimi K2.6
- GLM-5.1 754B
- DeepSeek V4 Pro 1.6T
- DeepSeek V4 Flash 284B
- GLM-5.2 753B
- Nemotron 3 Ultra 550B-A55B
- MiMo V2.5 Pro
- Kimi K2.5
- MiniMax M3
- Inkling
- Kimi K3
- Qwen3.8 2.4T-A95B
- DeepSeek V4 Pro 0813 1.6T
- Ornith 1.5 397B (MoE)
- GLM-5.3 753B
- GLM-5.3-Flash 320B
- DeepSeek V4.1 Flash 552B
Compare Apple M1 Ultra with other GPUs
Frequently asked questions
- How much memory does the Apple M1 Ultra have?
- The Apple M1 Ultra ships in 2 unified-memory configurations: 128 GB and 64 GB, all at 800 GB/s.
- Should I get the 64 GB or 128 GB Apple M1 Ultra?
- Both run everything that fits natively in 64 GB. The extra memory in the 128 GB configuration additionally fits DeepSeek V4 Flash 0731 284B, Qwen3 235B-A22B (MoE), MiniMax M2.5 229B, and 7 more models natively in VRAM, worth the upgrade if you plan to run any of those.
- How much VRAM does the Apple M1 Ultra (128GB) have?
- The Apple M1 Ultra (128GB) has 128 GB of LPDDR5 with 800 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M1 Ultra (128GB) best for?
- With 128 GB of unified memory, the Apple M1 Ultra (128GB) is a high-capacity workstation platform that runs 70B-class dense models and large MoE models natively, with plenty of room for long context.
- What LLMs can the Apple M1 Ultra (128GB) run locally?
- The Apple M1 Ultra (128GB) can run 73 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen3.8-Flash-Next at Q3_K_M, DeepSeek V4 Flash 0731 284B at UD-IQ3_XXS, Qwen 3.8 27B at BF16.
- Can the Apple M1 Ultra (128GB) run Gemma 4 31B?
- Yes. The Apple M1 Ultra (128GB) runs Gemma 4 31B natively in VRAM at BF16 quantization, achieving approximately 10.2 tokens per second.
Show 8 more questions
- Can the Apple M1 Ultra (128GB) run Qwen 3.6 27B?
- Yes. The Apple M1 Ultra (128GB) runs Qwen 3.6 27B natively in VRAM at BF16 quantization, achieving approximately 11.7 tokens per second.
- Can the Apple M1 Ultra (128GB) run Qwen3 8B?
- Yes. The Apple M1 Ultra (128GB) runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 37.2 tokens per second.
- How much VRAM does the Apple M1 Ultra (64GB) have?
- The Apple M1 Ultra (64GB) has 64 GB of LPDDR5 with 800 GB/s memory bandwidth (unified system memory, shared between CPU and GPU).
- What is the Apple M1 Ultra (64GB) best for?
- With 64 GB of VRAM, the Apple M1 Ultra (64GB) is ideal for running 70B-class models at Q4 quantization and large MoE models, a workstation sweet spot for local inference.
- What LLMs can the Apple M1 Ultra (64GB) run locally?
- The Apple M1 Ultra (64GB) can run 63 of the 99 open-weight models tracked by CanItRun natively in VRAM at 8k context. Top options include: Qwen 3.8 27B at Q8_0, Ornith 1.5 35B-A3B (MoE) at Q8_0, Qwen 3.5 122B-A10B (MoE) at Q2_K.
- Can the Apple M1 Ultra (64GB) run Gemma 4 31B?
- Yes. The Apple M1 Ultra (64GB) runs Gemma 4 31B natively in VRAM at Q8_0 quantization, achieving approximately 18.7 tokens per second.
- Can the Apple M1 Ultra (64GB) run Qwen 3.6 27B?
- Yes. The Apple M1 Ultra (64GB) runs Qwen 3.6 27B natively in VRAM at Q8_0 quantization, achieving approximately 21.9 tokens per second.
- Can the Apple M1 Ultra (64GB) run Qwen3 8B?
- Yes. The Apple M1 Ultra (64GB) runs Qwen3 8B natively in VRAM at BF16 quantization, achieving approximately 37.2 tokens per second.