Mistral Medium 3.5 128B
Mistral Medium 3.5 128B needs roughly 90.6 GB VRAM at Q4_K_M quantization (290.0 GB at FP16). 22 GPUs we track can run it fully in VRAM at 8k context.
22 GPUs run this natively · 6 with CPU offload
Mistral Medium 3.5 128B is a 128B parameter dense large language model developed by Mistral AI. Released in April 2026, it supports text and vision inputs with a 256K context window, released under the Apache 2.0 license, allowing commercial use. Dense successor that folds Mistral's dedicated reasoning model (Magistral) and dedicated coding model (Devstral 2) into a single configurable-effort model, replacing both as Mistral's default coding/agentic pick. Standard GQA attention (8 KV heads), not the MLA-style compressed attention used by Mistral's larger MoE releases.
To run Mistral Medium 3.5 128B locally, you need approximately 90.6 GB of VRAM at Q4_K_M quantization with 8k context. 22 of the GPUs we track can run it fully in VRAM, with a further 6 able to offload to system RAM. At Q4_K_M it requires 90.6 GB — more than any single consumer GPU. A multi-GPU server or 80 GB datacenter GPU is required. At Q8_K_M (155.7 GB), you get near-FP16 quality while still fitting on large server or multi-GPU setups. FP16 requires 290.0 GB, limiting it to datacenter-class hardware with 80 GB+ VRAM.
Mistral Medium 3.5 128B is designed for general-purpose instruction and chat. Its 128B parameters provide good performance on standard benchmarks when run locally.
VRAM at each quantization
Calculated at 8k context. Since KV cache scales linearly with context, longer sessions need more VRAM than shown here.
| Quant | Weights | KV cache | Total |
|---|---|---|---|
| FP32 | 512.0 GB | 2.95 GB | 576.8 GB |
| BF16 | 256.0 GB | 2.95 GB | 290.0 GB |
| FP16 | 256.0 GB | 2.95 GB | 290.0 GB |
| Q8_0 | 136.1 GB | 2.95 GB | 155.7 GB |
| Q6_K | 105.1 GB | 2.95 GB | 121.0 GB |
| Q5_K_M | 91.1 GB | 2.95 GB | 105.4 GB |
| Q4_K_Mrec | 78.0 GB | 2.95 GB | 90.6 GB |
| Q3_K_M | 61.6 GB | 2.95 GB | 72.3 GB |
| Q2_K | 48.8 GB | 2.95 GB | 57.9 GB |
| NVFP4cuda | 64.0 GB | 2.95 GB | 75.0 GB |
KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.
Benchmarks
GPUs that run Mistral Medium 3.5 128B natively (22)
- NVIDIA H100 80GBNVFP4 · 32.5 t/s
- NVIDIA A100 80GBNVFP4 · 19.8 t/s
- NVIDIA RTX Pro 6000NVFP4 · 13 t/s
- NVIDIA DGX Spark (128GB)NVFP4 · 2.7 t/s
- AMD Instinct MI300XQ8_0 · 24.8 t/s
Show 17 more
- AMD Strix Halo (128GB)Q5_K_M · 1.8 t/s
- AMD Strix Halo (96GB)Q3_K_M · 2.6 t/s
- Apple M5 Max (128GB)Q5_K_M · 5.2 t/s
- Apple M4 Ultra (384GB)BF16 · 3.4 t/s
- Apple M4 Ultra (192GB)Q8_0 · 6.3 t/s
- Apple M4 Max (128GB)Q5_K_M · 4.6 t/s
- Apple M4 Max (96GB)Q3_K_M · 6.8 t/s
- Apple M3 Ultra (512GB)BF16 · 2.5 t/s
- Apple M3 Ultra (256GB)Q8_0 · 4.7 t/s
- Apple M3 Ultra (96GB)Q3_K_M · 10.2 t/s
- Apple M3 Max (128GB)Q5_K_M · 3.4 t/s
- Apple M3 Max (96GB)Q3_K_M · 3.7 t/s
- Apple M2 Ultra (384GB)BF16 · 2.5 t/s
- Apple M2 Ultra (192GB)Q8_0 · 4.6 t/s
- Apple M2 Max (96GB)Q3_K_M · 5 t/s
- Apple M1 Ultra (128GB)Q5_K_M · 6.8 t/s
- Intel Data Center GPU Max 1550Q6_K · 19.7 t/s
Plus 6 GPUs that run it with CPU offload (slower)
- NVIDIA A100 40GBQ2_K · 1.7 t/s
- NVIDIA L40SQ2_K · 3.1 t/s
- NVIDIA RTX A6000Q2_K · 3 t/s
- NVIDIA RTX 6000 AdaQ2_K · 3.1 t/s
- AMD Radeon PRO W7900Q2_K · 3.1 t/s
- Intel Data Center GPU Max 1100Q2_K · 3.3 t/s
Notes
Dense successor that folds Mistral's dedicated reasoning model (Magistral) and dedicated coding model (Devstral 2) into a single configurable-effort model, replacing both as Mistral's default coding/agentic pick. Standard GQA attention (8 KV heads), not the MLA-style compressed attention used by Mistral's larger MoE releases.
Frequently asked questions
- What are the VRAM requirements for Mistral Medium 3.5 128B?
- Mistral Medium 3.5 128B requires approximately 90.6 GB of VRAM at Q4_K_M quantization, 155.7 GB at Q8, and 290.0 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
- How many parameters does Mistral Medium 3.5 128B have?
- Mistral Medium 3.5 128B has 128 billion parameters.
- Can Mistral Medium 3.5 128B run on a 16 GB GPU?
- No. At Q4_K_M, Mistral Medium 3.5 128B needs 90.6 GB of VRAM — more than 16 GB. You will need a multi-GPU server.
- Can Mistral Medium 3.5 128B run on a 24 GB GPU?
- No. Even at Q4_K_M, Mistral Medium 3.5 128B needs 90.6 GB. Consider a multi-GPU server with 80 GB+ total VRAM.
- What is the smallest quantization for Mistral Medium 3.5 128B that fits in 24 GB of VRAM?
- Mistral Medium 3.5 128B cannot fit in 24 GB of VRAM at any standard quantization level. The minimum needed is 57.9 GB at Q2_K.
- What GPU do I need to run Mistral Medium 3.5 128B locally?
- You need a multi-GPU server. At Q4_K_M, Mistral Medium 3.5 128B needs 90.6 GB VRAM, more than any single consumer GPU. Consider 2–4× H100 or A100 GPUs.