CanItRun Logocanitrun.

Mistral Medium 3.5 128B

Mistral Medium 3.5 128B needs roughly 90.6 GB VRAM at Q4_K_M quantization (290.0 GB at FP16). 22 GPUs we track can run it fully in VRAM at 8k context.

22 GPUs run this natively · 6 with CPU offload

Mistral AI128B params256k contextApache 2.0Commercial use ok

Mistral Medium 3.5 128B is a 128B parameter dense large language model developed by Mistral AI. Released in April 2026, it supports text and vision inputs with a 256K context window, released under the Apache 2.0 license, allowing commercial use. Dense successor that folds Mistral's dedicated reasoning model (Magistral) and dedicated coding model (Devstral 2) into a single configurable-effort model, replacing both as Mistral's default coding/agentic pick. Standard GQA attention (8 KV heads), not the MLA-style compressed attention used by Mistral's larger MoE releases.

To run Mistral Medium 3.5 128B locally, you need approximately 90.6 GB of VRAM at Q4_K_M quantization with 8k context. 22 of the GPUs we track can run it fully in VRAM, with a further 6 able to offload to system RAM. At Q4_K_M it requires 90.6 GB — more than any single consumer GPU. A multi-GPU server or 80 GB datacenter GPU is required. At Q8_K_M (155.7 GB), you get near-FP16 quality while still fitting on large server or multi-GPU setups. FP16 requires 290.0 GB, limiting it to datacenter-class hardware with 80 GB+ VRAM.

Mistral Medium 3.5 128B is designed for general-purpose instruction and chat. Its 128B parameters provide good performance on standard benchmarks when run locally.

VRAM at each quantization

Calculated at 8k context. Since KV cache scales linearly with context, longer sessions need more VRAM than shown here.

QuantWeightsKV cacheTotal
FP32512.0 GB2.95 GB576.8 GB
BF16256.0 GB2.95 GB290.0 GB
FP16256.0 GB2.95 GB290.0 GB
Q8_0136.1 GB2.95 GB155.7 GB
Q6_K105.1 GB2.95 GB121.0 GB
Q5_K_M91.1 GB2.95 GB105.4 GB
Q4_K_Mrec78.0 GB2.95 GB90.6 GB
Q3_K_M61.6 GB2.95 GB72.3 GB
Q2_K48.8 GB2.95 GB57.9 GB
NVFP4cuda64.0 GB2.95 GB75.0 GB

KV cache figures assume 8k context at FP16. NVFP4 quantization requires a CUDA-capable GPU. Enable TurboQuant in the calculator to see reduced KV cache estimates.

Benchmarks

GPUs that run Mistral Medium 3.5 128B natively (22)

Show 17 more
Plus 6 GPUs that run it with CPU offload (slower)

Notes

Dense successor that folds Mistral's dedicated reasoning model (Magistral) and dedicated coding model (Devstral 2) into a single configurable-effort model, replacing both as Mistral's default coding/agentic pick. Standard GQA attention (8 KV heads), not the MLA-style compressed attention used by Mistral's larger MoE releases.

Hugging Face ↗Released 2026-04-29

Frequently asked questions

What are the VRAM requirements for Mistral Medium 3.5 128B?
Mistral Medium 3.5 128B requires approximately 90.6 GB of VRAM at Q4_K_M quantization, 155.7 GB at Q8, and 290.0 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
How many parameters does Mistral Medium 3.5 128B have?
Mistral Medium 3.5 128B has 128 billion parameters.
Can Mistral Medium 3.5 128B run on a 16 GB GPU?
No. At Q4_K_M, Mistral Medium 3.5 128B needs 90.6 GB of VRAM — more than 16 GB. You will need a multi-GPU server.
Can Mistral Medium 3.5 128B run on a 24 GB GPU?
No. Even at Q4_K_M, Mistral Medium 3.5 128B needs 90.6 GB. Consider a multi-GPU server with 80 GB+ total VRAM.
What is the smallest quantization for Mistral Medium 3.5 128B that fits in 24 GB of VRAM?
Mistral Medium 3.5 128B cannot fit in 24 GB of VRAM at any standard quantization level. The minimum needed is 57.9 GB at Q2_K.
What GPU do I need to run Mistral Medium 3.5 128B locally?
You need a multi-GPU server. At Q4_K_M, Mistral Medium 3.5 128B needs 90.6 GB VRAM, more than any single consumer GPU. Consider 2–4× H100 or A100 GPUs.