CanItRun Logocanitrun.

Phi-4-mini Instruct

Phi-4-mini Instruct needs roughly 3.8 GB VRAM at Q4_K_M quantization (9.7 GB at FP16). 102 GPUs we track can run it fully in VRAM at 8k context.

102 GPUs run this natively · 1 with CPU offload

Microsoft3.8B params128k contextMITCommercial use ok

Phi-4-mini Instruct is a 3.8B parameter dense large language model developed by Microsoft. Released in February 2025, it is a text-only model with a 128K context window, released under the MIT license, allowing commercial use. Same size class as Phi-3.5 Mini but adds GQA (8 KV heads, versus Phi-3.5's full 32) for a much smaller KV cache, plus function calling and a 200K vocabulary.

To run Phi-4-mini Instruct locally, you need approximately 3.8 GB of VRAM at Q4_K_M quantization with 8k context. 102 of the GPUs we track can run it fully in VRAM, with a further 1 able to offload to system RAM. At Q4_K_M it needs just 3.8 GB, making it accessible even on mid-range 16 GB cards like RTX 4080 and RTX 5070 Ti. At Q8_K_M (5.7 GB), you get near-FP16 quality while still fitting on 8, 12, 16, 24, 32, 48 and 80 GB GPUs. FP16 requires 9.7 GB, limiting it to 48 and 80 GB GPUs.

With an MMLU-Pro score of 67.3, it delivers strong general reasoning for local deployment. The license allows commercial use.

VRAM at each quantization

Figures below assume 8k context; KV cache grows linearly as context length increases.

QuantWeightsKV cacheTotal
FP3215.2 GB1.07 GB18.2 GB
BF167.6 GB1.07 GB9.7 GB
FP167.6 GB1.07 GB9.7 GB
Q8_04.0 GB1.07 GB5.7 GB
Q6_Krec3.1 GB1.07 GB4.7 GB
Q5_K_M2.7 GB1.07 GB4.2 GB
Q4_K_M2.3 GB1.07 GB3.8 GB
Q3_K_M1.8 GB1.07 GB3.3 GB
Q2_K1.4 GB1.07 GB2.8 GB
NVFP4cuda1.9 GB1.07 GB3.3 GB

Shown at 8k context with FP16 KV cache. NVFP4 needs a CUDA GPU to run. Toggle TurboQuant in the calculator to view compressed KV cache numbers.

Benchmarks

GPUs that run Phi-4-mini Instruct natively (102)

Show 97 more
Plus 1 GPUs that run it with CPU offload (slower)

Notes

Same size class as Phi-3.5 Mini but adds GQA (8 KV heads, versus Phi-3.5's full 32) for a much smaller KV cache, plus function calling and a 200K vocabulary.

Hugging Face ↗Ollama ↗Released 2025-02-27

Frequently asked questions

What are the VRAM requirements for Phi-4-mini Instruct?
Phi-4-mini Instruct requires approximately 3.8 GB of VRAM at Q4_K_M quantization, 5.7 GB at Q8, and 9.7 GB at FP16. These numbers assume 8k context window; VRAM scales linearly with context length due to the KV cache.
How many parameters does Phi-4-mini Instruct have?
Phi-4-mini Instruct has 3.8 billion parameters.
How capable is Phi-4-mini Instruct?
With an MMLU-Pro score of 67.3, Phi-4-mini Instruct delivers solid general-purpose performance suitable for most everyday tasks and professional use.
Can Phi-4-mini Instruct run on a 16 GB GPU?
Yes. Phi-4-mini Instruct needs 3.8 GB at Q4_K_M, which fits in a 16 GB GPU like the RTX 4080 or RTX 5070 Ti.
What is the smallest quantization for Phi-4-mini Instruct that fits in 24 GB of VRAM?
At FP32, Phi-4-mini Instruct needs 18.2 GB — the highest-quality quantization that fits in 24 GB of VRAM.
What GPU do I need to run Phi-4-mini Instruct locally?
A 16 GB GPU is enough. At Q4_K_M, Phi-4-mini Instruct needs 3.8 GB VRAM. Good options: RTX 4080 (16 GB), RTX 5070 Ti (16 GB).