modelbenchmark.io

Phi-4

Microsoft · phi-4

Compare

Open-weight instruction model for adaptable chat and self-hosted production workloads

Specification

most-agreed values

Context
128K
Max output
4K
Released
2024-12-11
Knowledge cutoff
2023-10
Retires
Open weights
yes
Input
text
Output
text

Price

US dollars per million tokens · most-agreed

Input
$0.125
Output
$0.50

Quality

4 benchmarks · 1 source

Composite
35th
percentile of 334 scored models
Rank
218 / 334
±0.07 sd
Evidence
4 × 1
one source, or fewer than 3 benchmarks
Effort range
one configuration only
reasoning32nd2/4 bench
math39th2/7 bench

Every score

one row per source and configuration — nothing averaged away

BenchmarkScoreConfigurationSourceRun
MATH level 564.9 ±1.1defaultEpoch AI2025-01-31
GPQA diamond56.1 ±2.6defaultEpoch AI2025-01-31
Chess Puzzles1.0 ±1.0defaultEpoch AI2026-08-28
OTIS Mock AIME 2024-202513.8 ±3.7defaultEpoch AI2025-02-25

The leading figure for a benchmark is the median across its configurations, so one heroic high-effort run cannot set the number.

Available from 6 hosts

HostIn $/MOut $/MCache rdCache wrContextOutputRetires
deepinfra$0.07$0.1416K16K
kilo$0.07$0.1416K15K
openrouter$0.07$0.1416K15K
azure$0.125$0.50128K4K
azure_ai$0.125$0.5016K16K
azure-cognitive-services$0.125$0.50128K4K

More from Microsoft

most-hosted first