> Phi-4 is an open-weights AI model from Microsoft, released 12 Dec 2024. Index score 112 (GPT-4 = 100), rank 153 of 229 ranked models. API price $0.070 in / $0.14 out per million tokens, 16K context.

[Models](https://themodelindex.org/models/) /[Microsoft](https://themodelindex.org/labs/microsoft/)

# Phi-4

Microsoft

Released 12 Dec 2024

Open

[Compare](https://themodelindex.org/compare/phi-4-vs-o1/)

112100–124

Index score

#153 of 229 models

1/ 8

Core tests taken

1 of 6 areas · plus 5 anchor tests

−24

Versus the frontier at release

best then: o1

$0.070/ $0.14

Price per million tokens

input / output, list price

16K

Context window

tokens it reads at once

## Where it sits · on the frontier

_Chart: Model Index score of every scored model since late 2022, with the best closed and best open model over time_

o1-mini reached this score 3 months earlier.

## Test results · by area · tick = best by any model

Show 4 anchor tests

Reasoning & knowledge

evidence

GPQA Diamond

56%

Humanity's Last Exam

no result

Maths

evidence

FrontierMath T1-3

no result

Coding

no evidence yet

SWE-bench Verified

no result

Terminal-Bench

no result

Agents

evidence

APEX-Agents

no result

Novel problems

no evidence yet

ARC-AGI-2

no result

Long tasks

no evidence yet

METR time horizon

no result

Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.

## What others say · rank on each leaderboard

The Model Index

independent tests, our method

#153

of 219 · 112

Epoch Capabilities Index

overall capability from many benchmarks

#134

of 239 · 130.4

LMArena · Maths

maths questions in chat

#259

of 373 · 1265

LMArena · Coding

coding questions in chat

#280

of 383 · 1306

LMArena · Text

overall chat quality, judged by people

#293

of 388 · 1256

LMArena · Creative writing

creative writing

#292

of 386 · 1210

Artificial Analysis Intelligence Index

overall capability across ten evaluations

#177

of 207 · 6

Further left is better. Leaderboards measure different things: LMArena is people voting blind between two answers, the others are test scores. Click any row for the source.

## Best at · top three among current models

Not in the top three among current models on any category we track.

Where this model places in the top three of today's models, on category leaderboards and on our core tests.

## Where to use it · prices checked 1 Oct 2026

#### API providers · 1 · per million tokens, in / out

D

DeepInfra

$0.070 / $0.14 · bf16

cheapest

#### Run it yourself

🤗

Official weights

microsoft/phi-4

↗

Provider prices come from OpenRouter's live listing; going direct to a provider can differ. Subscriptions are the lab's own apps.

## Facts and sources

Released

12 Dec 2024

Announcement

not linked yet

Release entry

Phi-4

Scored as

Phi-4

Artificial Analysis index

6 (their estimate)

Output speed

40 tokens/s

Hugging Face

[microsoft/phi-4 ↗](https://huggingface.co/microsoft/phi-4)

## Microsoft releases · around this one

Phi-2

Open

12 DEC 2023

Phi-3

Open

23 APR 2024

index 94

Phi-4

Open

12 DEC 2024

index 112

Phi-4-reasoning

Open

30 APR 2025

MAI-1-preview

Closed

28 AUG 2025

---
Source: https://themodelindex.org/models/phi-4/ · The Model Index · data as of 2026-10-01
