> Qwen2.5 is an open-weights AI model from Alibaba, released 19 Sep 2024. Index score 110 (GPT-4 = 100), rank 159 of 229 ranked models. API price $0.36 in / $0.40 out per million tokens, 33K context.

[Models](https://themodelindex.org/models/) /[Alibaba](https://themodelindex.org/labs/alibaba/)

# Qwen2.5

Alibaba

Released 19 Sep 2024

Open

Date verified

[Compare](https://themodelindex.org/compare/qwen2-5-vs-o1-mini/) [Announcement](https://qwenlm.github.io/blog/qwen2.5/)

11099–122

Index score

#159 of 229 models

2/ 8

Core tests taken

2 of 6 areas · plus 10 anchor tests

−13

Versus the frontier at release

best then: o1-mini

$0.36/ $0.40

Price per million tokens

input / output, list price

33K

Context window

tokens it reads at once

## Where it sits · on the frontier

_Chart: Model Index score of every scored model since late 2022, with the best closed and best open model over time_

Claude 3.5 Sonnet (Jun 2024) reached this score 3 months earlier.

## Test results · by area · tick = best by any model

Show 10 anchor tests

Reasoning & knowledge

evidence

GPQA Diamond

49%

Humanity's Last Exam

no result

Maths

evidence

FrontierMath T1-3

no result

Coding

evidence

SWE-bench Verified

no result

Terminal-Bench

no result

Agents

evidence

APEX-Agents

no result

Novel problems

no evidence yet

ARC-AGI-2

no result

Long tasks

evidence

METR time horizon

5.2 min

Independent results from the Epoch AI benchmarking hub, best reasoning setting. The core tests are shown here; anchor tests also inform the score and can be shown above.

## What others say · rank on each leaderboard

The Model Index

independent tests, our method

#159

of 219 · 110

LMArena · Creative writing

creative writing

#152

of 386 · 1353

LMArena · Text

overall chat quality, judged by people

#171

of 388 · 1374

LMArena · Maths

maths questions in chat

#178

of 373 · 1363

LMArena · Coding

coding questions in chat

#184

of 383 · 1402

Epoch Capabilities Index

overall capability from many benchmarks

#128

of 239 · 132.5

Further left is better. Leaderboards measure different things: LMArena is people voting blind between two answers, the others are test scores. Click any row for the source.

## Best at · top three among current models

Not in the top three among current models on any category we track.

Where this model places in the top three of today's models, on category leaderboards and on our core tests.

## Where to use it · prices checked 1 Oct 2026

#### API providers · 2 · per million tokens, in / out

D

DeepInfra

$0.36 / $0.40 · fp8

cheapest

N

Novita

$0.38 / $0.40 · bf16

↗

#### Run it yourself

🤗

Official weights

Qwen/Qwen2.5-72B-Instruct

↗

Provider prices come from OpenRouter's live listing; going direct to a provider can differ. Subscriptions are the lab's own apps.

## Facts and sources

Released

19 Sep 2024

Announcement

[qwenlm.github.io ↗](https://qwenlm.github.io/blog/qwen2.5/)

Release entry

Qwen2.5

Scored as

Qwen2.5-72B

Artificial Analysis index

–

Output speed

–

Hugging Face

[Qwen/Qwen2.5-72B-Instruct ↗](https://huggingface.co/Qwen/Qwen2.5-72B-Instruct)

## Alibaba releases · around this one

Qwen1.5

Open

4 FEB 2024

Qwen2

Open

7 JUN 2024

index 103

Qwen2.5

Open

19 SEP 2024

index 110

Qwen2.5-Coder-32B

Open

12 NOV 2024

QwQ-32B-Preview

Open

28 NOV 2024

---
Source: https://themodelindex.org/models/qwen2-5/ · The Model Index · data as of 2026-10-01
