What capability costs today, what a task really costs once thinking is counted, and how fast each level gets cheaper.
Price is blended 3:1 input to output tokens. The line traces the best score you can buy at each price: anything under it is beaten by something cheaper. Dashed rings are estimated scores.
DeepSeek-V4.1-Flash lists 5.8× cheaper per token than Inkling-Small, yet costs 3.0× more per task on Artificial Analysis's test set. The bill depends on how many tokens a model uses, not just its rate.
Average cost of one task in Artificial Analysis's evaluation suite, at the setting shown on hover, counting every token the model uses (input, reasoning and answer). Token price is OpenRouter's list price, blended 3:1, for reference.
Each line steps down when a cheaper model reaching that score is released, priced at what it costs today. It shows how quickly cheaper models arrive, not what anyone paid at the time (GPT-4 itself listed at $37.50 per million tokens in March 2023).