| OpenAI o1 | OpenAI o3 | |
|---|---|---|
| arenaElo | 1380 | 1418 |
| humanEval | 89.1 | 92.4 |
| sweBench | 48.9 | 69.1 |
| mmlu | 88.7 | 90.2 |
| gpqa | 78 | 83.3 |
| aime | 83.3 | 91.6 |
| Price in/out ($/M) | $15 / $60 | $2 / $8 |
There is no trade-off to weigh here, which is rare. o3 beats o1 on every single published metric, including SWE-bench Verified 69.1% against 48.9%, GPQA 83.3% against 78%, Aider Polyglot 79 against 61, Terminal-Bench 48 against 33 and Humanity's Last Exam 20.3% against 8.8%, and it does it at $2 / $8 per M tokens against $15 / $60, roughly a seventh of the price. The context window is the same 200K and both take text and vision, so a migration changes nothing about how you call it. If you are still paying for o1 in 2026 you are paying seven times more for a measurably worse model. Move.