OpenAI o1 vs o3, the Upgrade That Costs a Seventh and Wins Everywhere

OpenAI o1OpenAI o3
arenaElo13801418
humanEval89.192.4
sweBench48.969.1
mmlu88.790.2
gpqa7883.3
aime83.391.6
Price in/out ($/M)$15 / $60$2 / $8

Our verdict

There is no trade-off to weigh here, which is rare. o3 beats o1 on every single published metric, including SWE-bench Verified 69.1% against 48.9%, GPQA 83.3% against 78%, Aider Polyglot 79 against 61, Terminal-Bench 48 against 33 and Humanity's Last Exam 20.3% against 8.8%, and it does it at $2 / $8 per M tokens against $15 / $60, roughly a seventh of the price. The context window is the same 200K and both take text and vision, so a migration changes nothing about how you call it. If you are still paying for o1 in 2026 you are paying seven times more for a measurably worse model. Move.

OpenAI o1 · OpenAI o3 · All AI model comparisons