Claude Mythos vs OpenAI o4: Reasoning Champions Compared

Claude MythosOpenAI o4
arenaElo14781438
humanEval98.794
sweBench93.972
mmlu94.292.5
gpqa94.687.1
aime97.694
Price in/out ($/M)$25 / $125$12 / $48

Our verdict

o4 was the GPQA king at 87.1%, until Mythos posted 94.6% on the same benchmark and 97.6% on USAMO 2026 (vs ~90% for o4). The catch: o4 is a public reasoning model you can call today, while Mythos is invitation-only. If you need a reasoning monster you can actually deploy this quarter, o4 is the answer. If you are inside Project Glasswing, Mythos is simply the most capable reasoner ever shipped.

Claude Mythos · OpenAI o4 · All AI model comparisons