| Claude Mythos | OpenAI o4 | |
|---|---|---|
| arenaElo | 1478 | 1438 |
| humanEval | 98.7 | 94 |
| sweBench | 93.9 | 72 |
| mmlu | 94.2 | 92.5 |
| gpqa | 94.6 | 87.1 |
| aime | 97.6 | 94 |
| Price in/out ($/M) | $25 / $125 | $12 / $48 |
o4 was the GPQA king at 87.1%, until Mythos posted 94.6% on the same benchmark and 97.6% on USAMO 2026 (vs ~90% for o4). The catch: o4 is a public reasoning model you can call today, while Mythos is invitation-only. If you need a reasoning monster you can actually deploy this quarter, o4 is the answer. If you are inside Project Glasswing, Mythos is simply the most capable reasoner ever shipped.