| Claude Mythos | GPT-5.5 | |
|---|---|---|
| arenaElo | 1478 | 1432 |
| humanEval | 98.7 | 94.2 |
| sweBench | 93.9 | 66 |
| mmlu | 94.2 | 93 |
| gpqa | 94.6 | 68.7 |
| aime | 97.6 | 92 |
| Price in/out ($/M) | $25 / $125 | $12 / $48 |
Mythos beats GPT-5.5 on every shared benchmark: SWE-bench Verified (93.9% vs 65.1%), GPQA Diamond (94.6% vs 68.7%) and HumanEval (98.7% vs 94.2%). GPT-5.5 still wins on availability: it is generally available with native audio, while Mythos is locked behind Project Glasswing. For consumer-facing or multimodal voice products, GPT-5.5 is the only realistic option. For defensive security and mathematical research where you can secure access, Mythos is in a different tier.