Claude Mythos vs GPT-5.5: Frontier Lab Showdown

Claude MythosGPT-5.5
arenaElo14781432
humanEval98.794.2
sweBench93.966
mmlu94.293
gpqa94.668.7
aime97.692
Price in/out ($/M)$25 / $125$12 / $48

Our verdict

Mythos beats GPT-5.5 on every shared benchmark: SWE-bench Verified (93.9% vs 65.1%), GPQA Diamond (94.6% vs 68.7%) and HumanEval (98.7% vs 94.2%). GPT-5.5 still wins on availability: it is generally available with native audio, while Mythos is locked behind Project Glasswing. For consumer-facing or multimodal voice products, GPT-5.5 is the only realistic option. For defensive security and mathematical research where you can secure access, Mythos is in a different tier.

Claude Mythos · GPT-5.5 · All AI model comparisons