Claude Mythos vs Grok 4 Heavy: Frontier Reasoning at Scale

Claude MythosGrok 4 Heavy
arenaElo14781391
humanEval98.788.4
sweBench93.955
mmlu94.289.3
gpqa94.674.5
aime97.692
Price in/out ($/M)$25 / $125$15 / $60

Our verdict

Grok 4 Heavy uses multi-agent debate to hit 74.5% GPQA at $15 / $60 per M tokens. Mythos posts 94.6% GPQA in a single forward pass at $25 / $125. On Humanity's Last Exam with tools Mythos leads by ~6 points. For most teams Grok 4 Heavy is the pragmatic ceiling. Mythos is the answer when you need maximum capability and have access.

Claude Mythos · Grok 4 Heavy · All AI model comparisons