Claude Mythos vs Claude Opus 4.7: The Capybara-Tier Jump

Claude MythosClaude Opus 4.7
arenaElo14781420
humanEval98.793.1
sweBench93.964.3
mmlu94.292.3
gpqa94.665.2
aime97.688
Price in/out ($/M)$25 / $125$5 / $25

Our verdict

Mythos pushes SWE-bench Verified from 53.4% (Opus 4.6) to 93.9% and USAMO from 42.3% to 97.6%, a 4.3× leap on the model-performance trendline. Opus 4.7 stays the practical default: it is generally available, costs 5× less ($5 / $25 vs $25 / $125 per M tokens) and ships through every major cloud. Pick Mythos only if you have a Project Glasswing seat or Anthropic green-lit your use case; otherwise stay on Opus 4.7.

Claude Mythos · Claude Opus 4.7 · All AI model comparisons