Claude Mythos vs DeepSeek V4: Closed Frontier vs Open-Weights Frontier

Claude MythosDeepSeek V4
arenaElo14781395
humanEval98.791
sweBench93.962
mmlu94.291.5
gpqa94.673
aime97.690
Price in/out ($/M)$25 / $125$1.74 / $3.48

Our verdict

DeepSeek V4 is the open-weights flagship of 2026: MIT-licensed, 1M context, ~14× cheaper input than Mythos, and self-hostable. Mythos retains the capability lead by 20+ points on GPQA and SWE-bench, plus its cybersecurity capabilities are unmatched. Pick V4 for any team that needs to deploy frontier-tier reasoning on its own hardware or under data residency requirements. Pick Mythos only if you have access and your use case justifies a 14× per-token cost.

Claude Mythos · DeepSeek V4 · All AI model comparisons