| Claude Mythos | Kimi K2 | |
|---|---|---|
| arenaElo | 1478 | 1384 |
| humanEval | 98.7 | 90.1 |
| sweBench | 93.9 | 65.8 |
| mmlu | 94.2 | 89.1 |
| gpqa | 94.6 | 64.7 |
| aime | 97.6 | 70 |
| Price in/out ($/M) | $25 / $125 | $0.5 / $2 |
Kimi K2 from Moonshot AI is the most capable open-weights model with a 2M context window and strong long-form writing. Mythos beats K2 by ~30 points on GPQA and ~9 on HumanEval, plus has cybersecurity capabilities K2 cannot match. K2 wins on price (~50× cheaper input), open weights and APAC-language coverage. Use K2 when you need self-hosted long context with frontier-grade reasoning; reach for Mythos when capability is the only axis that matters.