| Kimi K2 | DeepSeek R1 | |
|---|---|---|
| arenaElo | 1384 | 1389 |
| humanEval | 90.1 | 90.2 |
| sweBench | 65.8 | 49.2 |
| mmlu | 89.1 | 90.8 |
| gpqa | 64.7 | 71.5 |
| aime | 70 | 79.8 |
| Price in/out ($/M) | $0.5 / $2 | $0.55 / $2.19 |
R1 wins on reasoning (GPQA 71.5 vs 64.7); Kimi K2 wins on context (2M vs 128K) and general chat quality. For math and science, R1. For long-form writing, agents and APAC products, Kimi.