| Kimi K3 | Kimi K2 | |
|---|---|---|
| arenaElo | 1478 | 1384 |
| humanEval | 96.5 | 90.1 |
| sweBench | 75.5 | 65.8 |
| mmlu | 94 | 89.1 |
| gpqa | 87.5 | 64.7 |
| aime | 92 | 70 |
| Price in/out ($/M) | $3 / $15 | $0.57 / $2.3 |
K3 is a generational leap: 2.8T parameters vs 1T, roughly 2.5x scaling efficiency, always-on reasoning and frontier-class coding (SWE-Bench Pro 70.0 vs 56.0). K2 keeps two advantages. A 2M context window (vs 1M) and much cheaper tokens ($0.5/$2 vs $3/$15). For agents, coding and hard reasoning, K3. For bulk long-document pipelines on a budget, K2 still earns its keep.