| Kimi K2 | GPT-5 | |
|---|---|---|
| arenaElo | 1384 | 1412 |
| humanEval | 90.1 | 91.6 |
| sweBench | 65.8 | 65 |
| mmlu | 89.1 | 91.5 |
| gpqa | 64.7 | 63.8 |
| aime | 70 | 92 |
| Price in/out ($/M) | $0.5 / $2 | $10 / $40 |
GPT-5 has the richer ecosystem, native audio and tighter tool-calling. Kimi K2 wins on price and context length. For voice agents, GPT-5. For long-doc pipelines and self-host, Kimi K2.