| DeepSeek V4 | GPT-5 | |
|---|---|---|
| arenaElo | 1395 | 1412 |
| humanEval | 91 | 91.6 |
| sweBench | 62 | 65 |
| mmlu | 91.5 | 91.5 |
| gpqa | 73 | 63.8 |
| aime | 90 | 92 |
| Price in/out ($/M) | $1.74 / $3.48 | $1.25 / $10 |
On reasoning benchmarks DeepSeek V4 matches or beats GPT-5 (GPQA 73.0 vs 63.8) at a small fraction of the cost. GPT-5 retains the multimodal edge (audio) and tool-use ecosystem. The V4 release is the strongest case yet that frontier-tier intelligence is commoditizing: for any team that controls its own inference stack, V4 is the new default.