| DeepSeek V4 | DeepSeek R1 | |
|---|---|---|
| arenaElo | 1395 | 1389 |
| humanEval | 91 | 90.2 |
| sweBench | 62 | 49.2 |
| mmlu | 91.5 | 90.8 |
| gpqa | 73 | 71.5 |
| aime | 90 | 79.8 |
| Price in/out ($/M) | $1.74 / $3.48 | $0.7 / $2.5 |
V4 supersedes R1 on every axis: 1M context (vs 128K), native vision, +1.5 GPQA, +0.8 HumanEval, and a hybrid sparse attention that uses ~10% of R1's KV cache at 1M tokens. R1 still wins on raw input price ($0.55 vs $1.74), so keep R1 for batch reasoning where context fits in 128K and budget is tight. For everything else, upgrade.