| DeepSeek R1 | Claude Opus 4.7 | |
|---|---|---|
| arenaElo | 1389 | 1420 |
| humanEval | 90.2 | 93.1 |
| sweBench | 49.2 | 64.3 |
| mmlu | 90.8 | 92.3 |
| gpqa | 71.5 | 65.2 |
| aime | 79.8 | 88 |
| Price in/out ($/M) | $0.7 / $2.5 | $5 / $25 |
DeepSeek R1 actually beats Opus 4.7 on GPQA Diamond (71.5 vs 65.2) at a fraction of the price and with open weights. Opus 4.7 still wins on long-context coding, tool use and polish. DeepSeek is the value pick for pure reasoning.