DeepSeek R1 vs Claude Opus 4.7, reasoning Showdown

DeepSeek R1Claude Opus 4.7
arenaElo13891420
humanEval90.293.1
sweBench49.264.3
mmlu90.892.3
gpqa71.565.2
aime79.888
Price in/out ($/M)$0.7 / $2.5$5 / $25

Our verdict

DeepSeek R1 actually beats Opus 4.7 on GPQA Diamond (71.5 vs 65.2) at a fraction of the price and with open weights. Opus 4.7 still wins on long-context coding, tool use and polish. DeepSeek is the value pick for pure reasoning.

DeepSeek R1 · Claude Opus 4.7 · All AI model comparisons