| DeepSeek V4 | Claude Opus 4.6 | |
|---|---|---|
| arenaElo | 1395 | 1402 |
| humanEval | 91 | 91.4 |
| sweBench | 62 | 53.4 |
| mmlu | 91.5 | 90.8 |
| gpqa | 73 | 62.1 |
| aime | 90 | 84 |
| Price in/out ($/M) | $1.74 / $3.48 | $15 / $75 |
DeepSeek V4 (April 2026) effectively retires the case for Opus 4.6 as a budget-frontier choice: V4 ties or beats Opus 4.6 on MMLU and GPQA, doubles the context (1M vs 500K) and ships at ~12% of Opus 4.6's API cost ($1.74/$3.48 vs $15/$75 per M tokens). Opus 4.6 still wins on Chatbot Arena, agentic tool reliability and Anthropic's safety tuning. If you can self-host or accept Chinese-vendor cloud, V4 is the rational replacement; for regulated enterprise pipelines already on Anthropic, stick with Opus 4.6 (or upgrade straight to 4.7).