| Gemini 3 Deep Think | Claude Opus 4.7 | |
|---|---|---|
| arenaElo | 1429 | 1420 |
| humanEval | 92.4 | 93.1 |
| sweBench | 64 | 64.3 |
| mmlu | 92.1 | 92.3 |
| gpqa | 84.3 | 65.2 |
| aime | 95 | 88 |
| Price in/out ($/M) | $14 / $56 | $15 / $75 |
Deep Think outscores Opus 4.7 on GPQA (84 vs 65) and matches its 1M context. Opus 4.7 still wins on coding, tool-calling and agent reliability in the wild. Pick Deep Think for scientific reasoning over long sources, Opus for anything agentic.
Gemini 3 Deep Think · Claude Opus 4.7 · All AI model comparisons