| OpenAI o4 | Gemini 3 Deep Think | |
|---|---|---|
| arenaElo | 1438 | 1429 |
| humanEval | 94 | 92.4 |
| sweBench | 72 | 64 |
| mmlu | 92.5 | 92.1 |
| gpqa | 87.1 | 84.3 |
| aime | 94 | 95 |
| Price in/out ($/M) | $12 / $48 | $14 / $56 |
o4 is narrowly ahead on GPQA (87 vs 84) and HumanEval; Gemini 3 Deep Think wins on context (1M vs 256K) and pricing ($14 vs $12 input). For math olympiad-grade problems, both are usable: choose by ecosystem.