| Grok 4 Heavy | OpenAI o4 | |
|---|---|---|
| arenaElo | 1391 | 1438 |
| humanEval | 88.4 | 94 |
| sweBench | 55 | 72 |
| mmlu | 89.3 | 92.5 |
| gpqa | 74.5 | 87.1 |
| aime | 92 | 94 |
| Price in/out ($/M) | $15 / $60 | $12 / $48 |
o4 is cleaner on GPQA and has a tighter ecosystem. Grok 4 Heavy's multi-agent debate mode occasionally beats o4 on Humanity's Last Exam-style evals but costs more per answer. For production reasoning, o4. For research benchmarks, either.