| GPT-5.5 | OpenAI o4 | |
|---|---|---|
| arenaElo | 1432 | 1438 |
| humanEval | 94.2 | 94 |
| sweBench | 66 | 72 |
| mmlu | 93 | 92.5 |
| gpqa | 68.7 | 87.1 |
| aime | 92 | 94 |
| Price in/out ($/M) | $12 / $48 | $12 / $48 |
o4 destroys GPT-5.5 on GPQA (87% vs 69%) but thinks for seconds to minutes per reply. GPT-5.5 wins on everyday throughput, multimodality and cost per interaction. Route hard reasoning to o4, everything else to GPT-5.5.