| GPT-5.5 | GPT-4o | |
|---|---|---|
| arenaElo | 1432 | 1345 |
| humanEval | 94.2 | 90.2 |
| sweBench | 66 | 33 |
| mmlu | 93 | 88.7 |
| gpqa | 68.7 | 53.6 |
| aime | 92 | 42 |
| Price in/out ($/M) | $12 / $48 | $2.5 / $10 |
GPT-5.5 beats GPT-4o on every benchmark by double digits and has 4.5× the context. GPT-4o only stays relevant for pipelines that depend on the Realtime API voice stack. New builds should start on GPT-5.5.