| Claude Opus 4.8 | GPT-5.5 | |
|---|---|---|
| arenaElo | 1435 | 1432 |
| humanEval | 94.5 | 94.2 |
| sweBench | 67 | 66 |
| mmlu | 93 | 93 |
| gpqa | 67.5 | 68.7 |
| aime | 90 | 92 |
| Price in/out ($/M) | $5 / $25 | $5 / $30 |
On agentic coding Opus 4.8 leads at 69.2% SWE-Bench Pro vs 58.6% for GPT-5.5, and it tops the Super-Agent benchmark as the only model to complete every case end-to-end at parity on cost. Opus 4.8 is the pick for long-horizon autonomous agents and computer-use; GPT-5.5 remains a strong all-rounder where you are already invested in the OpenAI stack. As always, benchmark both on your real workload before committing.