| GPT-5 | Claude Opus 4.7 | |
|---|---|---|
| arenaElo | 1412 | 1420 |
| humanEval | 91.6 | 93.1 |
| sweBench | 65 | 64.3 |
| mmlu | 91.5 | 92.3 |
| gpqa | 63.8 | 65.2 |
| aime | 92 | 88 |
| Price in/out ($/M) | $1.25 / $10 | $5 / $25 |
Claude Opus 4.7 leads on long-context coding and agentic workflows; GPT-5 wins on multimodal breadth (native audio) and tooling maturity. For enterprise coding agents, pick Opus 4.7. For voice-first multimodal apps, pick GPT-5.