GPT-5 Codex vs Claude Opus 4.7: Best Coding Model of 2026?

GPT-5 CodexClaude Opus 4.7
arenaElo13951420
humanEval94.893.1
sweBench6864.3
mmlu87.292.3
gpqa58.465.2
aime8588
Price in/out ($/M)$8 / $32$15 / $75

Our verdict

GPT-5 Codex leads HumanEval by a nose (94.8 vs 93.1) and is cheaper. Opus 4.7 has a larger context (1M vs 400K) and better agent stability over multi-hour runs. For one-shot PRs, Codex. For repo-wide refactors and agents, Opus.

GPT-5 Codex · Claude Opus 4.7 · All AI model comparisons