GPT-5.4 Codex vs Claude Opus 4.7: The Code-Writing Decider

GPT-5.4 CodexClaude Opus 4.7
arenaElo14081420
humanEval95.593.1
sweBench7064.3
mmlu88.692.3
gpqa60.865.2
aime8688
Price in/out ($/M)$9 / $36$15 / $75

Our verdict

5.4 Codex holds the HumanEval crown (95.5 vs 93.1) and is cheaper per token. Opus 4.7 compensates with 2.5× the context and stronger long-horizon agents. Specialized code pipelines: Codex. General coding agents that also reason and plan: Opus.

GPT-5.4 Codex · Claude Opus 4.7 · All AI model comparisons