Claude Code vs GPT-5.4 Codex: Agent Wrapper vs Raw Model

Claude CodeGPT-5.4 Codex
arenaElo14201408
humanEval93.195.5
sweBench64.370
mmlu92.388.6
gpqa65.260.8
aime8886
Price in/out ($/M)$15 / $75$9 / $36

Our verdict

GPT-5.4 Codex scores higher on HumanEval in isolation, but Claude Code's orchestration, hooks and memory give Opus 4.7 the edge on long runs. If your workflow fits inside a CLI, Claude Code wins on wall-clock time; if you integrate into your own harness, pick 5.4 Codex directly.

Claude Code · GPT-5.4 Codex · All AI model comparisons