OpenAI o4 vs Claude Opus 4.7: Deep Reasoning vs Long-Horizon Agents

OpenAI o4Claude Opus 4.7
arenaElo14381420
humanEval9493.1
sweBench7264.3
mmlu92.592.3
gpqa87.165.2
aime9488
Price in/out ($/M)$12 / $48$15 / $75

Our verdict

o4 is state-of-the-art on GPQA (87% vs 65%) and wins one-shot math and science. Opus 4.7 runs faster, has 4× the context and beats o4 on multi-hour agentic workflows. Pick o4 for hard single answers; pick Opus 4.7 for agents that must ship.

OpenAI o4 · Claude Opus 4.7 · All AI model comparisons