| Claude Fable 5 | OpenAI o4 | |
|---|---|---|
| arenaElo | 1492 | 1438 |
| humanEval | 98 | 94 |
| sweBench | 77 | 72 |
| mmlu | 95.1 | 92.5 |
| gpqa | 92 | 87.1 |
| aime | 94 | 94 |
| Price in/out ($/M) | $10 / $50 | $12 / $48 |
o4 is a deep single-answer reasoner that thinks for seconds to minutes per reply. Fable 5 is built for long-horizon autonomous work (multi-hour agentic runs, end-to-end deliverables and self-verifying builds) with always-on thinking you tune via effort. For one-shot hard reasoning, o4 still competes; for agents that must plan, act and ship over a long session, Fable 5 is in a different category.