OpenAI o4
Last updated: April 22, 2026
Specs
| Vendor | OpenAI |
|---|
| Released | February 12, 2026 |
|---|
| Context window | 256K tokens |
|---|
| Input | $12 / M tokens |
|---|
| Output | $48 / M tokens |
|---|
Benchmarks
| mmlu | 92.5 |
|---|
| gpqa | 87.1 |
|---|
| humanEval | 94 |
|---|
| arenaElo | 1438 |
|---|
| sweBench | 72 |
|---|
| sweBenchPro | 60 |
|---|
| liveCodeBench | 85 |
|---|
| aiderPolyglot | 83 |
|---|
| terminalBench | 53 |
|---|
| mmluPro | 88 |
|---|
| hle | 28 |
|---|
| aime | 94 |
|---|
| math500 | 98 |
|---|
| mmmu | 80 |
|---|
Strengths
- Current state-of-the-art on GPQA (87%)
- Thinking budget you can tune per request
- Beats o3 on every reasoning benchmark
Weaknesses
- Slow: up to minutes per reply
- Reasoning tokens inflate real cost per answer
- Not suited for realtime chat UX
Best for
- Deep research assistants
- Math olympiad + science problem solving
- Code review on hard algorithmic bugs
Compare with other AI models