GPT-5.4 Codex
Last updated: April 22, 2026
Specs
| Vendor | OpenAI |
|---|
| Released | February 20, 2026 |
|---|
| Context window | 500K tokens |
|---|
| Input | $9 / M tokens |
|---|
| Output | $36 / M tokens |
|---|
Benchmarks
| mmlu | 88.6 |
|---|
| gpqa | 60.8 |
|---|
| humanEval | 95.5 |
|---|
| arenaElo | 1408 |
|---|
| sweBench | 70 |
|---|
| sweBenchPro | 64 |
|---|
| liveCodeBench | 84 |
|---|
| aiderPolyglot | 84 |
|---|
| terminalBench | 52 |
|---|
| mmluPro | 82 |
|---|
| hle | 18 |
|---|
| aime | 86 |
|---|
| math500 | 95 |
|---|
Strengths
- Current best HumanEval of any OpenAI model
- Handles multi-hour coding sessions without drift
- Fine-tuned for tool-calling via Codex CLI
Weaknesses
- Not positioned for general chat (use GPT-5.5)
- Still weaker than Opus 4.7 on agent reliability
- Pricing sits above GPT-5 Codex at modest quality gains
Best for
- Automated PR factories
- Long refactor runs where drift matters
- Enterprise Codex deployments on tier 5
Compare with other AI models