GPT-5 Codex
Last updated: April 22, 2026
Specs
| Vendor | OpenAI |
|---|
| Released | September 15, 2025 |
|---|
| Context window | 400K tokens |
|---|
| Input | $8 / M tokens |
|---|
| Output | $32 / M tokens |
|---|
Benchmarks
| mmlu | 87.2 |
|---|
| gpqa | 58.4 |
|---|
| humanEval | 94.8 |
|---|
| arenaElo | 1395 |
|---|
| sweBench | 68 |
|---|
| sweBenchPro | 60 |
|---|
| liveCodeBench | 82 |
|---|
| aiderPolyglot | 82 |
|---|
| terminalBench | 50 |
|---|
| mmluPro | 80 |
|---|
| hle | 16 |
|---|
| aime | 85 |
|---|
| math500 | 94 |
|---|
Strengths
- Specialized coding fork of GPT-5: top HumanEval
- Long-horizon agentic coding up to 7 hours
- Powers Codex CLI and Codex cloud by default
Weaknesses
- Not built for multimodal or general chat
- Lower GPQA than GPT-5 on non-code reasoning
- Overkill for simple autocomplete tasks
Best for
- Autonomous coding agents
- PR generation and multi-file refactors
- IDE copilots that need GPT-5-class coding
Compare with other AI models