| GPT-5 Codex | Codestral | |
|---|---|---|
| arenaElo | 1395 | 1254 |
| humanEval | 94.8 | 89.7 |
| sweBench | 68 | 35 |
| mmlu | 87.2 | 73.2 |
| gpqa | 58.4 | 41 |
| aime | 85 | 30 |
| Price in/out ($/M) | $8 / $32 | $0.3 / $0.9 |
Codestral is 25× cheaper and still clears 89% HumanEval, remarkable for inline autocomplete and refactors. GPT-5 Codex wins on agentic coding, multi-file edits and reasoning about unfamiliar code. Mix them: Codestral for suggestions, Codex for agents.