| Claude Mythos | Gemini 3 Deep Think | |
|---|---|---|
| arenaElo | 1478 | 1429 |
| humanEval | 98.7 | 92.4 |
| sweBench | 93.9 | 64 |
| mmlu | 94.2 | 92.1 |
| gpqa | 94.6 | 84.3 |
| aime | 97.6 | 95 |
| Price in/out ($/M) | $25 / $125 | $14 / $56 |
Both models target hard reasoning, but with different philosophies: Deep Think uses Gemini 3 Pro plus an inference-time thinking budget; Mythos is a from-scratch frontier model. Mythos leads GPQA by ~10 points and HumanEval by ~6, and beats Deep Think on competition math. Deep Think wins on availability: it ships under Google AI Ultra subscription with no Glasswing barrier.
Claude Mythos · Gemini 3 Deep Think · All AI model comparisons