| Claude Mythos | Gemini 3 Ultra | |
|---|---|---|
| arenaElo | 1478 | 1441 |
| humanEval | 98.7 | 91 |
| sweBench | 93.9 | 66 |
| mmlu | 94.2 | 93.4 |
| gpqa | 94.6 | 72.6 |
| aime | 97.6 | 93 |
| Price in/out ($/M) | $25 / $125 | $18 / $72 |
Gemini 3 Ultra has a 3× larger context window (3M vs 1M) and native video, which Mythos does not match. Mythos returns the favor on raw capability: ~23 GPQA points and ~8 HumanEval points ahead, plus saturated cybersecurity benchmarks. For whole-movie or 2000-page workloads on Vertex AI, Ultra is the right pick. For frontier reasoning and code where your data fits in 1M tokens, Mythos, if you can get access.