| Claude Mythos | Claude Sonnet 4.6 | |
|---|---|---|
| arenaElo | 1478 | 1358 |
| humanEval | 98.7 | 89 |
| sweBench | 93.9 | 60 |
| mmlu | 94.2 | 88.7 |
| gpqa | 94.6 | 59.4 |
| aime | 97.6 | 80 |
| Price in/out ($/M) | $25 / $125 | $3 / $15 |
Different jobs entirely: Sonnet 4.6 is the production default (generally available, $3 / $15 per M tokens, sub-second latency for chat UX). Mythos is a frontier research preview behind Project Glasswing at $25 / $125, mostly used for autonomous cybersecurity work and graduate-level reasoning. Almost no team should choose Mythos when Sonnet 4.6 will do the job; almost no team should choose Sonnet when the task actually needs Mythos.
Claude Mythos · Claude Sonnet 4.6 · All AI model comparisons