Claude Mythos
Last updated: May 9, 2026
Specs
| Vendor | Anthropic |
|---|
| Released | April 8, 2026 |
|---|
| Context window | 1M tokens |
|---|
| Input | $25 / M tokens |
|---|
| Output | $125 / M tokens |
|---|
Benchmarks
| mmlu | 94.2 |
|---|
| gpqa | 94.6 |
|---|
| humanEval | 98.7 |
|---|
| arenaElo | 1478 |
|---|
| sweBench | 93.9 |
|---|
| sweBenchPro | 88 |
|---|
| liveCodeBench | 90 |
|---|
| aiderPolyglot | 90 |
|---|
| terminalBench | 65 |
|---|
| mmluPro | 91 |
|---|
| hle | 35 |
|---|
| aime | 97.6 |
|---|
| math500 | 99.5 |
|---|
| mmmu | 85 |
|---|
Strengths
- Frontier of frontiers: 93.9% SWE-bench Verified, 97.6% USAMO 2026, 94.6% GPQA Diamond
- 4.3× jump over the previous performance trendline; saturates Cybench at 100% pass@1
- Long-context that actually holds: 80% on GraphWalks BFS at 1M tokens vs 21.4% from GPT-5.4
Weaknesses
- Gated research preview: only Project Glasswing partners and a handful of vetted critical-infrastructure operators get keys
- List price is 5× Opus 4.7 ($25 / $125 per million tokens) when access is granted
- No general API self-serve: Anthropic states it does not plan to release Mythos publicly
Best for
- Defensive cybersecurity research at nation-scale (zero-day discovery + patch authoring)
- Mathematical proof generation and graduate-level scientific reasoning
- Long-horizon coding agents that need to reason across an entire codebase in a single window
Compare with other AI models