Grok 4 Heavy
Last updated: April 22, 2026
Specs
| Vendor | xAI |
|---|
| Released | July 9, 2025 |
|---|
| Context window | 256K tokens |
|---|
| Input | $15 / M tokens |
|---|
| Output | $60 / M tokens |
|---|
Benchmarks
| mmlu | 89.3 |
|---|
| gpqa | 74.5 |
|---|
| humanEval | 88.4 |
|---|
| arenaElo | 1391 |
|---|
| sweBench | 55 |
|---|
| sweBenchPro | 50 |
|---|
| liveCodeBench | 70 |
|---|
| aiderPolyglot | 66 |
|---|
| terminalBench | 40 |
|---|
| mmluPro | 85 |
|---|
| hle | 24 |
|---|
| aime | 92 |
|---|
| math500 | 96 |
|---|
| mmmu | 74 |
|---|
Strengths
- Multi-agent reasoning mode: several Grok instances debate
- Strong on Humanity's Last Exam and competition math
- Realtime X (Twitter) grounding
Weaknesses
- Expensive per reply because of multi-agent debate
- Smaller developer ecosystem than OpenAI/Anthropic
- Not suited for realtime UX
Best for
- Hard one-shot reasoning with parallel debate
- Research workloads leaning on X data
- Edgy use cases where other models refuse
Compare with other AI models