Claude Mythos vs Gemini 3 Deep Think: Capybara Tier vs Deliberate Reasoning

Claude MythosGemini 3 Deep Think
arenaElo14781429
humanEval98.792.4
sweBench93.964
mmlu94.292.1
gpqa94.684.3
aime97.695
Price in/out ($/M)$25 / $125$14 / $56

Our verdict

Both models target hard reasoning, but with different philosophies: Deep Think uses Gemini 3 Pro plus an inference-time thinking budget; Mythos is a from-scratch frontier model. Mythos leads GPQA by ~10 points and HumanEval by ~6, and beats Deep Think on competition math. Deep Think wins on availability: it ships under Google AI Ultra subscription with no Glasswing barrier.

Claude Mythos · Gemini 3 Deep Think · All AI model comparisons