What was claimed
Grok 4.5 leads VulcanBench’s new coding benchmark with 91.3%, solving 21 of 23 real-world software tasks, beating Claude Fable 5 and GPT-5.6 Sol while owning the cost-efficiency frontier
Our verdict
InaccurateAcross neutral coding benchmarks, Claude Fable 5 consistently scores higher than Grok 4.5: for example, DeepSWE 1.1 (70% vs 53%), SWE Bench Pro (80.4% vs around mid‑field for Grok 4.5), and aggregate coding leaderboards where Fable 5 tops charts and Grok trails frontier models. These sources show Grok 4.5 does not generally beat Fable 5 in coding accuracy. Sources agree Grok 4.5 is highly cost‑efficient and more token‑efficient than many frontier models, with pricing at $2/M input and $6/M output and far fewer output tokens per SWE Bench Pro task than Opus 4.8. However, other models like GPT‑5.6 Luna and Terra have competitive or lower input/output prices, and analyses explicitly state that Grok 4.5 trails frontier models in accuracy, so claiming it "owns" the entire cost‑efficiency frontier over all competitors overstates and simplifies the situation.
Check your own claim
Paste any statement, headline, or AI answer — 3 independent AIs verify it in seconds, with sources.
Key findings
Grok 4.5 beat Claude Fable 5 and GPT-5.6 Sol on that VulcanBench report.
Grok 4.5 "owns the cost-efficiency frontier."
Grok 4.5 leads VulcanBench’s new coding benchmark with 91.3%, solving 21 of 23 real-world software tasks.