What was claimed

AGI is near after Opus 5 scores 30.2% on ARC-AGI-3

Our verdict

Needs caution

While the score is impressive, the benchmark was specifically designed to counter claims that AGI is near. Humans score 100%, the test is far from saturated, and experts note the unclear real-world applicability of the benchmark result.

2 of 3 AI systems agree16 sources citedChecked Jul 25, 2026

Check your own claim

Paste any statement, headline, or AI answer — 3 independent AIs verify it in seconds, with sources.

Key findings

AGI is near after Opus 5 scores 30.2% on ARC-AGI-3

Misleading95%
1 AI checked

AGI is near after Opus 5 scores 30.2% on ARC-AGI-3.

Verified97%
3 of 5 AIs agree·ChatGPT: Misleading, Perplexity: Misleading

Detailed Analysis

The numerical part of the statement is accurate and well supported by multiple independent sources. However, the conclusion that AGI is near based solely on this score is speculative and not supported by current expert consensus or by ARC-AGI-3 benchmark authors.

Why this verdict

  • The numerical part of the statement is accurate and well supported by multiple independent sources.
  • However, the conclusion that AGI is near based solely on this score is speculative and not supported by current expert consensus or by ARC-AGI-3 benchmark authors.

Claims checked

AGI is near after Opus 5 scores 30.2% on ARC-AGI-3

Misleading82%
ARC-AGI-3 is designed as a difficult fluid reasoning benchmark and prior commentary explicitly cautioned that even record scores like 1.5% were not grounds to declare AGI is near.[16] Extrapolating one benchmark result to a claim that AGI is near is speculative and not directly supported by the benchmark designers or broader expert consensus.[16]

This claim may mislead readers without added context.

Opus 5 scores 30.2% on ARC-AGI-3

Verified97%
Multiple independent benchmark reports and commentary state that Claude Opus 5 achieves about 30.2% on ARC-AGI-3, often noting this as three times the next-best model's score.[1] [2] [3] [4] [5] [6] [9] [11]

What other AI think

The first part of the claim (Opus 5's 30.2% score) is factually accurate. However, the conclusion that "AGI is near" based on this score is misleading and contradicts the benchmark's own purpose. The benchmark was designed to demonstrate that frontier AI systems still lack critical capabilities despite industry claims of near-AGI performance.

Why this verdict

  • The first part of the claim (Opus 5's 30.2% score) is factually accurate.
  • However, the conclusion that "AGI is near" based on this score is misleading and contradicts the benchmark's own purpose.
  • The benchmark was designed to demonstrate that frontier AI systems still lack critical capabilities despite industry claims of near-AGI performance.

Claims checked

AGI is near after Opus 5 scores 30.2% on ARC-AGI-3

Misleading95%
While the score is impressive, the benchmark was specifically designed to counter claims that AGI is near. Humans score 100%, the test is far from saturated, and experts note the unclear real-world applicability of the benchmark result.

This claim may mislead readers without added context.

Opus 5 scores 30.2% on ARC-AGI-3

Verified99%
Multiple verified sources confirm Opus 5 achieved a 30.2% score on ARC-AGI-3, representing a significant jump from previous models.
The numeric score claim is supported by multiple benchmark reports and system-card disclosures. The conclusion that "AGI is near" is a broad, causal inference that is not justified by a single benchmark result and is therefore misleading.

Why this verdict

  • The numeric score claim is supported by multiple benchmark reports and system-card disclosures.
  • The conclusion that "AGI is near" is a broad, causal inference that is not justified by a single benchmark result and is therefore misleading.

Claims checked

AGI is near after Opus 5 scores 30.2% on ARC-AGI-3.

Misleading70%
A single benchmark improvement does not by itself demonstrate nearing AGI; ARC-AGI-3 measures specific interactive problem-solving and does not capture the full set of capabilities, generality, safety, or independent agency that would constitute AGI.

This claim may mislead readers without added context.

Opus 5 scored 30.2% on ARC-AGI-3.

Verified95%
Multiple public reports and the Opus 5 system card/leaderboard entries list a verified ARC-AGI-3 result of about 30.16–30.2% at high effort.

Share this result