What was claimed

GPT-6 Astra surpasses human performance on 96% of ARC-AGI-3 levels and arc-agi-3 is now saturated.

Our verdict

Inaccurate

The official ARC Prize results still show non-perfect performance under the standard harness, including 62.7% on ARC-AGI-3, which is not saturation. The benchmark only approaches near-complete scores under a special provider adapter harness, not universally. (Only 2 of 3 AI systems responded.)

0 of 2 AI systems agree20 sources citedChecked Sep 4, 2026

Check your own claim

Paste any statement, headline, or AI answer — 3 independent AIs verify it in seconds, with sources.

Key findings

ARC-AGI-3 is now saturated.

Incorrect93%
1 of 2 AIs agree·Claude: Misleading

GPT-6 Astra surpasses human performance on 96% of ARC-AGI-3 levels.

Verified97%
All 2 AIs agree

Detailed Analysis

The claim is partly supported but overstated. ARC Prize says GPT-6 Astra surpassed the human baseline in action efficiency on 96% of ARC-AGI-3 levels, but it does not say the benchmark is saturated; the official ARC-AGI-3 results still show 62.7% with the standard harness and 99.95%/98.55% under the provider adapter setup, which is not the same as full saturation.[1] [2] [16]

Why this verdict

  • The claim is partly supported but overstated.
  • ARC Prize says GPT-6 Astra surpassed the human baseline in action efficiency on 96% of ARC-AGI-3 levels, but it does not say the benchmark is saturated; the official ARC-AGI-3 results still show 62.7% with the standard harness and 99.95%/98.55% under the provider adapter setup, which is not the same as full saturation.[1] [2] [16]

Claims checked

ARC-AGI-3 is now saturated.

Incorrect93%
The official ARC Prize results still show non-perfect performance under the standard harness, including 62.7% on ARC-AGI-3, which is not saturation.[1] [2] [16] The benchmark only approaches near-complete scores under a special provider adapter harness, not universally.

This claim may mislead readers without added context.

GPT-6 Astra surpasses human performance on 96% of ARC-AGI-3 levels.

Verified98%
ARC Prize explicitly says Astra used fewer actions than the median tested human on 96% of levels.[1] [4] [7]

What other AI think

The first claim about surpassing human performance on 96% of levels is accurate and well-documented. However, the second claim about ARC-AGI-3 being "saturated" is misleading and contradicted by official sources. The statement conflates different metrics and lacks important context about how the benchmark is measured.

Why this verdict

  • The first claim about surpassing human performance on 96% of levels is accurate and well-documented.
  • However, the second claim about ARC-AGI-3 being "saturated" is misleading and contradicted by official sources.
  • The statement conflates different metrics and lacks important context about how the benchmark is measured.

Claims checked

ARC-AGI-3 is now saturated

Misleading90%
While Greg Brockman (OpenAI) tweeted 'arc-agi-3 is now saturated,' this statement is misleading without context. The benchmark itself remains far from saturated for most frontier models. Astra's high scores depend heavily on specific harness configurations (Provider Adapter harness achieves 99.9%, but Standard harness only 62.7%). The benchmark was designed to be 'hard for AI' and remains challenging for most systems.

This claim may mislead readers without added context.

GPT-6 Astra surpasses human performance on 96% of ARC-AGI-3 levels

Verified95%
Multiple official sources confirm that Astra surpasses the human action-efficiency baseline on 96% of ARC-AGI-3 levels. This is measured by action efficiency (how quickly/efficiently the model solves tasks compared to median human performance).
ChatGPTDid not respond

Share this result