What was claimed

Prime Intellect is releasing a self-improving Prime Agent that scores 95.5% on ARC-AGI-3 (above human baseline) and the gains are not benchmark-specific

Our verdict

Needs caution

ARC-AGI-3 documentation and announcements state humans score 100% on the benchmark, so 95.5% would be below the published human baseline.

1 of 3 AI systems agree19 sources citedChecked Aug 6, 2026

Check your own claim

Paste any statement, headline, or AI answer — 3 independent AIs verify it in seconds, with sources.

Key findings

95.5% is above the human baseline on ARC-AGI-3.

Incorrect90%
1 AI checked

Prime Intellect is releasing a self-improving Prime Agent.

Verified90%
All 2 AIs agree

The gains are not benchmark-specific (i.e., improvements generalize beyond ARC-AGI-3).

Verified90%
1 of 2 AIs agree·ChatGPT: Can’t verify

Prime Agent’s 95.5% score on ARC-AGI-3 is above the human baseline

Verified95%
2 of 3 AIs agree·ChatGPT: Can’t verify

Detailed Analysis

All key factual elements of the statement match Prime Intellect’s own announcement and independent coverage. The only subtlety is that the 95.5% figure refers to ARC-AGI-3’s RHAE Best@1 metric, but that is the standard way scores are reported, and the claim is consistent with sources.

Why this verdict

  • All key factual elements of the statement match Prime Intellect’s own announcement and independent coverage.
  • The only subtlety is that the 95.5% figure refers to ARC-AGI-3’s RHAE Best@1 metric, but that is the standard way scores are reported, and the claim is consistent with sources.

Claims checked

The gains from Prime Agent are not benchmark-specific

Verified90%
Prime Intellect’s X post states that although Prime Agent scores 95.5% on ARC-AGI-3, “the gain is not benchmark-specific,” and other posts describe improvements across multiple models compared to proprietary harnesses.[2] [5]

Prime Intellect is releasing a self-improving Prime Agent

Verified95%
Prime Intellect’s blog explicitly introduces Prime Agent as “A self-improving RLM agent” and describes its release and capabilities.[1] [4] [5]

Prime Agent scores 95.5% on ARC-AGI-3

Verified95%
Prime Intellect’s technical blog states that Prime Agent with Opus 5 achieves 95.5% RHAE Best@1 on ARC-AGI-3, and this figure is repeated in their social posts and third-party coverage.[1] [2] [3]

What other AI think

Prime Intellect is releasing Prime Agent, a self-improving coding harness that scores 95.5% on ARC-AGI-3, surpassing the human-expert baseline, with gains that are not benchmark-specific. All major claims in the statement are directly supported by official sources from Prime Intellect and recent reporting.

Why this verdict

  • Prime Intellect is releasing Prime Agent, a self-improving coding harness that scores 95.5% on ARC-AGI-3, surpassing the human-expert baseline, with gains that are not benchmark-specific.
  • All major claims in the statement are directly supported by official sources from Prime Intellect and recent reporting.
Part of the statement matches public reporting about Prime Intellect releasing agent-training tools. However the specific numeric score and the claim it exceeds the human baseline lack independent confirmation and contradict ARC-AGI-3's published human baseline. The claim that gains are "not benchmark-specific" is unverified.

Why this verdict

  • Part of the statement matches public reporting about Prime Intellect releasing agent-training tools.
  • However the specific numeric score and the claim it exceeds the human baseline lack independent confirmation and contradict ARC-AGI-3's published human baseline.
  • The claim that gains are "not benchmark-specific" is unverified.

Claims checked

95.5% is above the human baseline on ARC-AGI-3.

Incorrect90%
ARC-AGI-3 documentation and announcements state humans score 100% on the benchmark, so 95.5% would be below the published human baseline.

This claim may mislead readers without added context.

The Prime Agent scores 95.5% on ARC-AGI-3.

Can’t verify45%
I could not find any authoritative report, benchmark entry, or ARC-AGI-3 result page attributing a 95.5% score to a "Prime Agent." No primary-source evidence confirms that exact numeric result.

The gains are not benchmark-specific (i.e., improvements generalize beyond ARC-AGI-3).

Can’t verify50%
There is no published evaluation showing the claimed improvements generalize across multiple independent benchmarks; broader generalization claims require evidence from diverse, peer-reviewed evaluations.

Share this result