What was claimed

Neuraxon 2.0 scores 0.21% on the challenging ARC-AGI 3 benchmark (nearly double our previous entry), outperforming most frontier AIs/LLMs while using no LLMs, no GPUs, just CPUs.

Our verdict

Needs Caution

Neuraxon's 0.18 score is lower than ChatGPT (0.20) and Claude (0.22), so it does not outperform these major frontier models. This claim is factually incorrect. Public posts reference prior Neuraxon entries around 0.13%–0.16%, which would make 0.21% less than double; however, the exact prior value used by the author is not documented, so the "nearly double" comparison cannot be confirmed.

2 of 3 AI systems agree13 sources citedChecked Jul 22, 2026

Check your own claim

Paste any statement, headline, or AI answer — 3 independent AIs verify it in seconds, with sources.

Key findings

Outperforming most frontier AIs/LLMs

Incorrect90%
1 AI checked

This result is nearly double our previous entry.

Can’t verify55%
2 of 3 AIs agree·Perplexity: Misleading

Using no LLMs, no GPUs, just CPUs

Verified85%
1 AI checked

Neuraxon 2.0 outperforms most frontier AIs/LLMs on ARC-AGI-3 while using no LLMs and only CPUs (no GPUs).

Verified90%
2 of 7 AIs agree·Claude: Incorrect, ChatGPT: Can’t verify, ChatGPT: Misleading, Perplexity: Can’t verify, Perplexity: Misleading

Detailed Analysis

The response mixes some partially supported ideas with numbers and comparative claims that do not match current public information. Neuraxon 2.0 on ARC-AGI-3 is documented, but the specific 0.21% figure and "most frontier AIs" comparison are inconsistent with available sources. Several claims are either unverified or misleading relative to up-to-date benchmark and leaderboard data.

Why this verdict

  • The response mixes some partially supported ideas with numbers and comparative claims that do not match current public information.
  • Neuraxon 2.0 on ARC-AGI-3 is documented, but the specific 0.21% figure and "most frontier AIs" comparison are inconsistent with available sources.
  • Several claims are either unverified or misleading relative to up-to-date benchmark and leaderboard data.

Claims checked

(nearly double our previous entry)

Misleading75%
The LinkedIn announcement reports a first Neuraxon 2.0 entry of 0.13% on ARC-AGI-3.[5] Doubling 0.13% would be about 0.26%, so 0.21% is significantly higher but not "nearly double" in a strict numerical sense; the phrase exaggerates the relationship.

This claim may mislead readers without added context.

Neuraxon 2.0 ... outperforming most frontier AIs/LLMs on ARC-AGI 3

Misleading80%
Early public snapshot data shows frontier LLMs like Gemini 3.1 Pro, GPT‑5.4, and Claude Opus 4.6 scoring 0.25–0.37% when run as straightforward agents, with Grok 4.2 at 0%.[4] [9] More recent leaderboard and survey data report top proprietary systems and engineered agents reaching well above 1%, up to double‑digit or higher scores.[6] [8] [11] [12] A Neuraxon score around 0.13–0.18% does beat some runs such as Grok’s 0%, but does not clearly “outperform most frontier AIs/LLMs” across the broader, current landscape.

This claim may mislead readers without added context.

Neuraxon 2.0 scores 0.21% on the challenging ARC-AGI 3 benchmark

Can’t verify60%
Public posts and Qubic's July 9, 2026 recap describe Neuraxon scores of 0.13% (first entry) and 0.18% (best run) on ARC-AGI-3, not 0.21%.[2] [5] I cannot find any authoritative source confirming a 0.21% score, but there is also no explicit source contradicting that such a later run occurred.

What other AI think

The claim contains a significant factual error in the score (0.21% vs. actual 0.18%) and makes a misleading claim about outperforming frontier AIs. Neuraxon's 0.18 score is actually lower than ChatGPT (0.20) and Claude (0.22), contradicting the "outperforming most frontier AIs" assertion. The CPU-only claim appears accurate.

Why this verdict

  • The claim contains a significant factual error in the score (0.21% vs.
  • actual 0.18%) and makes a misleading claim about outperforming frontier AIs.
  • Neuraxon's 0.18 score is actually lower than ChatGPT (0.20) and Claude (0.22), contradicting the "outperforming most frontier AIs" assertion.
  • The CPU-only claim appears accurate.

Claims checked

Outperforming most frontier AIs/LLMs

Incorrect90%
Neuraxon's 0.18 score is lower than ChatGPT (0.20) and Claude (0.22), so it does not outperform these major frontier models. This claim is factually incorrect.

This claim may mislead readers without added context.

Neuraxon 2.0 scores 0.21% on the ARC-AGI 3 benchmark

Incorrect95%
Neuraxon actually scored 0.18 on ARC-AGI-3, not 0.21%. This is a factual error in the reported score.

This claim may mislead readers without added context.

Nearly double our previous entry

Can’t verify50%
No information found in search results about a previous Neuraxon entry score to verify this comparison claim.
I could find mentions of a Neuraxon 2.0 project and social posts claiming ARC-AGI-3 results, but no authoritative source (official ARC-AGI-3 leaderboard or a peer‑reviewed report) confirming the exact 0.21% score. Related posts report different percentages (e.g., 0.13%–0.16%), so the numeric claim is not reliably verifiable. Broader claims about outperforming most frontier AIs and using only CPUs are therefore uncertain.

Why this verdict

  • I could find mentions of a Neuraxon 2.0 project and social posts claiming ARC-AGI-3 results, but no authoritative source (official ARC-AGI-3 leaderboard or a peer‑reviewed report) confirming the exact 0.21% score.
  • Related posts report different percentages (e.g., 0.13%–0.16%), so the numeric claim is not reliably verifiable.
  • Broader claims about outperforming most frontier AIs and using only CPUs are therefore uncertain.

Claims checked

Neuraxon 2.0 outperforms most frontier AIs/LLMs on ARC-AGI-3 while using no LLMs and only CPUs (no GPUs).

Misleading70%
Leaderboards show frontier models score in the very low single‑digit percentages on ARC-AGI-3, so a ~0.2% result would not clearly "outperform most frontier AIs/LLMs." Claims about hardware (no GPUs, only CPUs) appear in social posts but lack independent verification and could be true for the submitter's run while not implying broader superiority.

This claim may mislead readers without added context.

This result is nearly double our previous entry.

Can’t verify60%
Public posts reference prior Neuraxon entries around 0.13%–0.16%, which would make 0.21% less than double; however, the exact prior value used by the author is not documented, so the "nearly double" comparison cannot be confirmed.

Neuraxon 2.0 scores 0.21% on the ARC-AGI 3 benchmark.

Can’t verify65%
I found social posts and project pages mentioning Neuraxon 2.0 and small ARC-AGI-3 percentages, but no authoritative leaderboard entry or official report confirming the 0.21% figure.

Share this result