ARC-AGI-3
5 saves·entity·peaked week of Jul 27, 2026
Weekly volume
Shows up alongside
Topics tagged on the same saves.
Who’s driving it
Most-saved authors on this topic.
Recent saves
The items behind the chart, newest first.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversa…
@fchollet·Sep 3, 2026↗
GPT-6-Astra Benchmarks ARC-AGI-3 - 98.6% FrontierMath Tier 4 v2 - 97.6% DeepSWE v1.1 - 74.1% ExploitBench - 100%
@scaling01·Sep 3, 2026↗
I got 96.2% on ARC-AGI-3 with Opus 5, and 99.3% pass@2. The program is basically Claude Code + Opus 5 (high), one action command, and filesystem logs. Almost nothing ARC specific. https://t.co/NHyib…
@jeremyberman·Aug 12, 2026↗
GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember wha…
@OpenAI·Jul 30, 2026↗
lol did nobody at Anthropic stop for a second and wonder why the numbers looked this absurd before posting the “victory”-tweet? https://t.co/DPdPc04YZT
@steipete·Jul 30, 2026↗