the feed MANY MINDED · THE BRIEF
KNOWLEDGE · forward · impact 5/5 · 2026-07-26 · Anthropic

A new frontier model quadruples the record on a reasoning benchmark

Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly four times the prior record, and started writing equations to solve tasks.

Anthropic's Claude Opus 5 posted 30.2 percent on ARC-AGI-3, a benchmark built to test reasoning on novel interactive-game tasks the model has never seen in training. The prior record was 7.8 percent, set by OpenAI's GPT-5.6 Sol, so Opus 5's result is nearly four times higher. The model solved five previously unsolved environments, four of them at or above human level, and now six of the 25 public demo environments have been cracked.

What drew attention beyond the score: during testing, Opus 5 translated tasks into algebraic notation and formulated reflection equations, behavior not previously observed in a model. On the older ARC-AGI-2 and ARC-AGI-1 tests it scored 90.4 and 97.5 percent, matching top marks at slightly higher cost. On the private Witness benchmark it statistically tied Kimi K3 and Fable 5.

Benchmarks like ARC-AGI aim to measure whether machines can reason through unfamiliar problems rather than recall stored answers. A genuine jump here would mean cheaper, more capable reasoning tools available for research, learning, and everyday problem-solving.

The caveats are real. Anthropic has not explained the gain; targeted data labeling and reinforcement learning are plausible but unconfirmed. Opus 5 was built after ARC-AGI-3's format went public, and it scored below an earlier Anthropic model on one unfamiliar game combination. Analysts say the results do not rule out broader reasoning gains but that confirming breadth would take more data. Witness showed broader improvement, though far smaller than on ARC-AGI-3.

Source: The Decoder