Anthropic's Claude Opus 5 scores 30.2% on ARC-AGI-3 benchmark
Anthropic released Claude Opus 5, which scored 30.2 percent on the ARC-AGI-3 benchmark, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The model improves long-horizon reasoning, agentic coding, and professional knowledge work while making capabilities previously associated with the most expensive frontier systems more economically usable. The benchmark's developers say Opus 5 independently formulated reflection equations, a behavior they had never seen from another model, and attribute this to stronger logical reasoning. Frontier models are becoming better at sustaining complex work over time: understanding large systems, planning across many steps, using tools, revising decisions, and maintaining coherence throughout long tasks.