Opus 5 Scores 2.3x Higher Than GPT-5.6 on ARC-AGI-3 Framework
Despite OpenAI's framework adjustments boosting GPT-5.6 Sol to 38.3%, Opus 5 achieved 30.2% on the standard benchmark, highlighting a significant performance gap in reasoning capabilities.
Woofun AI reports that OpenAI attributed GPT-5.6 Sol’s initial low score of 13.3% on ARC-AGI-3 to context window limitations and memory loss within the general evaluation framework. After modifying the setup via the Responses API to preserve reasoning history, GPT-5.6 Sol’s score rose to 38.3% while reducing output tokens to one-sixth of the original volume.
However, Opus 5 achieved a 30.2% score using the unmodified general framework with High reasoning tier settings. This result indicates that Opus 5 outperforms GPT-5.6 Sol significantly under standard conditions, as the latter required framework alterations to improve its standing.
Comments
No comments yet.