OpenAI's GPT-5.6 Sol Surpasses Rival Opus 5 on ARC-AGI-3 Benchmark with Custom API Settings
OpenAI's GPT-5.6 Sol model has achieved a score of 38.3 percent on the ARC-AGI-3 benchmark, outperforming Anthropic's Opus 5 model, which scored 30.2 percent. This breakthrough was made possible by OpenAI's custom harness with retained reasoning and compaction, highlighting the importance of technical setup in AI model performance.
The latest results from OpenAI mark a significant milestone in the development of artificial intelligence models. By leveraging its custom Responses API with retained reasoning and compaction, OpenAI's GPT-5.6 Sol model has surpassed Anthropic's Opus 5 on the ARC-AGI-3 benchmark, a widely recognized test of logical reasoning. The score of 38.3 percent achieved by GPT-5.6 Sol represents a substantial improvement over the 7.8 percent score obtained using the official test harness, which discards the model's reasoning after each action.
This achievement is particularly notable given the recent performance of Opus 5, which had quadrupled the record score on the ARC-AGI-3 benchmark. The fact that OpenAI's model can now outperform Opus 5 using custom API settings underscores the complex interplay between model architecture, technical setup, and benchmark design. While the official ARC scores use a standardized approach to ensure fair comparisons, the use of custom settings by OpenAI raises important questions about the role of technical setup in determining model performance.
The ARC-AGI-3 benchmark is designed to test pure model performance, without the influence of provider-specific settings. However, OpenAI argues that benchmarks inevitably measure not just the model, but also the technical setup surrounding it. This perspective is supported by the significant performance boost achieved by GPT-5.6 Sol when using OpenAI's custom harness. The retained reasoning feature, which keeps the model's chain of thought between steps, and the compaction feature, which summarizes old context instead of truncating it, are critical to the model's improved performance.
The implications of this breakthrough are far-reaching, with significant consequences for developers, businesses, and everyday users. As AI models become increasingly integrated into various applications, the importance of technical setup and benchmark design will only continue to grow. The fact that different providers may use different settings to achieve optimal performance creates a potential parity issue, as noted by ARC Prize co-founder François Chollet. However, Chollet also emphasized that as long as the settings and costs are clearly reported, such differences can be acceptable.
In historical context, the development of GPT-5.6 Sol and its performance on the ARC-AGI-3 benchmark represent a significant step forward in the evolution of AI models. Previous versions of the GPT model have demonstrated impressive capabilities, but the latest iteration has clearly raised the bar. The ability of OpenAI's model to outperform Opus 5, a rival model from Anthropic, highlights the intense competition in the AI landscape and the rapid pace of innovation.