OpenAI is pushing back after Anthropic's Claude Opus 5 dominated the ARC-AGI-3 benchmark, publishing results that show its own GPT-5.6 Sol model hitting 38.3 percent — provided you use OpenAI's preferred settings rather than the official test harness.
The dispute, which played out on July 30, 2026, cuts to the heart of a growing problem in AI evaluation: benchmarks no longer measure just a model, but the entire technical scaffold wrapped around it. For deeper analysis of the latest AI developments and model releases, AI Buzz Wire is tracking the story.
The Numbers and the Catch
OpenAI reported that GPT-5.6 Sol reaches 38.3 percent on ARC-AGI-3, edging past Anthropic's Claude Opus 5, which scored 30.2 percent after having previously quadrupled the benchmark's record. The caveat is significant: that 38.3 percent figure depends on two specific features of OpenAI's Responses API.
The first is Retained Reasoning, which preserves the model's chain of thought between steps rather than discarding it after each action. The second is Compaction, which summarizes older context instead of truncating it, allowing the model to keep working on long problems without losing earlier reasoning.
Run through ARC Prize's official standardized harness — which intentionally avoids provider-specific settings to ensure apples-to-apples comparisons — GPT-5.6 Sol managed just 7.8 percent. The reason, OpenAI argues, is that the official setup discards the model's reasoning after every step, gutting performance on tasks that require sustained multi-step thinking.
A Benchmark Built for Purity
ARC-AGI-3 was designed by François Chollet and the ARC Prize team to probe a model's ability to solve novel reasoning puzzles it has never seen before — a test of general intelligence rather than memorized knowledge. To keep comparisons fair, the official harness deliberately strips away provider-specific features, measuring raw model capability in a neutral environment.
OpenAI's counterargument is that this neutrality can become a handicap. The company contends that the official harness relied on an older "OpenAI-style completions API" that lacked capabilities the Claude API already offered, effectively putting OpenAI at a structural disadvantage relative to Anthropic in the original comparison.
In essence, OpenAI is saying: you measured our model with one hand tied behind its back.
Chollet's Response
ARC Prize co-founder François Chollet responded by drawing a clear line between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. But general-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game.
In effect, Chollet conceded that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage. He noted that the organization has had "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed the company "starting to figure out the answer."
Chollet did flag one remaining concern. Different providers using different settings creates what he called "a potential parity issue" — but he considers that acceptable "as long as the settings and the cost are clearly reported."
Why This Matters Beyond the Leaderboard
The episode underscores a tension that is reshaping how the AI industry evaluates progress. As models are increasingly deployed through feature-rich APIs rather than as standalone systems, the line between a model's inherent ability and the engineering around it keeps blurring. A benchmark that ignores scaffolding may understate real-world performance; one that embraces it may reward companies for gaming the test.
The ARC-AGI-3 debate is unlikely to be the last. As frontier models cluster near the top of established benchmarks, disagreements over what counts as a fair measurement — and whose harness gets to set the standard — will only intensify.
Stay Ahead of AI
Want to follow the breaking AI news on models, benchmarks, and the labs building them? AI Buzz Wire has you covered.
Read more AI news