OpenAI says it reclaimed the ARC-AGI-3 lead, but the benchmark dispute is really about tooling
OpenAI says its GPT-5.6 Sol model reached 38.3% on ARC-AGI-3, a score that would put it ahead of Anthropic’s Claude Opus 5 at 30.2% on the same benchmark. On paper, that sounds like a straightforward reversal in a closely watched contest over abstract reasoning. In practice, the claim is more complicated. The higher result depends on running GPT-5.6 Sol with OpenAI’s own Responses API and two specific settings that are not part of the benchmark’s standard test harness.
That distinction matters because ARC-AGI-3 is designed to measure how well a model handles unfamiliar logic problems rather than how well a vendor can wrap its model in product-specific infrastructure. According to the supplied source text, GPT-5.6 Sol scored only 7.8% in the official harness, where the model’s reasoning is discarded after each action. OpenAI’s much higher score came from using “Retained Reasoning,” which preserves the model’s chain of thought between steps, and “Compaction,” which summarizes older context instead of dropping it.
The result is a debate that cuts deeper than one leaderboard update. It raises a now-central question in AI evaluation: when systems improve through memory handling, context management, and API design, where should the line be drawn between model capability and product scaffolding?
Why ARC-AGI-3 has become such a flashpoint
ARC-style benchmarks have unusual status in the AI industry because they aim to test generalization rather than memorization. The tasks are intended to probe whether a system can infer rules and solve novel puzzles, making them especially attractive to researchers and vendors looking for evidence of broader reasoning progress. That is also why disputes over setup become so sensitive. A benchmark built to isolate model performance becomes harder to interpret if one provider’s result depends on features outside the standard environment.
OpenAI’s argument, as described in the source text, is that benchmarks never measure only the model. They also measure the technical setup around it. That position reflects a broader industry reality. In deployed products, models rarely operate alone. They depend on orchestration layers, context windows, retrieval systems, tool use, and state management. If a benchmark ignores those components, vendors can reasonably argue that it understates real-world performance.
But ARC-AGI-3 was not designed to reflect every real-world deployment choice. It was designed to standardize the environment enough to make provider-to-provider comparisons meaningful. That is why the official benchmark harness matters so much to the conversation. When OpenAI posts a much stronger score outside that harness, it may be showing a more capable product stack, but it is not necessarily proving a stronger base model under the benchmark’s original rules.
The two settings that changed the outcome
The source text identifies two OpenAI features behind the jump. The first, Retained Reasoning, keeps the model’s chain of thought between steps. The second, Compaction, summarizes old context rather than truncating it. Together, those settings appear to address the weakness that held GPT-5.6 Sol back in the official harness, where intermediate reasoning was lost after each action.
That kind of persistence can be decisive on multi-step reasoning tasks. If a model can preserve earlier deductions and compress them into usable context, it may avoid repeatedly rebuilding its own internal state. In other words, the benchmark result changes not because the puzzle changed, but because the system is allowed to remember more effectively while solving it.

OpenAI’s claim therefore seems to rest on a broader interpretation of what the system being tested actually is. If the “system” includes the vendor’s API behavior and memory-management features, then the 38.3% score is relevant. If the “system” is supposed to be the model under a neutral standardized interface, the official 7.8% figure remains the more direct comparison point.
ARC Prize’s response leaves room for both views
The supplied source text says ARC Prize co-founder François Chollet responded by distinguishing between two classes of test setups. Custom harnesses built specifically to solve the benchmark or that contain knowledge about the benchmark format are off limits. General-purpose API settings available to all users, by contrast, are acceptable. Chollet also said ARC Prize had ongoing discussions with OpenAI about how best to test its models, especially around compaction, and welcomed the company “starting to figure out the answer.”
That response is notable because it does not reject OpenAI’s score outright. Instead, it suggests that ARC Prize sees some provider-specific settings as legitimate so long as they are general-purpose, broadly available, and transparently reported along with cost. At the same time, the source text notes that different providers using different settings creates a potential parity issue. In other words, the benchmark organizer appears willing to tolerate some asymmetry, even while acknowledging that it complicates direct comparisons.
This effectively shifts the argument from whether OpenAI’s result is valid to what kind of validity it has. It may be valid as a measurement of GPT-5.6 Sol used through OpenAI’s current API stack. It is less clearly valid as a clean apples-to-apples benchmark result against models tested under different assumptions.
What the episode says about AI benchmarking now
The dispute illustrates a broader transition in AI evaluation. As frontier systems become more agentic and more dependent on long-horizon workflows, benchmark performance will increasingly depend on runtime design choices, not just the underlying model weights. Memory retention, context compression, step orchestration, and interface design can all alter outcomes. That creates a problem for anyone trying to compare vendors through a single static leaderboard.
It also means benchmark consumers need to ask more precise questions. Was the model tested in a default configuration? Were vendor-specific features enabled? Did the harness preserve intermediate reasoning? Were costs and settings disclosed? A headline score without that context is becoming less informative than it used to be.
For OpenAI, the announcement shows the company is intent on contesting Anthropic’s recent benchmark momentum. For ARC Prize, the response suggests a pragmatic willingness to adapt the rules to modern API realities, provided the adaptations are not benchmark-specific hacks. And for the wider industry, the lesson is less about who is first on a particular day and more about how difficult it has become to define a “fair” test once model performance depends on product-layer intelligence.
That does not make the benchmark irrelevant. It makes disclosure more important. OpenAI’s 38.3% score is meaningful as evidence that GPT-5.6 Sol can perform much better when its reasoning is retained and its context is compacted instead of truncated. The official 7.8% score is meaningful as evidence that the same model fares far worse under the benchmark’s stricter standard harness. Both numbers describe something real. They just do not describe the same thing.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com







