GPT-6 Astra lands with strong numbers, conflicting scorecards, and a clearer efficiency story
OpenAI’s GPT-6 Astra is being received in unusually contradictory fashion: one benchmark aggregator places it clearly at the front of the model field, another says it is merely level with its predecessor overall, and a flagship reasoning test suggests a different kind of breakthrough altogether. The result is a revealing snapshot of the current AI race, where the answer to “which model is best?” increasingly depends on what is being measured and how much compute is required to get there.
According to The Decoder’s summary of newly published evaluations, Epoch AI combines more than 50 benchmarks into its ECI score and puts GPT-6 Astra in first place with 169 points, ahead of 267 models. Artificial Analysis, by contrast, gives Astra 61 points on its Intelligence Index, exactly matching Astra’s predecessor Sol and trailing Anthropic’s Claude Fable 5.1 at 66. The disagreement is not a small methodological footnote. It cuts to the heart of how the industry currently interprets frontier progress.
For readers trying to make sense of the split, the simplest explanation is that these benchmark suites emphasize different things. One broad composite can reward consistent gains across a wide range of tasks, while another may punish weaker performance more sharply in selected domains such as knowledge retrieval, coding, or long-context comprehension. The same model can therefore look like a decisive leader in one framework and only a lateral move in another.
Efficiency may be the more important signal
The more striking claim in the supplied report is not just about raw score placement. It is about efficiency. The Decoder says Astra uses only about a third of the compute steps used by Sol and a fifth of what Opus 5 uses. On coding tasks, Artificial Analysis reportedly found Astra scoring level with Claude Fable 5 while costing less than half as much per task. That kind of result matters because frontier competition is no longer only about reaching the best answer. It is about how economically a model can reach it.
That distinction becomes even sharper in the model’s pricing profile. OpenAI reportedly charges two and a half times as much per unit of processed text for Astra as for Sol, making a task roughly 75 percent more expensive than before. On its face, that looks like a substantial step up in cost. Yet The Decoder’s account suggests Astra’s thrift in compute changes the comparison depending on the competitor. Against its own predecessor, Astra appears costlier. Against rival models that consume much more computation, it can still look comparatively efficient for certain workloads.
In practical terms, this is the frontier market’s new tension. A model can be more expensive at the API level and still improve value if it completes work with fewer reasoning steps, lower token consumption, or better success rates on high-value tasks. For developers and enterprise buyers, those tradeoffs are likely to matter more than any single leaderboard position.
ARC-AGI-3 shifts the conversation
The report’s most headline-grabbing result comes from ARC-AGI-3, a benchmark built around unfamiliar game worlds. There, GPT-6 Astra reportedly reached 62.7 percent, far above Sol’s 7.8 percent and well ahead of Opus 5 at 30.2 percent. The article says Astra was, for the first time, more efficient than the average human on that test. François Chollet, ARC Prize chief, is described as calling the pace of progress about twice as fast as he had expected and moving his AGI forecast earlier.
Even with that framing, the broader lesson is less about one dramatic percentage and more about the kind of ability being rewarded. ARC-style evaluations are designed to probe abstraction and adaptation rather than simple memorization. Strong performance there carries symbolic weight because it is often treated as evidence that models are getting better at handling unfamiliar structures, not just reproducing patterns seen during training.
Still, the same supplied text makes clear that Astra is not uniformly stronger everywhere. Artificial Analysis reportedly shows the model losing about 80 Elo points on GDPval-AA v2 and slipping on banking support, SciCode, and long-context reasoning tasks. That uneven profile is important. It suggests Astra may be a more specialized or more aggressively optimized system rather than a flat improvement on every axis.
Cleaner outputs, uneven gains
One of Astra’s more practical improvements may be its reduced hallucination rate. On AA-Omniscience, The Decoder reports a drop from 92 percent to 51 percent. That is still a substantial hallucination rate, but the change is large enough to matter operationally in workflows where unsupported assertions create risk. For applied AI teams, reductions like that can be more valuable than marginal leaderboard movement, especially in research, coding, and business decision support where reliability gates adoption.
The report also points to rare performance on a difficult mathematics benchmark. Epoch AI says Astra was the only model to solve two of 68 open FrontierMath Erdős problems with Lean-verified proofs under a $300-per-attempt budget. Three additional solutions emerged in extra, non-standardized runs that used more than $220,000 in compute, but those did not count toward the score. That caveat is crucial. It shows both the ceiling of the model’s potential and the importance of cost-bounded evaluation when comparing systems fairly.
A model launch that exposes the limits of leaderboard thinking
The early Astra picture is therefore neither simple hype nor straightforward disappointment. It is a case study in why frontier AI evaluation has become harder to compress into a single ranking. One benchmark family says Astra is the leader. Another says it is merely tied with its predecessor overall. A reasoning benchmark suggests an unusually important leap in efficient problem-solving. Task-level measurements show gains in coding economy and lower hallucination, but also regressions in other areas.
That complexity is probably the real story. As models mature, the industry is moving past an era where one headline score can stand in for capability. GPT-6 Astra appears to be a model whose significance lies in selective advances, especially around efficiency and certain forms of abstract reasoning, rather than universal dominance. For buyers, researchers, and competitors, that makes it more interesting than a clean win would have been, because it hints at where the next phase of the AI race may actually be decided.
This article is based on reporting by The Decoder. Read the original article.
Originally published on the-decoder.com





