OpenAI's GPT-6 Astra has posted state-of-the-art results on ARC-AGI-3, the abstract reasoning benchmark designed to measure agentic intelligence, according to results published Wednesday by the ARC Prize Foundation. The model scored 62.7% on the benchmark's Semi-Private set under the foundation's Standard harness, and 99.9% when evaluated with a Provider Adapter harness that preserves the model's internal reasoning state between requests.
The results landed one day after OpenAI began the staged rollout of Astra, the flagship model the company has framed as the start of a new era. For readers trying to keep score as the frontier labs trade benchmark blows, the latest AI developments page tracks every major release as it happens.
Two Harnesses, Two Very Different Scores
The gap between Astra's two headline numbers is the most important detail in the report. ARC Prize evaluates agents under different harnesses, and each changes what the model is allowed to do with its own context.
Under the Standard harness, a model can carry forward notes it chooses to keep as it works through an environment. With that setup, Astra at maximum reasoning effort scored 62.7%, at a total cost of $26,098 for the evaluation run. At the opposite end, running the same harness with reasoning turned off produced just 35.2% while costing more — $49,791 — because the model needed many more actions to make progress.
The Provider Adapter harness is more generous: it preserves the model's opaque reasoning state between requests and uses compaction for longer conversations, letting Astra reuse prior work. Under that harness, Astra scored 99.9% at high reasoning effort for $18,817, and 98.6% at maximum effort for $17,332. Even the lowest reasoning setting reached 96.7%.
In other words, the difference between a partial score and a near-perfect one is not raw capability alone — it is how much of the model's working state survives between steps.
Beating Humans on Action Efficiency
ARC-AGI-3 is built around novel, abstract, turn-based game environments. Agents must explore without instructions, infer how the world works, identify goals from sparse rewards, and plan multi-step actions. The environments are calibrated so that humans can solve 100% of them, which makes human performance the baseline the benchmark measures against.
Astra did not just solve most levels — it solved them efficiently. According to ARC Prize, the model used fewer actions than the median tested human on 96% of levels, surpassing the human baseline in action efficiency across the set.
The economics of the comparison are stark. Human test participants were paid $115 per 90-minute session plus $5 per completed game, working out to roughly $12.78 per attempted game. A full Astra evaluation run costs five figures, depending on configuration. But ARC Prize noted a counterintuitive wrinkle: higher reasoning effort generally cost less, because Astra solved games in fewer actions, reducing the total number of model calls and tokens burned per run.
A Model That Builds Its Own Language
Beyond the scores, ARC Prize highlighted a qualitative behavior the team observed. Astra consistently turned unfamiliar environments into compact symbolic world models — representing game mechanics as logical rules, and developing its own domain-specific shorthand to track state and plan actions.
That is precisely the skill ARC-AGI-3 was designed to probe. Where earlier generations of the benchmark tested fluid reasoning over novel patterns, the third generation tests whether an agent can explore, model, set goals and execute plans in environments it has never seen — the components ARC Prize lists as exploration, modeling, goal-setting, and planning and execution.
The AGI-Era Backdrop
The benchmark results arrived amid maximal rhetoric from OpenAI itself. At a press briefing before Thursday's launch, company president Greg Brockman said it was "not unreasonable to feel that we are now in the AGI era," while stopping short of a formal declaration, according to The New Stack's coverage of the launch. He described AGI as a "mission concept or spiritual concept" rather than a contractual trigger.
OpenAI researcher Aidan Clark said Astra was the company's largest training run to date — the first pre-trained on more than 100,000 GPUs at the Stargate site in Texas — and the first OpenAI release in which earlier models played a significant role in supervising the training process. Access remains staged: the rollout begins with enterprise customers in the Daybreak program, with Plus, Pro, Business and Enterprise availability plus API and AWS access promised in the coming days.
The Asterisk on the Score
Caution is warranted on several fronts. Astra's ARC-AGI-3 performance depends heavily on harness configuration, as the two headline numbers show, and custom harnesses that preserve model state have been a recurring point of contention in benchmark disputes. The New Stack's analysis of the launch noted that Astra "posts big gains on specialized tasks, but it comes at a premium and does not clearly lead the coding pack" — a reminder that agentic puzzle benchmarks and everyday commercial workloads measure different things.
ARC Prize itself frames the exercise as measuring the "residual gap" between current AI and AGI, defining AGI as the ability to acquire any skill a human can, as efficiently as a human can. A near-perfect score under favorable conditions does not close that gap by itself. And at $17,000 to $50,000 per evaluation run, only well-resourced labs can afford to run the benchmark at this scale — though ARC Prize's cost data shows smarter reasoning can actually shrink the bill.
Why It Matters
Action efficiency is a genuinely new axis of comparison. A model that solves every level but wastes thousands of calls doing it is less useful — and more expensive — than one that acts like Astra did here: exploring deliberately, compressing what it learns, and acting on it. Whatever one concludes about the AGI debate, the ARC Prize data shows frontier agents are now matching humans not just on whether they can solve novel problems, but on how economically they solve them.
Stay Ahead of AI
Every benchmark result, model launch and safety debate, tracked as it happens — read more AI news →
