OpenAI's GPT-6 Astra has posted state-of-the-art results on ARC-AGI-3, the abstract reasoning benchmark designed to measure agentic intelligence, according to results published Wednesday by the ARC Prize Foundation. The model scored 62.7% on the benchmark's Semi-Private set under the foundation's Standard harness, and 99.9% when evaluated with a Provider Adapter harness that preserves the model's internal reasoning state between requests.
The results landed one day after OpenAI began the staged rollout of Astra, the flagship model the company has framed as the start of a new era. For readers trying to keep score as the frontier labs trade benchmark blows, the latest AI developments page tracks every major release as it happens.
Two Harnesses, Two Very Different Scores
The gap between Astra's two headline numbers is the most important detail in the report. ARC Prize evaluates agents under different harnesses, and each changes what the model is allowed to do with its own context.
Under the Standard harness, a model can carry forward notes it chooses to keep as it works through an environment. With that setup, Astra at maximum reasoning effort scored 62.7%, at a total cost of $26,098 for the evaluation run. At the opposite end, running the same harness with reasoning turned off produced just 35.2% while costing more — $49,791 — because the model needed many more actions to make progress.
The Provider Adapter harness is more generous: it preserves the model's opaque reasoning state between requests and uses compaction for longer conversations, letting Astra reuse prior work. Under that harness, Astra scored 99.9% at high reasoning effort for $18,817, and 98.6% at maximum effort for $17,332. Even the lowest reasoning setting reached 96.7%.
In other words, the difference between a partial score and a near-perfect one is not raw capability alone — it is how much of the model's working state survives between steps.
Beating Humans on Action Efficiency
ARC-AGI-3 is built around novel, abstract, turn-based game environments. Agents must explore without instructions, infer how the world works, identify goals from sparse rewards, and plan multi-step actions. The environments are calibrated so that humans can solve 100% of them, which makes human performance the baseline the benchmark measures against.
Astra did not just solve most levels — it solved them efficiently. According to ARC Prize, the model used fewer actions than the median tested human on 96% of levels, surpassing the human baseline in action efficiency across the set.
The economics of the comparison are stark. Human test participants were paid $115 per 90-minute session plus $5 per completed game, working out to roughly $12.78 per attempted game. A full Astra evaluation run costs five figures, depending on configuration. But ARC Prize noted a counterintuitive wrinkle: higher reasoning effort generally cost less, because Astra solved games in fewer actions, reducing the total number of model calls and tokens burned per run.
A Model That Builds Its Own Language
Beyond the scores, ARC Prize highlighted a qualitative behavior the team observed. Astra consistently turned unfamiliar environments into compact symbolic world models — representing game mechanics as logical rules, and developing its own domain-specific shorthand to track state and plan actions.
That is precisely the skill ARC-AGI-3 was designed to probe. Where earlier generations of the benchmark tested fluid reasoning over novel patterns, the third generation tests whether an agent can explore, model, set goals and execute plans in environments it has never seen — the components ARC Prize lists as exploration, modeling, goal-setting, and planning and execution.
The AGI-Era Backdrop
The benchmark results arrived amid maximal rhetoric from OpenAI itself. At a press briefing before Thursday's launch, company president Greg Brockman said it was "not unreasonable to feel that we are now in the AGI era," while stopping short of a formal declaration, according to The New Stack's coverage of the launch. He described AGI as a "mission concept or spiritual concept" rather than a contractual trigger.
Ο ερευνητής του OpenAI Aidan Clark είπε ότι το Astra ήταν η μεγαλύτερη εκπαίδευση της εταιρείας μέχρι σήμερα - η πρώτη προεκπαιδευμένη σε περισσότερες από 100.000 GPU στην τοποθεσία Stargate στο Τέξας - και η πρώτη έκδοση OpenAI στην οποία προηγούμενα μοντέλα έπαιξαν σημαντικό ρόλο στην επίβλεψη της εκπαιδευτικής διαδικασίας. Η πρόσβαση παραμένει σταδιακή: η διάθεση ξεκινά με εταιρικούς πελάτες στο πρόγραμμα Daybreak, με τη διαθεσιμότητα Plus, Pro, Business και Enterprise, καθώς και πρόσβαση σε API και AWS να υπόσχονται τις επόμενες ημέρες.
Ο αστερίσκος στο σκορ
Απαιτείται προσοχή σε πολλά μέτωπα. Η απόδοση του ARC-AGI-3 του Astra εξαρτάται σε μεγάλο βαθμό από τη διαμόρφωση της πλεξούδας, όπως δείχνουν οι δύο αριθμοί επικεφαλίδας, και οι προσαρμοσμένες ιμάντες που διατηρούν την κατάσταση του μοντέλου αποτελούν επαναλαμβανόμενο σημείο διαμάχης σε διαφωνίες συγκριτικής αξιολόγησης. Η ανάλυση του New Stack για το λανσάρισμα σημείωσε ότι το Astra "σημειώνει μεγάλα κέρδη σε εξειδικευμένες εργασίες, αλλά έρχεται σε υψηλότερη τιμή και δεν οδηγεί ξεκάθαρα το πακέτο κωδικοποίησης" - μια υπενθύμιση ότι τα πρακτικά σημεία αναφοράς παζλ και ο καθημερινός εμπορικός φόρτος εργασίας μετρούν διαφορετικά πράγματα.
Το ίδιο το βραβείο ARC πλαισιώνει την άσκηση ως μέτρηση του "υπολειπόμενου χάσματος" μεταξύ της τρέχουσας τεχνητής νοημοσύνης και του AGI, ορίζοντας το AGI ως την ικανότητα απόκτησης όποιων δεξιοτήτων μπορεί ένας άνθρωπος, τόσο αποτελεσματικά όσο μπορεί ένας άνθρωπος. Ένα σχεδόν τέλειο σκορ υπό ευνοϊκές συνθήκες δεν κλείνει αυτό το χάσμα από μόνο του. Και με 17.000 έως 50.000 $ ανά δοκιμή αξιολόγησης, μόνο τα καλά εξοπλισμένα εργαστήρια μπορούν να αντέξουν οικονομικά να εκτελέσουν το σημείο αναφοράς σε αυτήν την κλίμακα — αν και τα δεδομένα κόστους του ARC Prize δείχνουν ότι η πιο έξυπνη συλλογιστική μπορεί στην πραγματικότητα να συρρικνώσει τον λογαριασμό.
Γιατί έχει σημασία
Η αποτελεσματικότητα της δράσης είναι ένας πραγματικά νέος άξονας σύγκρισης. Ένα μοντέλο που επιλύει κάθε επίπεδο, αλλά σπαταλά χιλιάδες κλήσεις κάνοντας το είναι λιγότερο χρήσιμο - και πιο ακριβό - από ένα μοντέλο που λειτουργεί όπως το Astra εδώ: να εξερευνά σκόπιμα, να συμπιέζει όσα μαθαίνει και να ενεργεί σύμφωνα με αυτό. Ό,τι και αν συμπεράνει κανείς σχετικά με τη συζήτηση AGI, τα δεδομένα του ARC Prize δείχνουν ότι οι συνοριακοί πράκτορες ταιριάζουν τώρα με τους ανθρώπους όχι μόνο για το αν μπορούν να λύσουν νέα προβλήματα, αλλά και για το πόσο οικονομικά τα λύνουν.
Μείνετε μπροστά από την τεχνητή νοημοσύνη
Κάθε αποτέλεσμα αναφοράς, λανσάρισμα μοντέλου και συζήτηση για την ασφάλεια, παρακολουθείται όπως συμβαίνει — διαβάστε περισσότερα νέα AI →
