OpenAI's GPT-6 Astra may be struggling with user-trust problems on the coding front, but on a new robotics benchmark it is posting numbers that have researchers talking about a qualitative leap. On StationeryBench, a test built around ordinary desk objects, Astra fully completed 7 out of 100 tasks while Ai2's MolmoAct2, the model it was pitted against, completed zero.

The benchmark, described by The Decoder, pits the two models against each other across five desk-object tasks — uncapping a marker, pouring out paper clips, or passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials, a control design meant to isolate model capability from hardware differences. All results, videos and code have been published on GitHub. For more research coverage like this, check our AI news hub.

A 'Step Change,' According to Researchers

Yoav Artzi, an AI researcher at Cornell and Google DeepMind, called Astra's performance a "step change in spatial reasoning" — the kind of language benchmark results rarely earn. On a still-unpublished benchmark called REMAP, Astra reaches accuracy close to human level, though Artzi noted that "even Astra doesn't get to what humans do in other scenarios."

Astra's median progress score hit 46 out of 100, compared with 12 for MolmoAct2 — a gap that matters more than the headline completion numbers suggest. Robotics tasks are graded on partial progress as well as full success, and tripling a rival's median score indicates the model is consistently getting further through manipulation sequences rather than succeeding by luck on a handful of trials.

Why Spatial Reasoning Suddenly Matters

The result hints at how Astra was trained. Artzi suspects OpenAI trained the model on large amounts of 3D data, such as Blender scenes, which would line up with the model's particular strength on 3D tasks. If that reading is right, it suggests synthetic 3D environments — cheaper and safer than real-world data collection — are becoming a viable path to physical-world competence, one of the longest-standing bottlenecks in robotics.

The benchmark's design also says something about where robotics evaluation is heading. StationeryBench's tasks are deliberately mundane — uncapping a marker, pouring out paper clips, handing over a ruler — because everyday manipulation, not cinematic feats, is what separates lab demos from useful machines. Grading both models on the same dual-arm YAM hardware across 200 trials, and publishing every video, is an unusually transparent setup for a capability claim involving a frontier lab's flagship.

MolmoAct2, the comparison model from the Allen Institute for AI (Ai2), completing none of the tasks is its own kind of datum: it suggests the gap between general-purpose frontier models and purpose-built robotics models is no longer just closing but inverting, at least on reasoning-heavy manipulation.

It also fits with OpenAI's stated ambitions: the company has long-term plans to build its own consumer robots, according to reporting cited by The Decoder. A frontier language model that can reason about objects, hands and trajectories is the software half of that pitch. The hardware half remains unsolved, but benchmarks like StationeryBench are how that story gets measured.

Context: A Model With a Complicated Week

The result lands in a strange news cycle for OpenAI. The same model at the center of this benchmark has spent the week facing user complaints that it had grown dumber since launch — complaints OpenAI traced to three named defects, from misfiring skills to a misbehaving context experiment, before resetting Codex usage limits. A model that stumbles in production while posting step-change results in the lab is a précis of the current AI industry: capability and reliability are advancing on different clocks.

It also lands amid a safety debate that OpenAI's own leadership joined. The company's chief scientist has warned that no lab has solved control of increasingly autonomous systems, and Anthropic's CEO has called for the industry to slow frontier development altogether. Benchmarks showing models manipulating the physical world — even if the world is just a desk covered in stationery — are exactly the kind of result both executives cite when they argue the pace of progress is outrunning oversight.

The Bottom Line

Seven fully completed tasks out of 100 will not put a robot on your desk tomorrow. But going from a rival's zero to seven, with a median progress score three times higher, is the sort of discontinuity that redrew capability charts in language understanding two years ago. The open question is whether the same model can deliver that competence dependably, outside the benchmark, at a price developers are willing to pay.

Stay Ahead of AI

Benchmarks, breakthroughs and the race toward machines that act — followed every day.

Read more AI news →