OpenAI's GPT-6 Astra has added an unusual trophy to its record: a full playthrough of Valve's iconic first-person puzzler Portal, completed autonomously in under 24 hours at a total compute cost of $571.18. The Verge reported the milestone on September 6, 2026, crediting a user known as cozyblaze, who wired the frontier model up to the game and let it figure out the rest.

Portal is a demanding test for an AI agent despite its age. Released by Valve in 2007, the game requires players to reason about spatial puzzles using a portal gun — creating linked openings in walls, floors, and ceilings to move through test chambers that routinely defy intuition. There is no combat grind to lean on and no quest marker to follow; every chamber is essentially a physics riddle. For readers following how frontier models are being stress-tested, this run is a vivid entry in our AI developments coverage.

From Demo Videos to a Full Completion

The run didn't come out of nowhere. Earlier on September 6, a YouTube video showing GPT-6 Astra playing Portal circulated on Hacker News, and a follow-up post noting that the model had "autonomously completed" the game picked up attention as the day went on. The Verge's report confirmed the details: the full game, beaten in under 24 hours, for less than $600 in compute — with a condensed video of the run published for skeptics.

The Verge couldn't resist quantifying the environmental side of the bill, noting the run probably consumed roughly a small swimming pool's worth of water in data center cooling. It's a wry reminder that agentic gaming runs are cheap for the user but not free for the planet.

Why a 2007 Puzzler Is a Serious Benchmark

Beating Portal is not a scheduled benchmark like ARC-AGI or SWE-bench, and that's partly why it resonates. The game demands exactly the skills that have historically separated humans from AI systems:

  • Spatial reasoning. Solutions require predicting where a portal exit will launch the player's body — three-dimensional physics reasoning that has tripped up agents for years.
  • Long-horizon planning. Chambers chain multiple steps; failing the final step usually means redoing the sequence from the start.
  • Learning from failure without hand-holding. There are no hints. The agent must infer the game's rules through trial and error, then generalize them to new chambers.

Earlier this year, OpenAI's GPT-6 Astra topped ARC-AGI-3 benchmarks focused on agentic reasoning, and the model has drawn attention for computer-use capabilities — the ability to operate software interfaces the way a person would. The Portal run is a playful but genuine demonstration of that same skill set: perception, planning, and control, unrolled for hours without human intervention.

Part of a Broader Wave

The run is also a sign of where AI gaming has landed in 2026. Benchmark scores now share space with cultural milestones: models composing Bach chorales, beating handwritten-correction recognition tests, and completing games that once stood in for the phrase "computers will never do this." Each pass relieves an old certitude and raises the question of what the next one will be.

There are practical implications, too. The same loop that let cozyblaze's setup steer Portal — screenshots in, decisions out, at a cost of a few hundred dollars — is the loop being productized for software testing, robotics, and computer-use assistants. A model that can hold a game's physics in mind for a full playthrough is one that can, in principle, grind through long, messy software tasks that defeat shorter-context systems.

For now, though, the takeaway is simpler. An AI beat one of gaming's most beloved puzzles, start to finish, for about the price of a new GPU's deposit — and posted the video. The test chambers of Aperture Science were designed to make humans feel clever. This weekend, they made a model look fluent.

How the Run Was Received

The response tracked the arc of AI gaming milestones generally: disbelief first, then the video, then acceptance. On Hacker News, the run drew a steady stream of commenters dissecting how the setup worked and what it implied about agentic AI — with several noting that the total bill was lower than many assumed such a run would cost. The condensed video of the completion gave skeptics a way to verify the claim themselves, which is more than most benchmark results offer.

It also invited comparison with the milestones that came before. Gaming has served for decades as the field's public proving ground, from board games to real-time strategy to open-ended sandboxes. Each time a game falls, the argument shifts: today's skeptics insist the feat was narrow, specialized, or specialized in some way not generalizable. What makes the Portal run notable within that history is how ordinary the setup was — a commercially available frontier model, a consumer game, and a few hundred dollars of compute, rather than a bespoke research system trained specifically for the game.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →