OpenAI has fixed three defects that it says explain a week of user complaints that GPT-6 Astra, its newest frontier model, had grown dumber since launch — and it is resetting usage limits for Codex subscribers as an apology of sorts. The update came from Codex product lead Tibo Sottiaux, who said the team worked with users to trace the faults, then shipped fixes and a full usage reset before midnight local time, according to an update published Saturday.

The admission closes a rough stretch for what OpenAI has billed as its most capable model to date. Since Astra reached wide availability in early September, developers on X had been running identical prompts against launch-day Astra and the current version, keeping every setting the same, and posting side-by-side comparisons that appeared to show a quietly weakened model. For more on the models driving this debate, see our latest AI developments coverage.

The Three Defects, Named

According to Sottiaux, the first culprit was skills — prompt files written for earlier models that were triggering too often under Astra, sometimes stopping the model from checking its own work. The second was an optional context management experiment that could make the model quit early or reply to stale messages. OpenAI disabled the experiment after roughly 4,000 to 5,000 users were affected, the company said.

The third defect involved infrastructure rather than the model itself: some inference engines had been configured incorrectly and measurably degraded quality on the tail traffic routed through them. OpenAI removed those engines and made several smaller repairs. Sottiaux wrote that the model now tracks a user's latest message more accurately — a nod to complaints that Astra had been answering questions the user had already moved past.

Developers Had Already Moved On

The complaints built through the week. Developer Pranjal Paliwal, who had praised Astra days earlier, went back and read the code it wrote for him, then called the outcome a regression rather than progress. Dax Raad, who builds the coding tool Opencode, moved his team back to the older GPT-5.6 Sol after spending on Astra doubled for downsides he judged were not worth the money. One forum poster described Astra as now level with Sol on identical tasks, with the same retries.

Not everyone accepted that reading. A pseudonymous user writing as Antikythera argued the model had always been lazy and fond of bullet points, and that users had simply stopped being dazzled. That disagreement captures how hard model-quality claims are to settle: identical prompts, different verdicts.

Capacity Trouble in the Background

The quality row landed while capacity has shadowed the rollout. OpenAI has cut Astra usage limits for heavy ChatGPT users according to user reports, and paused new $200 Pro subscriptions — a step Sottiaux confirmed on September 10, citing demand. The model still costs $10 per million input tokens and $50 per million output tokens, which is 2.5 times what Sol charged when it launched.

There is also an uncomfortable precedent: this is the second OpenAI flagship in two months to draw the same accusation from paying users. In July, users said Sol's top reasoning mode had gone shallow overnight. Sottiaux denied weakening it deliberately at the time, while confirming the company had been experimenting with reasoning effort. Repeated cycles of "the model got worse" followed by "we found defects" are becoming a pattern the company will have to manage — transparency helps, but so would stability.

OpenAI's Own Advice: Use Fewer Guardrails

Alongside the fixes, OpenAI published migration guidance from engineer Eric Provencher that amounts to a counterintuitive recommendation: more capable models need less hand-holding. Overly long skill descriptions, blanket reading requirements and rigid approval rules can get in Astra's way, the guidance says. Instructions that have piled up over time can eat context or cause the model to stop work too early.

Provencher recommends that teams review their skills, AGENTS.md files and task prompts whenever they switch models, tie instructions tightly to specific tasks, and define clearly when a job is done — because even without restrictions, Astra may stop earlier than GPT-5.6 Sol did. Teams that locked things down after earlier models went rogue should revisit those rules: Astra has better judgment, the post argues, but it can interpret old restrictions so literally that it stops when you want it to keep going.

What Happens Next

The usage reset gives Codex subscribers a clean slate to re-evaluate the model, and OpenAI will be watching the same side-by-side tests its users run. The deeper question is trust: paying developers who saw a flagship degrade within a week of launch now have a documented explanation, three named defects, and a reset — but also a fresh reason to benchmark every model update against the day it shipped.

Stay Ahead of AI

Model launches, defect post-mortems, and the tools reshaping software work — tracked every day.

Read more AI news →