OpenAI released GPT-6.1 Sol on September 29, retiring GPT-6 Sol just seven days after that model's debut, according to independent benchmarking firm Artificial Analysis. The swap coincided with the company's DevDay developer conference, where CNET reported that attendees could start using GPT-6.1 Sol immediately alongside the newly launched Dots agents and a set of subscription changes.

A seven-day flagship-to-flagship replacement is unusual even by the accelerated standards of 2026, and the benchmarks suggest why OpenAI moved fast. For continuous coverage of model launches, benchmark fights, and the pricing war underneath them, AI Buzz Wire tracks the developments that matter as they happen.

A measurable jump in seven days

Artificial Analysis reports that GPT-6.1 Sol gains 4 points on the firm's Intelligence Index over GPT-6 Sol, and 5 points over GPT-5.6 Sol, landing just 1 point below GPT-6 Astra, OpenAI's most capable model. The firm's evaluation blends reasoning, knowledge work, math, and agentic tasks into a single comparative score.

The sub-scores show where the gains concentrate. GPT-6.1 Sol improved 4 points in AA-Briefcase v1.1 and 5 points in GDPval-AA v2.1, both of which test agentic knowledge work. A 12-point jump in Terminal-Bench 4.0 points to materially better performance in real terminal environments, while Humanity's Last Exam rose 5 points and GDP.pdf rose 6.

One number stands out for anyone who uses models for research: accuracy on AA-Omniscience climbed 8 points while the hallucination rate fell from 60 percent to 54 percent. Both figures remain high in absolute terms, but the direction of travel matters for workflows where a confident wrong answer is worse than no answer.

Same sticker price, cheaper in practice

On pricing, GPT-6.1 Sol keeps GPT-6 Sol's rate card of $2 per million input tokens and $10 per million output tokens, but the cache read discount rises from 90 percent to 95 percent. Artificial Analysis notes that GPT-6.1 Sol's blended price for agentic workloads is therefore slightly lower than GPT-6 Sol's, extending a streak of effective price cuts that began when GPT-6 Sol launched at a 50 percent discount to GPT-5.6 Sol.

The pricing trajectory is the real story here. Within the span of two releases, OpenAI has roughly halved the effective cost of its mid-tier frontier line twice, while nudging quality upward each time.

Cost per task: $0.72 versus $3.26

Cost per task is where the gap with GPT-6 Astra becomes stark. At maximum effort, Artificial Analysis pegs GPT-6.1 Sol at $0.72 per Intelligence Index task, less than a quarter of GPT-6 Astra's $3.26. Against its own predecessor the savings are 31 percent ($1.05 for GPT-6 Sol), and against GPT-5.6 Sol they reach 64 percent ($1.99).

According to the firm, every effort level of GPT-6.1 Sol pushes out the cost-efficiency Pareto frontier — meaning that for a given level of intelligence, no cheaper option exists among the models it tracks. The token-efficiency picture is more mixed: GPT-6.1 Sol consumes roughly 10 to 30 percent more output tokens than GPT-6 Sol across effort levels, though its low and medium effort settings remain Pareto-optimal for token efficiency once the higher intelligence is factored in. In coding, the model gained 3 points over GPT-6 Sol at max effort on the firm's Coding Agent Index.

Why the seven-day turnaround matters

DevDay's wider slate reinforces the same picture. CNET's report framed the conference around three things developers could use right away — GPT-6.1 Sol, the Dots agents, and reworked subscription plans — rather than a distant research preview. The message to developers was availability today, not possibility tomorrow.

Flagship cadences this short signal that the competitive constraint has shifted from training runs to post-training refinement and serving infrastructure. A point-level upgrade that ships in a week looks more like an optimized checkpoint and a new price sheet than a from-scratch foundation model, and rivals are behaving the same way: Google announced Gemini 4 Argon on September 30, a frontier model that is initially reaching trusted cyber defenders through its Fairwind Program at the same $2/$10 per million token rate.

डेवलपर्स के लिए, सत्यापित संख्याओं से तीन व्यावहारिक निष्कर्ष निकलते हैं। सबसे पहले, भारी कैश पुन: उपयोग वाले एजेंटिक वर्कलोड को 95 प्रतिशत कैश छूट से तुरंत लाभ मिलता है। दूसरा, बजट को माइग्रेट करने से पहले उच्च आउटपुट टोकन खपत का मॉडल बनाना चाहिए, क्योंकि प्रति-टोकन कीमत अपरिवर्तित रहती है, भले ही मॉडल लंबे समय तक सोचता हो। तीसरा, एस्ट्रा उन कार्यों के लिए उच्चतम सीमा बनी हुई है जहां बुद्धिमत्ता का अंतिम बिंदु 4x लागत प्रीमियम के लायक है - जीपीटी-6.1 सोल का मामला यह है कि अधिकांश कार्यों के लिए, यह नहीं है।

दोहराने लायक चेतावनी: ये आंकड़े एक एकल स्वतंत्र मूल्यांकनकर्ता से आते हैं, भले ही इसकी कार्यप्रणाली कितनी भी कठोर क्यों न हो, और सूचकांक स्कोर विशिष्ट कार्यभार मिश्रणों को पुरस्कृत करते हैं। अत्यधिक कार्यभार वाली टीमों को उत्पादन ट्रैफ़िक को पुनः रूट करने से पहले अपने स्वयं के मूल्यांकन को फिर से चलाना चाहिए।

---

एआई से आगे रहें

नवीनतम एआई समाचार, विश्लेषण और सफलताएँ प्राप्त करें - सभी एक ही स्थान पर।

अधिक AI समाचार पढ़ें →