Just one day after OpenAI made its most powerful model generally available, the company says the system has produced a complete proof of the Cycle Double Cover Conjecture — a celebrated open problem in graph theory that has resisted mathematicians for roughly half a century.

OpenAI researcher Ethan Knight announced the result on July 10, 2026, posting that GPT-5.6 Sol Ultra "produced a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in just under one hour." The claim quickly climbed to the top of Hacker News, drawing hundreds of comments from mathematicians and engineers debating whether the proof holds up.

For anyone tracking the latest AI developments, the result is a striking demonstration of where frontier models are heading — and a reminder that their mathematical claims still need to be checked the old-fashioned way.

What Is the Cycle Double Cover Conjecture?

The conjecture sits in a branch of mathematics called graph theory. In plain terms, it asks a deceptively simple question: can the edges of any "bridgeless" graph — one with no edge whose removal would split it apart — always be covered by a collection of cycles so that every single edge appears exactly twice?

The problem was posed independently by several mathematicians over the years, including Paul Seymour in 1979 and George Szekeres in 1973, with earlier roots in work by W. T. Tutte and by Alon Itai and Michael Rodeh. Despite its simple statement, it became one of the best-known unsolved problems in the field, and partial results were known only for special classes of graphs.

According to the proof note OpenAI published, the argument reduces the problem to loopless cubic graphs, then leans on a classical result — the nowhere-zero flow theorem over the group F32 (equivalent to Tutte's 8-flow theorem) — and a key linear-algebra step that converts an edge labeling into the structure needed for a cycle double cover. The writeup is only a few pages long and, notably, uses "no mathematics developed within the last 30 years," as one Hacker News commenter observed.

How GPT-5.6 Sol Ultra Approached It

What sets this result apart from ordinary chatbot answers is how the model was deployed. GPT-5.6 Sol Ultra ran in OpenAI's "multiagent v2" mode, spinning up a system of up to 64 cooperating subagents that explored different proof strategies in parallel. OpenAI released the full prompt alongside the proof, offering a rare look at the instructions given to the system.

The prompt instructed the model to "use multiagent v2 aggressively and dynamically," to maintain a "diverse portfolio of approaches," and to deploy adversarial agents to audit any candidate proof for common pitfalls — such as closed trails masquerading as cycles or accidental circular reasoning. It even told the system to "spend at least 8 hours on this before even thinking of returning or giving up," though the final run reportedly completed in under an hour.

In the proof note's "Statement of AI use" section, OpenAI is explicit: "The proof in this note is entirely due to GPT 5.6 Sol Ultra and the writeup with Codex (with GPT 5.6 Sol)."

Not Everyone Is Convinced Yet

The mathematical community has reacted with a mix of excitement and caution — and that skepticism is well placed. As of publication, the proof has not been formally verified in a proof assistant such as Lean, nor has it passed traditional peer review.

On Hacker News, several commenters noted that short, elegant proofs of famous conjectures deserve especially careful scrutiny. "It's a very short proof that uses no mathematics developed within the last 30 years," wrote one user. "Which doesn't necessarily make it wrong, but in the absence of mechanization in Lean or proper peer review, I think it is premature to post this."

Others pointed out that graph theory's leading formal proof libraries are not yet mature enough to check research-level results of this kind, meaning verification may need to come from human experts reading the argument line by line. OpenAI itself framed the announcement as part of a broader pattern: "More test-time compute leads to greater intelligence," the company said, signaling that the multiagent approach — not this single result — is the real story it wants to tell.

Why It Matters

Even if the Cycle Double Cover proof ultimately needs correction, the episode signals a meaningful shift in what large language models can do in pure mathematics. Frontier AI systems have previously produced novel research results in specialized domains — including other recent claims of solved problems in geometry and combinatorics — but attacking a named, decades-old conjecture with a coordinated swarm of reasoning agents raises the bar considerably.

For researchers, the takeaway is twofold. First, the multiagent recipe — many independent attempts, cross-pollination of ideas, and adversarial auditing — appears to be a genuinely useful template for hard problem-solving. Second, the verification bottleneck is now squarely on humans: the model can generate candidate proofs faster than the community can check them.

That tension is likely to define the next phase of AI-assisted mathematics. As one commenter put it, if this proof holds up, "someone is about to start a list" of other long-standing problems to throw at the next generation of models.

For now, mathematicians will be reading the three-page note carefully — pencil in hand — while the rest of the AI industry watches to see whether GPT-5.6's hour-long effort becomes part of the permanent mathematical record.

Stay Ahead of AI

For more on frontier model breakthroughs and AI research coverage, visit the AI Buzz Wire homepage.

Read more AI news →