John Schulman, the OpenAI co-founder who now serves as chief scientist at Thinking Machines, said he expects recursive self-improvement — AI systems accelerating their own development — in roughly three to four years, a timeline that places him among the more aggressive forecasters in the field. Schulman gave the estimate during a roundtable episode of the Dwarkesh Podcast published Friday, in which three frontier lab researchers openly debated how close the industry actually is to its ceiling.
The conversation offered a rare unguarded look at how researchers close to the frontier are thinking about progress, arriving amid a noisy week that saw Nvidia CEO Jensen Huang declare that "AGI has arrived" and a group of 25 Fields Medallists warn that AI benchmark-chasing is damaging mathematics. For readers tracking the latest AI developments, the episode is a useful counterweight: these are people building the systems, and they disagree with one another sharply.
Who was at the table
Host Dwarkesh Patel said he assembled the panel because the guests work at what he called "somewhat open-ish labs and companies," meaning they could speak on the record. Joining Schulman were Beren Millidge, chief technology officer at Zyphra, a startup developing open-source models, and Charlie O'Neill, head of model training at Baseten.
Schulman's credentials give his timeline talk particular weight. He co-founded OpenAI and led the reinforcement learning from human feedback research that produced the techniques behind ChatGPT, before departing for Thinking Machines, where he now serves as chief scientist.
The case against the singularity, steelmanned
Patel opened with a pointed question: if 2036 arrives without billions of superintelligent systems transforming the world, what will have been the technical reason? Millidge offered the strongest case for skepticism, describing what he called an almost Moravec's paradox-style pattern in which models ace whatever benchmarks researchers put in front of them, yet real-world impact disappoints.
"There's still some persistent sim-to-real gap which is somehow blocking everything," Millidge said, describing a world where AI becomes "extremely good at everything that people put into a benchmark" without generalizing further. He added that he considers this outcome unlikely, pointing to evidence that reinforcement learning already generalizes in practice.
Schulman largely agreed that progress is not automatic. Humans hold real advantages over today's models, he said, and each new model release catches up in some areas while remaining bottlenecked in others — the pattern that has defined the last several product cycles.
Rapid-fire timelines: two years, three to four, five to ten
When Patel pressed for numbers, the panel split. Asked when AI might deliver a tenfold productivity uplift for a competent white-collar worker, Schulman first refused to give a single figure, then said "I would say two years." Millidge said he could "kind of see that, actually," arguing the gain already exceeds 10x for some categories of work. O'Neill was the most conservative, answering "5 to 10" years.
Patel then pushed further, characterizing the scenario as "basically just ASI." Schulman's response: "I would say 3-4 years," adding that automating AI research itself is "not one of the hardest things for AI, because it involves a lot of" work that current techniques can already approach. The exchange was notable for how little pushback the estimate drew from the other researchers.
How automated researchers would actually be trained
A substantial portion of the episode dealt with the mechanics of the scenario Schulman was pricing in. Asked how AI systems capable of automating AI research would be trained, Schulman described "some combination of learning from human feedback to absorb the researchers' taste, and just creating a lot of practice environments which involve doing multi-step research."
That answer connects to where the industry's money is going. Google's recently completed deal for Mechanize, a startup that builds simulated working environments for training agents, reflects exactly this bet that better training grounds — not just bigger models — will drive the next round of capability gains. Schulman also flagged a core difficulty: reward functions built on natural data are hard to specify, and superficial signals like whether a user accepted a code edit can mislead training.
Distillation versus centralization
One of Schulman's most concrete predictions concerned industry structure. "I think distillation is the main thing that fights against the centralizing force," he said, arguing that capabilities learned through reinforcement learning can be transferred to smaller models relatively easily, spreading frontier-adjacent competence beyond the few companies that can afford frontier training runs.
On the question of what remains for humans, Schulman was categorical. "I would say that the last job for humans, or the role for humans that'll last the longest, is defining the objective and deciding what we actually want," he said. When Patel summarized "alignment is the final job," Schulman concurred: "Alignment is sort of the answer," while noting it decomposes into figuring out the right objective and then actually achieving it.
Where the episode fits in the timeline debate
The episode's subtitle — "We're nowhere near the ceiling" — captures its overall tenor: researchers who see substantial runway ahead, even as they dispute how bumpy the path will be. The discussion ranged across what is driving Chinese labs' progress, whether long-horizon reinforcement learning can elicit general intelligence, and why RL has started working as well as it has.
The public debate over timelines has intensified on all sides this month. Huang's "AGI has arrived" proclamation drew skepticism from analysts who read it as a business case rather than a scientific claim, while OpenAI CEO Sam Altman told staff the company is open to slowing cutting-edge development as safety concerns mount, according to Bloomberg. Schulman's three-to-four-year window for recursive self-improvement sits somewhere between marketing and moratorium — a working estimate from someone with a track record of building the techniques now driving the field.
Whether he is right will depend on questions the panel only circled: whether reinforcement learning keeps generalizing past its training environments, and whether the gap between benchmark performance and real economic value continues to close.
Stay Ahead of AI
Frontier researchers are debating AI's limits in public — follow the story as it develops.
Read more AI news on AI Buzz Wire