Google has not yet finished rolling out Gemini 4 Argon to the public, but according to internal documents reviewed by Business Insider, the company is already testing the next thing: a Gemini 4 variant code-named Carbon that is showing early signs of matching Anthropic's strongest coding model.

The reporting, published Friday and detailed by The Decoder, describes documents, screenshots and internal chats in which Google employees discuss three Gemini 4 variants under development: Argon, Barium and Carbon. For readers tracking the frontier model race, this is the latest sign that the gap between model generations is collapsing into a rolling stream of updates, and our AI model news coverage follows every step of it.

Carbon Is Already Running Inside Google

According to Business Insider, Carbon was deployed in recent days on Jetski, Google's internal coding platform, and early users say it is primarily stronger than Argon at programming tasks. One employee compared Carbon with Claude Opus 5.5, Anthropic's strongest coding model, while cautioning that the Google model still needs further testing before that comparison means much.

The same documents offer a rare look at how Google names and stages its frontier releases. Argon, the company's current frontier model, previously carried the codename "Barium-B," and an employee described Carbon internally as the "Gemini pro next model." That suggests Carbon and Barium are checkpoints or candidate updates within the Argon frontier family rather than separate product tiers, though the documents do not make clear whether Carbon will eventually ship as an Argon upgrade or as its own release.

Not everyone inside Google is convinced the new model is a clean leap forward. Early versions of Argon reminded another employee of the older Opus 5 on some coding tasks, even as the overall internal reception of the model was reported as positive.

Why Argon Is Not a Flash Model

When Argon was announced on September 30, speculation circulated that it might be a speed-optimized Flash-class model rather than a maximum-performance release. Google has since foreclosed that theory. In the announcement of its Gemini agent for Google Workspace, the company laid out its model lineup by purpose: Argon for frontier reasoning, Flash for speed and volume, Omni for generative media, and Gemma for lightweight open-weights work.

Argon is therefore Google's strongest reasoning model, positioned by the company alongside Anthropic's Opus line and OpenAI's Astra. It is currently available only to a limited group of trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens, and cached input discounted 95 percent.

'Recursive Self-Improvement' Jokes Write Themselves

The pace of internal iteration has not gone unnoticed. Vedant Misra, a Google DeepMind researcher, responded to the Business Insider report on X with a pointed quip: "Have you heard of recursive self improvement?" — a reference to the practice of using AI systems to help build better AI systems. OpenAI and Anthropic have both reported similar dynamics in their own pipelines, where successive model generations increasingly assist in producing their successors.

Whether or not Carbon is a product of that loop, the codename churn is telling. A fast succession of internal checkpoints, each tested against frontier rivals within days of deployment, suggests Google has industrialized its release process in a way that compresses the traditional six-to-twelve-month model cycle.

Google Is Already Preparing the Launch Machinery

There is still no official date for a broader Gemini 4 launch, but Google's own products are being rearranged to make room. In the Gemini app, users and testers have spotted a new "Automatic" mode for the 3-series models alongside an adjustable reasoning-intensity setting that ranges from low to high. In Google AI Studio, a new "Ultra" mode has appeared, promising access to "advanced skills and tools."

Google's Logan Kilpatrick has confirmed publicly that the team is working to get the maximum out of Argon, which is consistent with the pattern the internal documents describe: ship the frontier model to a small professional audience, keep iterating against competitor benchmarks, and expand access as guardrails mature.

What to Watch Next

Three questions will define the next few weeks. First, whether Carbon's early coding performance holds up in wider internal testing or fades the way promising checkpoints sometimes do. Second, whether Google expands Argon beyond cyber defenders to developers and consumers — the company has said only that broader availability will come "as soon as possible." Third, how Anthropic and OpenAI respond: Opus 5.5 is the reference point Google's own employees are using for coding comparisons, and a competitive Gemini update would restart a pricing and capability contest that has already compressed coding-model costs this year.

For now, Carbon remains an internal codename attached to screenshots and chat logs, not an announced product. But the direction of travel is clear. Google is no longer competing by launching models; it is competing by continuously replacing them.

Stay Ahead of AI

Frontier model leaks, launches and benchmarks land daily — we sort signal from noise.

Read more AI news →