Google launched Gemini 3.8 Flash on Tuesday, its newest workhorse model positioned as the company's best reasoning and coding system yet at the same price as its predecessor, alongside a specialized variant called Gemini 3.8 Flash Cyber built for defensive cybersecurity work. The announcement, published on Google's official blog by Tulsee Doshi, senior director of product management, and Raluca Ada Popa, Gemini security lead at Google DeepMind, marks Google's third Flash release in only six weeks.
The launch lands in the middle of an unusually crowded release season, and it is a direct answer to the pricing-and-performance pressure rivals have put on Google's model line. For a wider view of how the frontier labs are trading blows, the breaking AI news team tracks every major model release as it happens.
Two Variants, One Foundation
Gemini 3.8 comes in two versions built on the same foundational intelligence. Gemini 3.8 Flash is the general-purpose workhorse: Google says it delivers significant improvements over 3.7 Flash across software engineering, agentic tasks and multi-step reasoning in specialized domains, at the same introductory price of $0.75 per million input tokens and $3.75 per million output tokens.
Gemini 3.8 Flash Cyber is described as Google's most capable cybersecurity model, with what the company calls frontier-level performance in vulnerability detection and automated patching. Unlike the standard Flash model, it is not broadly available — Google is distributing it to a set of trusted defenders through a new program called Fairwind.
According to the blog post, the coding and reasoning gains across the shared core were driven partly by rigorous training in the demanding domain of cybersecurity, and both models were accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models.
Benchmarks: Working Harder, Not Just Getting Bigger
Google's benchmark claims center on a design philosophy the company sums up plainly: 3.8 Flash "works harder." On complex tasks, the model executes extra reasoning steps and calls tools iteratively, sometimes consuming more tokens to maximize performance at higher effort levels. Developers can dial effort down to cut token overhead, or stay on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.
The headline results Google cites:
- DeepSWE v1.1 (long-horizon software engineering): 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, at what Google describes as a fraction of the cost.
- Vals Finance Agent V2 and Harvey's Legal Agent Benchmark: 3.8 Flash outperforms 3.7 Flash and other frontier models in quantitative and professional fields requiring advanced analysis and reporting.
- HLE-Verified: 3.8 Flash achieves 54.9%, which Google presents as evidence of multi-step reasoning across STEM, humanities and professional fields.
Flash Cyber: Defense Over Offense
The Cyber variant is the more unusual release. Google says it deliberately prioritized vulnerability fixing over offensive capabilities like exploitation. On CWE-Bench, an external patching benchmark run by Collinear, Gemini 3.8 Flash Cyber scored a pass@1 of 47.2% — just behind a leading frontier model at 47.8%, according to Google, but at a significantly lower cost, putting it on the benchmark's Pareto frontier.
On CyberGym, a standard industry benchmark for finding vulnerabilities, Google says the Cyber model demonstrates frontier-level autonomous vulnerability discovery, surpassing its predecessor 3.5 Flash Cyber as well as significantly larger frontier models. On Google's internal benchmark spanning complex codebases in 20 programming languages, the model reached a success rate exceeding 70%.
Already Securing Google's Own Code
Google says the Cyber model is already deployed internally, and the examples it provides are specific:
- The Chrome Security team found that 3.8 Flash Cyber produced 2.6 times more correct patches to Chrome vulnerabilities than the best commercial models, which are much larger.
- At security firm Wiz, the model achieved 7.5 to 9.7 percentage points higher recall on an internal penetration-testing benchmark at 2.3 to 5.2 times lower cost compared with other leading frontier models.
- Google's Cloud Vulnerability Research team used the model to find a critical foundational vulnerability in under two hours — a class of vulnerability whose research and discovery, Google notes, usually takes months.
What It Means for Developers and the Market
For developers, the practical takeaway is price stability at the low end of the frontier. Holding 3.8 Flash at 3.7's introductory pricing keeps the cost of a competitive coding-and-agentic model at $0.75/$3.75 per million tokens, a segment where open-weight challengers and rival labs have been undercutting incumbents all year. Google's message is that buyers no longer need to choose between cost and capability in the workhorse tier.
The Fairwind Program distribution model for Flash Cyber is equally telling. Front-tier labs are increasingly treating elite cybersecurity capability as something to be gated rather than shipped broadly — providing state-of-the-art defensive tooling to vetted defenders while limiting the potential for offensive misuse. Google's decision to prioritize patching benchmarks over exploitation capabilities in the model's training follows the same logic.
The cadence is notable too. Three Flash releases in six weeks — 3.7 Flash three weeks ago, now 3.8 — shows Google iterating its mid-tier line far faster than the annual-style release cycles of two years ago. The competitive pressure from OpenAI's GPT-6 rollout and a wave of fast-moving open-weight models from China has compressed everyone's timelines, and Google is clearly choosing speed at the workhorse tier while its frontier models carry the research flag.
Whether 3.8 Flash's "works harder" approach — spending more tokens on diligent multi-step reasoning — proves more efficient in practice than raw model scale will show up in developers' bills over the coming months. Google's own benchmarks suggest the trade can favor diligence; the market will render its verdict one API bill at a time.
Stay Ahead of AI
Every model launch, benchmark and price change, tracked as it happens — read more AI news →
