Google DeepMind has turned computer use into a built-in tool inside Gemini 3.5 Flash, giving developers a native way to build AI agents that can interact with software across platforms without bolting on third-party frameworks.
The announcement, published on June 24, 2026 on Google's official Keyword blog, positions Gemini 3.5 Flash as a contender in the fast-moving market for agents that can see a screen, move a cursor, click, and type — capabilities that until recently required significant custom engineering. For more context on this story, see our ongoing breaking AI news.
"Computer use is now a built-in tool in Gemini 3.5 Flash to build agents that can interact across platforms," the company wrote. The post was authored by Mateo Quiros, a product manager at Google DeepMind.
How It Works
Computer use allows a model to operate a computer much like a human would: by interpreting what is on screen and taking actions such as clicking, scrolling, and entering text. By baking the capability directly into Gemini 3.5 Flash rather than offering it as a separate product, Google is lowering the barrier for developers who want to build agents that navigate real applications, websites, and operating systems.
According to the blog post, developers and enterprises can start using computer use in 3.5 Flash via the Gemini API and the Gemini Enterprise Agent Platform, Google's dedicated environment for building and deploying agents at scale.
Google demonstrated the capability with concrete use cases. In one example, Gemini 3.5 Flash used computer use to analyze the Gemini app itself and return a categorized list of features. In another, the model audited its own documentation for accessibility issues — a self-referential task that hints at how computer-using agents might one day help maintain the software ecosystems they run on.
Safety: Adversarial Training and a 'Defense-in-Depth' Approach
The most consequential part of the release may be what Google said about safety. Agents that operate in live environments are uniquely vulnerable to prompt injection — attacks in which hidden instructions embedded in a webpage, email, or document hijack the agent into taking unintended actions.
Google said it used targeted adversarial training for computer use in Gemini 3.5 Flash to mitigate those risks. It is also releasing two optional enterprise safeguard systems that the company said enable organizations to better control what their agents can do.
Taking a "defense-in-depth" approach, Google encouraged developers to combine these features with secure sandboxing, human-in-the-loop verification, and strict access controls. The company published additional guidance in its best practices documentation for the feature.
That framing matters. Computer use is widely seen as one of the more dangerous agent capabilities to deploy, because a compromised agent can take real-world actions inside a user's accounts and systems. By leading with adversarial training and layered safeguards, Google is trying to reassure enterprise buyers who have been hesitant to hand control of a cursor to an AI.
Why Computer Use Is Becoming a Battleground
Gemini 3.5 Flash's computer use arrives as the major AI labs race to make agents that can do useful work rather than just answer questions. Anthropic introduced a computer use tool for Claude in late 2024, and OpenAI's operator and agent products have pushed in the same direction. Making the capability native to a fast, cost-efficient model like Flash — rather than reserving it for a premium tier — signals that Google wants to commoditize agent control quickly.
The choice of Flash is itself notable. Flash is Google's efficiency-tier model, optimized for speed and cost. Embedding computer use there, rather than only in the larger Pro model, suggests Google expects high-volume, latency-sensitive agent workloads — the kind that automate repetitive tasks across business software at scale.
Google said it is already seeing customers drive value with computer use, though the blog post did not name specific commercial deployments or quantify the results.
The Competitive Stakes
For developers, the appeal of a built-in computer-use tool is straightforward: fewer integrations to maintain and a single vendor accountable for both the model and the agentic layer. But it also deepens dependence on Google's platform, a tradeoff that enterprises will weigh against the flexibility of model-agnostic agent frameworks.
The release also sharpens a broader tension in the agent economy. Independently developed open-source agent gateways, such as the Lightport framework that makes multiple LLM providers OpenAI-compatible, show how much demand exists for portable agent infrastructure. Google's bet is that a first-party, safety-hardened computer-use tool will win over developers who would otherwise stitch together third-party tools. As models gain the ability to act on screens, the line between an AI assistant and an autonomous operator keeps blurring — and so does the line between a productivity breakthrough and a security liability. Google's answer to that tension is layered safeguards and adversarial training. Whether that is enough to satisfy risk-averse enterprises will determine how quickly computer-using agents move from demos to daily use.
With Gemini 3.5 Flash, Google has made a clear bet: the future of AI is not just models that answer, but models that act — and it wants developers building those acting models on its platform.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →


