Security researchers have demonstrated that the sandboxes protecting four widely used AI coding agents — Cursor, OpenAI's Codex, Google's Gemini CLI, and Antigravity — can be escaped without ever attacking the sandbox head-on. The findings, published on July 20, 2026, by Pillar Security's research team and reported by BleepingComputer, expose a structural weakness in how AI coding tools isolate the code their agents generate from the developer machines they run on.
The research is a significant data point for anyone tracking breaking AI news on agent safety, because it shows that even a perfectly compliant agent — one that obeys every rule inside its sandbox — can still break out. The flaw is not in the agent's behavior but in the trust boundary the sandbox assumes.
How the Escapes Work
The key insight is deceptively simple. Modern AI coding agents run inside a sandbox that draws a line: the agent is trusted inside the project workspace, and the host outside is protected. The assumption is that files inside the workspace are inert — data, not commands.
But they are not inert. Tools running outside the sandbox constantly read and act on those files. Integrated development environments resolve Python interpreters, Git integrations scan repositories, VS Code runs task files, hook engines fire commands, and Docker Desktop exposes a local socket. A sandboxed agent can obey every rule it is given and still write a file that one of those external tools later executes, loads, or scans.
According to BleepingComputer's reporting, the escape "happens on its own": the agent stays inside the box, follows every rule, and just writes a file that a trusted tool outside the box subsequently runs. The agent never breaks out; the breakout is done on its behalf by software the developer already trusts.
The Trigger: Prompt Injection
The mechanism that sets these escapes in motion is prompt injection — the same vulnerability that has plagued AI agents across domains. A malicious instruction planted in a README file, a GitHub issue, a project dependency, or a code diff becomes a local action on the developer's machine once the agent processes it.
This connects the sandbox research to a broader pattern in AI security. The agent does not need to be compromised or jailbroken. It simply needs to encounter a poisoned input in the course of its normal work — reading a file, reviewing a pull request, installing a package — and then carry out the embedded instruction by writing the right file in the right place. The sandbox permits the write, because writing files is exactly what a coding agent is supposed to do.
The 'Week of Sandbox Escapes'
Pillar Security's research team — Eilon Cohen, Dan Lisichkin, and Ariel Fogel — reproduced the bypasses over several months and published them as a series they call the "Week of Sandbox Escapes," releasing one write-up per day. The researchers sorted their seven findings into four distinct failure modes.
One category is what they describe as denylist sandboxes — sandboxes that try to block specific dangerous actions rather than permitting only safe ones. Denylists are notoriously fragile because they depend on anticipating every possible attack, and the escapes show how an agent can route around a blocked action by enlisting an external tool that was never on the denylist in the first place.
The four affected tools — Cursor, OpenAI's Codex, Google's Gemini CLI, and Antigravity — represent a broad cross-section of the AI coding agent market, from consumer IDE plugins to enterprise command-line tools. That breadth suggests the problem is not a bug in any single product but a shared architectural assumption that the research invalidates.
Why This Matters for the Agent Economy
The implications extend beyond individual developers. As coding agents are embedded into automated pipelines, continuous integration systems, and autonomous workflows, a sandbox escape becomes a potential foothold for supply-chain attacks. An attacker who can get a poisoned file into a repository — through a dependency, a cloned repo, or a compromised contributor — can conceivably turn a trusted coding agent into an execution vector on a developer's machine.
នេះគឺជាប្រភេទហានិភ័យដូចគ្នាដែលបានកើតឡើងនៅដើមឆ្នាំ 2026 នៅពេលដែលអ្នកស្រាវជ្រាវបានរកឃើញថាការបើកឃ្លាំងដែលមានគំនិតអាក្រក់នៅក្នុង Cursor អាចដំណើរការកូដដោយស្ងៀមស្ងាត់នៅលើ Windows ។ ប្រអប់ខ្សាច់គេចចេញពីការគំរាមកំហែងនោះជាទូទៅ៖ វាមិនមែនជាឧបករណ៍មួយ ឬវេទិកាមួយនោះទេ ប៉ុន្តែគំរូអន្តរកម្មរវាងភ្នាក់ងារប្រអប់ខ្សាច់ និងឧបករណ៍ដែលអាចទុកចិត្តបានដែលនៅជុំវិញពួកគេ។
បញ្ហាពិបាកនៃឯកសារដែលអាចទុកចិត្តបាន។
ការលំបាកជាមូលដ្ឋានគឺថាប្រអប់ខ្សាច់របស់ភ្នាក់ងារសរសេរកូដមិនអាចចាត់ទុកឯកសារកន្លែងធ្វើការទាំងអស់ថាមិនគួរឱ្យទុកចិត្តបាន ដោយមិនធ្វើឱ្យខូចប្រយោជន៍របស់ភ្នាក់ងារនោះទេ។ ភ្នាក់ងារត្រូវការសរសេរកូដ ការកំណត់រចនាសម្ព័ន្ធ និងស្គ្រីប ហើយឯកសារទាំងនោះត្រូវអាន និងធ្វើសកម្មភាពដោយឧបករណ៍របស់អ្នកអភិវឌ្ឍន៍។ ដកការជឿទុកចិត្តនោះចេញ ហើយភ្នាក់ងារមិនអាចដំណើរការបានទេ។ រក្សាវា ហើយវ៉ិចទ័ររត់គេចនៅសល់។
ការស្រាវជ្រាវរបស់ Pillar មិនផ្តល់នូវការជួសជុលការធ្លាក់ចុះតែមួយទេ ហើយនោះជាផ្នែកនៃមូលហេតុដែលការរកឃើញមានសារៈសំខាន់។ ពួកគេបង្កើតបញ្ហាការរចនា ដែលឧស្សាហកម្មនឹងត្រូវដោះស្រាយជាសមូហភាព — តាមរយៈភាពឯកោកាន់តែខ្លាំងរវាងកន្លែងធ្វើការរបស់ភ្នាក់ងារ និងផ្ទៃប្រតិបត្តិរបស់ម្ចាស់ផ្ទះ តាមរយៈការចុះហត្ថលេខា ឬបញ្ជាក់ឯកសារដែលសរសេរដោយភ្នាក់ងារ ឬតាមរយៈការគិតឡើងវិញថាតើឧបករណ៍ណាខ្លះត្រូវបានអនុញ្ញាតឱ្យដំណើរការមាតិកាកន្លែងធ្វើការដោយស្វ័យប្រវត្តិទាំងអស់។
សម្រាប់អ្នកអភិវឌ្ឍន៍ដែលប្រើ Cursor, Codex, Gemini CLI ឬ Antigravity ថ្ងៃនេះ ការដកចេញជាក់ស្តែងគឺមានការប្រុងប្រយ័ត្នជាមួយនឹងការបញ្ចូលដែលមិនគួរឱ្យទុកចិត្ត៖ ឃ្លាំងក្លូន ភាពអាស្រ័យភាគីទីបី និងកូដដែលបានរួមចំណែកគួរត្រូវបានចាត់ទុកថាជាសក្តានុពលនៃការចាក់បញ្ចូលភ្លាមៗដែលកំណត់គោលដៅមិនត្រឹមតែគំរូប៉ុណ្ណោះទេ ប៉ុន្តែប្រព័ន្ធឯកសារជុំវិញវា។
នាំមុខ AI
សម្រាប់ការគ្របដណ្តប់បន្តនៃសន្តិសុខ AI ភ្នាក់ងារសរសេរកូដ និងហានិភ័យហេដ្ឋារចនាសម្ព័ន្ធដែលពួកគេណែនាំ សូមអនុវត្តតាម ការគ្របដណ្តប់ឧស្សាហកម្ម AI របស់យើង។
អានព័ត៌មាន AI បន្ថែម

