Frontier AI models differ wildly in how hard they are to break, according to a new benchmark from safety nonprofit FAR.AI — and the gap is far larger than many assumed.

On July 29, FAR.AI launched its AI Security Leaderboard, testing the safeguards of four leading models under identical conditions. Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol held firm against the attacks, while xAI's Grok 4.5 and Google's Gemini 3.1 Pro each broke for under $300, the organization said. FAR.AI described the difference between the most and least robust models as a "hundredfold gap" in safeguards. For more context on this story, see our ongoing breaking AI news.

Alongside the leaderboard, FAR.AI introduced what it calls the Minimal Standard for Safeguards, a shared security baseline assembled entirely from defenses that peer developers already run in production. The goal is to give labs a concrete floor to meet — and to make weak protections visible.

The findings echo FAR.AI's earlier stress tests. A May report found that DeepSeek-V4-Pro's safeguards had "collapsed almost completely," with low-skill attackers bypassing safety mechanisms 98–100% of the time across every domain tested. A publicly available jailbreak built for the model's predecessor worked without a single modification, suggesting known vulnerabilities went unpatched.

WIRED, reporting on the leaderboard, concluded that it remains "frighteningly easy to jailbreak some frontier AI models." As agents take on more sensitive tasks, the disparity raises pressing questions about which systems are safe to deploy.

For continuous coverage of AI safety and security research, readers can follow AI Buzz Wire.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →