Anthropic says its Claude models can now design protein binders from scratch as well as — or better than — leading human experts, a result the company validated in wet labs and published in a research post on Monday that marks one of the clearest demonstrations yet of AI accelerating real scientific discovery.

The findings come from two experiments. In the first, Anthropic tested Claude's ability to design protein binders — small proteins engineered to latch tightly onto target proteins — a task that historically takes a human specialist weeks or months per target and that sits at the heart of modern drug design. For more context on this story, see our ongoing latest AI developments.

Beating the Benchmarks

Working with external evaluators Adaptyv Bio and Twist Bioscience, which independently produced and tested Claude's designs in the lab, Anthropic ran a multi-arm design campaign against 15 protein targets. Claude succeeded against 14 of them, producing 354 confirmed binders from 1,320 designs. That alone represents a significant addition to the public corpus of de novo protein designs: the two largest existing collections, Anthropic noted, together contain roughly 770 binders out of 5,700 designs against 40 targets.

The hit rates are what caught biologists' attention. Anthropic's Mythos Preview model achieved an overall hit rate of 26.7% when designing against all targets simultaneously in a 48-hour session, with Claude Opus 4.8 close behind at 22.6%. Typical hit rates in protein design campaigns today run 10 to 15%, the company noted. When Mythos Preview focused on one target at a time in multiple 24-hour sessions, its hit rate climbed to 35.1%. The campaign drew its targets from benchmarks that are studied extensively — including all of Adaptyv Bio's BenchBB suite — plus two novel targets, 15-PGDH and GDF-8, chosen specifically to ensure Claude could not lean on pre-recorded successes from its training data. For each of the 15 targets, researchers asked Claude to design 30 protein binders.

On individual competition targets, the margins were wider still. Against RBX1, a protein involved in targeted protein degradation, Mythos Preview in single-target mode achieved a 40% hit rate — compared with 3.7% among human participants in the same competition — and its top-ranked design outperformed the winning human design on binding affinity.

There were instructive failures. A de novo-designed beta-barrel protein called BBF-14 and maltose-binding protein (MBP) both remained out of reach, targets whose novelty and smooth, flexible surfaces leave a binder little to grab onto. And in a surprise, it was Opus 4.8 — not the more powerful Mythos Preview — that succeeded on TNFα, a notoriously difficult target whose blockade underlies some of the best-selling drugs ever made, including Humira.

An Autonomous Researcher

Perhaps most striking is how little human involvement the campaign required. After an initial prompt, Anthropic says it provided no additional scientific, technical or operational guidance. Claude chose where on each protein to design against, generated candidate structures and sequences by orchestrating publicly available specialist protein design and co-folding models, and ran the campaign end to end — work that can take a human operator weeks.

The effort consumed up to 12,500 NVIDIA H100 hours of compute for the multi-target run, and up to 2,500 H100 hours per target in single-target mode. Claude even produced 15 confirmed binders containing significant beta-sheet structure, a fold that is harder to design than the alpha-helix bundles most computational tools default to.

From Designing Molecules to Reading Raw Data

The second experiment tested a different, more mundane bottleneck: analytical chemistry. Every time a chemist makes a molecule, they must confirm what it is and how pure it is, using nuclear magnetic resonance (NMR) spectroscopy and liquid chromatography–mass spectrometry (LC-MS). Interpreting those files is painstaking manual work.

Supplied with only a contract lab's raw instrument files and a short plain-language prompt — no vendor software, no operator — Claude Opus 5 returned fully processed NMR and LC-MS results in 23 and 19 minutes respectively, working in parallel. Its results matched the lab's own analysis. Along the way, the model reverse-engineered an undocumented vendor file format, verified its reading against the instrument's recorded totals for all 2,664 scans, and proposed the exact follow-up experiment — a heavy-water check — that the contract lab had independently run three days later.

Anthropic says the run also demonstrated genuine scientific judgment, not just automation, and notes that a chemist typically needs half an hour to an hour per sample for this kind of analysis.

Safety First

The company was explicit about the dual-use implications. The same autonomous research capabilities that could speed human therapies could, without robust safeguards, enable dangerous research such as bioweapons development, Anthropic warned. Life-science research tasks are currently blocked in its most capable model, and the company says it is preparing a trusted access program for scientists — one of its highest priorities. In the meantime, Claude Opus 5 remains its most capable generally available model.

Anthropic has shared the prompts used for the campaigns along with all in vitro and in silico data generated — a transparency measure that will let outside labs verify the claims. For a field still sorting hype from results, the post offers something rare: lab-validated evidence that a general-purpose AI model can do, in hours, work that once defined a specialist's career.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →