A new benchmark is challenging how the AI industry measures scientific capability, and its first results are humbling for the frontier labs. Terminal-Bench-Science 0.1, launched this week by researchers at Stanford University and the team behind the popular Terminal-Bench coding benchmark, evaluates AI agents on real scientific research workflows contributed by working scientists — and the strongest model tested, Claude Opus 5 running through Claude Code, completed just 30% of the tasks.
The benchmark's announcement page describes its guiding principle bluntly: scientists, not model developers or data vendors, should set the bar for scientific capability in AI. The project was built in collaboration with domain experts from research institutions around the world, and its first release includes 70 tasks spanning the life sciences, physics, Earth science, mathematics, and engineering.
How It Differs From Textbook Benchmarks
Most scientific AI evaluations test knowledge: multiple-choice questions, textbook problems, standardized exercises. Terminal-Bench-Science instead measures whether agents can execute the technically demanding, time-consuming workflows that fill actual researchers' days — the kind of work that appears in no exam. Tasks run in realistic environments, and grading is based on concrete artifacts: analyses, simulations, proofs, code, and data products, checked with reproducible task-specific tests rather than judge models or multiple-choice scoring.
That design reflects a theory of what AI assistants should eventually do for science. Rather than replacing scientific judgment, the benchmark's authors argue, agents should take over the execution layer — running the experiments-in-silico, processing the data, checking the proofs — to free scientists for the parts of research where human judgment matters most: defining questions, forming hypotheses, interpreting results, and communicating findings.
The project is also explicitly built as a moving target. Where many benchmarks are released once and abandoned — the announcement calls this the "papers to publish" problem — Terminal-Bench-Science is designed as a continuous benchmark that evolves alongside the frontier, with regular releases intended to keep the evaluation ahead of model progress and prevent contamination-driven score inflation.
The First Leaderboard: A Long Way From Reliable
The v0.1 leaderboard, which drew significant attention when it reached Hacker News this week, shows a steep drop-off from the top:
- Claude Opus 5 (Claude Code): 30.0%
- GPT-5.6 Sol (Codex): 22.4%
- Claude Fable 5 (Claude Code): 21.4%
- Claude Opus 4.8 (Claude Code): 10.5%
- GPT-5.6 Terra (Codex): 8.6%
- GLM 5.3 (Claude Code): 8.1%
- Kimi K3 (Claude Code): 7.1%
- Grok 4.6 (Grok Build): 7.1%
- GPT-5.6 Luna (Codex): 3.3%
The results suggest scientific workflows remain substantially harder for agents than software engineering, where the original Terminal-Bench and rival coding benchmarks have seen scores climb rapidly. A 30% resolution rate for the best available system means that on seven of ten real research tasks, even a top-tier agent paired with a strong harness fails to produce verifiably correct work.
The spread within model families is also instructive. The gap between Claude Opus 5 and Claude Opus 4.8 — nearly twenty points — and between the GPT-5.6 variants Sol, Terra, and Luna shows how much outcome quality depends on the interplay of model and agent harness, not raw model intelligence alone. Every high-scoring entry uses an agentic coding harness (Claude Code, Codex, Grok Build), reinforcing the emerging consensus that scaffolding is as decisive as the underlying weights.
Why Scientists Built Their Own Benchmark
The benchmark's framing carries a subtle institutional argument. "The stakes in science are too high," the announcement states, "and its benchmarks must reflect the scientific community's priorities rather than outside interests." The project gives working researchers a direct mechanism to encode the tasks they actually need automated — and to hold vendors' claims of "accelerating scientific discovery" against a verifiable standard.
That matters because the gap between marketing and reality has real consequences. If frontier labs claim their agents can accelerate research while independent, expert-designed evaluations show single-digit-to-30% resolution rates, funding decisions, hiring plans, and lab policies built on those claims inherit the error. Continuous, community-owned benchmarks are one corrective: they convert vague claims into reproducible numbers.
A Signal for Research Institutions
For universities, national labs, and R&D teams weighing when to deploy agents, the benchmark offers something vendor demos cannot: a community-maintained measurement of where the reliability line actually sits. A resolution rate in the teens or twenties means these systems are best treated today as assistants that require review, not as autonomous executors of experimental pipelines. The task categories — spanning life, physical, Earth, mathematical, and engineering sciences — also give domain specialists a template for contributing workflows from fields that generic benchmarks routinely overlook.
The broader industry context reinforces why the timing matters. Science has become one of the most heavily promoted use cases for frontier AI, with labs touting contributions to materials discovery, protein science, and mathematics. Those claims are difficult to audit: scientific output is slow to validate, and enthusiasm moves faster than replication. A benchmark that grades verifiable artifacts on real workflows gives research institutions a way to separate demonstrated capability from demonstration theater — and to track, release by release, whether agentic progress in science is compounding as quickly as it is in software engineering, where Terminal-Bench first made its name.
For the AI industry, the message of v0.1 is double-edged. The frontier is clearly advancing — Opus 5's 30% towers over the older Opus 4.8's 10.5% — but the ceiling remains far from the reliability real laboratories require. The benchmark's designers say regular releases will track exactly how fast that gap closes. Scientists watching the emerging AI trends in agentic research tools now have a yardstick they built themselves.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →