Anthropic published a detailed set of self-measurements on Saturday intended to let the public, journalists and governments track how fast frontier AI is actually being developed — including a first-of-its-kind index of how much of the company's own research is now performed by AI, and the disclosure that roughly 30,000 AI agents are doing research and engineering work inside the company at any given moment.

The post, titled "Measurements for understanding the pace of AI development inside frontier labs," arrives at the center of the most consequential argument in the industry: whether frontier development should be deliberately slowed. Dario Amodei has called for coordination to "pace the frontier," a proposal that has shaken markets, drawn a consumer lawsuit alleging pacing collusion among major labs, and prompted President Trump to announce an "AI Force" dedicated to preventing any slowdown. Anthropic's framing is explicit about that context: "As the world considers slowing the pace of frontier AI development, the public needs more information." For more on the story, see our ongoing AI industry coverage.

Three Measurements for Tracking AI Progress

The company laid out three measurement areas that it argues any frontier lab could publish regularly:

1. How much of AI R&D is performed by AI itself. Anthropic built a prototype "Anthropic R&D Automation Index" by cataloguing every kind of AI research and development work done at the company, rating how automated each task currently is, and aggregating the results. The ratings use Epoch AI's public "Automation Level" scale, which runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously with no human in the loop). On this scale, AL3 means AI "collaborates" — doing large chunks of work under close human direction — while AL4 means AI "leads," completing most of a task end-to-end from a high-level prompt while a human supervises.

The point of the index is to measure how close any lab is getting to recursive self-improvement — a model fully autonomously building its own successor. Anthropic notes that models accelerating AI development "could make it more challenging for humans to understand or control these systems."

2. How well AI agents' actions are overseen. This is where the 30,000 figure appears. As of August 2026, approximately 30,000 agents were doing research and engineering work at any one time inside Anthropic's most-used internal platform, delegating tasks to one another and working semi-autonomously for long stretches. The company measures three things about oversight of those agents: coverage (the share of agent actions that pass through a monitor before or after execution), review latency (how quickly flagged actions get reviewed, first by an automated monitor and then by a human), and escalation rate (the share of activities blocked, redirected, or flagged for review).

Anthropic's monitoring data so far suggests individual agents rarely misbehave — but the company warns that "when there are millions or billions of agents operating in the economy, even rare events can happen regularly."

3. How compute is allocated. The third measurement tracks how the lab's computing resources are distributed — the raw input that correlates with model capabilities. Together with the other two metrics, Anthropic argues this gives outsiders a way to correlate model inputs with model outputs, complementing the capability evaluations the company already publishes through its Responsible Scaling Policy risk reports.

Numbers Would Move Under a Slowdown

The most politically pointed line in the post is aimed directly at the pacing debate: Anthropic states that it "would expect these numbers to shift if there were coordination on pacing the frontier, as called for by Anthropic CEO Dario Amodei." In other words, if the industry genuinely slows down, these metrics are how the public could verify it — and if companies claim to be pacing while their numbers keep climbing, the measurements would expose that too.

Independent Verification Is the Missing Piece

Anthropic acknowledges the central weakness of self-reported metrics: it is using its own models to evaluate its own systems, which means the "judge" model could share the blind spots of the model being checked. The company also notes the lack of a common methodology across labs, making comparisons difficult.

Its proposed remedy is to embed independent third-party evaluators from multiple organizations inside Anthropic, giving them access to internal processes, systems and data comparable to what internal risk assessment teams see. Those third parties would verify safety practices, report incidents, and monitor the new metrics. The company says METR, the evaluation nonprofit, has previously independently red-teamed its offline monitoring platform, and that it plans to publish these measurements regularly in a form others can verify.

The measurements tie into Anthropic's broader Advanced AI Framework proposal, which lays out "rules of the road" for frontier model releases, including transparency obligations that governments could require — such as standardized risk reports.

A Test for the Whole Industry

The deeper significance of the post may be competitive. Anthropic is effectively challenging every frontier lab to publish the same numbers, arguing that "any frontier developer could publish these measures regularly, using a public methodology." If rivals decline, the contrast becomes its own data point in the safety debate; if they agree, the industry gains the first common yardstick for how fast AI is really advancing.

For now, the numbers are self-reported by a company with an obvious interest in the pacing conversation. But they represent something the debate has lacked: concrete, repeatable measurements that would show whether the frontier is accelerating, holding steady, or actually slowing down.

Stay Ahead of AI

Frontier lab disclosures like this one land weekly. For more breaking AI research and industry coverage, follow AI Buzz Wire.

Read the latest AI news →