Merit AC
2026-09-17

Anthropic publishes the numbers behind "pace": Claude now leads 26% of its own R&D

Three days after Dario Amodei called for the industry to slow down, Anthropic backed the ask with real internal metrics -- automation share, agent-oversight coverage, and safety compute allocation -- and is daring competitors to publish the same.

Anthropic published three new transparency metrics on September 17, arguing that if the industry is going to debate "pacing the frontier" -- the idea, which CEO Dario Amodei publicly floated days earlier, that AI labs should deliberately slow capability gains -- the public needs actual numbers to have that debate with, not just competing public statements. "As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows," the company wrote, and it backed the point by disclosing its own internal figures rather than only proposing a framework for others.

What the three metrics actually show

The headline number: Claude now leads roughly 26% of Anthropic's own AI research and development work, up from under 1% just six months earlier, measured on a 0-5 automation scale where "leads" (AL4) sits above "collaborates" (AL3) -- over 90% of the company's R&D work now clears the "collaborates" bar or higher. Anthropic is careful to note Claude isn't "operating fully autonomously" for any measured subset of that work. The second metric covers oversight: roughly 30,000 agents run research and engineering work on Anthropic's internal platform at any given time, with both real-time and after-the-fact monitor coverage, and a block rate on flagged decisions of about 0.002% -- roughly 1 in 47,000. The third is compute allocation: in a July 13-20 snapshot, about 6% of AI R&D compute went to safety work specifically.

Why a number beats a pledge

Anthropic is explicitly asking OpenAI, Google DeepMind, and every other frontier lab to publish the same three categories on a recurring basis, with independent third-party evaluators eventually verifying them -- turning "we take safety seriously" from a talking point into something a customer, regulator, or competitor can actually check quarter over quarter. That distinction matters directly for anyone trying to answer whether AI spend is producing real work or slop: a self-reported automation-share number is still self-reported, but it's a far better starting point than no number at all, and it sets a bar the next lab's own transparency push will now be measured against.

Sources

← All news