Anthropic’s pace scorecard shows Claude already leading a quarter of its AI research
Anthropic published three measurements of how fast it is automating its own research, how it oversees internal agents, and how it allocates research compute to safety, according to a new Anthropic Institute post. The company framed the numbers as a public scorecard for frontier labs and said it expects them to shift if the industry coordinates on pacing.
As of August 2026, Claude is not operating fully autonomously on any measured slice of Anthropic’s AI R&D. Claude “leads” 26% of that work. More than 90% of the work sits at or above “AI collaborates,” using Epoch AI’s Automation Level scale from AL0 to AL5.
“AI leads” (AL4) means the model can finish most of a task from a high-level prompt while a human supervises. “AI collaborates” (AL3) means large chunks of work under close human direction.
Recursive self-improvement would mean a model fully building its successor with little or no human in the loop. Anthropic says Claude has not reached that for the measured R&D subsets.
On the company’s most-used internal agent platform, about 30,000 agents were doing research or engineering at once as of August 2026. Online monitors catch 100% of actions before execution.
Of more than a billion decisions analyzed that month, 0.002% were blocked, about 1 in 47,000. Offline, 100% of transcripts are ingested after the fact. About 100,000 transcripts a week get flagged, and the highest-priority about 50 a week escalate to human review.
A July 13-20 compute snapshot put about 6% of compute that went to AI R&D toward safety, and about 12% of compute that went to AI-driven AI R&D toward safety. Anthropic calls those figures deliberately conservative. Dual-purpose tokens, compute that serves both product and safety work, were not counted as safety, and safeguard classifiers, filters that block risky outputs, were excluded.
The company says it plans to embed independent third-party evaluators with access comparable to internal risk teams. The snapshot is a company self-report until that happens.
The numbers still say the lab is accelerating. Claude already leads a quarter of the research, tens of thousands of agents are running, and the safety-compute slice stays thin.
Sources

Pingback: Claude’s scorecard, a stolen-labor memo, and the referee that never got hired