Anthropic reports Claude leads 26% of AI R&D in its new internal measurement, offering a rare numerical view of how a frontier laboratory uses artificial intelligence to build future models. The figure represents supervised, end-to-end task execution—not fully autonomous model development.

 

The September 17 disclosure also describes roughly 30,000 research and engineering agents operating at once on Anthropic’s most-used internal platform. Every action on that platform passes through an online monitor before execution and enters an offline monitoring system afterward.

 

Anthropic published three measurements intended to make frontier development easier to track:

  • Claude’s role across the laboratory’s AI research tasks
  • Monitoring coverage and escalation rates for internal agents
  • The share of AI R&D compute allocated to safety work

 

Related Research

 

How Claude Leads 26% of AI R&D at Anthropic

Anthropic created an internal R&D Automation Index using a six-level scale proposed by the research group Epoch AI. It ranges from AL0, where AI has no role, to AL5, where a system operates autonomously without a human in the loop.

 

Claude reached AL4, or “leads,” on 26% of the work measured in August 2026. At that level, an AI system can complete most of a task from a high-level prompt, but a person still supervises the work. More than 90% of measured tasks were at AL3, “collaborates,” or higher.

 

The boundary matters. Anthropic says Claude did not reach AL5 for any measured subset of research and development. The result therefore shows extensive automation inside a model laboratory, but it does not establish that Claude can independently design, train and validate its own successor.

 

The company argues that consistent reporting could help governments and outside researchers detect when model development is becoming substantially more automated. A shared metric could also support policy triggers, such as a fixed evaluation period before a new model is used for further AI R&D.

 

How Anthropic Built the R&D Automation Index

The index begins with a map of work inside Anthropic’s model-development loop. During each week of July, the company randomly sampled 20% of staff in participating departments and used work records to produce about 15,000 granular tasks.

 

Claude organized those tasks into a hierarchy containing 542 nodes and 378 leaf categories. A separate Claude-based judge assigned automation levels, while person-time estimates determined how heavily each type of work counted in the overall result.

 

Anthropic compared the judge’s ratings with assessments from staff responsible for the relevant work. Exact model-to-human agreement was 59%, while human reviewers agreed exactly with one another 35% of the time. Model and human scores fell within one level of each other in 97% of comparisons.

 

Those checks do not remove the index’s limitations. Anthropic is using its own models to evaluate automation inside its own organization, and reasonable reviewers can disagree about where collaboration ends and leadership begins. The company says independent evaluators should be able to inspect comparable internal processes and data.

 

30,000 Agents Operate Under Two Monitoring Layers

Anthropic’s most-used internal agent platform had approximately 30,000 agents performing research and engineering work at any one time in August. The figure covers that platform rather than every AI-assisted activity across the company.

 

Online monitors review actions before execution and can block or redirect behavior that could cause immediate, irreversible harm. Offline monitors analyze activity afterward for slower-developing patterns, including possible signs of misalignment.

 

Anthropic analyzed more than one billion agent decisions during August. Its online system blocked 0.002%, or about one in 47,000 decisions. The offline system flags roughly one or two transcripts per thousand for additional filtering, with about 50 high-priority cases reaching human review each week.

 

A low blocking rate is not a direct safety score. It could reflect reliable agents, narrow permissions, effective preventive controls or monitors that miss problematic behavior. Coverage, latency and escalation statistics become more informative when paired with independent tests of what the monitoring systems actually catch.

 

Safety Compute Metrics Need Independent Verification

Anthropic also examined compute use during the week of July 13 to July 20. About 6% of the compute assigned to AI R&D was classified as safety work, rising to 12% within AI-driven AI R&D.

 

The company describes those estimates as conservative. Work that advanced capabilities as much as safety was counted as general R&D, while separate safeguard classifiers were excluded. Anthropic also notes that compute is an imperfect proxy because many safety projects require more researcher time than accelerator time.

 

The disclosure’s longer-term value depends on repetition and comparability. A frozen task basket can show how the same work becomes more automated, but it may not capture new kinds of work created for people as older tasks shift to machines.

 

Regular publication by several laboratories, using common definitions and outside verification, would turn Anthropic’s snapshot into an industry benchmark. Until then, the 26% figure is best read as a significant internal measurement of supervised automation—and an early warning that oversight methods must scale alongside the agents building the next generation of models.