IA2 MIN

Claude now leads 26% of Anthropic's measured R&D work, but it still cannot build models alone

The company has published its first measurement of AI's role in developing the next generation of models. The result matters, but it is not an autonomous intelligence factory.

Anthropic chart showing Claude's role in its research and development work
Image: Anthropic
01

The 26% figure needs context

Anthropic has tried to quantify a question that frontier labs usually answer with anecdotes: how much does today's model contribute to its successor? In the company's first internal development report, Claude led 26% of a sample of research and development tasks in August 2026. It made a substantial contribution or took the larger role in at least 90% of them.

Sample is the important word. Anthropic studied work inside its own teams and used assessments from employees and models to classify the contribution. It is not claiming that Claude performed 26% of every R&D activity across the company. None of the measured tasks reached Anthropic's highest category, where AI completes the work without human involvement.

02

Thirty thousand internal agents, but no full autonomy

The operating scale is striking. Anthropic says its most widely used internal platform can run roughly 30,000 agents at once. An online monitor checks actions before execution and blocked about 0.002% of them, or one in 47,000. The highest-priority cases, roughly fifty each week, also receive human review.

An agent is not an independent digital employee. It is a model invocation connected to tools, instructions and boundaries. It may write code, investigate a bug or prepare an experiment, but the result still depends on its context and the person validating it. Our artificial-intelligence dictionary separates models, agents, tools and complete systems.

Anthropic also identifies a measurement weakness. Some classifications come from the model and the employees using it. That makes thousands of observations possible, but it is not an external audit and a large contribution is not automatically a correct one.

Anthropic visual identity used in its research report
Image: Anthropic
03

Safety is competing for resources too

Anthropic examined one week of compute allocation to see where effort was going. It estimates that safety accounted for 6% of total R&D compute and 12% of AI-driven research compute. Those are internal figures rather than an industry standard, but few labs publish a comparable reference at all.

The report does not show Claude improving itself alone. It shows that models already sit inside a growing share of the cycle, writing, testing and reviewing technical work. Faster iteration can follow, along with a greater need for logs, permissions and oversight. Our analysis of GPT-5.6 and delegated work explains why completing more tasks does not prove that the right decisions were made.

Anthropic plans to repeat the measurement. A trend will be more useful than one headline number, showing whether AI moves from collaboration into a leading role more often and whether any measured task eventually reaches full autonomy.

Cover of Anthropic's study on the pace of AI development
Image: Anthropic
00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

YOUR NEXT ROUTE

Keep following AI models and agents

If this story interests you, these three pieces are the best place to carry on.

OPEN THE FULL TOPIC
  1. 01Claude merges chat and Cowork, bringing documents into the conversationIA · 3 MIN
  2. 02Google Home opens up to other AI agents, but setup takes more than a chatIA · 3 MIN
  3. 03OpenAI opens Codex's machinery with an Agents API built for long-running workIA · 2 MIN

KEEP READING

You may also like

FRONT PAGE