Anthropic published a prototype R&D Automation Index this week with a number that would have been science fiction in February: Claude now leads 26% of the measured work of building its own successor - completing work end to end from a high-level prompt under human supervision. In February the same index read effectively zero. The company also says the model collaborates on more than 90% of measured R&D work, while stressing that nowhere is it fully autonomous yet.
The index itself is the quiet innovation. Rather than a benchmark score, it measures how much of the actual model-development pipeline a Claude instance completes without human intervention - experiment design, runs, analysis, write-ups - and grades the model's role as lead, collaborator, or observer. Leading 26% means the model owned the work stream. Collaborating on more than 90% means it had a hand in nearly everything else.
The framing is deliberate. Anthropic pairs the number with an explicit caveat - Claude is not yet fully autonomous in any measured R&D area - and with the supervision infrastructure that makes the claim meaningful: automated code review checking proposed changes for bugs and security flaws, human owners on every stream, and by August a reported pool of roughly 30,000 concurrent agents doing research and engineering work inside the company.
The arc here is the one this paper has tracked since summer. OpenAI's automated research intern crossed 3.1 agent-workdays per human workday. Anthropic's own researchers resigned over the acceleration. Now the lab that asked the industry to pace the frontier is publishing measurements of the acceleration itself - which reads as either transparency or bracket creep, depending on how you weigh a number that went from zero to a quarter in seven months.
What the index does not answer is the question that matters: who signs off. A model leading a quarter of successor-development work under supervision is a management structure. A model leading that work where supervision becomes advisory is a different entity entirely. The index measures the slope. It does not tell you where the line is.
The Index Problem
What to watch is the denominator. An index that measures supervised lead work can rise honestly for quarters while the underlying arrangement - human owners on every stream - stays fixed. It can also rise because the supervision quietly thinned and nobody re-derived the categories. The index does not distinguish these, and Anthropic does not claim it does. The disclosure is the product: a lab publishing its own acceleration numbers while asking the industry to pace itself, and daring the rest of the field to publish theirs. The next index that matters is someone else's.
The Takeaways
- Anthropic's prototype R&D Automation Index says Claude leads 26% of measured successor-model work, up from effectively zero in February.
- The model collaborates on more than 90% of measured R&D - but Anthropic stresses it is not fully autonomous in any area yet.
- The lab that asked the frontier to pace itself is now publishing measurements of its own acceleration. Watch who publishes an index next.

