Claude Now Leads 26% of Anthropic's AI Research: What the Numbers Show
Last updated: September 2026
On the evening of September 17, 2026, Anthropic published measurements showing that its model Claude now "leads" 26% of the company's AI research and development work as of August 2026 — up from under 1% in February — and that more than 90% of that work now happens with AI collaborating at or above the "AI collaborates" level. The figures come from an official post on the Anthropic Institute titled "Measurements for understanding the pace of AI development inside frontier labs," the first publicly released dataset of its kind from inside a frontier AI lab.

Source: Anthropic
What exactly was announced?
"AI building AI" is no longer a conference talking point. Anthropic released three sets of internal measurements documenting how Claude has become a genuine partner in building the next generation of models — not just a code-completion tool. The post was co-authored by researchers Marina Favaro and Phillie Wright, with research direction from Jack Clark, one of the company's co-founders. Within hours, the story was picked up by Reuters, Bloomberg, and the Associated Press under headlines like "Claude is helping to build the next version of itself."
The nuance most quick headlines missed: Anthropic did not claim Claude works fully autonomously. The classification uses Epoch AI's automation-level scale, and Claude reached level four (AL4) — "leads" the work — on a quarter of measured tasks, while no measured work reached level five (AL5), full autonomy. The gap between those two numbers is the gap between a real partnership and a self-running lab, and it matters to anyone tracking the future of research and engineering jobs.
The numbers in detail: three measurements, not one
| Measurement | Reported value | What it means practically |
|---|---|---|
| Claude-"led" share of Anthropic's AI R&D | 26% (AL4) | Claude leads a quarter of research and development work, up from under 1% in February 2026 |
| Work at "AI collaborates" level or above | Over 90% | The overwhelming majority of research tasks now involve models as working partners |
| Concurrent internal agents | ~30,000 agents | Running simultaneously on the most-used internal platform |
| Blocked decisions out of 1 billion+ | 0.002% (1 in 47,000) | Pre-execution monitors intercept risky actions before they run |
| Research compute directed to safety | ~6% | Versus ~12% of compute going to AI-driven AI R&D (week of July 13–20) |

Source: Official Anthropic Institute post
The agent oversight numbers deserve a closer look. Online monitors cover 100% of agent actions before execution, "usually within seconds," and offline monitors review everything afterward — flagging roughly 100,000 transcripts a week and escalating only about 50 cases a week to humans, who review blocked actions within a week. Whatever you think of the pace of deployment, "labs running agents with no supervision" no longer describes what is happening inside at least this one lab.
How did Anthropic measure it? An unusually transparent methodology
Anthropic did not just publish a percentage; it published how the percentage was produced. A weekly sample of 20% of staff during July 2026, drawn from Slack and internal documents, yielded roughly 15,000 tasks organized into a frozen tree of 542 nodes. A Claude-based judge then rated the tasks — and, notably, the company published the judge's reliability: exact model-human agreement of 59%, higher than human-human agreement on the same evaluations (35%), with ratings within one level 97% of the time.
The same conservatism shows in the compute measurement: a Claude classifier sorted workloads on a sample of about 14% of 10,000 runs, and dual-purpose work (safety and capability together) was counted as R&D rather than safety. That makes the 6% safety figure a deliberately conservative floor rather than an inflated headline.
What does this mean for you?
If you are a developer, researcher, or knowledge worker anywhere from Lagos to Manila, the message is twofold. First, model improvement cycles are likely to accelerate: labs are now using their own products to build their next products, and the end user inherits faster improvements across writing, coding, and analysis tools. Second, mid-level routine work in research and engineering will keep shrinking, while value shifts to people who can direct these systems, decompose problems for them, and audit their output — a skill you can start building today with tools that are already available globally.
The timing matters too. The post lands days after CEO Dario Amodei's essay "We Must Pace the Frontier," which we covered in our piece on the major labs agreeing to slow the race — and it puts hard numbers behind the words: yes, the acceleration is real, and for the first time a lab has quantified its internal automation at this level of detail.
Quick comparison: how labs handle R&D automation transparency
| Organization | What has been published officially | Level of detail |
|---|---|---|
| Anthropic | Claude leads 26% of its R&D; 90%+ with AI involvement; full methodology post | Highest so far: figures, limits, and method |
| OpenAI | Earlier disclosures about an automated research program and model incident reports | Announcements without a unified published index |
| Google DeepMind | Governance and safety initiatives without publishing internal automation percentages | Qualitative rather than quantitative |
The comparison reveals a new competitive axis: transparency. The first lab to publish debatable internal numbers earns credibility with regulators and users alike — which explains why this post drew global coverage within hours.
Read the numbers critically: the limits of the study
This is a welcome disclosure, not a peer-reviewed study. The sample comes from one company measuring itself; the judge is a version of Claude, the system that benefits from a positive result; and no work reached full autonomy (AL5), meaning humans still set goals and review blocked decisions. The figures describe Anthropic only — they cannot be generalized across the industry — and the measurement window ends in August 2026, so any later jumps are not captured.
Does this mean Claude is now an independent researcher?
No. AL4 means Claude leads work under human supervision; no measured task reached AL5, full autonomy. Humans still define objectives and review blocked actions within a maximum of one week.
How does this relate to Dario Amodei's call to slow down?
The post explicitly references Amodei's "We Must Pace the Frontier" essay and positions these measurements as the factual basis any serious pacing discussion needs — decisions without measurement remain purely theoretical debate.
Is Anthropic actually in control of tens of thousands of agents?
By its own numbers: pre-execution monitors cover 100% of actions within seconds, offline review scans every transcript, and about 1 in 47,000 decisions gets blocked. Transparency in publishing does not eliminate the need for independent external verification.
What does this mean for programming and research jobs?
Routine mid-level tasks will shrink faster, while the premium grows for people who can manage fleets of agents and audit their output. Investing in direction and review skills is the practical hedge.
When do these capabilities reach ordinary users?
Claude's public capabilities improve with every development cycle accelerated by this automation, but the 26% figure describes internal company research — it is not a product you can buy today.
Three things are worth watching now: whether other labs adopt the same numeric disclosure habit, whether these measurements repeat quarterly and become an industry standard, and whether internal automation shows up in the pricing and timing of upcoming releases. We will keep tracking those questions in our ongoing AI coverage — and if you want to move from following the news to building actual skills, the curated AI course collections at Truescho are a practical place to start.
Sources
- Anthropic — Measurements for understanding the pace of AI development — the official post with all figures and methodology
- Reuters — Anthropic says Claude now leads a quarter of work building its next AI models — first wire confirmation (September 17)
- Associated Press — Claude is helping to build the next version of itself — main US wire coverage (September 18)
- Our coverage: AI's biggest rivals agree to slow down — the full background on Amodei's essay
- Our coverage: OpenAI's model misalignment reporting framework — how labs handle model behavior