Anthropic Says Claude Now Leads 26% of Its R&D — While Across 30,000 Agents, Only One Action in 47,000 Gets Blocked

On September 17, Anthropic published internal figures: by its "Anthropic R&D Automation Index," 26% of its R&D work had reached level 4 as of August 2026 — humans set broad direction and AI performs most of the work on its own. The scale runs 0 to 5, where 0 is humans doing everything and 5 is full autonomy. The company defines "leading" as Claude completing most of a given task end-to-end from a high-level prompt while remaining under human supervision. The growth curve is steep: under 1% in February, 26% in August; on more than 90% of R&D work, Claude operates at or above the "collaborates" level, meaning it handles large chunks of work under close human direction. On scale: at any moment in August roughly 30,000 agents were writing code and running experiments on its busiest internal platform, with every action screened first; of more than a billion decisions that month, about one in 47,000 was blocked, and humans directly reviewed around 50 high-priority cases per week. Separately, about 6% of AI R&D compute went to safety-related research, based on a tally for the week of July 13–20. The company's stated reason for publishing is that it should do everything possible to minimize the gap between what frontier labs know and what the public knows.

Read the Scale Before You Read the 26%

This figure reads far too easily as "26% of R&D has been automated." It isn't. The 26% is the **share of work that reached level 4**, and level 4 is defined as humans setting broad direction while AI performs most of the work on its own. So the accurate phrasing is that 26% of R&D has moved into a mode where humans only handle direction. "End-to-end" and "under human supervision" can coexist in one announcement precisely because of that 0-to-5 grading. The general method for reading numbers like this is to go find the scale definition first — without one, words like "leading" and "collaborating" can mean anything. The figure about more than 90% of work at or above "collaborates" deserves reading alongside it, since that level is defined as handling large chunks of work under close human direction. Most R&D now has model involvement, but for the majority it sits in the tier where a person is still watching. The genuinely hands-off portion is that 26%.

One in 47,000 Proves Two Things at Once

Roughly 20,000 blocks out of a billion decisions means the screening system is clearly working. But put the other number next to it — humans directly review about 50 high-priority cases a week — and against a billion decisions a month, the human-seen fraction sits around 10⁻⁷. That is not a criticism; it is what scale forces. Thirty thousand concurrent agents cannot be watched by people. What it does is pin down what "human supervision" actually means at this volume: **automated screening plus a vanishingly small sample, not people looking.** Any team scaling its own agent fleet hits the same transition, and this data measures where the line falls: at a billion decisions, oversight has to be a program, and humans only handle the few dozen the program surfaces. The 6% figure needs its terms read too: that is safety-related research as a share of **AI R&D compute**, not of all compute, and the window is a single week (July 13–20), not a long-run average.

Put This Week's Three Stories Together

On September 14 this site covered Amodei calling for the industry to slow down, with Altman claiming the item about independent evaluators with employee-like access. On September 18 we covered OpenAI's misalignment disclosure framework and its six accompanying cases. Today Anthropic publishes the curve of AI's share of its own R&D rising from under 1% to 26%. The same set of companies, in the same week, publicly calling for deceleration and publicly documenting their own acceleration. The company's stated reason is closing the information gap — a reason that holds up, and this data does supply something that did not exist before: a public metric that can be tracked month over month. But it does not dissolve the tension; it puts the tension on the table. What to watch is the next update. With the scale, the level definitions and the measurement terms all public, this curve is now set up for continuous comparison: what follows 26%, whether that 90% tier shifts upward as a whole, and how the one-in-47,000 block rate moves as scale grows. Once the first set of numbers exists, every later set can be checked against it — which is worth more than any single statement.

via: The Washington Post, The Korea Times, Digital Today, Technology.org