Anthropic's Second Risk Report: The Internal Benchmark for "Can a Model Replace Our Own Researchers" Has Saturated

2 views

Anthropic published its second company-wide risk report on August 14 — 186 pages, and the first to assess unreleased internal models alongside shipped ones. It raises the rating for catastrophic harm from misalignment in high-stakes settings from "very low" to "low," and states plainly that CoBench, the internal benchmark tracking its AI R&D automation threshold, has saturated: its most concrete task-based evaluations no longer register capability gains. The report also discloses Model 2, an internal model it does not plan to release.

The Ruler Ran Out First

CoBench is Anthropic's internal benchmark for judging whether a model is approaching the point where it could substitute for the company's own research scientists and engineers — the AI R&D automation line drawn in version 3.4 of its Responsible Scaling Policy. The report's wording is blunt: the rating remains "low" and neither RSP criterion is considered met, but "we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have 'saturated.'" The report also notes early signs of acceleration inside the company: Claude now authors the large majority of code merged into Anthropic's production codebases, and AI-assisted R&D is meaningfully faster than unaided work, though not yet by a factor of two. In other words, the thing being monitored is speeding up while the instrument monitoring it has pinned at the top of its scale.

The Rating Went Up, but Not Because a Test Failed

Misalignment risk moved from "very low" in the first report in February to "low." The trigger was not a failed safety evaluation but recent cybersecurity-evaluation incident disclosures raising overall uncertainty, and Anthropic writes that its underlying argument likely still supports "very low." What was downgraded is really the company's confidence in its own judgment, not its conclusion about model behavior. The report covers February 24 through July 15, 2026.

Model 2, a Bio Classifier Gap, and the Incident Cut Entirely

As of the coverage date, the company held three unreleased frontier or near-frontier models internally: Claude Opus 5 (released since), a lower-usage Model 1, and Model 2. Model 2 scores 62.8% on CoBench against 50.3% for the released flagship Mythos 5, and there are no plans to release it publicly. Bio and chemical risk stays "low, but higher than our previous estimate": Anthropic found that between May 2025 and April 2026, roughly 133 million exchanges involving about 50,000 human-feedback contractors ran without its blocking biological classifiers in place — since remediated, with no evidence of concerning misuse. One incident from the covered period was redacted entirely from the public edition, and Anthropic had Mythos review the report before release; the model's feedback flagged that redacted incident as among the most consequential material withheld.

via: Anthropic, "Risk Report: August 2026", SiliconANGLE report, Unite.AI report; verified 2026-08-16