A Sentence Spoken Four Days Ago Is Now a Contract
On September 14 this site covered Amodei's call for the industry to slow down, and the specific item Altman claimed in his reply was exactly this: giving independent evaluators access comparable to an employee's. On September 18, Anthropic turned that sentence into a partnership. Four days. The speed is itself worth recording — it tells you this was not an idea floated that day, but something already under negotiation that simply got said out loud. And because it landed this fast, the standard for judging it should switch from vision to terms. The term worth reading is the access level. External evaluators only see a model's outputs; embedded evaluators can watch it take shape in training, follow how build and deployment decisions get made, and ask employees directly. **That is a substantive increase in evaluation capability, not a rhetorical one.** Seeing the process and seeing the result are two different kinds of evaluation.
Yesterday's Structural Problem Shows Up Here in Its Purest Form
Writing about AEF-1 on September 18, this site named a structural problem: the evaluator's revenue comes from the evaluated. Today's story takes it to the limit — the evaluator not only takes the assessed party's money, it sits in the assessed party's building, on access the assessed party grants. What has to be recorded alongside that is that Anthropic states the point itself in the announcement: under the framework it published in June, long-term funding for independent evaluation should come from pooled or government sources, and it is only absent such a system that it works with different evaluators under different arrangements. **This is an announcement that concedes it is a second-best arrangement**, which is not common in this genre and deserves credit. But conceding is not solving. There are no published conflict-of-interest terms — whether an evaluator who finds something can say so publicly, how much of it, and what happens when they are overruled. Those clauses are where the word "independent" actually lives, and none of them are written down yet. That is the thing to press on next. One distinction not worth blurring: Accenture is not a nonprofit. It is a consultancy selling AI implementation to a great many enterprises, which puts it in a different position from an organization like METR. Anthropic says it is in dialogue with METR and others to pilot embedded evaluation on their own funding — if that path materializes, it is more independent than this one, because the payer and the assessed would finally not be the same party.
Read the Terms on "At Least $1 Billion Each"
It is easy to read "each expects to invest at least $1 billion" as a billion dollars spent on buying evaluations. Per the announcement, this is what each company expects to put into **building capacity in this area** over five years — not the fee for this evaluation work. Faculty's position deserves the same care. It is an applied-AI company Accenture acquired, with prior experience testing and evaluating models for major AI laboratories and building systems for high-stakes settings including government, defense, healthcare and infrastructure. That résumé explains why it was chosen, and it also means this is a team that grew up doing vendor-side work now moving to the assessing side. Two things to watch over the coming weeks. Anthropic says more evaluators will be announced — whether any of them take no money from Anthropic is the first real signal of how much this arrangement is worth. And whether the METR pilot on its own funding actually happens. Until then, the accurate description of this partnership is: **a large step forward in evaluation capability, still missing the published terms that would institutionalize the independence.**