The AEF-1 Evaluator Standard Isn't New This Week — What Changed Is That the Labs Showed Up: OpenAI Matches Anthropic, and xAI Cosigns

AEF-1 is a voluntary standard from the AI Evaluator Forum, giving third-party evaluators a standard and checklist for demonstrating the conditions under which an evaluation was actually produced; conformance means completing the checklist and publishing it alongside the results. It covers five principles: sufficient access and resources, minimized conflicts of interest, analytic autonomy, transparent methods and results, and protection of sensitive information. On access specifically it recommends external evaluators obtain system prompts, information about the training process and data, pre-existing internal evaluation results, and knowledge of system vulnerabilities, plus a catch-all requirement for "sufficient technical access to assess the specific system characteristics being evaluated." One timeline correction first: the Forum was formed in December 2025, and the standard launched alongside a public statement signed by more than 40 voices from across the sector — roughly nine months before this week's headlines. What changed this week is the labs' posture: OpenAI matched Anthropic's embedded-evaluator pledge, and xAI cosigned. Anthropic's description of the access it is offering is notably concrete: desks in its offices, access badges and company laptops, with permissions roughly comparable to internal risk assessment teams. Organizer Conrad Stosz is a former Acting Director of the U.S. Center for AI Standards and Innovation.

Get the Timeline Straight First

This week's coverage reads easily as "a new standard has appeared." In fact AEF-1 and the forum behind it have existed since December 2025, launched with a statement signed by more than 40 people arguing that the conditions of third-party evaluation should be more transparent, and forum members said then that they would begin adopting it in forthcoming evaluations. So the news is not the standard. It is that **the labs walked over to it.** The distinction matters because it determines what to watch: a voluntary standard sitting unused for nine months and three frontier labs publicly aligning on the same standard are entirely different events.

It Answers Two of the Four Questions We Left Open Last Week

On September 14 this site covered Amodei's "We Must Pace the Frontier" and the item Altman claimed hours later — independent evaluators with employee-like access. We said that line carried weight because it decomposes into four concrete questions: who are the evaluators, how deep does the access go, how much of the result gets published, and starting with which model. AEF-1 answers two of them. **Who the evaluators are:** the Forum's member organizations. Commentary also flagged the other side of that — those members are likely to become the primary third-party auditors the major labs recruit. **How deep the access goes:** the standard's list is fairly specific — system prompts, information about the training process and data, pre-existing internal evaluation results, knowledge of known vulnerabilities, plus the catch-all requirement of sufficient technical access for whatever is being assessed. Anthropic's own description reads more like an onboarding list: desks, badges, company laptops, permissions roughly matching internal risk assessment teams. That is considerably more grounded than the adjective "employee-like." The other two questions remain open: **how much gets published**, and **starting with which model.** Those are the two to track over the next three to six months.

The Old Problem With Self-Regulation Is Not Solved

The criticism needs stating plainly: if AEF members both write the standard and get hired as auditors by the major labs, then "independent" becomes something requiring extra proof — the evaluator's revenue comes from the evaluated, a structure the auditing profession has handled repeatedly and repeatedly gotten wrong. AEF-1's design has a partial answer. Two of its five principles are minimized conflicts of interest and analytic autonomy, and conformance means publishing the checklist alongside the results. In other words it does not promise independence; it requires **laying the conditions out for readers to judge.** That is the right instinct, and it gives anyone reading an evaluation report a concrete new step: **check whether the checklist is attached, and what the access section actually says.** An evaluation with no checklist, and one whose access rows are largely blank, do not carry the same weight. The boundaries, stated: AEF-1 is voluntary, with no enforcement and no arranged consequences for departing from it, and the labs' commitments are likewise voluntary public statements. On September 14 this site noted that pacing proposals lack an enforcement mechanism; that still holds here. What changed is the move from a claim to a checkable list of conditions — not a small step, but not yet a constraint.

via: AI Evaluator Forum: minimum operating conditions for independent third-party evaluation, AI Evaluator Forum launch announcement, Latent Space / AINews on this week's lab cosigning