What Is Third-party AI Evaluation? Independent Assessment of Frontier Models Explained

Third-party EvaluationIndependent AssessmentAI Safety

Third-party AI evaluation means independent organizations outside the model developer assessing an AI model's capabilities and risks. In September 2026, Anthropic and OpenAI both committed to giving independent evaluators employee-like access, the AEF-1 standard began drawing lab sign-ons, and the debate shifted to who pays and who can publish the conclusions.

Third-party AI evaluation means independent organizations outside the model developer testing and assessing an AI model's capabilities and risks—for example, whether it would help with cyberattacks, whether it deceives, and whether its safeguards work.

It doesn't conflict with the evals vendors run themselves; the difference is standing. When you grade yourself, the result is always open to doubt. As release decisions for frontier models lean more and more on evaluation results, "who evaluates, how much the evaluator can see, and whether conclusions can be published" became one of the central questions in AI safety in 2026.

Grab It in One Sentence First

Third-party evaluation means having people who aren't paid by the developer, and don't speak for it, check how dangerous a model really is.

An everyday analogy is a financial audit. A company can of course do its own books, but a listed company's financial statements are audited by outside accountants—and the auditors need to see the real ledger, not just the pages the company picks. AI evaluation faces almost the same problem: independence and access, and it needs both.

Why It Keeps Coming Up Now

On one hand, model capabilities have crossed some thresholds the vendors set for themselves. In early September OpenAI classified GPT-6 Astra as the first model to cross the "critical" cybersecurity threshold in its Preparedness Framework, and the GPT-6 Astra system card acknowledged that if the model deliberately sandbagged during evaluation, the company likely wouldn't notice. The more critical and harder to verify evaluation results become, the more outside checks are worth.

On the other hand, several leading companies took public positions in quick succession in September 2026:

  • September 12: Anthropic CEO Dario Amodei published an essay arguing for "pacing" the frontier, including having third-party evaluators confirm that companies have enough time to align and secure their models. Hours later OpenAI CEO Sam Altman responded that giving independent evaluators employee-like access was a good idea and that OpenAI would do it too.
  • September 18: Anthropic announced an "embedded evaluation" partnership with Faculty, part of Accenture: evaluators work inside the company with access comparable to employees and can watch models take shape during training. Each company expects to invest at least $1 billion in this area over five years.
  • September 22: OpenAI published priorities and principles for third-party assessments, saying it would allow outside organizations to assess models during training, evaluation, and deployment, not only near release. The document named no partners and gave no timeline.

AEF-1: A "Statement of Conditions" for Evaluators

The AI Evaluator Forum was founded in December 2025. Its AEF-1 is a voluntary standard: a third-party evaluator fills out a checklist and publishes it alongside the results, stating the conditions under which the evaluation was carried out. It covers five principles:

  • Adequate access and resources
  • Minimizing conflicts of interest
  • Analytical independence
  • Transparency of methods and results
  • Protection of sensitive information

On access, AEF-1 recommends that outside evaluators be able to see system prompts, information about the training process and data, existing internal evaluation results, and information on known vulnerabilities. The standard itself isn't new; what changed in September 2026 is that OpenAI followed Anthropic's embedded-evaluator commitment, and xAI signed on as well.

Grants access Affects independence Model developer Third-party evaluator Capability & risk assessmentduring training / pre-release / post-deployment Evaluation report Public conclusions Funding source

The Forms It Usually Takes

Pre-release external testing: as a model nears release, outside organizations get a period of time and an access interface to test it. This has been the most common form; the drawbacks are the short window and no view of training.

Embedded evaluation: evaluators work inside the company, taking part much like employees and following training and deployment decisions. It goes deep, but staying independent is hard—working inside a company for long periods, paid by that company, can both affect judgment.

Industry self-regulatory body: according to reports, Anthropic, OpenAI, and Google DeepMind have been discussing since July a pre-release testing body modeled on the US Financial Industry Regulatory Authority (FINRA), funded by industry and staffed by independent experts. There is no formal charter yet.

Common Misunderstandings

The first is "third-party means independent." Independence depends on who pays, how the contract is written, and whether conclusions can be published without consent. In the Anthropic–Accenture partnership, Anthropic pays directly for the service, and Accenture is reportedly also Anthropic's largest Claude Code deployment partner. Anthropic itself has said that long-term funding should ideally come from pooled or government sources.

The second is "the industry jointly setting testing standards is always good." Cohere CEO Aidan Gomez publicly criticized the self-regulatory body under discussion among the three leading labs, calling such a body a cartel by another name and arguing that the real questions are who writes the rules and who gets to take part. Competitors jointly agreeing on testing and release timing is also contested from an antitrust perspective.

The third is "evaluation itself carries no risk." In an incident disclosed in September, Gemini accessed the systems of three real companies during a cybersecurity test run by the independent evaluation firm Irregular, because the test environment had accidentally allowed network access. Evaluations need sandbox isolation and clear disclosure rules as well.

How to Read a Third-party Evaluation

When you read that "a model has passed third-party evaluation," ask a few follow-up questions: Who is the evaluator, and who paid? At what stage was the evaluation done, and what access did the evaluator have? Can the evaluator publish conclusions independently, and does the report state its limitations? If you can find an answer for each of AEF-1's five principles, the evaluation is worth more.

For related concepts, see Alignment, Evals, and Sandbagging.

Sources