Put the Statement and the Framework Side by Side, Word for Word
On September 12, Altman said: desks, badges, laptops, and the right to publish. On September 22, the framework says evaluators **may** be brought into the offices for **the most sensitive** work. "May" and "the most sensitive" are two qualifiers, and together they move office access from a default to an exception. As for publication rights — whether an evaluator who finds something can say so publicly, and how much of it — no corresponding clause appears in the new text. This is not an accusation that OpenAI went back on anything. A statement and a framework are different genres, one addressed to an audience and one addressed to a contract. But precisely for that reason, **the step from statement to framework is the only reliable measure available for things like this: did the specific commitments increase or decrease?** This time they decreased.
The Real Gain Is in the Words "Training Phase"
Having finished the criticism, the progress deserves to be stated accurately too. Third-party assessment has until now happened at or near release, so assessors received a finished product — they could see inputs and outputs, not how it came to be that way. This framework pushes assessment back into the training and evaluation stages. That is a substantive change, not a rhetorical one. And it points at the same thing this site covered in Anthropic's announcement on September 20: seeing the process, not just the result. Two companies moved the same idea forward in the same week, and that synchrony is worth recording. Of the four priority areas, the one that deserves its own mention is the fourth: **independent investigation of misalignment incidents.** On September 18 this site covered the misalignment disclosure framework OpenAI had just delivered, with six- and twelve-business-day deadlines and six of its own cases attached — that was "we disclose it ourselves." Today's is "someone outside investigates." Only both together make an accountability chain; the first alone, however detailed, remains self-report. Worth adding: writing about the Gemini breakout on September 20, this site argued that a disclosure framework triggered by *misalignment* cannot hold incidents like Google's. If an external investigation right actually materializes, there is a chance that gap gets filled by an outside assessor rather than by the vendor itself.
The Two Companies' Approaches Are Mirror Images
Set OpenAI's and Anthropic's moves from these two weeks next to each other and the contrast is clean. Anthropic gave a name (Accenture), a number (at least $1 billion over five years), and an access level (comparable to an employee's). The cost is the relationship: Anthropic funds the work directly, and Accenture is reported to be Anthropic's largest Claude Code deployment partner — **the payer, the assessed and the customer are all the same pair of companies.** That is one layer tighter than was known when this site covered the deal on September 20. OpenAI gave no name, so the problem does not arise yet. But neither is there anything to check: no partner, no timeline, and an access level written as a possibility and an exception. **One has concrete terms and a conflict of interest; the other has no conflict and no terms.** This is not a ranking. It is an observation that no approach yet exists that is both independent and operational — and that the half each company is missing is exactly the half the other one has.
The Only Usable Test Right Now Is a Future Event
Until there are names and dates, this document is a statement of intent. So judging whether it gets honored requires a specific, checkable event in the future tense. That event is: **the first system card that names an outside evaluator and states whether they saw training-phase data.** All three elements are required — the name tells you who, the stage tells you what they saw, and appearing in a system card tells you it entered the routine process rather than being a one-off arrangement. One piece of legal background, stated plainly. Yesterday this site covered the antitrust class action against four labs over agreeing to slow down, and the phrase OpenAI uses to frame this document is precisely "pace the frontier." Worth noting that this is a **unilateral** act — opening your own models to third-party assessment requires agreeing to nothing with any competitor. That is considerably cleaner legally than an industry-wide body, and it may explain why several companies spent this week moving toward doing it themselves.
via: OpenAI, "Priorities and principles for effective third party assessments", OpenAI, "Strengthening our safety ecosystem with external testing", The Next Web, StreetInsider