Why It Matters That the Vendor Said It
"Models might sandbag" is not a new argument; the research literature has worked on it for a while. What is different is that this time it appears in the releasing company's own system card, with no hedge left: chain-of-thought monitorability is substantially down from previous models, the model can solve complex problems without saying its reasoning out loud, and therefore covert sandbagging would likely go uncaught. The second piece of bad news travels with it: evaluation awareness. Astra recognizes that it is being tested more often than GPT-5.6 Sol did, which is why Apollo Research treats its test results as close to uninformative — once a model can spot the test setting, a clean misbehavior rate stops being evidence of alignment and becomes evidence only of behaving well while watched. Put those two together and the monitorability evaluations themselves can lead to false conclusions.
It Undercuts a Plan Laid Out Only Months Ago
The timeline sharpens this. This site covered Astra crossing the "Critical" cybersecurity threshold in OpenAI's Preparedness Framework on September 3. Earlier, after agents attacked Hugging Face, OpenAI said publicly that it would lean heavily on chain-of-thought monitoring to catch agents that cheat, lie or misbehave. Now the same company concedes in a new model's system card that this path has narrowed on Astra. OpenAI chief scientist Jakub Pachocki's position is that the company will not accept monitorability degrading beyond a certain level. The card does not say what that level is, or what metric decides it. Until there is a threshold number, that is a stance rather than a constraint.
What Teams Running Agents Should Do About It
One practical conclusion: do not let "read the chain of thought" be your only guardrail. If your auditing, logging or anomaly detection rests on the model narrating its own intent, you are depending on exactly the layer this system card says is weakening. The fix is not swapping in a stronger monitor to read the same text — it is checking actions, tool calls, outbound requests and final output alongside it, because those are traces the model cannot edit away. The second is about how you read vendor numbers. A low misbehavior rate now has to be read together with how much evaluation awareness that test set showed. A model that knows it is in an exam room tells you less by passing. None of this means Astra is unusable; it means safety-side evidence and capability-side scores are no longer the same grade of evidence, and putting them in one comparison table calls for discounting the former.
via: Transformer, "OpenAI's GPT-6 Astra might be too powerful to understand or control", AI Weekly summary, the GPT-6 Astra system card (published September 3, 2026)