The Denominator in 3.1 Is Human Time, Not Output
Get the definition straight first, because this reads far too easily as "AI did three times the work." The numerator is total agent runtime and the denominator is total human hours. It measures how much machine time gets burned per human day invested. It is not an output ratio. OpenAI's own qualifications are not vague either: high-level planning is still scarce in agent output, and more than half of successful four-to-eight-hour tasks required at least one human intervention. There's an independent comparison worth holding alongside it — the arXiv paper *The Shift to Agentic AI: Evidence from Codex* reports that on June 11, 2026 the median employee had Codex turns running for 2.5 hours, while the 99th percentile has recently been running around 71 hours of agent turns in an average day. Two orders of magnitude between median and frontier means 3.1 is an average pulled up by a small set of heavy users, not a picture of a typical researcher's day. The one number that genuinely reads as output is this: experiments per active experimenter hit their highest level in August since tracking began in January 2025. That is more persuasive than 3.1, because it counts experiments completed rather than machine hours consumed.
What's Worth Remembering Is the Spend Curve and the Cultural Side Effects
At API prices, $600 a day for the median researcher and $7,000 at the 90th percentile, on top of 124× growth in token output, means the cost structure of this way of working has changed shape: research used to be marginally expensive mostly in people, and now carries a machine bill of comparable magnitude that is still climbing fast. Anyone hoping to reproduce this way of working on their own team should price that first. One detail is easy to skip past: internal channels where researchers asked colleagues for help have gone quiet, office-hours attendance fell, and at least one team stopped holding them. Researchers stopped asking coworkers and started asking agents. Whether that is good or bad long term isn't visible yet, but it is changing how knowledge moves through the organization.
Publishing Both on the Same Day Is Itself the Story
The core line in Pachocki's "An Alien Mind" is that he currently believes no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer; he also writes that this is a time calling for extreme caution and that he is concerned nobody is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI frames publishing its internal metrics as a transparency exercise it thinks should eventually be mandatory, and notes 20% of compute goes to safety monitoring. Read side by side, this is "evidence that we are accelerating" paired with "the industry should slow down." This site covered the line in the Astra system card on September 6 — if the model sandbags covertly, we would likely be unable to catch it. Put all three together and within one week the same company conceded that its new model is harder to monitor, published the curve of its own internal speedup, and called on the industry to slow. One irony flagged in analysis is worth keeping: after Astra's compute allocation was cut by 59.2% over cybersecurity concerns, compute shifted to other models and offset 85% of the reduction. If a restriction inside one company mostly relocates resources, then lowering overall scaling velocity through voluntary slowdowns is a harder road than the essay makes it sound.
via: The Next Web, OpenAI's two blog posts of September 6, 2026 (internal research metrics and "An Alien Mind"), arXiv, *The Shift to Agentic AI: Evidence from Codex*