OpenAI Releases 722 Math Manuscripts Written by an Unreleased Internal Model: Some Lean-Formalized, and Its Advisory Group Explicitly Does Not Endorse Them

On October 6 OpenAI published on GitHub 722 mathematics manuscripts produced by an unreleased, unnamed internal frontier model, grouped into 372 families of related results and drawn from an evaluation of about 4,000 research problems; OpenAI says each result used on average the equivalent of about three hours of ChatGPT Pro thinking compute. The repository includes Lean formalizations for some results and 10 abridged reasoning summaries, and OpenAI acknowledges that "some of the unformalized results could have issues" and that results are at different stages of verification. OpenAI says the release drew on advice from the Institute for Advanced Study's Advisory Group on Mathematics and AI, but that group had already said it does not endorse testing advanced mathematical problems on proprietary models, and most of its recommendations, including naming the model, publishing prompts, reasoning summaries for every result and compute costs, and using a community-controlled repository, were not met.

What was released

On October 6 OpenAI published "Sharing AI progress in mathematics" and put 722 mathematics manuscripts on GitHub at once, grouped into 372 families of related results. All of them come from an internal frontier model that is unreleased and unnamed: OpenAI used it in an evaluation of about 4,000 research problems, with each result consuming on average the equivalent of about three hours of ChatGPT Pro thinking compute.

Results named in the README include two-point correlations of multiplicative functions, the irrationality exponent of π, the symmetric and general Mahler conjectures, and Kaplansky's direct-finiteness conjecture in characteristic two. Some manuscripts come with Lean formalizations that a computer can check, but not all; OpenAI itself says the results are "at different stages of verification" and that "some of the unformalized results could have issues," with corrections recorded as new versions and old versions kept. Only 10 abridged summaries of the model's reasoning were published.

Where the advisory group stands

OpenAI says the release drew on advice from the Institute for Advanced Study's Advisory Group on Mathematics and AI (AGMAI). But when that group published recommendations in late September, based on responses from more than 600 mathematicians, it said plainly that it does not endorse testing advanced mathematical problems on proprietary models and asked labs to stop. Its recommended practice includes repositories not controlled by any AI lab, and publishing the model name, prompts, reasoning summaries, time taken and compute cost. According to Unite.AI and other outlets, this release still sits in OpenAI's own GitHub account, the model is unnamed, and reasoning summaries cover only 10 results, so most recommendations were not met. OpenAI says it is still looking for a community-hosted option that meets the group's guidelines, and will fund workshops around AI-produced mathematical results.

What it means

If a substantial share of these results survives peer review, this will be the largest public instance of AI contributing to mathematical research. Until then, it is first of all a batch of manuscripts awaiting review: only the Lean-formalized parts can be checked quickly by machine, the rest need mathematicians to read them one by one, and the model that produced them cannot be used or reproduced by anyone outside. For researchers, the safer approach is to start with the formalized results and avoid citing unformalized conclusions until they are independently verified.

via: OpenAI post, GitHub repository, OpenAI on the advisory group, Unite.AI report