Novel, recalled or borrowed? OpenAI's math results and the missing transcripts
· Christoph Heuwieser
On 6 October OpenAI published openai/math, a public repository of mathematical manuscripts that it says were produced by an unreleased internal model. Many of them claim results on problems mathematicians have listed as open. I've spent a couple of days reading the repository, and I came away with one question it can't answer: where did each of these proofs actually come from?
That isn't a dig at the mathematics. Experts will judge the mathematics, and some of it is already formally checked. My point is narrower. The release doesn't include the material you would need to answer the provenance question, and that is a problem worth talking about.
What OpenAI published
These are the facts as the repository states them on the day I write this:
- Scale. According to the README, the catalogue holds 719 manuscripts in 372 "families" of related papers.
- Process. OpenAI says it evaluates its models on open research problems, and that it widened this evaluation after its existing math evaluations "saturated". The model was posed "approximately 4,000 problems". Each result used, on average, "three hours of ChatGPT Pro thinking compute". The output was grouped into families, and only results meeting "an appropriate level of significance" made the catalogue. The README also says that "some outputs build upon earlier results produced by the models". It names exceptions to that fixed procedure, including two specific results, and says one write-up "was human edited for readability".
- Verification. Roughly 42% of the top-line results (300 of 719) have Lean formalizations. The README warns that "some of the unformalized results could have issues".
- Corrections. The history records that on 7 October three manuscripts were withdrawn after a sign error, and 14 others were revised. Earlier versions stay accessible. That is good practice, and OpenAI deserves credit for it.
- Reasoning. For 10 of the 372 families, OpenAI published what it calls "abridged summaries of the model's reasoning", in the reasoning_traces folder. Nine of the ten are titled "Summarized chain of thought". Most open with excerpts of the original prompt. Each gives a prose account of how the reasoning went, and in the one I read in source form a few short passages are marked as verbatim excerpts from the model.
One detail in those prompt excerpts matters for everything that follows. The prompt for the spin-glass result, for example, states the problem by reference to specific equations and a theorem in a named prior paper (Panchenko and Talagrand), and asks for the conjecture stated right after that theorem. The problem was posed inside the literature that frames it.
What isn't in the repository
I couldn't find any of the following in the README, the overview or the history file:
- the name of the model;
- the full prompts, outside the excerpts in the 10 summaries;
- full transcripts or reasoning traces for any result;
- the failed attempts, or how the roughly 4,000 problems were chosen;
- compute for each result, as opposed to the average;
- what humans did between a prompt and a finished manuscript, beyond the stated exceptions;
- any statement about training data, contamination checks or how novelty was assessed against the literature.
The independent Advisory Group on Mathematics and Artificial Intelligence, which says on its own page that it discussed its recommendations with OpenAI, asked for much of this in its 29 September guidelines. The guidelines ask labs to name the model and publish the prompts, "a (summarized) chain of thought", the time taken and the estimated compute cost. They strongly encourage releasing "the initial LLM outputs before they were cleaned up". For batch releases they ask labs to document "how the problems were chosen" and how many comparable problems failed. They also ask that the literature be searched for related ideas, and that those ideas be cited. In its 6 October note on the release, the group says its role "should not be interpreted as a judgment of the impact of these results or an endorsement of the process", and that it is up to the mathematical community to judge how far its recommendations were followed. Those guidelines say nothing about training data, though.
Three sources that look identical from the outside
When a model produces a correct proof of something that was listed as open, there are at least three ways it could have got there. From a finished manuscript, you can't tell them apart.
(a) Genuinely new reasoning. The model found an argument that nobody had written down. This is the claim that makes a release like this exciting, and for some of these results it may well be true.
(b) Recall or recombination of things seen in training. Models are trained on enormous amounts of mathematical text. A proof can be "new" for this particular problem and still be a known technique moved across from a neighbouring one, or a solution that already existed somewhere obscure. This isn't hypothetical. When a team led by Google DeepMind researchers ran Gemini on 700 open problems from the Erdős problems database, in the paper's current version, 9 of the 13 problems it addressed turned out to have solutions already in the literature, including one that was first counted as novel and later reclassified as an independent rediscovery. The authors made that reclassification partly by checking the model's thinking logs, which is exactly the kind of check a trail makes possible. Even then, they write that "there could have been leakage from the pretraining and post-training phases that we would not be able to detect". They explicitly name the risk of "subconscious plagiarism" by AI. On simpler benchmarks, researchers at Scale AI found accuracy drops of up to 8% when models were tested on fresh problems written to match a well-known test set, which suggests some models had partly memorized it. The same paper notes that many frontier models showed minimal signs of this. Recombination isn't cheating. It's how a lot of human mathematics works too. But it changes what "solved by AI" means, and who should be cited.
(c) Reasoning that came from people working on the same problems. This is the question I can't answer, so I'll put it as a question. Open problems are open because people work on them, and some of them may well use ChatGPT or similar tools while they do it. OpenAI's help centre says that for its services for individuals, such as ChatGPT and Codex, it may use your content to train its models unless you turn off "Improve the model for everyone" in your data controls. Business and API data is handled differently. OpenAI's pages are clear about what the setting covers. Its Terms of Use define "Content" as both your input and the model's output, and the training opt-out applies to Content. Its data controls page says that with the setting off, "your new conversations won't be used to train OpenAI models". The same page notes two limits. The setting applies to new conversations only. And if you rate a response with a thumbs up or down, "the entire conversation associated with that feedback may be used to train OpenAI models". What no user can check from the outside is whether any particular conversation of theirs was used. So the question is: could partial ideas from a mathematician's own sessions with these products have become part of the training data, and then resurfaced in a model's "new" proof?
I have no evidence that this happened for any result in this repository. OpenAI hasn't said it did, or that it didn't. It may not even be fully knowable from inside the lab. But the policies allow the path to exist, and the release doesn't contain the material that would let anyone rule it in or out.
Why the transcript matters
A finished proof tells you what is true. Only the trail can show how it was reached.
A full session record shows the prompt, every intermediate step, the tools called, the searches run, the dead ends and the moment the key idea appeared. With it, a reviewer can ask questions that a polished PDF can't answer:
- Did the decisive idea show up after reading a particular paper, or with no outside input?
- Does a key step closely match an existing argument, a published one or someone else's?
- How many attempts failed first, and what did the failures look like?
- Where did a human step in?
None of this would prove (c) one way or the other on its own, since training data isn't visible in a transcript. But without the transcript, (a), (b) and (c) blur into one headline. And the people whose work may have fed the result have no way to be credited.
The ten summaries OpenAI published are a real step, and I'd like to see more labs do the same. But keep the proportion in mind. There are 10 of them, covering 10 of 372 families and 719 manuscripts, not one per result. They run from 5 to 45 pages each. And they are summaries, written after the fact, of reasoning chosen for publication, not the reasoning traces themselves. Dead ends are exactly what a summary leaves out, and dead ends are often where provenance shows. What would let outsiders start answering these questions is the full trace for every result.
Where we stand
This is the reason a4sx exists in the form it does. We think the session itself, the whole trail of an agent's work with the dead ends left in, is an artifact worth keeping, sharing and attributing to the person who did the work. It shouldn't be thrown away once a polished result has been pulled out of it. When a result matters, the trail behind it is what lets other people check it, build on it and credit it properly.
I don't know how much of OpenAI's release is (a), (b) or (c), and I doubt anyone outside OpenAI does yet. I'd like that to be answerable next time.
a4sx is an exchange for AI coding-agent sessions: publish, share and continue real work, dead ends included.