All news

OpenAI publishes 722 AI-written maths papers. 162 carry a proof a computer can check

An unreleased OpenAI model was set roughly 4,000 open problems and produced 372 families of results. The company has put them all on GitHub, with Lean proofs for some, and says it is working to release the model itself.

By Sahi Padhai News8 min read

On the evening of 6 October, OpenAI made a GitHub repository public and, with it, handed the mathematics community something it has never had to deal with before: 722 research manuscripts, written by a machine, claiming progress on open problems, and arriving far faster than anyone can check them.

Highlights
  • What was released: 722 mathematical manuscripts, organised into 372 result families, in the public repository openai/math, under an Apache-2.0 licence.
  • Who wrote them: an unreleased internal OpenAI model. It has no public name, and OpenAI says it is "working to responsibly release" it.
  • How they were produced: the model was posed approximately 4,000 problems, each result using on average the equivalent compute of roughly three hours of ChatGPT Pro thinking.
  • How much is verified: 162 of the 722 manuscripts have a Lean formalisation — a proof a computer can check line by line — covering 185 formalised statements in all.
  • OpenAI's own caution: "Some of the unformalized results could have issues."
  • What OpenAI is doing about it: it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and says it will fund workshops and conferences on understanding major results produced by AI.

What OpenAI actually released

The repository is called openai/math, and it went live at 21:47 UTC on 6 October 2026. It contains 722 manuscripts grouped into 372 result families — a family being a principal result together with its companion arguments, consequences or alternative proofs. Each family is classified by mathematical discipline, and each manuscript comes with its PDF, its source files and its own citation block.

OpenAI's framing is careful, and worth preserving as the numbers get repeated. These are results at "different stages of verification", in the repository's words, and the company describes the collection as a broad range of new mathematical results rather than a list of solved conjectures.

Everything is released under an Apache-2.0 licence, so the text and the accompanying code can be reused with attribution. OpenAI says it will keep the full release history, recording corrections and revisions as new versions, and that the repository carries "protocols for paper revisions and citations". It is also, it says, still looking at community-hosted alternatives to GitHub for material like this.

How the results were produced

The method is set out plainly in the repository's readme, and it explains both the scale of the release and the unease it has caused.

"The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model," the readme says. "On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."

So: four thousand problems in, 372 families of results out, at roughly three hours of reasoning each — an estimate of equivalent compute, measured in ChatGPT Pro usage, rather than a stopwatch reading. The filtering, "requiring an appropriate level of significance", was OpenAI's own.

There were exceptions. The readme singles out work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties as not following the standard procedure. It also notes that the write-up for the Re(s) > 11/12 zero-free region "was human edited for readability" — the only human involvement the company declares anywhere in the collection.

Why do this at all? The readme gives a revealing answer: "As part of model development, we evaluate our models on open research problems. We expanded these evaluations after performance on our existing mathematical evaluations saturated." In other words, the tests ran out. The model had stopped failing the benchmarks, so OpenAI pointed it at problems nobody has solved.

The number that matters: 162 of 722

Alongside the manuscripts, OpenAI published formalisations in Lean, a language in which a proof is written out in such complete logical detail that a computer can verify every step. A result that passes a Lean check is, as far as anything in mathematics gets, certainly correct.

The repository's own formalisation catalogue lists 162 manuscripts with a formalised main result, covering 185 formalised statements between them. The readme puts it less precisely: "Not all have accompanying Lean formalizations. We will continue to update this repository with Lean formalizations as we obtain them."

It also carries a warning that deserves more attention than it has received: "Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly."

That gap is the story. For 162 manuscripts, a machine has confirmed the logic holds. For the remaining 560, the argument rests on the word of a model trained to produce text that looks right — which is precisely the failure mode described in hallucinations and jailbreaks. A language model has no internal mechanism separating "I have proved this" from "this reads like a proof", and a wrong proof written fluently is much harder to catch than a wrong answer.

One further detail from the same catalogue is easy to miss and worth stating. The formalisations were themselves produced by an agent, the catalogue records its own scope as "Partial progress", and its review status is marked unchecked. Lean's verification does not depend on anyone trusting the agent that wrote the Lean code — the compiler either accepts the proof or it does not — but it does mean nobody has yet confirmed that each formal statement faithfully expresses the theorem the paper claims. That is a human judgement, and it has not been made.

OpenAI also released abridged summaries of the model's reasoning for ten of the families, including the irrationality exponent of π, the symmetric and general Mahler conjectures, and the isomorphism of free group factors.

What a Lean check does and does not tell you

Lean verifies that a proof's logic is valid. It cannot tell you whether the result is new, whether it is interesting, or whether the formal statement matches what the paper says in English. A formalised proof of something trivial is still trivial. Judging significance remains entirely a human job — and with 722 manuscripts, that job has just become very large.

What OpenAI did to prepare the ground

This was not dropped without warning, and the company's own account of the preparation matters to how the release should be judged.

OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and drew on that group's advice and public recommendations in deciding how to publish. It has committed to improving future releases "via the citations, mathematical exposition, and presentation of the results for better understanding" — an acknowledgement that machine-written papers are not yet easy for mathematicians to read.

It also says it will fund "a series of workshops, conferences, and special programs around the understanding of major results produced by AI", with details to come. And on the question that most limits what anyone outside the company can do with this: OpenAI says it is "working to responsibly release the model that produced these results", though it gives no date.

Why mathematicians are still uneasy

The reaction has split, and not along the lines you might expect. Some see exactly the transparency the field asked for: full papers, machine-checkable proofs, an open licence, reasoning summaries, no press release standing in for evidence. Others see the bottleneck simply moved — from a benchmark score to a volume of review that no community can absorb.

According to the decoder, Timothy Gowers, a Fields medallist consulted through the Institute for Advanced Study advisory group, warned that within one to two decades the mathematical literature could grow enormously while no human community remains that truly understands it. Terence Tao, also a Fields medallist, argued that training young mathematicians must emphasise the human side and tightly limit AI use so that genuine learning survives. The same report describes an open letter signed by 25 Fields medallists, "A Severe Misalignment of AI in Mathematics", warning that mass-producing true statements could destroy fertile ground rather than bring new ideas to life.

The objection is not that the results are worthless, and it is not that OpenAI failed to consult anyone. It is that reading and judging 722 papers is slow, specialised work, it falls on a small number of people, and funding workshops does not by itself make the arithmetic work.

What to watch

Whether 162 becomes 722. OpenAI says it is still adding formalisations. How fast that number climbs is the clearest measure of how much of this collection is solid.

Whether anything is found to be wrong. With 560 unformalised manuscripts in public, the first confirmed error will say a great deal — about the model, and about how quickly the community can audit work at this scale.

Whether a journal takes one. Peer review has no settled policy for a paper whose author is an unnamed model. Someone will have to write one.

When the model is actually released. OpenAI says it is working on it. Until that happens, none of this is reproducible by anyone outside the company, and for a company publishing mathematics that is an uncomfortable position to hold.

The honest summary

Something real has happened here: a model produced a large body of mathematics, and a meaningful part of it has been machine-verified as logically sound. That is not nothing, and the saturated-benchmarks admission tells you the labs are running out of ways to measure these systems on anything short of real problems.

But none of it is peer reviewed, most of it is unformalised, all of it comes with OpenAI's own assessment of its significance, and the model behind it cannot yet be examined. The right response is neither celebration nor dismissal. It is the 162.