OpenAI Delivers Hundreds of AI-Generated Math Results, Prompting Years of Review
OpenAI recently released a substantial collection of AI-generated mathematical results, described by more than three dozen mathematicians in conversations with The Verge as "staggering," "overwhelming," "unprecedented," "surreal," and "pure insanity." Researchers anticipate that simply understanding this deluge of material could take years, raising anxiety about the future of the field.

The release includes nearly 400 AI-generated results spread across more than 700 manuscripts. These cover a diverse array of mathematical disciplines, including combinatorics, several branches of geometry, number theory, theoretical computer science, algebra, topology, probability and statistical mechanics, and mathematical physics. OpenAI published guidance to help navigate the sprawling GitHub repository, as the sheer volume makes even preliminary assessment difficult; some mathematicians reported spending the better part of an hour just reviewing the roughly 40-page table of contents and abstracts, with Álvaro Lozano-Robledo of the University of Connecticut calling the experience "overwhelming."
A significant concern among mathematicians is the varying degree of verification for these results. While some manuscripts include formalizations in Lean, a programming language and proof assistant, OpenAI acknowledged on GitHub that results are "at different stages of verification" and "many, but not all, of the manuscripts have been formalized." As of the release, only 300 top-line results out of 719 manuscripts, approximately 42 percent, had been formalized. OpenAI stated it "will update the repository with more formalizations as we obtain them," but researchers expressed frustration over the current lack of formalization and the inconsistent quality of even computer-verifiable proofs.
Kevin Buzzard, a mathematics professor at Imperial College London, identified numerous theorems in his field, algebraic number theory, but noted only around six "stood out" and few appeared formally verified. He voiced concerns about "possibly-not-correct slop," a term for low-quality, erroneous AI-generated material that has seen a "huge uptick." OpenAI's previous mathematical write-ups were criticized for their "sloppy nature" and poor attribution, leading some researchers to call the anticipated release a "slopocalypse." While early impressions of the new papers were "far better than they had expected," many noted this was "hardly a high bar." Without formal verification, the quality of accompanying papers becomes critical for scrutiny, a formidable task given OpenAI's models produce mathematics at a speed far outstripping its human staff's capacity for rigorous review.
What to watch: How quickly OpenAI formalizes the remaining results and the long-term impact on mathematical research standards.
Editor's note: The draft accurately synthesizes the extensive source material, capturing the scale of the release, the verification concerns, and expert perspectives.
AI-generated and fact-checked against the original report; claims the gate cannot verify are held back.