The OpenAI Math Repository: 719 Manuscripts and the Mathematicians Response

OpenAI's math repository on GitHub lists 719 manuscripts organized into 372 families, produced by an unreleased internal model that was posed roughly 4,000 problems. OpenAI published the collection on October 6. About 42 percent of the top-line results come with Lean formalizations, which let a computer check each step of a proof. The average result used the equivalent of about three hours of ChatGPT Pro thinking. The README describes the collection as results at different stages of verification and says, "Some of the unformalized results could have issues."
Ten families come with abridged summaries of the model's reasoning, including the irrationality exponent of pi. The release follows OpenAI's September 8 Navier-Stokes announcement, in which a group of about 10,000 concurrent agents produced a finite-time singularity for a forced fluid in roughly 88 hours. For this release, OpenAI said it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and plans to fund workshops and conferences. The catalogue is at github.com/openai/math.
A mathematician's reply
Po-Shen Loh answered the Navier-Stokes announcement on September 19 in a guest post on Terence Tao's blog. He cited the open letters that followed: the Leiden Declaration with more than 4,000 signatories, Math and AI with more than 7,000, and a letter opposing the Caltech Mathathon with more than 2,000. He also described objections from economists, including Tyler Cowen and Joshua Gans, who argued that mathematicians should adapt and cede control.
Loh proposed that every profession wishing to stay human-led adopt a stated axiom: "We (humans) should help humanity flourish." He knows of no case where a far more capable species hands decision-making control to a less capable one, and he describes frontier AI as a system whose decision processes nobody can read. From there he argues that every point where AI touches banking, water or software needs a domain expert in charge, and that someone must steer mathematical research from the frontier. He expects more such control points than qualified people to staff them. The post cites no labor-market figures for that shortfall. For mathematics, Loh lists changes the community could make: ending any stigma around using AI for discovery, building a pipeline of people able to oversee AI agents, counting teaching in hiring and tenure decisions, and mathematicians taking on government problems.
What the proofs look like
Scott Aaronson, a computer scientist at UT Austin, wrote on October 7 that the release includes a proof of Subhash Khot's Unique Games Conjecture, which his wife, complexity theorist Dana Moshkovitz, has worked toward for her whole career. He wrote that no human appears to have understood just about any of the proofs yet, and that Lean certificates exist for some of the results but not all.
He published texts in which Moshkovitz gave her first reading. The proof invents a recursive code with a noise test that matches neither of the standard codes in the field, which she called "some alien craziness." She described the paper as poorly written, with many irrelevant citations, and said it was unreadable without AI help. She asked an AI model called Astra for the completeness and soundness claims of one component, and it assembled them from across the paper. The paper also gives direct NP-hardness proofs for Max Cut and other constraint satisfaction problems that bypass the conjecture. In a later comment, Aaronson reported that she now mostly understands the proof and wants to give talks on it.
Aaronson's list of other results includes integer multiplication in less than n log n time and a positive answer to the unitary synthesis problem, which he and Greg Kuperberg posed in 2007. The unitary synthesis paper proves an oracle exists without giving a way to construct it. Asked whether the model could have exploited bugs in Lean, he replied that he cannot rule it out until the kernel bugs are fixed, but he considers it very unlikely, since researchers who reviewed the reasoning chains found no sign of it. P versus NP is absent from the release, as are P=BPP and NEXP not in P/poly, which Aaronson read as a sign that the hardest problems remain hard. On graduate training, he guessed students will pair heavy use of AI tools with learning material the old way, since otherwise no human minds would be left to do the first part.
He also described two ways of communicating AI results. OpenAI posted undigested proofs and set off a race among humans to explain them. After an Anthropic model supplied the key idea for Virginia Williams and Josh Alman's 3SUM result, Anthropic let them write the announcement, with compensation. The Advisory Group on Mathematics and Artificial Intelligence stated that its role is not an endorsement of how OpenAI obtained the results, and that the mathematical community must assess them.

