OpenAI Says a Secret AI Model Cracked Hundreds of Open Math Problems in One Prompt—Mathematicians Want Receipts
OpenAI published 722 math manuscripts on GitHub on Tuesday, all produced by an internal model the company has not released. An OpenAI spokesperson said almost everything came from a single prompt handed to a single AI agent, though some may have taken multiple attempts. It's a bold claim and a potentially significant breakthrough in the field of mathematics. But not everyone is a fan, or buying the hype. “Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified,” >Andrew Sutherland, a mathematician at MIT, told Scientific American. “We should ask for receipts,” he said. The papers are grouped into 372 “families” of related results, and a family can bundle a main theorem with companion arguments, consequences or alternative proofs. That makes 722 a count of manuscripts, not of solved problems. OpenAI says it posed roughly 4,000 problems to the model and kept the outputs it judged significant enough to publish. The average result used the equivalent of roughly three hours of ChatGPT Pro thinking compute, >per OpenAI. The Navier-Stokes claim last month looked very different, with 10,000 coordinating agents working for 88 hours. OpenAI released abridged reasoning summaries for 10 of the results. That said, only 162 of the 722 papers come with a computer-checked main result, according to a formalization catalog in the repository. That is about 22% of the collection, translated into Lean, software that checks every logical step mechanically. OpenAI itself says not all manuscripts have Lean formalizations and that “some of the unformalized results could have issues.” In other words, a lot of what they published could be wrong. A passing Lean check confirms only that the proof follows from the statement as written in Lean. It >does not show that the statement matches the original problem, or that the result is new or important, which is the part mathematicians now have to judge. And this is where research
AI Analysis:
Disclaimer: This information is from public sources for reference only. Traceless does not guarantee accuracy.