OpenAI's New Model Solves More Math Problems, Shocks Field
OpenAI’s latest large language model has reportedly resolved hundreds more open problems in mathematics and theoretical computer science. The claim arrives just a month after the company said it had cracked one of the six major open problems in mathematics, and this time OpenAI says it did so without the multi-million dollar computing costs of that earlier effort.

OpenAI announced 372 new results from its internal model, each one either solving or significantly advancing a major open question. The company released the findings via a GitHub repository, and mathematicians will need months to fully assess the proofs. Many have already been verified in Lean, a programming language built to validate logical proofs, which suggests a reasonable degree of correctness even before full peer review.

One Prompt, Hundreds of Proofs

A spokesperson for OpenAI told Scientific American that a single AI agent generated almost every result in response to one prompt. That stands in sharp contrast to the company’s earlier Navier-Stokes solution, which needed a 10,000-strong agentic swarm and millions of dollars in compute to pull off.

If the claim holds up, it would mark a genuine leap in efficiency, not just capability. But OpenAI’s track record for bold announcements paired with limited transparency has left plenty of mathematicians unconvinced.

Why Mathematicians Aren’t Celebrating Yet

Andrew Sutherland, a mathematician at MIT, has urged caution. He says claims about “one-shotting problems with a single agent” cannot be verified without the model itself being released and the results independently replicated. OpenAI’s own spokesperson has conceded that some of the 372 results likely took multiple attempts, not one clean shot.

The backlash over the Navier-Stokes claim pushed OpenAI to form an independent advisory group tasked with setting guidelines for how AI-generated math results should be published responsibly. Those recommendations call for three things to be made public:

  • The model used
  • The exact prompt given
  • The total compute time spent

So far, OpenAI has only shared average compute times and partial statistics. The specific prompts remain undisclosed. A company spokesperson told Scientific American the team takes the guidelines seriously but is not bound by them, a distinction that hasn’t done much to reassure critics.

A Field Split on What to Do Next

The advisory group singled out AI companies that rely on proprietary internal models not available for outside scrutiny, which is exactly the setup OpenAI is using here. The company maintains it is working to release the model “as quickly and responsibly as possible,” though no timeline has been given.

Mathematicians themselves are divided on whether that even matters. Daniel Litt at the University of Toronto sees real upside in releasing the results now rather than sitting on them, arguing there’s no good reason to keep answers to math questions secret. Terence Tao has taken the opposite view, criticizing what he calls the “insane” pace at which frontier AI labs are now generating unverified results. OpenAI has acknowledged that many of its own researchers don’t yet fully understand some of the 372 new proofs. The company shows no sign of slowing down regardless, treating unsolved math problems as one of the clearest tests of how far its models have actually advanced.

Hashlytics Take

The number that should give anyone pause here isn’t 372. It’s zero, as in zero independent replications so far. A verified Lean proof confirms internal logical consistency, not that the process used to get there matches what OpenAI is describing. Until the model, the actual prompt, and enough detail to replicate the run are public, “one agent, one prompt, 372 results” is a marketing claim wearing a math paper’s clothes. The advisory group OpenAI itself commissioned said as much, and the company is choosing to publish the headline while withholding the receipts.

Follow Hashlytics on Bluesky, Facebook, LinkedIn , Telegram and X to Get Instant Updates