There's a decent chance that by the time I publish this, the discourse cycle will have moved onto the next thing. Things to think about happen faster than I can think about them. It's one of those years. Last week we had the rare occurrence of research math escaping academic containment. Then a number of AI leaders signed on to a letter from Dario Amodei (CEO of Anthropic) advocating for creating a regulatory framework to slow the pace of model development. I could write a lot on both of these.
Let's start with math, because that's where my heart lives.
OpenAI Publishes Solution to Navier-Stokes Millennium Prize Problem
Disclaimer: I don't have any stake or specialist knowledge in PDEs or mathematical physics. In a past life, I did algebraic combinatorics. Please don't ask me to solve a differential equation, because it'll take me a while.
Exactly one week ago today, OpenAI announced that they used a swarm of LLM agents to generate a solution to the Navier-Stokes millennium problem. There are good explanations of what this actually means out there already, but the short version is
-
The Navier-Stokes equations are a set of partial differential equations (PDEs) used to show how motion in a viscous fluid evolves given an initial state. Given that state, they spit out a vector field that describes the velocity of a point in the fluid at any time after.
-
Mathematicians are troublemakers and like to try to break things. Given some sort of statement, they want to know if it always holds or if it's possible to quantify where and how it breaks.
-
Mathematicians like pretty things that can be described in the most general terms.
-
Smooth (infinitely differentiable) functions are well behaved and well understood. A smooth vector field also can correspond to the behaviour of fluid in the physical world.
-
A nonsmooth vector field contains undifferentiable points (singularities) where things go haywire. In the case of Navier Stokes solutions, such singularities correspond to points of flow approaching infinite velocity. These are often called 'explosions' or described as the field "blowing up." This doesn't correspond to the physical behaviour of fluid, unless you've ever stirred a pot of water and had a tiny stream shoot off into space at light speed, in which case I'd love to hear from you.
-
The Millennium Problem about the Navier-Stokes equations is (with some handwaving):
- For any well behaved initial state in three dimensional space, show there exists a smooth vector field that satisfies the Navier-Stokes equations corresponding to that initial state, OR
- Show that there exists some initial state for which a solution to the Navier-Stokes equations blows up.
This is an old problem. PDEs are notoriously hard to solve -- they only have closed form solutions in rare cases. In most real life physics applications, systems modeled by PDEs are solved by approximation algorithms. This is good enough for engineers, but underwhelming for mathematicians.
OpenAI (claims) they've shown the second thing. This is demonstrated via Lean code (a proof verification tool) and an 166 page LaTeX paper generated by the AI. Cool.
For math to be 'accepted' by the academic community, it has to be read and reviewed. Often this is through academic peer review before being published in a journal, but some results are so immediately potentially important that their pre-prints are scoured over by peers independently of a journal review process. As leaders in the field convince themselves that the result is valid and examine the papers' methods, that confidence trickles down across academia, and it becomes a known result. Sometimes this aligns with journal publication, but it's more of a social phenomenon on the distribution of knowledge and how you ascertain truth if you're not directly engaging with the material. This is something that takes time. It can take years.
OpenAI's paper hasn't been through that process yet. That's not to say that it's not true or not a solution -- between the Lean verification and no experts calling bullshit (that I'm aware of), it seems likely that it will hold up. It's just that verification of the result is going to be a long journey. It might be a particularly long journey if the paper is 166 pages of AI slop (166 pages is a very long paper, regardless of slop content). I don't understand the math well enough (or at all) to comment on the quality of the paper writing, but I'm super curious to hear what some experts think about how it's written, if they can follow along, and how it stacks up compared to human authors.
On the other hand, the math community does have a history of bending over backwards to understand terribly written papers if it seems plausible that they address an important unsolved question (see: Mochizuki and the abc-conjecture saga). Does that grace extend to AI authors?
I think it might, but in the case of AI it's less charity and respect and more that there's so much research and knowledge that is still up for grabs. Math research is a neverending process of streamlining and rewriting. How you construct an argument can have just as much value as the final result -- sometimes the internal mechanics lead to more research than the proven statement. My own thesis was constructed by looking at an existing result and thinking, "I can demonstrate this in a different (in that case, explicitly geometric) way that will generate value to the field." Research mathematicians at OpenAI have been open about not being experts in Navier-Stokes research. With experts around the world reviewing the paper it's assured that along with simply trying to understand it, they're looking at if anything can be simplified, changed or otherwise used in independent research.
This is a process that happens with any result but is particularly relevant to the release of an AI proof because it was built without the iterative knowledge that usually creates a math argument. Usually, a research paper building to a large proof is a bit like a car engine. You build it bit by bit by crafting and understanding each part, and that understanding lets you put it all together. By the time you've built the engine, there's not a lot of mystery to it. Meanwhile, an AI paper dropped in your lap is a bit like if you're a mechanic and a completely new type of engine you don't recognize is dropped in the shop. You know you can turn it on and make it run, but you have no idea how it works and how to maintain it. There's a lot to be learned and documented from taking it apart and reverse engineering it.
In short: there's still a lot to learn here. The problem might be 'solved,' but there are more open questions now on its inner workings than there would be on a human led solution. There's a certain amount of hand wringing going on about AI ending pure math research, but I think that belies an overconfidence in AI, a lack of imagination in posing research questions, and a misunderstanding of what math knowledge actually is. The end result was never the value driver.
I've also seen some thoughts along the lines of "AI proofs might drastically affect the profession, drive down our research value and lead to our departments being minimized." That would be a bummer. It also seems to be an endpoint of every department at a university other than computer science when research is funded based on economic return rather than knowledge, so I don't think that's really unique to the field. The English Lit department says hi.
Math is a liberal art. Get at me.
There's an apocryphal story about Gauss and a student. The student stops showing up to class, and Gauss asks a colleague if they know what happened. He's told that the student had a change of heart and has abandoned math and decided to become a poet. Gauss replies "That's a shame, but I didn't think they were creative enough to be a mathematician anyways."
More quick thoughts:
-
The lack of references in OpenAI's paper is a legitimate criticism. On publishing, it had 16 references, which is ludicrous for something of its length. I guess it's not impossible, but it's extremely unlikely. As of writing, it is now up to 22 references so I guess they're adding more as they find them. It's still crazy low. AI is not beating its reputation as a plagiarism machine here. Somewhere out there a grad student is going to go through the thing line by line, find every missed plausible attribution and dunk on the company.
-
On the other side of things, I think the drama with Buckmaster and Levent is overblown. Not to say that OpenAI handled it well with them because they didn't, but Buckmaster's statement has a conspiracy brained tinge to it that doesn't do it favours. OpenAI has added progressively more disclaimers on their page that their internal model was not trained on any conversations from the Buckmaster team nor were those conversations spied on, but I don't think those were particularly plausible scenarios to begin with.
-
OpenAI used millions of dollars of compute on this so even with the Millennium Prize (that they're not advocating for) they'd be profit negative on this. Still, their spend is orders of magnitude less than the combined prorated salary of every mathematician who has spent time on the problem over the last century.
-
Yes, I am aware that this can all be summarized as an OpenAI marketing stunt.