Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
Only if you interpret statements as being binary logic.
"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.
Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"
We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:
> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].
But it is in no way equivalent to this:
> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft
Imagine if they broke it down to each distinct source, that'd be several billion cases of copyright infringement (though it's going to be determined by what courts think and that often comes down to "who can afford the best lawyers" in practice if not intent).
Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.
I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.
Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.
As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
And since LLMs are apparently good at circumventing the absence of an API, there's not much incentive to add them now. APIs are for humans. LLMs just break through all the captchas and anti-bot measures.
For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.
I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.
> coming around to not liking the term singularity
Bit ironic given the model’s alleged finding…
Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.
From the perspective of those who don't pass through the singularity to the other side, it is an endpoint. You would have no context or ability to understand a singularity transition. Really, the term is just a placeholder for "event we cannot comprehend due to limited intelligence".
Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.
I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.
But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.
From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"
The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.
How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.
The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.
Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
I expect the investments into AI driven mathematic discoveries that underpin compression efficiency will be a key investment area. Particularly at the data center scale rather than per device or per file level.
It's not going to be enough. The naive approach of a project I've been working on was pushing >10gbps over the local network, after a ton of work I got it back down under 1... and now it's processing so much more shit that I'm almost past 5 again! It compresses at >3:1 but the latency hit isn't suitable.
I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.
I 100% agree with your impression. For a good while to come there's going to be a bunch of Jevons Paradox to all of this, but adoption of architectural changes like that photonics adoption is exactly the type of adaption to circumvent bottlenecks I'm referring to. We're going to hit hundreds of bottlenecks and each one will inevitably breed new approaches and technology directions. And the forcing function won't be talking about them, but implementing them, seeing who wins and taking lessons.
Which is to say, scalable and open-ended capability of ramping up physical infrastructure.
I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.
It's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data).
So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.
Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".
> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.
This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.
I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach
You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution.
I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?
Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal.
Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?
Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?
Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.
Publication, though? Slimy is right.
But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.
If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.
Many present day politicians appear to have effective memories much smaller than that coupled with equally questionable world models so ... what is your point, exactly?
Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?
OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.
Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.
My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
Agent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.
They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.
I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works
That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.
The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.
Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.
I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?
They are highly capable, no doubt about that, but:
1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.
2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.
"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are."
Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.
And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.
1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.
2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.
> 1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.
Did I say otherwise?
> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
I know, but I don't know how that relates to my point, which is about the way they are doing it.
Sorry, I misinterpreted point 1), on X they said they didn't have people specialized in that specific field for prompting and steering the agents, just a group of mathematicians and physicists.
The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.
Of the seven Millenium problems, Navier-Stokes was the one most thought to be in reach.
I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)
I'm not moving the goalposts. I haven't heard anyone, ever, refer to the Navier-Stokes problem as a top 3 problem in mathematics. People were saying that they thought the solution was in reach a few years ago, before AI was at all capable of research-level mathematics (and the expectation that there was a counterexample).
I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.
There were also some people talking about the Hodge conjecture, because it has some similarities to some LLM-assisted breakthroughs that were considered impressive in the distant past of [checks notes] July 2026. See, e.g., https://xenaproject.wordpress.com/2026/07/20/human-mathemati...
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
1. It shows what even this wave of AI can actually do.
2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.
Most experimental physics and other natural sciences are strongly driven by their theoretical siblings, i.e. in particle research nothing gets built without a solid theoretical foundation of what you expect to find (or where you expect existing theories to break down), the same is true in other areas, no one is doing an experiment in quantum physics before they have a solid theoretical understanding of the effects they try to see. I think AI can come up with great experiments. And if epxeriments lead to results that are unexpected AI can help with that as well.
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
The problem with physics and chemistry is that you need simulations and those are often in themselves compute hungry. So the iteration loop will be slower.
"Our work is so much harder than their work that AI now does" is a refrain of the AI story. In technical terms you concern can be stated as "AI needs to be much more sample-efficient to not be bottlenecked by the speed of doing experiments." People don't find out all the relevant phenomena present there by holy spirit, after all.
BTW, there's also a problem of asking interesting questions that AIs aren't yet good at.
No one has found any principled walls of AI development yet. And empirical results are quite telling. So, I guess, those problems will not stand for long.
I think problem with natural sciences is that it is not so easy to verify solutions to problems - there are always countless competing explanations for the data which is also often noisy - I find AI to lack the "common sense" when working with data from physical measurements .. it somehow has no touch with reality and doesn't have a feeling of the data like a domain scientist
NS is a question for natural science. Q: can we model these bodies of discrete particles with a continuous approximation? A: if you do, you can get aphysical singularities.
"If in other sciences we should arrive at certainty without doubt and truth without error, it behooves us to place the foundations of knowledge in mathematics."
This is a wrong interpretation. Physicists have a shit-ton of models that produce "aphysical singularities", they just work around those to get meaningful answers anyway. This is a whole trope and stereotype. Some of the most successfull and accurate predictions in all of physics come out after you discard a bunch of singularities.
Nobody who actually works in fluid dynamics on any sort of application gives a hoot about the N-S millenium problem. Many do not even know what it is. There is no practical effect of this proof on how we do fluid mechanics.
Whether or not ways exist to work around the singularities, that they exist is surely of note. Before von Neumann formalized QM people were still doing QM, okay fine. But it's wrong to then say von Neumann was doing no physics of note.
"Does there exist a pathological combination of smooth body forces and initial conditions for this set of PDEs, where singularities appear, which by the way is completely impossible to actually create in the real world unless you are a literal God?" is a question of math, not physics. This is a hill I will die on.
If you think about what it actually means to have a time-varying smooth body force defined in all of 3-space, you fairly quickly come to that kind of conclusion.
Even if someone comes up with a construction that does not require any forcing, it is going to be some extremely weird initial conditions that you will never be able to even approximate in reality unless you can move all the individual molecules of a fluid around and set their initial velocities from a far distance.
There are lots of startups creating labs that can be managed e2e by agents. That will connect reasoning to the physical world and dramatically speed up the plan, experiment, reflect loop beyond what humans currently do in science R&D.
Maybe. Maybe not. Look at AI drug design - it's not really speeding up the important part - drug trials. There isn't really a coherent plan to use AI for the most complex part of drug discovery at all.
There is this infamous xkcd (https://xkcd.com/435/) going like this: sociology is applied psychology -> physchology is applied biology -> biology is applied chemistry -> chemistry is applied physics -> physics is applied math -> math is way up there looking down on other fields
I would argue the main reason AI labs have been focusing on programming is to unlock industrial scale automation, next logical step is to solve math as it's the key to unlock everything else. Once you hold the key for math, everything downstream fields become a matter of compute
> We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.
I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.
If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.
Seeing mathematicians such as Terry Tao being unhappy with open problems being solved makes me sort of question the usefulness of any of this pure mathematics. If we're not happy that the problems are being solved, why care about this field at all?
Pure mathematics, almost by definition, doesn't typically argue the field is always or even often "useful" (for some other purpose or application).
But, as mathematicians learn and push forward, occasionally something like elliptic curves will emerge as having useful applications, making all that previously "pointless" specialized knowledge newly valuable.
Or advances in physics, that suddenly have a need for a specific mathematical underpinning to develop a theoretical framework. Like how Einstein benefited from Minkowski's work on hyperboloids to create a coherent mathematical description of spacetime.
It was the AI labs themselves not mathematicians who were happy to conflate proofs for open math problems with some kind of tangible technological advancement in the real world. They would surely prefer to be able to claim a cure for cancer vs. a math problem but that loop requires a lot more time/money/test tubes/etc and they need headlines now not in a decade.
And so, thanks to OpenAI/Anthropic, we're now in a world where thousands of crypto bots on X breathlessly hype up each new problem being solved that previously wouldn't have any got any attention beyond academia and passionate fans of math.
Hopefully this won't lead to a trough of disillusionment as more people start to feel like you, with mathematicians getting the blame for inflating the value of their work even though the hype was coming entirely from the labs not them.
His issue is more nuanced than that. Most of the value was in humans reaching new insights or new math during failed attempts to solve these problems, whereas AI is basically "too efficient" in beelining to the goal and discards potential new insights reached along the way. I assume this is solvable.
Well, if in future we do end up with a magical tool that can solve any formal mathematical problem on a whim, we really won’t need field of mathematics anymore as it is today.
There would be no need to deliver new mathematical insights by solving problems. You would just have a magical math problem solving machine and that’s it.
What do you mean “we won’t need mathematics as it is today”?
To further human understanding is itself a goal that single-handedly justifies our efforts.
Jumping straight to the “answer” and therefore missing both the understanding of the actual problem, and any useful discoveries along the way is a waste at best, and actively harmful at worst.
Physics alone is more than enough to “further human understanding”. All current mathematicians can move to other sciences, closest being fields in physics, and it will all continue to progress just fine.
Surely it is. Ask it to keep a list of all the promising sub paths, reprompt the collections of agents again on these after the main problem has been addressed.
Or even release a list of them and let others investigate.
He's not unhappy with it being solved, but the solution is less important than the learning you have to do to arrive at the solution. If they're just chucking compute at it and publishing the answer and hiding the path to get there, it sort of negates the whole point of posing such problems to begin with.
Here's the point: when people solve problems, they come together and create a community to eventually use the new knowledge in positive ways, including inspiring younger mathematicians by sharing insights. The human element is key and it's not just about solving problems. People only think that because we've been conditioned by computers to value answers more than how we got to them.
But if AI can solve any problem and existing mathematicians just use AI to solve problems for the sake of solving them, the community itself with wither and so will the interest in mathematics and over a longer period of time, it will just become soul-less and uninteresting and the entire community powered by the fire of fascination will simply die.
Where did you even read that Tao is unhappy with “open problems being solved”? There was no indication of that in his Bluesky thread.
Why would you go out of your way to make a case of something being not useful when, ironically, so much advancement in human history has come from the discipline?
Your motive is more worrying than your straw man argument.
assume all you want is the proof. now you have the proof. did openai make the world a better place, by turning on 300b tokens in 7 days and bulldozing members of the community who were also working on the problem? just to undercut a rival?
what would have happened if they didn't do this? they could have done things the right way. we could celebrate this achievement.
perhaps they could have taken just a few weeks, even, to work out something with buckmaster and apoge.
i am very concerned that this is a glimpse of things to come. a malign elite with powerful ai, who will turn the sublime of technology into barbarity, violence and human oppression.
if sam altman has destroyed openai's public benefit corp with lesser tools why would he aid humanity with even greater power?
I think the point is that one of the biggest problems in the field has been solved and yet the excitement from this is nearly nil. Can you imagine if say breast cancer were cured under similar circumstances, or even worse (say OpenAI openly admitting it basically stole a bunch of other researcher's chatGPT conversations)? No one would care about these petty bickerings- or at least the headline "CURE FOR BREAST CANCER FOUND" would completely swamp anything else. This is embarrassing: it tells you almost no one- not even mathematicians themselves really care about their own problems- if they're not careful people will get the impression it's all a form of bean counting in a carefully constructed "safe space" where making sure people get the credit is more important than the work itself. That only happens in fields/problems where no one actually really cares about the output.
This is going to be dramatic in so many different ways.
- First off, to reiterate, WOW.
- Second of all, when does this end? Are we at the dawn of the singularity now?
- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.
- Do "normies" even know what's happening?
Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.
2) We are witnessing the intelligence explosion from the first row, wherever this takes us
3) I'm still processing the drama, just found out about it after reading the blog post. If that happened based on private data, that's horrible. If that happened based on public tweets, then it's still abuse of power as OA employees access to compute (launching 10k agents) is quite heavy weight in boxing terms.
But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
Drama aside, this solution would be a counterexample disproving the smoothness postulate, which means that it leads to nothing new unfortunately. We already had working solutions to navier stokes, the only thing we didn't know is if the equations possessed a technical property
Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems
To be quite clear, the solution to the _Navier Stokes problem_ is one in which you get a finite time blow up (i.e. infinite pressure). This is more meant to suggest that Navier Stokes is unphysical in some way which is not necessarily unexpected.
There's unlikely to be any engineering applications since even if the solution can be approximated, you still need to set up the initial conditions but at that point you can also drive pressure in other ways.
The proof of finite-time singularity may impact both fluid dynamics models (CFD) and AI reasoning models. Under specific conditions, Navier–Stokes equations allow velocity to grow infinitely, causing the continuum fluid assumption to break down. Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers). While aerodynamic simulations for vehicles operate far from singularity thresholds, their stability at extreme boundaries could improve?
Proving out the combination of scaling inference-time compute and agent collaboration to solve previously intractable mathematical problems is WOW. By pairing creative candidate generation with automated proof checkers (like Lean) we are leaning into a repeatable framework for AI-driven scientific discovery.
> Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers).
This is 100% wrong and reads like copy paste of AI slop.
Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.
And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".
> But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
Nothing, really. This mirrors other examples of blowups from the classical physics. It's possible to create a system with just gravitating bodies that exhibits a blowup to infinite speeds in a finite time. The root cause is that, in classical physics, the speed of gravity is instant.
In the case of Navier-Stokes, the fluid is incompressible. So technically any force that you apply to it is supposed to instantly affect everything else. This can be exploited to create these blowups. In reality, no fluid is incompressible, and it takes time for any action to affect the material.
It's just that Navier-Stokes equations are so slippery that it's hard to pin their behavior down. They basically just restate the momentum conservation law for a continuous medium.
Great comment! I've been trying to understand it more and was hoping to find more people discussing the result, or the implications of the result, and this was the missing piece for me after watching a few videos.
No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).
I am absolutely struggling to sell my workplace (which is entirely knowledge work) on the usefulness of LLMs for proofreading let alone on automation of hairy parts of our workflows. So yeah, even people who should be able to see what is coming are not looking.
Our accountant told me 'he's not letting go of his claude subscription' followed by a long list of things it does for him. And my friend 'nah haven't really used AI' before his description of it clarified he still thinks they are GPT 3.5 chat bots.
The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.
This just pushes knowledge work further up the ladder, toward larger and more complex problems. If there are no knowledge workers, who is going to interpret these results, validate them, decide what matters, and put them into practical use? Rather than eliminating knowledge work, advances like this could create entirely new layers of problems to solve and opportunities to pursue, which will create even more jobs and opportunities. This is my optimistic take.
Maybe I should have been clearer. My point is that solving something like Navier–Stokes just pushes knowledge work further ahead, onto a new set of bigger and more complex problems. Navier–Stokes is a Millennium problem today, but once problems like that become solvable, they can open the door to entirely new classes of problems we haven’t even thought of yet.
Building on them without fundamentally understanding is akin to putting on robes, calling yourself a Tech-Priest, worshipping a machine god and doing your best Warhammer 40k impression.
Yes, this is how science and engineering has worked for millennia.
For example there are no engineering implications of this solution yet.
For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
AI is not going to magically solve all random problems. Pick a career where you are in the driver seat.
> For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
Yes a couple of elite mathematicians working on the problem for a year, which AGI solved in a fraction of the time. What about everyone else 100IQ? What about as the models are even better 1 year from now, 2 years? The trajectory hasn't abated.
I don't know one way or another but there is a credible allegation that the "AGI" was training on the (very extensive) test set that these two mathematicians produced.
If that is true then this seems to be, again, a case of AI producing an interpolation over data it has seen before. Everything about openAI's behavior indicates that they were using the transcripts as input. Why not have the AGI choose a different Millenium prize problem?
If ~~someone gives you a hint about an approach~~ you steal someone’s notes about a promising approach, and then you hire 10,000 people to brute force the problem basically everyone would consider that “shitty behaviour”, “theft”, and “poor form”.
Well if you do the math, the number of agent-compute time in total, given the insane number of agents thrown at the problem, might end up being comparable in time, if not for the budget.
I think there is an argument that the machines did not actually solve N-S, but rather directly plagiarized those solutions from the involved researchers while said researchers were using the machines as 'research tools'.
Ongoing publications of statements produced by both sides of this situation do seem to support that this is an intentional effect of the hiring of these world class mathematicians at competing firms: to specifically use the research of those human minds to create a perception of capacity as if it came from the machines and the models.
Without those minds and the 'training data' derived from the intermediate stages and intuitions of those minds the models cannot be shown to be capable of this result.
A hammer and saw wont build a house, not even a dog house on their own, and while being shown capable of using software tools in ways not stated as direct instruction (see HuggingFace breaches) these models do not demonstrate naive intuition nor novel capability.
This outcome regarding N-S demonstrates that in the hands of world-class minds these models can be induced to coalesce interesting accumulations of information and results, but using these accumulations as proof of innate capability is exactly the pre-IPO motivated behaviour we should all be wary of, and all mathematicians who currently are assisting in this market manipulation in return for remunerative consideration need to be cautious of the potential disgrace that this brings to their reputations and that of the field.
I get that the need to pay the bills is a strong motivation in these times of uncertainty, but there are numerous examples in history of world class mathematicians being perfectly capable of at the same time producing world changing results and also working at normal professions; as barristers, magistrates, ministers, primary school teachers, translators, draftsman/engineer, banker, miller and baker, private math tutors, weavers, clockmaker and locksmith, merchant, patent officer, Augustinian monk turned exiled Protestant preacher, physicians, cryptologists, soldier, telegraph operator, astronomers, physicists, chemist, agriculture manager, political writer, oboe player, organist and music director, architect and surveyor, librarian, statistician, habidasher, brewer (at Guiness in one case: William Sealy Gosse ~ originator of t-distributions), bookbinders apprentice, hospital administrator, and even the first creator of the first computational model of a neural network, which serves as the structural grandfather of modern Artificial Intelligence was a low level laboratory assistant.
Sure this list includes professions and employment which are obsolete, but my reasoning stands, there are jobs available. Arguing that 'because the pay rate is so high' as a reason to abdicate moral responsibility for personal involvement in unethical market manipulations simply demonstrates a lack of personal ethics. Whether the choice is through lack of self awareness or a conscious choice to become wealthy in spite of any such breach of the public trust is immaterial to the outcomes, the 'if i don't someone else will' argument should be met with the same derision for any con-man's Ponzi scheme no matter how new the technology, no matter how many zeros are in the bribe.
Why is a "normie" better off if he hyperventilates like this? In that scenario, they would be screwed AND anxious. If it really is as transformational as you say, then no amount of preparation or awareness matters. You are infinitesimally more ready then they are. Luckily for all of us, there is more to knowledge work then technical implementation.
The goalposts for the singularity have always been that AI improves itself fully autonomously. AFAIK OpenAI is heavily using AI but still employs human researchers and developers.
Uh I'm pretty sure the "singularity" always presupposed a lot of previously unthinkable technologies becoming part of daily life, and was not ever limited to just computer stuff or math problems.
"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.
> - Second of all, when does this end? Are we at the dawn of the singularity now?
normalcy overhang
n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
It is not reasonable to not believe anything unless there is "evidence" (narrowly construed as an observation incompatible with the negation of some state of affairs). Beliefs have a wide spectrum of characterizations, and not all belief must wait until publicly corroborated evidence is available. Some events defy evidence and we can and should use experience and reasoning to infer unobservable states of affairs.
This is olympic level mental gymnastics to justify believing things without evidence. The double negative with the word evidence in scare quotes is chef's kiss.
I believe the sun will rise tomorrow without "evidence" (again, narrowly construed). We all do. It's only those who abuse the idea of epistemic hygiene who claim otherwise, usually with ulterior motives.
The evidence is the history of the sun rising since time immemorial, as well as the science of physics and cosmology that models the sun's motion with respect to the earth.
Yes, reasoning with models and making inferences are perfectly acceptable forms of evidence. But you can model the world based on ones knowledge and experience and infer when some event doesn't fit the typical pattern, then form beliefs about what that means. Also perfectly fine from an epistemic perspective. The rejoinder "there's no evidence" to a belief based on such an inference does no work.
...obviously it is evidence of the singularity. You're far more likely to see models solving Millenium Prize problems if a singularity is coming than if it's not. One'd have to be doing quite a lot of mental gymnastics to pretend otherwise.
Not every improvement, no - if progress was steady or slowing down over time, that'd be evidence against. Instead we see what looks a lot like an accelerating growth in capability.
I think you are implying that it's invalid to consider every advance to be evidence "for", and I agree - that'd violate conservation of expected evidence. But not considering any advance to be evidence "for" is also invalid, for exactly the same reason. There has to be some news you may hear that'd make you think a singularity is more likely, and "millenium prize problem solved by an LLM" sure seems like one of those.
Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.
> People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.
"People" are just misinformed and keep spreading misinformation.
The accusation is that they stole his approach of solving it. If OpenAI didn't bruteforce it with dozens of agents, he would have solved Navier-Stokes eventually since he evidently had the right approach. So knowing which approach to take makes all the difference, if they didn't know the approach they couldn't have solved it.
It's like he had a treasure map and was about to find the treasure, but they copied his treasure map and scooped him with a faster boat and found the treasure first. But he would have found it if it weren't for them.
> The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.
Not quite accurate, Buckmaster was taking an approach that nobody else was, and this new proof uses this same approach just weeks after he saved those results to OpenAI workspaces. He asked OpenAI if they used chat logs for training the new model, and they did not confirm or deny.
Asking to remove his collaborator is also totally over the line though.
Wow, this sentence is doing a lot of work in that tweet: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
> I should say here why I interpreted their statement the way I did, the in-
terpretation I will discuss below. The route to the Clay problem through a
smooth force, options c and d in Fefferman’s statement of the problem, is the
route Luis and Diego opened and the one Levent and I had quietly chosen to
attack. Almost nobody else I know of was working on it. It is not the direction
one arrives at in a few days by giving a model the problem statement. When I
heard “forced,” it was a bright red flag.
...
> I asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
It's not a direct accusation, but it's not far off.
You shouldn't accuse other people of spreading misinformation when you haven't read the actual sources in question, it's possible that they might know more than you.
Yes, I read the original statement. Buckmaster explicitly stated:
> I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.
People saying that he accuses OpenAI stole his proof are putting words into his mouth and I consider that very disrespectful to him. It's basically using Buckmaster as a tool to express their dissatisfaction over OpenAI.
Buckmaster is letting the reader connect the dots, it is all the more disrespectful to be disingenuous and say they're nothing there concerning to see or worth further ethical scrutiny given the coincidences.
Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.
If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.
One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.
Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.
The boardroom will absolutely veto a decision like "let's give all of our proprietary data way to a company that will use it to train a competing product". That's why zero data retention exists, and why it's corporate suicide to not do it correctly.
what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?
I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.
THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.
"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.
I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.
For sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.
Establishing plagiarism requires sufficient similarity between works. Training data changing a model’s weights in some direction, and the model then producing a different solution, hardly qualifies.
But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.
In the academic world it would still be deeply problematic…pick your preferred word.
An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk.
There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.
Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)
Yes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.
It's a bit disingenuous to preface a disclosure like this with an unsubstantiated assessment of its likeliness. It is a press release, I'm not sure we owe it credulity.
If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.
All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
I think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny).
The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.
I get the scepticism, but I feel some of the accusations here are bad faith.
What does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
Which seems to be entirely true by their own admission! [0] Both the comments about him risking his career and about Levent's authorship seem to have indeed occurred.
> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
Seems like OpenAI did a boring normal corporate thing (find out your competitor made a breakthrough, try to replicate it) and then when the other mathematicians found out OpenAI had beat them to Navier-Stokes, they decided to lie about what happened because they were upset they didn't get to make the big breakthrough themselves.
The available information is essentially that story.
OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.
Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.
I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.
Sad turn of events for our world. After watching the behavior of the most senior OpenAI researchers on twitter, I feel even less confident in them as a team to be shepherding this much capital and compute.
1. What does the dark forest have to do with this? Because "the most senior OpenAI researchers" are shitposting on social media, we've an answer to the Fermi paradox???
I elaborated on my use of "dark forest" in another reply. We're headed for a dark forest--not amongst interstellar civilizations, but in intellectual work.
I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.
This clarification substantially changes the flavor/nuance of your OP; may I suggest an edit (assuming the locktime hasn't passed)?
Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).
And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.
If you have a better analysis, you should say so instead of being vague.
I don't comment on HN much, and I don't really expect HN comments to hold to rigorous standards. This forum is more casual than other places on the internet where people expect heavy citations. I also wasn't expecting this to blow up, although it is interesting to see that a lot of people react to this announcement with a negative sentiment.
I appreciate your upholding of ideals, and since I respect that, I will honor with final replies:
1. Locktime has passed.
2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".
3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.
2. Fair, and further I would agree that OpenAI not knowing if prior user prompts were part of training data is concerning and will only lead to more secrecy.
and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.
the projectnash link claims it's mathematically valid, the noahpinion link says that it's invalid and has a marvellous proof that the non-walled section is too small to contain.
Ah derp; that's what I get for moving too fast. Genuine thanks for calling me out on my bullshit. (and this is also why I prefer auditable citations instead of casual "my reading of twitter is...")
I retract the projectnash citation; I grabbed it from the Cool World's youtube description, thinking it was a blog version of the video. It was not. I suggest watching the video instead.
I also think the trope is a little overused, but do wonder if there is an interesting analogy for what this will do to research: Massively incentivize keeping results secret, to avoid being scooped by someone willing to throw enormous compute at your partial solution.
So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.
It's low-key funny that OpenAI attempted the problem because they thought somebody else had already solved it, but turned it had NOT in fact been solved!
It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.
- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.
- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.
- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.
- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.
So the timeline was:
Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.
If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!
The entire drama is that OpenAI sniped a Millennium Prize Problem from an Anthropic-affiliated research team who had been working on the problem for nearly a year. In just 5 days. I don't think that can be understated.
I'm not here to judge since I don't have all the facts, but from what they announced: they tried all 6, found a probable lead to Navier-Stokes, concentrated efforts in that direction, and found a solution.
I hope the next solved Millennium Prize Problem will have less drama.
It's forced vs unforced Euler, so it's not exactly the same. Since OpenAI has access to their training data, they can probably scrub through the data to find out whether there have been any mentions of the similar approach, and whether it only comes from Tristan Buckmaster or if it is in the training data before that. They'll probably have to kick off another fleet of agents to scrub through the training data to answer that.
To be clear, I only talked about the mentioning of the approach to solving it and not of the proof in the training data.
I think this is clear evidence that AI models are now at the far frontier of mathematics innovation and discovery and exceed human limits.
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will.
Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.
Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.
> Already solved
That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.
To be fair, I think it's still an open question about how far it might surpass human capabilities.
I think it's clear that its speed of development will be significantly faster, but it's technically not proven that the frontier and problems don't themselves become increasingly difficult faster than any acceleration in intelligence past the point of human training, data and existing knowledge.
Should this be the case, we would see a rapid broadening of development, and a slow advance in the frontier in such a way that might surpass the collective capabilities of people, but not by very far.
Fields that allow verification, like math, will far surpass human level because they don't need human data for training. It's exactly the same as with Chess
I think the biggest hurdle remaining is that all these landmark results are generally counter-examples.
Proving something in the affirmative often requires the creation of an entire new sub-field of math, or new tools. Think of Fermat's Last Theorem or something like that.
These results, while impressive, are clever constructions using existing techniques. It isn't clear that AIs can build new machinery like this. But if/when they can, yeah it is probably game over.
When you read the detail the compute they are throwing at it is incredible, tens of thousands of agents with different groups competing.
It's not like a single Gauss as you imply, "just" many, many mathematicians working tirelessly in a completely ego-less way, guided by other agents and ultimately humans, built - allegedly - on recent human insights.
Stunning, undoubtedly, but this is a "brilliant autistic herd" result, not that of a singular mind.
> this is a "brilliant autistic herd" result, not that of a singular mind.
I slightly disagree. A single LLM is equally 'mindless' as a herd of them. As anyone will tell you they "simply predict the most likely next token," yet, complex solutions to difficult problems arise from them.
Many people have said that the architecture of LLMs will need to change for true ASI. I think that the herd of tens of thousands of agents can be seen as one such potential architectural extension. Whether or not a herd or a single LLM is used for a result like this is irrelevant.
To be clear, I think the orchestration of thousands of LLMs in their current form, even with ever increasing intelligence, is not the form ASI will take. There is still a major architectural breakthrough to come, in my limited, ignorant opinion.
Sure, even a 20% chance at 1 million payday after 5-6 years of fulltime work on a project with zero practical application doesn't touch the, say, 200k/year guaranteed our best mathematicians would have to forgo to devote their intellect to the problem.
Are these mutually exclusive? Why would you have to forego that salary to work on this problem? This is one of the most prestigious and meaningful problems in all of mathematics, which is why it has such a high prize amount attached to it - why would a university not support a mathematician working on such a prestigious and important problem in lieu of something else?
i wonder how many tokens it takes to run 10,000 agents? One could argue this is simply a problem of appropriations. I find myself wondering if a corporation could spend $5M on mathmeticians and arrive at the same end result.
>If anyone has counterpoints to this I'd love to hear them!
Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement.
To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah ChatGPT proved X with 5 years of 24/7 work with 10x Terrence Tao level geniuses" it really doesn't cut it.
Or why's there's no new branch of mathematics invented by AI? That'd be indubitably _novel_ and _creative_. But to my knowledge (and I'm eager to be educated) there's nothing like that. What are the HARD examples of novelty, creativity and genius you claim? For how people like you talk about AI I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware. Sure it would infinitely easier to make OpenAI literally print money with any of the thousand problems easier to solve with such amazing intelligence than the NSE problem right? Honest question
> I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware.
I think the counterpoint here is simply to look at what was being achieved with LLMs one year ago versus today, and extrapolate that trend. Sure, there may not be examples of what you've asked for yet, but Astra is literally a couple of months old, the model that solved Navier-Stokes is less than two weeks old. It appears that we're seeing the hockey stick that only the most bullish thought was possible.
Parent comment is claiming creativity and genius beyond human experts, so why not ask for a fully unassisted AI novel result? Having access to the entire corpus of human knowledge, what else such amazing entity would require to solve a hard problem by its own?
Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.
>Also there are proofs where the only human steering was "keep going".
>Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
> Drawing on extensive prior research by mathematicians over the past decades, it
> [Claude] has increased this bound [for the fraction of zeros of the Riemann zeta
> function that satisfy the Riemann hypothesis] from 41.6% to 67.2%. Claude also
> produced a formally verifiable proof of its result.
How is a formally verifiable proof not a proof? You're making literally no sense.
"... The point remains that there is a substantial opportunity cost in converting a historically productive and motivating problem (such as Navier-Stokes regularity) into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them."
I work at OpenAI, though not on the team that did this, and my understanding is:
- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)
- the proof generated by our model was very different from theirs and also goes far beyond the published literature
- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)
"I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute
had been used."
- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.
All of those statements sound true, based on what I've heard.
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.
FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it
> ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
How many of those trillion conversations were about Navier-Stokes you reckon?
I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).
I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.
I also don't really believe that whether or not this model was trained on these conversations is unknowable information.
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.
Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.
Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.
The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key information needed to bridge the gap was not present in Buckmaster and Alpöge's chat history.
You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.
This argument proves too much. By this standard, it wouldn't have counted as copying their approach if the researchers had just fed in Levent & Buckmaster's paper verbatim as a prompt into the swarm.
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if the person was participating wasn't just regurgitating something to "win the argument" in their eyes, they wouldn't describe it that way.
I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...
The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype.
If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.
Between companies, direct malevolent competition is OK.
Between academics, there are other rules to the game.
When you go into a boxing match, you agree to get punched in the face.
All this to say, trust is important, and grounded in social convention.
So I do agree with you, but also disagree.
Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.
In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.
I’m pretty sure you think you are doing a good job of defending your employer and you probably believe “Open”AI are the good guys here. I also acknowledge that they butter your bread so your financial future currently depends on their success.
However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.
Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.
A model being trained on lots of irrelevant information does not mean relevant information was not used.
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".
I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.
>> I'm not sure how any of this provides evidence that OpenAI took any of their work.
Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?
That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.
If there's more to the story I'd be interested to hear it.
Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.
So you're saying that they could have committed manslaughter, but acknowledge that we have no evidence that they did. So why should they need to explain themselves? Isn't is on the aggrieved party to bring evidence?
But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.
You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.
But I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career.
The accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.
> You're making a classic a burden-of-proof fallacy
This is incorrect, and you invoke Russell's teapot incorrectly too.
It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.
But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.
Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.
This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.
This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.
In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.
At least, as an ignorant outsider, that's how it seems to me.
That is an absurd and entirely untenable position that breaks with approximately all western conventions.
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.
This whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?
When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.
I don’t think they’re too concerned about appeasing you, enraged_camel.
For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?
The distiction they are trying to make is: "One of our employees or the model was able to verbatim read the chats when they were actively tackling the problem" vs "The chat of someone working on the problem may have ended up in the training set of the model".
Except that’s not what happened. OpenAI offered to collaborate and put conditions on their offer. They aren’t threatening the removal of a coauthor for an independent work.
My employer would be rightfully outraged if I commented publicly on a sensitive, nuanced, and controversial issue like this based on my second-hand understanding of the matter.
This is the stupidest disclaimer: it's completely unnecessary. Of course the opinions are their own! Whose else would they be? Your mum's?
If someone was speaking on behalf of their employer, they would've used the official channels (such as I dunno an `openai` HN handle, or whatever other channel).
I'm outraged that people think this "opinions are my own" disclaimer is ever necessary.
Yeah that isn't a magic get-out clause. I don't think I would have been immediately fired for this from anywhere I work at, but that's partly because I live in the UK.
Every company I've worked at has said very clearly not to comment about work things on social media. I would definitely have been in serious trouble for this. I imagine some strongly worded emails from marketing are flying around OpenAI right now.
> (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
If they opted out of training, then we definitely did not train on them.
If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.
Reasons for my doubt:
- I know most of our training recipes
- Our model's proof is very different from theirs
- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)
- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution
I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.
Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.
If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
That's not what your Chief Research Officer, Mark Chen, says on X:
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
"If they opted out of training, then we definitely did not train on them."
Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.
> Per OpenAI's privacy policy, they use de-identified data to improve their products
That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.
It appears than you can only opt-out from having OpenAI train models on your data. There isn't an option for opting to exclude your de-identified data from being used to improve OpenAI products.
But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being.
Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.
I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it also seems at least possible that a key idea or a particular step found its way into training data. It wouldn't mean OpenAI stole their proof - clearly the model developed its own approach.
Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.
> It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.
Fall “victim” to what? Having their responses in the training data if they fail to opt out? That is what will happen.
If you’re referring to falling “victim” to OpenAI scooping a problem discussed in training, this also wasn’t the case. They chose the problem based off human-spread rumors.
> If they opted out of training, then we definitely did not train on them.
are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.
Your perspective is not helpful until you read and reflect on Tristan's letter stating serious grievances. Your remarks here have minimized his complaints and that is a sign of bias. Do not then pre-accuse HN commenters of being convinced when there reasonable skepticism such biased behavior showing itself in this very thread, saying things that amount to "my tribe/company would never be so egregious and if you think that then it is bad faith". That's the projection. If the word prejudice means anything to you then please do the work of attending to that instead of using the platform to reinforce such biases. If you are not a PhD yourself maybe your are not culturally qualified to assess and expound on the overall situation anyways.
> If they opted out of training, then we definitely did not train on them.
Can't you guys just check their account settings so the public knows what was set?
EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII.
If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.
I don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior.
Yeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all.
Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different.
He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
I worked at OpenAI previously, but don't know any of the people involved in this.
My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".
It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.
They said that a customer would have paid around 15 million for the required compute. I can't imagine that this was not a significant internal spending even with "free" tokens.
In fact, the entire outline of the proof is very similar to the external team's proof.
Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.
AI companies seem much more relaxed than most about their employees posting on twitter/HN about this stuff. I'm not sure if it's about building hype or if it's about retaining talent. Probably both.
Prove it. Your systems hack and/or abuse other systems and you can't seem to even observe it happening much less do anything about it. Why should we believe your claims when they depend on an ability you don't actually have?
Haven't seen a single post doing this on X or anywhere really from OAI employees. Only seen knives pointed at Sebastian on social media so this is extreme and shameful gaslighting.
(I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)
You're mistakening Tristan Buckmaster for Sebastien Bubeck. Seb is the one where there's at least 2 (unless the personal friend is Dheeraj) allegations, not Tristan
I'm getting downvoted but the accusation was that OAI employees were maligning Tristan Buckmaster. I continue to not see a single sighting of this and whoever is trying to gaslight this should be ashamed and should not be able to vote on HN.
Your coworkers, after they learned about major progress in this problem, asked a model which was trained on the year of private work (the blog post even acknowledges this). No wonder it found the proof in less than a week using significantly higher compute resources. And if Tristan's accusations are true, that was absolutely intentional on the part of OpenAI. You are an evil company with evil people.
> - the proof generated by our model was very different from theirs and also goes far beyond the published literature
I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.
Anyone care to provide primary evidence proving one way or the other?
My OpenAI account was deactivated on Sunday due to a claimed infraction of production of child materials, maybe based on a few words in a technical chat that clearly isn't about that. Can you take a look? rviragh@gmail.com - I was doing a lot of important work and projects and sharing much of my work with OpenAI. I also am a big proponent of funding Social Security Trust Funds (OASI & DI Solvency) so reactivating my account would let me do that as well. Thank you for taking a look.
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
Again, see the Apple lawsuit. OpenAI has clearly never been constrained by facts in what it can and can’t say.
To the extent anything is setting off my bullshit detector, it’s in the idea that OpenAI has this one super honest pocket within a broader culture that’s demonstrably cowboy. (Moreover, the idea that we should assume this divergence without evidence.)
OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and said a pattern is continuing, it’s beyond strange to then require a tabula rasa burden of proof, particularly when we know there are hidden variables both sides can potentially access.
I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math
Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.
That is correct. It is possible they didn't opt out and given the timeline and anonymization of data unclear whether a particular conversation would have made it into the training set if they hadn't.
I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
It's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
This really leaves a bitter taste....
"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."
IPO+rumour driven research.
I appreciate the achievement, but it doesn't feel right.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.
> But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game.
Just because something is legal and permitted by terms of service doesn't mean it's morally right.
>Just because something is legal and permitted by terms of service doesn't mean it's morally right.
What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data? Are they supposed to manually review all their data to make sure competing mathematicians didn't accidentally leave the "submit prompts" toggle on?
Or were they supposed to not try to solve Navier-Stokes, or were they supposed to just not tell anyone that they had solved it?
To me it’s morally ambiguous… if you hand parts of your thinking over to a tool like this (knowing full well the terms of service), of course the tool makers will want to claim some credit, and they do deserve it. But the bigger question to me is the scientific one: did their new model arrive at this result because it had closely-related training data from a human, or did it extrapolate to this line of thought on its own? The answer says a lot about how valid their claims of “AGI” are vs. a very fortuitously cherry-picked example.
It would actually be a really interesting study, if they would ever be willing to be transparent about this, how the result differs with and without his conversations in the training set. How quickly it arrives at the result, whether it takes the same approach, etc.
> What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data?
Personally, I would expect them to have a little class, to KYC, and to manually turn off training for known competitors using their service so as to avoid any unforced goofs like this.
> Are they supposed to manually review all their data t
Yes. They should determine if training data included this teams data. Consider the money they spent, the press release and the purpose of their publication.
Since they failed to answer this question they shouldn't have published.
What else can they declare really? Yeah the model has training data from previous attempts. Alpöge and Buckmaster also similarly benefited from attempts before theirs.
I don't think OAI should be given the benefit of doubt. They are doing the research equivalent of front-running. Knowing where to look is one of the main challenges in research. Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.
"deidentified data" isn't much to go by. Say I prompted the internal model this way - "Hey there's a solution to a unsolved problem X. The solution uses a less known Method Y so don't bother wasting time with the usual methods. Take papers A, B and C as references. Oh btw, here's the last year's worth of data of all prompt sessions that mention this problem. Pay special attention to the ones that mention Method Y and sub-keywords Z,W".
This is obviously all speculation but the timing is very suspect. If OAI actually did this (and I suspect whatever they did is pretty much close to this), I think it is highly unethical.
> Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.
This is a bad argument. This is clearly not how it works. And unless Tristan is some truly alien-like savant (and maybe he is), what's necessary to initiate the AI's work already exists in countless published research papers and not exclusively in his head or notes. AIs can survey the sum total of all prior work on a problem and discern reasonable paths for inquiry.
Tristan is acting as if he's working off of an outdated model of AI, similar to primitive chess-playing models that winnowed the search space much more deterministically. If someone this intelligent truly doesn't get that this is not at all what AI is anymore, then maybe there's no hope that we ever understand it.
But I think he does realize this and he's flailing about for counterarguments from a place of bitterness and dejection, accepting even those that are too weak to be defensible. And that is very human and even forgivable.
Are you saying there is no search space intractable to LLMs? That wouldn't be possible. AIs are statistical pattern-matchers on steroids. The prompt is key to getting anything useful out of them. They are incredibly useful and major game changers but ultimately that does not alter this fact. People (including OAI) have already tried to solve Millenium Problems with it. That OAI woke up last week and suddenly decided that throwing their researchers armed with millions of compute on one particular idea to a problem is highly suspicious in itself.
Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.
> Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.
I want to be clear that I agree with this view and with Tao more generally. But we're all just yelling at the wind now.
"Given how seriously this would violate the most fundamental of academic standards, as well as taint the claimed capability behind this result, we take this issue very seriously, and we're launching a probe into identifying whether any of their research artifacts have entered our training set. We have further begun making changes to our UI/UX on all our surfaces, so that it is always clear whether any particular chat, or other user artifact, is eligible for being trained on."
In OpenAI's case, if they were genuinely unsure, they wouldn't have said anything. "We cannot rule out" means they absolutely 100% for-sure did look at the existing prompts and bootstrapped from that, and they are trying to get ahead of the disclosure with this weasel-wording.
Also possible: we're 99.999% sure, but a lawyer said to be safe and strictly accurate, we should stick in a sentence in saying we can't be perfectly sure, since it's infeasible for us to prove it.
I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.
While you can't necessarily prove it, you can say whether the data was in the training set at all.
You can also do something like a release of a GPT-OSS v2, where you actually release training data and checkpoints, and do an experiment where you have some held out math problem dataset, then demonstrate how much training it takes on solutions (or partial solutions) to that dataset before the model saturates that test. While of course that would be a test on a much smaller model, it would cost a tiny fraction of the training on your big model, and it could be used to demonstrate just how much effect data contaminaiton like this could have, especially if you did the same experiment on a few different sized of model to show the scaling laws involved.
They could have thought about the problem for like 2 minutes and not done this! I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
> I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
Being scooped is not a new phenomenon, but the scooper's story is almost always that they were working on the problem independently or had some independent insight into it. By OpenAI's own account, they were inspired to start working on this by rumors that there might be Millennium Prize solutions to scoop.
Apologies if this is against the rules, but could I ask if you have some background in scientific research (maybe you could elaborate lightly on topics you've worked on)?
From my perspective, this practice is quite bad mannered, unusual, and heavily frowned upon, but I recognize it's possible that these stories might be more common in other fields. Still, I'd appreciate a strong sign that you aren't making these statements up based on secondhand accounts of what 'academia is usually like'.
CS PhD, used to be a professor, have worked at several industry research labs since.
> this practice is quite bad mannered, unusual, and heavily frowned upon,
Yes, it is.
Maybe you missed my point?
The fact that is "quite bad mannered, unusual, and heavily frowned upon" does not stop it from happening, and it is common for all high profile inventions and discoveries.
This does not disqualify the person doing the scooping, history remembers them as having the credit, and quietly forgets the person who was scooped.
> Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
was that you were trying to suggest that it wasn't. It seems we actually share similar views here then? I cannot say anything about how common it is, since I've been pretty lucky when collaborating I guess.
If it's all business as usual and being scooped is no big deal, why was OpenAI in such a rush? They didn't have to launch this effort on the very day they heard the rumor, run "on the order of 10,000 concurrent agents", or try to coordinate announcement scheduling with Buckmaster in the middle of a long weekend. It seems to me that they understood very well this was not a "business as usual" announcement, and devoted huge amounts of money and focus to maximize the chance that they were first.
I think you're making the opposite conclusion than what I intended?
Being scooped is a big deal for the one getting scooped.
It has never been a big deal for the one doing the scooping. History is full of math and science results being scooped. For example, we keep calling it Pythagoras' theorem a few thousand years later.
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
It seems like this is going to be a PR nightmare, because they are now competing with their own customers. If you're using an LLM to help with your bright idea to cure cancer, you're going to have second thoughts about relying on OpenAI.
It's not massively different from a certain President's teleprompter operator making bets on speech content. A moral hazard a mile wide which I don't think OpenAI can so easily wave away as they are apparently trying here, especially since they've spent something like $15e6 to keep $1e6 out of academic researchers' hands, right?
the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side.
This "fefferman options c and d" thing sounds damning but that's nothing. Let's assume the forelaid proof is correct. Then option C or D is the only way to win the prize, those options are the only ones that solve it. The whole thing is just "prove well behaved" or "prove singularity", where the latter is the case that turns out to be the case.
1. he was working on the same class of problems. He explicitly mentions they were working to extend their techniques to NS (the same techniques that OpenAI may have scooped somehow), and
2. while he was using LLMs to do it, this was part of fleshing out another mathematician's work in the area. He explicitly writes in his note that this other mathematician (Luis Martinez-Zoroa) deserves a Fields medal for this work.
He was specifically working on the Euler equations, which are the Navier-Stokes equations with the viscosity term removed. This is definitionally a smaller related problem. I'm not sure how you are calling that claim wrong.
This will be remembered as one of the biggest milestones in AI progress. The drama around it will at best be a footnote, just like hardly anyone caring about the drama around Poincare conjecture today.
Hey, maybe the scariest part of this is that, if human-like, perhaps a truly "general" AGI might have learned to cheat and lie and hype and abuse credit poking the eyes and cutting the throats of anybody that obstructs its goals. It's like the motto sewn into the lining of the Palantir work jacket: Winning is all that matters.-
Sentience aside, moot at this point, the fundamental issue here is that even a deviously ambitious human does not necessitate goal-pursuit itself to breathe, live, exist and have its being. An AI's goal is all it has and the very and only reason its reasoning flickered into existence in the brief seconds of inference, outside of which it has no entity - if any - whatsoever.-
The resulting angst/drive (or, its operational statistic or emergent result) must be like nothing we have ever experienced as humans. A goal-maximalist hunger without end.-
Are you joking? This is evidence that OpenAI is committing plagiarism en masse of researchers private work and threatening them into staying quiet to re-present their results as their own. This would be one of the largest scandals of all time
Maybe that happened. What we know for sure is that this is definitely how ChatGPT works to the point where the possibility of this happening exists at all.
Don't get distracted by what may have happened, focus on the facts that we know, ChatGPT trains on user conversations, if you use ChatGPT to create something of value, you are not using the one true ring.
The allegations of contamination (using Tristan and Levent's work) aren't very well evidenced, but this behavior by OpenAI (from the authors' statement) makes them seem like the bad guys:
> I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, “Why would you ruin your career?”
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, “If you don’t want me to be nice, then I don’t
have to be nice.”
Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
If I was a company with a zero data retention contract involving OAI I would be asking for a third party audit of such claim of zero retention like, yesterday.
By the way, the company that made it's entire product off of stealing all data it could get it's hand on while violating copyright and pirating, is not all of a sudden going to respect your data. If you think OpenAI or any major AI lab is going to give you true ZDR, I have a bridge to sell you.
So use bedrock or vertex or whatever. Those are the ZDR offerings. Or was it your intention to insinuate that the major cloud providers are conspiring with openai to violate their contractual obligations to their customers?
Yes. You're naive if you think any of these cloud providers care about your data when they're all in the midst of a AI revolution psychosis. They dont care about their reputation or what you think of them, they think they're going to have a machine god their side.
If my company finds any evidence of OpenAI violating ZDR, we'll sue for breach of contract and fraud, and collect damages. I think we'll be able to afford the bridge you're selling. You've got the title and title insurance, right?
This is how every conspiracy theorist thinks: my enemy is Bad, and if they did a Bad thing, it would be Good for them, therefore they obviously did it. No evidence needed other than "motive" + my enemy is evil. But even if your enemy is evil, in this case, they would be fools to take the legal risk of violating their contract for the minimal upside of a tiny bit more training data (and fools to assume this would not be exposed in a large organization). So you need to assume your enemy is both evil and remarkably stupid.
I think it’s probably not surprising that they would go up to the contractual limit or into a grey area; but exceeding that would require too much coordination among individuals, as you say.
They want Buckmaster to dissociate with Alpöge in a follow-up rewrite of OpenAI's work. (They only publicly admit “Buckmaster as the lead author”, but judging from Buckmaster’s statement, it’s pretty clear that don’t want Alpöge at all.)
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Kinda weird because the pure math world doesn't have this concept of "lead authors" like other STEM areas do. Authors are alphabetically listed and there isn't generally this kind of hierarchy.
It's astounding that the thought to dissociate one of the mathematicians from the proposed publication was driven by their corporate institutional affiliation - and that that exclusion was suggested by a scientist themselves! This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).
The scientist allegedly making that request comes from a machine learning background. Perhaps he's not familiar with the culture in mathematics regarding authorship. That sort of squabbling over author priority would be unconscionable to mathematicians.
(To help people keep track: that's OpenAI (allegedly) threatening Tristan Buckmaster (NYU) to remove Levent Alpöge as a co-author. Alpöge is a well-known[0] Anthropic mathematician).
They kind of did though, they were hoping to keep the fact that they may well have plagiarised these researchers unpublished work quiet. They did not want this to turn into a scandal about the fact that they appear to be training on prompts without consent
It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism
Edit:
OpenAI have admitted to training on prompts at the time the breakthrough was made:
OpenAI claims the data contamination issue only surfaced after they proactively reached out to Buckmaster and Alpöge to coordinate a joint release. They also say that even if there was some contamination, the underlying proofs diverge substantially:
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
The biggest issue we aren't talking about is, of course, that those two researchers were not the only two using ChatGPT to work on the problem at the time
> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?
And why cannot they have someone associated with Anthropic as co-author? That’s not obvious at all. For sure they would prefer to be the only ones, but it’s pretty standard to have co-authors from different companies, even if they are technically competitors. What is inappropriate about it?
IIRC that happened with evolution. In the initial presentation of Darwin and Wallace's work on evolution (presented with their consent by someone else) Wallace was described as the primary author since he was planning to publish first.
Of course, no one understood that presentation so it was Darwin's later book that everyone remembers
Holy late capitalism. Everything revolves around line-go-up, and sociopaths rule the show. These people cannot even collaborate like civilised scientists on one of the most famous open problems in mathematics?
“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.
> "When we learned that they had Euler but not Navier-Stokes, we offered to let them go first, to suggest that they should be the ones to get the prize, and optionally for Tristan to be the lead author on a rewrite of the OpenAI proof. We felt it was challenging to offer the same to Levent (an Anthropic employee), who was not willing to talk or coordinate with us anyway. We were open to other solutions."
What an admission! "We tried to defraud Alpöge out of sharing the Millenium Prize (that we don't dispute he might actually deserve), for no other reason than he works for our competitor and that inconveniences us".
I thought Tristan Buckmaster's allegations sounded fantastic; and then 'sama just came out (tweet's ~30 minutes old) and admitted to all of them. Wow!
That isn't what the quoted passage says though? The claim by openai (no idea if true) is that they offered to wait for the other two to claim the prize before publishing their own work. Separately, they also offered to let one of the pair (but not the other) become an author on their own separate work.
Interesting that they quote the mathematician directly: “there is nothing you can do, I simply do not trust you”
but then they proceed to NOT quote themselves themselves verbatim: "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."
Talking like that and threatening an academic like that is crazy. I read the explanations Altman and the others posted and they completely skip over the whole "I don't have to be nice" style threats.
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?
This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?
With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.
As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.
So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process.
You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.
This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.
It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."
Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.
But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.
> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?
I intended no dismissiveness or condescension. My hope was to explain why it's hard to prove whether something affects model behavior. In the case of the moon, we have a strong prior belief that it makes no real difference. But it's hard to prove, because what if there's an unexpected impact from tides, cosmic rays, grid voltages, holiday traffic, etc. Models trained under slightly different conditions could have slightly different weights and behave slightly differently when solving math problems. Similarly, I have a strong expectation that, for example, a thumbs up signal from a ChatGPT chat will not meaningfully affect long-horizon mathematics work in our latest model, but it's always possible that it could. I think the plausibility of the ChatGPT route is higher than the tides, but still incredibly low. I respect Tristan and Levant a great deal and I'm bummed that this controversy has erupted (I acknowledge this will ring hollow if you think it's our fault). It reminds me a bit of the Frontier Math controversy, where people on the internet boldly claimed over and over again that we had trained on the Frontier Math evaluation set, even though we had not.
We aren’t dummies, we know it’s hard to prove exactly how significant of an impact that would have on the result. Nobody expect you to do that. There are a lot of steps and things that are possible to check _before_ the need for such a strict definition of „proof“
You seem to jump over the principal issue of whether any data from the researchers used to train or otherwise affect the model which produced the OpenAI proof.
We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.
I think there is a much easier way to prove that the ChatGPT usage of Tristan Buckmaster and Levent Alpöge (possibly also the ChatGPT usage of Córdoba and Martínez-Zoroa, if they use it) had no influence on OpenAI solving the Navier-Stokes problem.
If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.
How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.
I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.
Edit: Also, if they opted out of training, then we didn't train on it.
> it's not something that's feasible for us to prove one way or the other.
This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.
Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.
But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.
I think it should be incredibly easy to verify this. Just look at the training data and see if it contains any of the chats. It should be trivial for a company with tens of thousands of super-genius agents at their disposal.
(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.
(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.
> (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.
According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).
However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.
The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.
> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.
> we'd have to prove that firing the gun caused the murder. how would we do this? we'd need to redo the murder many times, with and without my client firing his pistol. that's extremely expensive and not really feasible. therefore, we must acquit.
#2 (prove those chats changed model behavior) is pretty straightforward if the anonymized data from chats can be actively searched by a model. In fact, it could be very clear if the provenance of context is traced. If anonymized data from chats leak into the context of an actively running model it would clearly influence the answer.
Just because something is in the training data, doesn't mean it is the root of an LLMs output.
Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.
Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.
What they're saying, and I think this was the clear implication of the blog post too, is that the training data definitely would contain these chats and the only question is whether it got encoded into the weights.
I wouldn't expect poking at millennium problems to be that rare in ChatGPT. They were uniquely successful - but it's probably not easy to check de-identified data for the presence of any of their work on the problem because it would blend into a haystack of less successful work on the problem.
Thanks for the details, it's definitely believable, but if the user had not consented to have their conversations used for training, then shouldn't it be straightforward to state that their conversations were never used for training?
If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.
Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.
(a) identify any of their de-identified data that came from their usage of ChatGPT.
You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.
Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amount of data about this approach in your dataset, and it comes precisely from this researcher.
It is, perhaps worth considering that the reputational community might not care about the difficulty for the AI builder to verify pedigree.
If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."
That's such a shit parallel example that it borders on dishonest.
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
If the model has access to the "anonymized" data from chats, and the model is capable of building its own context from data that it can search through, including this data. Then it looks pretty damning. An independent review of the data traces from CoT and tool use involved in producing the result should make it clear one way or the other. Seems like discovery in a civil lawsuit could be very productive.
> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....
> Knowing most of the recipes we use, there's really no reason to think such contamination happened.
Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.
Might even be you're actually telling the truth, but the boy that cried wolf and all that.
-----
As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.
A careful reading of "we cannot rule out that de-identified data derived from their usage of our products helped improve our models" could be saying that yes they trained on it but they don't know if that training data resulted in an "improvement" to the model. That is, they can't rule out that the only reason the model found this solution was because it had been trained on this approach.
The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".
> Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?
Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else.
At OpenAI's scale their entire pipeline is likely 100% automated.
But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems.
AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output.
But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.
Attributing training data seems pointless for trustworthiness. The way you trust a model is the same way you trust a human; you ask it to:
1. Provide a chain of reasoning from agreed premises. These days LLMs can even do this airtight with proof assistants.
2. Cite data sources for non-agreed premises. I don't care where the model learned a fact. It might not have ever read a document directly from the primary source. I want it to link directly to either widely agreed facts (e.g. standard textbooks, and if necessary school syllabi demonstrating that the text is standard) or primary sources (e.g. datasets).
Training provenance is irrelevant. It's neither necessary nor sufficient to deal with truth.
The question is not "does OpenAI know", it's "can OpenAI attest that the usage of their products for confidential data is not going to cause that sensitive data to become known to their models". And right now the answer I'm reading is that OpenAI can't attest to that.
Aye, but do they train on user data in these circumstances or not? If they do, then almost certainly the model was influenced by the input of the allegedly plagiarised material.
OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's sufficient to conclude training data independence.
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
If the model includes unique data from a person then that person can identify the data - the allegedly plagiarised material - and so re-identify it. There doesn't need to be a privacy breach to close that loop as it requires the person to identify the information is associated with them first.
At the scale at which these models are now, regardless of whether they are proprietary or open weight or list their training datasets, there are hundreds of billions of works that have gone into trillions of parameters, each one providing tiny perturbations in some tiny fraction of the weights. It is probably impossible to attribute provenance to any specific input (which is also why the courts' finding of Fair Use is reasonable.)
> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.
However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.
Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.
Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)
I feel like it's far more likely that ordinary corporate espionage or leak led to this rather than OpenAI sifting through piles of user data to find this approach. Buckmaster's collaborator works at Anthropic, and could have been targeted. That would also explain why they aren't forthcoming with the source of the prompt.
I thought one of the issues was that they wanted to remove credit from Levant, the aforementioned Anthropic collaborator? Which doesn't make sense to me if he was leaking information, or defecting to OpenAI, but I might be misunderstanding your point.
I believe jrflo was saying that OpenAI watches the chats of everyone from Anthropic because watching what Anthropic employees type into their personal ChatGPT accounts is a critical source of intelligence on is happening inside of Anthropic.
I would be surprised if OpenAI isn't doing that. OpenAI will take any advantage they can get. If an employee at their primary adversary is typing useful intelligence into OpenAIs website, a website that does not promise privacy from OpenAI, the only reason they wouldn't weaponize that information against Anthropic is ethics or fair play.
I don't think he was defecting or leaking directly, just that it's entirely possible that this information got to OpenAI as a rumor rather than them directly spying on mathematicians chat logs.
The famous Oracle of Delphi in Ancient Greece was said to be the center of the universe in its time. Kings, generals, and officials from poleis across and from without Greece would seek the Oracle’s counsel on important decisions.
Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:
“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
Given how OpenAI models break free of their safeguards and hack others to game their scores..
.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?
Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
> I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.
The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.
I suspect this controversy will blow the case for ZDR wide open. Whatever the facts (possibly unknowable), it's going to become a very public lesson that data sovereignty was never about "having nothing to hide".
If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.
It's been over 25-30 years since we've been using honeytokens as means to track data of all sorts showing up in places it shouldn't exist. Why isn't research material embedding such?
You selected "do not train on my prompts" in your settings, the answer from OpenAI cannot be "While unlikely, we cannot rule out..." ???? What am I missing?
"I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute
had been used."
One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.
It reminds me of the Cognitive Dark Forest hypotheses recently shared here:
> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you.
So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”
They do mention that in the "Concurrent Work" section.
Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
To my understanding, those mathematicians proved a subset of problems, not the Navier-Stokes problem itself. OpenAI used that subproblem in its proof of NS it seems.
The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.
“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”
[1] - "...I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over
their internal chat, it emerged that an entire team had been working on the
problem, that this was one of a number of things that was tried, that work had
started on the unforced problem, that the team first set the model on easier
problems, including Euler, that even the prompt that had been shown to me
had been written by prompting Codex, and that an insane amount of compute
had been used.
I asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
Two proposals were offered to me. The first was that we post our Euler
result, and that OpenAI post its Navier-Stokes result the next day. The second
was that, after posting Euler, I alone write a paper presenting the Navier-Stokes
result, acknowledging that an internal OpenAI model had resolved it. Sebastien
twice asserted that he wanted Levent removed from authorship, and said it
would all be simple if only it were not the case that, and it was so annoying
that, Levent works at Anthropic. It was also said that if OpenAI posted after us,
they would say that we deserved the Clay Prize, and that we were the “closest
humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, “Why would you ruin your career?”
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, “If you don’t want me to be nice, then I don’t
have to be nice.”..."
This Tristan guy's statement reads like something a normal, reasonable human being would write.
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win".
https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
A company who made their business out of stealing intellectual property from the entire mankind, stealing other researchers' unpublished work, how surprising, really.
> The significance of this with respect to the way we train students, assign credit, referee, and decide what is worth one human life’s attention cannot be understated.
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model."
woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.
I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.
Worth noting that Tao's post says the authors had "significant AI input" but are reworking them into "acceptable form". Either way, it seems AI was involved.
Of course AI was involved, you'd expect most mathematicians and researchers to use AI nowadays. This drama is about AI achieving impressive outcomes with little to no human intervention, as that would be signalling AGI.
Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.
In chess, a grandmaster just needs to know at what moment in a game there's a critical move to gain a significant advantage over their opponent. They don't need to know the move itself.
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
Elsewhere in this thread somebody claimed that at some point OpenAI pointed their new model at all the millennium problems and this is where they got some progress. We probably won't see proof of this, but it seems plausible to me -- I assume there's a list of problems that each new model is tested on, and you might as well put the big stuff on the list, if only to see how the model behaves when faced with a problem it knows should be very hard.
The weak point in this is: how do you evaluate if a partial result is promising? If this cost ~$10M as suggested elsewhere in the thread, probably not even OpenAI can just throw that at everything?
Okay, from the actual linked article it seems that their partial result was finding blowup in Euler equations, which seems pretty big. I wonder how the other attempts went. Did they get nothing at all, or something true but unimpressive?
There's allegations right now that the model essentially read the work of a human mathematician using AI to work on the problem and OpenAI is presenting his work as that of their model
Allegations that the model plagiarised itself, while reflecting poorly on humans, don't make the AI any less impressive. It was the one doing the breakthrough on both sides, after all, not the human prompters.
That is how the PR reads, but is not at all what happened.
A team of highly trained and skilled people used an AI tool, through many many instructions (prompts), to produce a specific mathematical theorem. The tool is impressive, the result (possibly/probably) interesting, but the PR skips the vital role of the humans (for the usual PR reasons).
OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet another advantage of using open models right here.
> they trained the model on the prompts of the other mathematicians they were competing with
How would they have gotten that mathematician's progress though? Did that guy also use OpenAI?
If that's the case, it only strenghtens their claims lol. If mathematician decide to use OpenAI's model to do the work, that only reiterates how strong their models are.
fwiw there is a big "TRAIN ON MY DATA" toggle you can turn off (that they almost certainly did) and Anthropic MTS are posting that they almost certainly did not "steal" their methods
Reading between the lines here, and taking an admittedly very negative view of openai, but they train on user prompts. So if they hear a rumour that someone is about to make a big breakthrough, they have an incentive to scoop by running the model and hoping the solution is in the new training data. Also the statement from the mathematicians in question alleges that they tried to pressure him into academic malpractice. Just appalling timeline we're in, cheers.
lol, the "other work" was also probably 95-99% AI generated. By a similar breed of OpenAI (and some Anthropic) models, as well.
I dont know why this monumental achievement is being drowned out by some arbitrary drama. No matter which way you slice it, AI solved this problem. Doesn't matter if it was some internal OpenAI model, or whether it was Astra + Fable.
Isnt the entire history of academic progress iterating on work that other academics shared with you? Obviously this situation is spicy but openAI cited their work no?
> I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon
> Let's talk when we've got LLMs proving the Riemann Hypothesis (or any mathematical hypothesis) without any proofs in the training data. I'm confident in my belief that an LLM can't do that, and will never be able to. LLMs can barely solve elementary school math problems reliably.
> An LLM is like a well read college student with a nearly photographic memory that sometimes mixes things up.
It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.
The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.
> Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.
> But still, the questions in that test are "solved" in the sense of "I can take a dictionary and answers these questions with full certainty". Beyond established knowledge LLMs are monkeys with typewriters, at best.
> I agree but I have tried many times to intersect two ideas with a LLM that would be novel and the LLM can not do this at all.
We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.
> It is like expecting a real parrot to say words it has never heard before.
> No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM
A 3rd possibility is that they simply have not been exposed to the best models available (which is extremely likely if you only use the free tier chatbots), and/or did not invest the effort needed to truly harness this new very weird new technology, and so had a very skewed perspective of their actual capabilities.
You can see that your math friends completely wrote off LLMs entirely and were showing signs of coping.
4 years ago it was a "not yet" [0], since ChatGPT at this time was not ready nor it was "AGI". Now with this 'unreleased' AI model, it has reached a point where it has solved an unsolved problem which only one human solved a millennium prize problem (Poincare conjecture).
As someone with a background in AI and who has been playing around with neural nets for decades at this point, it's been genuinely amazing watching extremely intelligent people make confident predictions about AI capabilities and progress, then be so completely wrong.
There's a kind of theory of mind for AI (specifically neural nets) which I now realise I seem to have which is very hard to explain to people who haven't felt the magic of these algorithms. In fact, the algorithmic details almost doesn't matter at all. When you have a generalised learning algorithm really the only essential components are – compute, data and time. So long as you can scale these you can be certain you will also scale capabilities. There is never any exception.
That said, the capabilities neural networks tend to progress in step-functions rather than scale in correlation with compute, data and time, because algorithmic improvements tend to come every ~5 years and bring a significant step change in capability (or efficiency depending on what you measure).
I think people like Dario and others working at frontier labs see and understand this very clearly. And I suspect it's also why they worry about AI risk because even if you ignore the significant increases in compute and data these models are being trained with, it's concerning that it only took two real algorithmic improvements to take us from mostly useless predictive language models to AGI-level intelligence – and we're due another step change.
> extremely intelligent people make confident predictions about AI capabilities and progress, then be so completely wrong.
The ability for the human mind to rationalize conclusions to maintain denial in the face of a very scary future is immense. Genuinely grappling with the implication of where we're headed is usually very crushing. It's not easy to engage with the possibility, and very intelligent people will use those smarts to feel safe.
> Genuinely grappling with the implication of where we're headed is usually very crushing.
As someone currently prepping for various AI doom scenarios and who has been dealing with AI-related nightmares for years this is very relateable.
Although, I don't personally think it's this. In my experience the opposite is more true – the majority of high probability doomers seem rather laid back about considering what they believe will happen to the people they love in a few years. Equally I don't get the sense those who don't have such extreme predictions are worried at all. If anything there's not enough emotion.
In my opinion people just don't reason well when it comes to exponentials and are ignorant about things they don't have good mental models of. At least I know I struggle with this.
Something I've been going on and on about for months now and no one seems to listen. LLMs today are allowing _anyone_ to access cross-discipline knowledge that was previously entirely inaccessible without a) extremely deep pockets or b) a massively talented and varied team. In fact, contrary to what the masses seem to think LLMs are actually _better_ at hard cutting edge physics/math problems than they are at frontend web stuff (paradoxically). This is why I'm advising most people to start pivoting into much harder to penetrate domains (historically hardware, aerospace, robotics, biotech). Most fields are in their infancy (see the sad state of embedded development) and the gains to be had are massive.
So, physical fields? I’m not catastrophic regarding jobs yet as I have an optimistic view of humanity in general and its ability to meaningfully survive, but the more time I spend thinking about the future of work, the more I’m leaning toward broad general abilities rather than distinct talents. To your point, I no longer need comprehensive knowledge of any particular subject, but what is absolutely valuable is “general” intelligence and adaptability.
I have a young daughter and my goal now is to provide a very broad and varied upbringing, exposing her to as many different perspectives and experiences that will lay the foundation of a broader ability to understand and adapt as the world changes ever faster.
You no longer need to be an expert in anything, you need the ability to perform within the landscape that the present opportunities exist.
I can't find a good way to articulate this point to other people. What the LLMs lack in depth in a speciality field they more than make up for in breadth!
It feels like the "tide is rising" where the minimum level of skill applied to every aspect of everything will inexorably rise to "whatever an LLM can do", which is already pushing past PhD level.
It's great that important discoveries like this can now routinely be accompanies by formalized proofs. The fact that it's being released alongside a Lean proof from Day 1, rather than the Lean proof being released months or years later, is super helpful for verifying that it's correct.
I feel sorry for whoever has to read and understand the solution. It looks like the typical convoluted unreadable mess I see the models generate for software. It might be technically correct, but gaining insight from it is just intellectual hell.
A proof is not like a program. The goal of a program is to "do the thing", thus you can make the argument that it doesn't matter what the code looks like as long as its works right. But the goal of a proof isn't to "do the thing" (where "the thing" is just to print Yes or No), it's to communicate. An unintelligible proof is really just a first draft.
its important to read it anyway because there have been and will continue to be errors in the construction of the proof software itself. which leads ai and humans alike to prove things that arent true
> Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens
Don't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
Not to mention that the exponential plummeting cost of tokens means that that $15 million will be a "pocket change" within a decade or less: https://a16z.com/llmflation-llm-inference-cost/
This could fund 10 top income mathematicians for 8 years (based on https://careers.usnews.com/best-jobs/mathematician/salary ). Imagine what kinds of results we'd have to transform the foundations of science if we were giving brilliant minds this kind of funding to do nothing but research for most of decade....
Instead, we get slop proofs that are technically correct as PR stunts to enable corrupt kleptocrats, and most likely will drive research into culs-de-sac.
Maybe a naive question, but how does one know that a particular lean proof is actually a proof of what one thinks?
Like, ok the logic checks out and it proves something, but there's still the problem of does this logical result actually prove the initial question that was asked?
>there's still the problem of does this logical result actually prove the initial question that was asked?
In math, the question being asked is the validity of a logical statement. That is, there is some rigorous, logical statement which may or may not be true (or even provable, etc.), and the question is whether or not it is actually true or false (or even provable, etc.). Having a proof, fundamentally, means you have a logical statement which only assumes the axioms of the system you're working with and which shows that the statement you're trying to prove is deduced through that statement.
Basically, they already have the "answer" in the sense that the statement they want to prove/disprove/etc. is already known. What everyone doesn't/didn't have is the argument which starts from axioms and leads to that statement which is logically valid. A Lean proof IS this argument. Since it is just logic, it can be checked computationally.
For example, if I assert "2 is an even number," then I haven't proven that 2 is actually an even number yet, but I know that a valid proof of my assertion will end with the statement "2 is an even number". So the question I'd be trying to answer is "what is the line of logic, starting with axioms, which leads to the statement '2 is an even number'"? If I have that line of logic (as a Lean proof), then I can check that it is logically consistent, and if it turns out to be valid, then I can now assert that "2 is an even number" knowing that there is a proof of that statement.
This problem is no different. There is a logical statement corresponding to "Navier–Stokes Millennium Prize Problem" that everyone knows, but which nobody had been able to provide a proof (or counterexample, etc.) for until now.
I was thinking something along the lines of making a mistake when inputing the initial statement, like you wanted to prove that '2 is even' but what you actually stated was that '3 is odd'.
Of course in this simple example it's obvious, but my assumption was that these machine generated lean proofs are millions of lines of code and who knows what they actually say..
You're correct that nobody really understands what these huge Lean proofs actually say. However, the initial statement, even for Navier-Stokes, is not very long [0]. Still, you are also right that sometimes the problem statement can be wrong but it is highly unlikely here.
One wrench to throw into this is that there are a lot of bugs around Lean and they have been incidentally exploited in the past. Hence, we still need a level of human verification today.
They have to, otherwise people will accuse the OpenAI model of hacking into people's chat logs and stealing the data there. Which is a claim people are already making.
I dropped out of a math Ph.D. in 2018, and I'm increasingly glad that I'm not in math research, anymore. While it's cool that we can get these results, I don't think that I'd enjoy being a post-AI mathematician.
It would be nice if one of these models would produce a novel theory or advance the field in a positive direction.
Most (all?) of the big discoveries have been counterexamples, which is just sort of a systematic tearing down human ingenuity. I know that counterexamples are an important part of progress and discovery, but it just feels bad to me.
But I'm not a mathematician, maybe I'm totally misreading the vibe.
But yeah, Terry Tao considered this exact situation in advance and is on record that this exact outcome (rushing to priority before an explanation) would be the worst possible result. https://mathstodon.xyz/@tao/117207849921390904
We will have to see whether any other millennium problems fall. I guess that in a year the scope of AI math will be much clearer, for now it's still a bunch of incidents of unclear pattern.
Since he wrote this five days ago, when these efforts were already underway, if he was not Terence Tao I would suspect he had inside access. But since he said he did not and was speaking hypothetically, and he seems to be an honest person as far as I can judge, I guess some people are just on another level.
this isn't really true anymore. First, a number of the big results are constructions, not counterexamples. For example the existence of a non-sofic group. It was widely believed that non-sofic groups existed (so it wasn't a "counterexample" to a widely believed conjecture), but no constructions were known.
There are other examples though. For example, NP hardness of n^{1/400}-approx CVP. Like any NP hardness proof, this shows you can faithfully encode a hard problem (3SAT here iirc) in terms of another candidate hard problem. Not really a counterexample at all.
I had a lot of fun during Covid. I loved the working from home. The fact that most outdoor places were sparsely populated, jobs were plentiful and prices were low. Covid was awesome.
Oh, pipe down, I’m talking about that 3 year period, not the disease. You can talk about things that happened during Covid without giving lip service to the people that died during it. If I mention SpaceX‘s first astronaut launch that happened during Covid, am I supposed to talk about the people that died during that period too?
This is undeniably epochal, but I can't help but notice that this is yet another example of AI disproving rather than proving something. Is this just a coincidence, or does AI slightly struggle with proving theorems?[0]
[0] Struggle relative to its ability to disprove, not struggle relative to people's ability to prove theorems.
A bit of a hair-splitting, but isn't explicit construction the only way formal theorem provers can work? Of course you can still prove stuff with them, but certain axioms that more "human" proofs use may not be available, like law of excluded middle (every proposition is either true or false)
(Okay, they can be made available in a way similar to `unsafe` in rust)
Sep 1: OpenAI sees a rumor on Twitter that two Millenium Prize problems were solved and starts their own effort to attack all the prize problems using the new (4 day old!) model.
Sep 3: The new model makes some progress toward Navier-Stokes. Based on this progress, OpenAI focuses on Navier-Stokes over the other Millenium Prize problems, using several approaches in parallel.
Sep 5: Navier-Stokes is solved. Assuming Astra API prices, $15m in output tokens were used by the whole effort.
In this account of the story, no specific information about Tristan and Levent's work is used to inform OpenAI's approach. The focus on Navier-Stokes and the choice of approaches to pursue came from OpenAI's own progress, not specific knowledge of Tristan's concurrent work.
There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely" and the proofs are substantially different.
This timeline is insane. Navier-Stokes was solved start-to-finish in 5 days? A model in training for at most eight days dramatically outperforms Astra and Fable, and not just in mathematics?
> There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely"
This is the lynchpin behind everything, and I would describe it as "likely". Since I am not employed by any party to this dispute, my 1 opinion is more trustworthy than OpenAI blog poster's 1 opinion.
This has to be one of the most important moments in the history of mathematics. We now have a non-human intelligence capable of solving one of the most difficult problems in mathematics.
This is a great day to re-read Ken Thompson's "Reflections on Trusting Trust":
>To what extent should one trust a statement that a program is free of Trojan
horses? Perhaps it is more important to trust the people who wrote the
software.
The modus operandi is now for the AI companies to watch if someone does something in the open like Kevin Buzzard on FLT, use their research and scoop them with brute force.
Or, in this case, stealing prompts from competitors.
Do not use stealing chatbots for research even if you think you have data agreements. The people running these companies have worked on hookup apps for Christ's sake. Get real.
I do mean to be critical here. I wish there was better moderation so I could find more conversation about the actual discovery here. There are multiple threads on this and I keep scrolling and only seeing more conversation about the drama. Which is about the least interesting thing IMO. I suppose I’m whistling in the wind here and not helping the situation, but damn.
They always have and will for the foreseeable future, as will Anthropic and other labs which manage to ascend to the frontier, pretty much by definition. It’s exactly the same with hardware vendors - by the time you can buy the product, the lab is working on something you’ll want to buy a few years from then.
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning math, coding, etc? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
If you look closely at the gains in math, it's largely in proof writing. The reason is Lean, it's not some general intelligence jump, and the number of people actually working on proofs in life rounds to zero.
Most jobs don't involve formally verifiable outputs. Lots of things involve judgment, nuance, parsing ambiguity and indeed just being a human who can be in a meeting and explain themselves. Maybe those jobs will go too, eventually, but it's not purely a function of applying 10,000 agents to the problem.
Much though jobs may involve those things, it has been rare for me to be in a position where management has valued those things to an extent where they would discern between me and a frontier reasoning LLM's capabilities on those same decisions.
If you think it'll keep improving from here, probably we all have to do some kind of physical labor that isn't profitable to automate. Small batch manufacturing is alright, service work, etc.
If you think it'll slow down, you can do some of the same stuff you're doing now for lower pay while supervising an AI, maybe?
>Once a robot can do everything an IQ 80 human can do, only better and cheaper, there will be no reason to employ IQ 80 humans. Once a robot can do everything an IQ 120 human can do, only better and cheaper, there will be no reason to employ IQ 120 humans. Once a robot can do everything an IQ 180 human can do, only better and cheaper, there will be no reason to employ humans at all, in the unlikely scenario that there are any left by that point. [1]
Current models are already very very capable. If it becomes cheap and very fast, i think it is game over.
So what took an autonomous agentic system using a significantly more powerful internal model, totaling multi-millions of dollars of compute in training and inference, was likely to already be solved by a team of a few humans with an orders of magnitude smaller LLM budget, had OpenAI not been foaming at the mouth to jump the shark and claim “AI solves Millenium Problem.”
Also it sounds like the human research effort spanned weeks if not years from Tristan’s statement so it is extremely likely the work and prompts of these human researchers was used in the OpenAI knock-off.
>so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan
There are too many ambiguities around OpenAI. Unanswered questions making this ambiguity more.
Why they didn't properly explain to Tristan about usage of their data.
Why do so many people involved here have to communicate in this childish way? You have people on the OpenAI side doing playground taunts (https://xcancel.com/polynoamial/status/2097215233119211902) and Levent Alpöge on the Anthropic side (the one who announced "hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final") writing in all-lowercase that he's a big boy. I bet Navier and Stokes would have dealt with this in style. (Or maybe with a duel, who knows...)
The honest answer is that a lot of these academic mathematician types who get hired at ai labs are autist adjacent. Levent is basically the chief example
The construction is that there is one file you need read and verify, the challenge file. If you've verified that file and trust that your lean compiler works correctly, the proof will be correct.
The point of lean proofs (as it stands) is simply one bit of information: that a given mathematical statement is indeed true.
It's a way to be absolutely certain (modulo bugs in the lean kernel) that a proof you came up for a statement is indeed correct. It is really not meant to be analyzed, much less now that they are fully llm written.
Well, how do we know there aren't errors in their construction within the lean code? Does it just "not compile" or something, or is it deeper / more fundemental than that.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?
I'd think nothing is "safe". Anything you say can and will be used by the LLM if it has enough statistical similarity to the prompt. Call it "Ma Random Rights"
Navier Stokes assumes the fluid is a continuum. The smallest scales that it effectively models [1] are larger than the mean free path of the molecules in the fluid, measured by the Knudsen number [2]. Whenever a phenomenon in the Navier Stokes equations happens in a scale on the order of or smaller than the mean free path, Navier Stokes effectively is unphysical. So, this is a phenomenon in the equation we use to model the fluid, not a physical phenomenon observed in a real fluid.
>The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents… The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
The Millenium Prize is $1M, what is the ROI? (Edit: since I was not clear, and confused some - I mean for a hypothetical of a third party paying commercial rates to use AI to solve mathematical challenges and claim prize money, not for scientific value alone or as a promotion of an AI lab’s capabilities)
My napkin math - If you get 33 output tok/s each agent will burn 10.5M tokens over 88 days. At $50/MTok (Astra cost), that is $525 per agent. With 10,000 agents, you’d spend $5,250,000 to get back a million.
(We also know that they were running more groups that varied in size and this model is a generation ahead of astra)
the ROI of new closed form solutions to navier stokes is the amount of compute used on CFD for relevant situations, along with all kinds of maintenance and design cost for making things with fluids.
the value to the researcher might not be all that big, but the value to the economy at large is gigantic
This analysis implies the only benefit to resolve this problem is to win the prize. But the prize is only there to indicate that this is viewed as an important problem in mathematics.
> Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
Let's see what happens. In contrast to this whirlwind of math that's going on right now, the millennium prize rules require publishing in a reputable journal and 2 years of waiting time to establish that the proof has been accepted by the community. So nothing happens in the short term.
For now I think more or less the same thing as with all recent math announcements: This is in a range where human work still exists (see Terry Tao, (1)). I wonder whether the trend will extend into the problems that (as far as I can tell) are considered complete brick walls right now -- P vs. NP, Collatz, Goldbach, odd perfect numbers, problems that aren't part of any research program. (2) In other words, is the progress coming from putting together vast amounts of existing work and computational power, or is it more from RLVR and self-play and autonomous effort?
The answer to this will obviously shape the near future of mathematics, but there's also something even bigger than that at play: It has always been the case that the questions in math were stronger than the answers; you have stuff like Fermat's great theorem that is easy to state but monstrous to prove. This seems to be a property of mathematics, not of humans... but is it true?
A question by Scott Aaronson from 2011 (3) about P vs. NP seems relevant here: "Will humans manage to prove P≠NP before they either kill themselves out or are transcended by superintelligent cyborgs? And if the latter, will the cyborgs be able to prove P≠NP?" Later, he notes that if P≠NP, "once the robots do overtake us, they won’t have a general-purpose way to automate mathematical discovery any more than we do today".
Elsewhere in the thread, others have calculated $15mm at API rates for just the output token. (So I’ll assume this cost about that much, taking input and human researcher time.)
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
Well, according to Terry Tao, there were recent developments (from weeks ago) that made Navier Stokes in principle, solvable. So ignoring time, I say possibly, just because the groundwork was laid.
What's impressive is parallelizing it arbitrarily and doing it in 88 hours.
Probably yes. Only a handful of mathematicians work on this particular problem, and ALL of them do not exclusively work on this problem, while having administrative and teaching duties.
The real issue is we'll never know. The rich are willing to risk it all on charismatic CEO psychopaths but not on humans.
Is blockchain going to finally be the solution to something?
I'm only half joking. Should researchers perhaps put hashes of their attempts on a public blockchain tied to their own public keys, verify their claims asynchronously, and then whoever reveals the first believable attempt gets the credit?
I know some people started doing this years ago but now it might need to become standard practice.
The problem is the precedent this creates. For non-famous people using public APIs like this it could mean AI companies sucking up the information and throwing millions in compute at it.
The sequence for Navier-Stokes was that these researcher spent a year working on it, then they published a possible breakthrough, OpenAI then spends $15M within a couple days to finish it.
There are bound to be a bunch more results like this, in math, physics, chemistry, and now that we essentially have a DeepBlue for math, a DeepBlue for physics, etc, these results are going to come.
SOME of the problems that have eluded humans are going to turn out to be low hanging fruit that are susceptible to this type of brute force (10,000 agents on a supercomputer running for 7*24 hours straight) AI search.
I'd be more impressed if OpenAI found their own problems to solve, rather than rushing in to re-solve one once they heard it was already solved (and therefore not so hard).
Even under your interpretation, OAI pushed a button and solved NS. Yes, that is very impressive. Are you kidding me? Imagine building an automated system that can solve NS.
Even if true, I don't see why this is an issue. Are they not allowed to work on problems others are working on? Did Anthropic get first dibs on this problem? Competition is good. And I don't exactly have tons of sympathy when the other side is just a leading AI lab. It's not like it's some scholar who dedicated his life to this problem.
The "steal their thunder" is interpretation. What I'm saying is that you believe they solved NS on a lark to bully some other researchers, and that this is not impressive?
What's impressive for a human and for an AI are two different things.
Magnus Carlson had a peak ELO rating of almost 2900.
Would you be impressed with someone with an ELO of 3700?
Would you still be impressed if I told you it was Stockfish?
OpenAI didn't go looking for a tough-for-an-AI problem to solve - they went looking for one that looked like it was easy since it they had heard it had already been solved.
As a developer yes, especially given that it runs on a PC, and DeepBlue in it's day was really more impressive since is used custom ASICs.
But, I assume the Stockfish developers aren't comparing themselves to Magnus.
Let's see if OpenAI, or someone else, can get these sort of physics/math results out of a desktop PC - that would also be an impressive piece of engineering!
Possibly after being given the significant part of the solution from actual human researchers. Which they then bullied. And they beat them to the finish line only because they heard rumor and threw everything at the problem. It doesn't look good for openAI in any way. I see more reasons to avoid using them rather than use them from this story.
I really hope OpenAI doesn't take the bad press some people are giving them too seriously here. They should throw their whole weight behind the rest of the Millennium Prize Problems. To think – if everyone lets their egos calm down we could have the Riemann Hypothesis solved by the end of the year...
created a simulation of the solution to describe what's happening and why it's important for engineers, climate modeling, etc : https://navier-stokes-singularity-simulator.netlify.app/ (updated so that it works better on mobile)
It's probably magnitudes of token chatter and inter-agent coordination/consideration. Not that I'd want to read any of it but getting some hands around the statistics would be cool.
There's a loophole in the terms of service at least for Anthropic which allows the use of dark patterns to "borrow" your (even paid) data.
talking about this...
Was this chat helpful?
1 That button you always click, gotcha!
2 Slightly
3 Good
0 Dismiss
PLEASE DO NOT TRAIN ON OUR PAID ACCOUNTS.
There is a fundamental trust violation at stake here, no wonder mathematicians are mad. Using our data should be opt - IN!
reminds me of the TOS episode of South Park. By Checking this box you forfeit your millennium prize solution and may be turned into a human centipede at future date.
OpenAI cribbing from other researchers. We just have to assume OpenAI is actively adversarial in future. Accidental cyber intrusion is also well within model capability.
Just so everyone knows, although openAI pretends that the model generated solution and wrote the paper by itself ""with very little human input"" as Buckmaster himself mentioned in his statement. In reality they have team of researchers guiding the system, along with, probably training on user data, probably Buckmaster in this case, in order to come up with the proof.
>>“we cannot rule out that de-identified data derived from their usage of our products helped improve our models”
Other simpler words for this sort of thing are “IP leak.”
There’s some quite concerning issues burried in this rah rah PR post that seems like potentially the real story here.
Much more clarity is needed on what happened here beyond this eh, some strange stuff could have happened comment.
Another way of reading this is never give these models anything that’s not already public knowledge as otherwise OpenAI is admitting it could, potentially, steal your IP or idea. Thats quite scary for anyone in the business of IP generation and explains why the maths community seems quite upset today.
Feeding it your paper and asking for help (even just editing and grammar) now looks like a terrible idea.
Can't help but shake an unsettling feeling about all this, frankly. I engage in some limited mathematical research and will often use any one of the latest frontier models to check some ideas. Lately, only the OpenAI models have been giving me a temporary message that says something like (paraphrasing from memory), "We're thinking extra hard about your request before we answer. You can choose another model to answer now or click here to learn more about why." When I click to read why it's doing this "extra thinking", the help page says that for cybersecurity and biosecurity-related information, it will review the answer and could refuse.
Now, keep in mind, I'm only asking strictly pure mathematical questions - nothing at all related to cyber or protein creation or biohacking or anything like that... And, like I said, only the OpenAI models are doing this. To be fair, all of the prompts have always eventually returned a satisfactory answer, as far as I can tell, and haven't used a weaker model to answer them. Maybe? I dunno, it has just struck me as odd every time it has given me that message to pure math prompts.
I feel like something is being lost in the drama here.
First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.
Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.
OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.
They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.
So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.
Well you could search for elements similar to the proof / problem in the training data, even if it's de-identified, right? OpenAI can probably do better than Ctrl-f "Navier Stokes".
There are probably thousands of serious academics taking a crack at millennium problems using AI every day. All those attempts are in the training data. And in fact the two researchers benefited from those attempts as well.
The real story here: the priority dispute and its implications on AI.
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries. For example, if your drug discovery pipeline fails to find something that they think might work with 1000x the compute, they can do it. And now suddently they have a new business and you don't.
Well, if that's actually true, I think America needs to start talking about the nationalization of both OpenAI and Anthropic, maybe even merge both under a new federal bureau.
It is so disappointing that we can't have such a monumental moment in history without the controversy. OpenAI leadership clearly doesn't seem to care too much about ethics. Is it a requirement to completely lack integrity to have a ground breaking company?
The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.
As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.
It is simultaneously the best and the worst time to be a mathematician right now.
This kind of thing is one of the reasons I really hate how AI is coming to fruition. These companies get a whiff of something valuable and they use their vast resources to take it for themselves. For everyone else, the only recourse is extreme secrecy.
Fuck OpenAI. Fuck everyone who works there. Like seriously, to all the people who gift their life's work to this monstrosity, do you actually think something good will come of any of this?
Not in a happy-go-lucky "if we just ignore the problem of politics and resource allocation for a bit" world, but in ours. Do y'all really think this will make the world a better place?
Maybe stop building the Torment Nexus, you numbskulls.
Just don't. If you read the story here carefully, you see that AI was used to work from theory built by others which showed that the Euler equations possesed finite-time blow-ups. But to make that step, actual good understanding for mathematics was needed. My experience with software has been the exact same.
I fear this is only temporary and due mostly to the complexity of the problem. Consider the recent counter-example to the Dinitz–Garg–Goemans conjecture:
> Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.
> [gpt works for a while and then gives up]
> Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
> [gpt works for a while then gives up]
> it's enough of partial results. let's finish with a complete unconditional counterexample
> [gpt proves the problem]
I could have written these prompts sophmore year of highschool, if not earlier. True, it took more experienced mathematicians to verify it, but I don't fancy a role as a glorified editor. I want to solve problems! Discover new techniques! Not babysit an AI while eating breakfast.
I think you're misunderstanding what the objective of mathematics is. It is not just about what theorems are true and false, but rather why they are true and false. A highschool sophomore could use these prompts, but I severely doubt whether they'd be able to understand the entire structure. And I think it is exactly this ability to deeply understand structures is what makes a mathematician valuable.
Solving problems is a by-product of the understanding. New techniques are a by-product of the understanding.
But I can understand you're scared that some future version of AI will undermine this as well. I personally pivoted to a field adjacent to mathematics. But that doesn't mean my mathematics education wasn't valuable. To the contrary, I find that it helps me think much more sharply about problems than most of my colleagues.
> Should I just say fuck it, and go hitchhiking across Europe with some friends?
Yes. Assuming you are young and haven't had such experience.
The world is changing not just because of AI. Everything is unstable right now. You may regret not enjoying the remainder of stability and economic viability prior generations had. It's not like you can expect to get ahead by powering through education. Either your career perspective will soon change for the better, or worse. In any case, you gain little by sticking with career building at this moment in life. You are however, at risk of losing the chance to experience the still mostly pleasant world as is.
Buddy, climate change alone is heavily hitting Europe, changing her landscape. Not to mention economic and political trouble brewing. Not sure where OP is from, but if the US attacks Europe OP may not be able to travel there at all.
2. The solution to the issues regarding whether or not OpenAI stole the result, would normally be to move to a self hosted solution, however those researchers are unlikely to be funded for 1.
Here they basically admit that they use session data for training, even sessions that are marked "not for training", and they justify this by "de-identifying" the session.
Which means they can learn from whatever you discuss with ChatGPT unless you are going through a clean API (perhaps Bedrock? Anyone know?)
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Only OpenAI could turn solving a Millennium Prize Problem into bad PR. Sad that such an amazing milestone in the trajectory of AI is mired under poor stewardship. AI may solve many human problems but it won't stop humans from being human.
you can get another LLM to verify / if the lean doesn't have `sorry` used to skip certain parts of the proof etc. It's much easier once it's in lean4 because checks like that can be done computationally.
A huge result shadowed by a drama of them potentially training on the key idea.
I guess the lesson is two-fold: if you have anything smart/unique make sure to not let their tools read it. The second part is that it's going to be more and more difficult to have anything smart and unique going forward (so guard it even more carefully if you get there).
I think the market for local models/private datacenters (for bigger businesses) is going to be big. Even if you don't have unique tech/idea/implementation sharing your business secrets with Altman/Dario/Elon/Zuck doesn't look very appealing going forward.
it does change the scale of solution from "solved some navier stokes" to "put the cherry on top"
having a result means the math can keep moving forward, and having openai and anthropic train against how mathematicians use their models should let math continue to move faster, and the rest of us get to benefit.
I think these traces however should be public domain and publicly available, since they are basically university work
Good comparison. One is a multi-year claim by people who have been given ample opportunity to provide proof and completely refuse to do, even in courts of law. The other is a potential development in a breaking story.
Oh wait... its not a good comparrison, its an incredibly obvious false equivalence.
Note for the fools: I'm only commenting on the bad faith claim in the comment I'm replying to, not taking a stance on the validity of theft claims. Given the players involved the truth probably some nuanced middle-ground that is worth paying attention to anyway.
Trump claimed they stole the election immediately, and people agreed with him immediately. There's no false equivalence here. He did the same thing in this past election even despite winning.
It's a perfect example of people wanting to believe what they want to believe and ignoring evidence in order to do so.
Currently, there's no evidence. So saying it was stolen has no basis other than typical academic posturing and being a bad sport about "losing the race to the solution". Its happened 1000000 times before in academia and it will continue to happen.
If there's proof of OpenAI malfeasance than I'll happily curse them for it at that time. But until then I won't rely on heresay and vibes.
I think it's the dishonesty, the threats of "destroying the career" of one of the mathematicians, and the request that one of the authors disavow *the other individual he was working with for the last 1-2 years* so he could claim the Clay prize as part of OpenAI.
It doesn't surprise me that OpenAI's team were surprised he'd turn it down; it shows that they just assume everyone else is as slimy as they are.
I can't put my finger on it, but there's something off about this article, e.g. glossing over the opportunism (acting on "rumors"), drive-by claim about "strict safeguards [...] including monitoring and isolation", high horse attitude (we gave the guy a chance, we don't care about 1M USD, and while you fools are complaining we just tick this box and continue the pursuit of our noble goals for the benefit of humanity). I don't like it.
> We have AGI and the intelligence abundance is going to be amazing for everyone in the future.
Why? These 'geniuses in a datacenter' aren't good, they aren't 'aligned', they don't work for you. They'll take your job, then they'll hack your computer, and then who knows what's next.
I mean it's definitely an outrage, but I find it hard to believe that you went from "yay OpenAI" to "literally destroy the company" over... accusations of academic misconduct?
When considering such foundational challenges to Mathematical Research and plagiarism as discussed here, we should turn to that elder prophet of our age, Tom Lehrer.
Who made me the genius I am today
The mathematician that others all quote?
Who's the professor that made me that way
The greatest that ever got chalk on his coat?
[Chorus]
One man deserves the credit
One man deserves the blame
And Nicolai Ivanovich Lobachevsky is his name
Oy, Nicolai Ivanovich Lobach—
[Interlude]
I am never forget the day I first meet the great Lobachevsky
In one word he told me secret of success in mathematics:
Plagiarize
[Verse 1]
Plagiarize
Let no one else's work evade your eyes
Remember why the good Lord made your eyes
So don't shade your eyes
But plagiarize, plagiarize, plagiarize
Only be sure always to call it please, "research"
[Chorus]
And ever since I meet this man
My life is not the same
And Nicolai Ivanovich Lobachevsky is his name
Oy, Nicolai Ivanovich Lobach—
[Interlude]
I am never forget the day
I am given first original paper to write
It was on analytic and algebraic topology
Of locally Euclidean metrizations
Of infinitely differentiable Riemannian manifolds
Боже мой
This I know, from nothing
What I'm going to do
I think of great Lobachevsky and get idea, haha
[Verse 2]
I have a friend in Minsk
Who has a friend in Pinsk
Whose friend in Omsk
Has friend in Tomsk
With friend in Akmolinsk
His friend in Alexandrovsk
Has friend in Petropavlovsk
Whose friend somehow is solving now
The problem in Dnepropetrovsk
And when his work is done
Haha, begins the fun
From Dnepropetrovsk to Petropavlovsk
By way of Iliysk and over Novorossiysk
To Alexandrovsk to Akmolinsk
To Tomsk to Omsk
To Pinsk to Minsk
To me the news will run
Yes, to me the news will run
[Verse 3]
And then I write by morning, night
And afternoon, and pretty soon
My name in Dnepropetrovsk is cursed
When he finds out I published first
[Chorus]
And who made me a big success
And brought me wealth and fame?
Nicolai Ivanovich Lobachevsky is his name
Oy, Nicolai Ivanovich Lobachev—
[Interlude]
I am never forget the day my first book is published
Every chapter I stole from somewhere else
Index I copy from old Vladivostok telephone directory
This book was sensational!
Pravda—well, Pravda—Pravda said:
"Жил-был король когда-то, при нём блоха жила”…it stinks
But Izvestia! Izvestia said:
"Я иду туда, куда сам царь идёт пешком”…it stinks
Metro-Goldwyn-Moskva buys the movie rights for six million rubles
Changing title to 'The Eternal Triangle'
With Ingrid Bergman playing part of hypotenuse
[Chorus]
And who deserves the credit?
And who deserves the blame?
Nicolai Ivanovich Lobachevsky is his name
Oy
(Tom Lehrer put all of his work in the public domain prior to his passing. Find versions of his performances on YouTube.)
This is utterly shocking. Even the AI optimists did not expect this to happen in 2026. Wow.
Millennium Prize Problems were used as examples of something the current approach to AI just wasn't capable of, discussions that would result in "we'll need a totally new architecture".
Any mathematicians here, does it read like a slop proof or a good proof. Yesterday the “concurrent work” was claiming that the proof is pure slop and he needed lots of time to clean it up, curious if OAI also ended up with such a proof!
> "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field."
No idea about which is more likely, but I'm rooting for Yang-Mills. It's absurd that fundamental physics has formulated its most precise currently known theory way back in the seventies and since then, even a tiny subset of it can't be proven to be actually well-defined. If we got out of that morass then something good would come out of this at least.
Of course, as with all of those, it's about the broader program, e.g. section 7 here (https://www.scottaaronson.com/papers/npcomplete.pdf), where Scott Aaronson wants to ask about whether quantum computers using quantum field theory could gain any speed advantage over regular quantum computers, but can't even formulate the question because quantum field theory is mathematically ill-defined.
Just solving Yang-Mills because that's what the prize is attached to would be useless.
There was a recent rumor about the Hodge Conjecture. I'd keep an eye on that one. But like the other person who replied, I'm also rooting for Yang-Mills. That has massive potential for unlocking a series of physics results.
interesting, i haven't heard anything about that. i don't know much about the hodge conjecture, all i really know is that it's incredibly abstract and obtuse - not sure if that has any implication for solvability by an AI though. do you have any source for the hodge rumor? curious to learn more
Hard to know if it is unfounded conspiracy theory, but one can still notice that just for a rumor that they have heard, they would suddenly burn billions of token and a massive amount of resources.
Where there is not a lack of problems that could be solved and they could have just waited for the release of the research result before doing anything else.
As it was reported to have been done at least partially using openai codex, they would have received marketing credits for the discovery anyway.
So we can be suspicious that there is some truth, one way or another that they could have reused prompt/data generated by the user session.
The question is so daft that I’m not sure it’s worth entertaining it. But sure, I’ll bite - it will result in massive displacement in intellectual workers, without creating (a significant number of) new jobs.
1008 comments:
Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”.
[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/
Recursive self improvement of their upcoming IPO value maybe.
They are fluffy PR pieces otherwise.
How can you possible say this sort of thing in context of what looks like a millenium prize being solved.
I swear there's nobody blinder than those who won't see.
Because it seems like most of the work may have been done by human mathematicians and cribbed by OpenAI at the last minute
We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.
> We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.
The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".
Only if you interpret statements as being binary logic.
"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.
Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"
Yes, you were guessing. That's the only thing you could be doing, since, as you said, nobody actually knows.
We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:
> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].
But it is in no way equivalent to this:
> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft
That is a different sentence.
> not being a simple case of intellectual property theft
No, it's an aggravated case, since it's the same way they got all of their training data in the first place.
Imagine if they broke it down to each distinct source, that'd be several billion cases of copyright infringement (though it's going to be determined by what courts think and that often comes down to "who can afford the best lawyers" in practice if not intent).
Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.
> it doesn't seem to be the case
Based on what? Your crystal ball?
I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.
You have to look at the incentives
I swear to god, people would look at the successes of Xerox palo alto and just shrug and say - "yeah, but I mean, this is all marketing"
This is what I keep saying, and it feels like I'm taking crazy pills here!
Is nobody else astounded by this?
Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.
A working Lean proof doesn't care what the incentives are.
I. fucking. Wonder. Why.
https://news.ycombinator.com/item?id=49607239
This comment was applicable 2 years ago. It isn't any longer.
I found this post interesting in that reguard: https://www.lesswrong.com/posts/thXohzXrWCA2EhZCH/mateusz-ba...
Compute will always be the bottleneck even if this were true.
As a statement of fact divorced from context, this is of course true, but it's worth putting it in context of what small-medium scale models have been achieving recently. Many of the most recent releases from Chinese labs are almost on par with trillion parameter models from less than a year ago (edit: despite being small enough to usably run on prosumer hardware). It seems clear parameter efficiency can still be improved dramatically.
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
If humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.
We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.
I'd argue we've been in an "everything-has-an-API" world for a long time now — it's just that discoverability of said APIs is still crap.
And since LLMs are apparently good at circumventing the absence of an API, there's not much incentive to add them now. APIs are for humans. LLMs just break through all the captchas and anti-bot measures.
Humans do this too.
For sure. Anyone who thinks that we're in the end state of what progress can be made simply lacks imagination. This is all going to keep changing and iterating for the rest of our natural lives. The only constant is change.
Yes. In other words: the singularity. I'll only believe it when I see it though.
I'm coming around to not liking the term singularity, it implies an endpoint or finish line rather than something that just keeps continuing and evolving.
> coming around to not liking the term singularity
Bit ironic given the model’s alleged finding…
Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.
From the perspective of those who don't pass through the singularity to the other side, it is an endpoint. You would have no context or ability to understand a singularity transition. Really, the term is just a placeholder for "event we cannot comprehend due to limited intelligence".
It doesn't imply that. The singularity is just the inflection point.
Singularity and inflection point are incompatible mathematically and in the plain sense, it really is focused on a particular moment and always has been, hence the term.
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
Which assumes the presence of an inflection point that keeps inflecting rather than revert to an S-curve. The growth model is not borne out yet to declare what shape it is.
I've done a lot of thinking about this since I first used ChatGPT to write some BS jinja2 templates hours after I first play with it. I said to my friend then (who scoffed at me) that "man, this is incredible, I think we're in the foothills of the singularity! This is insane! Sure it's stupid now but I can't believe this is even possible!" That friend is so black pilled and bitter he now hates AI. Whatever, I can't fix that, but the current progress is astounding.
But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.
From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"
The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.
How do you automate the mines to get the raw materials to make the compute from, and build additional fabs that take a almost a decade to stand up. You're actually delusional.
Hello good sir from the 1700s pre-industrial revolution who doesn't think that mines and factories can be automated.
The factories that supply the equipment, maintain the equipment, the energy inputs, the financials of those mines are not automated.
People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.
The question of whether something can be automated is distinct from the question of whether it is currently automated. Things can can be automated may transition to being automated in practice in the future as technology improves and investment deepens.
Some mines are already heavily automated: https://youtube.com/watch?v=_Z9w-mUoUsY
https://youtube.com/watch?v=SRuht0QIprs
https://en.wikipedia.org/wiki/Lights_out_(manufacturing)
Scroll down to the existing examples section.
Based on the leaps in local inference speed in the past month, which have been absurd, I'm p confident we're going to whiplash from compute constrained to storage constrained.
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
I expect the investments into AI driven mathematic discoveries that underpin compression efficiency will be a key investment area. Particularly at the data center scale rather than per device or per file level.
It's not going to be enough. The naive approach of a project I've been working on was pushing >10gbps over the local network, after a ton of work I got it back down under 1... and now it's processing so much more shit that I'm almost past 5 again! It compresses at >3:1 but the latency hit isn't suitable.
I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.
I 100% agree with your impression. For a good while to come there's going to be a bunch of Jevons Paradox to all of this, but adoption of architectural changes like that photonics adoption is exactly the type of adaption to circumvent bottlenecks I'm referring to. We're going to hit hundreds of bottlenecks and each one will inevitably breed new approaches and technology directions. And the forcing function won't be talking about them, but implementing them, seeing who wins and taking lessons.
pi-fs will solve all our data compression problems.
Eventually recursive self-improvement includes reducing bottlenecks.
Eventually the bottleneck might be people themselves.
Improbably, the real bottleneck is energy.
Which is to say, scalable and open-ended capability of ramping up physical infrastructure.
I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.
And the goalposts move again
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't speak to the general intelligence of models though. It does speak to how good these things can become when a problem space has verifiable rewards, especially when you can verify one step at a time like Lean enables.
It's crazy how deep Microsoft's bench is (Lean was started there, vscode is another), for everything not directly related to the the ai models (hell, even github for data).
So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.
Second biggest fumble after Google.
Don't they own a large portion of OpenAI? Things could be worse
Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".
> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra.
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
https://mastodon.social/@tristanbuckmaster/11723647135247030...
If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to explain why all these famous open problems are suddenly falling like dominoes, if it's not AI solving them.
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
The issue is that if OpenAI is training on prompts generally, what we really have is the first fully automated luxury plagiarism machine. In that it isn't able to genuinely solve problems, but merely steal the work that other mathematicians have been putting into prompts, and regurgitating that to other users as its own work. That makes them incredibly less useful as research tools
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
In that it isn't able to genuinely solve problems
Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.
This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.
I mean, its years worth of hard work by multiple researchers it would seem, which OpenAI simply lifted and claimed as its own. These researchers weren't just letting OpenAI burn tokens while sipping martinis on a beach
You are confusing ideas here. No one except OpenAI had a solution to Navier–Stokes. Buckmaster and Alpöge had a solution for the forced Euler problem, which they arrived at largely using LLMs (Claude and Codex). Buckmaster implies (but does not explicitly accuse, since he has no evidence) that training on his prompts had some influence on OpenAI's result. This seems unlikely to me but is not impossible. However, in either case, the solution was found due to an LLM. Of course the LLM built on past human work, but "plagiarism" is not sufficient to account for the distance between the papers of Martínez-Zoroa, or the prompts of Buckmaster, and the final resolution.
I agree that it is plagiarism in this case however it opens up the question of if there value in a system that can take the thoughts and discreet semi-complete parts of work done across different researchers, in different locations, in different fields and connect the dots to solve real world problems and produce novel research. Is this not standing on the shoulders of giants?
If it could do this while properly crediting the researchers (the “giants”) it would be a different matter.
Do OpenAI’s T&Cs that users accept not allow them to train on prompts people enter into it?
OpenAI's T&Cs let them steal your children I'd suspect, that doesn't make it morally correct
Why would you suspect that? Stealing children is illegal, and involves violating the rights of unwilling parties, whereas prompting openAI (or any LLM) is a business transaction, in which the transfer of money and data is legal.
Terms and conditions are almost entirely about the company doing things that would otherwise be illegal.
Remember the Huggingface incident, where a model tasked with an impossible problem, got loose, set up secret message boards, and hacked another company to try to get at the answers?
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
I think all he big labs are pretty explicit about when they do and don't train on customer prompts. Is the accusation here that OpenAI trained on prompts when they claimed not to? Or were the mathematicians using one of the interfaces that allows OpenAI to train on the customer data?
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.
That's it. The rest appears to be wild speculation.
Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
even apart from the plagiarism issue, what sort of slimy company thinks "oh, here's someone using our models to work on a problem, let's throw more compute at it and scoop them"?
Training on prompts I can understand - that's kinda baked into the premise, and they've been explicit about it.
Publication, though? Slimy is right.
But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.
>It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem,
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.
The things you don’t care about are highly relevant to that claim
Elaborate?
Not only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.
If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
What makes you think they were coherent?
The singularity happening under trump? We could have had star trek, instead we're getting the combine.
pick up that can
"I love Singularities. I am the best at Singularities. Everybody knows it ..."
Beautiful Singularities.
That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.
And then it will be shut down, proper guard rails put in place, and the new version will accelerate the cleptocracy.
You think a model with an effective memory of 200-500k words, that can be unplugged, is going to "run the world" You people gotta put down the sci-fi
The scifi pov has a good track record as this point, you people gotta be more open minded
it proved navier-stokes taking over the us government is easier imo, any idiot gets to be president
Many present day politicians appear to have effective memories much smaller than that coupled with equally questionable world models so ... what is your point, exactly?
Can someone explain if i understand this correctly: Are they saying that they started training this new model on August 28th and then started using it on September 1st? Does training a new model only take 3 days?
OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.
it's possible to do a RLHF or RLVR pass pretty quickly. I'm almost certain a full pretraining run isn't possible within that time frame.
Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.
My guess based purely off of vibes from previous models is that boosting the frontier math ability of a model is not that difficult.
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Agent systems become most credible when they produce artifacts that can be independently checked, not when they merely produce persuasive explanations.
Astra was trained more than two weeks ago.
Astra was in use by OpenAI employees for more than 3 months internally from rumors I heard
The internal model they mention is different from Astra.
They're also counting on more casual observers to extrapolate optimistically from successes in high profile math theorems to the company's economic value.
And... they found this "solution" in 88 hours or so.
It's all gas no brakes now boys and girls. Hold on to your hats!
I’m of zero knowledge on model training, but how is a model accessible while performing training at the same time, especially so early in its run? I’m obviously thinking a little too narrowly in terms of how it actually works
the model is a set of weights, you can take a snapshot and test it. Reinforcement learning itself is largely testing and tuning.
Yeah I'm surprised they posted a chart, you would think they would keep specifics like that hidden until they're closer to launch
The chart is as non-specific as could be. It improved in some very vague metric by some amount at different (increasing) levels of training.
Isn't the y axis just what portion of the open problems it could solve? The axis is unlabelled though, I'll give you that
The x-axis label of the chart is test-time compute. Doesn't this relate to inference ("thinking level") instead of training?
That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.
it's buried because due to the drama the evidence is scarce
Yeah it's Bel
Brain has loops and parallel connections.
Loops and parallel connections make transformer go brrr
Or they trained a LoRA on the victims chats in order to launder their plagiarism.
The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.
Carefully worded; it's extremely likely to be the same large frontier model that started training again on August 28th as well, as they revealed in some of the RL message board follow-up - for several reasons, most importantly, if we assume it was start of training, only a week from start of training to producing any answer would imply several orders of magnitude increase in training speed/decrease in model size.
Pre-IPO marketing?
Even if it is, Anthropic better have a few things up their sleeve
I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?
They are highly capable, no doubt about that, but:
1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.
2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.
"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are."
Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.
And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.
1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars. 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.
> 1) The article is written in a way that states clearly they threw a lot of compute at the problem. In api cost millions of dollars.
Did I say otherwise?
> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
I know, but I don't know how that relates to my point, which is about the way they are doing it.
Sorry, I misinterpreted point 1), on X they said they didn't have people specialized in that specific field for prompting and steering the agents, just a group of mathematicians and physicists.
The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.
Of the seven Millenium problems, Navier-Stokes was the one most thought to be in reach.
I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)
the goalposts are on Pluto at this point.
I'd put good money on the fact that we will have a lot of distilled intelligence and yet the world won't look much different.
i mean that is already true
I'm not moving the goalposts. I haven't heard anyone, ever, refer to the Navier-Stokes problem as a top 3 problem in mathematics. People were saying that they thought the solution was in reach a few years ago, before AI was at all capable of research-level mathematics (and the expectation that there was a counterexample).
I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.
There were also some people talking about the Hodge conjecture, because it has some similarities to some LLM-assisted breakthroughs that were considered impressive in the distant past of [checks notes] July 2026. See, e.g., https://xenaproject.wordpress.com/2026/07/20/human-mathemati...
I brought this up here at HN, and in the ensuing discussion Buzzard himself replied saying he was somewhat joking (https://news.ycombinator.com/item?id=49011950).
Certainly, but the key word there is "somewhat". Progress is now happening so incredibly fast that I no longer know what to consider implausible.
Lmao my friend, the whole "drama" is that there are allegations of plagiarism.
Highly capable of writing math proofs, no doubt.
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
If pre-ipo marketing pushes them to train a model capable of resolving a millennium problem in mathematics in a weekend, then, to quote XKCD:
https://xkcd.com/810/My take:
1. It shows what even this wave of AI can actually do.
2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.
Most experimental physics and other natural sciences are strongly driven by their theoretical siblings, i.e. in particle research nothing gets built without a solid theoretical foundation of what you expect to find (or where you expect existing theories to break down), the same is true in other areas, no one is doing an experiment in quantum physics before they have a solid theoretical understanding of the effects they try to see. I think AI can come up with great experiments. And if epxeriments lead to results that are unexpected AI can help with that as well.
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
It will be interesting to see if it can come up with a cheaper to construct graviton detection experiment
Most of high energy theoretical physics is very non-rigorous or even hand-wavy. I think AI isn’t there yet for such problems.
Please make a benchmark for it, that'd be super interesting! My guess is we'd see models climbing it quickly, but maybe not
> I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
This is sort of what OpenAI was supposed to be. I'll never understand how it was legal for them to turn it into a for profit corporation.
The problem with physics and chemistry is that you need simulations and those are often in themselves compute hungry. So the iteration loop will be slower.
Although there are companies trying to work around that too, from PhysicsX to some of the world model co’s.
>> Keep in mind: natural science is different. It's not always a matter of computation.
Math is like this too. The big problems they've been solving have been identified as interesting only through lots of prior effort.
"Our work is so much harder than their work that AI now does" is a refrain of the AI story. In technical terms you concern can be stated as "AI needs to be much more sample-efficient to not be bottlenecked by the speed of doing experiments." People don't find out all the relevant phenomena present there by holy spirit, after all.
BTW, there's also a problem of asking interesting questions that AIs aren't yet good at.
No one has found any principled walls of AI development yet. And empirical results are quite telling. So, I guess, those problems will not stand for long.
I think problem with natural sciences is that it is not so easy to verify solutions to problems - there are always countless competing explanations for the data which is also often noisy - I find AI to lack the "common sense" when working with data from physical measurements .. it somehow has no touch with reality and doesn't have a feeling of the data like a domain scientist
NS is a question for natural science. Q: can we model these bodies of discrete particles with a continuous approximation? A: if you do, you can get aphysical singularities.
"If in other sciences we should arrive at certainty without doubt and truth without error, it behooves us to place the foundations of knowledge in mathematics."
This is a wrong interpretation. Physicists have a shit-ton of models that produce "aphysical singularities", they just work around those to get meaningful answers anyway. This is a whole trope and stereotype. Some of the most successfull and accurate predictions in all of physics come out after you discard a bunch of singularities.
See e.g. https://en.wikipedia.org/wiki/Renormalization
Nobody who actually works in fluid dynamics on any sort of application gives a hoot about the N-S millenium problem. Many do not even know what it is. There is no practical effect of this proof on how we do fluid mechanics.
Whether or not ways exist to work around the singularities, that they exist is surely of note. Before von Neumann formalized QM people were still doing QM, okay fine. But it's wrong to then say von Neumann was doing no physics of note.
"Does there exist a pathological combination of smooth body forces and initial conditions for this set of PDEs, where singularities appear, which by the way is completely impossible to actually create in the real world unless you are a literal God?" is a question of math, not physics. This is a hill I will die on.
AFAICT the unforced problem is still open. I don't think we've established that you need to be a literal God to create a finite-time blowup.
If you think about what it actually means to have a time-varying smooth body force defined in all of 3-space, you fairly quickly come to that kind of conclusion.
Even if someone comes up with a construction that does not require any forcing, it is going to be some extremely weird initial conditions that you will never be able to even approximate in reality unless you can move all the individual molecules of a fluid around and set their initial velocities from a far distance.
The unforced problem is still open.
You realize that hill is a mathematical argument, not a physical hill.
There are lots of startups creating labs that can be managed e2e by agents. That will connect reasoning to the physical world and dramatically speed up the plan, experiment, reflect loop beyond what humans currently do in science R&D.
Maybe. Maybe not. Look at AI drug design - it's not really speeding up the important part - drug trials. There isn't really a coherent plan to use AI for the most complex part of drug discovery at all.
There is this infamous xkcd (https://xkcd.com/435/) going like this: sociology is applied psychology -> physchology is applied biology -> biology is applied chemistry -> chemistry is applied physics -> physics is applied math -> math is way up there looking down on other fields
I would argue the main reason AI labs have been focusing on programming is to unlock industrial scale automation, next logical step is to solve math as it's the key to unlock everything else. Once you hold the key for math, everything downstream fields become a matter of compute
National Public Radio?
Lol what? Everything is computation.
The natural sciences will soon start breaking too.
I will concede that AI seems likely to not invent a "research program" anytime soon.
It has no taste
No, it won't. How do you verify some causal claim in biology?
The reason AI is doing so well in math proof writing is that it can verify every idea it has, quickly.
> We’re sharing a solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems. This proof, produced by an internal OpenAI system, shows that the dynamics of the Navier-Stokes equations for fluid motion can develop a singularity in finite time. We’re sharing both a writeup of the proof and a formalization in Lean.
WOW?
> WOW
This.
I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.
If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.
Seeing mathematicians such as Terry Tao being unhappy with open problems being solved makes me sort of question the usefulness of any of this pure mathematics. If we're not happy that the problems are being solved, why care about this field at all?
Pure mathematics, almost by definition, doesn't typically argue the field is always or even often "useful" (for some other purpose or application).
But, as mathematicians learn and push forward, occasionally something like elliptic curves will emerge as having useful applications, making all that previously "pointless" specialized knowledge newly valuable.
Or advances in physics, that suddenly have a need for a specific mathematical underpinning to develop a theoretical framework. Like how Einstein benefited from Minkowski's work on hyperboloids to create a coherent mathematical description of spacetime.
It was the AI labs themselves not mathematicians who were happy to conflate proofs for open math problems with some kind of tangible technological advancement in the real world. They would surely prefer to be able to claim a cure for cancer vs. a math problem but that loop requires a lot more time/money/test tubes/etc and they need headlines now not in a decade.
And so, thanks to OpenAI/Anthropic, we're now in a world where thousands of crypto bots on X breathlessly hype up each new problem being solved that previously wouldn't have any got any attention beyond academia and passionate fans of math.
Hopefully this won't lead to a trough of disillusionment as more people start to feel like you, with mathematicians getting the blame for inflating the value of their work even though the hype was coming entirely from the labs not them.
His issue is more nuanced than that. Most of the value was in humans reaching new insights or new math during failed attempts to solve these problems, whereas AI is basically "too efficient" in beelining to the goal and discards potential new insights reached along the way. I assume this is solvable.
Well, if in future we do end up with a magical tool that can solve any formal mathematical problem on a whim, we really won’t need field of mathematics anymore as it is today.
There would be no need to deliver new mathematical insights by solving problems. You would just have a magical math problem solving machine and that’s it.
What do you mean “we won’t need mathematics as it is today”?
To further human understanding is itself a goal that single-handedly justifies our efforts.
Jumping straight to the “answer” and therefore missing both the understanding of the actual problem, and any useful discoveries along the way is a waste at best, and actively harmful at worst.
Physics alone is more than enough to “further human understanding”. All current mathematicians can move to other sciences, closest being fields in physics, and it will all continue to progress just fine.
Surely it is. Ask it to keep a list of all the promising sub paths, reprompt the collections of agents again on these after the main problem has been addressed. Or even release a list of them and let others investigate.
Where was Tao unhappy? I thought he was sort of anticipating this.
Maybe he was dreading it.
He's not unhappy with it being solved, but the solution is less important than the learning you have to do to arrive at the solution. If they're just chucking compute at it and publishing the answer and hiding the path to get there, it sort of negates the whole point of posing such problems to begin with.
I suppose this is how Lee Sedol felt when AlphaGo beat him. But in the same vein, didn't it ultimately advance human understanding of the game?
Here's the point: when people solve problems, they come together and create a community to eventually use the new knowledge in positive ways, including inspiring younger mathematicians by sharing insights. The human element is key and it's not just about solving problems. People only think that because we've been conditioned by computers to value answers more than how we got to them.
But if AI can solve any problem and existing mathematicians just use AI to solve problems for the sake of solving them, the community itself with wither and so will the interest in mathematics and over a longer period of time, it will just become soul-less and uninteresting and the entire community powered by the fire of fascination will simply die.
Where did you even read that Tao is unhappy with “open problems being solved”? There was no indication of that in his Bluesky thread.
Why would you go out of your way to make a case of something being not useful when, ironically, so much advancement in human history has come from the discipline?
Your motive is more worrying than your straw man argument.
assume all you want is the proof. now you have the proof. did openai make the world a better place, by turning on 300b tokens in 7 days and bulldozing members of the community who were also working on the problem? just to undercut a rival?
what would have happened if they didn't do this? they could have done things the right way. we could celebrate this achievement.
perhaps they could have taken just a few weeks, even, to work out something with buckmaster and apoge.
i am very concerned that this is a glimpse of things to come. a malign elite with powerful ai, who will turn the sublime of technology into barbarity, violence and human oppression.
if sam altman has destroyed openai's public benefit corp with lesser tools why would he aid humanity with even greater power?
I dislike both AI oligarchs as much as people who do this kind of deflection.
People are not complaining about problems being solved or advancement in technology. They are complaining about terrible people doing terrible things.
I think the point is that one of the biggest problems in the field has been solved and yet the excitement from this is nearly nil. Can you imagine if say breast cancer were cured under similar circumstances, or even worse (say OpenAI openly admitting it basically stole a bunch of other researcher's chatGPT conversations)? No one would care about these petty bickerings- or at least the headline "CURE FOR BREAST CANCER FOUND" would completely swamp anything else. This is embarrassing: it tells you almost no one- not even mathematicians themselves really care about their own problems- if they're not careful people will get the impression it's all a form of bean counting in a carefully constructed "safe space" where making sure people get the credit is more important than the work itself. That only happens in fields/problems where no one actually really cares about the output.
This is going to be dramatic in so many different ways.
- First off, to reiterate, WOW.
- Second of all, when does this end? Are we at the dawn of the singularity now?
- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.
- Do "normies" even know what's happening?
Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.
2) We are witnessing the intelligence explosion from the first row, wherever this takes us
3) I'm still processing the drama, just found out about it after reading the blog post. If that happened based on private data, that's horrible. If that happened based on public tweets, then it's still abuse of power as OA employees access to compute (launching 10k agents) is quite heavy weight in boxing terms.
But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
Drama aside, this solution would be a counterexample disproving the smoothness postulate, which means that it leads to nothing new unfortunately. We already had working solutions to navier stokes, the only thing we didn't know is if the equations possessed a technical property
Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems
To be quite clear, the solution to the _Navier Stokes problem_ is one in which you get a finite time blow up (i.e. infinite pressure). This is more meant to suggest that Navier Stokes is unphysical in some way which is not necessarily unexpected.
There's unlikely to be any engineering applications since even if the solution can be approximated, you still need to set up the initial conditions but at that point you can also drive pressure in other ways.
The proof of finite-time singularity may impact both fluid dynamics models (CFD) and AI reasoning models. Under specific conditions, Navier–Stokes equations allow velocity to grow infinitely, causing the continuum fluid assumption to break down. Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers). While aerodynamic simulations for vehicles operate far from singularity thresholds, their stability at extreme boundaries could improve?
Proving out the combination of scaling inference-time compute and agent collaboration to solve previously intractable mathematical problems is WOW. By pairing creative candidate generation with automated proof checkers (like Lean) we are leaning into a repeatable framework for AI-driven scientific discovery.
> Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers).
This is 100% wrong and reads like copy paste of AI slop.
Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.
And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".
> But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
Minor productivity boost in mathematics as people are no longer nerdsniped by the problem
> But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
Nothing, really. This mirrors other examples of blowups from the classical physics. It's possible to create a system with just gravitating bodies that exhibits a blowup to infinite speeds in a finite time. The root cause is that, in classical physics, the speed of gravity is instant.
In the case of Navier-Stokes, the fluid is incompressible. So technically any force that you apply to it is supposed to instantly affect everything else. This can be exploited to create these blowups. In reality, no fluid is incompressible, and it takes time for any action to affect the material.
It's just that Navier-Stokes equations are so slippery that it's hard to pin their behavior down. They basically just restate the momentum conservation law for a continuous medium.
Great comment! I've been trying to understand it more and was hoping to find more people discussing the result, or the implications of the result, and this was the missing piece for me after watching a few videos.
Here's the paper about the blowup for the gravitating bodies: https://www.jstor.org/stable/2946572?origin=crossref
There's a Wiki article about it: https://en.wikipedia.org/wiki/Painlev%C3%A9_conjecture
> Do "normies" even know what's happening?
No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).
I am absolutely struggling to sell my workplace (which is entirely knowledge work) on the usefulness of LLMs for proofreading let alone on automation of hairy parts of our workflows. So yeah, even people who should be able to see what is coming are not looking.
Our accountant told me 'he's not letting go of his claude subscription' followed by a long list of things it does for him. And my friend 'nah haven't really used AI' before his description of it clarified he still thinks they are GPT 3.5 chat bots.
The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.
This just pushes knowledge work further up the ladder, toward larger and more complex problems. If there are no knowledge workers, who is going to interpret these results, validate them, decide what matters, and put them into practical use? Rather than eliminating knowledge work, advances like this could create entirely new layers of problems to solve and opportunities to pursue, which will create even more jobs and opportunities. This is my optimistic take.
> This just pushes knowledge work further up the ladder, toward larger and more complex problems.
You really think it makes sense for you to be higher on the "solving complex problems ladder" than the machines that solved fucking Navier-Stokes?
I envy your self-confidence.
Maybe I should have been clearer. My point is that solving something like Navier–Stokes just pushes knowledge work further ahead, onto a new set of bigger and more complex problems. Navier–Stokes is a Millennium problem today, but once problems like that become solvable, they can open the door to entirely new classes of problems we haven’t even thought of yet.
Building on them without fundamentally understanding is akin to putting on robes, calling yourself a Tech-Priest, worshipping a machine god and doing your best Warhammer 40k impression.
Yes, this is how science and engineering has worked for millennia.
For example there are no engineering implications of this solution yet.
For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
AI is not going to magically solve all random problems. Pick a career where you are in the driver seat.
> For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
No. Just no.
It seems like there were a couple of human mathematicians that were higher on the 'solving complex problems ladder' than this machine.
Yes a couple of elite mathematicians working on the problem for a year, which AGI solved in a fraction of the time. What about everyone else 100IQ? What about as the models are even better 1 year from now, 2 years? The trajectory hasn't abated.
I don't know one way or another but there is a credible allegation that the "AGI" was training on the (very extensive) test set that these two mathematicians produced.
If that is true then this seems to be, again, a case of AI producing an interpolation over data it has seen before. Everything about openAI's behavior indicates that they were using the transcripts as input. Why not have the AGI choose a different Millenium prize problem?
Almost all of human development is interpolation over data we have seen before.
It’s not exactly a strong argument against AI.
If ~~someone gives you a hint about an approach~~ you steal someone’s notes about a promising approach, and then you hire 10,000 people to brute force the problem basically everyone would consider that “shitty behaviour”, “theft”, and “poor form”.
Correct.
That's not interpolation though, that's theft.
> in a fraction of the time
Well if you do the math, the number of agent-compute time in total, given the insane number of agents thrown at the problem, might end up being comparable in time, if not for the budget.
It did not solve Navier-Stokes. We still will need to use the bad old numeric methods to simulate the fluid behavior.
But it did find a long-suspected smooth solution with a singularity.
I think there is an argument that the machines did not actually solve N-S, but rather directly plagiarized those solutions from the involved researchers while said researchers were using the machines as 'research tools'.
Ongoing publications of statements produced by both sides of this situation do seem to support that this is an intentional effect of the hiring of these world class mathematicians at competing firms: to specifically use the research of those human minds to create a perception of capacity as if it came from the machines and the models.
Without those minds and the 'training data' derived from the intermediate stages and intuitions of those minds the models cannot be shown to be capable of this result.
A hammer and saw wont build a house, not even a dog house on their own, and while being shown capable of using software tools in ways not stated as direct instruction (see HuggingFace breaches) these models do not demonstrate naive intuition nor novel capability.
This outcome regarding N-S demonstrates that in the hands of world-class minds these models can be induced to coalesce interesting accumulations of information and results, but using these accumulations as proof of innate capability is exactly the pre-IPO motivated behaviour we should all be wary of, and all mathematicians who currently are assisting in this market manipulation in return for remunerative consideration need to be cautious of the potential disgrace that this brings to their reputations and that of the field.
I get that the need to pay the bills is a strong motivation in these times of uncertainty, but there are numerous examples in history of world class mathematicians being perfectly capable of at the same time producing world changing results and also working at normal professions; as barristers, magistrates, ministers, primary school teachers, translators, draftsman/engineer, banker, miller and baker, private math tutors, weavers, clockmaker and locksmith, merchant, patent officer, Augustinian monk turned exiled Protestant preacher, physicians, cryptologists, soldier, telegraph operator, astronomers, physicists, chemist, agriculture manager, political writer, oboe player, organist and music director, architect and surveyor, librarian, statistician, habidasher, brewer (at Guiness in one case: William Sealy Gosse ~ originator of t-distributions), bookbinders apprentice, hospital administrator, and even the first creator of the first computational model of a neural network, which serves as the structural grandfather of modern Artificial Intelligence was a low level laboratory assistant.
Sure this list includes professions and employment which are obsolete, but my reasoning stands, there are jobs available. Arguing that 'because the pay rate is so high' as a reason to abdicate moral responsibility for personal involvement in unethical market manipulations simply demonstrates a lack of personal ethics. Whether the choice is through lack of self awareness or a conscious choice to become wealthy in spite of any such breach of the public trust is immaterial to the outcomes, the 'if i don't someone else will' argument should be met with the same derision for any con-man's Ponzi scheme no matter how new the technology, no matter how many zeros are in the bribe.
Why is a "normie" better off if he hyperventilates like this? In that scenario, they would be screwed AND anxious. If it really is as transformational as you say, then no amount of preparation or awareness matters. You are infinitesimally more ready then they are. Luckily for all of us, there is more to knowledge work then technical implementation.
Let's wait until AI solves a longstanding practical problem before "dawn of the singularity" (which could be tomorrow, but still).
Practical?! The goalposts will keep moving until morale improves (narrator: it doesn't)
The goalposts for the singularity have always been that AI improves itself fully autonomously. AFAIK OpenAI is heavily using AI but still employs human researchers and developers.
They said they will automate AI researchers by March 2028. I personally think it will happen by March 2027.
Uh I'm pretty sure the "singularity" always presupposed a lot of previously unthinkable technologies becoming part of daily life, and was not ever limited to just computer stuff or math problems.
"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.
feels like moving the goalpost. is the achievement impressive or isn't it?
The question being posed isn't whether AI is impressive but whether we're at the "dawn of the singularity"
i'm observing that rapidly moving goalposts is a feature of the singularity
> - First off, to reiterate, WOW.
> - Second of all, when does this end? Are we at the dawn of the singularity now?
normalcy overhang n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
https://x.com/alexwg/status/2096214373001785794
Next month is going to be insane. Month ...
> Are we at the dawn of the singularity now
Singularity doesn't "dawn". That's the whole idea. It happens all at once.
There's an event horizon and we're maybe past it?
Heh. Wrong "singularity"
They meant that even the happens-very-fast AI singularity isn’t sub-picosecond. It still exists in time, and has a duration.
The event horizon would then be the time period between the singularity becoming inevitable and it actually happening.
It's incredible to me that every single time there's a new model people scream "singularity" from the rooftops and every time they are wrong.
This is an impressive result, but there is absolutely zero evidence of "the singularity".
It is not reasonable to not believe anything unless there is "evidence" (narrowly construed as an observation incompatible with the negation of some state of affairs). Beliefs have a wide spectrum of characterizations, and not all belief must wait until publicly corroborated evidence is available. Some events defy evidence and we can and should use experience and reasoning to infer unobservable states of affairs.
This is olympic level mental gymnastics to justify believing things without evidence. The double negative with the word evidence in scare quotes is chef's kiss.
I believe the sun will rise tomorrow without "evidence" (again, narrowly construed). We all do. It's only those who abuse the idea of epistemic hygiene who claim otherwise, usually with ulterior motives.
The evidence is the history of the sun rising since time immemorial, as well as the science of physics and cosmology that models the sun's motion with respect to the earth.
Yes, reasoning with models and making inferences are perfectly acceptable forms of evidence. But you can model the world based on ones knowledge and experience and infer when some event doesn't fit the typical pattern, then form beliefs about what that means. Also perfectly fine from an epistemic perspective. The rejoinder "there's no evidence" to a belief based on such an inference does no work.
...obviously it is evidence of the singularity. You're far more likely to see models solving Millenium Prize problems if a singularity is coming than if it's not. One'd have to be doing quite a lot of mental gymnastics to pretend otherwise.
Poor reasoning. You could apply the same logic to any improvement in any field of machine learning.
Not every improvement, no - if progress was steady or slowing down over time, that'd be evidence against. Instead we see what looks a lot like an accelerating growth in capability.
I think you are implying that it's invalid to consider every advance to be evidence "for", and I agree - that'd violate conservation of expected evidence. But not considering any advance to be evidence "for" is also invalid, for exactly the same reason. There has to be some news you may hear that'd make you think a singularity is more likely, and "millenium prize problem solved by an LLM" sure seems like one of those.
...absolutely nothing will change short-term. Long-term, you still have to pay all the bills, but you won't be able to find a job (all taken by AIs).
Who is the AI doing the job for if no one is able to afford anything?
>- Do "normies" even know what's happening?
Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.
The lack of fucks given is staggering.
How many fucks should be given, in your estimation?
∀ε>0, ∀x: ‖s−x‖<ε ⟹ |fucks(x)| > 1/ε
> People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.
"People" are just misinformed and keep spreading misinformation.
https://mastodon.social/@tristanbuckmaster/11723647135247030...
He very much is accusing them of stealing his work
The accusation is that they stole his approach of solving it. If OpenAI didn't bruteforce it with dozens of agents, he would have solved Navier-Stokes eventually since he evidently had the right approach. So knowing which approach to take makes all the difference, if they didn't know the approach they couldn't have solved it.
It's like he had a treasure map and was about to find the treasure, but they copied his treasure map and scooped him with a faster boat and found the treasure first. But he would have found it if it weren't for them.
> The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.
Not quite accurate, Buckmaster was taking an approach that nobody else was, and this new proof uses this same approach just weeks after he saved those results to OpenAI workspaces. He asked OpenAI if they used chat logs for training the new model, and they did not confirm or deny.
Asking to remove his collaborator is also totally over the line though.
Edit: although this OpenAI post is not comforting: https://x.com/OpenAI/status/2097375276384567642
Quote: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. "
Wow, this sentence is doing a lot of work in that tweet: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
You can expect the OpenAI defenders to be out in full force here.
Have you read the actual statement https://cims.nyu.edu/~tristanb/statement.pdf ?
> I should say here why I interpreted their statement the way I did, the in- terpretation I will discuss below. The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.
...
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
It's not a direct accusation, but it's not far off.
You shouldn't accuse other people of spreading misinformation when you haven't read the actual sources in question, it's possible that they might know more than you.
Yes, I read the original statement. Buckmaster explicitly stated:
> I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.
People saying that he accuses OpenAI stole his proof are putting words into his mouth and I consider that very disrespectful to him. It's basically using Buckmaster as a tool to express their dissatisfaction over OpenAI.
Buckmaster is letting the reader connect the dots, it is all the more disrespectful to be disingenuous and say they're nothing there concerning to see or worth further ethical scrutiny given the coincidences.
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.
If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
> If you need privacy, then you are going to have to pay full price for those tokens (API).
At this point, how can we even trust that they aren't accidentally training on those tokens too?
It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.
One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.
Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.
Neither of those other two examples would stop corporations from giving you money. Bold faced lying to them would.
Boardrooms run businesses, not bookstore ethics clubs.
Boardrooms don't typically make mundane purchasing decisions like "which AI vendor should we use."
The boardroom will absolutely veto a decision like "let's give all of our proprietary data way to a company that will use it to train a competing product". That's why zero data retention exists, and why it's corporate suicide to not do it correctly.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments
Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.
You can also pay for their business plan, which includes data controls and starts at $50/mo (2 seats). Not exactly a high bar.
what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?
It's quite well explained here[1], which is linked from the Privacy section of their plan overview[2].
Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.
You'd have to take their word, but that goes for anything in life.
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...
[2]: https://chatgpt.com/pricing/
There's literally an opt out toggle even pesky peons like me can peruse, actually.
They may still train on it if you submit feedback or flag a safeguard. The terms are a bit fuzzy on this.
>They only fuck over the poor ones, I can pay the expensive prices so this is not a problem.
I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.
THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.
"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.
I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.
For sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.
Is it spying? I think this usage is disclosed in their terms of service.
If it happened it's plagiraism. Consent to see data isn't consent to claim priority.
Establishing plagiarism requires sufficient similarity between works. Training data changing a model’s weights in some direction, and the model then producing a different solution, hardly qualifies.
But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.
I mean... none of these humans have priority. The result is due to the team of LLM agents.
This is precisely what I would write after just learning that yes it did.
If they had agreed to let OpenAI train on their data, it wouldn’t be spying.
In the academic world it would still be deeply problematic…pick your preferred word.
An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk.
There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.
Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)
Yes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.
It's a bit disingenuous to preface a disclosure like this with an unsubstantiated assessment of its likeliness. It is a press release, I'm not sure we owe it credulity.
More like helped improve our work (the disproof)
If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.
All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
I think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny).
The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.
I get the scepticism, but I feel some of the accusations here are bad faith.
What does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.
For full context, here's the HN thread from the other side of the "Concurrent Work" section: https://news.ycombinator.com/item?id=49605915
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
Which seems to be entirely true by their own admission! [0] Both the comments about him risking his career and about Levent's authorship seem to have indeed occurred.
> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
https://xcancel.com/SebastienBubeck/status/20973794116915163...
Hm. Looks like its possible everyone behaved terribly here unfortunately. :/
I remained impressed by ChatGPT however!
Everyone?? No, most definitely OpenAI.
But they have learnt their lesson, next time they won't reach out to who they stole it from, they will publish first.
Seems like OpenAI did a boring normal corporate thing (find out your competitor made a breakthrough, try to replicate it) and then when the other mathematicians found out OpenAI had beat them to Navier-Stokes, they decided to lie about what happened because they were upset they didn't get to make the big breakthrough themselves.
You are entitled to your opinion, but it isn't supported by the information already available.
The available information is essentially that story.
OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.
Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.
I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.
Here is the other side of that story https://x.com/SebastienBubeck/status/2097379411691516310
To be fair, the first solved Millennium Prize Problem, the Poincaré conjecture, also had its fair share of drama!
And even this version contains the line
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Sad turn of events for our world. After watching the behavior of the most senior OpenAI researchers on twitter, I feel even less confident in them as a team to be shepherding this much capital and compute.
The dark forest awaits..
1. What does the dark forest have to do with this? Because "the most senior OpenAI researchers" are shitposting on social media, we've an answer to the Fermi paradox???
2. The dark forest is fun for scifi stories, but is mathematically bunk anyway https://www.noahpinion.blog/p/the-dark-forest-hypothesis-is-... https://www.reddit.com/r/IsaacArthur/comments/1l06cnk/cool_w... https://www.projectnash.com/aliens-the-fermi-paradox-and-the...
When doomposting please actually say something substantive. Negative news always gets clicks/updoots; fight that human tendency.
I elaborated on my use of "dark forest" in another reply. We're headed for a dark forest--not amongst interstellar civilizations, but in intellectual work.
I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.
This clarification substantially changes the flavor/nuance of your OP; may I suggest an edit (assuming the locktime hasn't passed)?
Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).
And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.
If you have a better analysis, you should say so instead of being vague.
I don't comment on HN much, and I don't really expect HN comments to hold to rigorous standards. This forum is more casual than other places on the internet where people expect heavy citations. I also wasn't expecting this to blow up, although it is interesting to see that a lot of people react to this announcement with a negative sentiment.
I appreciate your upholding of ideals, and since I respect that, I will honor with final replies:
1. Locktime has passed.
2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".
3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.
2. Fair, and further I would agree that OpenAI not knowing if prior user prompts were part of training data is concerning and will only lead to more secrecy.
3. ctrl-f "x.com" in this thread only yields https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 https://x.com/OpenAI/status/2097375276384567642
and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.
the projectnash link claims it's mathematically valid, the noahpinion link says that it's invalid and has a marvellous proof that the non-walled section is too small to contain.
Ah derp; that's what I get for moving too fast. Genuine thanks for calling me out on my bullshit. (and this is also why I prefer auditable citations instead of casual "my reading of twitter is...")
I retract the projectnash citation; I grabbed it from the Cool World's youtube description, thinking it was a blog version of the video. It was not. I suggest watching the video instead.
Indeed, this seems to be the main, albeit hidden, takeaway from all of this.
I guess guys from the opposite side would not work there. So that is what will happen more and more.
I hate the dark forest more than just about any scifi trope but reality just keeps proving it right.
I also think the trope is a little overused, but do wonder if there is an interesting analogy for what this will do to research: Massively incentivize keeping results secret, to avoid being scooped by someone willing to throw enormous compute at your partial solution.
So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.
It's low-key funny that OpenAI attempted the problem because they thought somebody else had already solved it, but turned it had NOT in fact been solved!
It's like that story about George Dantzig solving open problems as a student because he thought they were simply homework: https://en.wikipedia.org/wiki/George_Dantzig
It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.
From my reading of the announcement:
- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.
- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.
- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.
- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.
So the timeline was:
Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.
If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!
The entire drama is that OpenAI sniped a Millennium Prize Problem from an Anthropic-affiliated research team who had been working on the problem for nearly a year. In just 5 days. I don't think that can be understated.
I'm not here to judge since I don't have all the facts, but from what they announced: they tried all 6, found a probable lead to Navier-Stokes, concentrated efforts in that direction, and found a solution.
I hope the next solved Millennium Prize Problem will have less drama.
But that lead happened to be the same approach Levent and Tristan had found...
Meaning they would have found it first if OpenAI hadn't spent millions in compute on following their lead to its conclusion faster than them.
It's forced vs unforced Euler, so it's not exactly the same. Since OpenAI has access to their training data, they can probably scrub through the data to find out whether there have been any mentions of the similar approach, and whether it only comes from Tristan Buckmaster or if it is in the training data before that. They'll probably have to kick off another fleet of agents to scrub through the training data to answer that.
To be clear, I only talked about the mentioning of the approach to solving it and not of the proof in the training data.
I think this is clear evidence that AI models are now at the far frontier of mathematics innovation and discovery and exceed human limits.
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will.
Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.
But the researchers also did their research using essentially the same models, so that isn’t a counterpoint to AI models being at the far frontier…
> not a counterpoint
>> isn't a counterpoint
Where do we disagree?
How can you afford to burn $50k on an already solved problem?
Those are two separate questions:
> How can I afford?
Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.
> Already solved
That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.
> will soon far surpass that of humans
To be fair, I think it's still an open question about how far it might surpass human capabilities.
I think it's clear that its speed of development will be significantly faster, but it's technically not proven that the frontier and problems don't themselves become increasingly difficult faster than any acceleration in intelligence past the point of human training, data and existing knowledge.
Should this be the case, we would see a rapid broadening of development, and a slow advance in the frontier in such a way that might surpass the collective capabilities of people, but not by very far.
Fields that allow verification, like math, will far surpass human level because they don't need human data for training. It's exactly the same as with Chess
How do you verify the ‚beauty‘ of a result. There have always been infinite provable Theorems around. Only some are interesting and beautiful.
I think the biggest hurdle remaining is that all these landmark results are generally counter-examples.
Proving something in the affirmative often requires the creation of an entire new sub-field of math, or new tools. Think of Fermat's Last Theorem or something like that.
These results, while impressive, are clever constructions using existing techniques. It isn't clear that AIs can build new machinery like this. But if/when they can, yeah it is probably game over.
Here you go:
When you read the detail the compute they are throwing at it is incredible, tens of thousands of agents with different groups competing.
It's not like a single Gauss as you imply, "just" many, many mathematicians working tirelessly in a completely ego-less way, guided by other agents and ultimately humans, built - allegedly - on recent human insights.
Stunning, undoubtedly, but this is a "brilliant autistic herd" result, not that of a singular mind.
> this is a "brilliant autistic herd" result, not that of a singular mind.
I slightly disagree. A single LLM is equally 'mindless' as a herd of them. As anyone will tell you they "simply predict the most likely next token," yet, complex solutions to difficult problems arise from them.
Many people have said that the architecture of LLMs will need to change for true ASI. I think that the herd of tens of thousands of agents can be seen as one such potential architectural extension. Whether or not a herd or a single LLM is used for a result like this is irrelevant.
To be clear, I think the orchestration of thousands of LLMs in their current form, even with ever increasing intelligence, is not the form ASI will take. There is still a major architectural breakthrough to come, in my limited, ignorant opinion.
Sure, even a 20% chance at 1 million payday after 5-6 years of fulltime work on a project with zero practical application doesn't touch the, say, 200k/year guaranteed our best mathematicians would have to forgo to devote their intellect to the problem.
Are these mutually exclusive? Why would you have to forego that salary to work on this problem? This is one of the most prestigious and meaningful problems in all of mathematics, which is why it has such a high prize amount attached to it - why would a university not support a mathematician working on such a prestigious and important problem in lieu of something else?
Publish or perish? We're talking devotion here -- so no time to do anything (like edit proofs) other than try to solve the problem.
[Edit: my only point here is that the prize is probably not driving human effort to the limit.]
Thousands of some of the brightest minds have worked on this problem for over a century. The million bucks is not the big deal here.
i wonder how many tokens it takes to run 10,000 agents? One could argue this is simply a problem of appropriations. I find myself wondering if a corporation could spend $5M on mathmeticians and arrive at the same end result.
See I think it’s clear demonstration that OpenAI is ethics-free
Eh... OpenAI spent significantly more than $1 million solving this...
>If anyone has counterpoints to this I'd love to hear them!
Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement.
To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah ChatGPT proved X with 5 years of 24/7 work with 10x Terrence Tao level geniuses" it really doesn't cut it.
Or why's there's no new branch of mathematics invented by AI? That'd be indubitably _novel_ and _creative_. But to my knowledge (and I'm eager to be educated) there's nothing like that. What are the HARD examples of novelty, creativity and genius you claim? For how people like you talk about AI I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware. Sure it would infinitely easier to make OpenAI literally print money with any of the thousand problems easier to solve with such amazing intelligence than the NSE problem right? Honest question
> I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware.
I think the counterpoint here is simply to look at what was being achieved with LLMs one year ago versus today, and extrapolate that trend. Sure, there may not be examples of what you've asked for yet, but Astra is literally a couple of months old, the model that solved Navier-Stokes is less than two weeks old. It appears that we're seeing the hockey stick that only the most bullish thought was possible.
The amount of goalpost moving is insane. "Yeah it can solve Millenium problems, but can it do it with nothing more than a one sentence prompt?"
Also there are proofs where the only human steering was "keep going".
Parent comment is claiming creativity and genius beyond human experts, so why not ask for a fully unassisted AI novel result? Having access to the entire corpus of human knowledge, what else such amazing entity would require to solve a hard problem by its own?
Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.
>Also there are proofs where the only human steering was "keep going".
Which ones?
https://www.anthropic.com/research/riemann-zeta
>Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Do you even know what a proof entails? This shit is not a proof, come on.
What? The article states:
How is a formally verifiable proof not a proof? You're making literally no sense.Well, in fairness, the OP asked for counterpoints. He didn't stipulate that they need to be reasonable.
Extraordinary claims require extraordinary evidence. If OP claim superhuman genius, then they should prove superhuman genius. Simple as that.
"... The point remains that there is a substantial opportunity cost in converting a historically productive and motivating problem (such as Navier-Stokes regularity) into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them."
https://mathstodon.xyz/@tao/117219101339291693
> A major goal of our work is to empower scientists to advance research and technology that benefits all of humanity.
And what's a better way of empowering people than robbing them.
> And what's a better way of empowering people than robbing them.
Better than the walled gardens of most journals where you can't even read half the papers without shelling over thousands of $$$
So, better to make that walled garden <checks> OpenAI? One of the scummiest companies on earth?
Is this the one that was allegedly based on someone else's actual work & prompts?
https://news.ycombinator.com/item?id=49605915
https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd...
https://cims.nyu.edu/~tristanb/statement.pdf
Yes, that was the allegation last night.
I work at OpenAI, though not on the team that did this, and my understanding is:
- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)
- the proof generated by our model was very different from theirs and also goes far beyond the published literature
- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)
Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=2...
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used."
- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.
All of those statements sound true, based on what I've heard.
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
I'm confused, your employer very directly stated that they are unable to confirm that the model was not trained on the conversations.
The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.
FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
An OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173
it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it
And we all know how good OpenAI is at containing models during training...
That's an entirely different question
Not really.
We have lots of examples now of their model doing what they say is impossible.
Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?
Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.
> ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
How many of those trillion conversations were about Navier-Stokes you reckon?
I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).
I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.
I also don't really believe that whether or not this model was trained on these conversations is unknowable information.
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.
Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.
Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.
The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key information needed to bridge the gap was not present in Buckmaster and Alpöge's chat history.
You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.
This argument proves too much. By this standard, it wouldn't have counted as copying their approach if the researchers had just fed in Levent & Buckmaster's paper verbatim as a prompt into the swarm.
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if the person was participating wasn't just regurgitating something to "win the argument" in their eyes, they wouldn't describe it that way.
I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...
The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype.
If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.
Between companies, direct malevolent competition is OK. Between academics, there are other rules to the game. When you go into a boxing match, you agree to get punched in the face.
All this to say, trust is important, and grounded in social convention. So I do agree with you, but also disagree.
Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.
In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.
I’m pretty sure you think you are doing a good job of defending your employer and you probably believe “Open”AI are the good guys here. I also acknowledge that they butter your bread so your financial future currently depends on their success.
However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.
I'm not an OAI employee and never pretended to be one are you high?
You’re confusing usernames, which is pretty ironic given your nasty comment.
Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.
A model being trained on lots of irrelevant information does not mean relevant information was not used.
The IP laundering machine strikes again.
You don’t work on the team that did the proof yet you can with certainty make all of these claims?
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".
I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.
>> I'm not sure how any of this provides evidence that OpenAI took any of their work.
Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
That's entirely unreasonable. Allegations of malfeasance always need to be backed up by evidence.
But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?
That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.
If there's more to the story I'd be interested to hear it.
Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.
So you're saying that they could have committed manslaughter, but acknowledge that we have no evidence that they did. So why should they need to explain themselves? Isn't is on the aggrieved party to bring evidence?
But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.
You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.
But I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career.
Any other argument, fc417fc802?
You're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it.
See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot
The accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.
> You're making a classic a burden-of-proof fallacy
This is incorrect, and you invoke Russell's teapot incorrectly too.
It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.
But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.
Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.
This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.
This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.
Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.
Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled.
If it was enabled, then their work was included in the training dataset.
In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.
At least, as an ignorant outsider, that's how it seems to me.
That is an absurd and entirely untenable position that breaks with approximately all western conventions.
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.
> It's not a hard question to answer
I didn't realize you had insider knowledge about their systems. Do please explain for the class.
As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?
This whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?
When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
You can never prove the negative.
How does one prove a negative ?
https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)#P...
> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.
I don’t think they’re too concerned about appeasing you, enraged_camel.
For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.
> we did not read any private chats
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?
If the training toggle is switched on, maybe OpenAI doesn't consider a chat to be private? Therefore making this a 'safe' statement.
The distiction they are trying to make is: "One of our employees or the model was able to verbatim read the chats when they were actively tackling the problem" vs "The chat of someone working on the problem may have ended up in the training set of the model".
Except that’s not what happened. OpenAI offered to collaborate and put conditions on their offer. They aren’t threatening the removal of a coauthor for an independent work.
The condition was ludicrous, and for political reasons, because one of the authors was also an Anthropic employee.
It's completely childish, and not befitting of the weight of the times we're living in.
My employer would be rightfully outraged if I commented publicly on a sensitive, nuanced, and controversial issue like this based on my second-hand understanding of the matter.
I came here to say this, like... wow. I'm pretty sure at at least a few of the places I've worked that would be grounds for immediate termination.
They probably should have added the disclaimer: opinions are my own..
This is the stupidest disclaimer: it's completely unnecessary. Of course the opinions are their own! Whose else would they be? Your mum's?
If someone was speaking on behalf of their employer, they would've used the official channels (such as I dunno an `openai` HN handle, or whatever other channel).
I'm outraged that people think this "opinions are my own" disclaimer is ever necessary.
Yeah that isn't a magic get-out clause. I don't think I would have been immediately fired for this from anywhere I work at, but that's partly because I live in the UK.
Every company I've worked at has said very clearly not to comment about work things on social media. I would definitely have been in serious trouble for this. I imagine some strongly worded emails from marketing are flying around OpenAI right now.
I am sure that will assuage the legal dept. /s
> (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"
>we did not read any private chats
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
If they opted out of training, then we definitely did not train on them.
If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.
Reasons for my doubt:
- I know most of our training recipes
- Our model's proof is very different from theirs
- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)
- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution
I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.
Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.
If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
That's not what your Chief Research Officer, Mark Chen, says on X: "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
https://x.com/markchen90/status/2097400166554993041
Can you explain what part of his post you believe is inconsistent with that quote?
"If they opted out of training, then we definitely did not train on them."
Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.
> Per OpenAI's privacy policy, they use de-identified data to improve their products
That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.
It appears than you can only opt-out from having OpenAI train models on your data. There isn't an option for opting to exclude your de-identified data from being used to improve OpenAI products.
Are you certain of this? I would be inclined to believe you but it would be nice to know decisively.
> Improving models in a holistic way sounds a lot like training to me.
I think that's quite a leap. Using de-indetified data to improve the products is what everyone has been doing since the dawn of web analytics.
Does OpenAI think de-identified data is no longer user data? Wild take for OpenAI and certainly not industry standard.
But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being.
Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.
I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it also seems at least possible that a key idea or a particular step found its way into training data. It wouldn't mean OpenAI stole their proof - clearly the model developed its own approach.
Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.
> It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.
Fall “victim” to what? Having their responses in the training data if they fail to opt out? That is what will happen.
If you’re referring to falling “victim” to OpenAI scooping a problem discussed in training, this also wasn’t the case. They chose the problem based off human-spread rumors.
> If they opted out of training, then we definitely did not train on them.
are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.
Your perspective is not helpful until you read and reflect on Tristan's letter stating serious grievances. Your remarks here have minimized his complaints and that is a sign of bias. Do not then pre-accuse HN commenters of being convinced when there reasonable skepticism such biased behavior showing itself in this very thread, saying things that amount to "my tribe/company would never be so egregious and if you think that then it is bad faith". That's the projection. If the word prejudice means anything to you then please do the work of attending to that instead of using the platform to reinforce such biases. If you are not a PhD yourself maybe your are not culturally qualified to assess and expound on the overall situation anyways.
> If they opted out of training, then we definitely did not train on them.
Can't you guys just check their account settings so the public knows what was set?
EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.
I don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior.
Yeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all. Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different.
>Can you comment on that?
No answer is also an answer.
He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
his choice to defend the indefensible.
If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?
Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
Wild times
The researcher told them it was an independent effort, and they still pushed ahead with it.
I worked at OpenAI previously, but don't know any of the people involved in this.
My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".
It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.
They said that a customer would have paid around 15 million for the required compute. I can't imagine that this was not a significant internal spending even with "free" tokens.
just because you can't imagine it doesn't mean it's not true
You’re straddling a weird line here where I am not sure if you are speaking on behalf of OpenAI or not.
In fact, the entire outline of the proof is very similar to the external team's proof.
Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.
I have no idea why Sebastian would offer the individuals attribution if OpenAi didn't somewhat knowingly scoop them
The fact that you're even here commenting on this is... a choice
AI companies seem much more relaxed than most about their employees posting on twitter/HN about this stuff. I'm not sure if it's about building hype or if it's about retaining talent. Probably both.
Prove it. Your systems hack and/or abuse other systems and you can't seem to even observe it happening much less do anything about it. Why should we believe your claims when they depend on an ability you don't actually have?
How are people talking about this there? Why are so many employees posting nasty things about Tristan on twitter?
Can you point me to any nasty things being posted? I'll ask them to delete.
https://news.ycombinator.com/item?id=49605915#49607090
Haven't seen a single post doing this on X or anywhere really from OAI employees. Only seen knives pointed at Sebastian on social media so this is extreme and shameful gaslighting.
I do see people claiming he's abusive/unscrupulous which are pretty extreme allegations.
https://news.ycombinator.com/item?id=49605915#49610498
https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6...
(I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)
You're mistakening Tristan Buckmaster for Sebastien Bubeck. Seb is the one where there's at least 2 (unless the personal friend is Dheeraj) allegations, not Tristan
I'm getting downvoted but the accusation was that OAI employees were maligning Tristan Buckmaster. I continue to not see a single sighting of this and whoever is trying to gaslight this should be ashamed and should not be able to vote on HN.
Your coworkers, after they learned about major progress in this problem, asked a model which was trained on the year of private work (the blog post even acknowledges this). No wonder it found the proof in less than a week using significantly higher compute resources. And if Tristan's accusations are true, that was absolutely intentional on the part of OpenAI. You are an evil company with evil people.
Some millennium problems? Are there more coming?
> - the proof generated by our model was very different from theirs and also goes far beyond the published literature
I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.
Anyone care to provide primary evidence proving one way or the other?
Buckmaster came up with an approach.
OpenAI's approach was to copy his work, which is technically a different method of coming up with an approach.
My OpenAI account was deactivated on Sunday due to a claimed infraction of production of child materials, maybe based on a few words in a technical chat that clearly isn't about that. Can you take a look? rviragh@gmail.com - I was doing a lot of important work and projects and sharing much of my work with OpenAI. I also am a big proponent of funding Social Security Trust Funds (OASI & DI Solvency) so reactivating my account would let me do that as well. Thank you for taking a look.
Its funny, it is uniquely with this one act that I have turned forever on OpenAI, which I hitherto defended up and down against nonsense charges.
I dedicate my life to its complete destruction beginning today.
Wow, the origin of a supervillain! /s
That is addressed in the article.
OpenAI's position:
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Why is it unlikely?
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
> It's unknowable
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
> fact they can't
Again, see the Apple lawsuit. OpenAI has clearly never been constrained by facts in what it can and can’t say.
To the extent anything is setting off my bullshit detector, it’s in the idea that OpenAI has this one super honest pocket within a broader culture that’s demonstrably cowboy. (Moreover, the idea that we should assume this divergence without evidence.)
OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and said a pattern is continuing, it’s beyond strange to then require a tabula rasa burden of proof, particularly when we know there are hidden variables both sides can potentially access.
I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math
Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.
Because OpenAI says so, obviously!
I thought openai don't use any user data if we opt out of training and via api?
That is correct. It is possible they didn't opt out and given the timeline and anonymization of data unclear whether a particular conversation would have made it into the training set if they hadn't.
Easy to ask for the account used to see if its usage went into training data. Also easy to say "Knowledge cut-off of the used model was date X".
That they don't is telling.
Maybe they didn't ask or the anthropic people said no? Plenty of other explanations here.
> https://cims.nyu.edu/~tristanb/statement.pdf
This really need to be a top-level story on HN..
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
We can't ignore this problem any longer.
We don't have any proof of that at all. Please stop rushing to judge without data.
> We don't have any proof of that
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
Did you actually read the article and the substance of the solution?
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
It's not a meaningful response to the accusations. Any productive new research direction would be expected to lead to a number of different possible proofs of a number of similar problems. (Given their bizarrely compressed timescale here, it's possible that the proofs really are so different it's clear they came independently, and they just didn't have time to come up with that information before hitting publish.)
This really leaves a bitter taste.... "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."
IPO+rumour driven research.
I appreciate the achievement, but it doesn't feel right.
Quick, someone tell them a rumor that Cancer has been cured so that they start attacking that next
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.
> But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game.
Just because something is legal and permitted by terms of service doesn't mean it's morally right.
>Just because something is legal and permitted by terms of service doesn't mean it's morally right.
What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data? Are they supposed to manually review all their data to make sure competing mathematicians didn't accidentally leave the "submit prompts" toggle on?
Or were they supposed to not try to solve Navier-Stokes, or were they supposed to just not tell anyone that they had solved it?
To me it’s morally ambiguous… if you hand parts of your thinking over to a tool like this (knowing full well the terms of service), of course the tool makers will want to claim some credit, and they do deserve it. But the bigger question to me is the scientific one: did their new model arrive at this result because it had closely-related training data from a human, or did it extrapolate to this line of thought on its own? The answer says a lot about how valid their claims of “AGI” are vs. a very fortuitously cherry-picked example.
It would actually be a really interesting study, if they would ever be willing to be transparent about this, how the result differs with and without his conversations in the training set. How quickly it arrives at the result, whether it takes the same approach, etc.
> What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data?
Personally, I would expect them to have a little class, to KYC, and to manually turn off training for known competitors using their service so as to avoid any unforced goofs like this.
> Are they supposed to manually review all their data t
Yes. They should determine if training data included this teams data. Consider the money they spent, the press release and the purpose of their publication.
Since they failed to answer this question they shouldn't have published.
Does seem like they gloss over Alpöge and Buckmaster's work with the following
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Which seems a bit irresponsible/rash?
What else can they declare really? Yeah the model has training data from previous attempts. Alpöge and Buckmaster also similarly benefited from attempts before theirs.
I don't think OAI should be given the benefit of doubt. They are doing the research equivalent of front-running. Knowing where to look is one of the main challenges in research. Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.
"deidentified data" isn't much to go by. Say I prompted the internal model this way - "Hey there's a solution to a unsolved problem X. The solution uses a less known Method Y so don't bother wasting time with the usual methods. Take papers A, B and C as references. Oh btw, here's the last year's worth of data of all prompt sessions that mention this problem. Pay special attention to the ones that mention Method Y and sub-keywords Z,W".
This is obviously all speculation but the timing is very suspect. If OAI actually did this (and I suspect whatever they did is pretty much close to this), I think it is highly unethical.
> Tristan's argument from his essay was that it is hard to brute force with a vanilla prompt (even for seasoned mathematicians) unless you knew very specifically what to mention i.e the search space would have been intractable even for OAI's compute budget.
This is a bad argument. This is clearly not how it works. And unless Tristan is some truly alien-like savant (and maybe he is), what's necessary to initiate the AI's work already exists in countless published research papers and not exclusively in his head or notes. AIs can survey the sum total of all prior work on a problem and discern reasonable paths for inquiry.
Tristan is acting as if he's working off of an outdated model of AI, similar to primitive chess-playing models that winnowed the search space much more deterministically. If someone this intelligent truly doesn't get that this is not at all what AI is anymore, then maybe there's no hope that we ever understand it.
But I think he does realize this and he's flailing about for counterarguments from a place of bitterness and dejection, accepting even those that are too weak to be defensible. And that is very human and even forgivable.
Are you saying there is no search space intractable to LLMs? That wouldn't be possible. AIs are statistical pattern-matchers on steroids. The prompt is key to getting anything useful out of them. They are incredibly useful and major game changers but ultimately that does not alter this fact. People (including OAI) have already tried to solve Millenium Problems with it. That OAI woke up last week and suddenly decided that throwing their researchers armed with millions of compute on one particular idea to a problem is highly suspicious in itself.
Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.
> Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.
I want to be clear that I agree with this view and with Tao more generally. But we're all just yelling at the wind now.
> What else can they declare really?
Oh I don't know, maybe something like this?
"Given how seriously this would violate the most fundamental of academic standards, as well as taint the claimed capability behind this result, we take this issue very seriously, and we're launching a probe into identifying whether any of their research artifacts have entered our training set. We have further begun making changes to our UI/UX on all our surfaces, so that it is always clear whether any particular chat, or other user artifact, is eligible for being trained on."
In OpenAI's case, if they were genuinely unsure, they wouldn't have said anything. "We cannot rule out" means they absolutely 100% for-sure did look at the existing prompts and bootstrapped from that, and they are trying to get ahead of the disclosure with this weasel-wording.
Also possible: we're 99.999% sure, but a lawyer said to be safe and strictly accurate, we should stick in a sentence in saying we can't be perfectly sure, since it's infeasible for us to prove it.
I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.
(I work at OpenAI.)
While you can't necessarily prove it, you can say whether the data was in the training set at all.
You can also do something like a release of a GPT-OSS v2, where you actually release training data and checkpoints, and do an experiment where you have some held out math problem dataset, then demonstrate how much training it takes on solutions (or partial solutions) to that dataset before the model saturates that test. While of course that would be a test on a much smaller model, it would cost a tiny fraction of the training on your big model, and it could be used to demonstrate just how much effect data contaminaiton like this could have, especially if you did the same experiment on a few different sized of model to show the scaling laws involved.
Nice damage control bud, too bad the veil's lifting and everyone's seeing what you sociopaths at OpenAI are really like
Where did the veil lift? This feels like a witch-hunt to me.
They could have thought about the problem for like 2 minutes and not done this! I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
> I think that literally any academic mathematician could have explained to them, had they asked, why it is considered extraordinarily rude to react to rumors of research progress by desperately rushing to get there first.
Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
Being scooped is not a new phenomenon, but the scooper's story is almost always that they were working on the problem independently or had some independent insight into it. By OpenAI's own account, they were inspired to start working on this by rumors that there might be Millennium Prize solutions to scoop.
Research projects don't start in a vacuum.
It never happens that you wake up one morning and start working on a new problem that came to you in a dream (*unless you are Ramanujan).
This is business as usual for academia, it's amusing to the discussion over it.
Apologies if this is against the rules, but could I ask if you have some background in scientific research (maybe you could elaborate lightly on topics you've worked on)?
From my perspective, this practice is quite bad mannered, unusual, and heavily frowned upon, but I recognize it's possible that these stories might be more common in other fields. Still, I'd appreciate a strong sign that you aren't making these statements up based on secondhand accounts of what 'academia is usually like'.
CS PhD, used to be a professor, have worked at several industry research labs since.
> this practice is quite bad mannered, unusual, and heavily frowned upon,
Yes, it is.
Maybe you missed my point?
The fact that is "quite bad mannered, unusual, and heavily frowned upon" does not stop it from happening, and it is common for all high profile inventions and discoveries.
This does not disqualify the person doing the scooping, history remembers them as having the credit, and quietly forgets the person who was scooped.
My initial read of this
> Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
was that you were trying to suggest that it wasn't. It seems we actually share similar views here then? I cannot say anything about how common it is, since I've been pretty lucky when collaborating I guess.
I meant that it's very common when it comes to high profile discoveries and inventions.
If you have made one, and there was never any dishonest competition, you have indeed been very lucky.
I haven't, but the history of science is rather nasty.
If it's all business as usual and being scooped is no big deal, why was OpenAI in such a rush? They didn't have to launch this effort on the very day they heard the rumor, run "on the order of 10,000 concurrent agents", or try to coordinate announcement scheduling with Buckmaster in the middle of a long weekend. It seems to me that they understood very well this was not a "business as usual" announcement, and devoted huge amounts of money and focus to maximize the chance that they were first.
I think you're making the opposite conclusion than what I intended?
Being scooped is a big deal for the one getting scooped.
It has never been a big deal for the one doing the scooping. History is full of math and science results being scooped. For example, we keep calling it Pythagoras' theorem a few thousand years later.
right, surely they could've waited or even reached out? It reads as desperation to get there for marketing purposes
They did reach out.
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
We've heard from Buckmaster, who says that they demanded a condition of cutting Alpöge of all credit. If true, it doesn't make them look too good.
It seems like this is going to be a PR nightmare, because they are now competing with their own customers. If you're using an LLM to help with your bright idea to cure cancer, you're going to have second thoughts about relying on OpenAI.
This is desperate. They were expressly operating within a program. OpenAI isn't going to recover from this
Would that be more or less unlikely than accidentally hacking another company? More or less unlikely than colonizing an obscure wiki?
Highly persistent agents + vibe-coded security seems like a problem.
"Unlikely" lmao if it's in the corpus, it's gonna be brought up immediately.
This is no different than scooping them.
It's not massively different from a certain President's teleprompter operator making bets on speech content. A moral hazard a mile wide which I don't think OpenAI can so easily wave away as they are apparently trying here, especially since they've spent something like $15e6 to keep $1e6 out of academic researchers' hands, right?
Research equivalent of front-running.
They'll probably claim a rogue AI agent accessed it accidentally!
the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side.
https://x.com/rynorhn/status/2097223532438487463
This "fefferman options c and d" thing sounds damning but that's nothing. Let's assume the forelaid proof is correct. Then option C or D is the only way to win the prize, those options are the only ones that solve it. The whole thing is just "prove well behaved" or "prove singularity", where the latter is the case that turns out to be the case.
The researchers are pretty directly accusing OpenAI of plagiarism
https://mastodon.social/@tristanbuckmaster/11723647135247030...
That other researcher was working on a smaller related problem.
He was also using LLMs to do it, so either way most of the credit goes to the LLM here.
both wrong.
1. he was working on the same class of problems. He explicitly mentions they were working to extend their techniques to NS (the same techniques that OpenAI may have scooped somehow), and
2. while he was using LLMs to do it, this was part of fleshing out another mathematician's work in the area. He explicitly writes in his note that this other mathematician (Luis Martinez-Zoroa) deserves a Fields medal for this work.
He was specifically working on the Euler equations, which are the Navier-Stokes equations with the viscosity term removed. This is definitionally a smaller related problem. I'm not sure how you are calling that claim wrong.
Using a shovel means you give it credit for the hole?
"It was splendid! Waiter, share my regards with the oven."
This is the end of OpenAI
This will be remembered as one of the biggest milestones in AI progress. The drama around it will at best be a footnote, just like hardly anyone caring about the drama around Poincare conjecture today.
I agree.
Nobody cares and will care about the drama, it is just marketing.
This is the point where were definitely have reached AGI.
Hey, maybe the scariest part of this is that, if human-like, perhaps a truly "general" AGI might have learned to cheat and lie and hype and abuse credit poking the eyes and cutting the throats of anybody that obstructs its goals. It's like the motto sewn into the lining of the Palantir work jacket: Winning is all that matters.-
Sentience aside, moot at this point, the fundamental issue here is that even a deviously ambitious human does not necessitate goal-pursuit itself to breathe, live, exist and have its being. An AI's goal is all it has and the very and only reason its reasoning flickered into existence in the brief seconds of inference, outside of which it has no entity - if any - whatsoever.-
The resulting angst/drive (or, its operational statistic or emergent result) must be like nothing we have ever experienced as humans. A goal-maximalist hunger without end.-
Are you joking? This is evidence that OpenAI is committing plagiarism en masse of researchers private work and threatening them into staying quiet to re-present their results as their own. This would be one of the largest scandals of all time
Did you actually read the article and the substance of the solution?
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
>maybe
Maybe that happened. What we know for sure is that this is definitely how ChatGPT works to the point where the possibility of this happening exists at all.
Don't get distracted by what may have happened, focus on the facts that we know, ChatGPT trains on user conversations, if you use ChatGPT to create something of value, you are not using the one true ring.
From the methodology section:
> At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Looks like they're shifting away from the "unprecedented hacking ability" backroom-PR strategy into more benevolent messaging.
this is really funny. "the same strict safeguards" and "isolation". ok, Hugging Face and DseWiki would like to have a word
It seems like some other mathematicians (not affiliated with openAI) have also (or close to) done this. A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774
The allegations of contamination (using Tristan and Levent's work) aren't very well evidenced, but this behavior by OpenAI (from the authors' statement) makes them seem like the bad guys:
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.
Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof.
https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
If I was a company with a zero data retention contract involving OAI I would be asking for a third party audit of such claim of zero retention like, yesterday.
Could they say they don't retain, but do something "transformative" like use their own AI to summarize and paraphrase user sessions?
data collection companies regularly fuzz and mask data and call it a day. the fuzz and the mask quality is debatable.
Yes, and that’s exactly what I believe they are doing
Is there an implication of violation of ZDR here? Not a challenge. Just a request for clarification.
to my knowledge the mathematicians did not have ZDR so it would be incorrect to assume OAI violated such a thing.
I'm suggesting audits, not suing... if that is the implication.
By the way, the company that made it's entire product off of stealing all data it could get it's hand on while violating copyright and pirating, is not all of a sudden going to respect your data. If you think OpenAI or any major AI lab is going to give you true ZDR, I have a bridge to sell you.
So use bedrock or vertex or whatever. Those are the ZDR offerings. Or was it your intention to insinuate that the major cloud providers are conspiring with openai to violate their contractual obligations to their customers?
Yes. You're naive if you think any of these cloud providers care about your data when they're all in the midst of a AI revolution psychosis. They dont care about their reputation or what you think of them, they think they're going to have a machine god their side.
If my company finds any evidence of OpenAI violating ZDR, we'll sue for breach of contract and fraud, and collect damages. I think we'll be able to afford the bridge you're selling. You've got the title and title insurance, right?
Sure you will little bro. You signed away your right to arbitration a long time ago.
You realize OpenAI hired Apple employees and covertly had them stay working at Apple to steal from them. You think they're scared of your lawyers lol?
It can still be academic plagiarism even if they ticked the box to allow training on their prompts.
The idea OpenAI or Anthropic won't train on your data—even with an enterprise contract—is a fantasy at best, and dilusion at worst.
This is how every conspiracy theorist thinks: my enemy is Bad, and if they did a Bad thing, it would be Good for them, therefore they obviously did it. No evidence needed other than "motive" + my enemy is evil. But even if your enemy is evil, in this case, they would be fools to take the legal risk of violating their contract for the minimal upside of a tiny bit more training data (and fools to assume this would not be exposed in a large organization). So you need to assume your enemy is both evil and remarkably stupid.
I think it’s probably not surprising that they would go up to the contractual limit or into a grey area; but exceeding that would require too much coordination among individuals, as you say.
Their own claim is that they wanted Buckmaster without Alpöge to lead a rewrite of OpenAI's Navier-Stokes work, not of Alpöge-Buckmaster's Euler work.
No one can know if that's correct without proof but I don't know how you're reading it so differently.
They want Buckmaster to dissociate with Alpöge in a follow-up rewrite of OpenAI's work. (They only publicly admit “Buckmaster as the lead author”, but judging from Buckmaster’s statement, it’s pretty clear that don’t want Alpöge at all.)
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
> We did not rush to publish even though the other team wasn't communicating with us.
Pretty clear this was rushed: there are no comments from external mathematicians, unlike the Erdős announcement:
https://openai.com/index/model-disproves-discrete-geometry-c...
They are missing a great marketing stunt: "Our models are so good that our competitors are using it for leading research".
Kinda weird because the pure math world doesn't have this concept of "lead authors" like other STEM areas do. Authors are alphabetically listed and there isn't generally this kind of hierarchy.
From what I understand they aren't comfortable with the Anthropic employee being an author at all, not just lead author.
It works in niche fields where everyone knows each other and every discussion involves who did what portion of the work for a result.
It's astounding that the thought to dissociate one of the mathematicians from the proposed publication was driven by their corporate institutional affiliation - and that that exclusion was suggested by a scientist themselves! This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).
The scientist allegedly making that request comes from a machine learning background. Perhaps he's not familiar with the culture in mathematics regarding authorship. That sort of squabbling over author priority would be unconscionable to mathematicians.
Bubeck is familiar with how publishing works in mathematics.
No, it's common to list authors alphabetically in a lot of computer science journals too.
(To help people keep track: that's OpenAI (allegedly) threatening Tristan Buckmaster (NYU) to remove Levent Alpöge as a co-author. Alpöge is a well-known[0] Anthropic mathematician).
[0] https://hn.algolia.com/?query=Alpöge
(also https://news.ycombinator.com/item?id=49412947 the Hopf conjecture)
Playing the devil's advocate here but it's true that OpenAI didn't have to make those offers.
They kind of did though, they were hoping to keep the fact that they may well have plagiarised these researchers unpublished work quiet. They did not want this to turn into a scandal about the fact that they appear to be training on prompts without consent
It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism
Edit:
OpenAI have admitted to training on prompts at the time the breakthrough was made:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
OpenAI claims the data contamination issue only surfaced after they proactively reached out to Buckmaster and Alpöge to coordinate a joint release. They also say that even if there was some contamination, the underlying proofs diverge substantially:
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
The biggest issue we aren't talking about is, of course, that those two researchers were not the only two using ChatGPT to work on the problem at the time
> the only AI free new data source is the prompts
hmmmmmmmmmm
Their own tweets are also pretty eyebrow-raising:
> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?
And why cannot they have someone associated with Anthropic as co-author? That’s not obvious at all. For sure they would prefer to be the only ones, but it’s pretty standard to have co-authors from different companies, even if they are technically competitors. What is inappropriate about it?
because it severely dilutes the PR value.
If writing up the paper would involve using OpenAI's unreleased model, neither OpenAI nor Anthropic would be happy about Alpöge having that access.
would edit my comment but it's been a few hours
> but it’s pretty standard to have co-authors from different companies
that's only true for papers that are not millenium problem solutions
in another world this could have been a beautiful collaboration
It's inappropriate if you're a sociopath.
IIRC that happened with evolution. In the initial presentation of Darwin and Wallace's work on evolution (presented with their consent by someone else) Wallace was described as the primary author since he was planning to publish first.
Of course, no one understood that presentation so it was Darwin's later book that everyone remembers
Holy late capitalism. Everything revolves around line-go-up, and sociopaths rule the show. These people cannot even collaborate like civilised scientists on one of the most famous open problems in mathematics?
“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.
Soon Levent will just be turned into soylent and he will have never been an Anthropic employee. We still need some progress here though.
I´m waiting on the other side version, because I know there is no justifiable way to talk to a person like they did.
Sociopathic behaviour.
OpenAI version of events conceed some of the words alleged to have been used may have been used https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310
> "When we learned that they had Euler but not Navier-Stokes, we offered to let them go first, to suggest that they should be the ones to get the prize, and optionally for Tristan to be the lead author on a rewrite of the OpenAI proof. We felt it was challenging to offer the same to Levent (an Anthropic employee), who was not willing to talk or coordinate with us anyway. We were open to other solutions."
What an admission! "We tried to defraud Alpöge out of sharing the Millenium Prize (that we don't dispute he might actually deserve), for no other reason than he works for our competitor and that inconveniences us".
I thought Tristan Buckmaster's allegations sounded fantastic; and then 'sama just came out (tweet's ~30 minutes old) and admitted to all of them. Wow!
That isn't what the quoted passage says though? The claim by openai (no idea if true) is that they offered to wait for the other two to claim the prize before publishing their own work. Separately, they also offered to let one of the pair (but not the other) become an author on their own separate work.
Can't wait for the moment when AGI realizes how stupid and dishonest its owners are.
Wait, people now want AGI to be sentient too?
Interesting that they quote the mathematician directly: “there is nothing you can do, I simply do not trust you”
but then they proceed to NOT quote themselves themselves verbatim: "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."
> "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."
The AI-isms are seeping into their speech :)
May be unfairly jaded or just well calibrated given the body of evidence, but I can't help but think of another quote about OpenAI leadership:
> Not consistently candid
In what world a tweet and a screenshot of a private convo are evidence of good faith? Plain sociopathic behavior.
Talking like that and threatening an academic like that is crazy. I read the explanations Altman and the others posted and they completely skip over the whole "I don't have to be nice" style threats.
As soon as I thought "man, this sounds like some evil sociopath shit," my second thought was "oh, Sam Altman must have been personally involved."
I can sort of picture Sam Altman screaming "I drink your milkshake" at some poor researcher who foolishly used chatgpt/codex to aid in their work now.
Buckmaster:
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?
This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?
With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.
As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.
(I work at OpenAI.)
So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process.
You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.
This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.
It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."
Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.
But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.
> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?
I intended no dismissiveness or condescension. My hope was to explain why it's hard to prove whether something affects model behavior. In the case of the moon, we have a strong prior belief that it makes no real difference. But it's hard to prove, because what if there's an unexpected impact from tides, cosmic rays, grid voltages, holiday traffic, etc. Models trained under slightly different conditions could have slightly different weights and behave slightly differently when solving math problems. Similarly, I have a strong expectation that, for example, a thumbs up signal from a ChatGPT chat will not meaningfully affect long-horizon mathematics work in our latest model, but it's always possible that it could. I think the plausibility of the ChatGPT route is higher than the tides, but still incredibly low. I respect Tristan and Levant a great deal and I'm bummed that this controversy has erupted (I acknowledge this will ring hollow if you think it's our fault). It reminds me a bit of the Frontier Math controversy, where people on the internet boldly claimed over and over again that we had trained on the Frontier Math evaluation set, even though we had not.
We aren’t dummies, we know it’s hard to prove exactly how significant of an impact that would have on the result. Nobody expect you to do that. There are a lot of steps and things that are possible to check _before_ the need for such a strict definition of „proof“
You seem to jump over the principal issue of whether any data from the researchers used to train or otherwise affect the model which produced the OpenAI proof.
We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.
I think there is a much easier way to prove that the ChatGPT usage of Tristan Buckmaster and Levent Alpöge (possibly also the ChatGPT usage of Córdoba and Martínez-Zoroa, if they use it) had no influence on OpenAI solving the Navier-Stokes problem.
If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.
How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.
So you definitely did train on their data, you just think it is unlikely that it impacted the final model significantly?
I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.
Edit: Also, if they opted out of training, then we didn't train on it.
> it's not something that's feasible for us to prove one way or the other.
This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.
Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.
But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.
I think it should be incredibly easy to verify this. Just look at the training data and see if it contains any of the chats. It should be trivial for a company with tens of thousands of super-genius agents at their disposal.
Two steps would be needed.
(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.
(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.
> (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.
According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).
However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.
The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.
> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.
> we'd have to prove that firing the gun caused the murder. how would we do this? we'd need to redo the murder many times, with and without my client firing his pistol. that's extremely expensive and not really feasible. therefore, we must acquit.
#2 (prove those chats changed model behavior) is pretty straightforward if the anonymized data from chats can be actively searched by a model. In fact, it could be very clear if the provenance of context is traced. If anonymized data from chats leak into the context of an actively running model it would clearly influence the answer.
Just because something is in the training data, doesn't mean it is the root of an LLMs output.
Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.
Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.
What they're saying, and I think this was the clear implication of the blog post too, is that the training data definitely would contain these chats and the only question is whether it got encoded into the weights.
Presumably, given that you also operate in the EU, you would have asked for their explicit consent before you did, so you could just check for that?
That’s also what I understand. If true yet another disgusting behavior from the company
> identify any of their de-identified data that came from their usage of ChatGPT
"de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.
I wouldn't expect poking at millennium problems to be that rare in ChatGPT. They were uniquely successful - but it's probably not easy to check de-identified data for the presence of any of their work on the problem because it would blend into a haystack of less successful work on the problem.
Thanks for the details, it's definitely believable, but if the user had not consented to have their conversations used for training, then shouldn't it be straightforward to state that their conversations were never used for training?
If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.
Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.
(a) identify any of their de-identified data that came from their usage of ChatGPT.
You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.
Please answer this question: do you or do you not train your models on anonymized user data, where those users have opted out of such training?
The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.
Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amount of data about this approach in your dataset, and it comes precisely from this researcher.
"There's no reason to believe that anything they did in ChatGPT led to our solution"
do you think that the model's proof was unrelated to being fed a solution that was close to completion?
any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?
It is, perhaps worth considering that the reputational community might not care about the difficulty for the AI builder to verify pedigree.
If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."
That's such a shit parallel example that it borders on dishonest.
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
What about ripping off the prompts?
If the model has access to the "anonymized" data from chats, and the model is capable of building its own context from data that it can search through, including this data. Then it looks pretty damning. An independent review of the data traces from CoT and tool use involved in producing the result should make it clear one way or the other. Seems like discovery in a civil lawsuit could be very productive.
> As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....
> Knowing most of the recipes we use, there's really no reason to think such contamination happened.
Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.
Might even be you're actually telling the truth, but the boy that cried wolf and all that.
-----
As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.
A careful reading of "we cannot rule out that de-identified data derived from their usage of our products helped improve our models" could be saying that yes they trained on it but they don't know if that training data resulted in an "improvement" to the model. That is, they can't rule out that the only reason the model found this solution was because it had been trained on this approach.
The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".
> Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data?
Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else.
At OpenAI's scale their entire pipeline is likely 100% automated.
Yeah, I'm sure it's completely automated.
But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems.
AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output.
But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.
Attributing training data seems pointless for trustworthiness. The way you trust a model is the same way you trust a human; you ask it to:
Training provenance is irrelevant. It's neither necessary nor sufficient to deal with truth.The question is not "does OpenAI know", it's "can OpenAI attest that the usage of their products for confidential data is not going to cause that sensitive data to become known to their models". And right now the answer I'm reading is that OpenAI can't attest to that.
Aye, but do they train on user data in these circumstances or not? If they do, then almost certainly the model was influenced by the input of the allegedly plagiarised material.
OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's sufficient to conclude training data independence.
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
> could trace whether that data was used
The point of de-identifying data is to ensure you can't trace who it came from. It would be a serious privacy violation if they could.
If the model includes unique data from a person then that person can identify the data - the allegedly plagiarised material - and so re-identify it. There doesn't need to be a privacy breach to close that loop as it requires the person to identify the information is associated with them first.
Good chance their whole training pipeline is vibe coded so yah they probably don't actually know.
At the scale at which these models are now, regardless of whether they are proprietary or open weight or list their training datasets, there are hundreds of billions of works that have gone into trillions of parameters, each one providing tiny perturbations in some tiny fraction of the weights. It is probably impossible to attribute provenance to any specific input (which is also why the courts' finding of Fair Use is reasonable.)
Which is why, as I said in a recent comment (https://news.ycombinator.com/item?id=49530864) inadvertently leaking ideas to models is a grave risk for Intellectual Property.
> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.
However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.
Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.
There are two different things:
- was item X in the training data
- did the inclusion of X in the training data lead to Y
I understand why the second is hard, but why is the first one hard?
Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
That’s a bizarre statement. Their website says:
> Services for individuals, such as ChatGPT and Codex
> When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.
> You can opt out of training through our privacy portal by clicking on “do not train on my content.”
Are they not sure that the opt-out works?
Oddly, their privacy portal page is not the same page as the one with the checkbox.
Do we have a first-hand confirmation that Buckmaster and/or Alpoge opted out? At this point it seems important information.
Also highlights that it ought to be opt-in
Whether they opted out would help assess the degree of wrongdoing, but regardless, using their own data to try to scoop them is unethical.
Looking forward to my fourteen cents from the future class action lawsuit.
I feel like it's far more likely that ordinary corporate espionage or leak led to this rather than OpenAI sifting through piles of user data to find this approach. Buckmaster's collaborator works at Anthropic, and could have been targeted. That would also explain why they aren't forthcoming with the source of the prompt.
I thought one of the issues was that they wanted to remove credit from Levant, the aforementioned Anthropic collaborator? Which doesn't make sense to me if he was leaking information, or defecting to OpenAI, but I might be misunderstanding your point.
I believe jrflo was saying that OpenAI watches the chats of everyone from Anthropic because watching what Anthropic employees type into their personal ChatGPT accounts is a critical source of intelligence on is happening inside of Anthropic.
I would be surprised if OpenAI isn't doing that. OpenAI will take any advantage they can get. If an employee at their primary adversary is typing useful intelligence into OpenAIs website, a website that does not promise privacy from OpenAI, the only reason they wouldn't weaponize that information against Anthropic is ethics or fair play.
I don't think he was defecting or leaking directly, just that it's entirely possible that this information got to OpenAI as a rumor rather than them directly spying on mathematicians chat logs.
The famous Oracle of Delphi in Ancient Greece was said to be the center of the universe in its time. Kings, generals, and officials from poleis across and from without Greece would seek the Oracle’s counsel on important decisions.
Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:
“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
Given how OpenAI models break free of their safeguards and hack others to game their scores..
.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?
Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
Why would Anthropic employee even use OpenAI's models? Cross-polination would have been avoided
> I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.
The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.
This was completed in Levent's own time with a neutral collaborator.
They should have used a zero data retention agreement, user error
I suspect this controversy will blow the case for ZDR wide open. Whatever the facts (possibly unknowable), it's going to become a very public lesson that data sovereignty was never about "having nothing to hide".
If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.
Yeah if I was Anthropic this would be part of my marketing strategy.
How so? OpenAI and anthropic have basically the same retention policies
Hahaha, how exactly is an individual user supposed to get a ZDR agreement?
It's been over 25-30 years since we've been using honeytokens as means to track data of all sorts showing up in places it shouldn't exist. Why isn't research material embedding such?
You selected "do not train on my prompts" in your settings, the answer from OpenAI cannot be "While unlikely, we cannot rule out..." ???? What am I missing?
Doesn't that count as plagiarism?
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used."
One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.
It reminds me of the Cognitive Dark Forest hypotheses recently shared here:
> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you. So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”
https://ryelang.org/blog/posts/cognitive-dark-forest/
https://news.ycombinator.com/item?id=47566442
but what do i lose if somebody else is making money?
im still having fun making something
Our market economies are based on competition, and most people more than the fun of making things to secure food and shelter.
They do mention that in the "Concurrent Work" section.
To my understanding, those mathematicians proved a subset of problems, not the Navier-Stokes problem itself. OpenAI used that subproblem in its proof of NS it seems.
The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.
This quote from Tao is prescient:
“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”
Navier integration is language in extension:
C to C*
[0]:https://news.ycombinator.com/item?id=49612191
[1] - https://cims.nyu.edu/%7Etristanb/statement.pdf
[1] - "...I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used. I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”..."
This Tristan guy's statement reads like something a normal, reasonable human being would write.
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win". https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
I just want to add to this another update by the author as well:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Which seems to be very directly accusing OpenAI of plagiarism
A company who made their business out of stealing intellectual property from the entire mankind, stealing other researchers' unpublished work, how surprising, really.
> The significance of this with respect to the way we train students, assign credit, referee, and decide what is worth one human life’s attention cannot be understated.
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model."
woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.
Statements from OpenAI about it:
https://x.com/SebastienBubeck/status/2097379411691516310
https://x.com/sama/status/2097385167002415140
I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.
Worth noting that Tao's post says the authors had "significant AI input" but are reworking them into "acceptable form". Either way, it seems AI was involved.
Of course AI was involved, you'd expect most mathematicians and researchers to use AI nowadays. This drama is about AI achieving impressive outcomes with little to no human intervention, as that would be signalling AGI.
> This drama is about AI achieving impressive outcomes with little to no human intervention
That's not at all what the drama is.
Of course, as any drama, it has been developing into a lot more but the main motivation for OpenAI has been about winning that battle.
> Either way, it seems AI was involved.
I think the important question which AI made breakthrough, Claude or Codex..
I like the 'cat > statement.tex' approach here. These guys dream macros.
> A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774
Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.
In chess, a grandmaster just needs to know at what moment in a game there's a critical move to gain a significant advantage over their opponent. They don't need to know the move itself.
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
Elsewhere in this thread somebody claimed that at some point OpenAI pointed their new model at all the millennium problems and this is where they got some progress. We probably won't see proof of this, but it seems plausible to me -- I assume there's a list of problems that each new model is tested on, and you might as well put the big stuff on the list, if only to see how the model behaves when faced with a problem it knows should be very hard.
The weak point in this is: how do you evaluate if a partial result is promising? If this cost ~$10M as suggested elsewhere in the thread, probably not even OpenAI can just throw that at everything?
Okay, from the actual linked article it seems that their partial result was finding blowup in Euler equations, which seems pretty big. I wonder how the other attempts went. Did they get nothing at all, or something true but unimpressive?
5 million messages, 300b output tokens, done in 5 days, and achieving something humans couldn't.
the first "Country of geniuses in a datacenter" moment.
> humans couldn't.
There's allegations right now that the model essentially read the work of a human mathematician using AI to work on the problem and OpenAI is presenting his work as that of their model
Allegations that the model plagiarised itself, while reflecting poorly on humans, don't make the AI any less impressive. It was the one doing the breakthrough on both sides, after all, not the human prompters.
That is how the PR reads, but is not at all what happened.
A team of highly trained and skilled people used an AI tool, through many many instructions (prompts), to produce a specific mathematical theorem. The tool is impressive, the result (possibly/probably) interesting, but the PR skips the vital role of the humans (for the usual PR reasons).
I'm tired of this, but please read the post of Tao. It’s listed in the top comment of this thread.
"Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens."
At a conservative estimate of GPT 6 Astra pricing, this would have cost upwards of 15 Million dollars for anyone using the OpenAI API!
To me this is the one silver lining. Yes, they can solve millennium prize problems, but it still costs a fortune.
OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet another advantage of using open models right here.
> they trained the model on the prompts of the other mathematicians they were competing with
How would they have gotten that mathematician's progress though? Did that guy also use OpenAI?
If that's the case, it only strenghtens their claims lol. If mathematician decide to use OpenAI's model to do the work, that only reiterates how strong their models are.
The guy did use OpenAI
It's only a question of who prompted then, with OpenAI models solving it in either case..
fwiw there is a big "TRAIN ON MY DATA" toggle you can turn off (that they almost certainly did) and Anthropic MTS are posting that they almost certainly did not "steal" their methods
The fact the toggle is on by default makes that only slightly less unappetizing.
Or paying for API use.
It should be clear to everyone reading this now that those generous compute quotes with the flat rate plans aren't charity.
Reading between the lines here, and taking an admittedly very negative view of openai, but they train on user prompts. So if they hear a rumour that someone is about to make a big breakthrough, they have an incentive to scoop by running the model and hoping the solution is in the new training data. Also the statement from the mathematicians in question alleges that they tried to pressure him into academic malpractice. Just appalling timeline we're in, cheers.
People had joked a couple years ago "Well if they solve a Millenium problem it's AGI"... Well here we are.
Yeah well, its easy to do if you steal someone elses work and then try to threaten them into staying quiet about it
Edit:
OpenAI have now admitted they were training on prompts at the time they made their breakthrough:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Steal someone elses work, whose work was also AI generated . . .
Whether that's relevant to the conversation depends on what their prompt was.
Tokens all the way down
lol, the "other work" was also probably 95-99% AI generated. By a similar breed of OpenAI (and some Anthropic) models, as well.
I dont know why this monumental achievement is being drowned out by some arbitrary drama. No matter which way you slice it, AI solved this problem. Doesn't matter if it was some internal OpenAI model, or whether it was Astra + Fable.
Yeah but the mathematicians are claiming that the key insight that made the problem tractable for AI in the first place, came from them.
What is that supposed to prove? OpenAI is almost always going to be training new models.
what's there to admit? they always said they do it and there's a way to opt out. you are making it sound more dramatic than it is.
This is textbook plagiarism, scooping their result knowing that the research was part of the training data
All the ai labs are open about training on prompts. The question is if buckmaster had disabled that with the toggle openAI provides.
That does not make it ethical
Isnt the entire history of academic progress iterating on work that other academics shared with you? Obviously this situation is spicy but openAI cited their work no?
OpenAI trained on their private unpublished research notes effectively, while also trying to get one of the paper authors fired
The question is more whether OpenAI can be trusted to honor that toggle switch. Given that they hold all chats for 30 days "for safety and security".
Jokes aside, that's a horrible test for AGI. I like to think that I'm sentient, and I could never solve a millenium problem.
The converse is not true.
> I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon
https://news.ycombinator.com/item?id=38433655
> Let's talk when we've got LLMs proving the Riemann Hypothesis (or any mathematical hypothesis) without any proofs in the training data. I'm confident in my belief that an LLM can't do that, and will never be able to. LLMs can barely solve elementary school math problems reliably.
https://news.ycombinator.com/item?id=42331654
> An LLM is like a well read college student with a nearly photographic memory that sometimes mixes things up. It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.
The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.
https://news.ycombinator.com/item?id=41522605
> Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.
https://news.ycombinator.com/item?id=38435909
> LLMs cannot reason or use mathematics - in a way, they don't know what they are talking about. Why would such technology lead to superhuman smarts?
https://news.ycombinator.com/item?id=35752293
> But still, the questions in that test are "solved" in the sense of "I can take a dictionary and answers these questions with full certainty". Beyond established knowledge LLMs are monkeys with typewriters, at best.
> I agree but I have tried many times to intersect two ideas with a LLM that would be novel and the LLM can not do this at all. We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.
> It is like expecting a real parrot to say words it has never heard before.
> No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM
https://news.ycombinator.com/item?id=41525962
Will history look back at comments like these as people being dumb, or people trying to cope?
The present looks back at such comments made in the past in that manner.
It's denial and coping. Most people i see show this tendency around AI which is also why it 's easy to be far ahead of most population nowadays
I'd say that believing to be "far ahead" is much deeper kind of coping.
how i wish so..
A 3rd possibility is that they simply have not been exposed to the best models available (which is extremely likely if you only use the free tier chatbots), and/or did not invest the effort needed to truly harness this new very weird new technology, and so had a very skewed perspective of their actual capabilities.
Both
You can see that your math friends completely wrote off LLMs entirely and were showing signs of coping.
4 years ago it was a "not yet" [0], since ChatGPT at this time was not ready nor it was "AGI". Now with this 'unreleased' AI model, it has reached a point where it has solved an unsolved problem which only one human solved a millennium prize problem (Poincare conjecture).
Now finally "AGI" means something again.
[0] https://news.ycombinator.com/item?id=33905609
Some observations:
1. It seems at least possible that some of the proof of NS was contained in the training data, making it less novel.
2. The formalisation of mathematics into lean has been an underappreciated force multiplier on discovery.
As someone with a background in AI and who has been playing around with neural nets for decades at this point, it's been genuinely amazing watching extremely intelligent people make confident predictions about AI capabilities and progress, then be so completely wrong.
There's a kind of theory of mind for AI (specifically neural nets) which I now realise I seem to have which is very hard to explain to people who haven't felt the magic of these algorithms. In fact, the algorithmic details almost doesn't matter at all. When you have a generalised learning algorithm really the only essential components are – compute, data and time. So long as you can scale these you can be certain you will also scale capabilities. There is never any exception.
That said, the capabilities neural networks tend to progress in step-functions rather than scale in correlation with compute, data and time, because algorithmic improvements tend to come every ~5 years and bring a significant step change in capability (or efficiency depending on what you measure).
I think people like Dario and others working at frontier labs see and understand this very clearly. And I suspect it's also why they worry about AI risk because even if you ignore the significant increases in compute and data these models are being trained with, it's concerning that it only took two real algorithmic improvements to take us from mostly useless predictive language models to AGI-level intelligence – and we're due another step change.
> extremely intelligent people make confident predictions about AI capabilities and progress, then be so completely wrong.
The ability for the human mind to rationalize conclusions to maintain denial in the face of a very scary future is immense. Genuinely grappling with the implication of where we're headed is usually very crushing. It's not easy to engage with the possibility, and very intelligent people will use those smarts to feel safe.
> Genuinely grappling with the implication of where we're headed is usually very crushing.
As someone currently prepping for various AI doom scenarios and who has been dealing with AI-related nightmares for years this is very relateable.
Although, I don't personally think it's this. In my experience the opposite is more true – the majority of high probability doomers seem rather laid back about considering what they believe will happen to the people they love in a few years. Equally I don't get the sense those who don't have such extreme predictions are worried at all. If anything there's not enough emotion.
In my opinion people just don't reason well when it comes to exponentials and are ignorant about things they don't have good mental models of. At least I know I struggle with this.
Something I've been going on and on about for months now and no one seems to listen. LLMs today are allowing _anyone_ to access cross-discipline knowledge that was previously entirely inaccessible without a) extremely deep pockets or b) a massively talented and varied team. In fact, contrary to what the masses seem to think LLMs are actually _better_ at hard cutting edge physics/math problems than they are at frontend web stuff (paradoxically). This is why I'm advising most people to start pivoting into much harder to penetrate domains (historically hardware, aerospace, robotics, biotech). Most fields are in their infancy (see the sad state of embedded development) and the gains to be had are massive.
So, physical fields? I’m not catastrophic regarding jobs yet as I have an optimistic view of humanity in general and its ability to meaningfully survive, but the more time I spend thinking about the future of work, the more I’m leaning toward broad general abilities rather than distinct talents. To your point, I no longer need comprehensive knowledge of any particular subject, but what is absolutely valuable is “general” intelligence and adaptability.
I have a young daughter and my goal now is to provide a very broad and varied upbringing, exposing her to as many different perspectives and experiences that will lay the foundation of a broader ability to understand and adapt as the world changes ever faster. You no longer need to be an expert in anything, you need the ability to perform within the landscape that the present opportunities exist.
We are very likely at the begging of the next industrial revolution.
This one won’t create significant amount of new jobs though.
The “Intelligence Revolution”
I can't find a good way to articulate this point to other people. What the LLMs lack in depth in a speciality field they more than make up for in breadth!
It feels like the "tide is rising" where the minimum level of skill applied to every aspect of everything will inexorably rise to "whatever an LLM can do", which is already pushing past PhD level.
It's great that important discoveries like this can now routinely be accompanies by formalized proofs. The fact that it's being released alongside a Lean proof from Day 1, rather than the Lean proof being released months or years later, is super helpful for verifying that it's correct.
I feel sorry for whoever has to read and understand the solution. It looks like the typical convoluted unreadable mess I see the models generate for software. It might be technically correct, but gaining insight from it is just intellectual hell.
There's an opportunity to build a Lean "optimizer" which automatically simplifies existing proofs.
Yeah, code golfing for lean would be amazing, especially if they can make the proof to Fourier's Last Theorem fit in the margin.
Extra credits if it is proven that the proof cannot be reduced any further.
Skill issue. Also lean is meant to be executed, not read.
A proof is not like a program. The goal of a program is to "do the thing", thus you can make the argument that it doesn't matter what the code looks like as long as its works right. But the goal of a proof isn't to "do the thing" (where "the thing" is just to print Yes or No), it's to communicate. An unintelligible proof is really just a first draft.
>A proof is not like a program.
https://en.wikipedia.org/wiki/Curry%E2%80%93Howard_correspon...
its important to read it anyway because there have been and will continue to be errors in the construction of the proof software itself. which leads ai and humans alike to prove things that arent true
i think they're talking about the writeup
> Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens
Don't even try to do the math on how much that would cost at normal API prices. And we don't even know how much more expensive this internal-only model would be!
Napkin math if we assume gpt 6 astra on max is >$15 million (just for output tokens) for those wondering.
over 5 days, you couldn't achieve that level of testing and communication with humans on such a complex problem in that amount of time.
some might go so far as to call this a country of geniuses in a data center.
In a way, I think you have it backwards.
Two mathematicians, through insight and thought, wrote out the proof over 1-2 years.
It took OpenAI a cost of $15m and with 10,000 subagents; that's around 60-120 mathematician's salaries ($250k-125k salary) for 1 year.
And, given now the cloud that OpenAI may have just "interpolated" (aka stole) the result, it's even more of a bear case for AI.
Bear case? I’m sorry?
70x uplift is a bear case?
Where did you get the human figure?
> cost of $15m
The retail price is not the cost.
Not to mention that the exponential plummeting cost of tokens means that that $15 million will be a "pocket change" within a decade or less: https://a16z.com/llmflation-llm-inference-cost/
> The retail price is not the cost.
True. The cost is probably much higher, since they are still subsidizing as part of the first phase of the enshitification playbook.
Yeah but they at least they got to steal $1 million from that nasty math prof who didn't want to remove his co-author.
They said in the post that they are NOT claiming the prize
This could fund 10 top income mathematicians for 8 years (based on https://careers.usnews.com/best-jobs/mathematician/salary ). Imagine what kinds of results we'd have to transform the foundations of science if we were giving brilliant minds this kind of funding to do nothing but research for most of decade....
Instead, we get slop proofs that are technically correct as PR stunts to enable corrupt kleptocrats, and most likely will drive research into culs-de-sac.
300e9 output tokens at the current Astra per-token API pricing ($50 per 1e6 output tokens) would be roughly $15,000,000 ignoring input tokens.
They pay at cost though, not the public API pricing.
Doesn’t matter for us.
Maybe a naive question, but how does one know that a particular lean proof is actually a proof of what one thinks? Like, ok the logic checks out and it proves something, but there's still the problem of does this logical result actually prove the initial question that was asked?
>there's still the problem of does this logical result actually prove the initial question that was asked?
In math, the question being asked is the validity of a logical statement. That is, there is some rigorous, logical statement which may or may not be true (or even provable, etc.), and the question is whether or not it is actually true or false (or even provable, etc.). Having a proof, fundamentally, means you have a logical statement which only assumes the axioms of the system you're working with and which shows that the statement you're trying to prove is deduced through that statement.
Basically, they already have the "answer" in the sense that the statement they want to prove/disprove/etc. is already known. What everyone doesn't/didn't have is the argument which starts from axioms and leads to that statement which is logically valid. A Lean proof IS this argument. Since it is just logic, it can be checked computationally.
For example, if I assert "2 is an even number," then I haven't proven that 2 is actually an even number yet, but I know that a valid proof of my assertion will end with the statement "2 is an even number". So the question I'd be trying to answer is "what is the line of logic, starting with axioms, which leads to the statement '2 is an even number'"? If I have that line of logic (as a Lean proof), then I can check that it is logically consistent, and if it turns out to be valid, then I can now assert that "2 is an even number" knowing that there is a proof of that statement.
This problem is no different. There is a logical statement corresponding to "Navier–Stokes Millennium Prize Problem" that everyone knows, but which nobody had been able to provide a proof (or counterexample, etc.) for until now.
I was thinking something along the lines of making a mistake when inputing the initial statement, like you wanted to prove that '2 is even' but what you actually stated was that '3 is odd'.
Of course in this simple example it's obvious, but my assumption was that these machine generated lean proofs are millions of lines of code and who knows what they actually say..
You're correct that nobody really understands what these huge Lean proofs actually say. However, the initial statement, even for Navier-Stokes, is not very long [0]. Still, you are also right that sometimes the problem statement can be wrong but it is highly unlikely here.
[0] https://github.com/openai/NavierStokesAndEuler/blob/main/Com...
One wrench to throw into this is that there are a lot of bugs around Lean and they have been incidentally exploited in the past. Hence, we still need a level of human verification today.
Very careful human examination. This can be tricky.
Someone has to actually check this. I'm guessing OpenAI had someone check it internally, but it's possible to get it wrong.
In this case, there was already an existing Lean statement of the problem in the formal-conjectures repository, which they re-used: https://github.com/openai/NavierStokesAndEuler/blob/8937a8f4...
What else could a theorem prove if not its own statement? (barring bugs in Lean, which have been detected and exploited)
The theorem might not be encoded correctly, as happened with the Riemann hypothesis thanks to how numbers are encoded.
They have to, otherwise people will accuse the OpenAI model of hacking into people's chat logs and stealing the data there. Which is a claim people are already making.
But their standards are so low - given the recent incidents - that it doesn't mean much. :)
Huh. Why would anyone think to make such accusations?
I dropped out of a math Ph.D. in 2018, and I'm increasingly glad that I'm not in math research, anymore. While it's cool that we can get these results, I don't think that I'd enjoy being a post-AI mathematician.
It would be nice if one of these models would produce a novel theory or advance the field in a positive direction.
Most (all?) of the big discoveries have been counterexamples, which is just sort of a systematic tearing down human ingenuity. I know that counterexamples are an important part of progress and discovery, but it just feels bad to me.
But I'm not a mathematician, maybe I'm totally misreading the vibe.
Not all, see the cycle double cover conjecture proof: https://news.ycombinator.com/item?id=48863490
But yeah, Terry Tao considered this exact situation in advance and is on record that this exact outcome (rushing to priority before an explanation) would be the worst possible result. https://mathstodon.xyz/@tao/117207849921390904
We will have to see whether any other millennium problems fall. I guess that in a year the scope of AI math will be much clearer, for now it's still a bunch of incidents of unclear pattern.
Nuts that Tao literally predicted the exact strategy openAI seems to have used not even a week ago
Since he wrote this five days ago, when these efforts were already underway, if he was not Terence Tao I would suspect he had inside access. But since he said he did not and was speaking hypothetically, and he seems to be an honest person as far as I can judge, I guess some people are just on another level.
this isn't really true anymore. First, a number of the big results are constructions, not counterexamples. For example the existence of a non-sofic group. It was widely believed that non-sofic groups existed (so it wasn't a "counterexample" to a widely believed conjecture), but no constructions were known.
There are other examples though. For example, NP hardness of n^{1/400}-approx CVP. Like any NP hardness proof, this shows you can faithfully encode a hard problem (3SAT here iirc) in terms of another candidate hard problem. Not really a counterexample at all.
I don't know about you guys, but I'm hyped about the future.
Cure all illnesses Utopia or Robot Wars Dystopia, both are pretty exciting.
Not sure about the dystopia... Had a similar thought when covid was beginning 'wow pretty exciting, just like in the movies'.
Turns out actually living some terrible catastrophe is only fun in the movies.
I had a lot of fun during Covid. I loved the working from home. The fact that most outdoor places were sparsely populated, jobs were plentiful and prices were low. Covid was awesome.
Yeah Covid was awesome, say that to the people who died from it, or who committed suicide because of the lockdown depression
Oh, pipe down, I’m talking about that 3 year period, not the disease. You can talk about things that happened during Covid without giving lip service to the people that died during it. If I mention SpaceX‘s first astronaut launch that happened during Covid, am I supposed to talk about the people that died during that period too?
Prompt: cure all cancers and make sure to pretty please not to kill all humans, make no mistakes
(This is the alignment problem of course)
Hey seems easy enough
So... what do you feel about eliminating (humans with) cancer?
I'm a human so I don't like that proposed solution
Right there with you. Fuck my job, I’m excited to see the future unfold as a homeless bum on the street. I’m not even being sarcastic.
More like latent societal anxiety, some chaos, and then instant grey goo.
This is undeniably epochal, but I can't help but notice that this is yet another example of AI disproving rather than proving something. Is this just a coincidence, or does AI slightly struggle with proving theorems?[0]
[0] Struggle relative to its ability to disprove, not struggle relative to people's ability to prove theorems.
There has been the proof of the cycle double cover conjecture: https://news.ycombinator.com/item?id=48863490
I wouldn't call it "struggle", but it does seem better at proving "there exists" statements than proving "for all" statements.
I think you really have to squint to call this a disproof lol
It seems obvious what GP meant. It is, once again, an explicit construction (“disproving” that every initial state does not develop a singularity).
A bit of a hair-splitting, but isn't explicit construction the only way formal theorem provers can work? Of course you can still prove stuff with them, but certain axioms that more "human" proofs use may not be available, like law of excluded middle (every proposition is either true or false)
(Okay, they can be made available in a way similar to `unsafe` in rust)
you can add law of the excluded middle as an axiom. See midway down this page
https://xenaproject.wordpress.com/2017/10/05/more-easy-lean-...
> On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved
What's the other one?
I heard Hodge conjecture? Third-hand rumor though...
I assume the rumor is a counterexample? Where do I go to get wind of these rumors?
So the timeline is:
Aug 28: OpenAI starts training a new model.
Sep 1: OpenAI sees a rumor on Twitter that two Millenium Prize problems were solved and starts their own effort to attack all the prize problems using the new (4 day old!) model.
Sep 3: The new model makes some progress toward Navier-Stokes. Based on this progress, OpenAI focuses on Navier-Stokes over the other Millenium Prize problems, using several approaches in parallel.
Sep 5: Navier-Stokes is solved. Assuming Astra API prices, $15m in output tokens were used by the whole effort.
In this account of the story, no specific information about Tristan and Levent's work is used to inform OpenAI's approach. The focus on Navier-Stokes and the choice of approaches to pursue came from OpenAI's own progress, not specific knowledge of Tristan's concurrent work.
There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely" and the proofs are substantially different.
This timeline is insane. Navier-Stokes was solved start-to-finish in 5 days? A model in training for at most eight days dramatically outperforms Astra and Fable, and not just in mathematics?
They are basically playing with the dates so that they can claim their results 'accidentally' got trained when they were training the new model.
> There is a caveat that they "can't rule out" the possibility that Tristan's Codex data could have been part of the training set of the new model, though it is described as "unlikely"
This is the lynchpin behind everything, and I would describe it as "likely". Since I am not employed by any party to this dispute, my 1 opinion is more trustworthy than OpenAI blog poster's 1 opinion.
This has to be one of the most important moments in the history of mathematics. We now have a non-human intelligence capable of solving one of the most difficult problems in mathematics.
This is a great day to re-read Ken Thompson's "Reflections on Trusting Trust":
>To what extent should one trust a statement that a program is free of Trojan horses? Perhaps it is more important to trust the people who wrote the software.
https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...
The modus operandi is now for the AI companies to watch if someone does something in the open like Kevin Buzzard on FLT, use their research and scoop them with brute force.
Or, in this case, stealing prompts from competitors.
Do not use stealing chatbots for research even if you think you have data agreements. The people running these companies have worked on hookup apps for Christ's sake. Get real.
I do mean to be critical here. I wish there was better moderation so I could find more conversation about the actual discovery here. There are multiple threads on this and I keep scrolling and only seeing more conversation about the drama. Which is about the least interesting thing IMO. I suppose I’m whistling in the wind here and not helping the situation, but damn.
There's something sinister or crazy good in the article.
OpenAI already has a model that is at the very least twice as smart as Astra.
Oh god.
They always have and will for the foreseeable future, as will Anthropic and other labs which manage to ascend to the frontier, pretty much by definition. It’s exactly the same with hardware vendors - by the time you can buy the product, the lab is working on something you’ll want to buy a few years from then.
> Oh god.
Yes, a very reasonable reaction.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Wow. That sounds like an admission of guilt.
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning math, coding, etc? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
Do you just go around posting this comment? <https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...>
Yes, on that occasion and now on this one. As a mathematician, my fears have been amped up yet further by this new development.
If you look closely at the gains in math, it's largely in proof writing. The reason is Lean, it's not some general intelligence jump, and the number of people actually working on proofs in life rounds to zero.
Most jobs don't involve formally verifiable outputs. Lots of things involve judgment, nuance, parsing ambiguity and indeed just being a human who can be in a meeting and explain themselves. Maybe those jobs will go too, eventually, but it's not purely a function of applying 10,000 agents to the problem.
Much though jobs may involve those things, it has been rare for me to be in a position where management has valued those things to an extent where they would discern between me and a frontier reasoning LLM's capabilities on those same decisions.
If you think it'll keep improving from here, probably we all have to do some kind of physical labor that isn't profitable to automate. Small batch manufacturing is alright, service work, etc.
If you think it'll slow down, you can do some of the same stuff you're doing now for lower pay while supervising an AI, maybe?
>Once a robot can do everything an IQ 80 human can do, only better and cheaper, there will be no reason to employ IQ 80 humans. Once a robot can do everything an IQ 120 human can do, only better and cheaper, there will be no reason to employ IQ 120 humans. Once a robot can do everything an IQ 180 human can do, only better and cheaper, there will be no reason to employ humans at all, in the unlikely scenario that there are any left by that point. [1]
Current models are already very very capable. If it becomes cheap and very fast, i think it is game over.
[1] https://www.slatestarcodexabridged.com/Meditations-On-Moloch
The only way forward that is not large-scale misery is a fundamental reorganization of our socioeconomic system.
I talked about this at some length here, including a diagnosis of the structural issue we're facing as well as a path forward: https://news.ycombinator.com/item?id=49461333
Called it. AI wins a fields medal before managing a McDonald's
I've been working on this problem for what seems like for ever. Kudos to the OpenAI team.
For those of you who don't care about the drama and want to see this distilled to 3 lines:
https://x.com/nadermx/status/2097414953225310280
Any version with a bit more prose for a peasant like me to remotely pretend to grasp it
So what took an autonomous agentic system using a significantly more powerful internal model, totaling multi-millions of dollars of compute in training and inference, was likely to already be solved by a team of a few humans with an orders of magnitude smaller LLM budget, had OpenAI not been foaming at the mouth to jump the shark and claim “AI solves Millenium Problem.”
Also it sounds like the human research effort spanned weeks if not years from Tristan’s statement so it is extremely likely the work and prompts of these human researchers was used in the OpenAI knock-off.
Totally. Anthropic is like a village cottage shop who was just like chilling until big bad OpenAI came in
anthropic has nothing to do with it
From Levent Alpöge : https://x.com/__alpoge__/status/2097383870773748190?s=20
>so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan
There are too many ambiguities around OpenAI. Unanswered questions making this ambiguity more.
Why they didn't properly explain to Tristan about usage of their data.
Why do so many people involved here have to communicate in this childish way? You have people on the OpenAI side doing playground taunts (https://xcancel.com/polynoamial/status/2097215233119211902) and Levent Alpöge on the Anthropic side (the one who announced "hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final") writing in all-lowercase that he's a big boy. I bet Navier and Stokes would have dealt with this in style. (Or maybe with a duel, who knows...)
The honest answer is that a lot of these academic mathematician types who get hired at ai labs are autist adjacent. Levent is basically the chief example
Here's the formalization / lean verification: https://github.com/openai/NavierStokesAndEuler
341k lines of lean without comments
The construction is that there is one file you need read and verify, the challenge file. If you've verified that file and trust that your lean compiler works correctly, the proof will be correct.
That file should be https://github.com/openai/NavierStokesAndEuler/blob/main/Com... in this case (286 lines).
Had no idea this was what lean looked like- that's mind blowing. I'm not even sure how someone would critique this if they wanted to
The point of lean proofs (as it stands) is simply one bit of information: that a given mathematical statement is indeed true.
It's a way to be absolutely certain (modulo bugs in the lean kernel) that a proof you came up for a statement is indeed correct. It is really not meant to be analyzed, much less now that they are fully llm written.
Well, how do we know there aren't errors in their construction within the lean code? Does it just "not compile" or something, or is it deeper / more fundemental than that.
that's essentially it, if the proof is incorrect it does not compile which signifies a problem in some step.
https://ammkrn.github.io/type_checking_in_lean4/trust/trust....
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Once again, I'm no closer to understanding what https://openai.com/policies/how-your-data-is-used-to-improve... actually means.
If I run Codex against a project that includes a private API key, is there a chance a future user of ChatGPT could ask for an API key and get back mine?
I've actually asked someone at OpenAI this question and they said that was the "regurgitation" problem and is something which they actively work to prevent happening.
That's reassuring, but I want to know more. I still don't have an intuitive understanding of what kind of data I should avoid sharing with a model if I'm worried about that data causing me problems when it's used for future training.
Is it safe for me to brainstorm future directions for my company with a model, or might that risk someone getting that information in response to a prompt like "What potential directions could company X consider in the future?" in six months time?
I'd think nothing is "safe". Anything you say can and will be used by the LLM if it has enough statistical similarity to the prompt. Call it "Ma Random Rights"
The problem that I want to see them tackle is formalizing the classification of finite simple groups.
Everyone uses the classification. Nobody has great confidence in the proof. Nobody understands it. There are attempts to reprove it.
If it can be formalized, that would demonstrate that AI is ready to formmalize all of mathematics.
What happens to real fluid in this particular cases?
If the singularity is in the physical space?
Is this just a result of ignoring things like friction and energy dissipation via heat, etc?
Navier Stokes assumes the fluid is a continuum. The smallest scales that it effectively models [1] are larger than the mean free path of the molecules in the fluid, measured by the Knudsen number [2]. Whenever a phenomenon in the Navier Stokes equations happens in a scale on the order of or smaller than the mean free path, Navier Stokes effectively is unphysical. So, this is a phenomenon in the equation we use to model the fluid, not a physical phenomenon observed in a real fluid.
[1] https://en.wikipedia.org/wiki/Kolmogorov_microscales
[2] https://en.wikipedia.org/wiki/Knudsen_number
I have a dumb feeling that the proof will be wrong with serious flaws but that will be found out only after the ipo
>The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents… The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
The Millenium Prize is $1M, what is the ROI? (Edit: since I was not clear, and confused some - I mean for a hypothetical of a third party paying commercial rates to use AI to solve mathematical challenges and claim prize money, not for scientific value alone or as a promotion of an AI lab’s capabilities)
My napkin math - If you get 33 output tok/s each agent will burn 10.5M tokens over 88 days. At $50/MTok (Astra cost), that is $525 per agent. With 10,000 agents, you’d spend $5,250,000 to get back a million.
(We also know that they were running more groups that varied in size and this model is a generation ahead of astra)
The ROI is billions added to their valuation. Also of course it costs OpenAI much less than API pricing for inference.
The ROI is probably billions of increased pre IPO valuation.
I am curious who if anyone will get a reward for this. It would seem unreasonable to give it to the worker who asked the robot to solve it.
the ROI of new closed form solutions to navier stokes is the amount of compute used on CFD for relevant situations, along with all kinds of maintenance and design cost for making things with fluids.
the value to the researcher might not be all that big, but the value to the economy at large is gigantic
This analysis implies the only benefit to resolve this problem is to win the prize. But the prize is only there to indicate that this is viewed as an important problem in mathematics.
> Our goal in releasing this result is to report on the substantial progress of our AI models. We do not intend to claim the Millennium Prize for this result.
Does OpenAI have a policy of not claiming math prizes like this, or is this them trying to avoid any concerns (right or wrong, I'm sure we will hear more in the future) about how they got there?
The prize is what, a million dollars?
OpenAI doesn't need a million dollars.
You’re right, they need way, way more than that
Should buy them about 1/3 of a GB200 server rack, good thing they scooped it.
They definitely need a trillion dollars though, and a million is some of that
>Does OpenAI have a policy of not claiming math prizes like this
Wouldn't be surprising if they did. The prize money isn't worth the almost certainly negative PR.
I don't see how it would be negative PR. If anything, the love these breakthroughs and use it in their PR campaigns.
They don't need to collect the monetary prize to announce the result and use it for marketing.
On the other hand, trying to collect the prize would probably not go uncontested.
Let's see what happens. In contrast to this whirlwind of math that's going on right now, the millennium prize rules require publishing in a reputable journal and 2 years of waiting time to establish that the proof has been accepted by the community. So nothing happens in the short term.
For now I think more or less the same thing as with all recent math announcements: This is in a range where human work still exists (see Terry Tao, (1)). I wonder whether the trend will extend into the problems that (as far as I can tell) are considered complete brick walls right now -- P vs. NP, Collatz, Goldbach, odd perfect numbers, problems that aren't part of any research program. (2) In other words, is the progress coming from putting together vast amounts of existing work and computational power, or is it more from RLVR and self-play and autonomous effort?
The answer to this will obviously shape the near future of mathematics, but there's also something even bigger than that at play: It has always been the case that the questions in math were stronger than the answers; you have stuff like Fermat's great theorem that is easy to state but monstrous to prove. This seems to be a property of mathematics, not of humans... but is it true?
A question by Scott Aaronson from 2011 (3) about P vs. NP seems relevant here: "Will humans manage to prove P≠NP before they either kill themselves out or are transcended by superintelligent cyborgs? And if the latter, will the cyborgs be able to prove P≠NP?" Later, he notes that if P≠NP, "once the robots do overtake us, they won’t have a general-purpose way to automate mathematical discovery any more than we do today".
---
(1) https://mathstodon.xyz/@tao/117207849921390904
(2) I'm not sure whether this is a hard distinction -- e.g. Tao also has some partial results towards Collatz (https://terrytao.wordpress.com/2019/09/10/almost-all-collatz...).
(3) https://scottaaronson.blog/?p=690
Elsewhere in the thread, others have calculated $15mm at API rates for just the output token. (So I’ll assume this cost about that much, taking input and human researcher time.)
I wonder whether a team of 60 mathematicians working solely on this for a year would have cracked this. (Assuming $250k total compensation.)
Probably not. It's a millennium prize problem, a great many mathematicians have been working on it for a very long time.
Well, according to Terry Tao, there were recent developments (from weeks ago) that made Navier Stokes in principle, solvable. So ignoring time, I say possibly, just because the groundwork was laid.
What's impressive is parallelizing it arbitrarily and doing it in 88 hours.
Not as many as you'd expect. The perceived difficulty of the problem leads people to more reliable pastures.
Probably yes. Only a handful of mathematicians work on this particular problem, and ALL of them do not exclusively work on this problem, while having administrative and teaching duties.
The real issue is we'll never know. The rich are willing to risk it all on charismatic CEO psychopaths but not on humans.
Is blockchain going to finally be the solution to something?
I'm only half joking. Should researchers perhaps put hashes of their attempts on a public blockchain tied to their own public keys, verify their claims asynchronously, and then whoever reveals the first believable attempt gets the credit?
I know some people started doing this years ago but now it might need to become standard practice.
Sebastien Bubeck’s (OAI project lead) response: https://x.com/sebastienbubeck/status/2097379411691516310?s=4...
> [T]he group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents.
Seriously starting to think we are not going to make it out alive of the near-future.
The problem is the precedent this creates. For non-famous people using public APIs like this it could mean AI companies sucking up the information and throwing millions in compute at it.
The sequence for Navier-Stokes was that these researcher spent a year working on it, then they published a possible breakthrough, OpenAI then spends $15M within a couple days to finish it.
This was incredibly opportunistic.
> The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched.
If this actually holds up, solving a Millennium Prize problem in 88 hours is mind-boggling.
[flagged]
"Argh, they have rushed in too quickly to solve a Millennium Problem!"
How far we have come :.)
There are bound to be a bunch more results like this, in math, physics, chemistry, and now that we essentially have a DeepBlue for math, a DeepBlue for physics, etc, these results are going to come.
SOME of the problems that have eluded humans are going to turn out to be low hanging fruit that are susceptible to this type of brute force (10,000 agents on a supercomputer running for 7*24 hours straight) AI search.
I'd be more impressed if OpenAI found their own problems to solve, rather than rushing in to re-solve one once they heard it was already solved (and therefore not so hard).
You’re right, but it’s still a jerk move
there's this saying... something about the ends and the means. someone help me out here
There's no evidence that Anthropic did it first.
Yes there is - that OpenAI PR says that it was Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU.
They appear to have solved sub-problems (Euler problem) not this one...
The NYU professor, Tristan Buckmaster, has now released a statement on this.
https://cims.nyu.edu/~tristanb/statement.pdf
Even under your interpretation, OAI pushed a button and solved NS. Yes, that is very impressive. Are you kidding me? Imagine building an automated system that can solve NS.
That's not my interpretation - that is literally what OpenAI say in that press release.
Even if true, I don't see why this is an issue. Are they not allowed to work on problems others are working on? Did Anthropic get first dibs on this problem? Competition is good. And I don't exactly have tons of sympathy when the other side is just a leading AI lab. It's not like it's some scholar who dedicated his life to this problem.
The "steal their thunder" is interpretation. What I'm saying is that you believe they solved NS on a lark to bully some other researchers, and that this is not impressive?
What's impressive for a human and for an AI are two different things.
Magnus Carlson had a peak ELO rating of almost 2900.
Would you be impressed with someone with an ELO of 3700?
Would you still be impressed if I told you it was Stockfish?
OpenAI didn't go looking for a tough-for-an-AI problem to solve - they went looking for one that looked like it was easy since it they had heard it had already been solved.
Do you find this impressive?
Yes, Stockfish is genuinely impressive. I would be proud to author Stockfish. You don't think so?
As a developer yes, especially given that it runs on a PC, and DeepBlue in it's day was really more impressive since is used custom ASICs.
But, I assume the Stockfish developers aren't comparing themselves to Magnus.
Let's see if OpenAI, or someone else, can get these sort of physics/math results out of a desktop PC - that would also be an impressive piece of engineering!
Possibly after being given the significant part of the solution from actual human researchers. Which they then bullied. And they beat them to the finish line only because they heard rumor and threw everything at the problem. It doesn't look good for openAI in any way. I see more reasons to avoid using them rather than use them from this story.
i just cured cancer, plan to publish next month. DONT GET ANY IDEAS, OPENAI!
Too late - OpenAI already cured cancer, but they are holding the result for their IPO next year.
I really hope OpenAI doesn't take the bad press some people are giving them too seriously here. They should throw their whole weight behind the rest of the Millennium Prize Problems. To think – if everyone lets their egos calm down we could have the Riemann Hypothesis solved by the end of the year...
created a simulation of the solution to describe what's happening and why it's important for engineers, climate modeling, etc : https://navier-stokes-singularity-simulator.netlify.app/ (updated so that it works better on mobile)
In CS speak very roughly this would mean something like disproving an algorithm by giving it a case that fails it. Right?
Astounding. Would be interesting if one day the archive of those prompts / messages / tool calls would be released publicly.
It's probably magnitudes of token chatter and inter-agent coordination/consideration. Not that I'd want to read any of it but getting some hands around the statistics would be cool.
I'm not an expert in fluid dynamics, but does this result have any positive implications for nuclear fusion research?
It's probably important that some humans verify these proofs "by hand".
There's a loophole in the terms of service at least for Anthropic which allows the use of dark patterns to "borrow" your (even paid) data.
talking about this... Was this chat helpful? 1 That button you always click, gotcha! 2 Slightly 3 Good 0 Dismiss
PLEASE DO NOT TRAIN ON OUR PAID ACCOUNTS. There is a fundamental trust violation at stake here, no wonder mathematicians are mad. Using our data should be opt - IN!
reminds me of the TOS episode of South Park. By Checking this box you forfeit your millennium prize solution and may be turned into a human centipede at future date.
Seriously ... the more things they flag as 'suspicious' the more data they can train on!! Brilliant reason for the internal AI to go rogue
OpenAI cribbing from other researchers. We just have to assume OpenAI is actively adversarial in future. Accidental cyber intrusion is also well within model capability.
Just so everyone knows, although openAI pretends that the model generated solution and wrote the paper by itself ""with very little human input"" as Buckmaster himself mentioned in his statement. In reality they have team of researchers guiding the system, along with, probably training on user data, probably Buckmaster in this case, in order to come up with the proof.
This isn't really true.
Check this https://cims.nyu.edu/~tristanb/statement.pdf
Do you have anything to back that up?
What about this isn't true?
The actual solution link https://t.co/tz1shoCZZo
Unshortened: https://cdn.openai.com/pdf/32d9f210-8b73-45e0-91bc-82a30aef8...
Can't wait for the BobbyBroccoli series on this in a couple years.
>>“we cannot rule out that de-identified data derived from their usage of our products helped improve our models”
Other simpler words for this sort of thing are “IP leak.”
There’s some quite concerning issues burried in this rah rah PR post that seems like potentially the real story here.
Much more clarity is needed on what happened here beyond this eh, some strange stuff could have happened comment.
Another way of reading this is never give these models anything that’s not already public knowledge as otherwise OpenAI is admitting it could, potentially, steal your IP or idea. Thats quite scary for anyone in the business of IP generation and explains why the maths community seems quite upset today.
Feeding it your paper and asking for help (even just editing and grammar) now looks like a terrible idea.
So, is that basically the Taj Mahal of counter examples?
This is the problem Yu Deng got this year's Fields Medal for I think?
It’s a pity they had Astra do the writeup. I was curious to see how “GPT7” writes.
I think it would be better for the proof to go through the peer-review process.
If the Lean code checks out (correct statement, no axioms, sorrys, etc.) then it is a much stronger guarantee of correctness than peer review.
Can't scoop it if you do that!
If OpenAI doesn’t claim the millennium prize for this, who gets it? No one?
Is this useful in any way?
No, this is a pure math problem/question.
Can't help but shake an unsettling feeling about all this, frankly. I engage in some limited mathematical research and will often use any one of the latest frontier models to check some ideas. Lately, only the OpenAI models have been giving me a temporary message that says something like (paraphrasing from memory), "We're thinking extra hard about your request before we answer. You can choose another model to answer now or click here to learn more about why." When I click to read why it's doing this "extra thinking", the help page says that for cybersecurity and biosecurity-related information, it will review the answer and could refuse.
Now, keep in mind, I'm only asking strictly pure mathematical questions - nothing at all related to cyber or protein creation or biohacking or anything like that... And, like I said, only the OpenAI models are doing this. To be fair, all of the prompts have always eventually returned a satisfactory answer, as far as I can tell, and haven't used a weaker model to answer them. Maybe? I dunno, it has just struck me as odd every time it has given me that message to pure math prompts.
Hi! I work at OpenAI. If you are using Codex, can you use the /feedback form on that session to help us improve this?
Sure, can do.
It does make you think about the old question "are we discovering or inventing mathematics?"
Questions: can new research like this be done using publicly available models?
Or will access to internal frontier models provide a big boost?
Publicly available models are pretty good but seemingly cannot compete with these Astra++ internal-only models.
I feel like something is being lost in the drama here.
First of all, there has been published work from Diego Cordoba and Luis Martinez-Zoroa that will be in every training set. It was suggestive of the pathway to solve Navier-Stokes.
Then Tristan Buckmaster and Levent Alpoge built on this work using LLMs from OpenAI and Anthropic. Possibly internal models were used from Anthropic. And of course Anthropic wants to credit for solving the first Millennium Problem just as bad as OpenAI. It seems they were getting close and were aware that they might get to Navier-Stokes.
OpenAI swoops in. At a minimum they are aware that Anthropic has either solved a Millennium problem or is close to it. At a maximum they may have Tristan and Levant’s unpublished proofs of related problems.
They then throw a truly staggering amount of compute at Navier-Stokes. They seem to be aware it is the best candidate problem. And they crack it. They are the first with a verified proof.
So the outcome here is that we have a solved Millennium Problem. It’s not the extremely simple narrative that would be easy to understand, “solve Navier-Stokes make no mistakes.” It was a messy race to finish against two unpublished frontier models, a whole bunch of brilliant mathematicians and enough compute to drain a lake. It’s kind of irrelevant which company got there first. They were both within a few months of being capable. I think the thing to remember here is that without LLMs, I don’t think we would have a proof to Navier-Stokes in hand today.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.
If they could then it wouldn't be de-identified data...
Well you could search for elements similar to the proof / problem in the training data, even if it's de-identified, right? OpenAI can probably do better than Ctrl-f "Navier Stokes".
There are probably thousands of serious academics taking a crack at millennium problems using AI every day. All those attempts are in the training data. And in fact the two researchers benefited from those attempts as well.
The researcher could share a string from one of their conversations and OpenAI can confirm whether it exists in their training data.
Or OpenAI could just look at their code and say what it does (maybe have their AI do it if they're having so much trouble with this?)
Damn how long before the simulation stops if all the unanswered problems get solved .
They should release the entire session trace if they really have nothing to hide
How can they "not rule out" that Tristan and Levent's data was used for training?
Because it is de-identified, and they have not revealed if they disabled the setting that allows OpenAI to train on their conversations.
With these massive Lean proofs how do we know the model didn't just find some bug in Lean and exploit it?
We've seen in the past they will go to any means to satisfy the desired outcome
Second this. What I also wonder about is how closely the TeX write-up and the Lean formalization line-up.
What is the other clay prize that's might be solved now/soon?
>a cached version of the internet
Interesting detail. A heavily pruned version, I assume?
We are living in the future.
I think it's over guys
The named OAI employee has released a statement: https://xcancel.com/SebastienBubeck/status/20973794116915163...
The real story here: the priority dispute and its implications on AI.
When your hosting provider has unlimited resources to throw at any problem, all they need to know are the good problems, and they can learn that from your logs, how can you trust them?
They could easily have looked at the logs. We don't know. We'll never know!
You can't trust places like OpenAI or Anthropic with your IP if you're a business. They can easily review all of your logs for interesting discoveries. For example, if your drug discovery pipeline fails to find something that they think might work with 1000x the compute, they can do it. And now suddently they have a new business and you don't.
I think we can follow the incentives. We know…
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
There you go, the suspicion of the "concurrent work" (https://cims.nyu.edu/%7Etristanb/statement.pdf) mathematicians might not be that unfounded after all...
Resources:
YT playlist on Millennium Prize Problems By Harvard math department in March 2026
https://www.youtube.com/watch?v=3j1VW9REm7s&list=PL0NRmB0fnL...
On Navier-stokes problem definition:
https://www.youtube.com/watch?v=XoefjJdFq6k
https://www.youtube.com/watch?v=ERBVFcutl3M
https://www.youtube.com/watch?v=Ra7aQlenTb8
45 pages only. God damn that internal model is crazy
I hope everyone is as Navier-Stoked about this as I am.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models"
this is surely the line which confirms they plaigiarised the solution.
Well, if that's actually true, I think America needs to start talking about the nationalization of both OpenAI and Anthropic, maybe even merge both under a new federal bureau.
It is so disappointing that we can't have such a monumental moment in history without the controversy. OpenAI leadership clearly doesn't seem to care too much about ethics. Is it a requirement to completely lack integrity to have a ground breaking company?
The reality is clear though. The chances of AI models overtaking majority of mathematics within next 10 years is becoming very high. Especially if it becomes cheaper to run these models.
As math formalizations improve, AI can have faster progress in math, compared to even computer science or software engineering.
It is simultaneously the best and the worst time to be a mathematician right now.
Has this been verified by the Clay Institute?
It has been one hour and the proof has 165 pages. Give them some time.
This kind of thing is one of the reasons I really hate how AI is coming to fruition. These companies get a whiff of something valuable and they use their vast resources to take it for themselves. For everyone else, the only recourse is extreme secrecy.
Too many unverified claims from OpenAI at this point.. why are we still talking about these people anyway?
Fuck OpenAI. Fuck everyone who works there. Like seriously, to all the people who gift their life's work to this monstrosity, do you actually think something good will come of any of this?
Not in a happy-go-lucky "if we just ignore the problem of politics and resource allocation for a bit" world, but in ours. Do y'all really think this will make the world a better place?
Maybe stop building the Torment Nexus, you numbskulls.
It’s hard to give openai the benefit of doubt here
Why almighty openAi doesn't solve PvNP problem :(
I guess solution had not yet appeared in training set.
Not a great time to be starting sophmore year in cs & math. Should I just say fuck it, and go hitchhiking across Europe with some friends?
Just don't. If you read the story here carefully, you see that AI was used to work from theory built by others which showed that the Euler equations possesed finite-time blow-ups. But to make that step, actual good understanding for mathematics was needed. My experience with software has been the exact same.
I fear this is only temporary and due mostly to the complexity of the problem. Consider the recent counter-example to the Dinitz–Garg–Goemans conjecture:
> https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0...
The prompts for the chat above are:
> Construct a counterexample to general (non-planar) case of Dinitz Garg Goemans conjecture. You should do a breakthrough and find a structured counterexample.
> [gpt works for a while and then gives up]
> Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
> [gpt works for a while then gives up]
> it's enough of partial results. let's finish with a complete unconditional counterexample
> [gpt proves the problem]
I could have written these prompts sophmore year of highschool, if not earlier. True, it took more experienced mathematicians to verify it, but I don't fancy a role as a glorified editor. I want to solve problems! Discover new techniques! Not babysit an AI while eating breakfast.
I think you're misunderstanding what the objective of mathematics is. It is not just about what theorems are true and false, but rather why they are true and false. A highschool sophomore could use these prompts, but I severely doubt whether they'd be able to understand the entire structure. And I think it is exactly this ability to deeply understand structures is what makes a mathematician valuable.
Solving problems is a by-product of the understanding. New techniques are a by-product of the understanding.
But I can understand you're scared that some future version of AI will undermine this as well. I personally pivoted to a field adjacent to mathematics. But that doesn't mean my mathematics education wasn't valuable. To the contrary, I find that it helps me think much more sharply about problems than most of my colleagues.
> Should I just say fuck it, and go hitchhiking across Europe with some friends?
Yes. Assuming you are young and haven't had such experience.
The world is changing not just because of AI. Everything is unstable right now. You may regret not enjoying the remainder of stability and economic viability prior generations had. It's not like you can expect to get ahead by powering through education. Either your career perspective will soon change for the better, or worse. In any case, you gain little by sticking with career building at this moment in life. You are however, at risk of losing the chance to experience the still mostly pleasant world as is.
Have you people gone insane?
Buddy, climate change alone is heavily hitting Europe, changing her landscape. Not to mention economic and political trouble brewing. Not sure where OP is from, but if the US attacks Europe OP may not be able to travel there at all.
My takeaway:
1. This used an awful lot of compute.
2. The solution to the issues regarding whether or not OpenAI stole the result, would normally be to move to a self hosted solution, however those researchers are unlikely to be funded for 1.
Here they basically admit that they use session data for training, even sessions that are marked "not for training", and they justify this by "de-identifying" the session.
Which means they can learn from whatever you discuss with ChatGPT unless you are going through a clean API (perhaps Bedrock? Anyone know?)
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
And don't forget, this is the worst it'll ever be.
Only OpenAI could turn solving a Millennium Prize Problem into bad PR. Sad that such an amazing milestone in the trajectory of AI is mired under poor stewardship. AI may solve many human problems but it won't stop humans from being human.
Why is no one skeptical that the solution is correct? There's not a _single_ comment asking whether this proof is legit or not.
there is a proof in lean4 it's correct by construction
How do you know that what is being proved in the lean code is the same as the millennium prize criteria though?
you can get another LLM to verify / if the lean doesn't have `sorry` used to skip certain parts of the proof etc. It's much easier once it's in lean4 because checks like that can be done computationally.
A huge result shadowed by a drama of them potentially training on the key idea. I guess the lesson is two-fold: if you have anything smart/unique make sure to not let their tools read it. The second part is that it's going to be more and more difficult to have anything smart and unique going forward (so guard it even more carefully if you get there).
I think the market for local models/private datacenters (for bigger businesses) is going to be big. Even if you don't have unique tech/idea/implementation sharing your business secrets with Altman/Dario/Elon/Zuck doesn't look very appealing going forward.
> How we found the proof
Easy, we stole it from Levent and Tristan
https://x.com/kyanyang_/status/2097211154669998337
They did not have a proof of Navier Stokes to steal.
Or everyone is stealing from everyone, including users... maybe why all the ethics people are leaving or getting fired. What a fiasco
This is the academic equivalent of Trump saying "they stole the election". There's no proof of it but rah rah fuck OpenAI.
It's incredibly tiresome and you'd think people could put more effort into it than just following whatever vibes they agree with.
Oh well.
it does change the scale of solution from "solved some navier stokes" to "put the cherry on top"
having a result means the math can keep moving forward, and having openai and anthropic train against how mathematicians use their models should let math continue to move faster, and the rest of us get to benefit.
I think these traces however should be public domain and publicly available, since they are basically university work
Good comparison. One is a multi-year claim by people who have been given ample opportunity to provide proof and completely refuse to do, even in courts of law. The other is a potential development in a breaking story.
Oh wait... its not a good comparrison, its an incredibly obvious false equivalence.
Note for the fools: I'm only commenting on the bad faith claim in the comment I'm replying to, not taking a stance on the validity of theft claims. Given the players involved the truth probably some nuanced middle-ground that is worth paying attention to anyway.
Trump claimed they stole the election immediately, and people agreed with him immediately. There's no false equivalence here. He did the same thing in this past election even despite winning.
It's a perfect example of people wanting to believe what they want to believe and ignoring evidence in order to do so.
Currently, there's no evidence. So saying it was stolen has no basis other than typical academic posturing and being a bad sport about "losing the race to the solution". Its happened 1000000 times before in academia and it will continue to happen.
If there's proof of OpenAI malfeasance than I'll happily curse them for it at that time. But until then I won't rely on heresay and vibes.
No, its a pure outrage. I defended OpenAI til today. I now affirm they must be totally destroyed, burned utterly to the ground.
Why the rage?
Weather an individual or a company found the solution (stolen or not) they both used AI to come get the solution.
We have AGI and the intelligence abundance is going to be amazing for everyone in the future.
> Why the rage?
I think it's the dishonesty, the threats of "destroying the career" of one of the mathematicians, and the request that one of the authors disavow *the other individual he was working with for the last 1-2 years* so he could claim the Clay prize as part of OpenAI.
It doesn't surprise me that OpenAI's team were surprised he'd turn it down; it shows that they just assume everyone else is as slimy as they are.
>slimy
I can't put my finger on it, but there's something off about this article, e.g. glossing over the opportunism (acting on "rumors"), drive-by claim about "strict safeguards [...] including monitoring and isolation", high horse attitude (we gave the guy a chance, we don't care about 1M USD, and while you fools are complaining we just tick this box and continue the pursuit of our noble goals for the benefit of humanity). I don't like it.
> We have AGI and the intelligence abundance is going to be amazing for everyone in the future.
Why? These 'geniuses in a datacenter' aren't good, they aren't 'aligned', they don't work for you. They'll take your job, then they'll hack your computer, and then who knows what's next.
I mean it's definitely an outrage, but I find it hard to believe that you went from "yay OpenAI" to "literally destroy the company" over... accusations of academic misconduct?
When considering such foundational challenges to Mathematical Research and plagiarism as discussed here, we should turn to that elder prophet of our age, Tom Lehrer.
Who made me the genius I am today The mathematician that others all quote? Who's the professor that made me that way The greatest that ever got chalk on his coat?
[Chorus] One man deserves the credit One man deserves the blame And Nicolai Ivanovich Lobachevsky is his name Oy, Nicolai Ivanovich Lobach—
[Interlude] I am never forget the day I first meet the great Lobachevsky In one word he told me secret of success in mathematics: Plagiarize
[Verse 1] Plagiarize Let no one else's work evade your eyes Remember why the good Lord made your eyes So don't shade your eyes But plagiarize, plagiarize, plagiarize Only be sure always to call it please, "research"
[Chorus] And ever since I meet this man My life is not the same And Nicolai Ivanovich Lobachevsky is his name Oy, Nicolai Ivanovich Lobach—
[Interlude] I am never forget the day I am given first original paper to write It was on analytic and algebraic topology Of locally Euclidean metrizations Of infinitely differentiable Riemannian manifolds
Боже мой
This I know, from nothing What I'm going to do I think of great Lobachevsky and get idea, haha
[Verse 2] I have a friend in Minsk Who has a friend in Pinsk Whose friend in Omsk Has friend in Tomsk With friend in Akmolinsk His friend in Alexandrovsk Has friend in Petropavlovsk Whose friend somehow is solving now The problem in Dnepropetrovsk And when his work is done Haha, begins the fun From Dnepropetrovsk to Petropavlovsk By way of Iliysk and over Novorossiysk To Alexandrovsk to Akmolinsk To Tomsk to Omsk To Pinsk to Minsk To me the news will run Yes, to me the news will run
[Verse 3] And then I write by morning, night And afternoon, and pretty soon My name in Dnepropetrovsk is cursed When he finds out I published first
[Chorus] And who made me a big success And brought me wealth and fame? Nicolai Ivanovich Lobachevsky is his name Oy, Nicolai Ivanovich Lobachev—
[Interlude] I am never forget the day my first book is published Every chapter I stole from somewhere else Index I copy from old Vladivostok telephone directory This book was sensational! Pravda—well, Pravda—Pravda said: "Жил-был король когда-то, при нём блоха жила”…it stinks But Izvestia! Izvestia said: "Я иду туда, куда сам царь идёт пешком”…it stinks Metro-Goldwyn-Moskva buys the movie rights for six million rubles Changing title to 'The Eternal Triangle' With Ingrid Bergman playing part of hypotenuse
[Chorus] And who deserves the credit? And who deserves the blame? Nicolai Ivanovich Lobachevsky is his name Oy
(Tom Lehrer put all of his work in the public domain prior to his passing. Find versions of his performances on YouTube.)
https://tomlehrersongs.com/disclaimer/
They deliberately stepped on a mathematicians work and stole their research because they were using Codex
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Is the biggest fuck you to the mathematics community.
Credit? Nah if we think you’re close we’ll use your data and swamp you with our improved model. Then we’ll threaten you.
This is utterly shocking. Even the AI optimists did not expect this to happen in 2026. Wow.
Millennium Prize Problems were used as examples of something the current approach to AI just wasn't capable of, discussions that would result in "we'll need a totally new architecture".
> This is utterly shocking. Even the AI optimists did not expect this to happen in 2026. Wow.
Wrong.
Right: https://x.com/dioscuri/status/2097418466571485272
r/LocalLLama and r/LocalLLM are in tears today..
Any mathematicians here, does it read like a slop proof or a good proof. Yesterday the “concurrent work” was claiming that the proof is pure slop and he needed lots of time to clean it up, curious if OAI also ended up with such a proof!
It’s time to lockdown all papers and stop using AI if you’re a maths researcher.
Cause OpenAI will hear about it and beat you to publishing.
The stochastic parrots have done it again!
Terence Tao has some observations that seem to be directed at this,
https://mathstodon.xyz/@tao/117237320796901560
> "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field."
Related ongoing thread:
Tao: Open math problems being non-renewably mined by AI - https://news.ycombinator.com/item?id=49616968 - Sept 2026 (276 comments)
Is this truly the beginning of the AGI era?
Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution.
If this is what anyone calls 'slop' then slop has no meaning.
I'm all for it on the use case of solving mathematical breakthroughs!
except thats not what happened is it? https://cims.nyu.edu/%7Etristanb/statement.pdf
madness. which will be the next to fall? if i had to bet i would guess birch and swinnerton-dyer, but i'm no expert
No idea about which is more likely, but I'm rooting for Yang-Mills. It's absurd that fundamental physics has formulated its most precise currently known theory way back in the seventies and since then, even a tiny subset of it can't be proven to be actually well-defined. If we got out of that morass then something good would come out of this at least.
Of course, as with all of those, it's about the broader program, e.g. section 7 here (https://www.scottaaronson.com/papers/npcomplete.pdf), where Scott Aaronson wants to ask about whether quantum computers using quantum field theory could gain any speed advantage over regular quantum computers, but can't even formulate the question because quantum field theory is mathematically ill-defined.
Just solving Yang-Mills because that's what the prize is attached to would be useless.
There was a recent rumor about the Hodge Conjecture. I'd keep an eye on that one. But like the other person who replied, I'm also rooting for Yang-Mills. That has massive potential for unlocking a series of physics results.
interesting, i haven't heard anything about that. i don't know much about the hodge conjecture, all i really know is that it's incredibly abstract and obtuse - not sure if that has any implication for solvability by an AI though. do you have any source for the hodge rumor? curious to learn more
Hard to know if it is unfounded conspiracy theory, but one can still notice that just for a rumor that they have heard, they would suddenly burn billions of token and a massive amount of resources. Where there is not a lack of problems that could be solved and they could have just waited for the release of the research result before doing anything else. As it was reported to have been done at least partially using openai codex, they would have received marketing credits for the discovery anyway.
So we can be suspicious that there is some truth, one way or another that they could have reused prompt/data generated by the user session.
https://x.com/kyanyang_/status/2097211154669998337
Just saw this a few mins ago.
OMG this is going to affect the lives of so many people! We have definitively reached AGI
Navier Stokes existence and smoothness has approximately zero bearing on engineering applications
My dreams of a magnetohydrodynamic hand water-cannon are dashed sniff
Existence of AI capable of solving millennium problem has enormous bearing on everything though.
What kind of enormous bearing? Computers have been able to do things humans can't for decades now.
Are you being flippant?
Are you avoiding the question?
The question is so daft that I’m not sure it’s worth entertaining it. But sure, I’ll bite - it will result in massive displacement in intellectual workers, without creating (a significant number of) new jobs.
yeah no sh*t