I think, of course, skepticism around this "LLM discovers X" thing is warranted, and there have been plenty of more recent examples around questionable LLM "discoveries". Just stating this because the LK99 thing I believe was notable as a (supposed) room-temp _super_conductor while this is about a _semi_conductor.
I start my day with plenty of optimism, then I go back and forth in the CLI and find out most of whats posted online is fake, and then towards the end of the day 2h past my bed time I end up ed zitron maxxing, it is the way it is ig
"Debacle"? That was the most fun I've had on the Internet in years. When's the last time so many people engaged in so many arguments about materials science and electromagnetism? Sometime in the 1800s?
Yeah I remember going to my physics professor super excited about LK-99 to ask him if he heard about it, and him just telling me "yes but stuff like that happens twice per year, they will find something is off", and in fact it's what happened...
You probably meant "I'm taking this with a tiny pinch of salt". The amount of salt is directly proportional to how much of the claim you are willing to accept.
Edit: I stand corrected. According to Gemini:
Me: Does using more salt mean accepting more of that claim?
Gemini: No, it actually means the exact opposite.
If you say you need to take a claim with a huge pile of salt (or a shovel of salt), it means you believe the claim is highly unbelievable and you need an immense amount of skepticism to accept it.
How the Metaphor Scales
• A single grain of salt: "I am slightly skeptical, but it could be true."
• A pinch of salt: "I have a healthy amount of doubt about this."
• A grain of sand / A truckload of salt: "This sounds completely made up, and I barely believe a single word of it."
The salt represents your skepticism, not your belief. Therefore, the more unbelievable the claim, the more "salt" you need to swallow it.
I don't think this is right. https://en.wikipedia.org/wiki/A_grain_of_salt The "grain" isn't a single grain, it's an old English measure which is around 65mg, i.e. roughly how much there is in a pinch. I've also only ever heard people use larger amounts to mean more scepticism.
In a way, you can think of pretty much anything we express with language, especially things that are already modeled in scientific language, or logical language, or in equations or code; to be representable in a parametric/searchable space
Thus, you can build ai/ml models+agents to explore those spaces, at a speed and scope much larger than what any human can do
I can imagine findings like these are going to keep increasing in frequency to a point in which the bar for novelty goes a lot higher
Anecdata: over the weekend, on a whim, I decided to download a real fly’s brain’s weights [0], run it on a simulated task like finding food, then train a logistic classifier using the fly’s decisions as the expert, then use the trained classifier as a decision model to simulate the fly on a 3d environment, running in real time on a website
It took me (using Claude code and some codex), about 3 hours to put it together
And even though it was a cool demo, it seemed so easy, that it also felt like it wasn’t worth sharing
It is not at all obvious that merely because we have words for concepts, that a model should be able to do all these miraculous mathematical and scientific things.
You are correct. My comment is not so much about that this is something elementary. But rather an observation that, given the current state of technology, it seems like we are being able to model increasingly more things, in increasingly more efficient and automated ways, to the point that there seems to be a pattern to it
> We’re all used to two types of magnet. The common one, the fridge magnet, is ferromagnetic — its atomic magnets all point the same way (up or down), adding their magnetic effects. The less well known one, the antiferromagnet (AF), has neighbouring atomic magnets that point opposite ways and exactly cancel out magnetically.
This is a very bizarre introduction. People encounter diamagnets (e.g., copper) and paramagnets (e.g., aluminum) way more than they encounter antiferromagnets. I don't know why you'd ever cast magnetism as a false binary between ferromagnets and antiferromagnets, without even acknowledging any other types of magnetic order.
Yeah, and it's not even an accurate explanation either.
> The common one, the fridge magnet, is ferromagnetic — its atomic magnets all point the same way (up or down), adding their magnetic effects.
Ferromagnets typically have domains with magnetic moments that point in different directions. It can still have a net magnetic moment without every 'atomic magnet' pointing the same way.
I am not sure how this process looks like. When they "discover" these, what are they actually doing?
The agents ran quantum-mechanical simulations of each crystal with the standard method for this, density functional theory, at two levels of approximation: a faster one (PBE+U) and a slower, usually more accurate one (HSE06). The band gaps and spin windows below come from the more accurate one.
So the agent runs a classic simulation or I am missing something.
A lot of the public successes with agents is really LLM-driven local search against an objective function that is evaluated in more traditional ways. This one seems to fit the pattern.
From the little I understand about this topic, it looks similar to approaches used in the recent Navier-Stokes breakthrough. These physical systems are governed by partial differential equations (PDEs) which can be solved numerically using standard algorithms. So when we say "simulation" in this context we really just mean "numerical solution".
In the case of quantum mechanics, it's the Schrödinger equation, which is no different than any other PDE. Agents are getting very good at searching through the space of possible simulation parameters and initial conditions to find solutions with certain properties. Coarser simulations are less accurate but faster to run, so the search uses simulations at different scales to find promising directions, and then refines those to verify that the simulation converges on the expected result.
One of the potential applications of quantum computing is that it might speed these simulations up exponentially. But scientists can and do regularly simulate quantum systems on classical computers.
I'm under the impression that this kind of modeling is one of the applications that quantum computers are likely to be good at.
I'd imagine there's a lot of documented research which has attempted to find such things using classical computers.
Seems like there would be a lot of well structured context for somebody to use while directing agents to repeat that research, now with updated models once quantum computing is ready for that kind of task.
not a classic simluation- a quantum simulation. This means they put a lot more work into representing the wave function of the simulation and modelling quantum effects.
They ran Quantum Espresso which is ok, but by no means the 'state of the art' for DFT. And in case, any DFT computation has to be taken with a few pounds of grains of salt before getting too excited about it.
No offense to the person writing this (assuming they did at all), but I'm not sure they really understand what they're doing..
Frankly, there is no point in trying to "understand" what an LLM does. Their thought process is effectively undecipherable by humans (it's essentially information arising from information) so even such a "simple explanation" is almost certainly wrong. The agents might appear to have "used this method", but the actual method of computation is far beyond our grasp.
Why are people being so belligerent about this? I thought it's fairly obvious at this point that LLM reasoning is far beyond anyones understanding. Or does anyone have a refutation?
You're confusing the weights of a model and internal chain-of-thought with the output of the model. Yes, we don't know a lot about how the internal mechanisms work. But with the correct prompt, agents will produce a worklog that documents exactly what solutions were tried and how the result was obtained.
This is a strange attitude. When an agent is optimizing a piece of code, comes up with 2 variations, and runs benchmarks on them to figure out which one is faster, then selects one of them based on tradeoffs between performance and other things it reasons about, do you ignore its explanation and all experiment runs?
What are you on about? I have had Fable come up with new shit for me several times (I do research for a living, so actual new shit nobody knew before), and each time it was perfectly understandable.
Of course I don’t know how it got its ideas for what to try. But heck, I don’t even understand how I get my ideas half the time. But the process, like what code it wrote, simulations it ran etc can be understood by (some) humans just fine!
Yes I saw 3Blue1Brown say the same thing in his tutorial on how neural nets worked where he built a simple model to recognize a particular letter. Good reminder.
I've been dabbling with some of my own (tiny) models recently and it's actually shocking at what they can "learn" despite having _zero_ mention of it in it's training data.
One of the materials is most likely impossible to synthesize. The other already exists, so that may actually be capable of being tested. It's only been synthesized once, 27 years ago though.
Who is vals.ai and why they keep submitting eye-catching claims. A few weeks ago they said fable 5.1 solved some obscure cipher and now opus 5.5 found room temperature semiconductor candidates. Meanwhile they seem to be in the business of making benchmarks.
Okay? Aren't the semiconductors we use today room temperature? I certainly don't use helium to cool my phone.
I don't see any claims that this is better than the current silicon and gallium arsenide semiconductors that we use. And the use of "room temperature" seems a deliberate attempt to misconstrue this with superconductors
A lot of these ‘an agent invented’ or ‘an agent solved’ are actually the agent wading through a lot of info and finding something a human did that no one noticed or saw the relevance of at the time.
If ai becomes so prolific that we humans all stop doing those things then will they still work?
Which is somewhat ironic since neural networks were "discovered" back in the 1940s... then forgotten... then wait, they were discovered again! ... then forgotten, again... and now here we are.
Well its not just any old human doing these things in a general sense. Its typically academics or highly paid researchers who love doing work like this. So, I don't think it will just one day stop
This is frankly one of the best uses of LLMs (along with proposing and evaluating drug therapies), and I think it's (at least partially) because these are things that will only work in the hands of people who are already experts and motivated in the field. The proposed thing is validate (or not validated), and then everyone moves on (either using the cool new thing, or knowing that it doesn't work). I'd also throw robotics in here.
The fact that the major "uses" of LLMs have been contributing to the acceleration of the dead internet theory, and building millions of versions of the same apps that no one is going to maintain, is extremely sad.
Sounds interesting. Excited to see physical versions of this cooked up. Also, very excited for a world a few years from now where we can talk about accomplishments like this from the frame of the driver of the AI, rather than hype that AI helped.
I would image it's the data the researchers fed the agents and in which a discovery was likely. Especially since it's "candidates", so it's not like a proper discovery.
This doesn't sound like something that needed an LLM? It just brute forced a lot of combinations of elements until finding one with the right material properties in a simulation, or what am I missing? And the one that actually worked wasn't even a new invention? I suppose it's quite likely that whoever discovered the second one in 1999 also discovered the first one and didn't publish it, since it didn't work
This should probably read: "Researchers discover two room-temperature magnetic semiconductor candidates. They used Opus 5.5 agents to perform some checks."
This gave me the idea to actually create a full (QED accurate) atomic simulation software. Essentially would allow you to play around with things like this. At a glance my workstation _probably_ has enough compute to handle it. At least to fully simulate at least a few dozen atoms and compounds.
Ugh. Unless this has been actually experimentally verified to be a room-temperature and room-pressure superconductor, it's about as ground breaking as "Yet another promising nuclear fusion candidate theoretically described."
Reading the title I saw the words "room-temperature" and my mind auto-completed it to superconductor, and based on other comments I don't think i'm alone in that.
I agree that it is about as ground breaking as "Yet another promising nuclear fusion candidate theoretically described."
I'm not sure why you would consider new and promising avenues for research to not be ground breaking. If it's an idea worth trying, it's an idea worth trying. If it doesn't survive testing, then it was still worth trying.
Current frontier LLMs empower effectively anyone with limitless knowledge. Historically, if I wanted to hire an engineer to, say, create something like this I would have needed a multi-million dollar budget. Now, anyone with $200 (or less) can achieve it.
You are vastly overestimating what has been achieved here.
This is something a couple of materials science grad students can do in limited time for poor compensation as well. The expensive budget is for the part that comes next.
Of course, "LLM solves quantum gravity and proves existence of God",
"nah brah, that's easy brah any kid could have done this brah".
This is what you sound like. I also like how the goalposts keep moving on a daily basis, a year ago it was that LLMs can't even write a Hello World program without making an error, but now things like this are "so easy a minimum wage intern could do it."
Isn't there quite a bit of space between "so easy a minimum wage intern could do it" and your original claim that it would have cost millions of dollars to produce these results?
77 comments:
After the LK-99 debacle, I'm taking this with a truck load of salt.
I think, of course, skepticism around this "LLM discovers X" thing is warranted, and there have been plenty of more recent examples around questionable LLM "discoveries". Just stating this because the LK99 thing I believe was notable as a (supposed) room-temp _super_conductor while this is about a _semi_conductor.
I start my day with plenty of optimism, then I go back and forth in the CLI and find out most of whats posted online is fake, and then towards the end of the day 2h past my bed time I end up ed zitron maxxing, it is the way it is ig
> After the LK-99 debacle
"Debacle"? That was the most fun I've had on the Internet in years. When's the last time so many people engaged in so many arguments about materials science and electromagnetism? Sometime in the 1800s?
Maybe he meant 'debacle' in an endearing sense, not a derogatory one. I personally agree with you and loved this debacle.
Yeah I remember going to my physics professor super excited about LK-99 to ask him if he heard about it, and him just telling me "yes but stuff like that happens twice per year, they will find something is off", and in fact it's what happened...
You probably meant "I'm taking this with a tiny pinch of salt". The amount of salt is directly proportional to how much of the claim you are willing to accept.
Edit: I stand corrected. According to Gemini:
Me: Does using more salt mean accepting more of that claim?
Gemini: No, it actually means the exact opposite. If you say you need to take a claim with a huge pile of salt (or a shovel of salt), it means you believe the claim is highly unbelievable and you need an immense amount of skepticism to accept it. How the Metaphor Scales
• A single grain of salt: "I am slightly skeptical, but it could be true."
• A pinch of salt: "I have a healthy amount of doubt about this."
• A grain of sand / A truckload of salt: "This sounds completely made up, and I barely believe a single word of it."
The salt represents your skepticism, not your belief. Therefore, the more unbelievable the claim, the more "salt" you need to swallow it.
Hmm? I always thought it was how much you had to flavor the statement to swallow it.
I don't think this is right. https://en.wikipedia.org/wiki/A_grain_of_salt The "grain" isn't a single grain, it's an old English measure which is around 65mg, i.e. roughly how much there is in a pinch. I've also only ever heard people use larger amounts to mean more scepticism.
[delayed]
Inversely proportional
Citation needed.
In a way, you can think of pretty much anything we express with language, especially things that are already modeled in scientific language, or logical language, or in equations or code; to be representable in a parametric/searchable space
Thus, you can build ai/ml models+agents to explore those spaces, at a speed and scope much larger than what any human can do
I can imagine findings like these are going to keep increasing in frequency to a point in which the bar for novelty goes a lot higher
Anecdata: over the weekend, on a whim, I decided to download a real fly’s brain’s weights [0], run it on a simulated task like finding food, then train a logistic classifier using the fly’s decisions as the expert, then use the trained classifier as a decision model to simulate the fly on a 3d environment, running in real time on a website
It took me (using Claude code and some codex), about 3 hours to put it together
And even though it was a cool demo, it seemed so easy, that it also felt like it wasn’t worth sharing
0: ChessFly (not mine), uses the FlyWire connectome (the fly’s brain’s weights) to play chess https://huggingface.co/spaces/mlabonne/chessfly
Many people wouldn't find that easy, even with AI
I think ai certainly raises the bar for those with taste
Oh please do share!
It is not at all obvious that merely because we have words for concepts, that a model should be able to do all these miraculous mathematical and scientific things.
You are correct. My comment is not so much about that this is something elementary. But rather an observation that, given the current state of technology, it seems like we are being able to model increasingly more things, in increasingly more efficient and automated ways, to the point that there seems to be a pattern to it
Right, it also has to model a substantial fraction of reality (or at least a true simulation of it) to accomplish these things.
Yeah...?
That's the pitch of LLMs lol
> We’re all used to two types of magnet. The common one, the fridge magnet, is ferromagnetic — its atomic magnets all point the same way (up or down), adding their magnetic effects. The less well known one, the antiferromagnet (AF), has neighbouring atomic magnets that point opposite ways and exactly cancel out magnetically.
This is a very bizarre introduction. People encounter diamagnets (e.g., copper) and paramagnets (e.g., aluminum) way more than they encounter antiferromagnets. I don't know why you'd ever cast magnetism as a false binary between ferromagnets and antiferromagnets, without even acknowledging any other types of magnetic order.
(I did a PhD in magnetic materials)
This is what happens when Claude writes it for you (and you don't review it)
I also like how they explained ferromagnetism as being arranged atomic magnets. Magnets all the way down.
Yeah, and it's not even an accurate explanation either.
> The common one, the fridge magnet, is ferromagnetic — its atomic magnets all point the same way (up or down), adding their magnetic effects.
Ferromagnets typically have domains with magnetic moments that point in different directions. It can still have a net magnetic moment without every 'atomic magnet' pointing the same way.
https://en.wikipedia.org/wiki/Magnetic_domain
I am not sure how this process looks like. When they "discover" these, what are they actually doing?
So the agent runs a classic simulation or I am missing something.A lot of the public successes with agents is really LLM-driven local search against an objective function that is evaluated in more traditional ways. This one seems to fit the pattern.
From the little I understand about this topic, it looks similar to approaches used in the recent Navier-Stokes breakthrough. These physical systems are governed by partial differential equations (PDEs) which can be solved numerically using standard algorithms. So when we say "simulation" in this context we really just mean "numerical solution".
In the case of quantum mechanics, it's the Schrödinger equation, which is no different than any other PDE. Agents are getting very good at searching through the space of possible simulation parameters and initial conditions to find solutions with certain properties. Coarser simulations are less accurate but faster to run, so the search uses simulations at different scales to find promising directions, and then refines those to verify that the simulation converges on the expected result.
One of the potential applications of quantum computing is that it might speed these simulations up exponentially. But scientists can and do regularly simulate quantum systems on classical computers.
I'm under the impression that this kind of modeling is one of the applications that quantum computers are likely to be good at.
I'd imagine there's a lot of documented research which has attempted to find such things using classical computers.
Seems like there would be a lot of well structured context for somebody to use while directing agents to repeat that research, now with updated models once quantum computing is ready for that kind of task.
not a classic simluation- a quantum simulation. This means they put a lot more work into representing the wave function of the simulation and modelling quantum effects.
They used quantum espresso.. undergrads usually run this in certain classes: https://www.quantum-espresso.org They didn't do any work there.
They ran Quantum Espresso which is ok, but by no means the 'state of the art' for DFT. And in case, any DFT computation has to be taken with a few pounds of grains of salt before getting too excited about it.
No offense to the person writing this (assuming they did at all), but I'm not sure they really understand what they're doing..
Frankly, there is no point in trying to "understand" what an LLM does. Their thought process is effectively undecipherable by humans (it's essentially information arising from information) so even such a "simple explanation" is almost certainly wrong. The agents might appear to have "used this method", but the actual method of computation is far beyond our grasp.
Why are people being so belligerent about this? I thought it's fairly obvious at this point that LLM reasoning is far beyond anyones understanding. Or does anyone have a refutation?
You're confusing the weights of a model and internal chain-of-thought with the output of the model. Yes, we don't know a lot about how the internal mechanisms work. But with the correct prompt, agents will produce a worklog that documents exactly what solutions were tried and how the result was obtained.
This is a strange attitude. When an agent is optimizing a piece of code, comes up with 2 variations, and runs benchmarks on them to figure out which one is faster, then selects one of them based on tradeoffs between performance and other things it reasons about, do you ignore its explanation and all experiment runs?
>Their thought process is effectively undecipherable by humans (it's essentially information arising from information
Are you trying to say that human brains are incapable of inference?
What are you on about? I have had Fable come up with new shit for me several times (I do research for a living, so actual new shit nobody knew before), and each time it was perfectly understandable.
Of course I don’t know how it got its ideas for what to try. But heck, I don’t even understand how I get my ideas half the time. But the process, like what code it wrote, simulations it ran etc can be understood by (some) humans just fine!
Yes I saw 3Blue1Brown say the same thing in his tutorial on how neural nets worked where he built a simple model to recognize a particular letter. Good reminder.
I've been dabbling with some of my own (tiny) models recently and it's actually shocking at what they can "learn" despite having _zero_ mention of it in it's training data.
Interesting; but until actually made and tested, not worth getting excited over.
One of the materials is most likely impossible to synthesize. The other already exists, so that may actually be capable of being tested. It's only been synthesized once, 27 years ago though.
> One of the materials is most likely impossible to synthesize
Is this a "actual impossible because it's inherently contradictory", or "we just don't know how to do it yet but give us a year"?
We don't know a way to precisely place atoms in a checkerboard pattern like that, without getting it so hot that the arrangement is destroyed.
It maybe could be possible but beyond the reach of current material science.
Who is vals.ai and why they keep submitting eye-catching claims. A few weeks ago they said fable 5.1 solved some obscure cipher and now opus 5.5 found room temperature semiconductor candidates. Meanwhile they seem to be in the business of making benchmarks.
Are they a promoter / influencer for Anthropic?
Okay? Aren't the semiconductors we use today room temperature? I certainly don't use helium to cool my phone.
I don't see any claims that this is better than the current silicon and gallium arsenide semiconductors that we use. And the use of "room temperature" seems a deliberate attempt to misconstrue this with superconductors
Sounds like a good reason to hire a lab to make some, and then make a big deal about it if the results pan out.
I can think of worse uses of VC AI funding.
A lot of these ‘an agent invented’ or ‘an agent solved’ are actually the agent wading through a lot of info and finding something a human did that no one noticed or saw the relevance of at the time.
If ai becomes so prolific that we humans all stop doing those things then will they still work?
Which is somewhat ironic since neural networks were "discovered" back in the 1940s... then forgotten... then wait, they were discovered again! ... then forgotten, again... and now here we are.
Well its not just any old human doing these things in a general sense. Its typically academics or highly paid researchers who love doing work like this. So, I don't think it will just one day stop
Yes, as long as we are advancing to behavior and world models, so that agents can interact with the world themselves. Which we are.
I wouldn't describe them both as being newly-discovered. The second one, KV[Cr(CN)₆], had already been discovered.
While we should be skeptical until made in a lab or verified by others, this is a much better use of LLMs than solving math theorems/conjectures
This is frankly one of the best uses of LLMs (along with proposing and evaluating drug therapies), and I think it's (at least partially) because these are things that will only work in the hands of people who are already experts and motivated in the field. The proposed thing is validate (or not validated), and then everyone moves on (either using the cool new thing, or knowing that it doesn't work). I'd also throw robotics in here.
The fact that the major "uses" of LLMs have been contributing to the acceleration of the dead internet theory, and building millions of versions of the same apps that no one is going to maintain, is extremely sad.
The people who wrote this seem to be lacking in expertise, and its just a model benchmarking company..Whos every article is just hyperbole about llms.
Not sure why we're calling it a discovery, when they've literally been made before, by a human.
Sounds interesting. Excited to see physical versions of this cooked up. Also, very excited for a world a few years from now where we can talk about accomplishments like this from the frame of the driver of the AI, rather than hype that AI helped.
Discovered in whose data?
All research is built off the existing body of all research data done by other people
I would image it's the data the researchers fed the agents and in which a discovery was likely. Especially since it's "candidates", so it's not like a proper discovery.
Lmao, anything goes
This doesn't sound like something that needed an LLM? It just brute forced a lot of combinations of elements until finding one with the right material properties in a simulation, or what am I missing? And the one that actually worked wasn't even a new invention? I suppose it's quite likely that whoever discovered the second one in 1999 also discovered the first one and didn't publish it, since it didn't work
This should probably read: "Researchers discover two room-temperature magnetic semiconductor candidates. They used Opus 5.5 agents to perform some checks."
I could have gotten this in one prompt lmao
This gave me the idea to actually create a full (QED accurate) atomic simulation software. Essentially would allow you to play around with things like this. At a glance my workstation _probably_ has enough compute to handle it. At least to fully simulate at least a few dozen atoms and compounds.
You should consider patenting this very much novel idea, my friend!
No worries, your workstation is more than enough to run accurate quantum simulations!
:)
Anyone remember LK99 lol
This is semiconductors, not superconductors.
That was a room temperature superconductor, a bit different of a task.
Ugh. Unless this has been actually experimentally verified to be a room-temperature and room-pressure superconductor, it's about as ground breaking as "Yet another promising nuclear fusion candidate theoretically described."
Well, given that this is a semiconductor, and not a superconductor, I don't see how that is relevant?
Reading the title I saw the words "room-temperature" and my mind auto-completed it to superconductor, and based on other comments I don't think i'm alone in that.
I agree that it is about as ground breaking as "Yet another promising nuclear fusion candidate theoretically described."
I'm not sure why you would consider new and promising avenues for research to not be ground breaking. If it's an idea worth trying, it's an idea worth trying. If it doesn't survive testing, then it was still worth trying.
Current frontier LLMs empower effectively anyone with limitless knowledge. Historically, if I wanted to hire an engineer to, say, create something like this I would have needed a multi-million dollar budget. Now, anyone with $200 (or less) can achieve it.
You are vastly overestimating what has been achieved here.
This is something a couple of materials science grad students can do in limited time for poor compensation as well. The expensive budget is for the part that comes next.
Of course, "LLM solves quantum gravity and proves existence of God", "nah brah, that's easy brah any kid could have done this brah".
This is what you sound like. I also like how the goalposts keep moving on a daily basis, a year ago it was that LLMs can't even write a Hello World program without making an error, but now things like this are "so easy a minimum wage intern could do it."
Isn't there quite a bit of space between "so easy a minimum wage intern could do it" and your original claim that it would have cost millions of dollars to produce these results?
Why hyperbolize when I am commenting on something it has actually done and the vastly exaggerated claims related to this?
It hasn't done that though.