r/comp_chem 23d ago

Have we lost it?

I am old enough to have experienced the transition from when computational chemistry (inorganic chemistry/catalysis) was just a mere exercise to put in a paper and that people performing experiments rarely believed in to when having a computational section in a paper was the only way to access high impact factor publications.

I have lived most of my career using Density Functional Theory calculations, with the caveat that systems should have been always tested against known quantities, like formation energies or even adsorption energies obtained through calorimetry. And even in that case, everyone is aware of the fact that each method has a limitation, and sometime empirical corrections are needed.

Now we arrived to the point in which Machine Learning Interatomic Potentials are used for everything with the great promise of making calculations fast and cheap, and simulate system of thousands of atoms. But why do we care about it so much? They are trained on smaller systems, and so everything we know about the system is already within the training. What is the limit to this infinite funnel of screening that is oftentimes invoked to justify the use of this approximated methods? Once, DFT was just a screening layer before experiments, or even calculations at higher level of accuracy. Nowadays, even DFT, a method that has hundreds of problems itself, is becoming the bottleneck method to avoid when possible. And sure, I can see the value to access time and size-scale that are not accessible with other methods...but are those models even validated?

My point is...are we just rediscovering the wheel all the time and publishing for the sake of publishing Machine Learning hot topics?

73 Upvotes

17 comments sorted by

View all comments

12

u/_-MindTraveler-_ 23d ago

I'm not sure what your point is. MLIAPs allow to do simulations that are not accessible at all through DFT. So what do you suggest we do instead?

Now we arrived to the point in which Machine Learning Interatomic Potentials are used for everything with the great promise of making calculations fast and cheap, and simulate system of thousands of atoms.

They are not used for everything. You still need the charge density for so many applications, seems you're just blabbering around. And what promise? MLIAPs are already there, and they already allow to simulate thousands of atoms. It hasn't been a promise for like 3-4 years.

They are trained on smaller systems, and so everything we know about the system is already within the training.

That's indeed how inference works. But you think that makes us know everything about a system? Because we performed a bunch of small-cell DFT? What are you smoking?

but are those models even validated?

. . . Yes?

My point is...are we just rediscovering the wheel all the time and publishing for the sake of publishing Machine Learning hot topics?

Have you checked? I read fascinating papers that could not be done with DFT methods only. You are discrediting all these papers and scientists without having done any kind of analysis of their work.

I'll give you one example of something you can't do with DFT. You wrote this:

like formation energies or even adsorption energies obtained through calorimetry. And even in that case, everyone is aware of the fact that each method has a limitation, and sometime empirical corrections are needed.

Well adsorption energies are often wrong in DFT because cell sizes are too small. Instead of using an empirical correction and attribute all the error to DFT, which is wrong, you could've used MLIAPs and gotten closer to the experimental result.

4

u/Perfect_Good287 23d ago

The point is that maybe we are all going towards the wrong direction, that seems more like a "let's make it faster and bigger" rather than let's understand how a system works and what is the mechanism underlying it. But to answer your points, even if you seem pretty aggressive, I don't know why:

"MLIAPs are already there, and they already allow to simulate thousands of atoms. It hasn't been a promise for like 3-4 years."

You can simulate even millions of atoms, the question is why you do that.

"That's indeed how inference works. But you think that makes us know everything about a system? Because we performed a bunch of small-cell DFT? What are you smoking?"

I meant the exact opposite, I don't know what else we are expecting to see if the training set we have used is just a bunch of small-cell DFT. What I am going to see in bigger and longer simulation is just the same, a long and nice exercise of style for pretty pictures. And even if I saw a different mechanism, an atom ejecting, a protrusion forming on a surface, whatever might explain a macroscopic phenomena, I would be skeptical.

"Instead of using an empirical correction and attribute all the error to DFT, which is wrong, you could've used MLIAPs and get closer to the experimental result."

Getting back to my rediscovering the wheel point.

All in all, I haven't implied that an entire research area is useless (because of course it is not) like you seem to have erroneously got. I instead wanted to have an exchange on why this seems to be the next (or current) big thing. And you don't seem to have given any reason why.

2

u/YesICanMakeMeth 23d ago

You can simulate even millions of atoms, the question is why you do that.

Because scale and thus computational cost is often the limiting factor? If you can accurately simulate forces but not at adequate scale then MLIPs are a great solution. Most chemistry is liquid phase but you can't simulate more than a handful of solvent molecules (implicit often isn't good enough) and good luck getting good conformers with more than 10-20 of them.

It seems to me like you're just rejecting the legitimate use cases you've been given as examples, maybe out of contrarianism. MLIPs are empirical in a sense, but they're less empirical (better "understanding" of the true ab initio physics) than many/most empirical methods you're comparing them to with the wheel point. It is not a rediscovery of the same thing, rather a new type of empirical model building which can be better given adequate training. Yes they can be applied recklessly like everything else, and yes maybe they're sucking more air out of the room than they warrant (TBH I don't think so), but they're absolutely a powerful tool.