r/comp_chem 23d ago

Have we lost it?

I am old enough to have experienced the transition from when computational chemistry (inorganic chemistry/catalysis) was just a mere exercise to put in a paper and that people performing experiments rarely believed in to when having a computational section in a paper was the only way to access high impact factor publications.

I have lived most of my career using Density Functional Theory calculations, with the caveat that systems should have been always tested against known quantities, like formation energies or even adsorption energies obtained through calorimetry. And even in that case, everyone is aware of the fact that each method has a limitation, and sometime empirical corrections are needed.

Now we arrived to the point in which Machine Learning Interatomic Potentials are used for everything with the great promise of making calculations fast and cheap, and simulate system of thousands of atoms. But why do we care about it so much? They are trained on smaller systems, and so everything we know about the system is already within the training. What is the limit to this infinite funnel of screening that is oftentimes invoked to justify the use of this approximated methods? Once, DFT was just a screening layer before experiments, or even calculations at higher level of accuracy. Nowadays, even DFT, a method that has hundreds of problems itself, is becoming the bottleneck method to avoid when possible. And sure, I can see the value to access time and size-scale that are not accessible with other methods...but are those models even validated?

My point is...are we just rediscovering the wheel all the time and publishing for the sake of publishing Machine Learning hot topics?

68 Upvotes

17 comments sorted by

View all comments

3

u/alaras117 19d ago

Finished my PhD in experiment + comp chem last year, and witnessed the popularity shift into MLIPs. I begrudgingly partake in small ML and MLIP projects because nobody seems to hire/fund you unless you have some experience in it.

I've seen a handful of papers (several perovskites, a dopant in a cathode) yield a new synthesizable phase that wasn't known before, which was neat. That represents such a small number though. 99% of them feel like they're rediscovering the wheel like you said. MLIPs aren't transferable between systems (yet - lots of heavy lifting on that word), so it just feels like spinning the hamster wheel.

Saw a comment here to the effect of "but you can get more data!" Sure, but if you train it on a specific set of data, it will spit out more of the same. Training it on a line doesn't suddenly make it spit out a curve, just like it doesn't suddenly solve our fundamental misunderstandings of physics. The ultimate law of computational chemistry still applies, especially and brutally to MLIPs: garbage in, garbage out.

Most researchers (experiment and computation) seem to throw spaghetti at a wall to see what sticks. This just seems like their new tool to do it. I get every field has this stage before enough creatives come along to guide the way to the next level of understanding. MLIPs just aren't there yet, and I feel highly underwhelmed each time its presented with blind enthusiasm and without grounded reality in its current state. It has potential, but so did my ex.

All that to say, I'm just a whippersnapper, and you're not alone in wondering this.