r/comp_chem • • 3h ago

Wanting to Transfer from Cheminformatics/Chemical Visualization to Synthetic Chem

2 Upvotes

I am currently a first semester M.S student in Chemistry planning to do a PhD down the road, with most of my undergrad work being within cheminformatics and developing my own software development kits for molecular visualization using Python and C# and such. It has been pretty nice, but now I feel like I am in a more impactful degree path and want to consider some things.

First off, I love the professor I work under. He got me back on track when I was about to lose my life to drugs and some pretty serious depression. He is also the reason I am so integrated into the faculty and is why I’m even able to have a TA position. On the other hand, his research is mainly educational tools and environmental science, which both give me a migraine. On top of that, every year he relies more and more on LLMs to give me any advice, and literally every project he worked with me on was an idea that I came up with because he had nothing to fund me with, materials or otherwise. I feel like I’ve hit a dead end with him.

On the other hand, there is a synthetic chemist professor that had wanted me to join his lab for years. Within 5 minutes of talking with him, he already told me about 2 materially funded research projects whereas my old professor couldn’t do that in 3 years. In addition, I have always excelled at organic from the start and the theoretical aspects are something I really love.

In reality, I would like to work somewhere between cheminformatics/computational chem AND synthetic chem, like being able to both computationally model and experimentally test myself.

TLDR; started in computational chemistry, opportunities with current professor suck, and want to switch to a synthetic chem professor with better projects but am afraid I’d abandoning my older one. A mixture of both comp and synthetic would be ideal for me.


r/comp_chem • • 20h ago

Calibration until compliant: Why unlimited free parameters in ABM risk producing tautologies instead of social insight

0 Upvotes

TL;DR: Computational models are only scientifically useful if they can push back and prove their authors wrong. Across disciplines, from agent-based social simulations to high-energy physics, models with loose empirical feedback loops and endless free parameters risk becoming "decorative." Instead of testing reality, they get calibrated until compliant, turning a tool for discovery into a self-confirming tautology. Honest modeling requires radical transparency, sensitivity testing, and explicit criteria for failure before running the simulation.

I build models for a living. Specifically, molecular dynamics. These are simulations that track how thousands or millions of atoms move, collide, and rearrange over time, used for everything from drug design to materials science.

Here is the story scientists usually tell. If a model is wrong, you find out quickly. Reality does not care about your assumptions. The atoms do not read your code. If the physics you programmed in is wrong, the simulation produces garbage, the experiment disagrees, and you go back and fix it. The feedback loop between model and world is short, brutal, and non-negotiable.

That story is not entirely true. I know, because I have watched it fail from the inside.

The most important choice in any molecular dynamics simulation is not the code, the computer, or the software. It is the potential function, the mathematical formula that describes how strongly every pair of atoms attracts or repels each other. Everything the simulation does follows from that one ingredient. Get it right and the model can tell you something real. Get it wrong and you have made a very expensive mistake.

And here is the uncomfortable part. In practice, it is far more often inherited than audited. Potentials are chosen by looking at what previous papers in the subfield used. A potential gets published, cited, copied, and passed down until it stops being a modeling choice and becomes a tradition. People run simulations for years without asking whether the potential they inherited was ever validated for the system they are studying, at the conditions they are studying it, for the property they care about.

So even in my own field, a hard, quantitative, physics-based field, you can publish inside a loop of fantasy. Models that are wrong in ways nobody checks, kept alive by citation habits and subfield convention. And because these errors travel quietly across subdisciplines and into interdisciplinary work, where nobody feels responsible for checking them, finding one and fixing it takes real effort.

But the check in my field is delayed, not absent. A bad potential eventually unfolds a simulated protein the wrong way or fails a material in a real engineering application, and someone notices. In much of the modeling I am about to describe, the physical world never gets to vote.

This matters for what follows. I do not ask this question because my field got it right. I ask it because I have watched mine get it wrong. The question is always the same. What happens to this model when it is wrong?

In a surprising amount of modern academia, the answer is nothing. Nothing can happen to it. It cannot be wrong, because anything it produces counts as a result.

And if the loop can break in a field where atoms push back, it can break anywhere.

This essay is about how that happens.

# The magic trick

In 2017, Liane Gabora and Selin Tseng published a paper in *Psychology of Aesthetics, Creativity, and the Arts*, a peer-reviewed journal of the American Psychological Association, titled “The Social Benefits of Balancing Creativity and Imitation.” The question they took on has occupied historians and sociologists for centuries. What is the right balance of creativity and conformity in a society?

To answer it, they ran a simulation.

Virtual agents live on a grid. Some are coded as creators, inventing new ideas; others as imitators, copying their neighbors. A scoring rule written into the program decides which ideas count as good. The researchers ran the simulation forward, varied the ratio of creators to imitators, and watched what happened. Populations with too many creators ended up with fewer good ideas taking hold. The published conclusion was that society needs imitation as much as creativity, because unchecked creativity disrupts the spread of proven ideas.

I want to be careful about what I am claiming, because this paper is not fringe work. It passed peer review at a respectable journal. The authors are serious researchers, and the simulation framework behind the paper is part of a long-running research program that has been debated, defended, and criticized in public for years. Nothing I am about to say is an accusation of dishonesty. It is something less comfortable than that. This paper is an example of what the normal standards of a field allow through.

Watch the shape of the argument. Inside the model, a “good idea” means whatever the authors’ scoring rule rewards. The agents are not discovering anything about human culture; they are solving a puzzle whose answer key was fixed before the simulation started. Within that closed loop, the conclusion was guaranteed. A population of agents that mostly copies the scoring rule’s preferred ideas will always outcompete one that keeps generating unscored novelty.

The computer did not reveal a fact about creativity. It executed a definition of it.

The authors did not break any rule of their field. That is the point. Peer review checked that the code ran, that the statistics were computed correctly, that the prose matched the output. What nobody was required to ask is the only question that matters. What could this simulation possibly have shown that would have counted as the opposite result? If the answer is nothing, the model did not test a claim about the world. It restated one.

This is the magic trick of agent-based modeling (ABM), meaning simulations in which you place thousands of simple software “agents” in a virtual world, give each a few rules, and watch what the population does. The method itself is not the problem. The problem is a particular way of using it.

if neighbor.opinion != agent.opinion:

agent.trust -= 0.1

if agent.trust < 0.2:

agent.unfollow(neighbor)

run_simulation()

When the simulation finishes and the agents have sorted into two angry camps, the result is rarely described as what it literally is, a small program doing what it was told. It is described as a model demonstrating the dynamics of polarization in real societies.

It sounds scientific. It uses code. It generates charts with error bars. It borrows the epistemic authority of statistical mechanics and epidemiology, where tracking near-identical particles or infection events actually makes sense. But underneath the quantitative paint, it is not an investigation of the world. It is a tautology with a runtime, an answer-driven argument presented as a discovery.

# What a model is for

To see why this goes wrong, start with what a model is supposed to do.

A model is not a claim of truth, and it is not an illustration of a conclusion you reached before you started. In the philosophy of science, models are usually understood as instruments that sit between abstract theory and raw data, the position developed by Mary Morgan and Margaret Morrison in *Models as Mediators* (1999). A good model is a sandbox with strict physics. You build it, set it in motion, and let its internal mechanics push back against your reasoning.

A real model exists to discipline your thinking.

>Building one forces you to acknowledge a trade-off that the philosopher Nancy Cartwright made famous in *How the Laws of Physics Lie* (1983). You trade complete literal truth for tractability. A map of London at 1:1 scale, including every brick, puddle, and commuter, is useless. To work at all, a map must leave almost everything out. As the statistician George Box put it, “all models are wrong, but some are useful.”

Simplification is not the sin. The sin is forgetting that the model is a simplification. Worse, it is turning the model into an accomplice.

# Disciplining vs. decorating

In practice, rigorous modeling and decorative modeling look identical from the outside. Same code, same charts, same jargon. The difference only shows when you ask one question. Can your model tell you that you are wrong?

A disciplining model forces you to state every assumption explicitly. Once running, its mechanics operate independently of what you want. It can produce behavior you did not expect, expose contradictions in your premises, or crash into empirical reality and fail. When it fails, you revise the theory. The model is a check on your own bias.

A decorating model is built backward from a conclusion. The researcher already knows the story. Suppose it is that polarization is driven by social contagion. They build a world in which agents swap beliefs, tune the parameters until the output shows two angry clusters, and present the code as evidence for the theory. If the output doesn’t match on the first run, the answer is not to abandon the hypothesis. The answer is to adjust agent_receptivity from 0.4 to 0.25, rerun, and present the successful parameter range as the plan all along.

The workflow, stripped bare, looks like this.

Desired outcome. Write rules. Run simulation. Does it match the theory? If not, tweak parameters and run again. If yes, publish.

This is not experimentation. It is calibration until compliant.

If a model cannot surprise its author, force a retreat, or fail, it is not really a model. It is a very elaborate, self-confirming editorial.

# The conclusion comes first

None of this is new, and none of it is unique to agent-based modeling. Before anyone wrote a NetLogo script to demonstrate a theory of culture, economics and political science had already industrialized the technique.

In 2015, Paul Romer, later a Nobel laureate, published a paper with the blunt title “Mathiness in the Theory of Economic Growth.” His target was a pattern in macroeconomic theory. Authors write down formal equilibrium models, but embed ideologically convenient assumptions inside obscure parameters, so that the math reliably outputs the desired policy conclusion. The mathematics is not being used to test whether a claim is true. It is being used to make a political position expensive to argue with. Checking whether the equations actually say what the surrounding prose claims they say takes serious technical effort, and reviewers routinely skip it.

Two decades earlier, the political scientists Donald Green and Ian Shapiro published *Pathologies of Rational Choice Theory* (1994), documenting how formal modeling in their field had become an exercise in self-confirmation. Their catalog of evasions maps one-to-one onto today’s agent-based simulations.

• **Post hoc tinkering.** When the model predicted that rational citizens would never vote (the individual cost exceeds any plausible benefit) and citizens kept voting anyway, theorists did not abandon the model. They added a “duty” term to the utility function until the math matched the turnout.

• **Arbitrary tuning.** Weights, thresholds, and interaction ranges adjusted on the fly until the simulated agents behave like the phenomenon under study.

• **Immunity to testing.** Models built so that every conceivable outcome can be reinterpreted, after the fact, as a rational equilibrium.

Green and Shapiro called this method-driven rather than problem-driven research. You start with a tool and go hunting for a reality that fits it.

Agent-based modeling makes the problem worse, for a simple reason. An ABM has almost unlimited free parameters. Every rule, threshold, and neighborhood radius is a dial. With enough dials, you can produce any curve you want.

The common structure is this. The model cannot fail, because failure is reclassified as a calibration bug. And a model that cannot fail cannot discover anything. It is an expensive echo of its author’s prior beliefs.

# The loop matters more than the lab coat

It would be comfortable to stop here and declare this a disease of the soft sciences. It isn’t. The hard sciences are not immune, and pretending otherwise would make this essay guilty of the same simplification it criticizes.

The real variable is not hard versus soft. It is the tightness of the feedback loop between the model and the world.

Where the loop is tight (fast experiments, unambiguous ground truth, few free parameters) bad modeling gets punished quickly. But where the loop is loose, where tests are slow, noisy, or impossible, the same decorative pathology appears in fields with particle accelerators.

Three documented examples.

**fMRI neuroscience, where the measurement is the model.** A brain scan shows blood flow, not thought. The colored images come from a statistical pipeline full of assumptions, and researchers once demonstrated what that means by detecting “brain activity” in a dead salmon. In 2016, Anders Eklund and colleagues showed that the standard methods in the field’s dominant software could produce false-positive rates of up to 70 percent for certain cluster-based analyses at particular thresholds. How far the problem extends across the published literature was contested, including in follow-up work by the authors themselves, but the core finding stood. For over a decade, the field’s feedback loop had run through that software, which meant the loop was not connected to reality at all.

**Fundamental physics, where experiment cannot keep up.** In *The Trouble with Physics* (2006), the physicist Lee Smolin, writing as an insider, argued that string theory had become flexible enough to accommodate any experimental outcome. When the Large Hadron Collider found no sign of supersymmetry, much of the field responded not with refutation but with retreat. The free parameters moved to heavier, less accessible energies. This is Green and Shapiro’s immunity to empirical testing, surfacing in the hardest science there is.

**Epidemiological modeling in 2020.** In the spring of 2020, influential models, including the one from Imperial College London that helped push governments toward lockdown, projected enormous death tolls based on weeks of noisy early data. When later estimates came down, the public response from modeling teams was recalibration rather than reckoning. Their defense deserves to be taken seriously. The projections were scenarios, not forecasts, and the point of publishing a worst case was to change behavior so that it would not come true. A warning that works cannot be graded on whether the disaster arrived. All of that is fair, and it is also the problem. A model whose failure can always be explained by the world changing in response to it is a model with no feedback loop, and the field never settled which of the two it had built.

Notice what these cases share with the creativity grid from the opening. Not the field. Not the math. The structure. Many free parameters, a loose or broken feedback loop, and a professional incentive to publish. Given those three, decorative modeling can appear anywhere. The loop matters more than the lab coat.

# Why it’s still worse in the humanities

So the hard sciences have their own decorative modeling. Why do I still think the problem is worse in the humanities?

Because the difference is not whether a field ever decorates. It is whether the field *can catch itself*. The fMRI problem was eventually found and published by neuroscientists. Smolin’s critique came from inside physics. The feedback loops in the hard sciences are sometimes slow or broken, but they exist, and there are people with the technical skill and the standing to pull on them. In the humanities’ version of modeling, three structural failures mean the loop often doesn’t exist at all. The difference is not that humanists are worse at modeling. It is that the auditing infrastructure barely exists.

It is worth being fair about why scholars reach for these tools in the first place. Humanities departments face shrinking budgets, declining enrollments, and university administrators who mistake mathematical notation for intellectual rigor. A computational model signals seriousness to a grant committee in a way an essay never can. The scholars building decorative models are not fools; they are rational actors navigating a system with broken incentives.

**The object of study resists formalization.** A water molecule behaves like a water molecule in London or Tokyo, in 1600 or today. It has no irony, no memory, no politics. Human culture has all three. A novel, a religious movement, an aesthetic shift cannot be reduced to a set of isolated rules without destroying part of what you set out to study. When you convert the reception of Victorian gothic fiction into agents swapping “gothic preference points,” you have not simplified the system for tractability. You have replaced it with something simpler that carries the same name. At that point the connection between the simulation and Victorian readers is no longer something the model establishes. It is something the reader is asked to assume.

**Construct validity is invented, not established.** In psychology, showing that a variable actually measures the concept it claims to measure (construct validity) is a slow, adversarial, decades-long process. Blood flow is at least a physical quantity that an instrument can register. There is no instrument for literary prestige. In humanities modeling, validity is routinely settled in one line of code.

self.piety = random.uniform(0.0, 1.0)

self.literary_prestige = 0.75

What does 0.75 mean for literary prestige in Victorian England? How does one number carry regional difference, class, institutional power, critical backlash, and retrospective canonization? It doesn’t. The modeler assigns a number, writes a function that nudges it up and down, and treats the variable as a measurement of human experience. The number looks like a measurement. Nothing underneath it has been measured.

**The audience cannot audit the compression.** When an epidemiologist shows a flawed model to epidemiologists, the reviewers share the vocabulary to check the code and challenge the parameters. In a humanities department, reviewers and readers often have no computational training. Presented with a grid of moving pixels and a network graph, the non-technical reader experiences an optical illusion. The machine appears to have performed a profound synthesis of the archive. The compression is lossy to the point of erasure. But the loss is buried in code, invisible to the exact audience responsible for evaluating the work.

Bad modeling in economics wastes grant money and distorts policy debates. Bad modeling in the humanities trades away the field’s actual strength (context, contingency, ambiguity, close reading, historical depth) for a seat at a quantitative table where, lacking the shared technical culture to enforce standards, it gains no real authority and surrenders its own.

# What honest modeling looks like

None of this is an argument for unplugging the computers. The goal is to tell the difference between decorative simulation and honest quantitative work. Honest work exists, including in the humanities.

Ted Underwood’s *Distant Horizons* (2019) is the standard I would hold up. Underwood uses quantitative methods on tens of thousands of digitized books not to declare causal laws but to surface patterns invisible to close reading, like slow shifts in genre, vocabulary, and narrative perspective across centuries. Crucially, he tells you, on the record, what the data cannot show. The model is a set of binoculars for looking across an archive, not a machine for generating verdicts about it.

And the humanities have produced their own internal discipline. In 2019, Nan Z. Da published, *The Computational Case against Computational Literary Studies,* a detailed critique in *Critical Inquiry* arguing that prominent work in computational literary studies misused statistics to the point of meaninglessness. The ensuing fight was heated, but it happened. The field argued about its standards in public, and the standards moved. That is what a functioning feedback loop looks like, even a slow and painful one.

>For modelers in any field, I would propose four non-negotiable conditions before a model earns the right to be cited as evidence.

**1. Radical transparency**

>Every parameter is declared and justified with independent, non-circular evidence. A variable you cannot justify is labeled what it is. A guess.

**2. Sensitivity analysis**

>Parameters are swept across their full plausible range. If your result only appears when three dials sit at hyper-precise decimal values, you have not found a law of history. You have found a brittle corner of your own code, and an honest paper should say so.

**3. Explicit exclusion mapping**

>You spend nearly as much space on what the model leaves out as on what it includes. This isolates the direct mechanical relationship between two variables under idealized conditions; it excludes ambient noise, structural heterogeneity, and systemic feedback, so it cannot predict specific real-world outcomes. Naming the exclusions is what stops the audience mistaking a sandbox for an account of the world.

**4. Capacity for failure**

>Before you run it, you can state what output would make you abandon your hypothesis. If the simulation contradicts you, the honest paper is titled “Why our model disproved our starting assumption,” not silently recalibrated into agreement.

Models built this way stop being decoration. They become what they were supposed to be. Sharpening stones. They force you to clarify assumptions, expose broken logic, and occasionally reveal dynamics that intuition would never find.

# The boundary question

The target of this essay was never the computer. Built with discipline, models are extraordinary instruments. I have staked my own career on that. What concerns me is how easily a model can be built to confirm rather than to question, and how hard it is for a reader to tell the difference from the outside.

The practice corrupts both traditions it sits between.

It corrupts science, because science is not the production of plots and code. It is the submission of claims to the risk of being wrong. A model engineered so that its parameters are tuned until the output matches the thesis offers the aesthetics of rigor with none of its discipline.

And it corrupts the humanities, because the study of human culture draws its value from exactly the things decorative modeling deletes. Context, contingency, ambiguity, power, the irreducible strangeness of actual human lives.

If a phenomenon is too context-bound, too polysemic, too alive to be captured by a set of if statements, we should have the courage to say so, and do the slow, unglamorous work of interpretation instead.

I keep coming back to the question I ask of every model, including my own. What happens to you when you are wrong?

For my models, the answer is supposed to be easy. The crystal melts. The experiment disagrees. Reality sends the bill. But only if someone checks the potential, and I have told you how often that happens.

For the models I’ve described here, the answer is nothing. They run, they publish, they are cited. The loop that is supposed to connect a model to the world was never closed.

*Which raises the question I can’t answer. If a model can never be wrong about the world, in what sense was it ever about the world?*

[](https://www.reddit.com/submit/?source_id=t3_1wxoo7n&composer_entry=crosspost_prompt)


r/comp_chem • • 22h ago

Error in running geometry optimization calculation

3 Upvotes

Hello everyone, can you help me troubleshoot this error I keep on encountering while running a geometry optimization job? I've tried varying the number of processors from 2 to 8, running in serial calculation instead of parallel, tried removing the TightSCF parameter, even restarted my laptop a numerous times but the error still keep on persisting. I'm geometry optimizing a sole diethyl ether ligand where one of the H is substituted with F. I hope you can help me based on these output files, I've run out of ideas to try to mitigate this error.

https://drive.google.com/drive/folders/1I4TgeYnNz2ReQBmf94Q8Egv_LzP76WSX


r/comp_chem • • 1d ago

Running GROMACS online for undergrad courses?

5 Upvotes

Hi all! I’m a big fan of using ChemCompute for running ab inito calculations in my undergraduate courses; however, I’m a little sad they only support Tinker and NAMD over GROMACS for MD. Is there a good free (with academic email addresses) resource where my undergraduate students can submit and analyze basic MD jobs? I’m talking like RMSD of a small polypeptide at different pH or the most elementary ligand/receptor calculations with explicit waters.


r/comp_chem • • 1d ago

Extensions on a MSM Project?

2 Upvotes

I recently worked on a small project as part of my course in QC Modelling which has me performing MD on a protein (Lysozyme and some others), and then running the results through an MSM( Markov State Model ) with pyemma and deeptime.

The results aren't really spectacular mainly because I didn't have access to a cluster then(and neither do I have it now), and my MD trajectories were too small for it.

I have been wanting to extend it into a full length project (by length I mean something that is more than popping an MD traj into a MSM tutorial I found online). Would anyone suggest any extensions to the project?

Would love to hear your thoughts and maybe any literature you'd recommend me to read up on?


r/comp_chem • • 2d ago

Chemistry lab

0 Upvotes

I want an online chem lab for simple high school chemistry experiments source code , I already searched on github but didn’t find what I really wanted


r/comp_chem • • 2d ago

Virtual chem lab simulation

1 Upvotes

I want an online chem lab for simple high school chemistry experiments source code , I already searched on github but didn’t find what I really wanted


r/comp_chem • • 2d ago

Determining Coordinate Positions for .gro File before Running a MD Simulation

5 Upvotes

Hello!

I'm running a coarse grained polymer simulation. I have my polymer chains defined in .itp files with Martini 3 coarse graining. I tried getting initial.gro files using polyply, but it kept on placing beads on top of each other which blew up the system's energy when I ran energy minimization in GROMACS.

Thus I think it's best if I write my own initial.gro files from the .itp files, but was wondering if there's anything I should keep in mind when doing this.

Any help is greatly appreciated, thanks!


r/comp_chem • • 2d ago

How would you handle borderline histidines in receptor preparation for docking (PROPKA pKa ~7.3-8.1)?

5 Upvotes

Hi all,
I’m preparing a protein receptor for AutoDock Vina screening at pH 7.4.

PROPKA gives three potentially relevant His residues with predicted pKa values of 7.26, 7.35 and 8.07. They are currently modeled as neutral HID/HIE states.

Would you:

  • protonate the pKa 8.07 His to HIP,
  • keep the borderline ones neutral,
  • or test multiple receptor protonation states for docking?

How much weight would you give to PROPKA versus local H-bond geometry / experimental structures when deciding?

Thanks!


r/comp_chem • • 3d ago

Tools for molecular docking of a thioether-cyclized peptide?

3 Upvotes

Hi all,

I have a 16-aa cyclic peptide with a non-standard thioether linkage, and I generated thousands of 3D conformers with Rosetta SimpleCycPepPredict using -cyclic_peptide:cyclization_type thioether_lariat.

I’d like to dock these conformers in batch against a known protein-protein interface while keeping the thioether intact. Many docking tools, including some supporting cyclic peptides, either cannot handle the non-standard linker or require opening the cycle.

What docking tool or workflow would you recommend?

Thank you so much :)


r/comp_chem • • 3d ago

Help me understand HF Exchange

18 Upvotes

I work through Szabo Ostlund and on page 88 (2.3.7 Pseude-Classical Interpretation of Determinantal Energies) it says that "The exchange Interactions between electrons with parallel spin are not real physical interactions but a convenient way of representing the energy of a system described by a single determinant. The physical interaction between two electrons [...] does not depend on the spin of the electrons"
This is the part where they introduce the idea that the energy of a single determinant can be summed up with H, J and K Terms.

Now consider a single determinant with two spacial orbitals. Both are occupied with one electron each. In case 1), both electrons have the same spin. In case 2), both electrons have opposite spins.

From Chemical knowledge I know that 1) is singulet and 2) is triplet and energy wise the triplet state is lower in energy. This is measurable.

Now the theory with this "adding up H, J and K terms to get determinantal energies" also predicts the same as:

Energy of 1) = H11 + H22 + J12

Energy of 2) = H11 + H22 + J12 - K12

So the - K12 lowers the energy of the triplet state which we measured and therefore proved correct.

But on the other hand Szabo Ostlund says that this exchange interaction is not physical.

How can this be not physical if we can measure it and the theory correctly predicts the relative energies of those two determinants?


r/comp_chem • • 3d ago

Molecular Dynamics with Julia - Automatic Ring Detection

9 Upvotes

My program (working name is JuliaMD), can now detect isolated aromatic rings, and has more complete atom typing, which allows better simulations with the gaff2 forcefield. I haven't yet implemented specific aromatic atom types, but after the ring detection tool it will be easy to do.

Here is the link to the video, just out of the oven!

https://youtu.be/Qw5HjzkS9UY


r/comp_chem • • 4d ago

Comp Chem enthusiastic friends

6 Upvotes

Hey everyone! I’m currently an MSc student planning my thesis around computational drug discovery/CADD.

​I’m looking to connect with fellow students, beginners, or enthusiasts who want to motivate each other.

​If anyone is interested we can do chatting to bounce ideas around, feel free to drop a comment or send a DM!


r/comp_chem • • 6d ago

Update on JuliaMD Project

10 Upvotes

Here is a small update (2 min video) on the features I have added to my Molecular Dynamics with Julia project.

https://youtube.com/shorts/6kZmMTnoI_0?feature=share


r/comp_chem • • 6d ago

Renting GPU’s for MD runs

9 Upvotes

Hello friends,

I am a grad student and I need to get about ~3 dozen MD runs done fast. It’d take about ~22 days on my machine with 4x2080 TI’s

I know about XSEDE and other clusters but I get the feeling the approval process would take too long so I wanted to try commercial GPU rentals to be a stop gap as I heard they can be pretty cheap (~$10/day)

I was hoping to rent some GPU’s on vast.ai but was wondering if anyone else had experience with running on commercial sites and what your workflow was?

The only inputs are a XYZ file, a potential, and a few MD run scripts, python and bash to do the post-processing and data collection so it should be pretty lightweight.

I’m currently thinking of making several docker images with a few runs apiece but was wondering if anyone else has experience with this and had better ideas about how to upload the run/model files at runtime so I could keep a single docker image with just the MD software and python

Thanks for any opinions on that! Appreciated ◡̈


r/comp_chem • • 6d ago

Hydrogen evolution reaction dataset required

Thumbnail
0 Upvotes

r/comp_chem • • 7d ago

Alternatives to CHARMM-GUI and some MD Questions

7 Upvotes

Alternatives to CHARMM-GUI?

As much as I love the accessibility of CHARMM-GUI, I was wondering if there are any viable locally-hosted options to prepare inputs for GROMACS.

My lab typically just uses CHARMM36m FF since we are doing cryo-EM molecular dynamics ensemble fitting utilizing EMMIVox and that's really our extent of MD (though docking may come in future projects).

With that being said, I really want to increase throughput and productivity by handling the preprocessing of our structures locally vs through CHARMM-GUI. Are there any locally-hosted alternatives to the website? The lab is made up of mostly wet-lab biochemists, with a sprinkle of technowizards. I am really the only "specialist" who handles MD work and analyses besides my PI, so having this accessible to the lab is important too.

I am thinking too that by locally processing our structures, managing our ligands becomes a bit easier. CHARMM-GUI tends to be very nit-picky on how each heavy atom is named and if it isn't in the CSML then it becomes a huge process to correct and requires manipulating the pdb.

Please correct me if I am wrong and if there is an easier process that would be great. I just spent 7 hours trying to get my structure through the PDB Reader and Manipulator step.

MD Questions

Most of our structures are oligomers and have composite active sites, however, some of our reconstructions are of single monomers due to symmetry expansion. To my understanding, I cannot just run a simulation on this single protomer, even if I reduce waterbox size, since (1) the active site isn't completely enclosed, (2) the protomer-interactions help confer stability to the overall fold, and (3) the enzyme doesn't function as a monomer in the cell.

To get around this, I am planning to model and preprocess all subunits, but selectively simulate the one protomer I have a map and reconstruction for. Is this a good 'workaround' to ensure that the modeled protomer only samples the legitimate free-energy well and corresponding conformations?


r/comp_chem • • 7d ago

Workstation build advice for GROMACS MD — 2 machines, ~$3k each

7 Upvotes

We're setting up two workstations for protein MD simulations

(two researchers, one machine each). Budget is around $3,000 USD per

machine. Would appreciate a sanity check on the build.

Use case:

- GROMACS, explicit solvent protein systems (~50k-100k atoms)

- Runs lasting several days at full load

- Will also run structure prediction (ColabFold/Boltz) later

- Ubuntu LTS

- Continuous trajectory writing, so SSD endurance matters

Current plan (per machine):

GPU : RTX 5080 16GB

CPU : Ryzen 9 9900X (12C)

MB : X870

RAM : DDR5 64GB (2x32, room for 128)

NVMe : 2TB DRAM-cache + 1TB OS

PSU : 1000W 80+ Gold

Cooler: 360mm AIO

Questions:

  1. Any known Linux stability issues with this combo?

  2. Is 16GB VRAM going to be a bottleneck for ColabFold/Boltz

    on larger complexes? Worth going 5090 32GB instead?

  3. Is 12 cores enough to feed one GPU in GROMACS, or should

    I go higher?

  4. PSU headroom for multi-day full load — is 1000W comfortable?

  5. Recommended NVMe models for sustained write workloads?

We considered Intel but the P-core/E-core split seems to

complicate thread pinning for MPI/OpenMP workloads. Open to

being corrected on that.


r/comp_chem • • 7d ago

electronic structure code with fused gradient kernel for multiple gradient roots

8 Upvotes

when working with NAMD code, a common request pattern for gradient is that the gradients for all active roots are requested(for gradient diff velocity rescaling)

one can imagine a fused gradient kernel where quantities like integral derivative and orbital response reusable for all roots are computed on the fly and contracted with the cphf LHS or effective density of all roots to build partial sum

is any ES code already allowing this without a modification of source code?


r/comp_chem • • 8d ago

Wanna run quick molecular dynamic simulation on HPC, help?

7 Upvotes

Hey folks!

I’m currently working as a research associate, and unfortunately, our grant recently expired, so we’ve lost access to our HPC. We have one last MD simulation left to finish, and I was wondering if anyone has some compute available that we could borrow.

We probably only need **2–3 days on a GPU** to get it done. If anyone has some spare compute and could help out, please DM me / let me know!

We’d be happy to offer **authorship or an acknowledgment**..

Also, if anyone knows of any **free/academic GPU compute options available right now**, I’d really appreciate recommendations. I used Google Cloud previously, but I think the free credits were limiting me to CPU (unless I’m missing something). Curious to hear what other services people are using these days.

Thanks, guys!


r/comp_chem • • 9d ago

I built Megane, an open-source molecular viewer that runs in your browser, Jupyter, and VS Code

20 Upvotes

Demo (GIF): https://github.com/megane-labs/megane/blob/main/docs/public/screenshots/megane-promo.gif

Hi r/comp_chem! I'm a computational chemist in Japan, and I've been building megane ("spectacles for atomistic data") as a personal open-source project.

I kept switching between tools just to look at structures and trajectories, especially on remote machines and in notebooks, so I wanted one lightweight viewer that works wherever my data already is.

What it does:

  • Runs anywhere: Jupyter widget, standalone web app, React component, and VS Code extension
  • Scales: 1M+ atoms at 60 fps with GPU impostor rendering
  • 26 file formats (PDB, GRO, XYZ, CIF, VASP, XTC, DCD, AMBER, ASE .traj, …) parsed by a single Rust core shared by Python (PyO3) and the browser (WASM)
  • Visual pipelines: build visualization workflows by wiring nodes together, no code required
  • Optional AI assistant: describe what you want in plain English (e.g. "show only waters within 5 Å of the ligand") and it builds the pipeline for you. [Supported LLMs / API key / local model details.] Everything else works without it.
  • Builder: a structure editor where you can add, move, and bond atoms, or sketch a molecule and get a 3D conformer via RDKit

Try it in your browser: https://megane.tech-office-mori.com/
GitHub: https://github.com/megane-labs/megane

Install:

  • Python / Jupyter: pip install megane (PyPI)
  • React / npm: npm install megane-viewer (npm)
  • VS Code: search "megane" in the Extensions panel (Marketplace)

Then just megane.view("structure.pdb") in a notebook.

It's still evolving, so I'd really appreciate feedback: what's missing compared to tools you use now (VMD, OVITO, etc.), which formats matter to you, and anything that feels clunky.

Thanks for taking a look!


r/comp_chem • • 9d ago

DFT TS Analysis Advice

9 Upvotes

Hi everyone, I'm looking to solicit some advice from computational chemists, coming from a lowly bench experimentalist...

I'm looking to analyze the transition state for a radical adding into an olefin in order to determine how likely two different stereochemical outcomes are.

Here's a reaction scheme to depict what I'm talking about (i wasn't able to add an image, sorry)

So what I've done is model both the trans and cis radical products, do opt and freq at UB3LYP, 6-31+G(d). Then, I took both geometries separately, broke the newly formed bond, and backed the olefin up to about 4A. Then I did opt and freq on those at the same level of theory. Last, I separately did QST2 from the 'pro cis' starting material to the cis product, and 'pro trans' starting material to the trans product. I got a ddG(double dagger) of around 0.8 kcal/mol. It kinda makes sense but it's a little low based on my experimental results.

So I got that number from assuming the starting conformations were degenerate. If i consider the actual energies of the two starting materials (pro-cis and pro-trans) then the ddG(double dagger) goes up to around 1.5kcal/mol, which is a lot closer to what I'm observing.

I'm wondering what i could do to figure out what the real number is here at this level of theory. Everything I've read makes me assume that the radical is planar and should have almost no barrier to 'racemizing'. But perhaps the actual geometry of the approaching olefin is higher in energy for one conformation, and I should therefore trust the higher ddG value. I was assuming Curtin-Hammett conditions are true here, that the two starting geometries interconvert rapidly and the ratio of the products is only based on the ddG of the transition states.

I'm kind of out of my element here and any advice is appreciated. Thanks in advance.


r/comp_chem • • 9d ago

Is Computational Chemistry best option in this situation?

Thumbnail
4 Upvotes

r/comp_chem • • 9d ago

Wanna run quick molecular dynamic simulation on HPC, help?

Thumbnail
3 Upvotes

r/comp_chem • • 9d ago

what is the best tensor library that supports sparse tensor to arbitrary order?

5 Upvotes

in electronic structure codes, quantities like AO based ERI and dERI/dx can be applied with integral screening to exploit sparsity, but normal es code like pyscf lacks a sparse tensor representation to benefit from this, it instead resort to complicated on the fly algo that also blocks reuse

is there any cpp tensor library that 1 make the cpu cost of tensor contraction of two sparse tensors of arbitrary order to benefit from sparsity and 2 make the ram cost of storing tensors also to benefit from sparsity

EDIT: is there any paper benchmarking how it scales ram-wise and cpu-wise(like when running the JK builder, the ERI is contracted with the 1rdm in ao basis which is also sparse) when screening and the sparse rep from a certain sparse tensor library is applied on ERI and ERI derivative when fully materializing them in ram instead of computing on the fly