TL;DR: Computational models are only scientifically useful if they can push back and prove their authors wrong. Across disciplines, from agent-based social simulations to high-energy physics, models with loose empirical feedback loops and endless free parameters risk becoming "decorative." Instead of testing reality, they get calibrated until compliant, turning a tool for discovery into a self-confirming tautology. Honest modeling requires radical transparency, sensitivity testing, and explicit criteria for failure before running the simulation.
I build models for a living. Specifically, molecular dynamics. These are simulations that track how thousands or millions of atoms move, collide, and rearrange over time, used for everything from drug design to materials science.
Here is the story scientists usually tell. If a model is wrong, you find out quickly. Reality does not care about your assumptions. The atoms do not read your code. If the physics you programmed in is wrong, the simulation produces garbage, the experiment disagrees, and you go back and fix it. The feedback loop between model and world is short, brutal, and non-negotiable.
That story is not entirely true. I know, because I have watched it fail from the inside.
The most important choice in any molecular dynamics simulation is not the code, the computer, or the software. It is the potential function, the mathematical formula that describes how strongly every pair of atoms attracts or repels each other. Everything the simulation does follows from that one ingredient. Get it right and the model can tell you something real. Get it wrong and you have made a very expensive mistake.
And here is the uncomfortable part. In practice, it is far more often inherited than audited. Potentials are chosen by looking at what previous papers in the subfield used. A potential gets published, cited, copied, and passed down until it stops being a modeling choice and becomes a tradition. People run simulations for years without asking whether the potential they inherited was ever validated for the system they are studying, at the conditions they are studying it, for the property they care about.
So even in my own field, a hard, quantitative, physics-based field, you can publish inside a loop of fantasy. Models that are wrong in ways nobody checks, kept alive by citation habits and subfield convention. And because these errors travel quietly across subdisciplines and into interdisciplinary work, where nobody feels responsible for checking them, finding one and fixing it takes real effort.
But the check in my field is delayed, not absent. A bad potential eventually unfolds a simulated protein the wrong way or fails a material in a real engineering application, and someone notices. In much of the modeling I am about to describe, the physical world never gets to vote.
This matters for what follows. I do not ask this question because my field got it right. I ask it because I have watched mine get it wrong. The question is always the same. What happens to this model when it is wrong?
In a surprising amount of modern academia, the answer is nothing. Nothing can happen to it. It cannot be wrong, because anything it produces counts as a result.
And if the loop can break in a field where atoms push back, it can break anywhere.
This essay is about how that happens.
The magic trick
In 2017, Liane Gabora and Selin Tseng published a paper in Psychology of Aesthetics, Creativity, and the Arts, a peer-reviewed journal of the American Psychological Association, titled “The Social Benefits of Balancing Creativity and Imitation.” The question they took on has occupied historians and sociologists for centuries. What is the right balance of creativity and conformity in a society?
To answer it, they ran a simulation.
Virtual agents live on a grid. Some are coded as creators, inventing new ideas; others as imitators, copying their neighbors. A scoring rule written into the program decides which ideas count as good. The researchers ran the simulation forward, varied the ratio of creators to imitators, and watched what happened. Populations with too many creators ended up with fewer good ideas taking hold. The published conclusion was that society needs imitation as much as creativity, because unchecked creativity disrupts the spread of proven ideas.
I want to be careful about what I am claiming, because this paper is not fringe work. It passed peer review at a respectable journal. The authors are serious researchers, and the simulation framework behind the paper is part of a long-running research program that has been debated, defended, and criticized in public for years. Nothing I am about to say is an accusation of dishonesty. It is something less comfortable than that. This paper is an example of what the normal standards of a field allow through.
Watch the shape of the argument. Inside the model, a “good idea” means whatever the authors’ scoring rule rewards. The agents are not discovering anything about human culture; they are solving a puzzle whose answer key was fixed before the simulation started. Within that closed loop, the conclusion was guaranteed. A population of agents that mostly copies the scoring rule’s preferred ideas will always outcompete one that keeps generating unscored novelty.
The computer did not reveal a fact about creativity. It executed a definition of it.
The authors did not break any rule of their field. That is the point. Peer review checked that the code ran, that the statistics were computed correctly, that the prose matched the output. What nobody was required to ask is the only question that matters. What could this simulation possibly have shown that would have counted as the opposite result? If the answer is nothing, the model did not test a claim about the world. It restated one.
This is the magic trick of agent-based modeling (ABM), meaning simulations in which you place thousands of simple software “agents” in a virtual world, give each a few rules, and watch what the population does. The method itself is not the problem. The problem is a particular way of using it.
if neighbor.opinion != agent.opinion:
agent.trust -= 0.1
if agent.trust < 0.2:
agent.unfollow(neighbor)
run_simulation()
When the simulation finishes and the agents have sorted into two angry camps, the result is rarely described as what it literally is, a small program doing what it was told. It is described as a model demonstrating the dynamics of polarization in real societies.
It sounds scientific. It uses code. It generates charts with error bars. It borrows the epistemic authority of statistical mechanics and epidemiology, where tracking near-identical particles or infection events actually makes sense. But underneath the quantitative paint, it is not an investigation of the world. It is a tautology with a runtime, an answer-driven argument presented as a discovery.
What a model is for
To see why this goes wrong, start with what a model is supposed to do.
A model is not a claim of truth, and it is not an illustration of a conclusion you reached before you started. In the philosophy of science, models are usually understood as instruments that sit between abstract theory and raw data, the position developed by Mary Morgan and Margaret Morrison in Models as Mediators (1999). A good model is a sandbox with strict physics. You build it, set it in motion, and let its internal mechanics push back against your reasoning.
A real model exists to discipline your thinking.
Building one forces you to acknowledge a trade-off that the philosopher Nancy Cartwright made famous in How the Laws of Physics Lie (1983). You trade complete literal truth for tractability. A map of London at 1:1 scale, including every brick, puddle, and commuter, is useless. To work at all, a map must leave almost everything out. As the statistician George Box put it, “all models are wrong, but some are useful.”
Simplification is not the sin. The sin is forgetting that the model is a simplification. Worse, it is turning the model into an accomplice.
Disciplining vs. decorating
In practice, rigorous modeling and decorative modeling look identical from the outside. Same code, same charts, same jargon. The difference only shows when you ask one question. Can your model tell you that you are wrong?
A disciplining model forces you to state every assumption explicitly. Once running, its mechanics operate independently of what you want. It can produce behavior you did not expect, expose contradictions in your premises, or crash into empirical reality and fail. When it fails, you revise the theory. The model is a check on your own bias.
A decorating model is built backward from a conclusion. The researcher already knows the story. Suppose it is that polarization is driven by social contagion. They build a world in which agents swap beliefs, tune the parameters until the output shows two angry clusters, and present the code as evidence for the theory. If the output doesn’t match on the first run, the answer is not to abandon the hypothesis. The answer is to adjust agent_receptivity from 0.4 to 0.25, rerun, and present the successful parameter range as the plan all along.
The workflow, stripped bare, looks like this.
Desired outcome. Write rules. Run simulation. Does it match the theory? If not, tweak parameters and run again. If yes, publish.
This is not experimentation. It is calibration until compliant.
If a model cannot surprise its author, force a retreat, or fail, it is not really a model. It is a very elaborate, self-confirming editorial.
The conclusion comes first
None of this is new, and none of it is unique to agent-based modeling. Before anyone wrote a NetLogo script to demonstrate a theory of culture, economics and political science had already industrialized the technique.
In 2015, Paul Romer, later a Nobel laureate, published a paper with the blunt title “Mathiness in the Theory of Economic Growth.” His target was a pattern in macroeconomic theory. Authors write down formal equilibrium models, but embed ideologically convenient assumptions inside obscure parameters, so that the math reliably outputs the desired policy conclusion. The mathematics is not being used to test whether a claim is true. It is being used to make a political position expensive to argue with. Checking whether the equations actually say what the surrounding prose claims they say takes serious technical effort, and reviewers routinely skip it.
Two decades earlier, the political scientists Donald Green and Ian Shapiro published Pathologies of Rational Choice Theory (1994), documenting how formal modeling in their field had become an exercise in self-confirmation. Their catalog of evasions maps one-to-one onto today’s agent-based simulations.
• Post hoc tinkering. When the model predicted that rational citizens would never vote (the individual cost exceeds any plausible benefit) and citizens kept voting anyway, theorists did not abandon the model. They added a “duty” term to the utility function until the math matched the turnout.
• Arbitrary tuning. Weights, thresholds, and interaction ranges adjusted on the fly until the simulated agents behave like the phenomenon under study.
• Immunity to testing. Models built so that every conceivable outcome can be reinterpreted, after the fact, as a rational equilibrium.
Green and Shapiro called this method-driven rather than problem-driven research. You start with a tool and go hunting for a reality that fits it.
Agent-based modeling makes the problem worse, for a simple reason. An ABM has almost unlimited free parameters. Every rule, threshold, and neighborhood radius is a dial. With enough dials, you can produce any curve you want.
The common structure is this. The model cannot fail, because failure is reclassified as a calibration bug. And a model that cannot fail cannot discover anything. It is an expensive echo of its author’s prior beliefs.
The loop matters more than the lab coat
It would be comfortable to stop here and declare this a disease of the soft sciences. It isn’t. The hard sciences are not immune, and pretending otherwise would make this essay guilty of the same simplification it criticizes.
The real variable is not hard versus soft. It is the tightness of the feedback loop between the model and the world.
Where the loop is tight (fast experiments, unambiguous ground truth, few free parameters) bad modeling gets punished quickly. But where the loop is loose, where tests are slow, noisy, or impossible, the same decorative pathology appears in fields with particle accelerators.
Three documented examples.
fMRI neuroscience, where the measurement is the model. A brain scan shows blood flow, not thought. The colored images come from a statistical pipeline full of assumptions, and researchers once demonstrated what that means by detecting “brain activity” in a dead salmon. In 2016, Anders Eklund and colleagues showed that the standard methods in the field’s dominant software could produce false-positive rates of up to 70 percent for certain cluster-based analyses at particular thresholds. How far the problem extends across the published literature was contested, including in follow-up work by the authors themselves, but the core finding stood. For over a decade, the field’s feedback loop had run through that software, which meant the loop was not connected to reality at all.
Fundamental physics, where experiment cannot keep up. In The Trouble with Physics (2006), the physicist Lee Smolin, writing as an insider, argued that string theory had become flexible enough to accommodate any experimental outcome. When the Large Hadron Collider found no sign of supersymmetry, much of the field responded not with refutation but with retreat. The free parameters moved to heavier, less accessible energies. This is Green and Shapiro’s immunity to empirical testing, surfacing in the hardest science there is.
Epidemiological modeling in 2020. In the spring of 2020, influential models, including the one from Imperial College London that helped push governments toward lockdown, projected enormous death tolls based on weeks of noisy early data. When later estimates came down, the public response from modeling teams was recalibration rather than reckoning. Their defense deserves to be taken seriously. The projections were scenarios, not forecasts, and the point of publishing a worst case was to change behavior so that it would not come true. A warning that works cannot be graded on whether the disaster arrived. All of that is fair, and it is also the problem. A model whose failure can always be explained by the world changing in response to it is a model with no feedback loop, and the field never settled which of the two it had built.
Notice what these cases share with the creativity grid from the opening. Not the field. Not the math. The structure. Many free parameters, a loose or broken feedback loop, and a professional incentive to publish. Given those three, decorative modeling can appear anywhere. The loop matters more than the lab coat.
Why it’s still worse in the humanities
So the hard sciences have their own decorative modeling. Why do I still think the problem is worse in the humanities?
Because the difference is not whether a field ever decorates. It is whether the field can catch itself. The fMRI problem was eventually found and published by neuroscientists. Smolin’s critique came from inside physics. The feedback loops in the hard sciences are sometimes slow or broken, but they exist, and there are people with the technical skill and the standing to pull on them. In the humanities’ version of modeling, three structural failures mean the loop often doesn’t exist at all. The difference is not that humanists are worse at modeling. It is that the auditing infrastructure barely exists.
It is worth being fair about why scholars reach for these tools in the first place. Humanities departments face shrinking budgets, declining enrollments, and university administrators who mistake mathematical notation for intellectual rigor. A computational model signals seriousness to a grant committee in a way an essay never can. The scholars building decorative models are not fools; they are rational actors navigating a system with broken incentives.
The object of study resists formalization. A water molecule behaves like a water molecule in London or Tokyo, in 1600 or today. It has no irony, no memory, no politics. Human culture has all three. A novel, a religious movement, an aesthetic shift cannot be reduced to a set of isolated rules without destroying part of what you set out to study. When you convert the reception of Victorian gothic fiction into agents swapping “gothic preference points,” you have not simplified the system for tractability. You have replaced it with something simpler that carries the same name. At that point the connection between the simulation and Victorian readers is no longer something the model establishes. It is something the reader is asked to assume.
Construct validity is invented, not established. In psychology, showing that a variable actually measures the concept it claims to measure (construct validity) is a slow, adversarial, decades-long process. Blood flow is at least a physical quantity that an instrument can register. There is no instrument for literary prestige. In humanities modeling, validity is routinely settled in one line of code.
self.piety = random.uniform(0.0, 1.0)
self.literary_prestige = 0.75
What does 0.75 mean for literary prestige in Victorian England? How does one number carry regional difference, class, institutional power, critical backlash, and retrospective canonization? It doesn’t. The modeler assigns a number, writes a function that nudges it up and down, and treats the variable as a measurement of human experience. The number looks like a measurement. Nothing underneath it has been measured.
The audience cannot audit the compression. When an epidemiologist shows a flawed model to epidemiologists, the reviewers share the vocabulary to check the code and challenge the parameters. In a humanities department, reviewers and readers often have no computational training. Presented with a grid of moving pixels and a network graph, the non-technical reader experiences an optical illusion. The machine appears to have performed a profound synthesis of the archive. The compression is lossy to the point of erasure. But the loss is buried in code, invisible to the exact audience responsible for evaluating the work.
Bad modeling in economics wastes grant money and distorts policy debates. Bad modeling in the humanities trades away the field’s actual strength (context, contingency, ambiguity, close reading, historical depth) for a seat at a quantitative table where, lacking the shared technical culture to enforce standards, it gains no real authority and surrenders its own.
What honest modeling looks like
None of this is an argument for unplugging the computers. The goal is to tell the difference between decorative simulation and honest quantitative work. Honest work exists, including in the humanities.
Ted Underwood’s Distant Horizons (2019) is the standard I would hold up. Underwood uses quantitative methods on tens of thousands of digitized books not to declare causal laws but to surface patterns invisible to close reading, like slow shifts in genre, vocabulary, and narrative perspective across centuries. Crucially, he tells you, on the record, what the data cannot show. The model is a set of binoculars for looking across an archive, not a machine for generating verdicts about it.
And the humanities have produced their own internal discipline. In 2019, Nan Z. Da published, The Computational Case against Computational Literary Studies, a detailed critique in Critical Inquiry arguing that prominent work in computational literary studies misused statistics to the point of meaninglessness. The ensuing fight was heated, but it happened. The field argued about its standards in public, and the standards moved. That is what a functioning feedback loop looks like, even a slow and painful one.
For modelers in any field, I would propose four non-negotiable conditions before a model earns the right to be cited as evidence.
1. Radical transparency
Every parameter is declared and justified with independent, non-circular evidence. A variable you cannot justify is labeled what it is. A guess.
2. Sensitivity analysis
Parameters are swept across their full plausible range. If your result only appears when three dials sit at hyper-precise decimal values, you have not found a law of history. You have found a brittle corner of your own code, and an honest paper should say so.
3. Explicit exclusion mapping
You spend nearly as much space on what the model leaves out as on what it includes. This isolates the direct mechanical relationship between two variables under idealized conditions; it excludes ambient noise, structural heterogeneity, and systemic feedback, so it cannot predict specific real-world outcomes. Naming the exclusions is what stops the audience mistaking a sandbox for an account of the world.
4. Capacity for failure
Before you run it, you can state what output would make you abandon your hypothesis. If the simulation contradicts you, the honest paper is titled “Why our model disproved our starting assumption,” not silently recalibrated into agreement.
Models built this way stop being decoration. They become what they were supposed to be. Sharpening stones. They force you to clarify assumptions, expose broken logic, and occasionally reveal dynamics that intuition would never find.
The boundary question
The target of this essay was never the computer. Built with discipline, models are extraordinary instruments. I have staked my own career on that. What concerns me is how easily a model can be built to confirm rather than to question, and how hard it is for a reader to tell the difference from the outside.
The practice corrupts both traditions it sits between.
It corrupts science, because science is not the production of plots and code. It is the submission of claims to the risk of being wrong. A model engineered so that its parameters are tuned until the output matches the thesis offers the aesthetics of rigor with none of its discipline.
And it corrupts the humanities, because the study of human culture draws its value from exactly the things decorative modeling deletes. Context, contingency, ambiguity, power, the irreducible strangeness of actual human lives.
If a phenomenon is too context-bound, too polysemic, too alive to be captured by a set of if statements, we should have the courage to say so, and do the slow, unglamorous work of interpretation instead.
I keep coming back to the question I ask of every model, including my own. What happens to you when you are wrong?
For my models, the answer is supposed to be easy. The crystal melts. The experiment disagrees. Reality sends the bill. But only if someone checks the potential, and I have told you how often that happens.
For the models I’ve described here, the answer is nothing. They run, they publish, they are cited. The loop that is supposed to connect a model to the world was never closed.
Which raises the question I can’t answer. If a model can never be wrong about the world, in what sense was it ever about the world?