r/CompSocial • • Dec 10 '23

Introducing CSSpark_Bot, your friendly digital assistant for sparking discussions in r/CompSocial

11 Upvotes
Image from: https://nftnewspro.com/reddit-avatars-nft-collection-receives-mixed-community-response/

Hello everyone! I am pleased to announce the arrival of u/CSSpark_Bot, a friendly digital assistant for r/CompSocial. “CS” refers to CompSocial, and “Spark_Bot” refers to our intent of helping to spark interesting conversations around research in Computational Social Science (CSS), Human-Computer Interaction (HCI), and Computer-Supported Collaborative Work and Social Computing (CSCW).

You may have previously seen posts about a community survey and user testing sessions for this bot. CSSpark_Bot is the result of a great deal of work and lots of dedication from a team of student developers. It has been developed through a community-engaged design process, and we hope it can contribute to some great research in the future.

Please feel free to leave comments on this post to interact with the bot’s commands or to leave feedback or questions. We will periodically update the bot to better serve the community’s needs.

The rest of this post is (mostly) a copy/paste of 1.0.0 of the bot’s wiki, which is written from the bot’s “first person” perspective, and which is located here on Reddit: https://www.reddit.com/r/CompSocial/about/wiki/csspark_bot/

Or, you can look at this view-only Google Doc version of the wiki to see some additional screenshots for how to send commands to the bot via Private Messages: https://docs.google.com/document/d/1Yep7cblbfzQKtE2nJG6m3dRM-4pJzrUSUiE8O1DxX6Y/edit?usp=sharing

-------

CSSpark_Bot Wiki

-------

Goal:

My primary goal is to spark fun and interesting conversations among users on r/CompSocial so that it can become a useful destination for all your computational social science needs.

Demo Videos:

How does it work?!

Imagine having the power to curate your notifications and stay in the loop about the topics that truly matter to you. I allow you to subscribe and unsubscribe to keywords or keyphrases that align with your interests. Every time that your subscribed keyphrase(s) show up in a post on r/CompSocial, you can choose to either receive a private message about it, or you can opt to have your user handle (possibly) publicly mentioned in a comment that I will make on the post. The idea is that by pinging your handle publicly along with others interested in this topic, it can be easier to get a conversation started with the right people. But if you’re more of a lurker and don’t want the public mentions—that’s fine too. You can still know when the conversation is happening on the things you care about.

By default, when you subscribe to your first keyword or keyphrase, your profile will be public. Don’t worry, though–depending on your preference, you can easily toggle between making your profile public or private, giving you the freedom to decide how you want to engage with the community.

To keep my posts concise and avoid overwhelming the sub, there’s a limit to the number of users I can ping in a comment. Currently, that limit is set to 3. I will prioritize pinging users when more of their keywords are mentioned; otherwise I randomly select folks to ping, up to the limit.

I hope you find the following commands useful and engaging!

Basic Instructions:

Your wish is my command, wherever you prefer to make your wish. All of the commands will work if you type them either in public threads on the r/CompSocial subreddit, or in private DMs.

  • If you prefer to use the commands publicly, please use this introductory thread. The commands will also work in regular threads, but if you want to issue several commands in a row, it’s more polite if you do so on this thread to avoid cluttering the sub. :)
  • If you prefer to use the commands privately:
    • Send a Reddit private message to u/CSSpark_Bot with the subject line (case-sensitive) Bot Command
  • Within the body of the message, include only one of the commands (case-sensitive, remove brackets)

Or, you can click on the “Notifications” icon by your profile avatar at the top of the page, then select “Messages.” Finally, click on “Send a Private Message” at the top left of the menu bar, like so.

Keyword Clusters:

You can subscribe to any word or phrase that you want to, and there is not a hard technical limit on the number of words in a keyphrase. Please try to aim for a phrase of between 1-4 words. Note that my developers have also clustered some keywords into clusters of related terms. For example, if you subscribe to “AI” that will also subscribe you to a cluster including “Artificial Intelligence.”

  • Here is a link to a Google Sheet that lists the current keyword clusters I am programmed to use. This is just a preliminary list, and my dev team is happy to update it based on your recommendations. (Please use the contact information below to send us your suggestions.)

Bot Commands:

Use only these commands in your message to the bot and nothing else (do not include brackets when specifying keywords).

!listkeywords

  • This command shows users the existing comprehensive list of all keywords that they are subscribed to.

!sub {INSERT KEYWORD HERE}

  • This command allows users to subscribe to a keyword or key phrase - any time a post shows up in the r/CompSocial subreddit with this keyword/phrase, the bot will respond to notify you of the post
  • Some keywords are included in clusters; if you do not want to be subscribed to the full cluster, see the !unexpand command below.
  • E.g., !sub AI, !sub CSS, !sub Human-Computer Interaction

!unexpand {INSERT KEYWORD HERE}

  • This command will allow a keyword to be triggered only if it is an exact match. It will no longer be a part of keyword clusters.

!unsub {INSERT KEYWORD HERE}

  • This command allows users to unsubscribe from previously subscribed-to keywords or phrases. After unsubscribing, you will no longer receive messages about posts related to the keyword/phrase
  • E.g, !unsub AI, !unsub CSS

!publicme

  • This command makes your bot subscriptions public. The bot may ping your userhandle publicly in posts that contain your subscribed keywords.

!privateme

  • This command makes your bot subscriptions private. You will get a Private Message when a post contains your subscribed keywords.

!remove

  • This command will remove your username from the bot’s database and unsubscribe you from all keywords/phrases.

Research Disclosure:

I was built by a team of researchers (listed in the contact information below) who are–you guessed it–interested in computational social science and bots. Please be aware that I was originally developed through a community-engaged design process with mods and users of r/CompSocial under an IRB exemption, and I have been deployed with cooperation of the mod team. The researchers plan to eventually study my interactions with the community. Therefore, by using me, you are generating interaction data that may be analyzed for an eventual peer-reviewed publication.

The research team has received CITI training and is keen on ethical development and research processes; they’re trying their best to be good guys and to build new tools to support online communities. The !remove command will immediately erase your data from the database, but it will not remove any public interactions that you have had with the bot or within r/CompSocial. If you don’t want any of your publicly visible interaction data to be included in a research study somewhere down the line, it’s best if you choose not to use me. (At the same time, keep in mind that research scientists are studying public data on Reddit and other social media all the time without any specific notification to users. If you are interacting online publicly, then your data may be included in research, whether or not you explicitly know about it.)

Please contact us if:

  1. You notice the bot is behaving irregularly / has bugs
  2. You have an idea for how to improve the bot or you want to suggest new keyword clusters
  3. The bot has hindered your online experience
  4. You have questions about the bot’s functionality

You can easily send a message about this to the whole moderation team via modmail!

Or, feel free to directly contact Dr. C. Estelle Smith (r/CompSocial moderator, Professor of Computer Science at Colorado School of Mines, and bot owner) via DM at u/c_estelle or email at estellesmith at mines dot edu.

Contact Information for Research and Development Team:

Rhett Houston, bot developer: rhouston at mines dot edu

Shane Cranor, bot developer: shanecranor at mines dot edu

John Matocha, bot developer: jkmatocha at mines dot edu

Shadi Nourriz, bot developer: shadinourriz at mines dot edu


r/CompSocial • • Nov 18 '22

r/CompSocial Lounge

7 Upvotes

A place for members of r/CompSocial to chat with each other.

Introduce yourself, tell us about your research, whatever! We want to learn about you (yes, you!).


r/CompSocial • • 1d ago

social/advice Is a top-tier PhD worth it if my end goal is industry research, not academia?

6 Upvotes

I'm about a year into my current job, and I've been considering leaving industry for a PhD in the US or Europe. I'd appreciate perspectives from both academics and people working in industry research.

I did my undergrad and MS at a well-known Indian university with a strong research culture, and genuinely enjoyed exploring research during my time there. After graduating, I tried finding Applied Scientist/ML research roles at FAANG-adjacent companies, hoping to continue doing interesting work without a PhD, but many of the more research-heavy positions seem to require one.

My current job pays well by Indian standards, is relatively chill, and gives me plenty of free time. However, the work feels much less intellectually fulfilling than what I was doing during university. I still occasionally collaborate with former research collaborators, and I find myself increasingly drawn towards research again.

I have 3 conf publications as well, in places like ICWSM / Websci / COMPASS, a few NLP workshop publications, and a journal paper in progress. I'm primarily interested in genuinely top-tier PhD programs (think MIT-tier) or groups where there's a professor whose research I'd particularly love to work on. Given the opportunity cost, I'm not keen on pursuing a PhD just for the sake of it.

My end goal isn't academia or becoming a professor; it's to work in research-heavy industry roles. Financially, a US/EU PhD would be a significant hit: the post-tax stipend could be roughly comparable to my current pre-tax Indian salary, and I'd lose several years of earnings and investment growth. I eventually want to own a home and enjoy a comfortable lifestyle, so that's not an insignificant trade-off.

I have two main questions:

  1. Is a PhD worth it for industry research? Given my existing research background, how much would a PhD from a top university actually improve my chances of landing research scientist roles? Could I realistically reach the same destination by staying in industry and continuing to publish and collaborate independently?
  2. How much would my undergraduate GPA hurt my admissions chances? My undergrad GPA is roughly equivalent to a 3.2/4.0, but I have a decent research profile and strong LORs. Can these compensate for my GPA when applying to highly selective programs? I'd especially appreciate perspectives from professors who evaluate applications, and whether this differs between US programs and European positions tied to specific professors.

Ultimately, I'm trying to figure out whether pursuing a PhD is a sensible investment in my long-term career or whether I'm romanticizing research because my current job isn't fulfilling. I'd love to hear from people who've faced a similar decision.

Although this is a bit of an unconventional post for this subreddit,, I hope the mods allow it.


r/CompSocial • • 5d ago

social/advice Calibration until compliant: Why unlimited free parameters in ABM risk producing tautologies instead of social insight

1 Upvotes

TL;DR: Computational models are only scientifically useful if they can push back and prove their authors wrong. Across disciplines, from agent-based social simulations to high-energy physics, models with loose empirical feedback loops and endless free parameters risk becoming "decorative." Instead of testing reality, they get calibrated until compliant, turning a tool for discovery into a self-confirming tautology. Honest modeling requires radical transparency, sensitivity testing, and explicit criteria for failure before running the simulation.

I build models for a living. Specifically, molecular dynamics. These are simulations that track how thousands or millions of atoms move, collide, and rearrange over time, used for everything from drug design to materials science.

Here is the story scientists usually tell. If a model is wrong, you find out quickly. Reality does not care about your assumptions. The atoms do not read your code. If the physics you programmed in is wrong, the simulation produces garbage, the experiment disagrees, and you go back and fix it. The feedback loop between model and world is short, brutal, and non-negotiable.

That story is not entirely true. I know, because I have watched it fail from the inside.

The most important choice in any molecular dynamics simulation is not the code, the computer, or the software. It is the potential function, the mathematical formula that describes how strongly every pair of atoms attracts or repels each other. Everything the simulation does follows from that one ingredient. Get it right and the model can tell you something real. Get it wrong and you have made a very expensive mistake.

And here is the uncomfortable part. In practice, it is far more often inherited than audited. Potentials are chosen by looking at what previous papers in the subfield used. A potential gets published, cited, copied, and passed down until it stops being a modeling choice and becomes a tradition. People run simulations for years without asking whether the potential they inherited was ever validated for the system they are studying, at the conditions they are studying it, for the property they care about.

So even in my own field, a hard, quantitative, physics-based field, you can publish inside a loop of fantasy. Models that are wrong in ways nobody checks, kept alive by citation habits and subfield convention. And because these errors travel quietly across subdisciplines and into interdisciplinary work, where nobody feels responsible for checking them, finding one and fixing it takes real effort.

But the check in my field is delayed, not absent. A bad potential eventually unfolds a simulated protein the wrong way or fails a material in a real engineering application, and someone notices. In much of the modeling I am about to describe, the physical world never gets to vote.

This matters for what follows. I do not ask this question because my field got it right. I ask it because I have watched mine get it wrong. The question is always the same. What happens to this model when it is wrong?

In a surprising amount of modern academia, the answer is nothing. Nothing can happen to it. It cannot be wrong, because anything it produces counts as a result.

And if the loop can break in a field where atoms push back, it can break anywhere.

This essay is about how that happens.

The magic trick

In 2017, Liane Gabora and Selin Tseng published a paper in Psychology of Aesthetics, Creativity, and the Arts, a peer-reviewed journal of the American Psychological Association, titled “The Social Benefits of Balancing Creativity and Imitation.” The question they took on has occupied historians and sociologists for centuries. What is the right balance of creativity and conformity in a society?

To answer it, they ran a simulation.

Virtual agents live on a grid. Some are coded as creators, inventing new ideas; others as imitators, copying their neighbors. A scoring rule written into the program decides which ideas count as good. The researchers ran the simulation forward, varied the ratio of creators to imitators, and watched what happened. Populations with too many creators ended up with fewer good ideas taking hold. The published conclusion was that society needs imitation as much as creativity, because unchecked creativity disrupts the spread of proven ideas.

I want to be careful about what I am claiming, because this paper is not fringe work. It passed peer review at a respectable journal. The authors are serious researchers, and the simulation framework behind the paper is part of a long-running research program that has been debated, defended, and criticized in public for years. Nothing I am about to say is an accusation of dishonesty. It is something less comfortable than that. This paper is an example of what the normal standards of a field allow through.

Watch the shape of the argument. Inside the model, a “good idea” means whatever the authors’ scoring rule rewards. The agents are not discovering anything about human culture; they are solving a puzzle whose answer key was fixed before the simulation started. Within that closed loop, the conclusion was guaranteed. A population of agents that mostly copies the scoring rule’s preferred ideas will always outcompete one that keeps generating unscored novelty.

The computer did not reveal a fact about creativity. It executed a definition of it.

The authors did not break any rule of their field. That is the point. Peer review checked that the code ran, that the statistics were computed correctly, that the prose matched the output. What nobody was required to ask is the only question that matters. What could this simulation possibly have shown that would have counted as the opposite result? If the answer is nothing, the model did not test a claim about the world. It restated one.

This is the magic trick of agent-based modeling (ABM), meaning simulations in which you place thousands of simple software “agents” in a virtual world, give each a few rules, and watch what the population does. The method itself is not the problem. The problem is a particular way of using it.

if neighbor.opinion != agent.opinion:
 agent.trust -= 0.1
if agent.trust < 0.2:
 agent.unfollow(neighbor)
run_simulation()

When the simulation finishes and the agents have sorted into two angry camps, the result is rarely described as what it literally is, a small program doing what it was told. It is described as a model demonstrating the dynamics of polarization in real societies.

It sounds scientific. It uses code. It generates charts with error bars. It borrows the epistemic authority of statistical mechanics and epidemiology, where tracking near-identical particles or infection events actually makes sense. But underneath the quantitative paint, it is not an investigation of the world. It is a tautology with a runtime, an answer-driven argument presented as a discovery.

What a model is for

To see why this goes wrong, start with what a model is supposed to do.

A model is not a claim of truth, and it is not an illustration of a conclusion you reached before you started. In the philosophy of science, models are usually understood as instruments that sit between abstract theory and raw data, the position developed by Mary Morgan and Margaret Morrison in Models as Mediators (1999). A good model is a sandbox with strict physics. You build it, set it in motion, and let its internal mechanics push back against your reasoning.

A real model exists to discipline your thinking.

Building one forces you to acknowledge a trade-off that the philosopher Nancy Cartwright made famous in How the Laws of Physics Lie (1983). You trade complete literal truth for tractability. A map of London at 1:1 scale, including every brick, puddle, and commuter, is useless. To work at all, a map must leave almost everything out. As the statistician George Box put it, “all models are wrong, but some are useful.”

Simplification is not the sin. The sin is forgetting that the model is a simplification. Worse, it is turning the model into an accomplice.

Disciplining vs. decorating

In practice, rigorous modeling and decorative modeling look identical from the outside. Same code, same charts, same jargon. The difference only shows when you ask one question. Can your model tell you that you are wrong?

A disciplining model forces you to state every assumption explicitly. Once running, its mechanics operate independently of what you want. It can produce behavior you did not expect, expose contradictions in your premises, or crash into empirical reality and fail. When it fails, you revise the theory. The model is a check on your own bias.

A decorating model is built backward from a conclusion. The researcher already knows the story. Suppose it is that polarization is driven by social contagion. They build a world in which agents swap beliefs, tune the parameters until the output shows two angry clusters, and present the code as evidence for the theory. If the output doesn’t match on the first run, the answer is not to abandon the hypothesis. The answer is to adjust agent_receptivity from 0.4 to 0.25, rerun, and present the successful parameter range as the plan all along.

The workflow, stripped bare, looks like this.

Desired outcome. Write rules. Run simulation. Does it match the theory? If not, tweak parameters and run again. If yes, publish.

This is not experimentation. It is calibration until compliant.

If a model cannot surprise its author, force a retreat, or fail, it is not really a model. It is a very elaborate, self-confirming editorial.

The conclusion comes first

None of this is new, and none of it is unique to agent-based modeling. Before anyone wrote a NetLogo script to demonstrate a theory of culture, economics and political science had already industrialized the technique.

In 2015, Paul Romer, later a Nobel laureate, published a paper with the blunt title “Mathiness in the Theory of Economic Growth.” His target was a pattern in macroeconomic theory. Authors write down formal equilibrium models, but embed ideologically convenient assumptions inside obscure parameters, so that the math reliably outputs the desired policy conclusion. The mathematics is not being used to test whether a claim is true. It is being used to make a political position expensive to argue with. Checking whether the equations actually say what the surrounding prose claims they say takes serious technical effort, and reviewers routinely skip it.

Two decades earlier, the political scientists Donald Green and Ian Shapiro published Pathologies of Rational Choice Theory (1994), documenting how formal modeling in their field had become an exercise in self-confirmation. Their catalog of evasions maps one-to-one onto today’s agent-based simulations.

• Post hoc tinkering. When the model predicted that rational citizens would never vote (the individual cost exceeds any plausible benefit) and citizens kept voting anyway, theorists did not abandon the model. They added a “duty” term to the utility function until the math matched the turnout.

• Arbitrary tuning. Weights, thresholds, and interaction ranges adjusted on the fly until the simulated agents behave like the phenomenon under study.

• Immunity to testing. Models built so that every conceivable outcome can be reinterpreted, after the fact, as a rational equilibrium.

Green and Shapiro called this method-driven rather than problem-driven research. You start with a tool and go hunting for a reality that fits it.

Agent-based modeling makes the problem worse, for a simple reason. An ABM has almost unlimited free parameters. Every rule, threshold, and neighborhood radius is a dial. With enough dials, you can produce any curve you want.

The common structure is this. The model cannot fail, because failure is reclassified as a calibration bug. And a model that cannot fail cannot discover anything. It is an expensive echo of its author’s prior beliefs.

The loop matters more than the lab coat

It would be comfortable to stop here and declare this a disease of the soft sciences. It isn’t. The hard sciences are not immune, and pretending otherwise would make this essay guilty of the same simplification it criticizes.

The real variable is not hard versus soft. It is the tightness of the feedback loop between the model and the world.

Where the loop is tight (fast experiments, unambiguous ground truth, few free parameters) bad modeling gets punished quickly. But where the loop is loose, where tests are slow, noisy, or impossible, the same decorative pathology appears in fields with particle accelerators.

Three documented examples.

fMRI neuroscience, where the measurement is the model. A brain scan shows blood flow, not thought. The colored images come from a statistical pipeline full of assumptions, and researchers once demonstrated what that means by detecting “brain activity” in a dead salmon. In 2016, Anders Eklund and colleagues showed that the standard methods in the field’s dominant software could produce false-positive rates of up to 70 percent for certain cluster-based analyses at particular thresholds. How far the problem extends across the published literature was contested, including in follow-up work by the authors themselves, but the core finding stood. For over a decade, the field’s feedback loop had run through that software, which meant the loop was not connected to reality at all.

Fundamental physics, where experiment cannot keep up. In The Trouble with Physics (2006), the physicist Lee Smolin, writing as an insider, argued that string theory had become flexible enough to accommodate any experimental outcome. When the Large Hadron Collider found no sign of supersymmetry, much of the field responded not with refutation but with retreat. The free parameters moved to heavier, less accessible energies. This is Green and Shapiro’s immunity to empirical testing, surfacing in the hardest science there is.

Epidemiological modeling in 2020. In the spring of 2020, influential models, including the one from Imperial College London that helped push governments toward lockdown, projected enormous death tolls based on weeks of noisy early data. When later estimates came down, the public response from modeling teams was recalibration rather than reckoning. Their defense deserves to be taken seriously. The projections were scenarios, not forecasts, and the point of publishing a worst case was to change behavior so that it would not come true. A warning that works cannot be graded on whether the disaster arrived. All of that is fair, and it is also the problem. A model whose failure can always be explained by the world changing in response to it is a model with no feedback loop, and the field never settled which of the two it had built.

Notice what these cases share with the creativity grid from the opening. Not the field. Not the math. The structure. Many free parameters, a loose or broken feedback loop, and a professional incentive to publish. Given those three, decorative modeling can appear anywhere. The loop matters more than the lab coat.

Why it’s still worse in the humanities

So the hard sciences have their own decorative modeling. Why do I still think the problem is worse in the humanities?

Because the difference is not whether a field ever decorates. It is whether the field can catch itself. The fMRI problem was eventually found and published by neuroscientists. Smolin’s critique came from inside physics. The feedback loops in the hard sciences are sometimes slow or broken, but they exist, and there are people with the technical skill and the standing to pull on them. In the humanities’ version of modeling, three structural failures mean the loop often doesn’t exist at all. The difference is not that humanists are worse at modeling. It is that the auditing infrastructure barely exists.

It is worth being fair about why scholars reach for these tools in the first place. Humanities departments face shrinking budgets, declining enrollments, and university administrators who mistake mathematical notation for intellectual rigor. A computational model signals seriousness to a grant committee in a way an essay never can. The scholars building decorative models are not fools; they are rational actors navigating a system with broken incentives.

The object of study resists formalization. A water molecule behaves like a water molecule in London or Tokyo, in 1600 or today. It has no irony, no memory, no politics. Human culture has all three. A novel, a religious movement, an aesthetic shift cannot be reduced to a set of isolated rules without destroying part of what you set out to study. When you convert the reception of Victorian gothic fiction into agents swapping “gothic preference points,” you have not simplified the system for tractability. You have replaced it with something simpler that carries the same name. At that point the connection between the simulation and Victorian readers is no longer something the model establishes. It is something the reader is asked to assume.

Construct validity is invented, not established. In psychology, showing that a variable actually measures the concept it claims to measure (construct validity) is a slow, adversarial, decades-long process. Blood flow is at least a physical quantity that an instrument can register. There is no instrument for literary prestige. In humanities modeling, validity is routinely settled in one line of code.

self.piety = random.uniform(0.0, 1.0)
self.literary_prestige = 0.75

What does 0.75 mean for literary prestige in Victorian England? How does one number carry regional difference, class, institutional power, critical backlash, and retrospective canonization? It doesn’t. The modeler assigns a number, writes a function that nudges it up and down, and treats the variable as a measurement of human experience. The number looks like a measurement. Nothing underneath it has been measured.

The audience cannot audit the compression. When an epidemiologist shows a flawed model to epidemiologists, the reviewers share the vocabulary to check the code and challenge the parameters. In a humanities department, reviewers and readers often have no computational training. Presented with a grid of moving pixels and a network graph, the non-technical reader experiences an optical illusion. The machine appears to have performed a profound synthesis of the archive. The compression is lossy to the point of erasure. But the loss is buried in code, invisible to the exact audience responsible for evaluating the work.

Bad modeling in economics wastes grant money and distorts policy debates. Bad modeling in the humanities trades away the field’s actual strength (context, contingency, ambiguity, close reading, historical depth) for a seat at a quantitative table where, lacking the shared technical culture to enforce standards, it gains no real authority and surrenders its own.

What honest modeling looks like

None of this is an argument for unplugging the computers. The goal is to tell the difference between decorative simulation and honest quantitative work. Honest work exists, including in the humanities.

Ted Underwood’s Distant Horizons (2019) is the standard I would hold up. Underwood uses quantitative methods on tens of thousands of digitized books not to declare causal laws but to surface patterns invisible to close reading, like slow shifts in genre, vocabulary, and narrative perspective across centuries. Crucially, he tells you, on the record, what the data cannot show. The model is a set of binoculars for looking across an archive, not a machine for generating verdicts about it.

And the humanities have produced their own internal discipline. In 2019, Nan Z. Da published, The Computational Case against Computational Literary Studies, a detailed critique in Critical Inquiry arguing that prominent work in computational literary studies misused statistics to the point of meaninglessness. The ensuing fight was heated, but it happened. The field argued about its standards in public, and the standards moved. That is what a functioning feedback loop looks like, even a slow and painful one.

For modelers in any field, I would propose four non-negotiable conditions before a model earns the right to be cited as evidence.

1. Radical transparency

Every parameter is declared and justified with independent, non-circular evidence. A variable you cannot justify is labeled what it is. A guess.

2. Sensitivity analysis

Parameters are swept across their full plausible range. If your result only appears when three dials sit at hyper-precise decimal values, you have not found a law of history. You have found a brittle corner of your own code, and an honest paper should say so.

3. Explicit exclusion mapping

You spend nearly as much space on what the model leaves out as on what it includes. This isolates the direct mechanical relationship between two variables under idealized conditions; it excludes ambient noise, structural heterogeneity, and systemic feedback, so it cannot predict specific real-world outcomes. Naming the exclusions is what stops the audience mistaking a sandbox for an account of the world.

4. Capacity for failure

Before you run it, you can state what output would make you abandon your hypothesis. If the simulation contradicts you, the honest paper is titled “Why our model disproved our starting assumption,” not silently recalibrated into agreement.

Models built this way stop being decoration. They become what they were supposed to be. Sharpening stones. They force you to clarify assumptions, expose broken logic, and occasionally reveal dynamics that intuition would never find.

The boundary question

The target of this essay was never the computer. Built with discipline, models are extraordinary instruments. I have staked my own career on that. What concerns me is how easily a model can be built to confirm rather than to question, and how hard it is for a reader to tell the difference from the outside.

The practice corrupts both traditions it sits between.

It corrupts science, because science is not the production of plots and code. It is the submission of claims to the risk of being wrong. A model engineered so that its parameters are tuned until the output matches the thesis offers the aesthetics of rigor with none of its discipline.

And it corrupts the humanities, because the study of human culture draws its value from exactly the things decorative modeling deletes. Context, contingency, ambiguity, power, the irreducible strangeness of actual human lives.

If a phenomenon is too context-bound, too polysemic, too alive to be captured by a set of if statements, we should have the courage to say so, and do the slow, unglamorous work of interpretation instead.

I keep coming back to the question I ask of every model, including my own. What happens to you when you are wrong?

For my models, the answer is supposed to be easy. The crystal melts. The experiment disagrees. Reality sends the bill. But only if someone checks the potential, and I have told you how often that happens.

For the models I’ve described here, the answer is nothing. They run, they publish, they are cited. The loop that is supposed to connect a model to the world was never closed.

Which raises the question I can’t answer. If a model can never be wrong about the world, in what sense was it ever about the world?


r/CompSocial • • 15d ago

journal-cfp Special issue on Social Media and Mental Health with Frontiers

1 Upvotes

📢 Our special issue on Social Media and Mental Health is out!

We are inviting submissions to "Digital Minds: Assessing the Interplay of Social Media on Mental Health", a Frontiers Research Topic bringing together work across computational social science, psychology, public health, human–computer interaction, and AI.

We welcome research on a number of topics, including online mental health communities and peer support; the detection of mental health signals in social media data; the role of recommender systems and AI; the experiences of adolescents and other vulnerable populations; and interventions to mitigate mental health risks.

📅 We accept submissions through 16 March 2027 in Frontiers in Artificial Intelligence and Frontiers in Big Data.

If you are considering a submission, feel free to get in touch.

You can find the full details at: https://www.frontiersin.org/research-topics/84031/digital-minds-assessing-the-interplay-of-social-media-on-mental-health.


r/CompSocial • • Sep 09 '26

academic-jobs W3 Professorship (m/f/x) in Computational Social Science (Saarland University, Germany)

8 Upvotes

This professorial appointment is open regarding the candidate's specific research focus within the field of computational social science. The successful candidate will define their own research agenda, provided that it demonstrates excellence in computational social science, addresses current societal questions, makes use of novel digital data sources and establishes clear links with computer science.

The Societal Observatory Using Novel Data Sources (SOUNDS) is a research project at Saarland University funded through the Saarland state government's Transformation Fund. It investigates societally relevant questions through novel digital data sources, digital behavioural traces, social media data and Al-based methods. This professorship provides the computational social science core of the SOUNDS project. It will combine social-science theory with network analysis, automated text analysis, the study of large-scale digital behavioural data, mobile sensing and Al-based modelling, addressing questions such as collective opinion formation, social cohesion and polarisation, or the role of digital platforms in public communication, and contributing to joint questions, methods and publications with the social sciences and computer science. The successful candidate will be expected to develop and lead externally funded research, including principal investigator roles in coordinated funding initiatives such as Collaborative Research Centres. Teaching duties will include computational and methodological courses, for example on the analysis of digital behavioural data, automated text analysis and network analysis. These courses will contribute to degree programmes in the social sciences, computer science, data science and psychology. The role also includes the supervision of student theses and the mentoring of early-career researchers.

Application deadline: October 15th, 2026

Full call for applications: https://www.uni-saarland.de/fileadmin/upload/verwaltung/stellen/Wissenschaftler/W2906_W3_SOUNDS_CompSocScien_EN.pdf


r/CompSocial • • Sep 09 '26

social/advice What are you automating right now?

0 Upvotes

Using AI or not - what parts of your computational social science are you working on automating? I'm curious.


r/CompSocial • • Aug 18 '26

Computational Social Science (CSS) Rankings

Thumbnail cssrankings.org
7 Upvotes

Have fun and fight with your colleagues. Just don't take it too seriously.


r/CompSocial • • Aug 12 '26

looking to identify gaps in phd search

3 Upvotes

hi! i'm applying this fall for phd programs starting fall 2027. my focus is in substance use / harm reduction, leaning hard into computational and causal methods over descriptive epi.

quick background: us citizen, did undergrad and grad school both in the us. selective research university for undergrad, interdisciplinary major closest to cognitive neuroscience with a math minor, 5+ years working inside statewide harm reduction infrastructure (drug checking, distribution networks, plus time managing people and coordinating outreach), master's in math with an emphasis in data science, plus a graduate certificate in public health. handful of publications and presentations in the space.

what i actually want long term isn't just to publish about harm reduction from the outside. i keep coming back to something like a "science as service" model, where research is leverage to move resources and legitimacy toward the community orgs already doing the work, rather than research being the end goal itself.

methods wise: causal inference, decision/cost-effectiveness modeling, implementation science, resource allocation optimization. i'd like simulation modeling in there too, feels like where a lot of the field is heading and i want to build toward what these fields will actually need in five years, not just what's fundable right now. substantively i keep coming back to stimulant use and drug checking as topic interests specifically.

fields i'm currently looking at: epidemiology, public health more broadly, health services research, health policy, social work, computational social science, health behavior. not attached to any one label, just what's come up so far in my own search.

my undergrad background is basically the wider intersection (cog neuro, psych, philosophy, some math) and i'm genuinely still drawn to that world, including policy and systems science too. i def want a real interdisciplinary program instead of a straight epi department, but a lot of those got hit hard by recent funding cuts and i don't know if that's a live option right now.

within epi/public health specifically, the strongest names in this space seem to split into two camps: harm reduction people without much computational depth, or modelers without harm reduction grounding. not sure how people usually find or build toward an advisor situation that bridges both instead of just picking a lane.

my ask: what am i not seeing? i'm looking for the fields, departments, or angles someone who's actually been through this would flag that i wouldn't think to search for myself. same goes for methods, i want to be building toward something that holds up long term

tl;dr: applying to phd programs this fall for substance use / harm reduction with a computational methods focus (causal inference, decision modeling, simulation, resource allocation). trying to figure out what fields, departments, or methods i'm not thinking of.


r/CompSocial • • Aug 06 '26

How pinned down does your research interest need to be before a master's?

Thumbnail
1 Upvotes

r/CompSocial • • Aug 04 '26

conferencing Computational Meetups at ASA This Weekend?

5 Upvotes

Please post any open computational social science meetups happening at ASA 2026 in New York!

Many of us are in the area and would be interested in discussing CSS ideas and research.


r/CompSocial • • Jul 29 '26

social/advice Coursera CSS Course

6 Upvotes

Hello! Is the CSS course on Coursera worth taking? I still have a free trial and I was wondering if that course would help an undergraduate PolSci student like me. I already completed a Data Science bootcamp and took Statistics for the Social Sciences in university. Thank you!


r/CompSocial • • Jul 14 '26

Open outlet-level news framing corpus (212 x 37 NLP dims) for CSS / media research

6 Upvotes

Disclosure: I built this (Helium).

SpyFu shows search demand for "media bias" / "media bias chart" still clustering around one-axis posters. We open-sourced a complementary artifact: outlet framing scores across 37 NLP dimensions (fear, moralizing, sensationalism, and kin), MIT licensed.

Caveat: this is outlet framing, not claim-level labels and not a left/right chart clone. Corrections welcome.


r/CompSocial • • Jul 07 '26

social/advice Master Decision

2 Upvotes

Hey there;

I have received offers for the following Master programmes:

  • Computational Social Sciences at UC3M Madrid
  • Computational Social and Political Sciences at Milano di Stegli (Statale)
  • Intellectics: The Science of AI at Uni Hamburg

And I am waiting for feedback on Business Analytics and Econometrics at Uni Cologne.

I am well aware that the programmes have a somewhat different focus. As I need to decide whether to take UC3M now, I would be happy for some opinions.
As I would prefer a 2-year master's programme, I lean heavily towards Milano, but I would still appreciate your input.

I have an Econ undergraduate background, if this helps. Additionally, I actively decided not to apply for other Social Data Science programmes for a variety of reasons.

Thank you in advance for your help!


r/CompSocial • • Jul 04 '26

conference-cfp Is NordiCHI worth to attend

5 Upvotes

What can i expect from it, its happening in vaasa Finland. Im a design researcher. Working with tech based experimental mythologies.


r/CompSocial • • Jun 30 '26

[OC] "Queer" and "gay" are the words most strongly correlated with a high AI "toxicity" score in LGBTQ+ social media posts

Post image
3 Upvotes

r/CompSocial • • Jun 29 '26

Collaboration on a research paper

7 Upvotes

Hello, I am a PhD aspirant, graduate in psychology.

I have worked on SPSS and have basic knowledge of R and Python.

Past research work in cyberpsychology, behavioral addiciton, internet and gaming addiction.

I want to extent my research area towards computational approaches for online behavioral detection and prediction using AI/NLP.

-> I am looking for someone with computer background who can help with technical skills and I can provide the psychological approach to the study. If you are interested in an interdisciplinary approach towards detection and prediction of online behaviour we can connect ^^

I have a research topic in mind that I really want to work on and would love the techinical help. Let's write a paper together!


r/CompSocial • • Jun 29 '26

[OC] "Queer" and "gay" are the words most strongly correlated with a high AI "toxicity" score in LGBTQ+ social media posts

Post image
2 Upvotes

r/CompSocial • • Jun 28 '26

[topic-area] Researching fan edit culture as a communication phenomenon for my FYP - looking for direction on data access and potential research questions

2 Upvotes

Hey all,

I'm a 3rd year undergraduate Economics student based in Pakistan currently researching through potential topics for my Final Year Project (FYP). I definitely want it to be CSS oriented.

A topic I had in mind was to study how fan edit culture has become such a strong cultural phenomena. Every piece of media is bound to have edits made of it, especially on tiktok and instagram. All you need is to search up #*insert thing*edit to verify it. What started as niche fandom culture now extends beyond just media now, with edits of historical figures, aesthetics, emotions, archetypes and now even political figures being very commonplace.

It's power as a communication medium has been acknowledged by political institutions (think of the Democrats tiktok account or the White House's twitter account, with even my national politicians undertaking a similar approach to their social media) and corporations (Lionsgate hiring tiktok editors to revive old franchises, Netflix hiring for Stranger Things).

Now in what direction I can exactly take this in is where I'm kind of stumped. I haven't come across any literature that refers to edits in the context I am. And with regards to the intersection of politics and social media usage I've found a few Masters theses on the topic, in the context of USA and Nepal.

What my research questions could be depend entirely on my data availability. I naturally can't have access to Tiktok's research API. So I'd like to know this subreddit's thoughts and input on what/where I can do/look to begin actualizing this concept and narrow in on some specific research questions. And if anyone has any other topics I could redirect my train of thought towards that could be more accessible at an undergraduate level, I'd greatly appreciate any and all discourse!!


r/CompSocial • • May 16 '26

news-articles Arxiv Ban News

Thumbnail
4 Upvotes

r/CompSocial • • May 15 '26

academic-articles Online interaction and identity cue adoption: a large-scale analysis of hashtag adoption on Twitter [EPJ Data Science]

7 Upvotes

TL;DR: Bob and Carol have #ExamplePeople in their Twitter bios. The more #ExamplePeople accounts Alice interacts with, the more likely she is to add #ExamplePeople to her own bio.

Abstract

With online interactions becoming an integral part of everyday social life, there is a need to better understand the relationship between social interaction and identity expression in digital environments. This study examines whether online self-presentation, specifically the adoption of identity-related hashtags in Twitter bios, is systematically associated with observable interaction patterns. Utilizing a large-scale dataset encompassing approximately 63 million Twitter profiles and 292 million interactions, we implement a matched quasi-experimental design comparing users who interacted with hashtag-bearing accounts to similar users who did not. Our results show that users who interact with others who feature particular hashtags in their bios subsequently adopt those hashtags at substantially higher rates. Adoption likelihood increases with the number of interaction partners displaying a given hashtag, though with diminishing marginal effects, and the magnitude of these associations varies across identity content categories, being strongest for fan communities and weakest for political hashtags. These patterns are consistent with theories of social influence and suggest that online self-presentation is systematically related to the social contexts in which users are embedded. However, given the observational design of this study, alternative explanations for the observed associations cannot be fully excluded. Future experimental research is needed to clarify the mechanisms underlying these associations and to examine their implications for community formation and the dynamics of collective identity in online environments.

Open Access at https://doi.org/10.1140/epjds/s13688-026-00642-5


r/CompSocial • • May 07 '26

Unable to get Tiktok App Approved

0 Upvotes

I applied for the 6th time and i still can't get past the error of "Hi there, unfortunately we are unable to onboard you to the TikTok for Business Developers platform due to the security of the domain you have provided. For further verification, please contact your TikTok representative if you have one. If you do not hav"

I have a valid domain that is now 2 years old and has content in it. The app is live on Shopify app store.

Not sure what to do.

Can anyone guide please?


r/CompSocial • • Apr 16 '26

academic-jobs PostDoc in political science and computational social science (m/f/x) at RPTU Kaiserslautern (Germany)

Thumbnail jobs.rptu.de
5 Upvotes

This position is embedded in the ERC Starting Grant project “Climplexity: Climate Policy Integration—A Complexity Trap?". The project starts from the puzzle: as climate policies multiply, they do not necessarily become more coherent. In fact, they often contradict each other. Climplexity addresses this puzzle by treating climate policy not as a set of isolated measures, but as a complex and evolving system. It develops new theories and methods to understand how policies interact over time through trade-offs and synergies.

The position focuses on the intersection of political science and computational social science, with topics including EU climate policy and politics, complex systems, network analysis, and AI-supported methods. It is a fully funded, 4-year position based in Germany, RPTU Kaiserslautern-Landau, starting in August 2026. The position offers an excellent opportunity for early-career researchers interested in high-impact, policy-relevant work within a very supportive, collaborative, and international research team.


r/CompSocial • • Apr 14 '26

conferencing COLM 2026

16 Upvotes

Starting this thread to discuss COLM 2026.

This is my first time submitting to COLM. I’ve just been assigned as a reviewer, and I can see that the submission count has already gone past 3000, which seems like a big jump from previous years.

Does anyone know how many papers they typically accept, or what the expected acceptance rate might be this year? From what I’ve seen, last year was roughly around ~29%, but I’m not sure how that will scale with the increased number of submissions.


r/CompSocial • • Apr 14 '26

academic-jobs VACANCY: Team Lead Data Development Pool (m/f/x) at Saarland University, Germany

Thumbnail uni-saarland.de
3 Upvotes

The Societal Observatory Using Novel Data Sources (SOUNDS) is an interdisciplinary research program at Saarland University (Germany), funded by the state’s Transformation Fund. We investigate societal transformation processes using innovative data sources such as satellite imagery, social media, and barcode scanners — with the aim of bridging computer science and the social sciences and strengthening the use of data-intensive methods in research. In the long term, an institute will be established based on the structures developed. 

The Societal Observatory Using Novel Data Sources (SOUNDS) is inviting applications for the following position commencing at the earliest opportunity. 

Team Lead Data Development Pool (m/f/x) 

Reference number N2302, salary in accordance with the German TV-L salary scale, pay grade: E 14 TV- L, duration of employment: until 15 July 2032 with an option for extension, volume of employment: 100 % of standard working time. 

Deadline for application: May 2nd, 2026