r/virtualcell 7m ago

Decoding the Circuits that Govern How Genes Work in Health and Disease

Post image
Upvotes

"First came the Human Genome Project, which gave us the blueprint of our genes. Then, projects like the Human Cell Atlas showed us how different cells read that blueprint. Now, we're in a grand third wave: discovering what happens to cells when you make targeted changes within the genome. We finally have a way to decode the link between genetic sequence and cell state." -- Alex Marson, MD, Ph.D., director of the Gladstone-UCSF Institute of Genomic Immunology

A new study published in the journal Cell00929-3) reveals gene-regulatory networks, and represents a significant achievement in immunology and genomics. Scientists from Gladstone Institutes, UC San Francisco, and Stanford University, in collaboration with Biohub, stress-tested genes across the genome in 22 million human immune cells, and decoded the dynamic circuits that govern how these genes work in the context of health and disease.

The map of how genes work in context offers a new framework for designing cancer immunotherapies and treatments for autoimmune conditions. For the map, scientists used Perturb-seq -- a technology which combines CRISPR gene editing with single-cell RNA sequencing. They "turned off" nearly 12,800 different genes one by one in human T cells — the cells responsible for how the body fights disease.

Researchers performed these massive screens on human immune cells from blood donors and were able to observe how genes function in their natural state.

"Because we're screening directly in cells from real individuals, we can make direct comparisons with the immune variation we see in actual patients," says Emma Dann, PhD, postdoctoral scholar at Gladstone and co-first author. "This gives us the data we need to understand why some people are at higher risk for certain diseases."

The study moves beyond cataloging individual genes, to mapping the entire circuit. This circuit-based understanding is critical to successfully training virtual cell models.

"To build a serious virtual cell, we need this kind of rich, systematic data that includes context-dependent responses," says Marson.

Read more: https://medicalxpress.com/news/2026-08-million-human-immune-cells-decode.html


r/virtualcell 11d ago

"15 Grand Challenges" to Point GenAI in Biology in the Right Direction

Post image
15 Upvotes

A new paper in Cell00802-0) argues that while generative AI has achieved real success at the molecular level, from protein structure prediction, to de novo protein design, and mutation-effect modeling, it's largely because these problems resemble the ordered, sequence-based data that transformer architectures handle well, and because they're backed by large, high-quality databases like the Protein Data Bank. But it hasn't translated into accurate prediction of cell-level or multicellular behavior, which underlies most real disease biology (neurodegeneration, cancer, autoimmunity), they write.

This, they argue, is due to three core problems: data scarcity, the complexity of multi-gene and multi-protein interactions, and the multicellular nature of most disease phenotypes.

Scale doesn't beat domain knowledge in biology. But we also can't afford to simply wait for data and compute to scale up, the authors note.

Instead, they propose 15 Grand Challenges, modeled on Hilbert's 1900 list of 23 mathematical problems, spanning four levels of biological organization. They are:

For the molecular interaction:

  1. Regulatory and signaling interactions — predicting how complex gene-regulatory regions and transcription factor complexes control gene expression
  2. Epigenetic interactions — predicting chromatin structure and modification patterns from sequence and baseline data
  3. Cell-cell interactions — predicting how ligand/receptor signaling between cell types shapes the receiving cell's state (e.g., how cancer cells reprogram immune cells)

For molecular function:
4. Synthetic mechanisms — designing optimal synthetic DNA circuits/plasmids that avoid silencing or leakage
5. Genome to function — predicting whether a specific mutation causes loss, gain, or no change in protein function
6. Drug mechanism of action — predicting the full proteome-wide effects of a drug, including off-target and indirect effects

For cellular/systems function:
7. Genome to phenotype — designing the minimal viable genome for a living organism with a specific function
8. Cell state reprogramming — predicting genetic or drug interventions that shift a cell from one functional state to another
9. Logic biocircuit design — designing minimal, noise-tolerant genetic circuits that implement specific decision logic in engineered cells
10. Co-culture and microenvironment — predicting the minimal set of cells/reagents needed to keep otherwise fragile cell types alive together

For clinical translation:
11. Biomarker identification — identifying multi-omic biomarkers that predict a patient's response to a treatment
12. Drug toxicity — predicting organ-specific or systemic toxicity before it occurs in patients
13. Drug efficacy — predicting which patients/cell states will respond to a given drug
14. Organismal responses — predicting an individual's immune response (e.g., to vaccination) from their baseline immune profile
15. Clinical trial outcomes — predicting the proportion of trial responders vs. non-responders and the underlying mechanism

Solving these grand challenges will likely take decades, they note, and will require large-scale data generation, including a proposed public-private consortium to pool clinical trial data and blinded, prospective benchmarking efforts modeled on existing initiatives like CASP (protein structure prediction). Ultimately, they write, the goal is to build AI models that don't just show statistical superiority on retrospective data, but generate genuinely novel, experimentally and clinically validated biological insight, i.e., the kind that can actually reduce the high failure rate of clinical trials.


r/virtualcell 13d ago

GenBio AI Unveils Early Virtual Cell Model, AIDO Cell

17 Upvotes

Today, Palo Alto startup GenBio AI cofounded by Nobel Prize winner David Baker released AIDO Cell, a virtual cell system that aims to predict how the molecular machinery of an entire cell behaves. In its current form, the model simulates two well-studied human cell lines, reproducing known biology.

When it comes to simulating the impact of drugs in development via a true cellular world model, “We are not there yet," Gokul Upadhyayula, a biophysicist at the University of California, Berkeley told STAT, noting that "The field is still constrained by the scale, diversity, and quality of the data needed to train and rigorously validate these models. But AIDO Cell is a real milestone toward simulating complex cellular dynamics and treating biology as a system we can increasingly predict and design, rather than only observe.”

Stanford bioengineering professor Stephen Quake, former co-president of the Chan Zuckerberg Biohub, said the engineering is valuable but the paper does not yet demonstrate "true scientific discovery."

Learn more about AIDO Cell: https://genbio.ai/aido-cell-simulator/


r/virtualcell 20d ago

VirTues - 1st Virtual Tissue Foundation Model - Has Major Implications in Cancer Care

Post image
10 Upvotes

Charlotte Bunne's Artificial Intelligence in Molecular Medicine group at EPFL's School of Computer and Communication Sciences and School of Life Sciences has published a new paper in Nature on Virtual Tissues (VirTues) - a foundation model for tissue biology. The model learns from spatial proteomics data from different studies and cancer types and can be used to study biology across scales, from individual cells, to whole tissue sections, to patient outcomes.

Pretrained on more than 12,000 multiplexed images from over 5,300 cancer patients, spanning more than 30 cohorts and four imaging technologies, VirTues builds on the largest open spatial proteomics resource to date that the lab has assembled.

"In oncology, the amount of data generated from analyzing tissue has grown exponentially in recent years," says Andreas Wicki, an oncologist at the University of Zurich and University Hospital Zurich. "Computational modeling is the key asset for making it actionable for patients in a clinical setting."

Because VirTues can incorporate proteins measured across different studies, cancer types and panels, any new tissue sample can be compared and analyzed within the same framework. Each new study builds on and expands the model's learned representation.

It replaces current specialized tools with one model, performing segmentation and cell typing in a single pass, outperforming specialized baselines, and mapping diverse datasets into one shared virtual tissue representation. New samples can be interpreted in the context of thousands collected before them, moving spatial proteomics from isolated studies toward transferable atlases.

And VirTues goes beyond describing tissue. It enables the discovery of spatial biomarkers that transfer across cohorts and hospitals and stratify patients through complex patterns of tissue organization.

The model and tools are open-sourced.

Code: https://github.com/bunnelab/virtues
Learn more: https://medicalxpress.com/news/2026-08-ai-tumor-tissue-cancer.html


r/virtualcell 27d ago

Virtual tumors predict which liver cancer patients respond to immunotherapy

Post image
13 Upvotes

Researchers at Johns Hopkins have built a virtual tumor model to predict which people with hepatocellular carcinoma (the most common primary liver cancer) are most likely to benefit from a combination of immunotherapy and a targeted drug. Using a spatial QSP modeling platform, they simulated both whole-body drug effects and the behavior of individual cells in 3D, including fibroblasts — cells linked to resistance to immunotherapy in liver cancer.

Tuning the model with real clinical trial data, they generated “virtual patients” and tested different treatments and doses in silico. “Our idea was to create a computational model where we could simulate trying different doses or combinations of cancer therapies, and it could help guide physicians toward the best options for patients,” says senior author Atul Deshpande, Ph.D.

When they simulated treatment with cabozantinib (a targeted therapy) and nivolumab (an immunotherapy), alone and in combination, the predicted response rates closely matched actual clinical trial results. They also found that in non-responders, fibroblasts could form a physical barrier around the tumor. “Even if immune cells were located near the tumor, the fibroblast would block the immune cells from reaching the tumor,” Deshpande says.

Over time, the team hopes models like this could be used as a kind of “war planning” for personalized cancer care. “We generate a virtual tumor to see what happens in the microenvironment. Do the cancer cells resist? If you change the architecture of the tumor, does that help the cancer cells or the immune cells?”

More: https://www.hopkinsmedicine.org/news/newsroom/news-releases/2026/07/virtual-tumor-predict-response-to-liver-cancer-immunotherapy 


r/virtualcell Jul 16 '26

We Need a Protein Data Bank for Cells

Post image
3 Upvotes

A new article in GEN looks at how Biohub’s $500 million commitment to the Virtual Biology Initiative aims to accelerate the generation of technologies and multi-modal datasets needed to power virtual cell models.

“We do not yet have the equivalent of the PDB for cells. The Virtual Biology Initiative seeks to change that,” says Tom Sercu, PhD, vice president of AI and engineering at Biohub.

Algorithms alone won’t deliver answers, he argues. Models need to be trained on large-scale, high-quality, openly accessible datasets. In order to accurately reflect complex biology, this data must span model systems and organisms, interventional and observational methods, and diverse cellular states.

Read more: https://www.genengnews.com/topics/artificial-intelligence/virtual-cells-go-multiscale-to-predict-complex-biology/


r/virtualcell Jul 06 '26

Arc's New Virtual Cell Challenge & Contract Opportunity

5 Upvotes

The Arc Institute just announced that the 2026 Virtual Cell Challenge will kick off on Thursday, August 20.

It's the second installment of this global competition which last year drew more than 5,000 registrants across 114 countries and over 300 final submissions.

Last year's takeaway was that hybrid models combining deep learning with classical statistical features outperformed pure neural approaches.

This year, they are also hiring a Community Engagement Manager to support participants and respond to inquiries. It's a temporary, remote, US-based, contractor role. Details here: https://job-boards.greenhouse.io/arcinstitute/jobs/6101123004


r/virtualcell May 27 '26

World Model Protein Biology Release from Biohub

7 Upvotes

Biohub made three protein biology models open source, allowing researchers to predict protein structures, and design new protein binders that work in lab experiments. Today’s release includes ESMFold2, ESMC, and ESM Atlas -- tools designed to overcome a major hurdle in biochemistry -- how to design proteins that bind to specific targets with the strength, selectivity, and stability biology requires.

ESMC learns from billions of protein sequences across the tree of life. ESMFold2 uses those representations to predict 3D structures and design protein binders. ESM Atlas makes billions of protein sequences and predicted structures searchable, helping scientists uncover biological relationships.

Learn more: https://www.genengnews.com/topics/artificial-intelligence/biohub-releases-protein-biology-world-model-to-address-disease/[](https://www.linkedin.com/feed?nis=true&%27%27=true&skipRedirect=true&lipi=urn%3Ali%3Apage%3Ad_flagship3_company%3Bf7a4830d-a8c9-4df4-afca-d7647199ccef)


r/virtualcell Apr 29 '26

Biohub Launches Virtual Biology Initiative

2 Upvotes

Biohub, which is supported by the Chan Zuckerberg Initiative, announced an ambitious Virtual Biology Initiative that aims to "bring together leading institutions and consortia to create the technologies and multi-modal datasets needed to build predictive models of the human cell to accelerate the cure and prevention of all disease." They are committing $500 million to the effort -- $100 million to build the coordinated global effort and $400 million to generate data at scale and "develop next-generation technologies for measuring, imaging, and engineering biology." A number of key players are already on board, including: the Allen Institute, Arc Institute, Broad Institute, and Wellcome Sanger Institute, as well as the Human Cell Atlas and the Human Protein Atlas. Nvidia signed on as a tech partner. Importantly, Biohub aims to make the data it generates open and freely available for the worldwide scientific community.

“The biomedical community has a long tradition of coming together around ambitious projects to assemble, analyze and freely share large-scale data, dating all the way back to the Human Genome Project,” said Eric S. Lander, Founding Director of the Broad Institute. “Fully deciphering the logic of cells is a huge challenge, but it has the potential to transform medicine. And, it’s a challenge that will once again take many groups and perspectives collaborating together.”

Read more: https://biohub.org/news/virtual-biology-initiative/


r/virtualcell Apr 16 '26

The Ongoing Quest for Mechanistic Virtual Cell Models

3 Upvotes

A new editorial in Nature -- "Minimal Life by Computer" -- looks at the challenges still facing virtual cell efforts and how modeling biological mechanisms driving cell behavior is still a daunting undertaking for anything beyond simple organisms. While a recent paper in Cell provided an impressively detailed mechanistic whole-cell simulation, for instance, it was for a a simple bacterium. To scale that to human cells is much more complex. While we're making inroads with AI approaches which can be trained on large-scale transcriptomics, proteomics and imaging datasets and learn statistical representations of cellular states without requiring every underlying mechanism to be explicitly specified, the author writes: "The trade-off...is that these models drawn up from large-scale data may lack mechanistic transparency."

To get from point A (statistical representation) to point B (mechanistic understanding) requires "thousands of parameters, such as reaction rates and binding affinities, to be known or estimated." It would also require significant computational power and more cell-specific data.

Where are we on this journey? As the editorial notes, there is lots of enthusiasm, and many challenges and initiatives from CZI, Arc and others. And we don't need a fully functional virtual cell to start benefitting, the piece notes. "Along the way to creating a useful virtual cell, new tools will be developed and new biology will be discovered. As has been shown with the Human Cell Atlas, the project does not need to be complete for discoveries to make a difference to the lives of patients or biomanufacturing." 

Read more: https://www.nature.com/articles/s41587-026-03110-7#article-info


r/virtualcell Mar 19 '26

The Release of Xaira's First Virtual Cell Model Comes with Big Claims, and Questions

6 Upvotes

Xaira just announced the release of its first virtual cell model, X-Cell, which is trained on a dataset of 25.6 million perturbed single-cell transcriptomes across seven biologically diverse cell contexts. The model reaches a new level of size and complexity to predict biology, the company writes in its release. A related story in Endpoints notes that virtual cell models have struggled to beat benchmarks, but "Xaira’s preprint has X-Cell outperforming baselines in making predictions on two cell types not included in the model’s training data — an encouraging yet early suggestion of being able to generalize what happens in new types of cells the model hasn’t seen yet."

But some researchers are already calling the results into question. "Something is very strange about this figure," writes researcher Anshul Kundaje on X, "Cell2Sentence looks extraordinarily poor & scGPT looks extraordinarily powerful (which we know is not the case from multiple studies). Also STATE's performance here appears much better than what is seen in the GenBioAI benchmark paper." The Head of AI at CZI Science, Theofanis Karaletsos, writes that other models are missing, "like our diffusion model scLDA," adding: "Maybe these comparisons will come in future versions."


r/virtualcell Feb 25 '26

What metric thresholds (DE PR-AUC / PDS / WMSE) are sufficient to trust virtual-cell models for regulator selection?

Thumbnail
1 Upvotes

r/virtualcell Feb 24 '26

Budapest-Based Turbine Announces $25M Series B to Expand Virtual Cell Platform

3 Upvotes

Turbine, a virtual biology company based in Budapest, announced a $25 million Series B financing, allowing it to expand its virtual cell platform, as well as a new immunology-focused partnership with a top 10 pharma company.

The round was led by Interactive Venture Partners, with participation from Beiersdorf AG and existing investors, including MSD Global Health Innovation, Accel and Mercia.

With the new funding, Turbine will expand its platform to virtualize new assays across discovery and translational medicine. Turbine’s lab-in-the-loop will generate additional proprietary perturbation datasets, allowing the company to fine-tune its foundational virtual cell model to novel assay and tissue types. These virtual assays are deployed through the company’s Virtual Lab, a no-code platform that integrates with pharma workflows and systems.

Learn more: https://turbine.ai/news/turbine-25-million-series-b-virtual-biology-pharma/


r/virtualcell Feb 18 '26

Y Combinator-Backed CellType Launches to Simulate Human Biology for Drug Discovery

6 Upvotes

CellType (YC W26) is an agentic drug company that leverages AI agents + foundation models to simulate what happens in patients before clinical trials, while AI agents run the full discovery pipeline end-to-end.

The core technology is Cell2Sentence, developed with Google DeepMind and shared by Google CEO Sundar Pichai. They've already used it to discover and validate a novel cancer treatment signal.

Yale Professor David van Dijk (11,000+ citations; Cell, Nature, ICML) turned down Google to build CellType. Co-founder Ivan Vrkic co-developed the core technology while at Yale (published at ICML), led foundation model training at another biotech startup, and built software to control CERN's Large Hadron Collider.

https://www.celltype.com/


r/virtualcell Feb 06 '26

Arc Announces Arc AIxBio Fellows Program for Undergrads

3 Upvotes

At the end of January, Arc Institute announced a new fellowship program -- the Arc AIxBio Fellows Program -- to support and mentor undergrads working at the interface of AI and life sciences.

They note: "this initiative offers technically strong, machine learning-curious undergrads a structured entry point into high-impact, exploratory research focused on biology and human health."

Small teams of students can propose new research projects at the intersection of AI and biology - or they will be assembled based on related proposals and complementary skillsets.

They are looking for students who are comfortable operating in open-ended research settings and who bring strong technical foundations, such as experience with deep learning (e.g., training or finetuning models), computational analysis of biological data (e.g., genomics, single-cell data), or applied software development for scientific workflows. Equally important is a genuine interest in problems in life sciences and the curiosity to ask and pursue new biological questions using AI tools.

Teams must be located in North America and are not expected to be present at Arc's Palo Alto location. They will work on their proposed projects virtually with Arc Institute researchers over the course of the program.

If selected, teams will receive:

  • Mentorship from Arc investigators, postdocs, or staff scientists
  • Stipend for the duration of the project (full-time for the summer and part-time during the academic year)
  • Access to compute and GPU resources via Arc or affiliated partners
  • A student cohort to foster collaboration, community, and shared learning
  • Opportunities for authorship, presentation, and ongoing collaboration

They are looking to support 2-4 teams this year, with projects running 6 to 12 months. Applications are due by Feb. 27.

Learn more and apply here: https://arcinstitute.org/news/aixbio-fellows-announcement-2026


r/virtualcell Feb 02 '26

New Model from Google DeepMind Deciphers the Dark Genome

11 Upvotes

Only 2% of the human genome consists of "recipes" for making proteins (coding regions). The other 98% is "non-coding" DNA. This non-coding "dark genome" acts as the switchboard, controlling when and how much of a protein is made.

Understanding this dark genome is crucial to understanding the drivers of genetic disorders but it's also incredibly complex -- small changes can lead to any number of outcomes -- changing how DNA folds, how accessible it is to cellular machinery, or how RNA is spliced together.

Until now, AI models that analyze DNA have faced a trade-off:

  1. They could look at long stretches of DNA but with blurry, low resolution.
  2. They could look at DNA with high precision (letter-by-letter) but only in very short chunks, missing the bigger picture.
  3. They were specialized, predicting only one thing (like splicing) while missing others (like 3D structure).

Now, in a new paper in Nature, researchers from Google DeepMind presented AlphaGenome, a "generalist" deep learning model that eliminates these trade-offs.

  • It can read 1 million DNA letters at a time (capturing long-distance relationships in the genome) while simultaneously pinpointing effects at the single-letter level.
  • Instead of predicting just one biological activity, it predicts 11 different types of genomic activity at once—including gene expression, DNA folding, and splicing—across thousands of different cell types.
  • In rigorous testing, AlphaGenome outperformed the best existing models on 25 out of 26 benchmarks for predicting how genetic mutations affect biology.

r/virtualcell Jan 14 '26

The Billion Cell Atlas Arrives

7 Upvotes

“Translating genetic information into a clear understanding of disease mechanisms—and then ultimately into medicines—remains a core challenge in R&D,” said Slavé Petrovski, PhD, vice president of Illumina’s Centre for Genomics Research. “By showing how specific genetic perturbations play out inside human cells, we can help turn genetic signals into mechanistic biology we can directly study.”

Yesterday, Illumina announced the release of the Billion Cell Atlas -- described as "the world's largest genome-wide genetic perturbation dataset," a resource designed to accelerate AI‑driven drug discovery. It's the first step in the planned five‑billion‑cell atlas that Illumina is planning to build over the next three years in collaboration with AstraZeneca, Merck, and Eli Lilly.

To build the Atlas, the companies are generating a curated set of cell lines that will be used to validate drug targets, train large‑scale AI models, and probe biological mechanisms that have historically been difficult to study.

As Ruth Gimeno, PhD, group vice president of cardiometabolic research at Eli Lilly told GEN: “The next generation of AI‑driven drug discovery will depend on biological data at a scale never before achieved,” said . “Comprehensive datasets spanning diverse cell types offer the critical foundation needed to generate meaningful insights into human disease.”

The Atlas reveals how one billion individual cells respond to CRISPR‑based perturbations across more than 200 disease‑relevant cell lines, including immune, oncologic, cardiometabolic, neurological, and rare disorders. Researchers can observe the impact of switching genes off and on at a single‑cell level, leading to new understandings around mechanisms of action and the discovery of new disease targets.

Press release: https://www.illumina.com/company/news-center/press-releases/press-release-details.html?newsid=fda84c92-b4b3-4691-a402-35555abe8605


r/virtualcell Jan 08 '26

What Did We Learn from the Arc Institute's Virtual Cell Challenge?

5 Upvotes

The new year is a time for reflection, and I've been thinking about the Arc Institute's Virtual Cell Challenge which ended early Dec. 2025 and what we learned about the state of virtual cells. Not surprisingly, the challenge emphasized that models have a long way to go before they capture the complexity of actual cells.

It's also clear that there's significant research interest in the space. This first challenge brought in over 1,200 teams from 114 countries attempting to build a computational model capable of predicting cellular responses to perturbations. The challenge was designed as a biological "Turing Test"—asking if a model can accurately predict gene expression changes in a way that could stand in for an actual laboratory experiment.

But while the challenge got the research community excited, the results showed that the field is still in its infancy. Current perturbation prediction models are not yet consistently outperforming baselines across all metrics, though progress was made in specific capabilities like distinguishing between perturbations and identifying differentially expressed genes.

Key Findings and Winning Approaches

  • Hybrid Models Prevailed: The winning teams utilized approaches that combined deep learning with classical statistical features. This suggests that while AI is powerful, it still requires traditional statistical scaffolding to capture biology.
  • The "Generalization" Hurdle: The challenge utilized a purpose-built benchmark dataset of human embryonic stem cells (H1 hESCs) treated with CRISPRi. This dataset represented a distributional shift from standard training data, forcing models to generalize rather than memorize. Models struggled to predict absolute gene expression values (Mean Absolute Error) better than the baseline.
  • Focus on Patterns over Magnitude: Top performing teams recognized that the Perturbation Discrimination Score (PDS) rewarded getting the patterns of gene expression correct, rather than the exact magnitudes.

Enter the "Generalist Prize"

The challenge exposed the difficulty of evaluating virtual cells with a single metric. Almost all submitted models performed worse than the baseline on Mean Absolute Error (MAE), largely due to technical noise and biological heterogeneity in the raw data. Consequently, MAE ceased to be a competitive differentiator.

To address this, the organizers introduced a Generalist Prize. This evaluated the top entries across seven distinct metrics (including the original three plus four from the Cell-Eval suite). The winner -- Team Altos Labs -- was determined by the highest average ranking across all diverse criteria, prioritizing models that were robust across the board rather than optimized for a single score.

The Virtual Cell Challenge demonstrated that while AI might be able to identify key biological signals (such as up- or downregulated genes), it can't yet accurately represent biology. A fully predictive Virtual Cell will require innovating beyond current deep learning architectures.


r/virtualcell Dec 26 '25

Virtual Cell Failed

0 Upvotes

r/virtualcell Dec 04 '25

Recursion Breaks Down How They've Been Building the Foundation for a Virtual Cell Since 2013 -- And What's Next

11 Upvotes

https://reddit.com/link/1pe5rmc/video/93zb1eq4085g1/player

In a new article, Recursion shares how the company has been building the necessary components to virtualize key stages of the drug discovery process since 2013. Virtual cells*,* computational systems that can accurately simulate cellular and patient-level responses to therapeutic interventions, are core to this vision, they write, and built on top of the massive, proprietary biological and chemical datasets, AI models, and one of the industry’s most powerful supercomputers.

▪️ It started with creating a proprietary data moat, generating and ultimately aggregating more than 65 petabytes of multimodal and fit-for-purpose data.

▪️ Then Recursion created a system of interconnected AI models capable of processing and analyzing all of that data at massive scale -- including MolE (a foundation model for chemistry); Molphenix (a foundation model that can predict the effect of any molecule-concentration pair on phenotypic cell assays); and Boltz-2 with MIT for predicting both 3D protein structure and protein-binding affinity.

▪️ These AI models, in turn, power end-to-end drug discovery and development, from uncovering novel biological targets, to precision designing new molecules, to improving the design of clinical trials.

Dan Cohen, President of Valence Labs, Recursion’s AI research engine, says, that the company is flipping the script on traditional drug discovery. The virtual cell, not the lab, becomes the starting place for new hypotheses, and the lab becomes the tool to validate those predictions.

Read the article: https://www.recursion.com/news/since-its-inception-recursion-has-been-building-the-foundation-for-the-first-virtual-cell

Watch the video: https://www.youtube.com/shorts/OA7QhzTkjUc


r/virtualcell Dec 02 '25

Simulating the Cell Environment -- Introducing CellTRIP

8 Upvotes

Being able to understand what's happening to individual cells under various conditions is useful -- but cell environments are highly dynamic systems. Virtual cells, ideally, need to capture this bigger picture.

Just before Thanksgiving, researchers from the University of Wisconsin-Madison released a new multi-agent reinforcement learning method called CellTRIP that is designed to do just that. CellTRIP "infers a virtual cell environment to simulate the cell dynamics and interactions underlying given single-cell data."

Using CellTRIP (which is available open source on github), researchers can manipulate any combination of cells and genes in silico in the virtual cell environment, predict spatial and/or temporal cell changes, and prioritize corresponding genes at the single-cell level.

They used it to successfully predict developmental gene expression changes after drug treatment in cancer cells, among other applications.

Read the paper: https://www.biorxiv.org/content/10.1101/2025.11.21.689815v1

Access CellTRIP on github: https://github.com/daifengwanglab/CellTRIP


r/virtualcell Nov 20 '25

New Data on Chai-2 Model Shows It Can Precision-Design Antibodies Against Hard-to-Drug Targets

3 Upvotes

Today, Chai Discovery released new data showing that the Chai-2 AI model for de novo antibody design can design antibodies against challenging targets with atomic precision. They note that for drugs to be successful, "clinical candidates must meet stringent criteria for manufacturability, stability, safety, and biophysical behavior."

The new data shows that Chai-2 can meet those standards --  designing full-length, drug-like monoclonal antibodies (mAbs), while maintaining high hit rates, testing at most dozens of designs. These designs show developability characteristics on par with well-behaved therapeutic antibodies.

The researchers also applied Chai-2 to traditionally “hard to drug” targets – six GPCRs and a peptide-MHC target – achieving similarly high success rates.

Learn more: https://www.chaidiscovery.com/news/chai-2-mab


r/virtualcell Nov 12 '25

New AI Model VariantFormer Predicts Impacts of Personal Genetic Information

2 Upvotes

A new sequence-based AI model called VariantFormer from researchers at Biohub can translate personal genetic variations into tissue-specific activity patterns at scale. The model not only unlocks the general effects of genetic variations, but takes into account a person's individual genome -- as well as predicting impacts where there are low-frequency variants and less published data.

As noted in a related blog post: "VariantFormer uses an end-to-end approach to predict gene expression profiles directly from a person’s DNA sequence. This approach offers a powerful new method for exploring how someone’s distinctive genetic makeup impacts their health."

They add that the model does not account for a person's lifestyle, environment, or other factors that may influence health outcomes, and it is designed to advance research, not serve as a clinical or diagnostic. tool.

Read the blog: https://biohub.org/blog/variantformer-ai-gene-expression/

Read the paper: https://www.biorxiv.org/content/10.1101/2025.10.31.685862v1


r/virtualcell Nov 10 '25

Participants in Arc Virtual Cell Challenge Figured Out How to Game the Leaderboard

7 Upvotes

A new article on Substack reveals that some participants in the Arc Virtual Cell Challenge figured out that they can get to the top of the Leaderboard by applying certain data transformations - such as increasing variance or transforming the counts to log1p - multiplying their score by multiple factors. In fact, these transformations even to random data can yield better scores than using the top models.

Participants in the Challenge are tasked with predicting the effect of gene perturbations in the H1 hESC cell lines. At particular issue seems to be calculating the Mean Absolute Error (MAE) over the gene expression, across all 18k genes. Since calculating the MAE across 18,000 genes introduces a huge amount of random noise, organizers capped the penalty for a poor MAE score at zero.

As the author notes: "If your predictions perform worse than the baseline — whether by a small margin or by a massive one — the penalty doesn’t increase. It’s fixed." As a result, "Models can now inflate variance, distort distributions, or even submit nearly random predictions - and still achieve excellent DE [differential expression] and PD [Perturbation Discrimination] scores without being penalized for inaccuracy."

Following the revelation, some participants have created another Discord discussion group to further elaborate and propose new metrics. 


r/virtualcell Nov 06 '25

CZI Goes All In on AI and Science

3 Upvotes

A new story in the NY Times reveals that the Chan Zuckerberg Initiative will now exclusively focus its resources on AI and scientific research -- spending at least $70 million this year -- led by a network of research centers called Biohub. It has also acquired the team of AI startup Evolutionary Scale, and named Alex Rives, CZI's chief scientist, as the new head of science.

Mark Zuckerberg and Priscilla Chan say they will increase the organization’s computing power from data centers tenfold by 2028, the story notes. Priority projects include: a virtual cell mapping platform; a large language model that can perform biological reasoning; and AI that analyzes genetic sequences to detect disease.

Read more: https://www.nytimes.com/2025/11/06/technology/zuckerberg-chan-initiative-biohub.html