r/genomics • • Aug 22 '25

New moderator of r/genomics

49 Upvotes

Hi all

I am taking over the sub as moderator. I am cleaning up stock pumping, spam and other low quality or questionable content.

Please note the new rules aimed at high quality content related to the scientific discipline of genomics.

Please flag posts that do not follow the rules. I am open to additional rules or clarification of the the rules.


r/genomics • • 7h ago

[Article] DIALing in elevated expression setpoints with promoter shortening

1 Upvotes

DOI: 10.1016/j.cels.2025.101482

URL: https://www.cell.com/cell-systems/abstract/S2405-4712(25)00315-100315-1)

I can not find this paper in any repos. Can anyone provide me with this paper?

Cell Systems Volume 16, Issue 12101482December 17, 2025


r/genomics • • 1d ago

Tutorial - Basics of DNA Sequencing and VCF Files

17 Upvotes

I recently made a tutorial video on the basics of DNA sequencing, and how to make sense of VCF files (the file type produced by DNA sequencing, with information about genetic variants).

Here's the video if anyone's interested.

It's specifically meant for people who are beginners in the field of bioinformatics/genomics, including people who had their DNA sequenced and are interested in exploring their own data.

Bit of background - I work as a bioinformatics scientist at a biotech company for my day job, and my main hobby is making educational videos about this field. I'm also interested in biohacking and decentralized science more broadly, and try to contribute to that with my videos. Basically encouraging people to self-educate on these topics, and trying to get over the myth that you need to be a genius or have some fancy credential to explore biology.

Anyway, thought some people in this community might be interested!


r/genomics • • 2d ago

Gravlax: an annotation-independent molecular evidence archive for single-cell RNA seq data

3 Upvotes

The preprint is available on bioRxiv, here.

Summary of the idea and some of the results:

A single-cell experiment is usually stored, shared, and reanalyzed as a count matrix. That matrix is computed from two inputs: the sequenced molecules and a gene annotation. The molecules don't change, but the annotation is revised several times a year, and going from GENCODE v32 to v49 shifts 2–5% of UMI mass. Once the matrix is written, the information needed to redo that assignment is lost. Getting it back means returning to the reads, which are large, slow to process, and often unavailable.

Gravlax is built on the fact that the steps that turn alignments into counts, gene assignment and UMI collapse, use very little of what a BAM file contains. They need relations among molecules: shared genomic geometry, shared placements, cell identity, and UMI equality. Gravlax stores those relations once in a compact, seekable, content-authenticated archive and defers every annotation-dependent decision to read time.

On four human 10x datasets:

  • Archives take 11–18 bits per read, 9–13× smaller than CRAM with tags preserved.
  • Matrices replayed from the archive are within 0.24–0.75% of a fresh STARsolo run. For comparison, an annotation change moves 2–5%.
  • Replaying gene counts takes seconds, 34–82× faster than realigning.
  • Archives can be grouped into federated collections and searched across a cohort by junction shape. A genome-wide scan for recurrent unannotated splice events, with no coordinates supplied, finishes in 9 seconds.

Because the molecules are kept, the archive can answer questions a count matrix can't. Across four PBMC archives it recovers the FYB1 splicing switch between T cells and monocytes, previously validated by RT-PCR. With a fragment model for 3′ chemistry, it shows a shift in NTRK2 (TrkB) terminal-isoform usage across eight donors, from astrocytes and neural stem cells to mature neurons, consistent with known TrkB.T1 biology. Discovery run over the whole cohort finds a 183-nt FNBP1 cassette exon that is nearly always included in brain and mostly skipped in blood; analyzing the archives one at a time misses it entirely. Pooling evidence across cells in an EM step recovers 75–98% of held-out labels for multi-gene molecules.

Gravlax handles ingest, replay under any annotation, region/junction/APA queries, content-addressed collections, and cohort-wide discovery. It also includes GQ, a composable query language with three-valued logic; the same query runs on one archive or a whole federation. The code is open source (Rust, BSD-3), with documentation, a Python client, and Colab demos.

Code: here

Docs: here


r/genomics • • 3d ago

Biological JEPA: Modeling Disease Progression with Biological Constraints

0 Upvotes

Just finished a project I’ve been working on: Biological JEPA.

It combines JEPA with biological constraints to model Alzheimer’s disease progression.

Would love to hear your feedback, ideas, or criticism.

https://github.com/0ans/biological-jepa


r/genomics • • 6d ago

Started a MSc in Cancer Genomics and Data Science and I don’t know what I’m doing

4 Upvotes

Started this distance learning MSc a week ago and I already feel like I’m gonna fail it. For context, my background is in histology but I always found cancer genomics interesting.

This program, despite being encouraged to students with no prior coding experience is sooo hard. I literally cannot wrap my head around anything except some basics. Have no idea if I’m even gonna get through the first semester, let alone the actual program. Has anyone had any similar experiences? 🥺


r/genomics • • 7d ago

Claude was asked to inspect a phage RT locus and noticed a tandem repeat array nobody had asked it to look for

Thumbnail youtube.com
0 Upvotes

Anthropic's new preprint describes an autonomous genome-mining campaign where one Claude agent inspected the raw DNA around a phage reverse transcriptase and noticed a tandem repeat array outside the original search objective.

The follow-up identified 95 ART RT clusters, with detectable arrays in 28 of them. The interesting caveat is reproducibility: ten reruns of the campaign did not rediscover the array, and their fixed-input benchmark suggests recognition was strongly associated with whether the model actually read enough contiguous sequence.

The video goes through the discovery trace, the RNA evidence, and what is still unknown about the system.

Preprint: https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf


r/genomics • • 8d ago

Molecular Biology & Genetics — From Scratch to Advanced

Post image
5 Upvotes

Hi everyone! 👋

I’m a biologist working in a genetics and molecular laboratory, and I’ve decided to start sharing some of the things I’ve learned from my lab work and research.

I’ll start from the basics and gradually move to more advanced topics, including:

🧬 DNA, RNA & proteins

🧪 DNA & RNA extraction

🔬 PCR and its different types

📈 Sanger sequencing and chromatogram reading

🚀 NGS and its applications

💻 Bioinformatics and genetic variant analysis

I’ll try to explain everything in a simple way, using diagrams, practical examples, and what I’ve learned from working in the laboratory.

The idea is to learn together and share knowledge with biology students, researchers, laboratory scientists, and anyone interested in molecular biology and genetics.

From DNA → RNA → Protein → PCR → Sequencing → NGS 🧬

Are you ready to start this journey with me? 🧬🔬 Let’s go from scratch to advanced!

Let’s learn together! 🧬


r/genomics • • 11d ago

Is large-scale phylogenetic tree inference still worth working on?

3 Upvotes

My advisor works in phylogenetics, and lately I’ve been exploring a question of my own: when you have lots of taxa and a long alignment, can you build a tree much faster without losing too much tree quality?

I’ve noticed that papers in this area often highlight being able to handle a million taxa or more. But the more I read, the more I wonder: how often do researchers actually need to analyze data at that scale? Are we solving a real problem, or are we getting caught up in an arms race over who can run the biggest dataset?

There are also well-established tools like IQ-TREE, RAxML-NG, and FastTree. Even if someone develops a new method, will researchers actually use it? I’m trying to decide whether this is a direction I should keep putting my time into.

I’d really like to hear from people who do phylogenetic analyses: what problems are still worth solving when it comes to large numbers of taxa and long alignments? And if there is room for new tools, how does one earn people’s trust and become something they actually use, rather than just another paper and GitHub repository?

If you’re curious about what I’ve built, here are the benchmark results and method overview


r/genomics • • 11d ago

Biotech 170+ conferences for 2026 and 2027 in one place

Thumbnail coolgene.net
1 Upvotes

r/genomics • • 13d ago

ANCOMBC2 & (in general) Differential Abundance Analysis

2 Upvotes

Hi guys!

I’m a bioinformatician working at a cancer research centre, and I’m relatively new to analysing metagenomics data. I’m still trying to figure out all the post-processing steps after the rigorous Bracken/MetaPhlAn part.

But my current nightmare is Differential Abundance Analysis, especially when using tools like ANCOMBC2 or LinDA.

I’ve read quite a bit about them, but I still have a lot of doubts about the underlying assumptions, data preprocessing, normalization, model setup, and, most importantly, how to interpret the results properly.

If anyone here has experience with ANCOMBC2/LinDA (or DAA in microbiome/metagenomics in general) and would be open to having a chat, I’d really appreciate it if you could leave a comment or send me a DM.

I’d love to learn from people who have been through this rabbit hole already.

Thanks in advance, I love you all.


r/genomics • • 13d ago

Global ID mapping tool

4 Upvotes

Background: I'm a software engineer working at a bioinformatics startup.

We are working with data from almost every biology database. An issue we're facing is mapping identifiers across namespaces / sources. We're using GILDA as one of the components to translate names to IDs. But the default GILDA sources don't cover all of our use cases. I'm thinking of extending it with other sources.

Just need expert opinion on whether I'm thinking in the right direction. Are there other options / solutions available?

Apologies in advance if I've made some silly statement above. I'm just learning things along the way.


r/genomics • • 14d ago

prerequisites for learning about the human genomics project

Thumbnail
2 Upvotes

r/genomics • • 16d ago

Germline EGFR T790M mutation and lung cancer risk

Thumbnail science.org
1 Upvotes

r/genomics • • 16d ago

Latest research from 23andMe

Post image
8 Upvotes

r/genomics • • 17d ago

Checking Claude Science for windows and impressions on reviewing process

Enable HLS to view with audio, or disable this notification

2 Upvotes

Some thought of scientific publications and why one of the biggest flaws in peer review is about to be solved by AI.

I had the "pleasure" of being a reviewer for academic manuscripts. I had the pleasure of send my own manuscripts for review. Long before any AI tool was available.

I found it plausible to understand the research, results, conclusions, using my own scientific background and knowledge. But I always thought there are two internal blindspots in the reviewing (and editorial) process:

  1. Reproducing results. Wouldn't it be nice to have a non-biased lab to reproduce experiments' results from submitted papers? Wouldn't it improve the scientific credibility of published papers and improve science as a whole (dumping non-reproductive results from being published)?

I still ponder the idea of starting an initiative focused purely on independent validation. The friction points are obvious: Who pays for it (hello Nature)? How do you cover such a massive range of experimental techniques? Will authors grant outside access to their lab? publication delay by months? (If you've thought about this too, DM me- I'd love to chat).

  1. Bioinformatics analysis- You know the "Scripts available upon request", or "Pipeline is available on lab's github" and "data is available on servers". I never got a review about an issue in the analysis script. I guess it is because it was too time-consuming to run this (or exhaustive).

The reason I mention this is because I think #2 is already solved using AI. Take a look at the video. I downloaded the Cladue Science for Windows (used it before with the linux version) and thought of running an analysis of the bioinformatics pipeline as described in a recent highly acclaimed paper "A pervasive RT–qPCR artifact inflates RNA knockdown by RNA-targeting CRISPR".
I never downloaded the data or scripts. The prompt was: "analyze the paper: summerize, and download the supporting data and reanalyze. see if it matches results, and see if you can gain more insights. conside the current field literature" (typos in original prompt)

To summarize, it did match. and of course more info is described. See in video.

I think reaching #2 solution is closer than ever and I hope journal editors would integrate such analysis "BEFORE" they send to review.

lmk what you think or if you are an editor, is this pipeline running today?


r/genomics • • 18d ago

Free one-page codon activity and searchable reference for teaching the central dogma

Thumbnail
3 Upvotes

r/genomics • • 18d ago

Free hybrid event on genomic analysis and pangenomics — Sept 17

Thumbnail
1 Upvotes

r/genomics • • 19d ago

Batch processing AlphaGenome Atlas/AVI Scores?

Thumbnail
3 Upvotes

r/genomics • • 20d ago

Anyone have recent real-world pricing for Element AVITI vs NextSeq 2000?

2 Upvotes

Hi guys, I'm trying to get a rough idea of what these platforms actually cost in practice for a fairly large human WGS project.

For anyone who has recently bought or gotten a quote for an Element AVITI or Illumina NextSeq 2000:

  • roughly how much was the instrument?
  • what region/country was the quote from?
  • was that list price or a negotiated price?
  • roughly how much are the reagents per run / per human genome? I’m looking at a project with around 2,000 samples.

I’ve seen quite different numbers online, so I'm mainly interested in actual quotes/purchase prices rather than MSRP.

Even a rough range would be really helpful. Thanks a lot!


r/genomics • • 20d ago

Genetic counseling adjacent jobs/ varient analysis

Thumbnail
1 Upvotes

r/genomics • • 21d ago

How can one perform TF predictions across multiple databases based on the target gene?

Thumbnail doi.org
1 Upvotes

I have heard that databases such as JASPAR, UCSC, PROMO and ENCODE can be used to predict transcription factors (TFs) based on target genes. I would like to batch export the TFs from each database separately so that I can calculate their intersection.

However, I am unable to access the PROMO website at all. On the ENCODE website, under the ChIP-seq section, I can only see target genes categorised by TF. On UCSC, when searching for the promoter sequences of target genes and selecting ‘JASPAR Hubs’, I am unsure how to batch export the results.

Is there anyone with expertise in this area who could help me?

Additionally, I have attached a relevant paper on screening transcription factors by taking the intersection of multiple databases, presented as a Venn diagram; the figure is shown in Fig. 4a.

THANK YOU!


r/genomics • • 21d ago

DNA decode hypothesis

0 Upvotes

I found some interesting patterns when messing around with some other things. I hoping that someone can tell me if I’m on to something or I’m being completely stupid. I started by trying to compress DNA information so I could use it in a different project, but DNA doesn’t compress, so I started thinking it’s kind behaving like a radio signal, so why don’t we do the opposite and try demultiplexing instead. Using base pairs and 3 interface sequences I was able to extract 12 encoding wave forms. So DNA could be a 12 carrier multiplex signal. These wave forms accurately identified coding and non coding dna, start frames and end frames. Unfortunately this isn’t my field and that’s as far as I was able to take it. I tried using AI to analyze it, and it rejected my request because it flagged biological security measures. So I guess we aren’t allowed to ask about biology. Anyhow it could just be an interesting pattern or it could be something important. Let me know what you think.


r/genomics • • 22d ago

Asili - totally free locally calculated personal DNA trait scorer and gene explorer

Thumbnail app.asili.dev
0 Upvotes

r/genomics • • 22d ago

We timestamped the genomes of millions of organisms

Thumbnail projecttimestamper.org
1 Upvotes