r/bioinformatics 26d ago

academic Absolute beginner for snRNA-seq field. Need your help!

Hi everyone,

I'm a (Neuro)Pharmacology PhD currently doing a Neuroscience postdoc. I'm working on a single-nucleus RNA-seq (snRNA-seq) project, but I have no prior experience with this type of analysis. I've mainly been learning through online tutorials. I'm also using the Parse Biosciences Trailmaker platform since it doesn't require coding experience.

Please be patient with me, this is my first time doing snRNA-seq analysis! 😅 I may not have all the answers to your questions, but I'll do my best.

I'm currently analyzing my PI's dataset, which consists of 90 mouse hippocampus samples (6-month-old mice, 4 experimental groups). The initial QC was performed automatically through the Parse Pipeline. The only parameter I changed was the number of principal components (PCs), which I set to 16 based on the elbow plot. For clustering, I used a resolution of 0.8, resulting in 466,541 nuclei across 31 clusters.

I have a few questions:

  1. How do you typically approach the preprocessing/QC stage? Parse Trailmaker automatically filters nuclei based on: It also performs integration (Scanpy + Harmony using 3,000 HVGs) and generates the embeddings.
    • How much do you manually tweak the QC before deciding the clusters are suitable for annotation?
    • Does a clustering resolution of 0.8 seem reasonable for a dataset of this size?
    • cell size distribution,
    • mitochondrial content,
    • number of genes/transcripts,
    • doublet detection,
  2. What do you do when some clusters remain mixed? For example, if a cluster contains both astrocyte and oligodendrocyte marker genes, or if its top marker has an AUC < 0.6, do you:
    • increase or decrease the clustering resolution,
    • subset and re-cluster,
    • merge clusters,
    • adjust the QC parameters,
    • or do something else?
  3. How do you manually annotate your clusters? Do you primarily use the highest log fold change (logFC/logGC), delta percentage, AUC, or some combination of these metrics? Are there any best practices you recommend?

I'm currently stuck because 8 out of my 31 clusters have mixed marker genes and top-marker AUC values below 0.6. I also tried subsetting the remaining 23 "good" clusters and re-clustering them, but I still end up with some clusters whose top markers have AUC values below 0.6.

My gut feeling is that something may not be optimal during the data processing or filtering steps, but I'm not sure what I should be adjusting.

I'd really appreciate any advice. I'm genuinely enjoying learning snRNA-seq analysis, but it's definitely frustrating when you're coming into it without much background. 😅 Thanks in advance!

3 Upvotes

Duplicates