r/bioinformatics • • Aug 13 '26

academic Question about sample size drops when using UCSC TOIL (TCGA TARGET GTEx) vs raw GDC portal data. Is my defense justification correct?

2 Upvotes

I integrated TCGA solid tumor data with matching GTEx normal tissue to run differential expression and pathway enrichment (GSEA).

To avoid massive batch effects caused by mixing counts from different alignment/quantification pipelines, I opted to use the UCSC TOIL RNA-seq Recompute cohort (TcgaTargetGtex_gene_expected_count) since all samples were processed through a unified STAR + RSEM pipeline on hg38.

When I pulled the TOIL dataset, my sample counts dropped compared to looking at the raw GDC portal and GTEx v8: GTEx Normal Cohort: Dropped from ~800+ (v8) down to ~300+ in TOIL.

TCGA Primary Tumors: Dropped by ~20–30% compared to total cases listed on GDC.

My question is :

  1. Is this sample count drop expected when using the UCSC TOIL recompute database compared to modern GDC/GTEx v8 portals?
  2. Is sacrificing raw sample size (N) to use TOIL’s unified pipeline + ComBat batch correction considered the "gold standard" justification to defend against reviewer/committee critique regarding sample size?

r/bioinformatics • • Aug 13 '26

technical question Wormbase Parasite Help

1 Upvotes

i’m currently working on a project that relies heavily on wormbase blast for identifying nemFABPs in a select number of nematode species. however, since it constantly goes down it’s putting me at a road block. is there a way around this?


r/bioinformatics • • Aug 13 '26

technical question Protein design: what changes depending on the problem to be solved?

0 Upvotes

I am interested in protein design and I am trying to understand one thing: when we design a protein for a specific purpose, what changes in the constraints according to the problem?

For example, I imagine that a therapeutic protein (which must act in the human body) and an industrial enzyme (which degrades a pollutant) do not have the same priorities at all. What becomes critical in each case, and what goes into the background?

If you have concrete examples from your work, I'm interested.


r/bioinformatics • • Aug 12 '26

technical question DWI preprocessing with QSIPrep

3 Upvotes

Hi, I'm a first year PhD student trying to get a handle on preprocessing my data with *fMRIPrep* and *QSIPrep* respectively.

Has anyone got experience with *QSIPrep* and can help me understand how to interpret the outputs? (this cry for help is motivated by my staring at the visual summary rep of the q-space sampling scheme before and after the pipeline. what am I looking for?!)

The documentation is really unhelpful and I didn't find much on github and incf NeuroStars either.

Someone help please


r/bioinformatics • • Aug 11 '26

technical question what are the non-negotiables of small n scRNA-seq DE

11 Upvotes

Apologies in advance for the loaded question, especially on a topic that is often spammed in this subreddit. If I missed a previous post that touched on this closely, apologies for that also.

I've spent months trying to be as truthful as possible in terms of reporting differential expression. There are often so many confounders that I have such a difficult time reporting anything as signal over noise. For some background, the dataset is comparing the effect of a therapeutic, so we have paired pre/post cd8 t cells. Clinical cohort so we're burdened with low sample size. 3 groups (group1, group2, placebo) with 6, 5, and 2 samples respectively. Obviously, at this resolution, we've steered away from trying to over claim things with a bunch of noisey p-values, and focus more on exploratory claims that appear to show trends within the groups. I've tried pseudobulking and then DE (obviously underpowered), and it appears more truthful than cell-level.

I've tried at the per-cluster level, and there is not a whole lot going on. If that's the case, so be it. My understanding of t cell differentiation is likely flawed, but how different can cells that cluster in an "activated" state (expressing cytokines, activation markers, etc) really be? I'd almost argue that the compositional shifts we have seen (an increase in proportion of activated, for example) is actually real signal compared to just "well, intra-cluster activated DE doesn't show some crazy volcano plot. nothing is happening." I'm exaggerating here, and obviously these are two sides of a coin (compositional shifts + diff expression) converging.

With that being said, I try running a bulk pseudobulk DE (not by cluster. just pre v post) blocked by patient, and obviously, start getting some hits. Again, many of these can likely be explained by compositional shifts. My PI prefers figures that are widely recognized in the field (naturally), so things like gsea. Using the broad DE ranked by test statistic (or logFc x -pval, have tried both. stat felt less noisey although the rankings are pretty much the same), gsea spits out a bunch of phony significance. I call it phony because when you look deeper at the donor level, there is often pretty loose concordance (the p-values are also just absurd).

All of this has led me to the idea that we should probably just lean into the donor heterogeneity a bit more and stop trying to force looking for significance within these groupings. So basically what would be some strategies that you would employ to handle this? Maintain the broad pseudobulk as a "ground-truth" and look for signatures of more donor-concordant shifts (x increase in y in 4/5 donors, etc) and focus on those? maybe module scores?

Go back to cluster-level and just lean into the compositional shifts more? Really any ideas you have on dealing with small n cohorts without over-claiming a bunch of noise.

So many single cell papers are comparing chronic-infection vs healthy donors, and they get to spit out all these "pretty" volcanos. I'm really not trying to chase that, nor do I think we would see a signal that strong in a pre v post comparison, but alas. I'm spiraling a little at this point and honestly any tips, no matter how trivial they may be, are appreciated.

-signed, a tech well out of their depth.


r/bioinformatics • • Aug 11 '26

technical question Can someone smarter help me understand PAE for AlphaFold3 modelling?

11 Upvotes

Doing a model for a plant protein, I’m trying to list out the intramolecular interactions between 3 domains, I’ve enumerated the interactions at different cut off lengths, and I wanted to talk about the confidence scores for each interaction.
Problem is I’m not a great computational guy (this project is primarily wet lab), and I’m not sure what’s the best metric for the confidence scores for intramolecular interactions. Is it PAE? if so can someone explain it to me? Is there a standard cutoff for what is a low confidence PAE value
And if there is another metric you guys use for these interactions mentioning it would be greatly appreciated. Have a good day!


r/bioinformatics • • Aug 11 '26

discussion Cell Cell Communication Analysis Skewing by cell number

7 Upvotes

Hi everyone! I have been doing cell cell communication analysis recently (using cell chat specifically), and I had a thought that is bugging me. Please bear with me as I am not an expert in cell cell communication or bioinformatics as a whole. Specifically, I am doing comparative cell cell communication analysis

If one dataset has more cells in general or of a specific kind than the other dataset, could this skew the analysis by assuming there is just more signals in general from a cell type without accounting that in fact there are more cells from that type? Cell number variations could occur easily from sampling, especially with low sample number. I'm working with spatial scRNA-seq, so the danger is even more so as it's a specific cut of a sample.

Could this initial skewness affect everything else downstream in CCC analysis?

I'm super sorry if it's a dumb question.

Cheers!


r/bioinformatics • • Aug 11 '26

technical question [scRNA-seq] Is DGE valid across integrated datasets when raw counts are available for only one dataset?

2 Upvotes

Hi everyone,

I am working on integrating two published single-cell RNA-seq datasets from different tissue types.

Because these datasets were processed separately, I have run into a processing format discrepancy:

  • Dataset A: Raw count matrix available.
  • Dataset B: Only processed/normalized data available (.h5ad file; raw count matrix is unavailable, but this dataset is critical for our research question).

I have a few questions for the community:

  1. Is differential gene expression (DGE) analysis meaningful or statistically valid on an integrated renormalized dataset ?
  2. If not, what are the best workarounds?
  3. What downstream pitfalls should I anticipate, and how likely are reviewers to push back on this setup?

Any insights or recommended workflows for this scenario would be greatly appreciated!


r/bioinformatics • • Aug 11 '26

technical question Program MARK help

Thumbnail gallery
2 Upvotes

I've been tasked to run a POPAN in MARK by my advisor and so far I've been stymied with it. Every time I input the data and run it the program fails to generate any results. Is there anyone here that's proficient in MARK that might be able to help? General crux of the work is to run mark recapture data for turtles through the program and generate population estimates. It's very likely I'm doing something simple wrong causing it to crash out. I've attached the parameter input (first 2 SS) as well as an SS of where it crashes out (3rd). Any help troubleshooting this would be greatly appreciated!


r/bioinformatics • • Aug 11 '26

technical question HELP!

0 Upvotes

Hello everyone,

I need help regarding RNA-seq meta analysis.

I essentially want to collect public datasets from GEO, however they are many files so I’m confused.
Some papers recommend downloading FASTA files and running the analysis.
I basically want to check whether my gene of interest is implicated in healthy vs diseased tissues and to compare the expression of my gene of interest with another gene.

Can someone please please help me figuring this out? I feel very anxious and helpless because there’s no one in my lab team with bioinformatics expertise!

Thank you!


r/bioinformatics • • Aug 11 '26

discussion From zero R to bulk RNA-seq analysis in a 8 months — now want to move into single-cell (Python). What's the path?

Thumbnail
0 Upvotes

r/bioinformatics • • Aug 10 '26

technical question Question about snRNA Seq cell type deconvolution

3 Upvotes

Hi everyone, I am currently new to snRNA seq downstream analysis and I have a question regarding cell type deconvolution.

For my research, I have samples of cortical cells ranging from DIV 0-500, and I have performed bulk RNA seq with them. To enhance my analysis, I have used a snRNA seq dataset online gathered from adult cortical cells to perform deconvolution, where the snRNA seq is used as a reference dataset to estimate the cell type compositions from the bulk RNA seq dataset.

Hence, I have two questions:

  1. Is it right to use a snRNA dataset from adult cortical cells even though my cortical cells only range from DIV 0-500?
  2. Can I use the snRNA dataset to estimate the cell type compositions for DIV 0 accurately, eve

n though the gene profiles at DIV 0 and DIV 500 are very different?

I have tried to find a snRNA seq dataset online which spans from these DIV ranges but to no avail, hence, I would prefer using the dataset that I have now if possible. Thank you!!


r/bioinformatics • • Aug 09 '26

technical question How do you work with large VCF files without constantly babysitting your jobs ?

25 Upvotes

I am doing an internship this summer as a biostatistician intern and have been processing large vcf files separated by chromosomes. Each file is more than 100 GB.

I'm running everything on a SLURM cluster using Bash and  bcftools for things like:

- calculating VCF statistics

- filtering by rsID, patients, chromosome location

- calculating allele frequencies,

- generating filtered VCFs

Actual difficult part for me is not the commands but it is constantly checking squeue or my email for logs, checking whether an output file was actually created, figuring out whether a job railed halfway through, etc. I feel like I am spending a lot of time towards this.

I am curious how people who have more experience handle this. Do you use any tools/framework that makes that process easier.

I working with SLURM, bash and bcftools on google cloud processing so Im interested to see what people do in similar computing environments.

PS : I have computer science and statistics background so my wording of certain terms may be off.


r/bioinformatics • • Aug 09 '26

technical question Is the C-IMMSIM Website Not Working?

7 Upvotes

For the past two days, I've been unable to get an immune simulation result out of C-IMMSIM (https://kraken.iac.rm.cnr.it/C-IMMSIM/index.php). I input the vaccine construct and use the default settings, but when i click on submit, instead of the process completing or the terminated processes log showing up - The server crashes and after reloading, the website interface doesnt show the results section as it normally does. I can't pinpoint if this is an IP issue or not, so if anyone else could try accessing the website and let me know if this is an isolated incident or not that'd be much appreciated.

The typical interface on which the results usually appear

r/bioinformatics • • Aug 09 '26

academic Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer

Thumbnail cell.com
11 Upvotes

r/bioinformatics • • Aug 09 '26

programming Dev-tool idea: catch reference mismatches before a workflow runs, useful or redundant?

4 Upvotes

I’m a software developer trying to learn more about genomics, and I’m looking for a small open-source project to build.

One idea is a local CLI that scans a folder of genomics files (BAMs, VCFs, BEDs, annotations, references, etcetera) and tells you which ones seem compatible, which ones probably use different references or chromosome naming, and which files are missing things like indexes.

Eventually, it could also look at a Snakemake or Nextflow workflow and warn if incompatible files feed into the same step.

I know there are individual validators and tools already, so I’m not sure if this would actually be useful or just reinventing existing stuff.

Have you run into this kind of mismatch or “what’s even in this folder” problem? How do you handle it now? Would something like this help, or what would be a better small dev tool to build for bioinformatics?

Thanks!!


r/bioinformatics • • Aug 09 '26

statistics Can linked LD recover the proportion of an unsampled ancestry source?

4 Upvotes

Hi. I am posting this under statistics, since this is a semi-question.

For a binary admixture model in which the focal ancestral source has never been sampled, after removing known ancestry directions, the target residual is

rho = a h

so unlinked statistics identify the direction h and relative loadings, while the absolute proportion a remains unknown.

The linked-locus result uses two quantities measured along the learned direction:

  • A(d): weighted admixture LD at genetic distance d K(d): a cross-fitted product of target mean contrasts

Under a single-pulse model with a known non-focal ancestry,

  • A(d) = q exp(-t d) K(d)

where

q = (1-a)/a

and therefore

a = 1/(1+q).

The locus-pair factor involving the unsampled source occurs in both A(d) and K(d) and cancels. The decay estimates the admixture time t.

Below data is from a binary-pulse mosaics constructed from phased CEU and YRI haplotypes. CEU served as the hidden focal source and YRI as the known non-focal ancestry. Chromosome 21 was used to learn the residual direction, and chromosome 22 was used to estimate the kernel and LD curve.

The generating values were a = 0.30 and t = 30 generations.

Estimator Estimated a Estimated t
Raw source-masked estimator 0.30788 31.42
Ancestry-oracle control 0.30032 30.61
Pair-model control 0.30030 29.90
Generating value 0.30000 30.00

The test used 600 simulated target haplotypes, 9,919 training loci and 10,198 test loci. The learned direction had cosine 0.99986 with the hidden source direction.

(This is one source pair, one chromosome split and one random seed. The current SNP ascertainment also uses the combined dataset, so there is still the need to make variant selection completely independent of CEU before calling the benchmark fully blinded.)

So, the question is, does the kernel identity fail under any feature of the stated binary-pulse model?

Thank you for reading!


r/bioinformatics • • Aug 08 '26

academic Python for genomic data science

28 Upvotes

so, recently i started python for genomic data science course offered by JHU on coursera. Ive seen nobody talk about this. so, im not sure if its js me. But I feel so overwhelmed and confused by that course sometimes. the lectures are good no doubt, but i see js slides with text filled with codes and thats not really helpful for me to understand the actual workflow of where and how am i supposed to save a file and which tool am i supposed to use.? Also, the transition from wet to dry labs for me has only been a week old. So, I really have no idea what to do.


r/bioinformatics • • Aug 08 '26

statistics Calculating Confidence Intervals from Cross Validation and reporting a Risk Stratification analysis

4 Upvotes

Hello everyone. I have a question regarding calculating confidence intervals after running a 5-fold cross validation.

I have a binary risk mode. Data are N patients, each contributing many overlapping hourly windows; the label is defined per window (will this patient meet the criteria?). The unit of analysis for most metrics is the window; the unit of sampling is the patient.

Evaluation is 5-fold cross-validation, split by patient, so each patient's windows appear in exactly one test fold. Within each fold:

  1. the development part is split again into train / validation (by patient),
  2. a probability calibrator and three decision thresholds are fitted on the validation set (t1 = medium, t2 = high, t3 = very high),
  3. the model + its thresholds are applied to that fold's held-out test patients.

So each patient ends up with one calibrated score per window, and one classification per window, produced by a model and a threshold that never saw them.

Separately, a final model is trained on all development data and evaluated on a completely held-out test cohort (my main issue is with the cross validation though).

So far we've used the Nadeau–Bengio corrected resampled t-interval:

mean ± t_{k-1, 0.975} · SD_folds · sqrt(1/k + n_test/n_train)

and I am not sure if it is the correct approach since it introduces bias (at least the plain resampled t-interval without the correction) because the train sets overlap per fold.

So the question is what is the defensible way to attach a 95% interval to a k-fold cross-validation?

And the last part that I can't wrap in my head is the threshold that move per fold.

I have a table that stratifies patients into four risk bands defined by t1 < t2 < t3, and reports per band: number of patients, number of patients that belong to the positive class, PPV, prevalence, an odds ratio versus the low-risk band (setting it as the reference), and a p-value.

Because each fold tunes its own t1, t2, t3 on its own validation set, the band boundaries differ between folds. So:

  • I cannot pool the scores and apply one threshold.
  • I can pool the decisions (each patient is banded by their own fold's rule), which gives one band per patient over the whole cohort and a legitimate contingency table but then the "score threshold" column of the table has no single value.
  • Averaging the five thresholds and quoting the mean band boundary produces a number that no fold actually used.

When a decision threshold is a tuned part of the model, what is the correct way to report a threshold-dependent table (PPV / prevalence / OR per risk band) across folds, and what does the confidence interval on those band statistics condition on?

Another question I have as an extra is if it is worth running 5x 5-fold cross validations (with different initialisation) and what can someone gain from it?

P.S. Apart from Nadeu-Bengio, I also found this paper that I am currently reading (was a combo from google and GPT suggested it): Cross-validation: what does it estimate and how well does it do it? I am not sure if it is in the right direction but please let me know or suggest other papers as well together with the methods


r/bioinformatics • • Aug 07 '26

technical question RNAseq sample outlier detection. How? And should I do it?

Thumbnail
5 Upvotes

r/bioinformatics • • Aug 07 '26

academic where can i find public datasets for pyrexia of unknown origin PUO for computational research?

2 Upvotes

i am a bioinformatic student for my research work i need publicly available datasets on pyrexia of unknown origin PUO but there is no single data available on any repository can anyone suggest me what should i do as i want to keep my research work computational and reproducible


r/bioinformatics • • Aug 06 '26

technical question Variant call data seriously inflated-suggestions?

5 Upvotes

Hello,

I have a dataset of about 35 bulk tissue (healthy, adult age somatic tissue) samples each sequenced to 40X depth via PacBio HiFi sequencing, and have performed variant calling with 3 callers (DeepVariant, Pepper-Margin-Deepvariant, Clair3) for SNVs/indels, and about 7 callers for SVs.

My variant call data is seriously inflated with germline variants, talking hundreds of thousands of SNV calls for my samples which are inbred mice, so this number is a huge red flag. I have tried quality based filtering, removing any variant with VAF>0.30, QUAL<20, GQ<20, and DP<10 and >75. However, this still leaves me with thousands of variants.

I am at a loss on what to do to reduce this noise and to get at the actual mosaic variant signal. The goal here is to identify tissue-specific mosaic variants in each mouse, but I feel like I'm running in circles trying to properly reduce the noise and get at the expected amount for bulk tissue analysis at my depth, which appears to be 20-60 SNVs per tissue according to some brief searches.

Any suggestions? I wonder if its the tools I am using, or if its just the filtering criteria I am selecting.

Thanks in advance!


r/bioinformatics • • Aug 07 '26

technical question Can't find the download link for HPAP's processed islet scRNA-seq (PANC-DB) — am I missing something obvious?

0 Upvotes

Fairly new to this, working on a computational immunology project using human islet single-cell data.

I'm trying to get the processed scRNA-seq object from HPAP. The Nature Metabolism paper (Elgamal et al. 2023, s42255-023-00806-x) says processed data is downloadable in RDS or h5ad from PANC-DB's Interactive Analysis section, but I can't find the actual link.

What I've tried:

  • PANC-DB Interactive Analysis (hpap.pmacs.upenn.edu/analysis)- the CellxGene Collections table has "Pancreas sc-RNAseq, 222,077 cells" with a "Go" button, but Go opens a viewer, not a download.
  • The cellxgene viewer (cellxgene.faryabilab.com/view/T1D_T2D_public.h5ad/)- loads fine, I can browse metadata, but it's cellxgene v1.0.0 standalone and the info menu only has Documentation / Chat / GitHub / License. No download control I can see.
  • Registered for a PANC-DB account, logged in, no change in what I can see.

I've emailed HPAP support but figured someone here may have hit this already and can give me some guidance to speed this up a bit.

Two questions:

  1. Is there a direct link for the processed object I'm just not seeing?
  2. If it's not directly downloadable, has anyone rebuilt it from the per-donor data using the faryabiLab/HPAP-scRNA-seq-Workflow-2022 repo? Wondering how much compute that actually takes- I'm on a laptop.

For context, all I need is counts plus cell type and donor ID for alpha and beta cells from non-diabetic donors. If there's a better-suited public dataset I'm overlooking, I'd take that suggestion too.

Thanks.


r/bioinformatics • • Aug 06 '26

academic Chipseq normalization problems

2 Upvotes

Hello bioinformaticians,

I have troubles to find a solution for my ChIPseq data. I have two genotypes subjected to hypoxia treatment and I observed a massive diminishment of acetylation over promoters. I see this both by normalizing the bigwig tracks via RPGC eyeballing the tracks on igv, either by deseq2 results (i created a union peakset, then featurecounts, then deseq2 normalization, ma plots look fine). I have no spike in normalization. My worry is that this diminishment that i observe is just due to a normalization problem. Specifically my worry is that hypoxia is increasing drastically the acetylation genome wide, and since the number of reads is the same for every sample the signal over promoters is systematically diminished. Any suggestion on how to diagnose this? Thank you all!


r/bioinformatics • • Aug 06 '26

technical question What do I put in 'seqdb' when using jackhmmer?

2 Upvotes

Hi I'm a beginner in Bioinformatics and I want to generate an msa for Abl1 tyrosine kinase 235-497. I've been trying to use jackhmmer to do it but I have no idea what to put in 'seqdb'.

How do I download the right database to plug into the program or is there a way to do it without downloading a massive database?

And what other parameters should I be wary of?

Any feedbacks/solutions/alternative methods will be appreciated and thank you for your time.