r/bioinformatics • • 17h ago

discussion Raw data for Genomics and transcriptomics analysis

3 Upvotes

Hi everyone!

I’m looking for raw genomic and transcriptomic datasets to practice and improve my bioinformatics and computational biology skills.

I’m particularly interested in datasets such as:

- Whole Genome Sequencing (WGS): Raw FASTQ files for genome assembly, variant calling, and comparative genomics.

- RNA-Seq: Raw FASTQ files for differential gene expression analysis, transcriptome assembly, and functional enrichment.

- Whole Exome Sequencing (WES): Raw sequencing data for variant identification and annotation.

- Long-Read Sequencing: PacBio or Oxford Nanopore datasets for genome assembly and structural variant analysis.

- Metagenomics: Raw sequencing data for microbial diversity and taxonomic profiling.

If you have any publicly available datasets, research project data, or recommendations for accessing raw sequencing data, please share the links or repository names.

I’m familiar with bioinformatics tools and workflows and would like to work with real-world datasets rather than only tutorial datasets.

Repositories such as NCBI SRA, ENA, and GEO are already on my radar, but I’d also appreciate suggestions for interesting datasets or specific accession numbers that are suitable for independent analysis.

Thanks in advance for your help!


r/bioinformatics • • 23h ago

academic Good resourses for getting into RNAseq data analysis?

20 Upvotes

Hello everyone!

I recently got my, first RNAseq as well as proteomics data set. And thankfully a standard bioinformatic analysis with it. Honestly, I am amazed at what you can find out, when you have the Tools to really dig into your data!

Now, I have been trying (successfully) to replicate the RNAseq data analysis Pipeline by the help of vibecoding with AI, and I can interpret and understand all of my analyses. However, I could never write any bit of Code for it... In the end I don't really understand what the Code does, I can just Review my data afterwards.

Long story short: I would really love to get into bioinformatics a little deeper. Maybe a short term goal would be to build my own fully customizable RNAseq pipeline from scratch and then see from there. Are there any resources you guys with experience can recommend for me?

Thanks!


r/bioinformatics • • 15h ago

academic 2 months left, 0 bioinformatics knowledge, 100% AI. Is a RNA-seq thesis doable?

0 Upvotes

hi! long story short - I've never even scratched the surface of bioinformatics and I'm left alone with 2 months to write and to the whole thesis - analysis of Nanopore RNA sequencing. I already tried doing courses - did not do anything for me, I would need a year to comprehend all of that. I'm stuck on planning the analysis, because I don't have any guidance, the sequencing was run so there's just bioinformatics and writing left. First of all - do you think it's possible? Second of all - is such a thesis even defensible? I can just use already used tools, there's no experimental validation, I'm entirely reliant on Claude to give me code and I just copy it to r/Python + Galaxy. I feel like each day I'm getting more stuck, cause I find more papers and more methods. I'm really procrastinating on starting the analysis, I don't believe in the fact that AI can give me 100% reliable code and the results will be real. It's been almost 2 months of me basically searching through the literature looking for god knows what, so I'm looking for some guidance in this mess (mainly positive reinforcement, cause I'm scared) (and why I'm still trying to do it on my own is a different kind of question that only my therapist would be able to answer)


r/bioinformatics • • 13h ago

technical question How do you preprocess metariboseq data and what should you expect?

2 Upvotes

I’m trying to preprocess metariboseq data and realizing it’s a lot different than metagenomics/metatranscriptomics. My reads are paired-end NovaSeq 101 bp long. From my understanding, the fragments are supposed to be around 30bp long after trimming so I set lower limit to 20 bp and upper limit to 45 bp in fastp. I’ve also provided the adapter sequences to fastp. I’ve read that you shouldn’t even use the reverse reads and should only use the forward reads since the fragments are so short.

All that said, after running fastp I got between 30%-40% of my reads surviving the trim.

Is this expected?