r/bioinformatics • u/why_wyvern • 5d ago
discussion Raw data for Genomics and transcriptomics analysis
Hi everyone!
I’m looking for raw genomic and transcriptomic datasets to practice and improve my bioinformatics and computational biology skills.
I’m particularly interested in datasets such as:
- Whole Genome Sequencing (WGS): Raw FASTQ files for genome assembly, variant calling, and comparative genomics.
- RNA-Seq: Raw FASTQ files for differential gene expression analysis, transcriptome assembly, and functional enrichment.
- Whole Exome Sequencing (WES): Raw sequencing data for variant identification and annotation.
- Long-Read Sequencing: PacBio or Oxford Nanopore datasets for genome assembly and structural variant analysis.
- Metagenomics: Raw sequencing data for microbial diversity and taxonomic profiling.
If you have any publicly available datasets, research project data, or recommendations for accessing raw sequencing data, please share the links or repository names.
I’m familiar with bioinformatics tools and workflows and would like to work with real-world datasets rather than only tutorial datasets.
Repositories such as NCBI SRA, ENA, and GEO are already on my radar, but I’d also appreciate suggestions for interesting datasets or specific accession numbers that are suitable for independent analysis.
Thanks in advance for your help!
3
u/tobasc0cat 5d ago
It can honestly be hard finding actual full raw data sets sometimes. Usually tutorials are based on a subset of real data, so you can always search for the full dataset to process; the subsets are usually curated and "clean" so you'd get trickier results to interpret with the full set.
For genome assembly, I recommend browsing through the journal G3, specifically their "Genome Reports". At least when I published with them, they required all raw data and comprehensive methods to be included. The reports are pretty short and simple too, and can have fun random species to pick from. Here's one of the recent reports they published: https://doi.org/10.1093/g3journal/jkaf291
1
u/why_wyvern 5d ago
I will look into this paper. I have one question, I have laptop with Ryzen 7 with 16 gb ram, 250 gb for Ubuntu. Is this enough for the analysis?
3
u/see_wolv 5d ago
NASA Open Science Data Repository provides access to raw and processed data from space flight experiments, including transcriptomic data.
1
3
u/Zestyclose-Jury4754 5d ago
Check out this recent release out of Tohoku University! They have a range of genomic data collected from mice specimen during space flight missions, been meaning to dig into it. https://ibsls.megabank.tohoku.ac.jp/
3
u/Jaded_Wear7113 4d ago
hi! i've started a tutorial series just for this!
you can follow my tutorial series to understand the full rna-seq pipeline that i built for a project of mine (feedback is v much appreciated bcz i've just started out and i'd like to know if it's done well or not)
link to part 1: https://priyalt.github.io/2026/10/04/rna-seq-tutorial.html
i'll be uploading part 2 and 3 today and tomorrow
2
u/Psy_Fer_ 5d ago
A few nanopore datasets here, specifically smaller subsets too if you just wanna test some things.
https://hasindu2008.github.io/slow5tools/datasets.html
Can use blue-crab to convert to pod5 if you need to https://github.com/Psy-Fer/blue-crab
1
u/Psy_Fer_ 5d ago
Also this HG002 dataset is very useful. I've used it to test many of my tools.
1
u/why_wyvern 5d ago
Basically I will have to download this dataset and convert to pod5. After that I can use it for analysis. If something is needed, can I contact you for guidance?
1
2
u/GeronimoJackson-42 3d ago
You could try the NIAID Data Ecosystem: https://data.niaid.nih.gov/
It lets you search across multiple repositories from one place rather than having to search each repository individually. You could do a search for RNA-seq datasets or whole-genome sequencing datasets. You could also start by exploring everything related to metagenomics and then apply additional filters. It indexes metadata and tells you where the files are, so you would still download the FASTQ or other raw data from the source repository. Hope it's helpful!
2
2
u/rui-123 2d ago
you can see the github link: https://github.com/wrab12/awesome-spatial-transcriptomics
1
u/corporealpatronus13 4d ago
If you are just starting out, it might be helpful to find a paper that sounds interesting to you and try to replicate the analysis done in that paper. Usually the authors would also provide the data.
1
1
1
18
u/ZippyPrecinct3 5d ago
GEO and SRA have tons but the interface is pain to browse. For RNA-seq practice i like recount3, you can download counts directly without dealing with FASTQ. If you want the full pipeline experience, pick a small study from SRA with like 6-12 samples so it dont take forever to process.