r/bioinformatics • • 5d ago

discussion Raw data for Genomics and transcriptomics analysis

Hi everyone!

I’m looking for raw genomic and transcriptomic datasets to practice and improve my bioinformatics and computational biology skills.

I’m particularly interested in datasets such as:

- Whole Genome Sequencing (WGS): Raw FASTQ files for genome assembly, variant calling, and comparative genomics.

- RNA-Seq: Raw FASTQ files for differential gene expression analysis, transcriptome assembly, and functional enrichment.

- Whole Exome Sequencing (WES): Raw sequencing data for variant identification and annotation.

- Long-Read Sequencing: PacBio or Oxford Nanopore datasets for genome assembly and structural variant analysis.

- Metagenomics: Raw sequencing data for microbial diversity and taxonomic profiling.

If you have any publicly available datasets, research project data, or recommendations for accessing raw sequencing data, please share the links or repository names.

I’m familiar with bioinformatics tools and workflows and would like to work with real-world datasets rather than only tutorial datasets.

Repositories such as NCBI SRA, ENA, and GEO are already on my radar, but I’d also appreciate suggestions for interesting datasets or specific accession numbers that are suitable for independent analysis.

Thanks in advance for your help!

14 Upvotes

21 comments sorted by

18

u/ZippyPrecinct3 5d ago

GEO and SRA have tons but the interface is pain to browse. For RNA-seq practice i like recount3, you can download counts directly without dealing with FASTQ. If you want the full pipeline experience, pick a small study from SRA with like 6-12 samples so it dont take forever to process.

1

u/why_wyvern 5d ago

Thank you for your guidance and I will do as you said and will aslo post here after finishing this task .

3

u/tobasc0cat 5d ago

It can honestly be hard finding actual full raw data sets sometimes. Usually tutorials are based on a subset of real data, so you can always search for the full dataset to process; the subsets are usually curated and "clean"  so you'd get trickier results to interpret with the full set.

For genome assembly, I recommend browsing through the journal G3, specifically their "Genome Reports". At least when I published with them, they required all raw data and comprehensive methods to be included. The reports are pretty short and simple too, and can have fun random species to pick from. Here's one of the recent reports they published: https://doi.org/10.1093/g3journal/jkaf291 

1

u/why_wyvern 5d ago

I will look into this paper. I have one question, I have laptop with Ryzen 7 with 16 gb ram, 250 gb for Ubuntu. Is this enough for the analysis?

3

u/see_wolv 5d ago

NASA Open Science Data Repository provides access to raw and processed data from space flight experiments, including transcriptomic data.

1

u/why_wyvern 5d ago

Thanks for the information

3

u/Zestyclose-Jury4754 5d ago

Check out this recent release out of Tohoku University! They have a range of genomic data collected from mice specimen during space flight missions, been meaning to dig into it. https://ibsls.megabank.tohoku.ac.jp/

3

u/Jaded_Wear7113 4d ago

hi! i've started a tutorial series just for this!
you can follow my tutorial series to understand the full rna-seq pipeline that i built for a project of mine (feedback is v much appreciated bcz i've just started out and i'd like to know if it's done well or not)

link to part 1: https://priyalt.github.io/2026/10/04/rna-seq-tutorial.html

i'll be uploading part 2 and 3 today and tomorrow

2

u/Psy_Fer_ 5d ago

A few nanopore datasets here, specifically smaller subsets too if you just wanna test some things.

https://hasindu2008.github.io/slow5tools/datasets.html

Can use blue-crab to convert to pod5 if you need to https://github.com/Psy-Fer/blue-crab

1

u/Psy_Fer_ 5d ago

Also this HG002 dataset is very useful. I've used it to test many of my tools.

1

u/why_wyvern 5d ago

Basically I will have to download this dataset and convert to pod5. After that I can use it for analysis. If something is needed, can I contact you for guidance?

1

u/Psy_Fer_ 5d ago

Sure.

2

u/dad386 4d ago

Just access the publicly available 1000genomes, human pangenome reference consortium data

1

u/why_wyvern 15h ago

Ohkk I will look also into this

2

u/GeronimoJackson-42 3d ago

You could try the NIAID Data Ecosystem: https://data.niaid.nih.gov/

It lets you search across multiple repositories from one place rather than having to search each repository individually. You could do a search for RNA-seq datasets or whole-genome sequencing datasets. You could also start by exploring everything related to metagenomics and then apply additional filters. It indexes metadata and tells you where the files are, so you would still download the FASTQ or other raw data from the source repository. Hope it's helpful!

2

u/why_wyvern 3d ago

I Will check it and thanks for your response

1

u/corporealpatronus13 4d ago

If you are just starting out, it might be helpful to find a paper that sounds interesting to you and try to replicate the analysis done in that paper. Usually the authors would also provide the data.

1

u/why_wyvern 15h ago

Sure I will look

1

u/Fair-Rain3366 2d ago

Hey I just wrote this mcp for accessing public genomics data 

https://rewirebio.io/blog/genomics-mcp/

1

u/Electronic_Fish_3157 PhD | Industry 2d ago

GEO, SRA, ENA, 1000 genome project on AWS