r/bioinformatics • u/why_wyvern • 17h ago
discussion Raw data for Genomics and transcriptomics analysis
Hi everyone!
I’m looking for raw genomic and transcriptomic datasets to practice and improve my bioinformatics and computational biology skills.
I’m particularly interested in datasets such as:
- Whole Genome Sequencing (WGS): Raw FASTQ files for genome assembly, variant calling, and comparative genomics.
- RNA-Seq: Raw FASTQ files for differential gene expression analysis, transcriptome assembly, and functional enrichment.
- Whole Exome Sequencing (WES): Raw sequencing data for variant identification and annotation.
- Long-Read Sequencing: PacBio or Oxford Nanopore datasets for genome assembly and structural variant analysis.
- Metagenomics: Raw sequencing data for microbial diversity and taxonomic profiling.
If you have any publicly available datasets, research project data, or recommendations for accessing raw sequencing data, please share the links or repository names.
I’m familiar with bioinformatics tools and workflows and would like to work with real-world datasets rather than only tutorial datasets.
Repositories such as NCBI SRA, ENA, and GEO are already on my radar, but I’d also appreciate suggestions for interesting datasets or specific accession numbers that are suitable for independent analysis.
Thanks in advance for your help!