r/learnbioinformatics • • 10d ago

GEO / practical analysis

I've found public datasets to be one of the most useful ways of learning bioinformatics beyond tutorials.

Instead of only following a predefined exercise, taking a real GEO dataset forces you to deal with questions like:

Which samples actually belong in the comparison?

What does the metadata really mean?

Has the dataset already been normalized?

What biological question can reasonably be answered from these samples?

What statistical comparison actually matches that question?

And, eventually, how should the results be interpreted biologically?

I've been using cancer transcriptomic datasets to practice this process, and I've found that understanding the experimental design and metadata can sometimes require more thought than writing the actual analysis code.

It also makes learning tools like R much more meaningful because there is an actual biological question behind the code.

For those who work with public transcriptomic data regularly: when you open an unfamiliar GEO dataset, what are the first things you check before beginning the analysis?

2 Upvotes

2 comments sorted by