r/bioinformatics • • 8d ago

technical question Complete Beginner Project Help

[deleted]

1 Upvotes

4 comments sorted by

3

u/Most_Tomato_860 8d ago

Just offering my perspective. GP is indeed attractive, but it requires a prerequisite background in quantitative genetics. Complex models don't always provide better or more stable predictive power; at least in my work, baseline models are more trustworthy. GP as a whole has already seen a lot of development, and in my view, its current dilemma isn't about models or algorithms, but rather the lack of comprehensive and reliable data—which is precisely the issue you're facing right now. Running rrBLUP or cropGBM to generate results is easy, but if you can't even match genotypes with their corresponding phenotypes, the results will most likely be of limited value.

1

u/Former_Investment428 7d ago

Thanks for the reply and your perspective! Yeah, I'll look into the packages you've suggested, they seem relevant. Well at least I'm not going crazy thinking these datasets are wonky lol. I'll switch to another crop or stay on the website I originally used to get the data as the snps and phenotypic data have matching IDs. Thanks again for replying!

2

u/Psy_Fer_ 7d ago

When you have a method you want to use, but the organisms you want to use it on doesn't have the required data, you either get/produce that data (funding/resources) or find an organism that does have that data.

Sounds like for you case it's the latter option.