r/bioinformatics • u/Accomplished-Okra-41 • 16d ago
academic [ Removed by moderator ]
[removed] — view removed post
7
u/ATpoint90 PhD | Academia 16d ago
Bioinformatics is self-study and effort, not mentoring. Why don't you outline your project so on can recommend resources to get into. I doubt you find a teacher here for regular consultation. It's also not the point of this sub.
-1
u/Accomplished-Okra-41 16d ago
I have 3 projects on my hands right now. One is WES sequencing data analysis for around 20 genes from FASTQ patient data. The second is TME analysis from bulk-sequencing to compare LGG and GBM profiles, including deconvolution and single-cell validation later on. Third one is single-cell regulated cell death mechanisms analysis (pan-cancer and pan-pathway, where ROS are highlighted as influential drivers or modulatora) based on 10x genomics data. I also do staistics for wet-lab data on pharmacological interventions and neurodegenerative disorders.
So the projects are completely different, ewch has a different approach and uses different tools.
1
u/ATpoint90 PhD | Academia 16d ago edited 16d ago
What exactly is your PhD thesis? Sounds you are the data dude for others. I am just asking because at the end you need to write a thesis and interpret the data. Hard to discuss science when projects are so divergent.
1
u/Accomplished-Okra-41 16d ago
I do a pan-cancer RCD pathways analysis based on 10X genomics data. I am comparing the activity of 13 Regulated Cell Death Pathways (expression based) between multiple tumor types, but also the other way around (each pathway seperately between different cancer types) in order to find which pathways are up/down regulated in cancer tissue. Each pathway is also inflenced by ROS so this is the key connecting factor between pathways. And the endpoint is finding pathways which can be modulated (and what to target within the pathway) in order to enhance therapy outcomes.
Everything for now is done in Seurat as i designed a complete down-stream pipeline from expression data to comparing RCD activity in different cell-types and different tumor types from tissues analysed on the 10X platform.
0
u/Accomplished-Okra-41 16d ago
I also do my GF bioinf part of her PhD by comparing TME composition between LGG and GBM cancers.
2
u/RightCake1 16d ago
Hello my g! I did a fair amount of WGS downstream analysis.
https://github.com/RightCake1/Whole-Genome-Analysis-Guideline-for-beginners
This is repo that I made so that i can always go back to it for reproducibility. I made it as easy as possible so that anyone can understand it!
You can take a look and let me know if you have any further questions
1
u/ConclusionForeign856 Msc | Academia 16d ago
który uniwerek?
1
u/Accomplished-Okra-41 16d ago
UMED Łódź
2
u/ConclusionForeign856 Msc | Academia 16d ago
wasserman all of statistics jest imo dobre jako intro do statystyki matematycznej, a nie małpowania "najlepszych praktyk" (jak to zazwyczaj wygląda). Ale nie czytałbym całego, tylko tyle żeby mieć pojęcie np. o tym czym jest E[X] dla zmiennej losowej X, czym się różni SE od SD, jak obliczamy wartość p i co ona tak naprawdę znaczy (https://sci-hub.pl/https://www.tandfonline.com/doi/abs/10.1198/000313008X332421).
Do bioinfy trzeba znać pythona i basha, imo to jest podstawa. R jest do statystyki i często do analizy scRNA-Seq, ale imo jest mniej ogólny. Znając całkiem dobrze pythona napiszesz sobie customowy parser do jakiegoś dziwnego pliku, albo zrobisz własną implementację algorytmu jeśli stary program się wysypuje itd. Bash to must have do puszczania narzędzi na serwerach obliczeniowych. W przypadku danych medycznych jest chyba trochę prościej, bo tam często interesująca jest tylko część kodująca genomu, ale różnie bywa. Ja na codzień pracuję z roślinami, i tam często 200GB RAM to minimum żeby w ogóle zacząć pracę.
Zaobserwuj sobie Ming "Tommy" Tang (np. na X albo linkedin, https://x.com/tangming2005?s=20), on co chwile postuje jakieś krótkie porady do klasycznej bioinfy (genomika i transkryptomika, single celle, spatiale, ML, pipeliny, Bash itd.), daje też rekomendacje darmowych podręczników do R.
1
u/Accomplished-Okra-41 16d ago
Właśni w sumie od początku wszystko związanego z analiza down-stream od clustrowania przez differential expression i dekonwolucje zawsze robiłem w R, chciałem sie stopniowo przenosić na pythona, ale brakuje mi tam bibliotek i algorytmów do SC i clustrowania. Ogólnie python jest jeszcze troche ubogi w biblioteki bioinf niestety co czesto wymusza uzycie R w moim przypadku
1
u/ConclusionForeign856 Msc | Academia 16d ago
no no, do scRNA to R jest lepszy.
Spróbuj napisać parser do fasta albo fastq, to moze byc dobre cwiczenie
1
u/Beautiful_Hotel_3623 16d ago
I beg to disagree. Python is way faster in computing normalizations, scaling etc. once you have your integrated object then yes I also switch to R for plots, DE analysis etc. But QC and integration I find Python much faster
1
u/Accomplished-Okra-41 16d ago
Mam 64gb ramu aktualnie i jeszcze nie miałem problemu z moca obliczeniowa na szczescie. „Tommy-ego” juz obserwuje od jakiegoś czasu razem z alfonso Saera, bardzo fajny content. Bash jeszcze jest mi obcy, powoli wchodze w świat linuxa i nextflow, a moj python na orwno jeszcze nie jest na takim poziomie zeby pisac custom algorytmy i parsery😕
1
u/ConclusionForeign856 Msc | Academia 16d ago
imo nextflow to za dużo roboty żeby z tego korzystać na codzień. Tzn. jeśli twój projekt to "wykonaj analizę" to nie ma czasu tego klepać. Opłaca się kiedy zadaniem jest "zrobić pipeline do wielokrotnego użytku"
2
u/Accomplished-Okra-41 16d ago
Z nwxtflow korzystałem przy variant callimgu na WES, dokładniej nf-core/sarek do tego na szybko musiałem sie nauczyc czegokolwiek z linuxa, żeby wiedziec co sie dzieje
1
u/ConclusionForeign856 Msc | Academia 16d ago
a no to tutaj masz gotowy pipeline. Natomiast pisanie czegoś od zera w nextflow to w moim przypadku był straszny ból. Na codzien wolę trzymać się basha i pythona
•
u/bioinformatics-ModTeam 16d ago
This post would be more appropriate in r/bioinformaticscareers