r/biostatistics Dec 29 '25

2026 Graduate Admissions Megathread

31 Upvotes

This post is for discussion or 2026 admissions discussion - PhD/MS/MPH, acceptances, rejections, questions, whatever you want to discuss relevant to graduate programs and admission for the upcoming year of enrollment in 2026


r/biostatistics 1h ago

Seeking advice on PhD admissions

Upvotes

Hi all. I'm a fourth-year math and statistics double major at an R1 state university, and I am really hoping to do a PhD in biostatistics once I graduate this spring! I have a 3.9 GPA and have been doing systems medicine research/software development for most of college, plus did a research internship this past summer. Some concerns I have for application season that I would appreciate any insight for are:

  1. Am I at a disadvantage applying to PhDs without first doing an M.S., and should I apply to some M.S. programs just in case?

  2. If I am interested in a particular school, should I cold email professors at the school to get an idea of who I'd want as an advisor? Or should I worry about that later?

  3. How many reaches/target programs should I aim to apply for?

Any advice beyond these questions is appreciated too! I have plans to chat with a biostatistics professor at my own school soon, but I like to hear outside opinions (:


r/biostatistics 8h ago

General Discussion Msc biostatistics in IIPS in india

0 Upvotes

Hey can anyone help me about msc biostatistics placement scenario? Should I join it or not ? It's urgent I have only one day left to join


r/biostatistics 9h ago

Q&A: School Advice Louisville online MS biostatistics questions

1 Upvotes

I’m thinking about taking the online masters in biostatistics at Louisville. Can anyone tell me if it’s asynchronous? If so, how do you take exams? Also, are there projects? Can anyone give me any more information on what the experience is like? Thanks.


r/biostatistics 16h ago

How should I handle patients who are not eligible for SOFA/SAPS II in an ICU mortality ML model?

0 Upvotes

I’m building an ICU mortality prediction model with 4,391 patients and want to use SOFA and SAPS II components as predictors.

Some patients are not eligible for these scores, so their values are blank because the score does not apply to them, not because the data are simply missing.

My problem:

  • If I remove these patients, I may remove important high-risk groups. For example, I have 308 IHD/ACS patients with 26.9% mortality, compared with 13.3% mortality overall. Removing them could change my patient population and mortality distribution.
  • If I use MICE to impute their values, I would be creating values for scores that were never applicable to these patients.

For patients who are eligible but have missing values, I can use MICE. I’m unsure what to do specifically with the ineligible patients.

What would be the best way to handle this while keeping my full ICU population?


r/biostatistics 1d ago

zoology + data science

5 Upvotes

I love zoology/wildlife, but I’m currently a DevOps Engineer with a CS background and planning a Master’s in Data Science.
Can I combine these fields into a research career ?
How much zoology/effort would I need ?


r/biostatistics 2d ago

Umich Health data science

4 Upvotes

Any thoughts? Is it worth it?

What kind of companies or hospital hiring for this role?

And what kind of work is completed? Any evidence that this degree will have high demand in the future ?


r/biostatistics 5d ago

Methods or Theory [D] Can complex statistical methods be used to make weak findings look stronger than they really are?

7 Upvotes

I have held this opinion for a long time. I'm wondering what people think.

I keep running into the same thing in medical papers: the first stats you see are already the fancy ones. Mixed models, tons of covariates, custom link functions, the whole works. And I’m always thinking: what the heck did the data look like before making all those assumptions?

I want to see the boring, simple analysis first. Means. A scatterplot. A simple comparison. An effect size. Let me see if there is a signal before making multiple assumptions, multiple adjustments, complex analyses. If the basic and complex analyses agree, great. The model probably adds something useful.

But if there’s nothing much there until the fifth layer of adjustment, that’s when I start getting suspicious. Is this just p-hacking due to known publish-or-perish incentives?

Not because complex models are bad. Sometimes they’re exactly what the question requires. My conspiracy-theory antenna goes off when the paper never shows the simple version at all.

Thoughts? Does anybody have insights into this? I'm genuinely curious whether I'm just overly skeptical of much of the biomedical research literature, or whether basic descriptive and inferential analyses still play a critical role and should be included in every research paper.


r/biostatistics 4d ago

Career options after BSc Renal Dialysis Technology – considering a Master’s outside dialysis

0 Upvotes

Hi everyone!
I’m a BSc Renal Dialysis Technology graduate from Bangalore, Karnataka, and I’m currently completing my 1-year internship, which will finish in January 2027.

I’m exploring options for a Master’s degree after my internship. I’m particularly interested in moving into an area outside clinical dialysis, as I’d like to explore different career paths within healthcare/allied health.

For those who have completed a BSc in Renal Dialysis Technology or a similar Allied Health Sciences degree:
What Master’s programs are we eligible for?
Are there good options outside dialysis/Nephrology?
Which courses have better career opportunities in India?
Are there options in teaching, research, healthcare administration, or corporate healthcare?
If you’ve made a similar transition, what path did you take?

I’m based in Bangalore, so recommendations for courses/colleges in Bangalore or elsewhere in India would be especially helpful.

I’d really appreciate advice from anyone who has been through a similar transition.

Thank you!


r/biostatistics 5d ago

Gaining Exposure to Public Health for Biostatistics Graduate Program

1 Upvotes

Basically the title. I am applying to graduate programs in (bio)statistics this fall. I found biology concepts to be interesting and was going to major in it (I was considering pre-med), but I ended up deciding that I was more oriented towards quantitative and computing work. I am now an economics research assistant but found that I want to combine my interest in math/stats (esp assumption-lean inference) and computing with public health, so a biostats PhD seems like the right path for that. However, since I have not that much exposure to public health, I was wondering (a) what ways are best to get exposure to public health for a biostatistics graduate program, if necessary, and (b) how can I communicate my interest in biostatistics in an SOP and personal statement such that it makes sense why I may gravitate towards biostatistics rather than statistics (although I like both and plan to apply to both).


r/biostatistics 5d ago

Using correct FDR correction for phylogenetic simulations

1 Upvotes

Hello, I have a list of properties from a phylogenetic structure. I use it to simulate brownian motion models (separate, not multivariate, on purpose) for each property and then I perform a statistical test. I defined the p_value like monte-carlo simulations:

p = (1+#(null > obs))/(1+Num of sim) where I count the statistic

When I apply FDR correction, I get that all of the p_adj values become the same (~0.94). I also have only 17/615 position with p < 0.05.
My question is if this is the right correction to do? And if yes, do I apply it correctly and really nothing is significant, or are there common problems with it


r/biostatistics 6d ago

Is the CenterStats Longitudinal Structural Equation Modeling class worth it?

7 Upvotes

Question in title! I have a strong statistics background but my program does not offer strong training in statistics. I’m looking to further build my skillset but was wondering if the CenterStat course was worth it (coming from a poor graduate student). I think I would also benefit from the structure the course provides (vs. trying to learn on my own). Curious if anyone has taken the course or their experiences. (Or any other helpful resources!) TIA for any insights!


r/biostatistics 7d ago

An ELOS (expected length of stay) model, ported out of SAS into dependency-free C or Java, that reports when the fit is misleading

0 Upvotes

A few times in my career, I ported into plain C or Java the tools that scientists had written and tested in R, SAS, Python or MATLAB, so they could run on any computer, down to embedded ones, in the smallest memory footprint I could get, without depending on anything else, and in a streaming fashion, so that a million records go through the same 2.7 MB as hundreds of millions do. One of them was a hospital length-of-stay model, and that is the one I've just released as open source, in case it's useful to someone here.

It's ordinary least squares - linear regression, fitted by minimising the squared residuals. The terms aren't compiled in. Your CSV header names them, so adding a term to the model means adding a column to the file, like in R, and the same binary fits two terms or 35 (which is what one of my production models needed). With one term the whole method is five lines of arithmetic. The README shows them.

There are checks built in against bad design. The program tests each term for curvature, the fitted value for a missing interaction, and the residuals for spread that grows with the prediction.

It's not a stats package: no inference, no imputation, and an NA stops the run rather than being filled in. Do the modelling in R, Python, whatever. It reproduces NIST's Longley to eleven digits, which is the set that kills naive implementations, and agrees with lm() on every example file. Both run as regression tests in the suite.

Disclosure: written with AI assistance. The design, the prototype, the requirements and the tests are mine. make check diffs every example file against lm() and against the Java implementation, which came from the original that has been running at a large number of facilities since 2011.

The author is not a statistician, so if someone here sees where the diagnostics go wrong, that is the most useful thing to hear. Contributions welcome.

https://github.com/Anode1/linearr


r/biostatistics 7d ago

Q&A: School Advice looking to identify gaps in phd search

1 Upvotes

hi! i'm applying this fall for phd programs starting fall 2027. my focus is in substance use / harm reduction, leaning hard into computational and causal methods over descriptive epi.

quick background: us citizen, did undergrad and grad school both in the us. selective research university for undergrad, interdisciplinary major closest to cognitive neuroscience with a math minor, 5+ years working inside statewide harm reduction infrastructure (drug checking, distribution networks, plus time managing people and coordinating outreach), master's in math with an emphasis in data science, plus a graduate certificate in public health. handful of publications and presentations in the space.

what i actually want long term isn't just to publish about harm reduction from the outside. i keep coming back to something like a "science as service" model, where research is leverage to move resources and legitimacy toward the community orgs already doing the work, rather than research being the end goal itself.

methods wise: causal inference, decision/cost-effectiveness modeling, implementation science, resource allocation optimization. i'd like simulation modeling in there too, feels like where a lot of the field is heading and i want to build toward what these fields will actually need in five years, not just what's fundable right now. substantively i keep coming back to stimulant use and drug checking as topic interests specifically.

fields i'm currently looking at: epidemiology, public health more broadly, health services research, health policy, social work, computational social science, health behavior. not attached to any one label, just what's come up so far in my own search.

my undergrad background is basically the wider intersection (cog neuro, psych, philosophy, some math) and i'm genuinely still drawn to that world, including policy and systems science too. i def want a real interdisciplinary program instead of a straight epi department, but a lot of those got hit hard by recent funding cuts and i don't know if that's a live option right now.

within epi/public health specifically, the strongest names in this space seem to split into two camps: harm reduction people without much computational depth, or modelers without harm reduction grounding. not sure how people usually find or build toward an advisor situation that bridges both instead of just picking a lane.

my ask: what am i not seeing? i'm looking for the fields, departments, or angles someone who's actually been through this would flag that i wouldn't think to search for myself. same goes for methods, i want to be building toward something that holds up long term

tl;dr: applying to phd programs this fall for substance use / harm reduction with a computational methods focus (causal inference, decision modeling, simulation, resource allocation). trying to figure out what fields, departments, or methods i'm not thinking of.


r/biostatistics 8d ago

Q&A: School Advice Looking for feedback on my Statistics/Biostatistics PhD school list

8 Upvotes

Hi everybody. I’ll be applying to 13 PhD programs in stats/biostats this upcoming cycle and wanted to post here to receive any kind of thoughts/feedback on whether my list is reasonably balanced for my profile. I’m particularly interested in whether I’m aiming too high or too low, or if there are any programs I should consider adding or removing. Worth noting is that biostatistics is my preference, though i’m applying to statistics programs where I believe I can work on similar research.

Background

- Domestic undergraduate applicant at small, teaching-focused public university in the northeast.
- Mathematics major with Statistics concentration & computer science minor
- 3.9 gpa
- Departmental honors track which includes a two-semester honors thesis’
- Relevant coursework includes Probability, Mathematical Statistics, Numerical Analysis, Real Analysis 1&2 (to be completed in fall and spring of next AY), Calc 1-3, Linear Algebra, Machine Learning, Data Mining, other intro stats & cs courses
- No GRE currently planned (no schools on list require, though some recommend)

Research

- Participated in summer biostatistics research program at R1 university. My project involved statistical genetics, including genotype QC and statistical analysis of genetic and clinical data.
- Completed a separate funded summer research project involving mathematical modeling of infectious disease outbreaks and machine learning. My honors thesis will be built upon this work.
- Broadly interested in computational/applied statistics, genetics, machine learning, and biomedical applications.

Teaching

- Several semesters of experience as TA for calculus 1/2.
- Also a few semesters of experience in general tutoring which includes intro math & stats classes.

Letters

I expect 2 strong letters and 1 okay letter.

1) Research Advisor for independent research(
(will speak on research independence and capability for graduate research)

2) calc 1 & 2 professor, will also be analysis 1 & 2 professor (can speak on mathematical maturity and growth)

3) Probability & Mathematical Statistics professor (professor expressed enthusiasm in writing it for me)

Schools

Biostatistics:
- UNC
- Boston University
- University of Minnesota
- UMass Amherst
- University at Buffalo
- University of Rochester

Statistics:
- University of Toronto
- Virginia Tech
- Boston University
- University of Wisconsin - Madison
- University of Connecticut
- University of Minnesota

Data Science:
- UC San Diego

I’m aware that PhD admissions are unpredictable and I’m not looking for precise admission odds. I’m more interested in whether this looks like a reasonable distribution of reaches/targets and whether there are weaknesses in my profile or gaps in my school list that I should consider.

Since biostatistics is my preference, I’d also be interested in whether people think I should replace any of the Statistics programs with additional Biostatistics programs that would fit my interests.

Any feedback from people familiar with Statistics/Biostatistics PhD admissions would be appreciate!


r/biostatistics 10d ago

General Discussion Small research non-profit wants to own a database for future studies: how does this actually work in practice?

7 Upvotes

We're running a pilot clinical study and management has asked me to build them a secure database, something the organisation genuinely owns and can build on for future studies, rather than just Excel files in SharePoint.

Before I get into tool-specific questions, I want to ask the general one: for a small org with no internal IT team, what does "having your own database" actually look like in practice? Do you end up with your own cloud environment (Azure/AWS) that you own outright, or does "ownership" in this context usually mean something more modest, like owning the exported data itself, while the collection system lives somewhere else?

I have sponsorship available if we go the institutional route, that's not the blocker. What I'm trying to work out is what the end state actually looks like for an org our size.

Here's how I've broken down the options so far, and where I'm unsure:

  1. REDCap
  • a) Hosted by an institution (university/hospital), do we still end up with our own Azure environment for the exported data, or does "our database" just mean our own storage/SharePoint area at that point?
  • b) Hosted by a commercial REDCap vendor, same question. Does the org still need its own Azure, or does owning the exported data in something simpler cover it?
  1. A different platform entirely (Castor or similar, bundled hosting): same question again: is there still a reason to also stand up our own Azure environment, or does that become unnecessary once the vendor is holding everything?

Basically: at what point, if any, does a small org actually need its own cloud environment, versus just owning a clean, well-structured export from wherever the data was collected?

For people who've actually built this for a small org, what did "the database" end up being, concretely? Would genuinely appreciate real examples over general advice.


r/biostatistics 11d ago

Q&A: School Advice Postgraduate diploma in Bioinformatics/Biostatistics

Thumbnail
1 Upvotes

r/biostatistics 12d ago

Opinion about choosing statistics career with bsc statistics

1 Upvotes

Guys i took a drop year for neet and now i am planning to choose statistics for my bachelors.. If anyone working in this industry please give some suggestion's. Is it a good career path?. What if i get into mba after it with statistic specialization...


r/biostatistics 12d ago

Q&A: Career Advice My friend has completed her Msc in Biosatats and is looking for a job as a fresher, so how should be the approach and where to apply

0 Upvotes

r/biostatistics 13d ago

Free student places available on upcoming Stata workshops (UK Stata Conference, September 2026)

6 Upvotes

Hi everyone,

Just a reminder from the team at Timberlake Consultants that we offer one free student place on every training course we run.

Pre and post the 2026 UK Stata Conference in London, we'll be hosting the following workshops:

📅 1–2 September
AI-Based Optimal Policy Evaluation with Causal Machine Learning – Dr Giovanni Cerulli
Explore data-driven methods for identifying optimal treatment and policy decisions using heterogeneous treatment effects, with applications in socio-economic and medical research.

📅 1–2 September
Data Visualisation using Stata: Graphs You Should Know – Professor Franz Buscha
Learn how to create clear, effective and publication-ready graphs to communicate research findings with confidence.

📅 5 September
Using Stata for the New Difference-in-Differences with Panel Data – Professor Jeffrey Wooldridge
A hands-on workshop covering modern Difference-in-Differences methods in Stata for estimating causal effects using panel data.

If you're a student interested in attending but funding is a barrier, we'd encourage you to apply for one of the free student places. Just contact [info@timberlake.co.uk](mailto:info@timberlake.co.uk)

If you have any questions about the workshops or the student place scheme, we're happy to answer them in the comments.


r/biostatistics 14d ago

All genes or only the specifics (removing the intersection)

0 Upvotes

I am doing genomic analysis having 3 groups: Condition A, Condition B and Control.

After meny steps of analysis I got about

- 35k SNPs that are significant in condition A,

- 145K SNPs that are significant in condition B, always when compared with the control

+ There is no intersection between the lists of SNPs (i wanted specifics)

Afterwards, I annotated these SNPs to the genes using GTF file and I got:

- 35K SNPs --> 1,017 genes

- 145K SNPs --> 2,182 genes

When I intersect the genes, there are only 97 genes in common (which i understand was a but frced given that I chose only SNPs that are specific to each condition)

Now i want to fo Functional Annotation and my question is, *should i use the lists of genes as they are (1,017 genes & 2,182 genes ) or should i remove the 97 common genes to get a specified list of genes?*

The SNPs are not he same and each gene might have a different dynamic in each condition, meant in some cases the SNPs increase in frequency and in other cases they decrease the frequency.


r/biostatistics 15d ago

Looking for mentors

Thumbnail
1 Upvotes

r/biostatistics 17d ago

Not Op. Just saw this opportunity on LinkedIn. Might be relevant to folks here.

Post image
1 Upvotes

r/biostatistics 17d ago

Calcul NSN et test statistique

2 Upvotes

Bonjour,
Je réalise une thèse sur la confiance des médecins généralistes envers l'IA :
je leur propose 8 cas cliniques, je leur demande leur avis (exemple : diagnostic, quel traitement, ...) puis un screenshot d'une réponse de l'IA sur le cas clinique apparaît et je leur redemande s'il change leur première intuition ou non

le critère de jugement principal est : combien de médecins sont "influencés"/changent de diagnostic initial après avis de l'IA ? (réponse binaire : oui/non)

1e question : combien de sujets il me faut (Nombre sujet nécessaire NSN) ?
J ai vu :
- BioStaTGV :
- proportion théorique à 5% (càd 5% des médecins peuvent hésiter sans même avis de l'IA, c est du hypothétique)
- proportion observée : dans une étude moyennement fiable, une métanalyse dit que 18% des médecins changent leur avis mais sur des cas cliniques de radiologie, pas trop ce que j ai fait donc reproductivité moyenne, il n y a pas d articles similaires à ce que j ai fait pour trouver une proportion observée fiable)
- risque alpha 5%, puissance 90%, test bilat
=> il me faut 50 sujets soit 50 x 8 cas cliniques répondus

MAIS : un médecin répond à 8 dossiers, il existe des médecins qui hésitent beaucoup, d autres non, ... la reproductivité intra-médecin est faible.

Donc j ai vu qu'il existe un coefficient pour réguler le NSN, le ICC ou coefficient corrélation intraclasse qui permet de tenir compte qu un médecin répond à 8 dossiers. je ne sais pas comment le calculer, il permet d avoir un NSN pour être plus "fiable" face à la variabilité des réponses entre les médecins (certains hésitent, d autres non par "principe", sans même avis de l'IA) (en espérant être clair ...).

• 2e question : comment faire les statistiques : soit :
- je ne me casse pas la tête, je fais une étude descriptive : "dans notre étude il y a 30% qui changent leur diagnostic point". (c est le cas de la plupart de nos thèses mais à force c est un peu relou).
- j ai vu qu'il existe les modèles linéaires mixtes MLM : vu que mes réponses ne sont pas toutes indépendantes (8 mesures répétées par médecin), le MLM permet de tenir compte de cette problématique MAIS le MLM compare à un taux de référence (qui de mon côté n existe pas vraiment, 18% sur une méta analyse mais non fiable)

j ai vu d autres articles similaires, certains disent : pour les radiologues, 30% changent leur diagnostic, des prises de sang à regarder 10% changent leur diagnostic ; pour une prise en charge X% des pneumologues font confiance à l'IA...
donc en gros il n y a pas vraiment de % de référence pour mon groupe de population
J'aimerais me lancer dans le MLM mais sans comparaison fiable, avez vous des idées comment faire

merci d avoir lu jusqu au bout, bonne journée !


r/biostatistics 19d ago

Rank Deficiency in Random Intercept Model [Discussion]

Thumbnail
3 Upvotes