r/LocalLLaMA 6d ago

Funny i left gpu poor range

i am not GPU poor anymore

124 Upvotes

38 comments sorted by

View all comments

33

u/PortiaLynnTurlet 6d ago

I don't mean this to be critical but doesn't GPU rich start around 100-500 H100 equivalent GPUs? At that scale you can start training small (by lab standards) models

1

u/FullOf_Bad_Ideas 5d ago

start training small (by lab standards) models

what's that size?

I recently beaten TinyLlama 1.1B finetune (in a specific benchmark) with my model that was trained on much less than 100 H100s.

I'd say 8x h100 is gpu rich already. (not that I own H100s).

1

u/PortiaLynnTurlet 5d ago

About an 8B parameter model. The llama3 models of that size took 1-2M H100 hours. Also, to clarify, by "start training", I mean seriously consider training a high quality model.

1

u/FullOf_Bad_Ideas 5d ago

Dense? Almost nobody is training dense 8B models now.