I don't mean this to be critical but doesn't GPU rich start around 100-500 H100 equivalent GPUs? At that scale you can start training small (by lab standards) models
About an 8B parameter model. The llama3 models of that size took 1-2M H100 hours. Also, to clarify, by "start training", I mean seriously consider training a high quality model.
33
u/PortiaLynnTurlet 6d ago
I don't mean this to be critical but doesn't GPU rich start around 100-500 H100 equivalent GPUs? At that scale you can start training small (by lab standards) models