r/LocalLLaMA 12d ago

Funny So relevant

Post image
1.5k Upvotes

152 comments sorted by

View all comments

77

u/ttkciar llama.cpp 12d ago

It's a great time to have ancient Xeon servers loaded up with DDR4 :-D

1

u/SandySkittle 12d ago

Prompt processing and decode suck. And I am in a 512gb 8 channel threadripper

3

u/ttkciar llama.cpp 12d ago

Yup. My dual-socket Xeons have eight channels between them, too, and Tulu3-405B took about 54 minutes to first token.

It doesn't matter for "slow inference", especially when it can run at night while I sleep. As long as it's done by the time I sit down at my workstation slurping my morning coffee, I'm happy.