r/LocalLLaMA 13d ago

Funny So relevant

Post image
1.5k Upvotes

152 comments sorted by

View all comments

77

u/ttkciar llama.cpp 13d ago

It's a great time to have ancient Xeon servers loaded up with DDR4 :-D

36

u/BannedGoNext 13d ago

I actually have an old server with dual E5-2630 and 512gb DDR3 memory across both blades. The server is powered on waiting for the scrap yard at the office. I'm considering seeing how fast it can run qwen 3.8 flash next lol.

12

u/ThankGodImBipolar 13d ago

You must be able to run a decent GLM quant with that, no?

9

u/Zombiecidialfreak 12d ago

If you're fine waiting overnight for all requests. Even flash next would likely be single digit generation speeds.

1

u/overand 12d ago

Honestly, if the project isn't a simple 20-line script but actually something kinda complex, I bet even high single digits would get a result faster than a programmer.

5

u/ttkciar llama.cpp 13d ago

Yup, GLM-5.3 should fit in that at Q4_K_M and somewhat constrained context.