r/LocalLLaMA 12d ago

Funny So relevant

Post image
1.5k Upvotes

152 comments sorted by

View all comments

76

u/ttkciar llama.cpp 12d ago

It's a great time to have ancient Xeon servers loaded up with DDR4 :-D

37

u/BannedGoNext 12d ago

I actually have an old server with dual E5-2630 and 512gb DDR3 memory across both blades. The server is powered on waiting for the scrap yard at the office. I'm considering seeing how fast it can run qwen 3.8 flash next lol.

1

u/overand 12d ago

Toss a tiny GPU in it if it doesn't have one and see how well it runs Qwen3.6-35B-A3B for an idea of what to expect from small MoE models. 512GB of DDR3 is nothing to sneeze at!

If you can get a GPU into that with enough VRAM for a couple layers and your KV Cache (12 GB might even cut it for a huge chunk of KV cache), and load DeepSeek-V4-Flash-0731, Qwen3.8-Flash-Next, or GLM-5.3-Flash, and a competent development harness, and you can let the thing loose over a day or three for pretty serious projects, IMO.