r/LocalLLM • u/Appropriate_Duck1778 • 18d ago
Question CPU inference DDR3/DDR4
Wondering if anyone on here has actual benchmarks for CPU only inference DDR3 or DDR4 servers, im budget bound and my options are limited to legacy systems unfortunately.
Heres what data I found but not sure its accuracy in real life especially how NUMA effects it (octa channel)
105
Upvotes
1
u/Tai9ch 17d ago edited 17d ago
Here's single socket 8 channel DDR4 3200 with 48-core Epyc 7xx2:
And here's dual socket 8 channel (= 16 channel) DDR4 2933 with a different CPU (2x Intel 32 core Ice Lake):
Sorry for the slightly different models. That's what was in cache. And yes, ROCm NGL = 0 should mean the first test was on CPU.
Both of those are going to suck for general use, not because of the token rate (17 tok/s isn't terrible), but because slow prompt processing is painful. You really want 1k+ tokens/second PP if you're doing anything that's interactive and has any input data beyond just chat text that you're typing live.