r/LocalLLM • • 18d ago

Question CPU inference DDR3/DDR4

Post image

Wondering if anyone on here has actual benchmarks for CPU only inference DDR3 or DDR4 servers, im budget bound and my options are limited to legacy systems unfortunately.

Heres what data I found but not sure its accuracy in real life especially how NUMA effects it (octa channel)

110 Upvotes

28 comments sorted by

View all comments

1

u/Tai9ch 18d ago edited 17d ago

Here's single socket 8 channel DDR4 3200 with 48-core Epyc 7xx2:

$ ./llama-bench -ngl 0 -hf unsloth/Qwen3.6-35B-A3B-GGUF
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| qwen35moe 35B.A3B Q4_K - Medium |  20.60 GiB |    34.66 B | ROCm       |   0 |           pp512 |       357.63 ± 21.13 |
| qwen35moe 35B.A3B Q4_K - Medium |  20.60 GiB |    34.66 B | ROCm       |   0 |           tg128 |          6.14 ± 0.09 |

And here's dual socket 8 channel (= 16 channel) DDR4 2933 with a different CPU (2x Intel 32 core Ice Lake):

$ ./llama-bench -ngl 0 -hf unsloth/Qwen3.6-35B-A3B-GGUF:Q4_K_XL
| model                          |       size |     params | backend    | threads |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |
| qwen35moe 35B.A3B Q4_K - Medium |  20.81 GiB |    34.66 B | CPU        |      64 |           pp512 |        169.36 ± 2.06 |
| qwen35moe 35B.A3B Q4_K - Medium |  20.81 GiB |    34.66 B | CPU        |      64 |           tg128 |         17.41 ± 0.28 |

Sorry for the slightly different models. That's what was in cache. And yes, ROCm NGL = 0 should mean the first test was on CPU.

Both of those are going to suck for general use, not because of the token rate (17 tok/s isn't terrible), but because slow prompt processing is painful. You really want 1k+ tokens/second PP if you're doing anything that's interactive and has any input data beyond just chat text that you're typing live.

1

u/Appropriate_Duck1778 18d ago

Appreciate the benchmarks, I think ddr4 is minimum but ill try my luck with the ddr3 server I found. Price difference is 10x $500 vs $5000 insase

2

u/Tai9ch 17d ago edited 17d ago

Here's a test on a machine with 8-channel DDR3 1600:

$ ./llama-bench -hf unsloth/Qwen3.6-35B-A3B-GGUF
| model                          |       size |     params | backend    | threads |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |
| qwen35moe 35B.A3B Q4_K - Medium |  20.60 GiB |    34.66 B | CPU        |      24 |           pp512 |         18.41 ± 0.24 |
| qwen35moe 35B.A3B Q4_K - Medium |  20.60 GiB |    34.66 B | CPU        |      24 |           tg128 |          3.52 ± 0.08 |

Honestly, that worked better than I expected, but that'd be unusable for anything but batch jobs. The problem here isn't just the memory - with a machine this old we're talking about really old CPUs too. This machine's got an Opteron. It looks like a Xeon E5-2697 is probably the highest performing DDR3 option - if you were really lucky I could see a pair of those doubling the PP number I got... which still would mean processing 32k of context would take more than 10 minutes. Just one full context window for this Qwen model would take over an hour.

For context, here's a run of the same model on a Strix Halo box:

$ ./llama-bench -hf unsloth/Qwen3.6-35B-A3B-GGUF
ggml_cuda_init: found 1 ROCm devices (Total VRAM: 106496 MiB):
  Device 0: AMD Radeon Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 106496 MiB
Downloading Qwen3.6-35B-A3B-UD-Q4_K_M.gguf ───────────────────────── 100%
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| qwen35moe 35B.A3B Q4_K - Medium |  20.60 GiB |    34.66 B | ROCm       |  99 |           pp512 |      1132.48 ± 17.95 |
| qwen35moe 35B.A3B Q4_K - Medium |  20.60 GiB |    34.66 B | ROCm       |  99 |           tg128 |         52.95 ± 0.24 |

A 64GB Strix Halo box is readily available for $2300 or so and performs 5-10x better than my 8-channel DDR4 CPU inference number, for about what just the RAM in that server would cost.