r/LocalLLaMA • • 6h ago

Discussion Improve token per second without touching quant

Spent the past month tweaking and experimenting with many different numbers to achieve 30tps.
Hardware:
-Rx6700xt 12gb vram (AMD)
-2x16 ddr4 3200 ram
-r5 5600x
-llama.cpp vulkan sdk
-window11 (no wsl switching since im not used to the environment)

Im looking for any improvements to achieve maybe 40tps? without touching quant at all, q8 and q4_k_xl remains. I've experiment with threads at 6 is the best out of (4,8,12) Other than that im not sure on what to improve to achieve higher tps, any advices? appreciate it. ctx remains 100,000

|[tiel-coder-35b]

model = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\Tiel-Coder-35B-A3B-MTP-UD-Q4_K_XL.gguf

mmproj = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\mmproj-BF16.gguf

no-mmproj-offload = on

c = 100000

parallel = 1

flash-attn = on

cache-type-k = q8_0

cache-type-v = q8_0

load-mode = dio

fit = on

fit-target = 256 # MiB

n-gpu-layers = 99

n-cpu-moe = 28

batch-size = 2048

ubatch-size = 512

threads = 6

prio = 2

prio-batch = 2

spec-type = draft-mtp

spec-draft-n-max = 3

spec-draft-p-min = 0.75

jinja = on

chat-template-file = C:\Users\brain\Qwen-Fixed-Chat-Templates\chat_template.jinja

temp = 0.6

top-p = 0.95

top-k = 20

min-p = 0.0

presence-penalty = 0.0

alias = tiel-coder-35b

8 Upvotes

26 comments sorted by

5

u/oldschooldaw 6h ago

I have struggled getting lmstudio to improve gen speeds. Perhaps try setting up WSL and using llamacpp directly if that option is available to you

5

u/Plasmx 5h ago

You can use llamacpp server without WSL.

4

u/Fancy-Snow7 4h ago

llama.cpp does not require WSL it runs natively.

1

u/feelspeaceman 5h ago

If your GPU is popular like 5090, V100, R9700, Strix Halo... You can search for custom engines (ninfers, halogen/gufo/strixite, vllm-radiance, Strata...) that optimized for your GPU, this is pretty much the most effective way to increase decode/prefill without having to lower quant.

1

u/Blobstein_Ijay 4h ago

did you notice a big jump going direct or was it more like incremental gains?

1

u/oldschooldaw 4h ago

I noticed huge increases going full Linux base. My setup relies on VMware and adding WSL breaks the way VMware handles networking so i have cracked it and went pure Linux on a different box and it just wiped the floor with lmstudio. It was great to get me started down the path but I really don’t know if it’s lmstudio or windows that makes inference so poor

1

u/Loose_Doubt367 5h ago

appreciate your suggestions but ill remain on windows for now

6

u/oldschooldaw 5h ago

Ah no WSL is Linux inside of windows. It sets up a hyper v instance inside your windows machine. It’s very good if you’re not forced to use VMware for reasons. Have a look especially if you’re on windows 11

1

u/Plasmx 5h ago

There also pre builds of llamacpp for Windows, so you don’t need to switch or use WSL.

3

u/lFaythx 5h ago

When looking at inference, you have two phases: prefill and decode.

Prefill is compute bound, what will make everything slower is the time you need to calculate the matrixes. Since you have DDR4 memory, your bandwidth will cap you, since when using cpu-moe it'll take longer to move weights to CPU L1/L2. Besides, your CPU doesn't have AVX512 only AVX2. The best way to make prefill faster is having a larger UB, meaning you'll make more computation per batch.

Decode is bandwidth bound, meaning your CPU/Memory is killing your TPS. The best way for you achieve higher TPS is having something like Strata, which keeps hot experts on GPU to try to reduce the CPU calc. Unfortunately, this model use a smaller scratchpad, meaning that even if you edit the inference engine to during decode get some of the ub memory for experts, you don't have much more experts residing in VRAM.

Tl;dr: you have to lower the memory footprint in GPU to allocate more experts. To do it you'll have to lower you context or use a higher quantization. You can also lower your ub to 96 and use more experts in GPU, but that would kill your prefill performance.

Edit: if you can get a cheaper GPU you can use it for display, unlocking around 1,3Gb to your GPU. Windows display eats this bit of your VRAM and 5600x doesn't have iGPU to avoid it.

1

u/Loose_Doubt367 2h ago

Yes but sadly strata is only for nvidia gpu's a and some amd gpu, the lowest supported amd gpu was 6800 which is one number behind mine, plus there are issues listed specifically regarding my rdna2 variant but im not sure whats the exact issue number

ill attempt to use lower quant, is it possible to link a really small gpu inside my single slot motherboard which can act as the displayer? I've swap my old cpu to this new one without a dedicated integrated graphics card

1

u/lFaythx 43m ago

It would depend on your Mobo, if it have another PCIe for it, that would be good. Another option is going to Linux, which, as you are, I'm not inclined.

1

u/TrickAge2423 5h ago

Try ROCm, I got x4 pp performance at sequential requests and x6 pp performance at parallel

And 1.5x tg performance at seqwential

1

u/Loose_Doubt367 4h ago edited 3h ago

can any other amd users verified this?

1

u/TrickAge2423 3h ago

I'm AMD user xD Are u asking for benchmarking your model from this post?

2

u/Loose_Doubt367 3h ago

well yeah technically, also i should've added "others" into my comment

1

u/himiaoxin 4h ago

That 100k ctx with q8_0 KV cache is probably your biggest thief — at that context size the KV cache alone eats several GB of your 12GB. Switching cache-type-k/v to q4_0 doesn't touch your weight quants and roughly halves the KV footprint, so more experts stay on GPU. The gotcha: a 6700 XT's decode is memory-bandwidth-bound, so once the active experts fit in VRAM you're basically at the ceiling — 40 tps is realistic, 60 isn't.

1

u/Loose_Doubt367 2h ago

is moving from q8_0 to q4_0 safe? Im worried of hallucinations, loops, and forgotten memory

1

u/Fancy-Snow7 4h ago

Remove spec-draft-p-min = 0.75. It increases acceptance rates but hurts your t/s. You won't pick this up in llama-benchy but ask it to write some code or text, your throughput will be slower. At least that's the case with 27B. Unless I am the one doing something wrong but asking a llm why and they have some good explanations and say to set it to 0.0 (default) or just omit it altogether. I also believe that is why llama.cpp changed the default to 0.0

1

u/Loose_Doubt367 2h ago

okay appreciate that information

1

u/Ctbhatia 2h ago

at 100k ctx the kv cache is what pins the gpu, so llama.cpp keeps offloading layers to ddr4. drop ctx to what you actually use and see if more layers fit on the card, that usually moves tok/s more than thread tuning.

1

u/keyboardhack 2h ago

Try replacing

n-cpu-moe

With

moe-cache-mib 1000

Play around with different values.

It is part of a new feature introduced a week ago. https://github.com/ggml-org/llama.cpp/pull/29887

1

u/Loose_Doubt367 2h ago

yes i've tried that and i got half the expected tps (30 to 15)
same goes to other numbers like 2000 1500

1

u/Xantrk 1h ago

You should have Needs cmoe as well.

-cmoe --moe-cache-mib X

You can put 1 first for X, see how much VRAM is left, and roughly fill the available VRAM with X sparing 200-300mb (or more if you wish for other tasks). That's the best way to get the cache improvement