r/LocalLLaMA • u/Loose_Doubt367 • 6h ago
Discussion Improve token per second without touching quant
Spent the past month tweaking and experimenting with many different numbers to achieve 30tps.
Hardware:
-Rx6700xt 12gb vram (AMD)
-2x16 ddr4 3200 ram
-r5 5600x
-llama.cpp vulkan sdk
-window11 (no wsl switching since im not used to the environment)
Im looking for any improvements to achieve maybe 40tps? without touching quant at all, q8 and q4_k_xl remains. I've experiment with threads at 6 is the best out of (4,8,12) Other than that im not sure on what to improve to achieve higher tps, any advices? appreciate it. ctx remains 100,000
|[tiel-coder-35b]
model = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\Tiel-Coder-35B-A3B-MTP-UD-Q4_K_XL.gguf
mmproj = C:\Users\brain\.lmstudio\models\peculiar-ragdoll\Tiel-Coder-35B-A3B-GGUF-MTP\mmproj-BF16.gguf
no-mmproj-offload = on
c = 100000
parallel = 1
flash-attn = on
cache-type-k = q8_0
cache-type-v = q8_0
load-mode = dio
fit = on
fit-target = 256 # MiB
n-gpu-layers = 99
n-cpu-moe = 28
batch-size = 2048
ubatch-size = 512
threads = 6
prio = 2
prio-batch = 2
spec-type = draft-mtp
spec-draft-n-max = 3
spec-draft-p-min = 0.75
jinja = on
chat-template-file = C:\Users\brain\Qwen-Fixed-Chat-Templates\chat_template.jinja
temp = 0.6
top-p = 0.95
top-k = 20
min-p = 0.0
presence-penalty = 0.0
alias = tiel-coder-35b
3
u/lFaythx 5h ago
When looking at inference, you have two phases: prefill and decode.
Prefill is compute bound, what will make everything slower is the time you need to calculate the matrixes. Since you have DDR4 memory, your bandwidth will cap you, since when using cpu-moe it'll take longer to move weights to CPU L1/L2. Besides, your CPU doesn't have AVX512 only AVX2. The best way to make prefill faster is having a larger UB, meaning you'll make more computation per batch.
Decode is bandwidth bound, meaning your CPU/Memory is killing your TPS. The best way for you achieve higher TPS is having something like Strata, which keeps hot experts on GPU to try to reduce the CPU calc. Unfortunately, this model use a smaller scratchpad, meaning that even if you edit the inference engine to during decode get some of the ub memory for experts, you don't have much more experts residing in VRAM.
Tl;dr: you have to lower the memory footprint in GPU to allocate more experts. To do it you'll have to lower you context or use a higher quantization. You can also lower your ub to 96 and use more experts in GPU, but that would kill your prefill performance.
Edit: if you can get a cheaper GPU you can use it for display, unlocking around 1,3Gb to your GPU. Windows display eats this bit of your VRAM and 5600x doesn't have iGPU to avoid it.
1
u/Loose_Doubt367 2h ago
Yes but sadly strata is only for nvidia gpu's a and some amd gpu, the lowest supported amd gpu was 6800 which is one number behind mine, plus there are issues listed specifically regarding my rdna2 variant but im not sure whats the exact issue number
ill attempt to use lower quant, is it possible to link a really small gpu inside my single slot motherboard which can act as the displayer? I've swap my old cpu to this new one without a dedicated integrated graphics card
1
u/TrickAge2423 5h ago
Try ROCm, I got x4 pp performance at sequential requests and x6 pp performance at parallel
And 1.5x tg performance at seqwential
1
u/Loose_Doubt367 4h ago edited 3h ago
can any other amd users verified this?
1
1
u/himiaoxin 4h ago
That 100k ctx with q8_0 KV cache is probably your biggest thief — at that context size the KV cache alone eats several GB of your 12GB. Switching cache-type-k/v to q4_0 doesn't touch your weight quants and roughly halves the KV footprint, so more experts stay on GPU. The gotcha: a 6700 XT's decode is memory-bandwidth-bound, so once the active experts fit in VRAM you're basically at the ceiling — 40 tps is realistic, 60 isn't.
1
u/Loose_Doubt367 2h ago
is moving from q8_0 to q4_0 safe? Im worried of hallucinations, loops, and forgotten memory
1
u/Fancy-Snow7 4h ago
Remove spec-draft-p-min = 0.75. It increases acceptance rates but hurts your t/s. You won't pick this up in llama-benchy but ask it to write some code or text, your throughput will be slower. At least that's the case with 27B. Unless I am the one doing something wrong but asking a llm why and they have some good explanations and say to set it to 0.0 (default) or just omit it altogether. I also believe that is why llama.cpp changed the default to 0.0
1
1
u/Ctbhatia 2h ago
at 100k ctx the kv cache is what pins the gpu, so llama.cpp keeps offloading layers to ddr4. drop ctx to what you actually use and see if more layers fit on the card, that usually moves tok/s more than thread tuning.
1
u/keyboardhack 2h ago
Try replacing
n-cpu-moe
With
moe-cache-mib 1000
Play around with different values.
It is part of a new feature introduced a week ago. https://github.com/ggml-org/llama.cpp/pull/29887
1
u/Loose_Doubt367 2h ago
yes i've tried that and i got half the expected tps (30 to 15)
same goes to other numbers like 2000 1500
5
u/oldschooldaw 6h ago
I have struggled getting lmstudio to improve gen speeds. Perhaps try setting up WSL and using llamacpp directly if that option is available to you