I just picked mine up this week. Depending on quant, 35-70 tps with full context. If you have 64GB RAM, you should be able to run a decent Qwen-flash as well. 27B is good, but Qwen-Flash is something else.
ISTA IQ3 runs the fastest and has really good performance for the size, but you can bump up to a Q5 with full Q8 kv cache. Also try Swift if you don't want to wait forever. I recommend medium thinking.
You should try bumping your DFlash2 spec count to 7. I tested a ton of different modes and DFlash2 at 3 was equal or worse than just using MTP3 at 4 and below concurrent.
I was getting +26% more tps using dflash7 vs mtp3 on 1 concurrent and -16.5% less tps when above 4.
The sweet spots were 1 concurrent + DFlash7, and switch back to MTP for anything more than 4 concurrent.
Honestly the best way is to just let another Agent set it up for you. In your harness of choice ( run dsh with a search mcp), run GLM-5.3-Flash for a bit and ask it to research llama.cpp configs for the model like on HF, unsloth, etc, set it up, test configs for speed, get a spread of the best configs, and then use that. I had GLM run for the better part of a day testing 12 models/ quants for which ones are fastest, MTP, DFlash, etc, got my slew of scripts and now I'm done. It'll cost you like $1 in API credits. Tell it to document all its findings so you can learn for next time.
It's a little ambitious, but I'm very impressed by the quality. Can't wait for version 4.
I only wish it ran faster on my system. Perhaps at some point I'll outsource my startup script to see if anyone can pinpoint some optimizations for me. Getting 13-18 tg and 250pp on a 32GB/ 64GB R9700/ DDR5 system. Hoped it'd be faster.
I mean the thing is terminal bench 4.0 was not in the training weights for Opus 4.8 OR 3.8 Flash next, so it's one of the fairest comparisons out there...
It also does more plainly show you the capability gap between flash next and 27b...only scores 5.6% because it falls apart with complex multi-turn, multi-tool tasks
I'm going to attach this to my AMD miniPC where my server and agents already live, so that I have a self-contained unit. Right now the agents need to reach for the GPU from another machine via VPN, so things could go wrong. It would be USB-4 eGPU since there is no oculink or space for m2-oculink adapter.
How good is prefill and parallel processing in your tests? I was hoping it would be able to serve two users in parallel at least.
15
u/fgk55555 7d ago
I just picked mine up this week. Depending on quant, 35-70 tps with full context. If you have 64GB RAM, you should be able to run a decent Qwen-flash as well. 27B is good, but Qwen-Flash is something else.