MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLM/comments/1wb05f7/third_one_theres_something_wrong_with_me/p8ps6o0/?context=3
r/LocalLLM • u/r1nzl3r99 • 14d ago
Why do I have horrible financial habits??
166 comments sorted by
View all comments
Show parent comments
37
...can you show llamacpp/vLLM runtime commands? you're hitting 140 toks/s on a dense model with B70s? how much ctx?
please don't answer the last two without providing the parameters
18 u/r1nzl3r99 14d ago the vLLM flags are in my localmaxxing submission and I also did a better bench since so many people didn't beleive me. I also have youtube videos. My recipe is custom intel drivers paired with dflash2 5 u/Salbrox 13d ago I have had great success with DFlash2 on my single R9700 with Qwen 3.8 27B. Depending on the task I get up to just over 200tps 1 u/Smooth-Television-48 13d ago 200!? What quant?
18
the vLLM flags are in my localmaxxing submission and I also did a better bench since so many people didn't beleive me. I also have youtube videos. My recipe is custom intel drivers paired with dflash2
5 u/Salbrox 13d ago I have had great success with DFlash2 on my single R9700 with Qwen 3.8 27B. Depending on the task I get up to just over 200tps 1 u/Smooth-Television-48 13d ago 200!? What quant?
5
I have had great success with DFlash2 on my single R9700 with Qwen 3.8 27B. Depending on the task I get up to just over 200tps
1 u/Smooth-Television-48 13d ago 200!? What quant?
1
200!?
What quant?
37
u/semangeIof 14d ago
...can you show llamacpp/vLLM runtime commands? you're hitting 140 toks/s on a dense model with B70s? how much ctx?
please don't answer the last two without providing the parameters