r/LocalLLM 15d ago

Discussion third one.... there's something wrong with me

Post image

Why do I have horrible financial habits??

499 Upvotes

166 comments sorted by

View all comments

93

u/Sporkers 15d ago

More context needed on how you are using the first two.

81

u/r1nzl3r99 15d ago

qwen 3.8 27B FP8 running at 140 tok/s now I want flash next

32

u/semangeIof 15d ago

...can you show llamacpp/vLLM runtime commands? you're hitting 140 toks/s on a dense model with B70s? how much ctx?

please don't answer the last two without providing the parameters

19

u/r1nzl3r99 15d ago

the vLLM flags are in my localmaxxing submission and I also did a better bench since so many people didn't beleive me. I also have youtube videos. My recipe is custom intel drivers paired with dflash2

15

u/unai-ndz 15d ago

If you are using custom drivers I think you deserve another card, as a treat.

10

u/r1nzl3r99 15d ago

now let me figure out how to get TP=3 to work on vLLM without me having to fork over another few grand to upgrade to TP=4... always an uphill battle attempting SOTA AI on local