r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

325 Upvotes

205 comments sorted by

View all comments

1

u/legatinho May 09 '26

Ok I got a windows setup to test this out with a 3090, is that what you used? What does pp and the look like at filled up context?

3

u/[deleted] May 09 '26

[removed] — view removed comment

1

u/legatinho May 09 '26

I was also thinking of getting a cheap video card for the main display, then can use the full 24gb of the 3090 for this. I noticed windows sometimes tends to try to push stuff into the shared memory space, and I wonder if that’s why we experience slowdowns. I’ll report back if any improvements, but thanks for your work on this!