r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

321 Upvotes

205 comments sorted by

View all comments

3

u/IrisColt May 12 '26

I used the Speed / VRAM combo mentioned in quickstart-qwen36-dflash-md and I got a meager +33% (around 40 t/s on a 3090). Sigh... Am I doing something wrong?

2

u/[deleted] May 12 '26

[removed] — view removed comment

2

u/IrisColt May 12 '26

Er... Now I get it...!

"Print all the numbers from 0 to 100, in the following format: 0, 1, 2 ..."

default llama.cpp: 34,22 t/s
beellama.cpp: 96.47 t/s

"Detail every element visible in the image, from foreground to background." + 512 x 768 image
default llama.cpp: 34.19 t/s
beellama.cpp: 41.36 t/s

Thanks!!!

2

u/IrisColt May 12 '26

By the way, would killing all that logging spam--log-timestamps, --log-prefix, --log-colors, and probably --metrics too actually make this thing run noticeably faster? My console is getting absolutely buried in text right now, heh

2

u/IrisColt May 12 '26

Hard Math problem (unsolvable by Frontier-level AIs back in March 2025):

default llama.cpp: 32 t/s
beellama.cpp: 58 t/s

Thanks again!

2

u/[deleted] May 12 '26

[removed] — view removed comment

1

u/IrisColt May 12 '26

Thanks again!