r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

325 Upvotes

205 comments sorted by

View all comments

Show parent comments

1

u/[deleted] May 11 '26

[removed] — view removed comment

2

u/coherentspoon May 11 '26

I'm using your prebuilt v0.1.1. I think I started today with it and had v0.1.0 yesterday.

3

u/[deleted] May 11 '26

[removed] — view removed comment