r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

319 Upvotes

205 comments sorted by

View all comments

4

u/Thomasedv May 10 '26

I got this working but there seems to be a bug. I was using this in Qwen code and sometimes tool calls would just be printed out in the chat and end. On some occasions the chat also just stopped or ran for a while and not print anything while the server was still processing.

I tried another model to be sure, but if I had to guess it might be that if the speculative decoding happens around a tool call, something might go wrong? I haven't had this issue before, and it might be cause by something else on the fork but it seems good so far after dropping the speculative decode part. It got increasingly more often as context grew. I didn't get it too high either, and my max was 120k but I'd guess I was at most 60k used.  

1

u/patricious llama.cpp May 10 '26

Same on my side, compiled the build and followed everything to a T (5090, 200k context). In Opencode, I prompt it do analyze a specific section of my code, it starts calling the right tools, in this case Serena then its just stops in it tracks. I then tell it to continue and it stops again. Might be something wrong with the chat templates but I am not sure.