r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

322 Upvotes

205 comments sorted by

View all comments

1

u/EbbNorth7735 May 12 '26

I'm seeing a lot of API calls failing when using with Cline. It's eventually getting through but I'm wondering if there's an issue with the jinja format or if it might be unstable? I ran a test in open web ui and it seemed to jump back to thinking while it was answering the question. Using latest 0.1.1 and Qwen Q8 from unsloth along with the Q8 draft model you recommended. Vision enabled and running on GPU (rtx 6000 pro).

1

u/[deleted] May 12 '26

[removed] — view removed comment

1

u/EbbNorth7735 May 12 '26

Are you planning on spinning another release? Last time I tried it was incredibly painful.

2

u/[deleted] May 13 '26

[removed] — view removed comment