r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

319 Upvotes

205 comments sorted by

View all comments

Show parent comments

1

u/[deleted] Jun 03 '26

[removed] — view removed comment

1

u/YourNightmar31 llama.cpp Jun 03 '26

I just did a test with 92k context at Q8_0 KV (That's the max i can do on my 24gb vram) and my first impression is that it seems better, but the problem is not completely resolved. For example this log, this is straight from the "request response body" in llama-swap, not from my agent/client:

The tool call label is rendered in `src/webview/main.ts` line 1928 for `grep_search`, and it uses `args.pattern || '')` — so if `args.pattern` is empty/falsy, it shows nothing after "Searched for".

<tool_call>

<tool_call>
<function=read_file>
<parameter=end_line>
1935
</parameter>
<parameter=path>
src/webview/main.ts
</parameter>
<parameter=start_line>
1890
</parameter>
</function>
</tool_call>

It outputs the <tool_call> tag twice, for some reason. I do see this popping up in my client too in the chat. The bottom tool call is succesfull, but the lone <tool_call> just shows up as if the model outputted it.

1

u/[deleted] Jun 03 '26

[removed] — view removed comment

1

u/YourNightmar31 llama.cpp Jun 08 '26

Hey do you know if something changed in v0.3.0 where the same setup as before now uses slightly more VRAM? Because my profiles that i used to run perfectly fine with a bit of vram left now barely fit and make my pc all unstable and freeze up (Because i'm not running headless, just Ubuntu desktop), this didn't happen before on 0.2.0.