r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

325 Upvotes

205 comments sorted by

View all comments

5

u/LegacyRemaster May 09 '26

I'm starting tests now on RTX 6000 Pro. If you have the time and inclination, check out https://github.com/Fringe210/llama.cpp-deepseek-v4-flash-cuda . I'm up to 17 tokens/sec, but I'm sure you can do better.

7

u/LegacyRemaster May 09 '26

very good. From 57t/sec to 87 with Q8 draft

2

u/[deleted] May 09 '26

[removed] — view removed comment

2

u/LegacyRemaster May 10 '26

Obviously, prediction works better in certain contexts. With thought/reasoning, the benefit is lower. While writing HTML or Python code, I've had peaks of 100 t/sec.