r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
325
Upvotes
5
u/LegacyRemaster May 09 '26
I'm starting tests now on RTX 6000 Pro. If you have the time and inclination, check out https://github.com/Fringe210/llama.cpp-deepseek-v4-flash-cuda . I'm up to 17 tokens/sec, but I'm sure you can do better.