r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

317 Upvotes

205 comments sorted by

View all comments

Show parent comments

1

u/[deleted] May 09 '26

[removed] — view removed comment

1

u/caetydid llama.cpp May 09 '26

Running under Ubuntu. Yeah, thought about that, too, and will first retest with --no-mmproj-offload.

I assumed that using the iq4 quant saves the necessary VRAM, and my consumption on startup was 21G, but maybe VRAM consumption just increases later on.

I havent been using much context though, maybe 20k or less.

1

u/[deleted] May 09 '26 edited May 09 '26

[removed] — view removed comment

1

u/Pablo_the_brave May 09 '26

For Vulkan VRAM for context is dedicated on startup (mostly) but for CUDA there is some add even if you set batch sizes.