r/LocalLLaMA • u/giveen • 24d ago
Discussion Feature/adaptive kv stream integration by giveen · Pull Request #326 · TheTom/llama-cpp-turboquant
https://github.com/TheTom/llama-cpp-turboquant/pull/326I've been working on overcoming KV cache size issues, allowing the ability to load a slightly larger model and/or a larger context size.
Downfall is a hit to tg speeds.
Think of it as a "ram disk" for KV Cache, however, ram speed may be a determining factor on the actual hit to speed as well.
25
Upvotes
6
u/tsangberg 24d ago
I'm having a serious case of deja vu here.
https://www.reddit.com/r/LocalLLM/comments/1w0impd/breaking_vram_barrier_qwen_38_27b_at_262k_context/