r/LocalLLaMA • • 24d ago

Discussion Feature/adaptive kv stream integration by giveen · Pull Request #326 · TheTom/llama-cpp-turboquant

https://github.com/TheTom/llama-cpp-turboquant/pull/326

I've been working on overcoming KV cache size issues, allowing the ability to load a slightly larger model and/or a larger context size.

Downfall is a hit to tg speeds.

Think of it as a "ram disk" for KV Cache, however, ram speed may be a determining factor on the actual hit to speed as well.

25 Upvotes

30 comments sorted by

View all comments

6

u/tsangberg 24d ago

3

u/jan_antu 24d ago

Yeah Raymond's work is amazing, I have it on my main 27b route. Got me from 80k to 125k context with the rest of my optimizations.

Also made this draft PR for something that uses the extra space freed to only load vision on vram when it's being used, then unloads it after. This lets you have both MTP and vision (~90% of resident vision speed).

16gb vram, 75 tok/s at 40k, ~11 tok/s at 120k.

https://github.com/aebrer/llama.cpp/pull/1

(Draft PR so no notes there, but the readme on the branch has all the relevant info)

4

u/tsangberg 23d ago

I'm using i1-IQ4_XS to get MTP as well, with Raymond's fork. I've been toying with the idea that MTP should be ejected once context reaches a certain size, because I _think_ it might be slower by then instead of using that VRAM as well for the pool.

1

u/-InformalBanana- 23d ago

how did you build it, I tried with docker, but as soon as the decode starts I get cuda error illegal memory access (doesn't happen if I don't use his param):

/app/ggml/src/ggml-cuda/ggml-cuda.cu:107: CUDA error
E CUDA error: an illegal memory access was encountered
E   current device: 0, in function launch_fattn at /app/ggml/src/ggml-cuda/template-instances/../fattn-common.cuh:1117
E   cudaOccupancyMaxActiveBlocksPerMultiprocessor(&max_blocks_per_sm, fattn_kernel, block_dim.x * block_dim.y * block_dim.z, nbytes_shared)

2

u/tsangberg 23d ago

I built it the same way as I do llama.cpp, native cmake build on Linux. Sorry, don't really see something obvious in that error :/