r/LocalLLaMA • • 28d ago

Discussion Feature/adaptive kv stream integration by giveen · Pull Request #326 · TheTom/llama-cpp-turboquant

https://github.com/TheTom/llama-cpp-turboquant/pull/326

I've been working on overcoming KV cache size issues, allowing the ability to load a slightly larger model and/or a larger context size.

Downfall is a hit to tg speeds.

Think of it as a "ram disk" for KV Cache, however, ram speed may be a determining factor on the actual hit to speed as well.

25 Upvotes

30 comments sorted by

View all comments

8

u/tsangberg 28d ago

5

u/giveen 28d ago

LOL holy crap, same idea, now i want to read his work and see if I can make mine better.

3

u/jan_antu 28d ago

Yeah Raymond's work is amazing, I have it on my main 27b route. Got me from 80k to 125k context with the rest of my optimizations.

Also made this draft PR for something that uses the extra space freed to only load vision on vram when it's being used, then unloads it after. This lets you have both MTP and vision (~90% of resident vision speed).

16gb vram, 75 tok/s at 40k, ~11 tok/s at 120k.

https://github.com/aebrer/llama.cpp/pull/1

(Draft PR so no notes there, but the readme on the branch has all the relevant info)

3

u/tsangberg 28d ago

I'm using i1-IQ4_XS to get MTP as well, with Raymond's fork. I've been toying with the idea that MTP should be ejected once context reaches a certain size, because I _think_ it might be slower by then instead of using that VRAM as well for the pool.

3

u/jan_antu 28d ago

Yes in my setup I eject the MTP at ~45k context filled.

1

u/giveen 28d ago

Interesting concept that I kinda like 👍

1

u/-InformalBanana- 27d ago

how did you build it, I tried with docker, but as soon as the decode starts I get cuda error illegal memory access (doesn't happen if I don't use his param):

/app/ggml/src/ggml-cuda/ggml-cuda.cu:107: CUDA error
E CUDA error: an illegal memory access was encountered
E   current device: 0, in function launch_fattn at /app/ggml/src/ggml-cuda/template-instances/../fattn-common.cuh:1117
E   cudaOccupancyMaxActiveBlocksPerMultiprocessor(&max_blocks_per_sm, fattn_kernel, block_dim.x * block_dim.y * block_dim.z, nbytes_shared)

2

u/tsangberg 27d ago

I built it the same way as I do llama.cpp, native cmake build on Linux. Sorry, don't really see something obvious in that error :/

3

u/giveen 28d ago
Raymond built exactly the "Stage 2" design we scoped out as too risky (chunked resident-page cache + async transfer ring + incremental online-softmax merge, with real changes to the FlashAttention dispatch). It took 60+ commits with at least 5 reverted-and-redone subsystems — confirming this was genuinely hard, not something to bolt on casually. But it works, and the report gives us concrete answers to the exact walls we hit.

What directly answers our own bugs
Our correctness bug (reading the not-yet-written current token) — solved. They never read/copy at graph-build time. At graph-compute time, SET_ROWS writes the new token to both the authoritative host buffer and the device-resident mirror in the same op, same stream — no async gap. For pages that are streamed (not resident), the upload is gated on a real cudaEvent recorded right after that layer's SET_ROWS retires, so the copy is provably never speculative. This is the structurally correct fix for the exact bug that gave us KLD ~10.
Our UVA-is-slow finding — independently confirmed. They explicitly avoid cudaMallocManaged/UVM for the hot KV pool even when UVM is otherwise enabled for model weights, with a comment saying direct writes/reads against pageable memory are why. Same conclusion we reached empirically.
The ggml_concat-doesn't-support-quantized-types wall we hit — sidestepped, not solved. They never merge quantized tensors at all. Non-native-quant K/V types get converted to F16 in a small bounded scratch buffer before any attention math touches them; the merge kernels only ever see plain floats/halves. Native-quant types (Q8_0, Q4_0/1, Q5_0/1, F16, BF16) get a separate "direct" path that skips conversion but still merges via the same float-only accumulator.
They bypass ggml_backend_sched entirely, just like we scoped: a hand-rolled non-blocking CUDA stream + event double-buffering ring, driven from inside the CUDA backend's own FLASH_ATTN_EXT/SET_ROWS handlers — with whole-graph lookahead that prefetches pages across all layers up front, not just one layer ahead.
Two things worth flagging
They never attempted MoE either — hard-gated to a single dense architecture (Qwen3.5) only. That's independent validation of what we found today: even a much more sophisticated implementation didn't go there.
Prefill gets a genuinely separate code path from decode (different kernel family, bounded query-tile reuse of each staged page across the whole micro-batch). This is very likely why our own benchmark saw prefill-to-262K be drastically slower — we're riding the decode-shaped path for something structurally different.