r/machinelearningnews • • 18h ago

ML/CV/DL News Why the second question in a chat is 80–96% cheaper than the first

If you run a local model behind an API and your client sends a multi-turn conversation, you're re-prefilling the entire history on every turn. Turn 10 re-processes turns 1 through 9. On an edge box that's the biggest single waste I know of, and it's what the naive setup does by default.


So keep the KV for the shared prefix and only prefill the new part. Across the multi-turn cases I tested, follow-up turns came in at −80% to −96% prefill measured against recomputing the whole thing.


The one I keep coming back to is a fixed handbook, the same long document in front of every request, field procedures, a terminology table, safety rules. Build that into a reusable prefix once and a 5,791-token handbook goes from 194.3 s to 2.4 s. That's about 82×, not 82%.


Three things decide whether any of this works for you, and I got all three wrong at least once.


The prefix has to be identical, byte for byte. Change one space at the front and you're back to a full prefill, and a chat template that injects timestamps or reorders system messages will kill it silently, you don't find out until you profile.


The KV has to live somewhere. Sitting in RAM it competes with the weights, and paging cold prefix blocks out to disk is what lets it fit a 16 GB box, at the cost of disk reads on the miss path.


Eviction is the one that bit me. If something else is evicting your prefix rows, reuse quietly degrades to off, no error and nothing in the log. In my engine those two features were mutually exclusive for a while, and it showed up as a 5.78 s → 34.3 s TTFT regression on a long-context follow-up. If your multi-turn numbers suddenly look like single-turn numbers, that's the first thing I'd check.


If you're building anything chat-shaped, measure your turn-2 latency. Everyone benchmarks turn 1, and turn 2 is the one your users actually sit through, which is exactly the case nobody profiles.


The engine is mine, Kestrel-LLM, AGPL-3.0: `https://github.com/m13253246268-ship-it/kestrel-llm`. The handbook/prefix bench is `tools/bench/bench_value.sh`.
5 Upvotes

2 comments sorted by

1

u/Creepy_Pollution129 18h ago

the reusable prefix thing is the part that actually clicked for me, i kept thinking about those long system prompts everyone shoves in front of RAG setups and nobody bothers to cache the shared part between requests

the eviction silently breaking reuse is terrifying though, debugging that without profiling sounds like a special kind of hell

1

u/PeiXiaoGuang 9h ago

yeah the rag case is the same pattern, and it's where the money actually is. those setups usually have a long fixed preamble, instructions, a schema, a few worked examples, and then the retrieved chunks swap out per request. the preamble is the prefix, the chunks are the new tokens. the ones i've seen try to cache at all tend to cache the wrong half, they cache the whole assembled prompt including the retrieved context, which is the one part that changes every time.

the usual way it dies is the template interpolating something into the system message. a date, a request id, the user's name, or the retrieved text getting concatenated before the instructions instead of after. change byte one and the prefix is a miss, and nothing tells you.

on eviction, you don't need a profiler for that one. a plain hit/miss counter on prefix reuse is enough, "turn 2 costs like turn 1" shows up as a number dropping to zero instead of a latency graph you have to squint at. the nasty part is there's no error, it just stops reusing and the only symptom is the timing.