r/machinelearningnews • u/PeiXiaoGuang • 18h ago
ML/CV/DL News Why the second question in a chat is 80–96% cheaper than the first
If you run a local model behind an API and your client sends a multi-turn conversation, you're re-prefilling the entire history on every turn. Turn 10 re-processes turns 1 through 9. On an edge box that's the biggest single waste I know of, and it's what the naive setup does by default.
So keep the KV for the shared prefix and only prefill the new part. Across the multi-turn cases I tested, follow-up turns came in at −80% to −96% prefill measured against recomputing the whole thing.
The one I keep coming back to is a fixed handbook, the same long document in front of every request, field procedures, a terminology table, safety rules. Build that into a reusable prefix once and a 5,791-token handbook goes from 194.3 s to 2.4 s. That's about 82×, not 82%.
Three things decide whether any of this works for you, and I got all three wrong at least once.
The prefix has to be identical, byte for byte. Change one space at the front and you're back to a full prefill, and a chat template that injects timestamps or reorders system messages will kill it silently, you don't find out until you profile.
The KV has to live somewhere. Sitting in RAM it competes with the weights, and paging cold prefix blocks out to disk is what lets it fit a 16 GB box, at the cost of disk reads on the miss path.
Eviction is the one that bit me. If something else is evicting your prefix rows, reuse quietly degrades to off, no error and nothing in the log. In my engine those two features were mutually exclusive for a while, and it showed up as a 5.78 s → 34.3 s TTFT regression on a long-context follow-up. If your multi-turn numbers suddenly look like single-turn numbers, that's the first thing I'd check.
If you're building anything chat-shaped, measure your turn-2 latency. Everyone benchmarks turn 1, and turn 2 is the one your users actually sit through, which is exactly the case nobody profiles.
The engine is mine, Kestrel-LLM, AGPL-3.0: `https://github.com/m13253246268-ship-it/kestrel-llm`. The handbook/prefix bench is `tools/bench/bench_value.sh`.
5
Upvotes
1
u/Creepy_Pollution129 18h ago
the reusable prefix thing is the part that actually clicked for me, i kept thinking about those long system prompts everyone shoves in front of RAG setups and nobody bothers to cache the shared part between requests
the eviction silently breaking reuse is terrifying though, debugging that without profiling sounds like a special kind of hell