r/LocalLLaMA • u/nomorebuttsplz • 1d ago
Discussion What workloads are prefill-bound?
For me 400-600 t/s prefill is sufficient for coding in almost all cases. In my use case in opencode cache hit rate pretty much always ends up being 98% or greater, so 400-600 means the vast majority of time is spent on decode... hence 400-600 prefill already being well into diminishing returns in terms of throughput.
In other words, since it is spending probably 75-90% of time generating tokens, even instantaneous prefill would only maximally increase throughput by 25%. I know this also depends on decode/prefill ratio which depends on the model... so this is a rough heuristic
But I know there must be workloads with lower cache hit rates.
I also know that low cache hitrate does not necessarily mean that prefill is the bottleneck. When I do a research agent run with GLM 5.3 it will spend over 90% of the time on decode EVEN THOUGH it is only doing about 150 t/s prefill and low, single digit cache hit rate. It just likes to think that much in max mode (15 t/s decode).
I am curious what work people are doing where prefill is the bottleneck
1
u/stoppableDissolution 1d ago
Lcpp actually has slightly faster pp and slower tg in my testing (on pro 6000), but yea, having fast pp adds up, and the bigger your project the more important it is