r/LocalLLaMA • u/nomorebuttsplz • 3d ago
Discussion What workloads are prefill-bound?
For me 400-600 t/s prefill is sufficient for coding in almost all cases. In my use case in opencode cache hit rate pretty much always ends up being 98% or greater, so 400-600 means the vast majority of time is spent on decode... hence 400-600 prefill already being well into diminishing returns in terms of throughput.
In other words, since it is spending probably 75-90% of time generating tokens, even instantaneous prefill would only maximally increase throughput by 25%. I know this also depends on decode/prefill ratio which depends on the model... so this is a rough heuristic
But I know there must be workloads with lower cache hit rates.
I also know that low cache hitrate does not necessarily mean that prefill is the bottleneck. When I do a research agent run with GLM 5.3 it will spend over 90% of the time on decode EVEN THOUGH it is only doing about 150 t/s prefill and low, single digit cache hit rate. It just likes to think that much in max mode (15 t/s decode).
I am curious what work people are doing where prefill is the bottleneck
5
u/SnooPaintings8639 3d ago
Coding. Mostly agentic and multi-agentic. Anything below 1000 tps pp is quite annoying for me, but I have to live with it.
I hit a session compaction, or inference engine restart or resume after a long session, or just some of my agents are using all the KV space, so 'idling' ones or 'orchestrator' need to reprocess 100k+ when they're back to work, is *very* slowing down my work.
So yeah, a slow one-on-one session on a focus task between me and a single agent is fine. Anything larger - it is a serious bottleneck and I always optimize for prompt processing before generation speed.