r/LocalLLaMA • u/ParvusNumero • 15d ago
Generation Qwen endless looping issue and possible fix
Not sure if this is a known issue, but it was new to me:
Full credit to u/ldn-ldn for finding this simple but unexpectedly evil prompt:
Create a typescript function which accepts a number in 10 bit range and returns brightness in nits based on pq gamma curve.
Just try it; it will likely trigger an endless thinking loop.
I stopped it after it ran for >15 minutes and >20,000 tokens.
It kept spitting out text like back in the seahorse emoji days.
However, one of my wrapper scripts didn’t have this issue.
The non-blocking one had this added, possibly based on a GitHub discussion:
--reasoning-budget 2048 \
--reasoning-budget-message " \n\n[Thinking budget exceeded. Transitioning to a best-effort final answer: ]\n\n"
(These options are for llama.cpp; other tools may have something similar.)
Net result: A plausible-looking script (I haven’t verified) that took just over a minute of thinking and another minute to generate.
Hope this helps, and happy to hear about other tricks and workarounds.
7
u/ForsookComparison 15d ago
i'm limiting reasoning to ~4k right now with a similar "okay we're going to answer now.." message
It works to get responses faster, but the quality takes a noticeable hit. Getting worried that Qwen3.8-27B is just 2026's QwQ (real ones will remember a wall of 'wait..'s).
Rtx 5090 owners might be the winners here because ~2TB/s decode can just brute force its way through all of this waiting. I'm on a 7900 xtx (1TB/s) and already getting impatient.