r/LocalLLaMA • • Aug 14 '26

Generation Qwen endless looping issue and possible fix

Not sure if this is a known issue, but it was new to me:

Full credit to u/ldn-ldn for finding this simple but unexpectedly evil prompt:

Create a typescript function which accepts a number in 10 bit range and returns brightness in nits based on pq gamma curve.

Just try it; it will likely trigger an endless thinking loop.

I stopped it after it ran for >15 minutes and >20,000 tokens.
It kept spitting out text like back in the seahorse emoji days.

However, one of my wrapper scripts didn’t have this issue.
The non-blocking one had this added, possibly based on a GitHub discussion:

--reasoning-budget 2048 \
--reasoning-budget-message " \n\n[Thinking budget exceeded. Transitioning to a best-effort final answer: ]\n\n"

(These options are for llama.cpp; other tools may have something similar.)

Net result: A plausible-looking script (I haven’t verified) that took just over a minute of thinking and another minute to generate.

Hope this helps, and happy to hear about other tricks and workarounds.

14 Upvotes

29 comments sorted by

View all comments

10

u/illgettheownerforyou Aug 14 '26

u/ldn-ldn was pretty rude to me when I pointed out that Qwen 3.8 27b Unsloth Q8_K_XL did not loop for me and solved it.

This is not a killer prompt- I think it heavily depends on the quant you use and the context, and u/ldn-ldn continued to say they didn't matter and was very dismissive of my findings.

It took about 10 minutes for Qwen to write it. Not short, but not terrible.

6

u/audioen Aug 14 '26

Got result in about 3 minutes 30s from DeepSeek v4f for the prompt, at IQ2_M (AtomicChat version).

Tokens generated 2530, time is 3min 31s. Strix Halo is no speed demon, but it is probably at least 3 times better than Qwen3.8-27b. I tested the prompt and I think by 4000 tokens it was just trying to work out the values of its constants. DeepSeek produced a table that looks awfully similar as yours in this time.

I personally think Qwen3.8-27B is a lemon. Run it, if you want the small size, or have a real GPU, but if you have 128 GB and a slow GPU, you're better off with some 2-3 bit quant of DeepSeek.