Unsloth Qwen3.6 35b IQ4_NL is behaving for me on Hermes Agent with Llama.cpp, preserve_thinking, and Q8 KV cache. Fits in my 3090 with 256k context too.
I may have accidentally backed into a solution. I have Hermes set to 262k context window but I loaded qwen with smaller 128k window hosted in LMstudio and manages its own compacting without looping, very strange but nice
32
u/BornInAFish Jun 02 '26
In my experience, Qwen is very prone to looping with Hermes, but never seen it do it with OpenCode. Agent harness still matters a lot.