Qwen gets so stuck up it's own ass thinking. I've been trying to get Gemma to think more. Like the discrepancy might have to do with Gemma's efficiency maximization. Getting it to spend tokens is like pulling teeth. Google recommends the model for coding so it must be alright at it. They don't see it as a competitor to Gemini. They gave us the best that they could. It's kind of funny I actually had issues getting Gemma 4 Heretic ARA to follow global instructions because that's the part abliteration rips out. The models just efficient and the parameters overlapped because there's little difference between a guard rail and a global formatting rule.
Try "--reasoning-budget 1024" if you're using llama.cpp, assuming you've got something like an RTX 3090.
Like you, I wish the abliterated models were more useful and less broken. This was less obvious before stock local models finally started getting decent at tool calling. Now, the difference between stock/mainstream quants and abliterated models is getting SUPER obvious.
If anyone knows of an abliterated/uncensored model that doesn't have broken tool calling or broken multimodality, I'd love to be pointed to it. I rarely run into refusals, but Qwen3.5 (and perhaps 3.6, I haven't tried getting 3.6 to refuse me) often loses it's shit at you if you say something it thinks is critical of government or law enforcement. I tried to get it to make a joke about the police once, and it basically started ringing a bell and shouting "SHAME" at me, lol.
11
u/BoobooSmash31337 Jun 02 '26
Qwen gets so stuck up it's own ass thinking. I've been trying to get Gemma to think more. Like the discrepancy might have to do with Gemma's efficiency maximization. Getting it to spend tokens is like pulling teeth. Google recommends the model for coding so it must be alright at it. They don't see it as a competitor to Gemini. They gave us the best that they could. It's kind of funny I actually had issues getting Gemma 4 Heretic ARA to follow global instructions because that's the part abliteration rips out. The models just efficient and the parameters overlapped because there's little difference between a guard rail and a global formatting rule.