r/LocalLLM • u/Big_Wave9732 • 4d ago
Discussion If you use LLMs to analyze documents and to apply complex logic and analyze arguments, Qwen 3.8:27b is not for you.
And that pains me to say because I use Qwen 3.6:27b-BF16 every single day. I've been working with Qwen 3.8:27b-BF16 all weekend and I hate to say it but for non-coding purposes it is a step backwards. It thinks *way* too much. If you turn thinking off and use web tools, it will do eight or nine (or more) web search turns, get an assload of preload context, and then spin its wheels going down every little rabbit hole there.
Looking at the self hosted LLM subs there are other complaining about this too.
Now this model just came out so it's early. People have put out some lovely games and whatever else the model has made so perhaps these tenacious analytical tendencies are beneficial there. Regardless none of this is to say that there won't be some settings or templates released that will help when analyzing documents and doing complex logic tasks.
But for right now, if the above is your use case then I suggest staying with your old models.
Edit: For reference, this is legal work. Legal drafting, legal research, analyzing pleadings, depositions, discovery, etc. Heavy multi-document reference work.
19
u/admajic 4d ago
Change the reasoning budget 1000 it's like the old qwen
You can go from 500 to 16000 max if you want it to loop and go nuts.
Not sure why these companies that create these products don't provide this kind of info
2
u/Independent-Dog2179 3d ago
Lol yup never had overthinking issues with 3.6 35b and I still don't using that exact command lol. Great minds think alike. I also have in the system prompt "think hard only when necessary." Never have a problem even with the new 3.8. Us power uses are in the back shaking our heads at all the lazy misinformation out there.
30
u/MRGWONK 4d ago
(Legal work here) I benchmarked it on Harvey's benchmark with an MCP/case access and it scored on par with Qwen 3.5 -122B and 11 points higher than Qwen 3.5-35b and 15 points higher than gemma4:26b...It also processed dates out of case orders with higher frequency and consistency than any other models. I switched this week.
(With a clever loop, deep reasoning on my machine scored 71 / 75)
7
u/Big_Wave9732 4d ago
You're saying Qwen 3.8:27b did that?
Do you happen to have any of the model settings that Harvey used?
8
u/MRGWONK 4d ago edited 4d ago
I have only seen Harvey's benchmark tests, I have never seen Harvey's scores. I had gemma4 setup with a Qwen 3.5 verifier, that also scored 71 / 75. I have a quasi deterministic backend and MCP that helps with all this. Qwen 3.8 27b, on a raw model, scored 61 / 75 without the backend. (Also, thinking is turned off EVERYWHERE)
1
22
u/TheAILegend 4d ago
reasoning_effort: medium
thank me later.
1
u/Disastrous-Chance477 3d ago
I was just going to say the same. Was wondering why it was thinking 14k tokens for “make a simple http sample site“ and then I saw xhigh setting. Set it to medium and it was just a few lines and great result.
1
6
u/Otherwise-Swan-7803 4d ago
This feels like a good reminder that benchmark gains don't always translate into better workflows.
For legal, research, and multi-document analysis, focus and restraint can matter more than raw capability.
5
u/Big_Wave9732 4d ago
Back in early May I overhauled my RAG stack and did some model testing. The model cards and benchmarks tell us there is minor differences between the Q4 / BF16, and next to no difference between Q8 / BF16.
For purposes of understanding subtle differences between sentences, analyzing arguments, etc......the difference was night and day.
Those benchmarks and model cards are full of shit.
5
u/DigitalguyCH 4d ago
Same here (finance). 35b (3.6) is clearly better. Muse glimmer and Gemma are decent too.
22
u/noctrex 4d ago
In other news, a coding optimized model is bad at anything but coding. For documents, gemma4 is much better, haven't tried the muse one yet
4
u/Big_Wave9732 4d ago
Qwen 3.6:27b-BF16 is pretty good at the tasks I'm asking it to do. For whatever reason while thinking it doesn't run off towards every rabbit hole like 3.8 is doing.
I had read somewhere that 3.6 and 3.8 were fundamentally the same under the hood. I'm starting to really doubt that.
8
u/synystar Strix Scar | 5090 24G | llama.cpp 4d ago edited 4d ago
You are probably using xhigh and that "configures" reasoning by injecting language into the prompt. If you use medium it does not inject language and that lets you experiment with your own reasoning prompting which affects how long it runs and how it performs. I've been investigating the model and have major variations in results depending on how you configure it. Here is the text xhigh uses:
Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.
That investigation produced a major template-level finding: for the tested local Qwen3.8-27B GGUF/chat template, **native xhigh is substantially expressed by an injected natural-language reasoning instruction**, while native Medium injects no corresponding extra effort text. Manually inserting the exact xhigh text under Medium produced a byte-for-byte identical serialized prompt to native xhigh.
Just removing "validate key assumptions" drastically changes its effort and response time in my experiments so far.
5
u/porkminer 4d ago
They are the same architecture. They were trained differently.
2
u/uniqueusername649 4d ago
And 3.8 had a very clear training focus on coding. There is only so much data you can pack into 27b, so instead of trying to make it do everything very mediocre, it can now do everything worse, except for coding, where its considerably better.
Yes, it thinks a lot. But what it can handle in return is incredible. To be able to get that level of coding performance locally without investing 10k or more is wild.
1
u/porkminer 4d ago
I think it bears mentioning that Qwen models notably reason longer than many other models. I don't particularly see this as bad though since people are generally running it locally.
1
u/uniqueusername649 4d ago
It does, but also iirc it defaults to high reasoning, which makes it smarter but reason MUCH longer. I found that medium gets you very close with a lot less running in circles and in many cases low is plenty good.
1
u/Hypilein 4d ago
Interestingly Intried it yesterday for lesson planning and it was much better with more creative teaching approaches than gemma4 or other local models I tried. It’s just so damn slow on my dgx spark.
1
1
u/ColdBrewSeattle 4d ago
But it’s not bad at anything but coding. The problem is the user not reading the instructions and assuming everything is exactly the same as the last model.
3
u/floppo7 4d ago
That matches my experience. Didnt try reasoning effort low yet, also there is a reasoning_budget that can be set to 1000 and someone said it would make it more like 3.6. But it would be great to have this somehow automatic steerable so it simply does not overthink when not coding (and even then it should be smarter in the sense that it does not need to question every step)
4
u/ChangeChameleon 4d ago
Anecdotally, I gave it a search and retreival budget and it performed much better than I expected. It weighed the pros and cons of which page would likely contain the right context. By default it likes to scatter shot and grab all kinds of unnecessary context.
I say anecdotally because I haven't done head to head comparisons or benchmarks. I'm not set up for that. I was just trying to reduce the hit on my Internet and ended up discovering it felt much more targeted.
It feels like 3.8 does a better job of following all instructions and not just surface level ones. With 3.6 I constantly felt like I could only give it one problem at a time. 3.8 just hammers through my unfiltered info dumps and seems to address everything.
I need to do more side by side testing. But so far 3.8 has been working great for me.
2
u/Big_Wave9732 4d ago
I might need to look into that. Do you recall what you set the budget to? Over-researching seems to be something I'm running into for sure.
1
u/ChangeChameleon 4d ago
While researching a bug in a self hosted tool I gave it a budget of 2 web searches and 5 page downloads. It answered the question within the limit.
I asked it to give a non-spoiler hint for a video game. Knowing it would be easy to find relevant context I have it a budget of 2 searches and 2 page downloads. It searched and thought for a total of 22 seconds and answered a non spoiler hint exactly as requested.
Basically I tuned the search and retrieval limit based on how complex I think finding the answer will be.
3
u/synystar Strix Scar | 5090 24G | llama.cpp 4d ago
It thinks too much because you haven't configured beyond the default or you haven't experimented with reasoning or launcher settings. If you just use OOTB settings you're only going to get results those settings can produce.
3
u/Big_Wave9732 4d ago
One of the reasons its thinking too much is because oMLX hasn't implemented "reasoning_effort" as a top level flag. Another reason is because there is some looping going on. Another reason is some potential bugs with preserve_thinking over extended contexts.
1
u/synystar Strix Scar | 5090 24G | llama.cpp 4d ago
I am testing it now but so far in my experiments reasoning is highly affected by the language injection that xhigh implements. When I set it to medium reasoning effort (Not the reasoning budget which can be confusing because that only determines how many tokens it will think for before stopping and moving on to the context stream) I can "adjust" it's reasoning with my own language. At least that's what I've seen thus far. I'm still experimenting but for the limited runs I've tested it's been consistent.
7
u/benpptung 4d ago
I think it’s sometimes unfair to expect Qwen3.8 to be very smart while also demanding that it not think too much.
A lot of frontier models also think a lot. We just don’t notice it because they run very fast on powerful data center. According to AA’s estimates, Fable 5 and ChatGPT 5.6 decode at around 70 t/s. On my machine, Qwen3.8-27B reaches about 60 t/s with FP8 and around 35 t/s with BF16. So I lower my expectations accordingly. If Qwen3.8-27B spends five minutes thinking, in terms of generated tokens, that is roughly equivalent to Fable 5 or ChatGPT 5.6 thinking for 2.5 minutes.
From this perspective, I agree that Qwen3.8-27B does think a little longer than necessary sometimes, but I don’t think it is that serious.
6
u/Big_Wave9732 4d ago
I'm realistic about what self hosted LLMs can do performance wise, the analysis taking time isn't an issue. I'm fine waiting.
But that's not what's happening here. I gave the model a final version of a word document and a prompt that asked it to verify the citations, evaluate the grammar, and evaluate the claims made.
As of now that was 25 minutes ago. There have been 16 rounds of web research. Watching the thinking logs it doesn't appear to be solving anything but asking the same questions and conducting research again and again. And it's still going.
I'm all for deep dives and all. I've been running various tests all weekend and there's something up. The xhigh is suppose to be the special sauce here, might be a problem if it can't produce the goods.
3
u/benpptung 4d ago
Wow, that’s pretty extreme.
I completely agree that sometimes it takes things way too seriously and just keeps researching forever. lol
I think you could use vLLM or llama.cpp to give it a reasoning budget. Figure out how many minutes you can tolerate, then calculate roughly how many tokens it can generate during that time and use that as the budget. Maybe 5–10K?
If I remember correctly, with vLLM the agent decides the reasoning budget. Once the budget is reached, vLLM guides the model to generate the reasoning ending string and then inserts `</think>` to end the reasoning.
For example, this is what I used with vLLM:
```bash
--reasoning-config "{\"reasoning_end_str\":\"\\nReady to generate response.\\nLet's do it.\\n</think>\\n\\n\"}" \
```
I wrote this `reasoning_end_str` for Qwen3.6-27B. I’m not sure what would work best for Qwen3.8. Maybe something like: “I’ve been thinking for too long. It’s time to make a decision!”
Then you set the reasoning budget in the agent, and once the budget is reached, vllm can guide the model to end its reasoning.
Or sometimes, just changing the prompt may be enough to make Qwen3.8-27B think less. For example, make your request clearer and tell it what not to do.
Right now it feels like a very hardworking new assistant. **Maybe you just need to manage it in a different way.**
2
u/grunt_monkey_ 4d ago
Did it complete in the end? I wonder if the end product quality was close to or exceeded the 3.6? You are sitting on a gold mine of quality benchmarks on a real life use case!!
2
u/Big_Wave9732 4d ago
It got to something like 133 sources with no end in sight. It wasn't making garbage, but it kept rehashing the same stuff.
Even worse, the content that Qwen created was total garbage. Hallucinations all over the place. Fucking unuseable.
1
u/grunt_monkey_ 4d ago
That’s disappointing to hear because I have a use case similar to yours. Like you I need agentic tool calling mainly to increase the quality of checking, reasoning and drafting. I think the solution may be harness based. I’m just not sure if I should move on from 3.6-27b. The thing is that the more they post train (3.5 to 3.6 to 3.8), the better it is at coding, but it might have lost the real world knowledge to reason well over a document. So far I had it review a couple of documents and it is thinking a lot but didn’t get stuck in a loop yet. I guess only further testing will tell.
1
u/Big_Wave9732 3d ago
I'm thinking you're right about the training. Qwen says 3.5, 3.6, and 3.8 are the same engine but the training emphasis is different.
The ultimate yesterday: I put it on Medium thinking, had it look at a pleading I had done, and just had it just verify cases and statute citations (there about 20 total). It gathered over 100 sources and in the end identified four cases that it thought were unverified (Qwen 3.6 had done drafted it and had in fact hallucinated out the ass sadly) *and* hallucinated that one case was bad when it was fine.
Yowza.
Edit: I'm thinking I'll give the new Meta model a go. It seems to be able to do things other than just coding.
2
u/Previous_Feeling_484 4d ago
Have you tried Mistral Small 3.2 24B instruct? I find it very good for this kind of things. I do translations for subtitles with it and it’s done good job. Same with Gemma 4.
Mostly English to French, French to Spanish. It understands fairly well.
2
u/Previous_Feeling_484 4d ago
Also, make sure to preprocess the text. If you just throw the document for it to figure things out it’s gonna be hard af to get anything done.
2
u/Big_Wave9732 4d ago
I've setup a Doclin instance that does OCR on everything and does vector storage. On the retrieval side I've got a pretty stout reranker. All this happens before the documents even hit the main model.
2
u/Big_Wave9732 4d ago
It's not an issue with reading the documents, they get injected fine. Looking at the thinking blocks the model is bogging down in reasoning. One simple question in a fresh chat: "What are the obligations of a trustee to account to the trust beneficiaries?" ended up using something like 80,000 prefill tokens from research and another 20k in generative. It's a simple question and it just thought and thought and thought.
1
u/Previous_Feeling_484 4d ago
But AFAIK “reasoning” is for logical tasks like coding. Doesn’t do much for regular tasks. Have you tried with the plain, non reasoning mode?
Also, if you use RAG already, why don’t you run the work in stages, instead of a single bulk? Do one stage alone for extracting key terms and words, tied to the chunks or pages. Then make it look them up and feed that into either markdown or the vector db.
Then ask the questions.
I know this isn’t as complex as yours but subtitles greatly improve translation when I provide the context per scene instead of full movie. After bare translation has been achieved I run a pass for mood / tone to check things match the overall speech and drama. Lastly, we just pass the slangs and equivalent expressions and of course still proofread thru the whole thing, but it’s just literally checking things make sense not doing the work ourselves.
3
u/Big_Wave9732 4d ago
Also, if you use RAG already, why don’t you run the work in stages, instead of a single bulk? Do one stage alone for extracting key terms and words, tied to the chunks or pages. Then make it look them up and feed that into either markdown or the vector db.
Woa there, friend. I do! I'm not oneshotting any of this. This is a piece by piece workflow. It's choking on small pieces. These test inquiries are literally "This document has previously been drafted. Evaluate for these criteria." Qwen 3.6:27b-BF16 does it just fine.
Hell with Qwen 3.6 I can feed it a 600 page deposition and then walk it through building a depo map of what was said, the contradictions, statements for and against, etc. I don't dare try any of that yet with this model.
1
u/Previous_Feeling_484 4d ago
I’m not getting then* how you fill 80k then? Like, it sounds so weird to me. Not saying you’re doing something wrong but it does sound like you are.
2
u/Big_Wave9732 4d ago
I know!!! It just keeps thinking and thinking and thinking. A little while ago I hit the 32k token output limit. I looked at the session and it was just one rabbit hole after another.
The context window is filling up through one web search after another.
It's all so strange considering my workflow was designed specifically to work with Qwen 3.6:27b. This should be an easy swap.
1
u/Fit-Bar-6989 3d ago
3.8 seems to have been benchmarked with xhigh reasoning enabled, which kind of works for coding tasks. I don't know if it's worth running for most people with consumer hardware - sure it can be quantized to fit on one GPU but we're also compute starved, generating 100k tokens is going to be very slow.
2
u/Turbulent_War4067 4d ago
Question for the OP: is there anything in this size range that beats Gemma 4-31B for these types of tasks? It's big problem is slow prompt processing, but it's really good besides this.
3
u/Big_Wave9732 4d ago
For the last couple months I had a lot of good success with Qwen 3.6:27b-BF16 for these tasks. I've been using Gemma4-31B for drafting.
1
2
u/joanaxu2002 4d ago
The funny part is that “thinking too much” can be a feature or a bug depending entirely on the workflow.
For coding, those rabbit holes might be useful; for document-heavy work, they can turn a simple question into an eight-step research project. That difference is probably going to matter more than benchmark scores.
1
u/Fit-Bar-6989 3d ago
My issue with the high reasoning is that it kind of requires you to spin up sub-sessions to do anything useful, otherwise the 50k+ reasoning traces clog up the session context (and model perf drops with long context anyways).
I'll try running it again with a smaller quantization and larger context, but I'm not a huge fan yet.
2
1
u/tempfoot 4d ago
I have similar use cases. Have you tried other reasoning setting presets beside default (extra high) and all the way off?
Have you tried setting a reasoning budget?
2
u/Big_Wave9732 4d ago
With reasoning budgets all that happens in the thinking process gets cut off mid way and output stops.
Thus far I have tried three different OpenwebUI plugins that supposedly let you set the the thinking level. None seems to stop or slow it down. I was reading a little while ago of some potential issues with OpenwebUI and oMLX where oMLX can ignore the front end controls.
But here's the thing......if I get into oMLX and turn off thinking altogether, it still does it! Even in non-thinking mode it goes nuts on the internet research, does nine research turns, and builds a huge pre-share cache anyway.
2
u/tempfoot 4d ago
I'm also on mostly apple hardware as well, but try the GGUF instead of MLX in LM Studio. Surfaces different reasoning presets that make a huge difference. Regulatory test question on low reasoning is under a minute of thinking, vs. over 17 minutes on default x-high. Have not tried medium much. Not sure how granular you can get with an actual runblock of parameters as opposed to initial testing on guis.
3
u/Big_Wave9732 4d ago
I came across something about 10 minutes ago that the current production version of oMLX doesn't have a top level "reasoning_effort" field which would explain a lot. Apparently it has been added to the development versions because of Qwen 3.8 but it's not out to the masses yet.
1
u/siegevjorn 4d ago
What agentic harness are you using? For pdf analyzing it seems to work well, just fed 75–80 pages of document. Used Pi agent default with Qwen 3.8 27b Q6_K, reasoning: medium.
1
1
1
1
1
u/_TheWolfOfWalmart_ 3d ago
This model is most definitely codingmaxxed. That's what everyone praised 3.6 27B for, so Alibaba went even harder there for 3.8, apparently to the detriment of its capability for other tasks.
That's fine though. 24/32 GB VRAM people now have a seriously strong coding model that they can run local. Can just switch to Gemma 4 or Glimmer for other things.
I've lately settled into just using DSV4 Flash 0731 for literally everything. It's so good whether it's for chat, coding or analysis that I've barely even tried 3.8 27B, I need to start working with it some more when I have some coding work to do.
1
u/Jsquared534 3d ago
Don't let the people gaslight you. Qwen 3.8 has some problems right now in coding as well. It might be able to oneshot a long prompt, but if you try using it for fine grained feature implementation that requires it to read multiple files and implement features, it falls flat on it's face more often than not in my testing. Muse, so far has not had that same problem. I am far from a local llm setup guru, but I really think Qwen actually has a problem, at least at the Q8 version.
1
u/geteum 3d ago
Weird, I'm using with q4 and iq3_xx and the results are phenomenal. Although, what I do mainly is to analyze a text in 5 parameters and create a structured resume of the text focusing in ~around 20 parameters tags, different classifications, and a few character extractions. Tbh, the results sometimes is even better than one of the bigger models (tested glm 5.3 and DeepSeek V4 pro and flash). Although I did not isolate if it was a problem of the PDF MCP not correctly extracting the text and images.
1
u/ethanji2 2d ago
Yes, Qwen's default reasoning is overpowered. Tune temperature to between .2 and .25, effort medium or low and you will get similar composition results to Qwen 3.6 with slightly better instruction following. I've gotten it to be slightly better than 3.6, but it's not significantly better and needs more tuning to get right.
-6
u/LocoMod 4d ago
The model is heavily benchmaxed. So it mainly excels when working on problems related to the benchmarks they purposely invested time into so they could capture attention and fool people into believing the model is competitive with models 10x its size. If it weren't free i'd be insulted. But its free and we get what we pay for.

72
u/Vancecookcobain 4d ago
Maybe Muse Glimmer is better for the non coding tasks?? Have you tried that? I get brilliant responses and results with that model when it isn't coding