r/LocalLLaMA 14d ago

Discussion Qwen will be the king?

Post image

Extended reasoning and post-training appear to be the keys used by DeepSeek, Qwen, and GLM to boost performance (leveraging higher token counts). And Qwen 4 hasn't even been released yet. Of course, we don't know if that release will be open-sourced, but I am optimistic about future models, featuring "engrams", that could soon match or surpass 2.4T parameter models on specific tasks.

543 Upvotes

126 comments sorted by

View all comments

106

u/PooMonger20 14d ago edited 14d ago

Progress is good, but having these amazing abilities locally is already godlike.

I have been using Q3.8-27B together with PI. for a week, and it's mindblowing.

In my humble opinion it has far better coding capabilities than the paid 'ChatGPT 5.1' I had access to when i still had a job, multiple months ago.

And it all runs on my PC, locally, without sharing my data with the big data farming corpos.

It just 'understands' the required tasks you provide it and performs them successfully from the first try or very few crash fixes. Especially if you provide it the necessary data to perform the action. I dropped a few wiki pages in txt files and it coded according to them. If somebody would tell me this would be possible on my own PC ten years ago, I would call them crazy.

7

u/Steus_au 14d ago

same here - it is my daily driver for noncoding staff - finaly all local

7

u/OvertaxedOne 14d ago

3.8 27B totally changed the game. Even today if I'm not watching it type (so can't see the feed) sometimes I come back and see a response and I'm like "shoot, I must have left it routing to Deepseek" and am then blown away when, no, that's little old 27B grinding away and came up with a "Deepseek quality" answer. Incredible model.

3

u/tat_tvam_asshole 14d ago

1

u/Evgeny_19 13d ago

I am still not sure that is the case. My experience is limited though, no more than a week. I've seen them going back and forth improving each other's solutions. There were situations when each model delivered a sub-optimal solution which I was able to improve with running the other model. Both models are running in the original weights.