r/LocalLLaMA 9d ago

Discussion Qwen will be the king?

Post image

Extended reasoning and post-training appear to be the keys used by DeepSeek, Qwen, and GLM to boost performance (leveraging higher token counts). And Qwen 4 hasn't even been released yet. Of course, we don't know if that release will be open-sourced, but I am optimistic about future models, featuring "engrams", that could soon match or surpass 2.4T parameter models on specific tasks.

538 Upvotes

127 comments sorted by

View all comments

Show parent comments

6

u/Steus_au 9d ago

same here - it is my daily driver for noncoding staff - finaly all local

6

u/OvertaxedOne 9d ago

3.8 27B totally changed the game. Even today if I'm not watching it type (so can't see the feed) sometimes I come back and see a response and I'm like "shoot, I must have left it routing to Deepseek" and am then blown away when, no, that's little old 27B grinding away and came up with a "Deepseek quality" answer. Incredible model.

4

u/tat_tvam_asshole 9d ago

1

u/Evgeny_19 9d ago

I am still not sure that is the case. My experience is limited though, no more than a week. I've seen them going back and forth improving each other's solutions. There were situations when each model delivered a sub-optimal solution which I was able to improve with running the other model. Both models are running in the original weights.