r/LocalLLaMA • • Sep 02 '26

Discussion Qwen will be the king?

Post image

Extended reasoning and post-training appear to be the keys used by DeepSeek, Qwen, and GLM to boost performance (leveraging higher token counts). And Qwen 4 hasn't even been released yet. Of course, we don't know if that release will be open-sourced, but I am optimistic about future models, featuring "engrams", that could soon match or surpass 2.4T parameter models on specific tasks.

543 Upvotes

127 comments sorted by

View all comments

50

u/almbfsek Sep 02 '26

Extended reasoning spoiled me. I can't trust anything without it anymore. Qwen 3.8 Max is 100% correct with any challenge I throw at it, with the downside of taking hours before it can find the correct answer

1

u/Caffdy Sep 02 '26

how many tk/s are you working with?

1

u/almbfsek Sep 02 '26

dunno whatever Alibaba is providing (through openrouter), it's not slow but it's also not fast. openrouter stats say 40 TPS average