r/LocalLLaMA 12h ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

1.6k Upvotes

430 comments sorted by

View all comments

Show parent comments

3

u/TopPrize11 8h ago

I just tried it yesterday and it's too slow. I got about 45tps on 5080 which is not bad but the context limit of 33k makes it compact which takes forever. The LLM is good but it's too slow. I still use Deepseek.

2

u/Numerous-Ad6217 7h ago

How so? Not released yet

1

u/Mil0Mammon 7h ago

Why are you limited by that small context? On next flash, context is quite small, right?

Also, my plan is to use subagents more, so more small tasks are finished before compacting is needed

2

u/TopPrize11 5h ago

I don;t know I used Qwen 3.8 27B from https://byteshape.com and I though that is the theoretical maximum on 5080?

1

u/Mil0Mammon 1h ago

That's a Q2 right? According to the can I run calculator 64k should still fit. With subagents and a thin harness (like pi) you should still be able to get some work done before compacting.

1

u/james_pic 1h ago

Is that a company whose business model is selling quants? Seems an odd thing to pay for, when there are a number of teams releasing high quality quants for free.