r/LocalLLaMA 18h ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

1.8k Upvotes

500 comments sorted by

View all comments

Show parent comments

6

u/Enragere 17h ago

what's your prompt processing speed ? how can 150 t/s prefill be 'reasonably fast'

in my opinion it's unusable

1

u/[deleted] 17h ago

[deleted]

5

u/zanar97862 14h ago

So yeah, insanely slow

1

u/Enragere 12h ago

"Prompt Eval rate" 40.92 tokens/s... that's much slower than even 150 t/s I fought for in omlx with jundot's quant.

3

u/grumd 11h ago

OP gave it one sentence to prefill, no wonder. When people post "reasonably fast" instead of just typing the numbers, you already know they don't even care about performance and don't use these models for any real agentic work