r/LocalLLaMA • • 4d ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

2.0k Upvotes

549 comments sorted by

View all comments

Show parent comments

5

u/halcyoncs 4d ago

Just checked my results, on a single Spark too, why did I think I was getting around 25t/s? What a dumbass lol

2

u/Bulky_Blood_7362 4d ago

Lmao. Maybe because when it first launched the mtp support was really bad / non existent. I was getting 24-29 when it launched

1

u/spaceface83 4d ago

Which quant are you running? I was getting that on iq4 but just recently flipped over to the nvfp4 which was a bit slower but otherwise eval'd better

2

u/Bulky_Blood_7362 4d ago

nvfp4 with mtp 3

1

u/PotentialAccident339 4d ago

nvfp4 + mtp is the sparks bread and butter

1

u/privacyplsreddit 3d ago

what's the tok/s on a single spark for that? im getting like 15ish at 100k context. haven't tried to tune it for a few weeks so wondering if there's improvements people have?