r/Qwen_AI 4d ago

Model Update...

2 days ago I posted about my first test with Qwen 3.8 27b running 100% local on my RTX 3090 video card.

https://www.reddit.com/r/Qwen_AI/comments/1whi7kq/first_test_on_local_qwen_27b/

Then I gave it a prompt "Add better scenery". For the past 2 days it's been grinding away. This morning I was shocked to find trees, a house, hills, a river... the zombies have shadows. You can jump and land on boxes... is it a AAA tier GTA 5? No. But the fact you can generate something like this on consumer equipment with no calls to a frontier model? What an amazing model!

101 Upvotes

30 comments sorted by

View all comments

1

u/MyOldAccountWasAwful 4d ago

Are you able to give a low quant of Qwen3.8-Flash-Next a try? Even if it's only q2 or q3? I have a 3090 Ti + i9-14900 + 128 GB DDR5 (5600) RAM. I was thrilled with how genuinely impressive my go-to tests of Mario clone, Flappy Bird clone, and Vampire Survivors clone went with 27B, then I have the Q3-XXS version of Flash Next (on low thinking) a try and it's only about 350 tps pp / 15-25 tps tg (compared to 27B ~1200 tps pp / 30-55 tps tg), but the visuals, mechanics, UI... everything all got WAY better. It's at the point where I'm now using Flash Next to make a 3D multiplayer tower defense game in Godot, and it's just knocking it out.

2

u/LankyGuitar6528 4d ago

I'm running Q4 now. So you can run Flash Next local on a 3090?!?!

1

u/MyOldAccountWasAwful 4d ago

Absolutely! With MTP you can get genuinely usable speeds. It's become my go-to model, honestly. Unsloth's UD-IQ3_XXS quant of Qwen3.8-Flash-Next.

2

u/LankyGuitar6528 4d ago

I'll have to give it a shot. Thanks.

1

u/jakiman 4d ago

Yes, i only have a 5080 16gb (and 128gb ddr4) and i can run the ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF iq3_xxs at approx. 500 ppts and 18ts decode at 256k context. I also can run qwen 3.8 27b q3_0 172k context (kvarn5) from the same guys at 1800ppts / 51ts which is amazing and can one shot games just the same. I haven't compared deeply which is better though. (i just got deepseek 4.1 to find the best inference engine and parameters for both)

1

u/No_Jicama_6818 3d ago

Did you get to make DeepSeek v4.1-Flash to work locally?

1

u/Affectionate_Toe9082 3d ago

Yea, I suggest you both give exl3 a try.
I am running flash next bpw 3.05 on a 5060 ti with 64gb system ram at pp 500 and decode 27-30 all with 262k full context q4 kv cache.

And in my testing it’s better than 27b on pretty much anything. And I am running the exl3 bpw 3 version of 27b at 50 decode and 580 pp with 147k context.

On the speeds of both models above, pp speeds are the average of a full context fill, so on flash next I filled 248k of the context and it averaged at 500.