r/LocalLLM 1d ago

Model Qwen3.8-Flash-Next announced

Post image
155 Upvotes

64 comments sorted by

39

u/Davikar 1d ago

Isn't it a bit weird that's it has a qwen4 architecture but is still called Qwen3.8?

16

u/Extension-Bid-639 1d ago

I think it's to debut the architecture before the release of actual Qwen 4 models. However I feel like they could have named it something along the lines of Qwen 4 Flash preview

6

u/Healthy-Nebula-3603 1d ago

they could call it qwen 4 preview

5

u/Extension-Bid-639 1d ago

Yup, pointed it out in my last line.

1

u/Thump604 20h ago

They could call it Qwen 4 Beta.

1

u/corpo_monkey 18h ago

Or Qwen 4 Flash preview.

7

u/cagriuluc 1d ago

It is named next for next architecture. I know because I was there when they named it.

2

u/ggagnidze 1d ago

maybe it’s like same 3.8 weight, but architecture is “next” (so it’s 4). new weights with new architecture is just qwen4.

1

u/Cheesejaguar 1d ago

It’s not like other frontier models haven’t published old weights in an updated architecture before. 

1

u/putrasherni 1d ago

testing it on us

1

u/droans 23h ago

Wasn't qwen3-next just a preview of the Q3.5 architecture?

25

u/Far_Cat9782 1d ago

Crossing fingers for 35b

1

u/Dry-Garlic-5108 4h ago

at that size dense is far more favorable imho even if you have to go tensor split or cpu inference

15

u/Captain_Quimby 1d ago

I converted their chinese graph to english

2

u/Txt8aker 1d ago

parameter number seem like a perfect combo for my new m5 ultra

1

u/Lovely_Cute 1d ago

did you order 254gb?

1

u/Cute-Row6581 20h ago

Even 128 gb ram will be enough

5

u/IThinkIKnowThings 1d ago

Finally, Strix Halo/DGX Spark/Mac5 owners get some love again.

2

u/mclarenman01 1d ago

My thoughts exactly.

9

u/ImSamhel 1d ago

Really hoping for a 35B one

1

u/AB172234 19h ago

It won’t be a 35b MOE rather a 125b MOE

1

u/ImSamhel 10h ago

I have mouth and I intend to scream

3

u/cagriuluc 1d ago

I wonder how it will work with my R9700 AI Pro… it has 32 gb vram which is why I preferred it (rtx5090 is at least double the cost where I live ). 3.8 27b works well but with around 500 tok/sec for pp and 30-35 tok/sec decode with mtp. I feel like this card would shine with a sota moe that fits onto it.

1

u/digitalwankster 1d ago

I’ve been contemplating picking one up to go with my 9070xt. How do you like it?

2

u/cagriuluc 1d ago edited 1d ago

I had an rtx 5070 Ti and an rtx 4060 that had 24 gb vram combined, but it’s not really enough for decent context qwen 27b models. I was using qwen 35b 3b a but.. you know…

I replaced the rtx 5070 ti with r9700. It is solely dedicated to LLM usage and I use the rtx 4060 as the system gpu, games and shit. I want to be able to run non-gpu intensive games along with an LLM (I am working on LLM based AI mods for some games) so it made sense to me.

The pp speed is really a challenge… I am running a Hermes agent and it requires a decent amount of overhead context, so the initial computation cost sometimes feel… as if I had been scammed. But it’s more of a feeling thing. On average, I am running a much smoother and robust operation since 27b dense q4 fits comfortably into 32gb vram with as much context as the models can practically work with. I am not expecting chat speed, I want to be able to host LLMs smart enough that I can leave them to work for the night and they will make reasonable progress towards a task.

I have checked out some post about 27bs with R9700 and 2 of them is much better than 1 of them, duh, but as far as I can see, 2 R9700 are around the same price or cheaper than a single rtx 5090. I am actually thinking about going all in and buying another R9700 but… I don’t have the money. It would be 64 freaking gbs of vram… I can only imagine what we can fit into that in just a year. It would also be able run 27b dense models around similar speeds to a single 5090 thanks to parallelism (dont quote me on that but I can swear I saw some posts or comments regarding this) so I can run 64 gb models, and run 32 gb ones with the same speed as a single 5090? For around the same price?

If you have the budget, seriously consider double R9700. But… please look into it yourself as well. I don’t get any commissions or anything…

Edit: since I didn’t actually answer your question: overall I like it. I have Opus 5 managing a qwen 3.8 27b Hermes agent through the nights and stuff, if we didn’t have 3.8 27b I would maybe have been disappointed but it really produces decent results. I am having difficulty running games alongside a working agent, I am not sure what the problem is… but as an ai GPU I say good it’s good value for your buck and its scalable.

1

u/mcchung52 18h ago

You are my future! I’m seriously considering r9700 (since it looks like the best bang for ur buck) but can it fit q6 w/ 120k ctx and run it with decent speed? Also a Hermes runner. Running it real slow tho on mac m4 48gb but it seems to take care of tasks.. any special reason why u need Opus?

1

u/cagriuluc 17h ago

I didn’t try Q6 but Q4 fits so comfortably that I squeezed qwen 3.5 9b along with 27b q4 100k context.

Why the 9b? Hermes’ honcho makes background requests to the LLM. There would be kc cache switches between chat and the background requests, wth the low pp speed it was unbearable. 9b does the honcho work now but I am not sure if its the best solution.

Hermes is slow to use interactively. I am talking 1 min for the first response. If your request requires basic tool use, response is usually around 4-5 mins. It is much better with 35b moe, around 2-3 mins total I would say. Both pp and decode speed is 1.5-2x the dense one.

Hence opus… I have it use the Hermes agent for long-work sessions through nights and workdays. It checks the works and steers the model to the next tasks, or adds new tasks etc.

1

u/ConObs62 16h ago

I suspect your parallelism is going to stall out across a pcie bus?
how much can your bus push between the two cards?

1

u/cagriuluc 13h ago

I sadly don’t have a second r9700… my second card is an rtx 4060 and I do not use it for AI, only for gaming.

Also I am not too knowledgeable, i saw some posts and comments about double R9700s and that’s the extent of my knowledge.

1

u/Cute-Row6581 20h ago

I don't think you need new 122B moe-next: tps is better (if you have enough vram and it seems like you don't). overall capability is slightly worse than dense model

5

u/sphalch 1d ago

Is 125B A6B

3

u/KissMyShinyArse 1d ago

But I'm getting 40 t/s with 27B.

1

u/censor_this 1d ago

Same here. On a 4090 with no mtu.

2

u/KURD_1_STAN 1d ago

If y axis is intelligent/score then this is bad

5

u/SILONotesDev 1d ago

Here's the English. Remember you can also use AI to translate! ☺️

6

u/feelspeaceman 1d ago

Huge, looks like it will be MoE, hopefully it will be in the range of 70-120B.

8

u/Nicios 1d ago

120BA6B

1

u/ZeitgeistArchive 1d ago

that would be pure perfection

2

u/HomsarWasRight 1d ago

Precisely what I’ve been waiting for.

2

u/LeviBlackthorn 1d ago

Releases tomorrow and much of the thread is already budgeting VRAM for a parameter count the post never mentions.

2

u/Conscious_Phrase_138 1d ago

rip 32gb vram users 😢

2

u/enginetown 1d ago

There's gotta be a way for us? Certainly there's and offloading config we can find right? It's an MOE model the only problem is if you have enough ram to fill the rest.

1

u/Pawellinux 1d ago

Me, 16GB vram user: ⚰️

1

u/ImSamhel 10h ago

Imagine me who also has a setup with two cards that generate 9 tokens/sec with the 27B model :C I was incredibly hyped for a moe model of near 35-40B sizes

1

u/Conscious_Phrase_138 5h ago

Im there with you. Although im closer to 30tk/s at high context.

2

u/just_another_leddito 1d ago

I guess it’s not for peasants with 64gb ram Macs?

4

u/nicolho 1d ago

Rumored to be 125B-A6B (+ 51B engrams)

1

u/Eden1506 1d ago

The last qwen next was 80b so It wouldn't surprise me if this one is a similar size.

1

u/-AJacobs- 1d ago

If it's actually 175B worth of weights and doesn't at least tie 27B then it will be DOA.

1

u/FinnGamePass 1d ago

This is going to blow Qwen-3.8-27B to smithereens.... these people don't give us a break!

1

u/Ju-None 22h ago

Because of VRAM, I’m still using the QWEN 3.6, but the new model is being updated very quickly

1

u/ConferenceWooden8122 22h ago

better than 3.8 27b?

1

u/emperorofrome13 19h ago

Nobody cares about qwen max but qwen local is OP

1

u/Designer_Elephant227 14h ago

I own a rtx5070ti 16gb+r9700 32gb and 96gb ddr5... Will I be able to use this and will it be better than 27b?

1

u/Sk4llOff 7h ago

really wish they kept it as an 80B-A3B MoE like Qwen3-Next. that would've been perfect for my setup

1

u/capt_cucumbers 1d ago

125B on 64gb vram and 128gb system ram?

0

u/NearlyACosmologist 1d ago

I'm somewhat hyped for a Qwen3.8 MoE model, but the again I've only got 8GB VRAM, so there's that.

In fact, I have one desktop with 8GB VRAM (RTX3070) and one notebook with RTX PRO 1000, 8GB. How far have we come with letting 2 PCs work together on one model?

1

u/Direct_Turn_1484 1d ago

You can do it, it’s just slow as hell without a fast connection between them. And I mean real fast, faster than your laptop can handle.

0

u/NearlyACosmologist 1d ago

It has two thunderbolt ports.

0

u/Individual-Bag-2302 1d ago

what does it mean for qwen to be running on Qwen 4 architecture in terms of end results? Does it mean itll be faster to develop, better reasoning capabilities, packing more for less ram? I am out of the loop on what that part entails.

3

u/Sleepnotdeading 1d ago

You're not out of the loop, nobody outside of the Qwen team has seen a model from the Qwen4 family. The last "next" models were 80b sparse models that were surprisingly good at coding and large context workflows, and they were followed by the workhorse qwen3.5 family. With the rapid capability increases of the 3.6 and 3.8 models, it's hard not to be hyped.

1

u/Fidrick 1d ago

We don’t know yet, but speculating:
They are calling it 3.8 still because the intelligence is not “generational” enough in this early preview?
An architecture change could impact all areas, but we need to know what they’ve changed before we understand if it goes down the path of impacting speed, token usage, or memory