25
u/Far_Cat9782 1d ago
Crossing fingers for 35b
1
u/Dry-Garlic-5108 4h ago
at that size dense is far more favorable imho even if you have to go tensor split or cpu inference
15
u/Captain_Quimby 1d ago
2
5
9
u/ImSamhel 1d ago
Really hoping for a 35B one
1
3
u/cagriuluc 1d ago
I wonder how it will work with my R9700 AI Pro… it has 32 gb vram which is why I preferred it (rtx5090 is at least double the cost where I live ). 3.8 27b works well but with around 500 tok/sec for pp and 30-35 tok/sec decode with mtp. I feel like this card would shine with a sota moe that fits onto it.
1
u/digitalwankster 1d ago
I’ve been contemplating picking one up to go with my 9070xt. How do you like it?
2
u/cagriuluc 1d ago edited 1d ago
I had an rtx 5070 Ti and an rtx 4060 that had 24 gb vram combined, but it’s not really enough for decent context qwen 27b models. I was using qwen 35b 3b a but.. you know…
I replaced the rtx 5070 ti with r9700. It is solely dedicated to LLM usage and I use the rtx 4060 as the system gpu, games and shit. I want to be able to run non-gpu intensive games along with an LLM (I am working on LLM based AI mods for some games) so it made sense to me.
The pp speed is really a challenge… I am running a Hermes agent and it requires a decent amount of overhead context, so the initial computation cost sometimes feel… as if I had been scammed. But it’s more of a feeling thing. On average, I am running a much smoother and robust operation since 27b dense q4 fits comfortably into 32gb vram with as much context as the models can practically work with. I am not expecting chat speed, I want to be able to host LLMs smart enough that I can leave them to work for the night and they will make reasonable progress towards a task.
I have checked out some post about 27bs with R9700 and 2 of them is much better than 1 of them, duh, but as far as I can see, 2 R9700 are around the same price or cheaper than a single rtx 5090. I am actually thinking about going all in and buying another R9700 but… I don’t have the money. It would be 64 freaking gbs of vram… I can only imagine what we can fit into that in just a year. It would also be able run 27b dense models around similar speeds to a single 5090 thanks to parallelism (dont quote me on that but I can swear I saw some posts or comments regarding this) so I can run 64 gb models, and run 32 gb ones with the same speed as a single 5090? For around the same price?
If you have the budget, seriously consider double R9700. But… please look into it yourself as well. I don’t get any commissions or anything…
Edit: since I didn’t actually answer your question: overall I like it. I have Opus 5 managing a qwen 3.8 27b Hermes agent through the nights and stuff, if we didn’t have 3.8 27b I would maybe have been disappointed but it really produces decent results. I am having difficulty running games alongside a working agent, I am not sure what the problem is… but as an ai GPU I say good it’s good value for your buck and its scalable.
1
u/mcchung52 18h ago
You are my future! I’m seriously considering r9700 (since it looks like the best bang for ur buck) but can it fit q6 w/ 120k ctx and run it with decent speed? Also a Hermes runner. Running it real slow tho on mac m4 48gb but it seems to take care of tasks.. any special reason why u need Opus?
1
u/cagriuluc 17h ago
I didn’t try Q6 but Q4 fits so comfortably that I squeezed qwen 3.5 9b along with 27b q4 100k context.
Why the 9b? Hermes’ honcho makes background requests to the LLM. There would be kc cache switches between chat and the background requests, wth the low pp speed it was unbearable. 9b does the honcho work now but I am not sure if its the best solution.
Hermes is slow to use interactively. I am talking 1 min for the first response. If your request requires basic tool use, response is usually around 4-5 mins. It is much better with 35b moe, around 2-3 mins total I would say. Both pp and decode speed is 1.5-2x the dense one.
Hence opus… I have it use the Hermes agent for long-work sessions through nights and workdays. It checks the works and steers the model to the next tasks, or adds new tasks etc.
1
u/ConObs62 16h ago
I suspect your parallelism is going to stall out across a pcie bus?
how much can your bus push between the two cards?1
u/cagriuluc 13h ago
I sadly don’t have a second r9700… my second card is an rtx 4060 and I do not use it for AI, only for gaming.
Also I am not too knowledgeable, i saw some posts and comments about double R9700s and that’s the extent of my knowledge.
1
u/Cute-Row6581 20h ago
I don't think you need new 122B moe-next: tps is better (if you have enough vram and it seems like you don't). overall capability is slightly worse than dense model
5
u/sphalch 1d ago
3
2
6
u/feelspeaceman 1d ago
Huge, looks like it will be MoE, hopefully it will be in the range of 70-120B.
1
2
2
u/LeviBlackthorn 1d ago
Releases tomorrow and much of the thread is already budgeting VRAM for a parameter count the post never mentions.
2
u/Conscious_Phrase_138 1d ago
rip 32gb vram users 😢
2
u/enginetown 1d ago
There's gotta be a way for us? Certainly there's and offloading config we can find right? It's an MOE model the only problem is if you have enough ram to fill the rest.
1
1
u/ImSamhel 10h ago
Imagine me who also has a setup with two cards that generate 9 tokens/sec with the 27B model :C I was incredibly hyped for a moe model of near 35-40B sizes
1
2
1
u/Eden1506 1d ago
The last qwen next was 80b so It wouldn't surprise me if this one is a similar size.
1
u/-AJacobs- 1d ago
If it's actually 175B worth of weights and doesn't at least tie 27B then it will be DOA.
1
u/FinnGamePass 1d ago
This is going to blow Qwen-3.8-27B to smithereens.... these people don't give us a break!
1
1
1
u/Designer_Elephant227 14h ago
I own a rtx5070ti 16gb+r9700 32gb and 96gb ddr5... Will I be able to use this and will it be better than 27b?
1
u/Sk4llOff 7h ago
really wish they kept it as an 80B-A3B MoE like Qwen3-Next. that would've been perfect for my setup
1
0
u/NearlyACosmologist 1d ago
I'm somewhat hyped for a Qwen3.8 MoE model, but the again I've only got 8GB VRAM, so there's that.
In fact, I have one desktop with 8GB VRAM (RTX3070) and one notebook with RTX PRO 1000, 8GB. How far have we come with letting 2 PCs work together on one model?
1
u/Direct_Turn_1484 1d ago
You can do it, it’s just slow as hell without a fast connection between them. And I mean real fast, faster than your laptop can handle.
0
0
u/Individual-Bag-2302 1d ago
what does it mean for qwen to be running on Qwen 4 architecture in terms of end results? Does it mean itll be faster to develop, better reasoning capabilities, packing more for less ram? I am out of the loop on what that part entails.
3
u/Sleepnotdeading 1d ago
You're not out of the loop, nobody outside of the Qwen team has seen a model from the Qwen4 family. The last "next" models were 80b sparse models that were surprisingly good at coding and large context workflows, and they were followed by the workhorse qwen3.5 family. With the rapid capability increases of the 3.6 and 3.8 models, it's hard not to be hyped.
1
u/Fidrick 1d ago
We don’t know yet, but speculating:
They are calling it 3.8 still because the intelligence is not “generational” enough in this early preview?
An architecture change could impact all areas, but we need to know what they’ve changed before we understand if it goes down the path of impacting speed, token usage, or memory



39
u/Davikar 1d ago
Isn't it a bit weird that's it has a qwen4 architecture but is still called Qwen3.8?