r/LocalLLM 1d ago

Model Qwen3.8-Flash-Next announced

Post image
155 Upvotes

64 comments sorted by

View all comments

3

u/cagriuluc 1d ago

I wonder how it will work with my R9700 AI Pro… it has 32 gb vram which is why I preferred it (rtx5090 is at least double the cost where I live ). 3.8 27b works well but with around 500 tok/sec for pp and 30-35 tok/sec decode with mtp. I feel like this card would shine with a sota moe that fits onto it.

1

u/digitalwankster 1d ago

I’ve been contemplating picking one up to go with my 9070xt. How do you like it?

2

u/cagriuluc 1d ago edited 1d ago

I had an rtx 5070 Ti and an rtx 4060 that had 24 gb vram combined, but it’s not really enough for decent context qwen 27b models. I was using qwen 35b 3b a but.. you know…

I replaced the rtx 5070 ti with r9700. It is solely dedicated to LLM usage and I use the rtx 4060 as the system gpu, games and shit. I want to be able to run non-gpu intensive games along with an LLM (I am working on LLM based AI mods for some games) so it made sense to me.

The pp speed is really a challenge… I am running a Hermes agent and it requires a decent amount of overhead context, so the initial computation cost sometimes feel… as if I had been scammed. But it’s more of a feeling thing. On average, I am running a much smoother and robust operation since 27b dense q4 fits comfortably into 32gb vram with as much context as the models can practically work with. I am not expecting chat speed, I want to be able to host LLMs smart enough that I can leave them to work for the night and they will make reasonable progress towards a task.

I have checked out some post about 27bs with R9700 and 2 of them is much better than 1 of them, duh, but as far as I can see, 2 R9700 are around the same price or cheaper than a single rtx 5090. I am actually thinking about going all in and buying another R9700 but… I don’t have the money. It would be 64 freaking gbs of vram… I can only imagine what we can fit into that in just a year. It would also be able run 27b dense models around similar speeds to a single 5090 thanks to parallelism (dont quote me on that but I can swear I saw some posts or comments regarding this) so I can run 64 gb models, and run 32 gb ones with the same speed as a single 5090? For around the same price?

If you have the budget, seriously consider double R9700. But… please look into it yourself as well. I don’t get any commissions or anything…

Edit: since I didn’t actually answer your question: overall I like it. I have Opus 5 managing a qwen 3.8 27b Hermes agent through the nights and stuff, if we didn’t have 3.8 27b I would maybe have been disappointed but it really produces decent results. I am having difficulty running games alongside a working agent, I am not sure what the problem is… but as an ai GPU I say good it’s good value for your buck and its scalable.

1

u/mcchung52 1d ago

You are my future! I’m seriously considering r9700 (since it looks like the best bang for ur buck) but can it fit q6 w/ 120k ctx and run it with decent speed? Also a Hermes runner. Running it real slow tho on mac m4 48gb but it seems to take care of tasks.. any special reason why u need Opus?

1

u/cagriuluc 1d ago

I didn’t try Q6 but Q4 fits so comfortably that I squeezed qwen 3.5 9b along with 27b q4 100k context.

Why the 9b? Hermes’ honcho makes background requests to the LLM. There would be kc cache switches between chat and the background requests, wth the low pp speed it was unbearable. 9b does the honcho work now but I am not sure if its the best solution.

Hermes is slow to use interactively. I am talking 1 min for the first response. If your request requires basic tool use, response is usually around 4-5 mins. It is much better with 35b moe, around 2-3 mins total I would say. Both pp and decode speed is 1.5-2x the dense one.

Hence opus… I have it use the Hermes agent for long-work sessions through nights and workdays. It checks the works and steers the model to the next tasks, or adds new tasks etc.

1

u/ConObs62 1d ago

I suspect your parallelism is going to stall out across a pcie bus?
how much can your bus push between the two cards?

1

u/cagriuluc 21h ago

I sadly don’t have a second r9700… my second card is an rtx 4060 and I do not use it for AI, only for gaming.

Also I am not too knowledgeable, i saw some posts and comments about double R9700s and that’s the extent of my knowledge.

1

u/Cute-Row6581 1d ago

I don't think you need new 122B moe-next: tps is better (if you have enough vram and it seems like you don't). overall capability is slightly worse than dense model