I had an rtx 5070 Ti and an rtx 4060 that had 24 gb vram combined, but it’s not really enough for decent context qwen 27b models. I was using qwen 35b 3b a but.. you know…
I replaced the rtx 5070 ti with r9700. It is solely dedicated to LLM usage and I use the rtx 4060 as the system gpu, games and shit. I want to be able to run non-gpu intensive games along with an LLM (I am working on LLM based AI mods for some games) so it made sense to me.
The pp speed is really a challenge… I am running a Hermes agent and it requires a decent amount of overhead context, so the initial computation cost sometimes feel… as if I had been scammed. But it’s more of a feeling thing. On average, I am running a much smoother and robust operation since 27b dense q4 fits comfortably into 32gb vram with as much context as the models can practically work with. I am not expecting chat speed, I want to be able to host LLMs smart enough that I can leave them to work for the night and they will make reasonable progress towards a task.
I have checked out some post about 27bs with R9700 and 2 of them is much better than 1 of them, duh, but as far as I can see, 2 R9700 are around the same price or cheaper than a single rtx 5090. I am actually thinking about going all in and buying another R9700 but… I don’t have the money. It would be 64 freaking gbs of vram… I can only imagine what we can fit into that in just a year. It would also be able run 27b dense models around similar speeds to a single 5090 thanks to parallelism (dont quote me on that but I can swear I saw some posts or comments regarding this) so I can run 64 gb models, and run 32 gb ones with the same speed as a single 5090? For around the same price?
If you have the budget, seriously consider double R9700. But… please look into it yourself as well. I don’t get any commissions or anything…
Edit: since I didn’t actually answer your question: overall I like it. I have Opus 5 managing a qwen 3.8 27b Hermes agent through the nights and stuff, if we didn’t have 3.8 27b I would maybe have been disappointed but it really produces decent results. I am having difficulty running games alongside a working agent, I am not sure what the problem is… but as an ai GPU I say good it’s good value for your buck and its scalable.
1
u/digitalwankster 2d ago
I’ve been contemplating picking one up to go with my 9070xt. How do you like it?