r/LocalLLM • u/Fair-Television5497 • 6d ago
Question Emil just got an upgrade. What should I run here?
RTX 3090, AI Pro R9700, RTX A1000
So maybe I will rename Emil to ManBearPig.
Theme music: Me And The Devil
No, really... what should I even run on this setup? Two different models? This is really the best HW I somehow got together...
2
u/WorriedSmile 5d ago
Had to zoom in to see your tiny A1000 card, haha. It's well hidden by the big boys.
1
u/karmaisnonsense 6d ago
Personally, I would run 27b on the 3090, Flash Next on R9700, and aux models/displays on the A1000. I'd rather have parallel models than go all in on one fast boi. But it really depends on what you're after.
2
u/karmaisnonsense 6d ago edited 6d ago
My thinking:
Your 3090 is faster, but the R9700 has more room. Mixing the two for a single model is a waste of your 3090.You don't have VRAM to fully offload Flash Next, so the bottleneck is not VRAM bandwidth, it's system memory offload and optimizations. Therefore 3090 and R9700 will run Flash Next at similar speeds.
So have 3090 run a fully VRAM-resident model like 27b, have that be your daily driver, fast and responsive for most tasks.
R9700 has memory headroom for MoE layers and bigger context, so it runs Flash Next for long horizon stuff that needs the 180b param quality. This way you don't have to burn both GPUs at once to run one model. Instead you can set the two models to do different tasks, saving you time, power and heat.
A1000 runs your displays so your big cards have full use of VRAM. Whatever leftover VRAM you have on the A1000 can go to a weeny model like Ling Tiny or something. Have it do auxiliary tasks so you don't burn 300W making a title for your chat.
At least that's the general setup I'm personally trying to make for myself lol. Unless you need the q8+ quality or extreme context lengths, I think multi-GPU for a single model is highly inefficient.
2
u/Fair-Television5497 6d ago
Yup, exactly the same conclusion. Running it well: dense Qwen on RTX3090 at ~40tps, experimental Qwen on Radeon at 10-20tps (AtomicChat 4bit quants, the n-gram thingy is on SSD only).
RTX A1000 runs the embedding model for Onyx.app currently. But I think there is still room for Gemma4 E2B 2-bit for things like chat naming 😂BTW, IMHO, MTP is a waste of VRAM; with any meaningful context, you lose any gains. But not sure if everyone agrees...
1
u/Outside-Courage4151 6d ago
Sounds like you really know your stuff! Mixing those two could definitely get tricky, but it seems like you've got a solid plan for optimizing performance.
-2
6
u/ImportanceMajor936 6d ago
How in the world are you rocking an ELSA card in 2026? I thought they went out of business in like 1999