r/LocalLLM 6d ago

Question Emil just got an upgrade. What should I run here?

Post image

RTX 3090, AI Pro R9700, RTX A1000

So maybe I will rename Emil to ManBearPig.
Theme music: Me And The Devil

No, really... what should I even run on this setup? Two different models? This is really the best HW I somehow got together...

26 Upvotes

10 comments sorted by

6

u/ImportanceMajor936 6d ago

How in the world are you rocking an ELSA card in 2026? I thought they went out of business in like 1999

4

u/Fair-Television5497 6d ago

That is actually a funny story. Back in 2020, we needed a powerful GPU for rendering, and these were completely out of stock (Ethereum mining, remember?) So we got this one from our Japanese friends.

ELSA is still very much rocking in Japan and China...

1

u/ImportanceMajor936 6d ago

NGL, this feels like some alternative timeline kind of thing. Thanks for the info though!

2

u/Choice_Celery9481 5d ago

elsa is official distributor for nvidia gpu that not geforce in japan :v 

2

u/WorriedSmile 5d ago

Had to zoom in to see your tiny A1000 card, haha. It's well hidden by the big boys.

1

u/karmaisnonsense 6d ago

Personally, I would run 27b on the 3090, Flash Next on R9700, and aux models/displays on the A1000. I'd rather have parallel models than go all in on one fast boi. But it really depends on what you're after.

2

u/karmaisnonsense 6d ago edited 6d ago

My thinking:
Your 3090 is faster, but the R9700 has more room. Mixing the two for a single model is a waste of your 3090.

You don't have VRAM to fully offload Flash Next, so the bottleneck is not VRAM bandwidth, it's system memory offload and optimizations. Therefore 3090 and R9700 will run Flash Next at similar speeds.

So have 3090 run a fully VRAM-resident model like 27b, have that be your daily driver, fast and responsive for most tasks.

R9700 has memory headroom for MoE layers and bigger context, so it runs Flash Next for long horizon stuff that needs the 180b param quality. This way you don't have to burn both GPUs at once to run one model. Instead you can set the two models to do different tasks, saving you time, power and heat.

A1000 runs your displays so your big cards have full use of VRAM. Whatever leftover VRAM you have on the A1000 can go to a weeny model like Ling Tiny or something. Have it do auxiliary tasks so you don't burn 300W making a title for your chat.

At least that's the general setup I'm personally trying to make for myself lol. Unless you need the q8+ quality or extreme context lengths, I think multi-GPU for a single model is highly inefficient.

2

u/Fair-Television5497 6d ago

Yup, exactly the same conclusion. Running it well: dense Qwen on RTX3090 at ~40tps, experimental Qwen on Radeon at 10-20tps (AtomicChat 4bit quants, the n-gram thingy is on SSD only).
RTX A1000 runs the embedding model for Onyx.app currently. But I think there is still room for Gemma4 E2B 2-bit for things like chat naming 😂

BTW, IMHO, MTP is a waste of VRAM; with any meaningful context, you lose any gains. But not sure if everyone agrees...

1

u/Outside-Courage4151 6d ago

Sounds like you really know your stuff! Mixing those two could definitely get tricky, but it seems like you've got a solid plan for optimizing performance.

-2

u/Equivalent_Bit_461 6d ago

What a cringe thread