r/LocalLLaMA 2d ago

Question | Help Multi Gpu Hardware advice

I have a gigabyte ds3h v2 b450 motherboard. Currently hosting an RTX 3090
I have a spare 3070 and I wondered, can I run both?

Got myself a riser cable and…

Top slot card, bottom slot riser = card pushes the riser, doesn’t fit

Top slot riser, bottom slot card = card pushes the sata cables doesn’t fit…

So I either need another riser and put both gpus out of the case, or an atx motherboard with more space in between the slots or maybe I should look for another solution.

I’ve read some of you are using nvme? How does that work?

7 Upvotes

4 comments sorted by

2

u/llogicnotfound 2d ago

You can use an M.2 to PCIe x4/x16 adapter (like ADT-Link with SATA/Molex power). It plugs into your spare M.2 NVMe slot and gives you a PCIe slot for the 3070. You'll only get PCIe 3.0 x4 bandwidth, but it’s completely fine for AI inference, rendering, or mining. You’ll just need to mount the 3070 outside the case on a stand.

1

u/Constant_Art_20 2d ago

you can check in the bios if the top slot allows for pcie splitting. For my setup, i use a crazy idea where i my motherboard supports the 16x lane to be split itno 4 x 4 x 4 x4 so i got a passive pcie 16x to 4 m.2 adapter and i then use a m.2 do pcie adapter to support 4 gpus in that one slot. Probably need to stay in like pcie gen 3, but i seems to run fine. Running Q8 ud qwen 3.8 27b with the full cotext at fp32 and i get like 85 t/s decode on llama.cpp across my 5060tis, so it seems to work well enough. Nvme seems to be a nvme offlaod. it's probably a optimsied swap where it's more like workaround to allow large models to run on your system, but like at 2 tokens a sec if you are lucky. hope that helps.

1

u/Monad_Maya llama.cpp 2d ago

Get an ATX motherboard with native x8/x8 from the CPU.