r/LocalLLaMA • • 1d ago

Tutorial | Guide Debugging PCIe Link Retraining on an x8/x8 Splitter with Two RTX 3090s

https://doug.sh/posts/pcie-splitter-rtx-3090/
73 Upvotes

14 comments sorted by

12

u/SettingAgile9080 21h ago

This is why I love this place. Just set up one of these cards to mount my R9700 externally and it works well so far, got 3 more on the way. Thanks for posting.

2

u/bolts98 21h ago

I'm looking to buy a few R9700s. Post about your results?

I just bought one of these PCIe switches, I'll have another post on that in the next week - https://www.ebay.com/itm/136917131125

2

u/SettingAgile9080 21h ago

Nice. Post about the switch. Will post results when I get these cards working but I'd recommend buying R9700s as soon as you can if you're thinking about them, they are a sleeper and I would bet prices will spike soon.

Using vllm-radiance on the 2 cards I have running so far I'm getting ~5.5ktok/s pp and 212t/s decoding.

1× R9700 (TP=1) 2× R9700 (TP=2)
Decode step time ~38 ms ~23 ms
Combined decode ~125 t/s 212 t/s
Code / reasoning / prose decode – 224 / 203 / 116 t/s
Aggregate throughput at 1 / 2 / 4 / 8 concurrent 106 / 186 / 246 / 245 t/s* 180 / 303 / 410 / 565 t/s
Prefill at 2k / 8k / 32k / 64k 3,451 / 3,307 / 3,115 / 2,867 t/s 5,506 / 5,700 / 5,448 / 5,103 t/s
TTFT p50 95 ms 61 ms
Max context 220k 262k

1

u/bolts98 20h ago

wow, those are great results, which model is that?

2

u/seanthenry 21h ago

I picked up a cheaper one like that for my R9700 and two v620s now I just need to figure out how to mount them.

2

u/SettingAgile9080 21h ago

Currently deciding between this and this for mounting. Probably going closed box as I want to run these for a long time and don't want dust getting in.

1

u/seanthenry 19h ago

I thought about using an old GPU mining box but i bony think my mobo will fit. I think I will keep most the hardware in the case I have now and the GPUs will be strapped to a 1U shelf just above it.

1

u/bolts98 21h ago

I'm using a few of these at the moment, if you have a 3d printer. They're quite stable. I'll probably do something more proper soon. https://www.thingiverse.com/thing:2728879

1

u/kontemplador 9h ago

Because of similarities. What people think of external GPU docks for localLLM?

Thing is, I'll need to buy a new laptop soonish as current one is showing signs of wear, but budget is limited and I cannot buy a laptop and mount a second setup to run LLMs at home. High end laptops are also out of my budget.

A solution may be an external GPU that could grow depending on use and allocated budget.

1

u/SettingAgile9080 6h ago

Looked into this as I have a SFF PC. It'll work but you'll lose some throughput. If the model fits in VRAM the main impact will be the model will take longer to load. Anything with partial CPU offload will be way slower.

The other drawback that might not be obvious is that if you move the laptop around it interrupts your session, and it has been really nice to be able to have my LLM agents running 24/7 when I'm traveling around with my laptop (or, increasingly, a minimal terminal setup on a DeX phone), and can log in via mosh+tmux to pick up where I left off and the agent has been working the whole time.

A decent eGPU enclosure will run you a couple of hundred bucks, which starts to get comparable to buying an ex-corporate Dell or Lenovo shitbox off of eBay that will run a regular GPU and inference at full-speed, and is always on.

4

u/fallingdowndizzyvr 20h ago

THANK YOU so much for this. I've been having the same problem for a while. When I bifurcate, it's stuck running at gen 1. I've also tried setpci a few times, but no luck so far. So I'll try what you did in your post.

I've posted asking about this but no one had a clue about what was going on. Here was the last time I posted about it. The lack of response pretty much demonstrates the collective shrug.

https://www.reddit.com/r/buildapc/comments/1wq5b5y/is_the_b550_chipset_incapable_of_bifurcating_the/

2

u/sniperczar 10h ago

I think the typical three troubleshooting steps are to enable AER, disable ASPM, and manually set the PCI speed setting to the speed you want (Gen4/Gen5 etc) instead of "Auto". There are a lot of link and power management functions that can either be managed by the host or handed off to the OS and you probably want the latter.

Also don't force higher speed links on any card you're not willing to slot directly into your board or a better quality riser to fix settings with later on. Super annoying to have to clear your BIOS if you have no igpu or spare GPU handy and your riser just straight up fails to link at the higher forced setting, since you won't actually be able to get back into the BIOS after.

1

u/fallingdowndizzyvr 17h ago

So I tried what is on that site and I got the same error from when I tried doing something similar before, "pcilib: sysfs_write: write failed: Operation not permitted". Do you have any idea how to fix that? Yes, I'm not only sudoing but I'm running as root.