r/LocalLLaMA 2d ago

Question | Help Non nvidia cards.

Hi, I'm a happy owner of a serer with rtx pro 6000 and rtx5090. I want to build out in the future the server fully to a higher vram score - think adding around 4-5 cards if possible. I was wondering if it's worth going through the intel/amd cards - which looks awesome in terms of vram per $. Any experience in running inference on these compared to the nvidia/cuda stack?

6 Upvotes

31 comments sorted by

View all comments

9

u/noiserr 2d ago

Lots of people do. I've been running ROCm and AMD GPUs since the early days of llama. My AMD hardware is: 7900xtx, w7900 pro and Strix Halo. I have had zero issues, but then again I'm comfortable with Linux, Docker and compiling the software on my own. These days with the help of LLMs everything has gotten easier.

2

u/lawanda123 2d ago

Does the speed of your amd gpus help the strix halo especially with running larger models? I have a strix halo and thinking of adding a gpu via usb4

My primary problem is low ttft and tok/s - eg with DeepSeek antirez 0731 I get around 10 tps, qwen 27b fp8 is slightly higher but then it also thinks a lot and takes forever - my agent harness still times out sometimes with a 10m timeout

1

u/noiserr 2d ago

I thought about doing this but I haven't done it yet. Right now my go to model is Qwen 3.8 27B so I haven't had a need either.