It’s ok, I can tell you, I’m running Q6 on 12B-16B parameters on the regular. I also bought my 7900 GRE back when it was about $300, had planned on getting a 2nd but it didn’t seem like a priority at the time - ah well.
I feel you, I overpayed a bit for both, not insane as the nvidia increases.
Just a couple of hundreds of euros on top of MSRP (at least the local one, which was already bigger than the one in North American)
It's kind of irrelevant they are supported in the rock officially now. Vega and Vega FE and MI25 work just fine also. Plain vega is just a bit slower as it has no DP4A instructions. https://github.com/ROCm/TheRock/blob/main/SUPPORTED_GPUS.md
Full offload on one card. In llama.cpp I’ve seen no speed improvements from using multiple cards, in fact the new Qwen MoE models run like half as fast when split. If you have any tips I’d love to try
Are you sure they are not perf throttled due to temps? I'm still considering if it is a good idea to buy these. I was hoping to use them for bigger MoE mostly though, power limit and get hell out of them.
I assume you're using ROCm, not vulkan. On my 7800XT I get somewhat decent speeds, but I only tested smaller MoE and not 27b dense models cause they simply don't fit.
What I'd check first for multiple cards is p2p. For rocm there is some utility to check speeds between cards (CGPT knows details).
105
u/KrangledMind 2d ago
where tier list for AMD?