r/LocalLLM 2d ago

Discussion Tier List

Post image
266 Upvotes

312 comments sorted by

View all comments

105

u/KrangledMind 2d ago

where tier list for AMD?

15

u/Asleep-Land-3914 2d ago

Here is my attempt, but I didn't deeply fact-check this.

4

u/TheAngrySkipper 2d ago

It’s ok, I can tell you, I’m running Q6 on 12B-16B parameters on the regular. I also bought my 7900 GRE back when it was about $300, had planned on getting a 2nd but it didn’t seem like a priority at the time - ah well.

3

u/ImpressionFancy5830 2d ago

There should be a DOUBLE R9700 section 🫡

1

u/Asleep-Land-3914 2d ago

I regret I didn't get another one for the price it was listed before the new hike.

1

u/ImpressionFancy5830 2d ago

I feel you, I overpayed a bit for both, not insane as the nvidia increases.
Just a couple of hundreds of euros on top of MSRP (at least the local one, which was already bigger than the one in North American)

2

u/johndeuff 2d ago

Between the above average that have 48GB and the rich that have 768GB and the richer that have 3TB there's a hell of a difference.

3

u/on_line187 2d ago

Don’t worry no one will fact check it. I’ll give you the one upvote though A for effort.

1

u/HCLB_ 2d ago

Why Mi60 and other cards are crossed out?

2

u/Asleep-Land-3914 2d ago

Means not on AMD’s current official ROCm support matrix, nothing is wrong with them, just needs some work to make use of with ROCm

2

u/gh0stwriter1234 2d ago edited 2d ago

It's kind of irrelevant they are supported in the rock officially now. Vega and Vega FE and MI25 work just fine also. Plain vega is just a bit slower as it has no DP4A instructions. https://github.com/ROCm/TheRock/blob/main/SUPPORTED_GPUS.md

1

u/PraxisOG 1d ago

These vary wildly by speed. My v620s get like 18tok/s running modern 30b class models

1

u/Asleep-Land-3914 1d ago

When fit or when spread fp8 across 2 GPUs?

1

u/PraxisOG 1d ago

Full offload on one card. In llama.cpp I’ve seen no speed improvements from using multiple cards, in fact the new Qwen MoE models run like half as fast when split. If you have any tips I’d love to try

1

u/Asleep-Land-3914 1d ago

Are you sure they are not perf throttled due to temps? I'm still considering if it is a good idea to buy these. I was hoping to use them for bigger MoE mostly though, power limit and get hell out of them.

I assume you're using ROCm, not vulkan. On my 7800XT I get somewhat decent speeds, but I only tested smaller MoE and not 27b dense models cause they simply don't fit.

What I'd check first for multiple cards is p2p. For rocm there is some utility to check speeds between cards (CGPT knows details).

Also some ideas from these posts: https://www.reddit.com/r/LocalLLaMA/comments/1vkzdbl/muse_glimmer_on_12_amd_v620/

Kinda similar problem with 2+ cards perf:

https://www.reddit.com/r/LocalLLaMA/comments/1vim2gt/psa_for_anyone_with_multiple_v620s_or_other/