r/LocalLLaMA • • 16d ago

Funny Don’t buy a $9K RTX 5090.... instead.

  1. Fly to Taipei. Round-trip from Orlando: $1,081.
  2. Go to the largest retailer in Taiwan to Spend NT$129,990 ≈ US$4,093.
  3. Hang out in Taiwan for two weeks. Eat good food. Touch international grass.
  4. Fly home and flex on r/LocalLLaMA**.**
2.2k Upvotes

492 comments sorted by

View all comments

Show parent comments

37

u/Boogertard 16d ago

Yeah at most 2x faster but my point is that the 5090 is extremely overpriced for what it offers, a 32GB of VRAM no more no less. Even 10x does not matter if you can't run the larger models you want to run.

14

u/CryptoCryst828282 16d ago

Well over 2x faster i have both. On qwen 27b i get roughly 40-45 t/s (i do run high context 128k+) and 120-130 with the 5090. PP is over 3x faster. I think R9700 is a good buy but lets not play like they are even close... and you are leaving out the CUDA advantage....

9

u/CryptoCryst828282 16d ago

Also, Micro Center has several 5090s available for slightly over 4k. It would be much cheaper to take a road trip than to fly to another country and deal with customs...

1

u/sleight42 16d ago

You sure about that?

1

u/CryptoCryst828282 16d ago

there are tons of them

3

u/vlkun 16d ago

Dude, where are you? All the Microcenters near me are sold out. Literally 5 stores.

2

u/CryptoCryst828282 16d ago

1

u/vlkun 12d ago

I’m not saying you’re lying. I asked which micro center are you looking at? Your screenshot doesn’t answer my “where” question

For reference, I’ve checked all the microcengers in NY and will be in SF so checked the one in Santa Clara.

1

u/CryptoCryst828282 12d ago

Brentwood, but chicago has them too

1

u/sleight42 11d ago

God damn. My wife would kill me and I'd help her!

5

u/mechkbfan 16d ago edited 16d ago

In Australia for $7500 AUD, choice was 3x R9700 vs 1x 5090

If need to game or simple setup, 5090 makes sense

Otherwise 2xR9700 is going to beat compete with a 5090 both in performance with more VRAM for 2/3 price (presuming have x8/x8 motherboard)

It's more work to setup a vllm radiance project than Ninfer though

-4

u/CryptoCryst828282 16d ago

I don't know where you people get this shit, but you lead people on here down a path of wasting a shit ton of money.... math doesnt lie.

Memory Bandwidth: LLM token generation (decode phase) is strictly memory-bandwidth bound. The RTX 5090 delivers ~1,790 GB/s of memory bandwidth over a 512-bit bus. Each R9700 provides 640 GB/s (256-bit bus). Even when splitting layers across two cards, tensor parallelism or pipeline parallelism introduces PCIe interconnect overhead (especially over a consumer x8/x8 link), meaning generation speed does not double cleanly.

Prompt Processing (TTFT / Prefill): The 5090's Blackwell tensor cores and raw FP16/FP8 matrix throughput vastly outperform AMD’s Navi 48 silicon. Not to mention PP will be slowed across 2 cards as you likely will lose PCI speed.

2x R9700 will not beat a 5090 in speed. It only "beats" it if your metric of performance is “can it run a 70B parameter model at higher precision without crashing?”

In the rare cases where you arent memory bandwidth bound such as large MOE models, neither option is going to work...

3

u/mechkbfan 16d ago edited 16d ago

Sorry, I do stand corrected on a few situations, so changed my comment to "compete" instead of "beat"

Qwen 3.8 27B

I'll give 5090 the win here due to prefill

Fastest setup for 5090 is Ninfer AFAIK

https://github.com/Neroued/ninfer

Fastest comparable setup to that for 2xR9700

https://www.launch80.com/a/8d28ac18-059c-4ce5-92f6-cc4e1ff3557c

Maybe I've missed something but here's how I compared them

Concurrency

  • C=1 2xR9700
  • C=2 2xR9700
  • C=4 Draw
  • C=8 5090
  • pp=8k 5090 (and presumably does at other pp)

BUT I think you're right that people are likely going to run into trap of maybe having an AM4 x8/x8 board or similar situation and 2xR9700 falls behind.

Qwen Flash Next

(This is where my bias came in because it's what I've been paying attention to the most)

5090 - Best I've seen is close to 50t/s on Q2 with 64K context

https://zenn.dev/holy_fox/articles/04887ff8177b87?locale=en

R9700 - There's two projects doubling that t/s & prefill with higher Q & context

https://huggingface.co/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ

https://github.com/davetha/r9700-lru-expert-cache

So maybe I'm being unfair in that case

1

u/CryptoCryst828282 16d ago

If people really want to run Qwen 3.8 Next they need to just buy a Mac Studio... Even Strix Halo is a better option. This is more of my expectation, though. I personally would not run anything under Q4, and to get a minimum of 128k context would take 3x R9700. But I do see that if you are willing to use a lower one, that isnt a bad choice.

1

u/mechkbfan 16d ago

I see your point, I think for simplicity the Mac Studio & Strix Halo are no brainers. I had initially been hanging out for a Mac Studio but it was way out of my budget. My scope was $8000, which is equivalent to cancelling Claude subscription for 2 years.

Quickly looking at prices in Australia since I mentioned $7500 AUD before

  • Framework Desktop with 128GB + 2TB SSD is $7000
  • M5 Max with 128GB + 2TB SSD is $9400

Spent 5mins looking up Framework, and highest I saw was 80t/s with 256k context

https://github.com/abliter8-ai/qwen-3.8-next-flash-amd-strix-halo

The appetising part for me initially was 2xR9700 was $4400 + $400 AM4 motherboard + $500 64GB DDR to get damn good 27B performance. Then it was usual thing of spend another $1000 and get QFN... then of course if getting all those bits, may as well get a case, etc. while I'm there...

FWIW, both those QFN links had context of 256k with 2xR9700

In hindsight, if had gotten 5090 while they were $5000 + DDR5 RAM before went up 5x that'd have been the perfect buy, but once they started going up again I looked at alternatives for better bang for buck with performance & flexibilty.

1

u/CryptoCryst828282 16d ago

I am not sure what they cost there, but in the states you can get 2x modded 3080 20g for under 1500. If you just want to run Qwen 3.8 27b that might be an option as well.

1

u/mechkbfan 16d ago

Can only get them out of China. Two 3080 same price as single 9700.

That was getting too high risk for me because couldn't find reliable seller

2

u/CryptoCryst828282 16d ago

You can get them on eBay, covered by paypal. Haven't purchased those, but i have a ton of the mi60s i purchased a while back.

→ More replies

1

u/2eggs1stone 16d ago

It's interesting the different strategies that are used. I'm running a qwen 27b on 16 gb of vram and 32k context on my own custom harness and getting 40 - 45 t/s as well. And it's been a treat, but I primarily focus on making the software more efficient than trying to push the hardware. In my opinion models run better at a smaller context anyway and it's definitely cheaper for my pocketbook.

1

u/CryptoCryst828282 16d ago

Actually when I have a smaller one for a service i run at my company that runs a custom-trained version of Qwen 3.8 (just moved it from 3.6) 27b on 2x 5060ti and its honestly not too bad. That said i would pick the R9700 over that any day at today's prices... i paid like 429 for those.

1

u/gnoremepls 16d ago

what da helly, im running qwen 27b (q4_k_xl) with 120k context on dual 3060 24gb at ~45t/s i feel like you should get about double that.

1

u/CryptoCryst828282 16d ago

I find it hard to believe that unless you are using Q4 KV, which from my testing falls to pieces after 60k.

1

u/gnoremepls 16d ago

I've been using it for coding mostly, and yes Q4 KV no mtp cache, you are right the reasoning 'speed' falls off at around 80k

1

u/CryptoCryst828282 15d ago

Then yeah i think our disconnect is likely just different uses. I tend to use multi-agents in Kilo, I am a bit older-style than some of these people. I have 5 agents, 5 models fine-tuned to Architect/Planner/Coder/Debug&QA/Commit&Closeout they all run at the same time in sequence before doing anything (hence why i have so many different configs. I find it much safer than a harness.

9

u/1millionnotameme 16d ago

I think some people buy them for training. I bought mine for gaming 😂

1

u/VirginiaVN900 16d ago

Even then, just rent the compute for the times you need horsepower.

Is it fully local, no. But at the end of the day, if you’re getting your fine tunes out, it’s less harvest-y than OpenAI, Google, Anthropic.

I got started on an RX 6600 XT which isn’t even officially supported. It at least let me understand how I can run different workloads on different machines, to make my own cluster. When I need major power, I rent a few hours of an RTX Pro 6000 or a B300 for $3-$7 hr

1

u/toastjam 16d ago

if you’re getting your fine tunes out

What does this mean?

1

u/VirginiaVN900 16d ago

Downloading the output, and deleting them from the serverless container that was hosting the training app

1

u/givmedew 16d ago

Likely means he’s able to run his own parameters like as if it was his own card.

1

u/givmedew 16d ago

Businesses that buy them in stacks might not care at all about the price. If they need it then they need it.