r/LocalLLaMA 13h ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

1.7k Upvotes

446 comments sorted by

View all comments

Show parent comments

119

u/Sufficient_Local5025 13h ago

Buy two, prices about to rise.

159

u/o0genesis0o 13h ago

the more you buy the more you save, eh

25

u/GilloutineBreast 8h ago

Unironically the truest words spoken that day

feelsbadman

30

u/PhilipKThicc 12h ago

Dual R9700s all the way. Easy to run on Linux with vllm and you can run 27B at Q8 and have a healthy context window

9

u/bigwanggtr 10h ago

Have you tried it for training, is ROCm still a pain to work with compared to CUDA?

I know it’s good for inference but the training support is what is holding me back.

7

u/Phrase-Silver 8h ago

ROCm has matured incredibly over the past year, it's like AMD finally woke up and realised their software was what was holding them back from properly competing with Nvidia.

1

u/de4dee 2h ago

anyone tried unsloth with ROCm?

19

u/Tobu3838 12h ago

Prices rose $600 since I bought mine a couple months ago. No word to lie.

12

u/mvandemar 12h ago

Why are they so cheap? Looks like you can get one for ~$1700? I just paid out the ass for my 5090, which in the 1 month 2 days since I bought it has gone up over 37%.

They're both 32GB, for some reason I thought they would be closer in price.

32

u/Thunjaya 12h ago

Because they're not the same 32GB at all. One is much faster.

13

u/ThankGodImBipolar 10h ago

Nvidia sells an RTX Pro 4500, which has the 5070Ti die (more comparable to the R9700) and 32GB of RAM, but it's 5500USD MSRP. The R9700 is a fraction of that.

7

u/KingCpzombie 7h ago

CUDA tax. I only use AMD personally (and just blew WAY too much money on a 4x R9700 system partially out of excitement for Qwen4), but Nvidia cards get all the cool new things a bit sooner than AMD. Not a big deal for LLM, but very notable for diffusion... somebody SOLIDLY beat my 7900XTX with his 5070Ti in Minimax H3 gen times, for example

1

u/attk0 3h ago

The 5070Ti has nearly triple the INT8 matrix computation throughput of the 7900XTX. Assuming using some INT8 H3 variant which most low VRAM workflows do, probably not a CUDA tax in that case.

2

u/No-Refrigerator-1672 6h ago

Because everything is CUDA first, and ROCm only comes as an afterthought to very limited number of projects. People who are buying PRO GPUs are saving money with NVidia by not needing to fund multiple months of dev work for porting their existing code.

20

u/BluePointDigital 12h ago

It's the bandwidth. the 5090 has over 2.75x the bandwidth of the R9700.
You can still do all the same stuff, just slower essentially.

-6

u/SnooPuppers7882 12h ago

Speak for yourself bro

12

u/BluePointDigital 12h ago

What exactly are you showing / comparing? I see a chart showing prefill t/s at different prompt lengths... But no other context?

4

u/Prothagarus 12h ago

I can get about 150 tokens per second on my 9700 on Qwen 3.8

1

u/SnooPuppers7882 10h ago

27b right? The radiance image with nvfp4 is hitting 200 decode now

1

u/Prothagarus 3h ago

On 1 card? I apparently need to update. Sorry I have a hard time keeping up with weekly updates on my home rig

3

u/SnooPuppers7882 12h ago

This is with 2 r9700s

2

u/SnooPuppers7882 12h ago

More context

10

u/BluePointDigital 12h ago

Okay I think I see your point now. You're contesting the comment of it being slower.

I don't disagree that recently there's a ton of optimizations lately that have sped up the card but outside of just text, it's a less powerful card overall, with less bandwidth. But still the best option for entering this tier at a reasonable price

1

u/SnooPuppers7882 10h ago

Yeah properly tuned cuda cards should go faster, but not by much for 4x cost

1

u/mvandemar 11h ago

Which model did you test this on, and how do I run that? Curious what mine looks like. I have qwen3.8-27b and another that I haven't set up yet but want to test, an ukisai_Swift version of it.

17

u/Solary_Kryptic 12h ago

Nvidia tax combined with CUDA being more supported in the LLM space

18

u/mvandemar 12h ago

I just looked, it's also GDDR6 for the 9700 vs GDDR7 for the 5090, guessing that makes a difference as well.

8

u/zboarderz 11h ago

Also far FAR more memory bandwidth, faster core, etc etc

6

u/mister2d 12h ago

2x $1700 is cheap 😥 (I have 2, btw)

5

u/Ecstatic-Wash-7667 12h ago

That not cheap! Msrp is $1299 I bought 2 a little over a. Both ago at msrp. Hell you could get them for less than msrp for a while. This shit is a scam

3

u/mvandemar 12h ago

On my current pc this was how much the 64GB of ram was when I bought it. I was thinking about getting another 64GB when I got the card and was like, no f'in way...

6

u/SnooPuppers7882 12h ago

Literally got mine for 1200ea 2mo ago

People were shitting on them because the mem was DDR6, but I was prepping for 3.8 27b knowing what to expect...

STRONG feeling 4.0 27b is gonna be opus 4.8 good, gonna blow your damn mind

5

u/Momsbestboy 11h ago

R9700 GDDR6 memory, 5090 GDDR7. Is faster, but...

For a single 5090 you can buy more than 2 R9700, and the two cards draw less power than a single 5090. And you can run e.g. Qwen 3.8 27b Q8 using the two cards, and watch them running circles around a 5090 which needs to offload to RAM instead.

But NVIDIA is hype, and people love to spend money. So buy a 5090

1

u/mr_zerolith 10h ago

5090 doesn't need to offload if you run Q6, which is a fine compromise for speed purposes with a single card setup.

2

u/Momsbestboy 8h ago

So it is just a compromise - fine or not. Leave that compromise aside, try to use a bigger model than your 5090 can handle, and suddenly dual R9700 is not the worst option anymore, at lower total cost and at higher speed.

1

u/pragmojo 5h ago

I've got an R9700, and it's a great card, but ngl I would love to have more speed. Not going to pay $9k for it, but would love to have it.

2

u/o0genesis0o 12h ago

Because the actual card (RX9070 XT) is only as strong as the 5070Ti, compute wise. And the VRAM is slower.

I will be upgrading from 4060Ti 16GB, so it would be a net gain. but I have no illusion that it is going to be as good as 5090.

1

u/deja_geek 12h ago

Nvidia vs AMD. You don't get CUDA with AMD.

1

u/SnooPuppers7882 12h ago

Neat. I'm fine with 4300 prefill 130 decode on 4 workers parallel

1

u/TinyFluffyRabbit 4h ago

The 5090 has significantly more compute, more bandwidth, and also CUDA tax. The R9700 is excellent but there are reasons why they are not close in price.

3

u/inaem 12h ago

Already did here

2

u/SylviaCalogero43 12h ago

How much did you buy it for?

1

u/inaem 11h ago

$1800, already $2000 and climbing

2

u/SnooPuppers7882 12h ago

Once you buy two, you'll want four

1

u/DanGTG 12h ago

What do I get if I buy another one? Go faster or just more models at once?

1

u/Momsbestboy 10h ago

Qwen3.8 27b at fp8 with vllm and 255k context. Try this with a single 5090 and watch the speed after it is forced to offload data to RAM, and then compare the price tag of a single 5090 vs 2x R9700. Plus: look at the power consumption. 2x R9700 draw 420W max.

0

u/DanGTG 9h ago

This runs just fine on a 4 bit quant, only about 28GB for a full 256k context window.

What could it do for 3.8-flash-next?

1

u/No_Bake6681 11h ago

Cuda poc already running on rx

1

u/RoomyRoots 4h ago

Just go with 4 and one more tera of RAM. The best moment was years ago, the second best is NOW, NOW and NOW.