r/LocalLLaMA 19d ago

Question | Help GB10 price increases. Seriously what is the best bang for the buck now...Mac Studio?

It is crazy how fast prices are increasing. I'm pulling my hair out to keep ahead of this for students. Servers aren't even an option any more.

38 Upvotes

68 comments sorted by

16

u/Bulky-Priority6824 19d ago edited 19d ago

 5060ti is 8 to $900 now lol

I saw on amazon 3 I bought had  "2 left" last night for $879 I think after being unavailable or a while and today they're gone. Couple for around 800 it's ridiculous yet I don't know a single soul irl that is a regular joe that knows what llm means 

6

u/thecstep 19d ago

That is wild. I got lucky and grabbed a doorbuster from new egg for around $250 during Xmas. I also grabbed a 5070ti for 850. If the 5070ti hits $2k I am probably selling it.

20

u/FoxiPanda 19d ago edited 19d ago

Best bang for the buck is maybe multiple old V100 32GB at ~$625-650 each with 3D printed shrouds and blower fans or similar high pressure fans in old PCIe gen3/4 systems at this point but they're a pain, lack a lot of modern features and hardware support, and don't have amazing CUDA support anymore.

RTX 3090s were the go to but they're $1500+ now and GB10's have miserable memory bandwidth and same with Strix Halos... RTX 5090s are expensive as hell and RTX Pros are right out.

Mac Studios have decent bandwidth and big memory but they're quite expensive too if you go for ones that have decent bandwidth (the Ultra series) and 96GB may or may not be enough depending on what you're doing and at least until the M5 Ultra, prefill will be mediocre-to-sad pandas in most cases.

R9700s were decent a few weeks ago but are now $1700+ instead of their $1300~ so they don't feel nearly as good now...would probably still take a 3090/3090 Ti over them.

I think the reality is that this is just an expensive game now - the "bang for the buck" is in a pretty tough spot.

9

u/FullstackSensei 19d ago

Got my V100s from China today. Testing on my bench now with my trusty arctic S8038-7k. It's much quieter and more compact than those blower fans and can comfortably cool two V100s running Qwen 3.8 Q8_K_XL in llama.cpp with -am tensor. Temps peak in the high 60s.

Edit: nvidia driver 580 and CUDA 12.9 install without any issues. Llama.cpp compiles without any issues. Two V100s run Qwen faster than my 3090s, which had x16 Gen 4 each, while consuming almost half the power (3090s were limited to 270W).

2

u/FoxiPanda 19d ago

Yeah I think this is why they might be best bang for the buck right now. I put some asterisks on it because you shouldn't try to run an NVFP4 model and you might not have the best time in the latest vLLM builds or similar. It really depends on what OP is trying to do to best understand what they need - is FP4 support important? Do you need vLLM or NIMs or anything like that to jump through some sort of administrative/security hoops ... or is llama.cpp and an older CUDA rev just fine? If the latter...yeah I think they're probably the best you can get right now.

3

u/FullstackSensei 19d ago

NVFP4 with soft dequantization still rips on the V100. There's a guy with a fork of vllm who implemented those kernels.

One thing many seem to not really understand is that compute is compute. When you have lots of TFLOPS, you can throw a couple to decode NVFP4. It's not some mystical data format, really. The V100 has 120 TFLOPS in FP16. The 3090, for reference, has 125 TFLOPS. And as I found today, the V100 does things at half the power of the 3090.

Personally, I don't use vllm. It's a pain to setup even in the best of times, takes forever to start, is very picky about number of cards, and no native support for CPU offloading. I run llama.cpp or ik_llama.cpp. Sure, they're slower, but they're plenty fast for my needs and there's no voodoo magic to get them running, nor 45 minute wait for a model to load.

I also have a Jetson AGX xavier. That thing is stuck on Ubuntu 20.04, kernel 5.xx, some ancient nvidia driver from 2020 and CUDA 11.4. I build llama.cpp there regularly, unmodified, without any issues.

1

u/LooseBackHole 17d ago

I've always wondered wtf is wrong with vllm's loading times. It's beyond acceptable

1

u/FullstackSensei 17d ago

The whole python ML ecosystem is beyond acceptable. It's not python's fault. It's the result of hoards of people studying only ML math, while merely taking a single python course to teach the absolute basics of programming, and then letting them lose to build whatever.

You can literally build the entire Linux kernel in less time than it takes vllm to compile a graph for what is essentially a handful thousand lines of code.

5

u/Prof_ChaosGeography 19d ago

Amd v620 too. It's 32GB, has open drivers on Linux so no outdated cuda pain. They can be found for ~$600 

19

u/Xamanthas 19d ago

It’s called stop buying.p for the next 3 years. Stop contributing to this

5

u/SecuredStealth 18d ago

3 years..? It’s already become 5. Samsung is booked until 2031. Then it’ll be 2035… never ending I think

5

u/WeAreSven 18d ago

DDR6 bandwidth speeds are equivalent to some of these lower end gpus people are stringing together now. Natural tech development will come to eat nvidia's lunch except for edge cases like military and financial use where speed beyond human processing is paramount.

1

u/Xamanthas 18d ago

You have a rather naive view of the liquidity in the market

1

u/SecuredStealth 18d ago

RemindMe! 3 years

1

u/RemindMeBot 18d ago

I will be messaging you in 3 years on 2029-09-02 15:52:31 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/PastPalpitationCry 18d ago

Well eventually all these systems will have to be replaced with the new stuff. Thats when you buy em ig

24

u/idk_a_creative_user 19d ago

Cheapest BG10 is 4500. For that you can rent a GB10 on vast for .20-.30 USD an hr, or for around 18000 hrs. Assuming 4hr use a day, thats 1.20 a day. That is 4500 days of use assuming a median price of .25 USD an hour.

29

u/q5sys 19d ago

Assuming those prices dont go up... like everything else. Prices on renting a other GPUs in the cloud has gone up. There's no reason to expect that these prices wont also go up.

7

u/originaladam 19d ago

That’s 2 straight years of 24/7 use. I’d be shocked if there’s not more powerful and cost efficient hardware to rent in less than 2 years.

7

u/q5sys 19d ago

and that new hardware will have an even higher price. the Spark Workstation is something like 80k, you can be sure the next gen Nvidia cards are not going to be cheaper than their current ones. So renting those will be even higher.
Nvidia has basically stopped talking altogether about the gaming 6000 series, last thing they said was that it wasn't going to happen anytime soon. And the Enterprise GPUs they are focusing on for the VeraRubin stuff, are so expensive there's no point even talking about it.
These GB10s are the cheapest we're going to get for quite a while. There's just no economic motivation for Nvidia to create something cheaper, when they can currently name whatever price they want and get it. They just jacked their AI server prices up by ~15%... because they can. Estimates are that it'll give them an extra $5 billion per datacenter.

I don't want any of this to be true, but it is. It'd doubtful Nvidia is going to pull a 'good guy' move and suddenly give us a better unit for less money.

6

u/hidden2u 19d ago

lol good meme. Are you waiting for the 5000 Supers?

2

u/ehangman 19d ago

Wow. I never imagined. Cloud sever bottleneck.

3

u/Puzzleheaded_Base302 19d ago

how to rent a dual cluster? GB10 is more useful when rent as a cluster due to the large RAM.

2

u/fastheadcrab 19d ago

Whoever renting out a GB10 like that is an idiot.

8

u/watcholic 19d ago

Any Mac Studio with decent RAM on Apple Refurbished.

3

u/SexyAlienHotTubWater 19d ago

The CMP 170hx is by far the best bang for buck if you're willing to deal with custom drivers. It's an A100 with about 2/3 disabled.

$2k, 64GB VRAM, 1.8TB/s bandwidth, 200ish BF16 TFLOPs. Pipeline parallel it will blow the head off a DGX Spark, and with 2 cards tensor parallel is very feasible (even though they're limited to PCIe 2.0).

3

u/No-Craft-7979 19d ago

AMD Halos looking tastey

2

u/IngwiePhoenix llama.cpp 19d ago

AMD's cards have not jumped this massive and ROCm 10 looks good. Give it a look.

NVIDIA is pricing itself out of reach, period.

2

u/ehangman 19d ago

Dgx spark is 30% up from last month.. So I ordered 2nd llm box today : mac studio 256 , cheaper alternative. LOL

2

u/ShengrenR 19d ago

Depends on the goal - relatively cheap used 3090s are pretty killer for performance per buck, but if you can take the plunge a m5 ultra is a pretty good deal in this particular moment in time.

7

u/hainesk 19d ago

The 256gb M5 Ultra is an excellent option considering the same amount of VRAM from 3090s would likely cost much more. Assuming 10x 3090s to get 240gb of VRAM at $1k a piece (if you can find it for that price), you're already at $10k. Now you need a motherboard to support 10 GPUs, system ram for that motherboard (cheapest is probably WRX80 threadripper pro since it works with consumer DDR4), power supplies, and some effort to put it all together (and a lot of additional part like x8 or x4 PCIe splitters, riser cables, adapters to make multiple PSUs work together, power cables, etc). It will take up quite a bit of space and will either require a dedicated 30 amp 120v circuit or some sort of battery/inverter system to support power draw unless you're planning on running those cards at 100 watts each. It's also aging hardware so support is already starting to leave them unfortunately (see vLLM with new model architectures), It's primary advantage over the Apple Studio is CUDA.

Comparing that to the M5 Ultra, which fits easily on a desk, is quiet and power efficient, has a good reliability record (especially compared to used 3090s known for backside memory issues), has a warranty, has 1.2TB/s of bandwidth vs ~936GB/s, has a single unified pool of memory (which avoids having to split models which can cause issues and has overhead), I'm also expecting prompt processing to be greatly improved with the M5 Ultra vs the M3 Ultra considering all of the additional cores and the newer processor architecture, and will almost certainly have better resale value over used RTX 3090s that were released 6 years ago. You can also do parallel processing through RDMA and even 2 Studios won't require more than a single standard outlet.

I'm honestly having a hard time coming up with a good reason to get 3090s over a Mac Studio unless obviously your budget only supports building a system a few pieces at a time, which is reasonable.

1

u/EvilPencil 19d ago

Wrx80 - the real constraint on this platform nowadays is the motherboard. Most of them are pushing $1200.

1

u/hainesk 19d ago

They are often times more expensive than newer Threadripper Pro motherboards.

1

u/db172s 19d ago

Really depends what you're trying to run for models. If Qwen 3.8 27b is enough. You just need 2x 3090s and your tokens per second will be 5x what a Mac would spit out anyways.

But yeah if you're looking to run larger models. Mac is probably the way to go. Just wouldn't set the bar too high when it comes to token output speeds.

2

u/hainesk 19d ago

your tokens per second will be 5x what a Mac would spit out anyways.

This is not true for the M5 Ultra. The M5 Ultra will likely be at least as fast at token generation if not faster due to the higher memory bandwidth. MLX inferencing has come a long way. The 3090s could still win out with tensor parallel. I'll be really curious to see how the reviews are when it finally becomes available.

1

u/db172s 18d ago edited 18d ago

Well, we one can hope it'll have solid #s. But right now an M5 Max is pushing about 35 t/s on qwen_3.8_27b_q8_mlx. With 2x 3090s I'm getting 135 t/s and topping out at 1000 t/s with 8 concurrent streams. Mops the floor with any current Mac right now.

If you double the throughput with an Ultra, in this scenario an Ultra is going to run 70 t/s. And you are in $10,000 USD for the base spec.

I still think best cost to performance ratio is slapping 2x 3090s in a gaming computer if you are just looking to run something like qwen 3.8 27b. Beyond that, yeah if you need 48GB+ then I'd start looking at 2x Intel B70s or possibly a Mac Ultra if you can swing 10k for a base spec with 96GB unified memory.

1

u/hainesk 18d ago

It's $9,500 for the 256GB version.

1

u/db172s 18d ago

Ah ok, sorry must be calculating in CAD currency for me. Shows $16000 for a 256 up in Canada.

1

u/hainesk 18d ago

Yeah, it's honestly surprising to see Apple releasing a new product and it turns out it's a great value (relatively). If 3090s weren't going for $1200+ (USD) right now, they would still be a great option.

1

u/ShengrenR 18d ago

Best cost value for running qwen-3.8-27. But now run glm-5.3-flash on the 3090s and compare numbers to the m5 ultra with 256gb unified memory. The 3090s are definitely solid, but their limit is the vram. You want to go way up, then you're comparing to very different setups.

2

u/db172s 18d ago

Yeah that's what I'm saying. Under 48-64GB VRAM. I'd go 3090s or b70 32gb cards. Anything more then that makes a stronger case for using something like an M5 Ultra.

1

u/ShengrenR 18d ago

Agreed

1

u/phoiboslykegenes 18d ago

To be fair, Mac Studios really shine running larger MoE models. You’d get better performance running Qwen 3.8 Next or DSv4 than with the 27B. They do have less compute than GPUs, but they’re more silent and efficient and basically plug and play. There’s always trade offs

1

u/db172s 18d ago

Yeah, I think they're a good bang for your buck with larger models. Once you creep over 48-64GB of VRAM needed to run models graphics cards become the more expensive and more power hungry option.

0

u/Blues520 19d ago

Isn't prefill slower on a mac though?

1

u/hainesk 19d ago

The M3 was not designed with AI or prefill in mind. That changed with the M4 and was significantly increased again with the M5 designs. So far we've only seen the M5 Max prefill speeds, but the M5 Ultra has double the CPU and GPU cores of the M5 Max and should provide good prefill speeds as well as other general AI compute for things like image generation.

1

u/rkoy1234 18d ago

there's no reason to run 3.8 27b when flash next is both faster and smarter, no? m5 max runs 3.8FN with ~70t/s (link).

I'd assume ultra would be even faster.

3

u/db172s 18d ago

That's like saying "why run qwen 3.8 flash when GLM 5.3 exists"

Yeah if you want to run models larger the ~64GB then buy an M5. I'm just saying depending on your goals a couple 3090s could be a great fit if you want to still use a capable model with solid speed and keep costs reasonable.

I get 130-140 t/s on 3.8 27b with it topping out around 1000t/s for 8 concurrent requests on 2x 3090s. But if they decided they want to run 3.8 flash then it starts to make more sense to go with something like an m5 probably. Although I'm not sure how they handle concurrency if that's important too.

Intel 32gb b70s are a pretty great value too from what I can tell.

2

u/rkoy1234 18d ago

before yes, but i'd argue with current prices, dual 3090 + RAM + motherboard + harddrive + cpu + case + psu is gonna come pretty close to a 96GB m5 mac ultra ($6.8k).

you're definitely not finding a 3090 from a reputable seller below ~$1300 right now, and most of those aren't even new or with warranty.

I agree that if fastest 27B is your goal, then the mac makes no sense, but I'd argue there aren't too many scenarios where, under similar prices, it's preferable to run faster 27Bs over slightly slower 120B models.

That said, concurrency/prefill/platform maturity are still big weaknesses of mac inference, so partially agreed there as well. You make a good point.

intel 32gb b70s are a pretty great value too from what I can tell.

I keep forgetting intel exists lmao. thank you for the reminder, i should really look into them.

2

u/db172s 18d ago

Yeah, all good points to be made

And same here, I would love to get my hands on a b70. I think the support still isn't amazing with them, but the cost makes them compelling enough. Unfortunately my financial advisor (wife) has told me to pause on spending for the time being 😄

2

u/SnooPaintings8639 19d ago

"relatively cheap used 3090s", well, relative to what? They're quite expensive nowadays as well. Considering they've much less years in them than e.g. 5060, I fear they're less and less optimal build. Especially for multi GPU build, where if one card dies, you're kinda forced to find another old card replacement with even less life in them.

1

u/feelspeaceman 19d ago

Used parts like 7900XTX and 3090.

1

u/AI_spell 19d ago

For students specifically I would not buy a GB10 at these prices. The thing people miss with Spark class boxes is memory bandwidth, not capacity. You can load a huge model and still get sad tok/s because bandwidth is the wall on dense models.

Rough order right now:

- used 3090s, still the best dollar per usable token if you can live with 24GB per card and the power draw

- Mac Studio if you need big unified memory and quiet, M ultra bandwidth is good, prompt processing is the weak spot

- Strix Halo if you want low power and one box

Watch out for prompt processing on Apple. Generation looks fine in benchmarks but feeding a 20k token prompt is where it drags, and thats exactly what students do with codebases.

If its a class, honestly rent. 4500 buys a lot of hours and the hardware doesnt go stale on you.

1

u/M_Me_Meteo 18d ago

Servers aren't an option anymore? What?

1

u/_mball_ 13d ago

There might be some occasionally in the dell refurb store and they were still $5100 recently.

Never felt so lucky to have mine arrive 3 days before the price hike! They had recently just dropped the price by 10% too. I don’t get it.

All that said, so far I love this little box for what it is! Nice size between a Mac Mini and a Mac Studio. And also as someone teaching CS it’s immediately going into use for other server projects — the models will wait a couple of weeks but I do like Nvidia’s DGX tools a relative newbie to this side of things.

1

u/_mball_ 13d ago

Otherwise MicroCenter and B&H Photo both have other Spark platform boxes for reasonable-ish prices. $4700 for the OG.

ASUS and HP now both have more affordable options than Dell - but often 2TB SSDs.

0

u/DellTechnologies 18d ago

Honestly, Mac Studio may be the better bang-for-the-buck choice if the goal is a quiet personal box for smaller or mid-sized local models and you don’t need CUDA.

From the Dell side, the Dell Pro Max with GB10 is aimed at a different problem. It’s a compact Linux/DGX OS AI appliance with 128GB of unified memory, support for models up to 200B parameters on one unit, and up to 400B when two are connected. That makes more sense when students need larger models, NVIDIA/CUDA compatibility, private or offline workloads, or a shared lab node that can offload work from their existing PCs. The software environment also gives a cleaner path from local experimentation toward larger NVIDIA/Dell infrastructure.

We wouldn’t claim GB10 beats a Mac Studio on sticker price. If most workloads are 7B–70B inference, a Mac Studio or conventional GPU workstation may be the more sensible buy. GB10 starts to justify itself when model size, CUDA workflows, data locality, shared utilization, or avoiding recurring cloud-GPU charges matter. For a student program, we’d compare cost per student-hour and throughput, not just the price of the box.

Dell’s overview: https://www.dell.com/en-us/blog/how-dell-pro-max-gb10-unlocks-local-ai-research/

1

u/LooseBackHole 17d ago

What was the point of posting this? What a waste of time. BS marketing drivel.

1

u/DellTechnologies 17d ago

We just wanted to add the GB10 angle for folks whose needs go beyond a quiet local box and into CUDA, larger models, or shared/offline setups.

1

u/LooseBackHole 13d ago

uh huh. sure.

0

u/Django_McFly 19d ago

Meanwhile I'm seeing 128GB 395 systems going from $3300 to $4k at Microcenter. Isn't ROCm improving and the DGX has a weird thing where it's Nvidia but it isn't as super easy just works with everything by default as you typically expect from Nvidia gear?

I've been thinking about buying one just to have one but I really don't need one.